SCRAPE-11: Implement runScrape() main orchestration function
Some checks failed
CI/CD Pipeline - Apartment API / Send Webhook Notification (pull_request) Failing after 2s
CI/CD Pipeline - Apartment API / Build & Push Image (pull_request) Has been skipped
CI/CD Pipeline - Apartment API / Scan Dependencies (pull_request) Successful in 13s
CI/CD Pipeline - Apartment API / Run Linting (pull_request) Successful in 9m37s
CI/CD Pipeline - Apartment API / Run Tests (pull_request) Successful in 9m56s
CI/CD Pipeline - Apartment API / Deploy to Production (pull_request) Has been skipped

Add the top-level runScrape() function that coordinates the full scraper
pipeline: fetch HTML (with retry), parse units, convert data types, upsert
units, insert prices, mark stale units, and update daily summary.

Features:
- dryRun mode skips all database writes while still parsing/validating
- htmlContent parameter allows injecting HTML directly (bypasses fetch)
- Calculates newUnitsCount and rentedUnitsCount by diffing against prior state
- Records every run to scraper_runs history (success or failure)
- Structured logging with jobId correlation throughout the pipeline
- Graceful error handling at each pipeline stage

Includes 15 tests covering full workflow, dry run, error handling,
trigger types, new/rented unit calculation, and empty HTML edge case.
This commit is contained in:
2026-02-06 01:35:15 -07:00
parent b4978caf31
commit 5ef5d5af73
2 changed files with 652 additions and 0 deletions

View File

@ -7,7 +7,9 @@
const axios = require('axios');
const cheerio = require('cheerio');
const crypto = require('crypto');
const config = require('../config/scraper');
const { createLogger } = require('./scraperLogger');
/**
* Sleep utility for retry delays
@ -619,7 +621,200 @@ async function updateDailySummary(db, summaryData, logger) {
}
}
// ============================================================
// Additional Helpers for runScrape Orchestration
// ============================================================
/**
* Get today's date in YYYY-MM-DD format (UTC).
* @returns {string} Today's date string
*/
function getTodayUTC() {
return new Date().toISOString().split('T')[0];
}
/**
* Get the set of unit codes that had price records yesterday.
* Used to calculate new and rented units by comparison.
* @param {Db} db - MongoDB database instance
* @param {string} today - Today's date in YYYY-MM-DD format
* @returns {Promise<Set<string>>} Set of unit codes from yesterday
*/
async function getYesterdayUnitCodes(db, today) {
const yesterday = getYesterday(today);
const collection = db.collection(config.COLLECTIONS.PRICES);
const yesterdayRecords = await collection
.find({ date_checked: yesterday }, { projection: { unit_code: 1 } })
.toArray();
return new Set(yesterdayRecords.map(r => r.unit_code));
}
// ============================================================
// Scraper Run History
// ============================================================
/**
* Record scraper run to history collection.
* This function intentionally catches errors and returns null
* rather than throwing, because recording history should not
* break the main scraper workflow.
*
* @param {Db} db - MongoDB database instance
* @param {Object} runData - Run data to record
* @returns {Promise<Object|null>} Insert result or null on error
*/
async function recordScraperRun(db, runData) {
try {
const collection = db.collection(config.COLLECTIONS.SCRAPER_RUNS);
const result = await collection.insertOne({
...runData,
recordedAt: new Date()
});
return result;
} catch (error) {
// Log but don't throw - recording history should not break scraper
console.error('Failed to record scraper run:', error.message);
return null;
}
}
// ============================================================
// Main Orchestration Function
// ============================================================
/**
* Execute a complete scrape operation.
* Orchestrates the full workflow: fetch -> parse -> convert -> DB ops.
*
* @param {Db} db - MongoDB database instance
* @param {Object} options - Scrape options
* @param {string} [options.trigger='manual'] - Trigger type ('scheduled' | 'manual')
* @param {string} [options.jobId] - Optional job ID (generated if not provided)
* @param {boolean} [options.dryRun=false] - Skip database writes for safe testing
* @param {string} [options.htmlContent] - Use provided HTML instead of fetching
* @returns {Promise<Object>} Scrape result with status and metrics
*/
async function runScrape(db, options = {}) {
const jobId = options.jobId || crypto.randomUUID();
const trigger = options.trigger || 'manual';
const dryRun = options.dryRun || false;
const htmlContent = options.htmlContent || null;
const logger = createLogger(jobId);
const startTime = Date.now();
let result = {
jobId,
trigger,
dryRun,
status: 'running',
startedAt: new Date().toISOString(),
completedAt: null,
duration: null,
unitsProcessed: 0,
pricesInserted: 0,
newUnitsCount: 0,
rentedUnitsCount: 0,
staleUnitsCount: 0,
errors: []
};
try {
logger.info('Scrape started', { trigger, dryRun, usingProvidedHtml: !!htmlContent });
// Step 1: Fetch HTML (or use provided content for testing)
const html = htmlContent || await fetchPage(config.TARGET_URL, logger);
// Step 2: Parse units
const rawUnits = parseUnits(html, logger);
if (rawUnits.length === 0) {
logger.warn('No units found in HTML - possible structure change');
result.errors.push('No units found in HTML');
}
// Step 3: Convert data types
const units = rawUnits.map(unit => convertDataTypes(unit));
result.unitsProcessed = units.length;
// Step 4: Database operations
const today = getTodayUTC();
// Get yesterday's unit codes for comparison
const yesterdayUnits = await getYesterdayUnitCodes(db, today);
// Determine new and rented units
const currentUnitCodes = new Set(units.map(u => u.unit_code));
const newUnits = units.filter(u => !yesterdayUnits.has(u.unit_code));
const rentedUnits = [...yesterdayUnits].filter(code => !currentUnitCodes.has(code));
result.newUnitsCount = newUnits.length;
result.rentedUnitsCount = rentedUnits.length;
// Database operations (skip if dryRun)
if (dryRun) {
logger.info('Dry run mode - skipping database writes', {
wouldUpsert: units.length,
wouldInsertPrices: units.filter(u => u.price !== null).length
});
result.pricesInserted = 0;
result.staleUnitsCount = 0;
} else {
// Upsert units
await upsertUnits(db, units, logger);
// Insert prices
const pricesResult = await insertPrices(db, units, today, logger);
result.pricesInserted = pricesResult.insertedCount;
// Mark stale units
const staleResult = await markStaleUnits(db, currentUnitCodes, today, logger);
result.staleUnitsCount = staleResult.modifiedCount;
// Update daily summary
await updateDailySummary(db, {
date: today,
newUnits: newUnits.map(u => u.unit_code),
rentedUnits,
staleUnitsCount: result.staleUnitsCount,
totalAvailable: units.length
}, logger);
}
result.status = 'success';
} catch (error) {
logger.error('Scrape failed', {
errorType: error.name,
errorMessage: error.message
});
result.status = 'failed';
result.errors.push(error.message);
} finally {
result.completedAt = new Date().toISOString();
result.duration = Date.now() - startTime;
// Record run to history (always runs, even on failure)
await recordScraperRun(db, result);
logger.info('Scrape completed', {
status: result.status,
duration: result.duration,
unitsProcessed: result.unitsProcessed,
pricesInserted: result.pricesInserted
});
}
return result;
}
module.exports = {
runScrape,
fetchPage,
parseUnits,
convertDataTypes,
@ -627,7 +822,10 @@ module.exports = {
insertPrices,
markStaleUnits,
updateDailySummary,
recordScraperRun,
// Export helpers for testing
getTodayUTC,
getYesterdayUnitCodes,
getYesterday,
parseInteger,
parsePositiveInteger,