Add scraper indexes and switch to validation collection names
All checks were successful
CI/CD Pipeline - Apartment API / Scan Dependencies (pull_request) Successful in 13s
CI/CD Pipeline - Apartment API / Lint & Test (pull_request) Successful in 43s
CI/CD Pipeline - Apartment API / Send Webhook Notification (pull_request) Successful in 2s
CI/CD Pipeline - Apartment API / Build & Push Image (pull_request) Has been skipped
CI/CD Pipeline - Apartment API / Deploy to Production (pull_request) Has been skipped

Add createScraperIndexes() to create indexes on the scraper_runs
collection: a compound index on status+startedAt for active job queries
and a descending index on startedAt for recent run lookups.

Change default collection names from production tables
(units_migration_test, unit_prices_migration_test) to dedicated
validation collections (units_scraper, unit_prices_scraper). This
enables the Node.js scraper to run in parallel with the existing Python
scraper during validation without interfering with production data.

Update all scraper test files to reference the new default collection
names.
This commit is contained in:
2026-02-07 12:04:37 -07:00
parent 80d44bca04
commit 2e3ef0580c
8 changed files with 101 additions and 20 deletions

View File

@ -30,10 +30,11 @@ module.exports = {
// Graceful shutdown timeout (how long to wait for running job before force-stopping)
SHUTDOWN_TIMEOUT: parseInt(process.env.SCRAPER_SHUTDOWN_TIMEOUT, 10) || 30000,
// MongoDB collection names (environment variable overrides for development isolation)
// MongoDB collection names - defaults to validation collections for scraper validation;
// switch to production collections via env vars (SCRAPER_UNITS_COLLECTION, SCRAPER_PRICES_COLLECTION) when ready
COLLECTIONS: {
UNITS: process.env.SCRAPER_UNITS_COLLECTION || 'units_migration_test',
PRICES: process.env.SCRAPER_PRICES_COLLECTION || 'unit_prices_migration_test',
UNITS: process.env.SCRAPER_UNITS_COLLECTION || 'units_scraper',
PRICES: process.env.SCRAPER_PRICES_COLLECTION || 'unit_prices_scraper',
DAILY_SUMMARIES: process.env.SCRAPER_SUMMARIES_COLLECTION || 'daily_summaries',
SCRAPER_RUNS: process.env.SCRAPER_RUNS_COLLECTION || 'scraper_runs'
}