Implement a status endpoint that returns the current state of the
scraper system. The response includes mutex state (idle/running with
job ID), last run history from scraper_runs collection (status,
timing, unit/error counts), next scheduled run timestamp, and
cron schedule expression.
Protected by requireAuth + requireAdmin middleware. Returns 503
on database errors for graceful degradation. Includes 13 tests
covering auth, response structure, edge cases, and error handling.
## Summary
Implements the top-level runScrape() orchestration function that coordinates the entire scraper pipeline end-to-end.
### What it does
- Full pipeline orchestration: Calls fetchPage, parseUnits, convertDataTypes, upsertUnits, insertPrices, markStaleUnits, updateDailySummary in sequence
- dryRun mode: When enabled, parses and validates HTML but skips all database writes
- htmlContent injection: Accepts raw HTML directly, bypassing the fetch step
- New/rented unit calculation: Diffs currently scraped units against previously active units to determine newUnitsCount and rentedUnitsCount for the daily summary
- Run history recording: Every scrape (success or failure) is recorded to the scraper_runs collection via recordScraperRun()
- Structured logging: All pipeline stages log with jobId correlation for traceability
- Error resilience: Catches and handles errors at each stage, ensuring partial failures are logged and recorded
### Test coverage (15 tests)
- Full workflow with mocked dependencies
- Result structure validation and jobId generation
- dryRun mode skips DB writes
- htmlContent bypasses fetch
- Success and failure history recording
- Fetch error handling with retry exhaustion
- Database operation error handling
- New/rented unit count calculation
- Default and scheduled trigger types
- Empty HTML (no units) edge case
Reviewed-on: #16
Co-authored-by: Stephen Minakian <stephenminakian@gmail.com>
Co-committed-by: Stephen Minakian <stephenminakian@gmail.com>