The fixture file was updated with real website data (10 units) but
the tests still referenced old synthetic data (5 units with CCT-*
codes). Updated tests to use real unit codes from the fixture (W2707,
E3205, W2603) and moved edge case tests (image extraction, missing
attributes, unavailable units) to use inline HTML for isolation.
Create three HTML fixture files in __tests__/scraper/fixtures/ to support
deterministic testing of the scraper's HTML parsing logic:
- sample-listing.html: Contains 10 real unit articles extracted from the
live listings page, covering studios, 1BR, and 2BR floor plans with
varied pricing and availability dates.
- sample-listing-empty.html: Minimal page structure with an empty units
section, for testing graceful handling of pages with no listings.
- sample-listing-call.html: Contains units with "Call for pricing" instead
of numeric rent values, for testing the parser's handling of non-numeric
price fields.
## Summary
Implements the top-level runScrape() orchestration function that coordinates the entire scraper pipeline end-to-end.
### What it does
- Full pipeline orchestration: Calls fetchPage, parseUnits, convertDataTypes, upsertUnits, insertPrices, markStaleUnits, updateDailySummary in sequence
- dryRun mode: When enabled, parses and validates HTML but skips all database writes
- htmlContent injection: Accepts raw HTML directly, bypassing the fetch step
- New/rented unit calculation: Diffs currently scraped units against previously active units to determine newUnitsCount and rentedUnitsCount for the daily summary
- Run history recording: Every scrape (success or failure) is recorded to the scraper_runs collection via recordScraperRun()
- Structured logging: All pipeline stages log with jobId correlation for traceability
- Error resilience: Catches and handles errors at each stage, ensuring partial failures are logged and recorded
### Test coverage (15 tests)
- Full workflow with mocked dependencies
- Result structure validation and jobId generation
- dryRun mode skips DB writes
- htmlContent bypasses fetch
- Success and failure history recording
- Fetch error handling with retry exhaustion
- Database operation error handling
- New/rented unit count calculation
- Default and scheduled trigger types
- Empty HTML (no units) edge case
Reviewed-on: #16
Co-authored-by: Stephen Minakian <stephenminakian@gmail.com>
Co-committed-by: Stephen Minakian <stephenminakian@gmail.com>