Implement the manual scraper trigger endpoint with the following:
- Protected by requireAuth and requireAdmin middleware (401/403)
- Mutex lock check via isScraperRunning to prevent concurrent runs (409)
- Generates UUID jobId for tracking async scrape execution
- Returns 202 Accepted immediately without blocking on scrape completion
- Releases mutex lock in .finally() to ensure cleanup on success or failure
- Logs ADMIN_TRIGGER_SCRAPE activity with jobId, dryRun, and
usingProvidedHtml metadata via activityLogger
- Supports dryRun option (defaults to false) and htmlContent for
testing with pre-fetched HTML
- Passes trigger: 'manual', jobId, dryRun, and htmlContent to runScrape
Includes 20 tests covering auth, mutex, async execution, activity
logging, dryRun/htmlContent options, and response structure.
## Summary
Implements the top-level runScrape() orchestration function that coordinates the entire scraper pipeline end-to-end.
### What it does
- Full pipeline orchestration: Calls fetchPage, parseUnits, convertDataTypes, upsertUnits, insertPrices, markStaleUnits, updateDailySummary in sequence
- dryRun mode: When enabled, parses and validates HTML but skips all database writes
- htmlContent injection: Accepts raw HTML directly, bypassing the fetch step
- New/rented unit calculation: Diffs currently scraped units against previously active units to determine newUnitsCount and rentedUnitsCount for the daily summary
- Run history recording: Every scrape (success or failure) is recorded to the scraper_runs collection via recordScraperRun()
- Structured logging: All pipeline stages log with jobId correlation for traceability
- Error resilience: Catches and handles errors at each stage, ensuring partial failures are logged and recorded
### Test coverage (15 tests)
- Full workflow with mocked dependencies
- Result structure validation and jobId generation
- dryRun mode skips DB writes
- htmlContent bypasses fetch
- Success and failure history recording
- Fetch error handling with retry exhaustion
- Database operation error handling
- New/rented unit count calculation
- Default and scheduled trigger types
- Empty HTML (no units) edge case
Reviewed-on: #16
Co-authored-by: Stephen Minakian <stephenminakian@gmail.com>
Co-committed-by: Stephen Minakian <stephenminakian@gmail.com>
- Update all test files to use response.body.data.* instead of
response.body.* to match the API response convention
- Skip SEC-4.3 rate limiting tests (moved to Phase 5)
- All 202 tests now pass