Commit Graph

18 Commits

Author SHA1 Message Date
ef15beb622 Add GET /api/admin/scraper/history endpoint with pagination
All checks were successful
CI/CD Pipeline - Apartment API / Scan Dependencies (pull_request) Successful in 13s
CI/CD Pipeline - Apartment API / Lint & Test (pull_request) Successful in 42s
CI/CD Pipeline - Apartment API / Send Webhook Notification (pull_request) Successful in 2s
CI/CD Pipeline - Apartment API / Build & Push Image (pull_request) Has been skipped
CI/CD Pipeline - Apartment API / Deploy to Production (pull_request) Has been skipped
Implement the scraper history endpoint that returns past scrape run
records from the scraper_runs collection. Features include:

- GET /api/admin/scraper/history protected by requireAuth + requireAdmin
- Pagination via limit (1-100, default 30) and offset (default 0) params
- Input validation returning 400 for invalid limit/offset values
- Results sorted by startedAt descending (newest first)
- Graceful error handling returning 503 on database failures
- 20 new tests covering auth, pagination, validation, and error cases
2026-02-06 22:33:01 -07:00
9d0b9debbb SCRAPE-17: Implement GET /admin/scraper/status endpoint (#22)
Co-authored-by: Stephen Minakian <stephenminakian@gmail.com>
Co-committed-by: Stephen Minakian <stephenminakian@gmail.com>
2026-02-06 22:20:06 -07:00
daa234428e SCRAPE-16: Implement POST /admin/scraper/run endpoint (#21)
Co-authored-by: Stephen Minakian <stephenminakian@gmail.com>
Co-committed-by: Stephen Minakian <stephenminakian@gmail.com>
2026-02-06 22:07:10 -07:00
c6d480a870 SCRAPE-15: Add getNextScheduledRun() test coverage (#20)
Co-authored-by: Stephen Minakian <stephenminakian@gmail.com>
Co-committed-by: Stephen Minakian <stephenminakian@gmail.com>
2026-02-06 18:21:29 -07:00
0c07baa977 SCRAPE-14: Add graceful shutdown handling (#19)
Co-authored-by: Stephen Minakian <stephenminakian@gmail.com>
Co-committed-by: Stephen Minakian <stephenminakian@gmail.com>
2026-02-06 18:15:24 -07:00
9d789d38fd SCRAPE-13: Create node-cron job initialization (#18)
Co-authored-by: Stephen Minakian <stephenminakian@gmail.com>
Co-committed-by: Stephen Minakian <stephenminakian@gmail.com>
2026-02-06 17:28:53 -07:00
dd0269eb30 SCRAPE-12: Implement in-process mutex (isScraperRunning flag) (#17)
Co-authored-by: Stephen Minakian <stephenminakian@gmail.com>
Co-committed-by: Stephen Minakian <stephenminakian@gmail.com>
2026-02-06 16:27:29 -07:00
0ad4c98abe SCRAPE-11: Create main runScraper() orchestration function (#16)
## Summary

Implements the top-level runScrape() orchestration function that coordinates the entire scraper pipeline end-to-end.

### What it does

- Full pipeline orchestration: Calls fetchPage, parseUnits, convertDataTypes, upsertUnits, insertPrices, markStaleUnits, updateDailySummary in sequence
- dryRun mode: When enabled, parses and validates HTML but skips all database writes
- htmlContent injection: Accepts raw HTML directly, bypassing the fetch step
- New/rented unit calculation: Diffs currently scraped units against previously active units to determine newUnitsCount and rentedUnitsCount for the daily summary
- Run history recording: Every scrape (success or failure) is recorded to the scraper_runs collection via recordScraperRun()
- Structured logging: All pipeline stages log with jobId correlation for traceability
- Error resilience: Catches and handles errors at each stage, ensuring partial failures are logged and recorded

### Test coverage (15 tests)

- Full workflow with mocked dependencies
- Result structure validation and jobId generation
- dryRun mode skips DB writes
- htmlContent bypasses fetch
- Success and failure history recording
- Fetch error handling with retry exhaustion
- Database operation error handling
- New/rented unit count calculation
- Default and scheduled trigger types
- Empty HTML (no units) edge case

Reviewed-on: #16
Co-authored-by: Stephen Minakian <stephenminakian@gmail.com>
Co-committed-by: Stephen Minakian <stephenminakian@gmail.com>
2026-02-06 09:45:32 -07:00
d1f717891a SCRAPE-10: Implement recordScraperRun() for scraper_runs (#15)
Co-authored-by: Stephen Minakian <stephenminakian@gmail.com>
Co-committed-by: Stephen Minakian <stephenminakian@gmail.com>
2026-02-06 02:00:19 -07:00
b4978caf31 SCRAPE-9: Implement updateDailySummary() aggregation (#14)
Co-authored-by: Stephen Minakian <stephenminakian@gmail.com>
Co-committed-by: Stephen Minakian <stephenminakian@gmail.com>
2026-02-06 00:15:30 -07:00
d8adfbf90c SCRAPE-8: Implement markStaleUnits() (#13)
Co-authored-by: Stephen Minakian <stephenminakian@gmail.com>
Co-committed-by: Stephen Minakian <stephenminakian@gmail.com>
2026-02-05 23:48:57 -07:00
2b53288fbe SCRAPE-7: Implement insertPrices() bulk operation (#12)
Co-authored-by: Stephen Minakian <stephenminakian@gmail.com>
Co-committed-by: Stephen Minakian <stephenminakian@gmail.com>
2026-02-05 22:32:02 -07:00
4863dc1824 SCRAPE-6: Implement upsertUnits() bulk operation (#11)
Co-authored-by: Stephen Minakian <stephenminakian@gmail.com>
Co-committed-by: Stephen Minakian <stephenminakian@gmail.com>
2026-02-05 22:10:54 -07:00
0cbeaaf9d5 SCRAPE-5: Implement convertDataTypes() with type parsing helpers (#10)
Co-authored-by: Stephen Minakian <stephenminakian@gmail.com>
Co-committed-by: Stephen Minakian <stephenminakian@gmail.com>
2026-02-05 21:33:58 -07:00
072e80d069 SCRAPE-4: Implement parseUnits() with cheerio selectors (#9)
Co-authored-by: Stephen Minakian <stephenminakian@gmail.com>
Co-committed-by: Stephen Minakian <stephenminakian@gmail.com>
2026-02-05 17:14:01 -07:00
c6beaf8333 SCRAPE-3: Implement fetchPage() with axios, timeout, retry (#8)
Co-authored-by: Stephen Minakian <stephenminakian@gmail.com>
Co-committed-by: Stephen Minakian <stephenminakian@gmail.com>
2026-02-05 15:58:07 -07:00
2993d019c5 SCRAPE-2: Add scraper config constants (#7) 2026-01-31 20:54:43 -07:00
9dba679221 Add scraperLogger.js with credential redaction (#6)
Co-authored-by: Stephen Minakian <stephenminakian@gmail.com>
Co-committed-by: Stephen Minakian <stephenminakian@gmail.com>
2026-01-28 10:52:00 -07:00