Commit Graph

19 Commits

Author SHA1 Message Date
7e796295cb fix: replace ESM-only uuid package with built-in crypto.randomUUID
All checks were successful
CI/CD Pipeline - Apartment API / Scan Dependencies (pull_request) Successful in 13s
CI/CD Pipeline - Apartment API / Lint & Test (pull_request) Successful in 41s
CI/CD Pipeline - Apartment API / Send Webhook Notification (pull_request) Successful in 2s
CI/CD Pipeline - Apartment API / Build & Push Image (pull_request) Has been skipped
CI/CD Pipeline - Apartment API / Deploy to Production (pull_request) Has been skipped
The uuid v13 package is ESM-only, which causes Jest to fail when
loading routes/admin.js -> jobs/scraperJob.js -> uuid without mocks.
This broke all admin dashboard tests (phase1-4) with silent 404s.

Switched to Node.js built-in crypto.randomUUID() and updated all
test mocks from uuid to crypto accordingly.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-06 22:05:31 -07:00
c30b02681e feat: add POST /api/admin/scraper/run endpoint
Some checks failed
CI/CD Pipeline - Apartment API / Scan Dependencies (pull_request) Successful in 13s
CI/CD Pipeline - Apartment API / Lint & Test (pull_request) Successful in 41s
CI/CD Pipeline - Apartment API / Send Webhook Notification (pull_request) Failing after 1s
CI/CD Pipeline - Apartment API / Build & Push Image (pull_request) Has been skipped
CI/CD Pipeline - Apartment API / Deploy to Production (pull_request) Has been skipped
Implement the manual scraper trigger endpoint with the following:

- Protected by requireAuth and requireAdmin middleware (401/403)
- Mutex lock check via isScraperRunning to prevent concurrent runs (409)
- Generates UUID jobId for tracking async scrape execution
- Returns 202 Accepted immediately without blocking on scrape completion
- Releases mutex lock in .finally() to ensure cleanup on success or failure
- Logs ADMIN_TRIGGER_SCRAPE activity with jobId, dryRun, and
  usingProvidedHtml metadata via activityLogger
- Supports dryRun option (defaults to false) and htmlContent for
  testing with pre-fetched HTML
- Passes trigger: 'manual', jobId, dryRun, and htmlContent to runScrape

Includes 20 tests covering auth, mutex, async execution, activity
logging, dryRun/htmlContent options, and response structure.
2026-02-06 21:20:42 -07:00
c6d480a870 SCRAPE-15: Add getNextScheduledRun() test coverage (#20)
Co-authored-by: Stephen Minakian <stephenminakian@gmail.com>
Co-committed-by: Stephen Minakian <stephenminakian@gmail.com>
2026-02-06 18:21:29 -07:00
0c07baa977 SCRAPE-14: Add graceful shutdown handling (#19)
Co-authored-by: Stephen Minakian <stephenminakian@gmail.com>
Co-committed-by: Stephen Minakian <stephenminakian@gmail.com>
2026-02-06 18:15:24 -07:00
9d789d38fd SCRAPE-13: Create node-cron job initialization (#18)
Co-authored-by: Stephen Minakian <stephenminakian@gmail.com>
Co-committed-by: Stephen Minakian <stephenminakian@gmail.com>
2026-02-06 17:28:53 -07:00
dd0269eb30 SCRAPE-12: Implement in-process mutex (isScraperRunning flag) (#17)
Co-authored-by: Stephen Minakian <stephenminakian@gmail.com>
Co-committed-by: Stephen Minakian <stephenminakian@gmail.com>
2026-02-06 16:27:29 -07:00
0ad4c98abe SCRAPE-11: Create main runScraper() orchestration function (#16)
## Summary

Implements the top-level runScrape() orchestration function that coordinates the entire scraper pipeline end-to-end.

### What it does

- Full pipeline orchestration: Calls fetchPage, parseUnits, convertDataTypes, upsertUnits, insertPrices, markStaleUnits, updateDailySummary in sequence
- dryRun mode: When enabled, parses and validates HTML but skips all database writes
- htmlContent injection: Accepts raw HTML directly, bypassing the fetch step
- New/rented unit calculation: Diffs currently scraped units against previously active units to determine newUnitsCount and rentedUnitsCount for the daily summary
- Run history recording: Every scrape (success or failure) is recorded to the scraper_runs collection via recordScraperRun()
- Structured logging: All pipeline stages log with jobId correlation for traceability
- Error resilience: Catches and handles errors at each stage, ensuring partial failures are logged and recorded

### Test coverage (15 tests)

- Full workflow with mocked dependencies
- Result structure validation and jobId generation
- dryRun mode skips DB writes
- htmlContent bypasses fetch
- Success and failure history recording
- Fetch error handling with retry exhaustion
- Database operation error handling
- New/rented unit count calculation
- Default and scheduled trigger types
- Empty HTML (no units) edge case

Reviewed-on: #16
Co-authored-by: Stephen Minakian <stephenminakian@gmail.com>
Co-committed-by: Stephen Minakian <stephenminakian@gmail.com>
2026-02-06 09:45:32 -07:00
d1f717891a SCRAPE-10: Implement recordScraperRun() for scraper_runs (#15)
Co-authored-by: Stephen Minakian <stephenminakian@gmail.com>
Co-committed-by: Stephen Minakian <stephenminakian@gmail.com>
2026-02-06 02:00:19 -07:00
b4978caf31 SCRAPE-9: Implement updateDailySummary() aggregation (#14)
Co-authored-by: Stephen Minakian <stephenminakian@gmail.com>
Co-committed-by: Stephen Minakian <stephenminakian@gmail.com>
2026-02-06 00:15:30 -07:00
d8adfbf90c SCRAPE-8: Implement markStaleUnits() (#13)
Co-authored-by: Stephen Minakian <stephenminakian@gmail.com>
Co-committed-by: Stephen Minakian <stephenminakian@gmail.com>
2026-02-05 23:48:57 -07:00
2b53288fbe SCRAPE-7: Implement insertPrices() bulk operation (#12)
Co-authored-by: Stephen Minakian <stephenminakian@gmail.com>
Co-committed-by: Stephen Minakian <stephenminakian@gmail.com>
2026-02-05 22:32:02 -07:00
4863dc1824 SCRAPE-6: Implement upsertUnits() bulk operation (#11)
Co-authored-by: Stephen Minakian <stephenminakian@gmail.com>
Co-committed-by: Stephen Minakian <stephenminakian@gmail.com>
2026-02-05 22:10:54 -07:00
0cbeaaf9d5 SCRAPE-5: Implement convertDataTypes() with type parsing helpers (#10)
Co-authored-by: Stephen Minakian <stephenminakian@gmail.com>
Co-committed-by: Stephen Minakian <stephenminakian@gmail.com>
2026-02-05 21:33:58 -07:00
072e80d069 SCRAPE-4: Implement parseUnits() with cheerio selectors (#9)
Co-authored-by: Stephen Minakian <stephenminakian@gmail.com>
Co-committed-by: Stephen Minakian <stephenminakian@gmail.com>
2026-02-05 17:14:01 -07:00
c6beaf8333 SCRAPE-3: Implement fetchPage() with axios, timeout, retry (#8)
Co-authored-by: Stephen Minakian <stephenminakian@gmail.com>
Co-committed-by: Stephen Minakian <stephenminakian@gmail.com>
2026-02-05 15:58:07 -07:00
2993d019c5 SCRAPE-2: Add scraper config constants (#7) 2026-01-31 20:54:43 -07:00
9dba679221 Add scraperLogger.js with credential redaction (#6)
Co-authored-by: Stephen Minakian <stephenminakian@gmail.com>
Co-committed-by: Stephen Minakian <stephenminakian@gmail.com>
2026-01-28 10:52:00 -07:00
ccda2be551 Update tests to expect wrapped response format
- Update all test files to use response.body.data.* instead of
  response.body.* to match the API response convention
- Skip SEC-4.3 rate limiting tests (moved to Phase 5)
- All 202 tests now pass
2026-01-22 14:33:01 -07:00
81e5682ab1 Admin Dashboard Backend (Phases 1-4) (#5)
Some checks failed
Deploy Apartment API / deploy (push) Failing after 5m1s
Co-authored-by: Stephen Minakian <stephenminakian@gmail.com>
Co-committed-by: Stephen Minakian <stephenminakian@gmail.com>
2026-01-22 14:11:43 -07:00