SCRAPE-16: Implement POST /admin/scraper/run endpoint #21

Merged
stephen merged 2 commits from scraper/api-run into dev 2026-02-06 22:07:11 -07:00
Owner

Summary

Adds the POST /api/admin/scraper/run endpoint for manually triggering scraper runs from the admin dashboard.

Endpoint Details

  • Route: POST /api/admin/scraper/run
  • Auth: Protected by requireAuth + requireAdmin middleware

Key Features

  • Mutex check: Returns 409 Conflict if a scrape is already running (isScraperRunning)
  • UUID job tracking: Generates a unique jobId for each scrape run
  • Async execution: Fires off the scrape asynchronously and returns immediately
  • 202 Accepted response: Returns { jobId, status: 'started', message, dryRun } without waiting for scrape completion
  • Activity logging: Logs ADMIN_TRIGGER_SCRAPE with metadata (jobId, dryRun, usingProvidedHtml)
  • Lock cleanup: Releases mutex in .finally() ensuring cleanup on both success and failure
  • dryRun support: Optional boolean flag (defaults to false) passed through to runScrape
  • htmlContent support: Optional pre-fetched HTML for testing without live fetching

Response Codes

Code Condition
202 Scrape triggered successfully
401 Missing or invalid authentication
403 User is not an admin
409 Scraper is already running

Test Coverage

20 tests covering:

  • Authentication and authorization (401/403)
  • Successful scrape trigger and response structure
  • Mutex lock acquisition, release on success, and release on failure
  • Activity logging with correct metadata
  • dryRun and htmlContent option handling
  • Async execution behavior (non-blocking 202 response)
## Summary Adds the `POST /api/admin/scraper/run` endpoint for manually triggering scraper runs from the admin dashboard. ### Endpoint Details - **Route**: `POST /api/admin/scraper/run` - **Auth**: Protected by `requireAuth` + `requireAdmin` middleware ### Key Features - **Mutex check**: Returns 409 Conflict if a scrape is already running (`isScraperRunning`) - **UUID job tracking**: Generates a unique jobId for each scrape run - **Async execution**: Fires off the scrape asynchronously and returns immediately - **202 Accepted response**: Returns `{ jobId, status: 'started', message, dryRun }` without waiting for scrape completion - **Activity logging**: Logs `ADMIN_TRIGGER_SCRAPE` with metadata (jobId, dryRun, usingProvidedHtml) - **Lock cleanup**: Releases mutex in `.finally()` ensuring cleanup on both success and failure - **dryRun support**: Optional boolean flag (defaults to false) passed through to runScrape - **htmlContent support**: Optional pre-fetched HTML for testing without live fetching ### Response Codes | Code | Condition | |------|-----------| | 202 | Scrape triggered successfully | | 401 | Missing or invalid authentication | | 403 | User is not an admin | | 409 | Scraper is already running | ### Test Coverage 20 tests covering: - Authentication and authorization (401/403) - Successful scrape trigger and response structure - Mutex lock acquisition, release on success, and release on failure - Activity logging with correct metadata - dryRun and htmlContent option handling - Async execution behavior (non-blocking 202 response)
stephen added 1 commit 2026-02-06 21:21:21 -07:00
feat: add POST /api/admin/scraper/run endpoint
Some checks failed
CI/CD Pipeline - Apartment API / Scan Dependencies (pull_request) Successful in 13s
CI/CD Pipeline - Apartment API / Lint & Test (pull_request) Successful in 41s
CI/CD Pipeline - Apartment API / Send Webhook Notification (pull_request) Failing after 1s
CI/CD Pipeline - Apartment API / Build & Push Image (pull_request) Has been skipped
CI/CD Pipeline - Apartment API / Deploy to Production (pull_request) Has been skipped
c30b02681e
Implement the manual scraper trigger endpoint with the following:

- Protected by requireAuth and requireAdmin middleware (401/403)
- Mutex lock check via isScraperRunning to prevent concurrent runs (409)
- Generates UUID jobId for tracking async scrape execution
- Returns 202 Accepted immediately without blocking on scrape completion
- Releases mutex lock in .finally() to ensure cleanup on success or failure
- Logs ADMIN_TRIGGER_SCRAPE activity with jobId, dryRun, and
  usingProvidedHtml metadata via activityLogger
- Supports dryRun option (defaults to false) and htmlContent for
  testing with pre-fetched HTML
- Passes trigger: 'manual', jobId, dryRun, and htmlContent to runScrape

Includes 20 tests covering auth, mutex, async execution, activity
logging, dryRun/htmlContent options, and response structure.
stephen added 1 commit 2026-02-06 22:05:37 -07:00
fix: replace ESM-only uuid package with built-in crypto.randomUUID
All checks were successful
CI/CD Pipeline - Apartment API / Scan Dependencies (pull_request) Successful in 13s
CI/CD Pipeline - Apartment API / Lint & Test (pull_request) Successful in 41s
CI/CD Pipeline - Apartment API / Send Webhook Notification (pull_request) Successful in 2s
CI/CD Pipeline - Apartment API / Build & Push Image (pull_request) Has been skipped
CI/CD Pipeline - Apartment API / Deploy to Production (pull_request) Has been skipped
7e796295cb
The uuid v13 package is ESM-only, which causes Jest to fail when
loading routes/admin.js -> jobs/scraperJob.js -> uuid without mocks.
This broke all admin dashboard tests (phase1-4) with silent 404s.

Switched to Node.js built-in crypto.randomUUID() and updated all
test mocks from uuid to crypto accordingly.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
stephen merged commit daa234428e into dev 2026-02-06 22:07:11 -07:00
stephen deleted branch scraper/api-run 2026-02-06 22:07:11 -07:00
Sign in to join this conversation.
No Reviewers
No Label
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: stephen/apartment-dashboard-api#21
No description provided.