SCRAPE-24: Create HTML fixture files for scraper tests #28

Merged
stephen merged 2 commits from scraper/fixtures into dev 2026-02-07 10:58:14 -07:00
Owner

Summary

Create three HTML fixture files in tests/scraper/fixtures/ to support deterministic, offline testing of the scraper HTML parsing logic.

Files Added

  • sample-listing.html: Contains 10 real unit articles from the live listings page, covering studios, 1BR, and 2BR floor plans with varied pricing and availability dates.
  • sample-listing-empty.html: Minimal page structure with an empty units section, for testing graceful handling of pages with no listings.
  • sample-listing-call.html: Contains units with Call for pricing instead of numeric rent values, for testing the parser handling of non-numeric price fields.

Test Plan

  • Verify all three fixture files are well-formed HTML
  • Confirm sample-listing.html contains 10 unit articles with expected data attributes
  • Confirm sample-listing-empty.html has the correct page structure but no unit entries
  • Confirm sample-listing-call.html contains units with Call for pricing text
## Summary Create three HTML fixture files in __tests__/scraper/fixtures/ to support deterministic, offline testing of the scraper HTML parsing logic. ### Files Added - sample-listing.html: Contains 10 real unit articles from the live listings page, covering studios, 1BR, and 2BR floor plans with varied pricing and availability dates. - sample-listing-empty.html: Minimal page structure with an empty units section, for testing graceful handling of pages with no listings. - sample-listing-call.html: Contains units with Call for pricing instead of numeric rent values, for testing the parser handling of non-numeric price fields. ## Test Plan - Verify all three fixture files are well-formed HTML - Confirm sample-listing.html contains 10 unit articles with expected data attributes - Confirm sample-listing-empty.html has the correct page structure but no unit entries - Confirm sample-listing-call.html contains units with Call for pricing text
stephen added 1 commit 2026-02-07 10:39:51 -07:00
Add HTML fixture files for scraper unit tests
Some checks failed
CI/CD Pipeline - Apartment API / Scan Dependencies (pull_request) Successful in 13s
CI/CD Pipeline - Apartment API / Lint & Test (pull_request) Successful in 43s
CI/CD Pipeline - Apartment API / Send Webhook Notification (pull_request) Failing after 2s
CI/CD Pipeline - Apartment API / Build & Push Image (pull_request) Has been skipped
CI/CD Pipeline - Apartment API / Deploy to Production (pull_request) Has been skipped
1080121386
Create three HTML fixture files in __tests__/scraper/fixtures/ to support
deterministic testing of the scraper's HTML parsing logic:

- sample-listing.html: Contains 10 real unit articles extracted from the
  live listings page, covering studios, 1BR, and 2BR floor plans with
  varied pricing and availability dates.

- sample-listing-empty.html: Minimal page structure with an empty units
  section, for testing graceful handling of pages with no listings.

- sample-listing-call.html: Contains units with "Call for pricing" instead
  of numeric rent values, for testing the parser's handling of non-numeric
  price fields.
Author
Owner

❌ Pipeline Failed

Test: ❌ Failed

> apartment-api@1.0.0 test
> jest --runInBand

  console.log
    {"timestamp":"2026-02-07T17:40:10.060Z","level":"error","message":"Async scrape failed","jobId":"scraper-admin","context":{"errorType":"Error","errorMessage":"Scrape failed"}}

      at log (services/scraperLogger.js:34:13)

  console.log
    {"timestamp":"2026-02-07T17:40:11.048Z","level":"error","message":"Error fetching scraper status","jobId":"scraper-admin","context":{"errorType":"Error","errorMessage":"Database connection lost"}}

      at log (services/scraperLogger.js:34:13)

  console.log
    {"timestamp":"2026-02-07T17:40:11.283Z","level":"error","message":"Error fetching scraper history","jobId":"scraper-admin","context":{"errorType":"Error","errorMessage":"Database connection lost"}}

      at log (services/scraperLogger.js:34:13)

PASS __tests__/scraper/scraperRoutes.test.js
  Scraper Routes
    POST /api/admin/scraper/run
      authentication and authorization
        ✓ should return 401 without auth token (195 ms)
        ✓ should return 403 for non-admin user (30 ms)
      successful scrape trigger
        ✓ should return 202 with jobId for admin user (17 ms)
        ✓ should return response with correct structure (jobId, status, message, dryRun) (14 ms)
        ✓ should return a UUID-formatted jobId (12 ms)
      scraper already running
        ✓ should return 409 when scraper is already running (12 ms)
        ✓ should check isScraperRunning before starting (11 ms)
      mutex lock behavior
        ✓ should acquire lock before starting scrape (11 ms)
        ✓ should release lock after scrape completes (116 ms)
        ✓ should release lock even when scrape fails (157 ms)
      activity logging
        ✓ should log ADMIN_TRIGGER_SCRAPE activity (15 ms)
        ✓ should log dryRun flag in activity metadata (15 ms)
      dryRun option
        ✓ should handle dryRun option in request body (15 ms)
        ✓ should pass dryRun option to runScrape (112 ms)
        ✓ should default dryRun to false when not provided (12 ms)
      htmlContent option
        ✓ should pass htmlContent option to runScrape (115 ms)
        ✓ should log usingProvidedHtml in activity metadata when htmlContent provided (15 ms)
      scrape execution
        ✓ should call runScrape with trigger "manual" (115 ms)
        ✓ should call runScrape with the generated jobId (115 ms)
        ✓ should return 202 immediately without waiting for scrape to finish (116 ms)
    GET /api/admin/scraper/status
      Authentication and Authorization
        ✓ should return 401 without authentication (58 ms)
        ✓ should return 403 for non-admin user (9 ms)
      Successful status response (admin)
        ✓ should return 200 with status for admin (12 ms)
        ✓ should return currentStatus "idle" when scraper is not running (10 ms)
        ✓ should return currentStatus "running" with runningJobId when scraper is running (10 ms)
        ✓ should return lastRun as null when no runs exist (10 ms)
        ✓ should return lastRun object with correct fields when a run exists (21 ms)
        ✓ should return lastRun with errors array when last run had errors (11 ms)
        ✓ should return the most recent run when multiple runs exist (11 ms)
        ✓ should return nextScheduledRun as ISO timestamp (10 ms)
        ✓ should return nextScheduledRun as null when scheduler is disabled (10 ms)
        ✓ should return schedule cron expression (10 ms)
      Error handling
        ✓ should return 503 on database error (10 ms)
    GET /api/admin/scraper/history
      Authentication and Authorization
        ✓ should return 401 without authentication (9 ms)
        ✓ should return 403 for non-admin user (8 ms)
      Successful history response (admin)
        ✓ should return 200 with data array for admin (11 ms)
        ✓ should return empty array when no runs exist (10 ms)
        ✓ should return run records with correct fields (12 ms)
        ✓ should sort results by startedAt descending (newest first) (13 ms)
        ✓ should default limit to 30 records (16 ms)
        ✓ should respect custom limit query parameter (13 ms)
        ✓ should cap limit at 100 even if higher value requested (10 ms)
        ✓ should apply offset to skip records (11 ms)
        ✓ should return empty array when offset exceeds total records (10 ms)
        ✓ should accept offset of 0 as valid (10 ms)
        ✓ should accept limit of 1 as valid (11 ms)
        ✓ should accept limit of 100 as valid (9 ms)
      Validation errors
        ✓ should return 400 for limit less than 1 (13 ms)
        ✓ should return 400 for negative limit (9 ms)
        ✓ should return 400 for non-numeric limit (9 ms)
        ✓ should return 400 for negative offset (9 ms)
        ✓ should return 400 for non-numeric offset (10 ms)
        ✓ should return specific error message for invalid limit (9 ms)
        ✓ should return specific error message for invalid offset (9 ms)
      Error handling
        ✓ should return 503 on database error (9 ms)
    Authentication & Authorization Middleware
      missing JWT token (401)
        ✓ should return 401 for POST /api/admin/scraper/run without auth token (4 ms)
        ✓ should return 401 for GET /api/admin/scraper/status without auth token (5 ms)
        ✓ should return 401 for GET /api/admin/scraper/history without auth token (5 ms)
      invalid JWT token (401)
        ✓ should return 401 for POST /api/admin/scraper/run with invalid token (5 ms)
        ✓ should return 401 for GET /api/admin/scraper/status with invalid token (6 ms)
        ✓ should return 401 for GET /api/admin/scraper/history with invalid token (4 ms)
      expired JWT token (401)
        ✓ should return 401 for POST /api/admin/scraper/run with expired token (8 ms)
        ✓ should return 401 for GET /api/admin/scraper/status with expired token (6 ms)
        ✓ should return 401 for GET /api/admin/scraper/history with expired token (6 ms)
      token signed with wrong secret (401)
        ✓ should return 401 for POST /api/admin/scraper/run with wrong-secret token (7 ms)
        ✓ should return 401 for GET /api/admin/scraper/status with wrong-secret token (6 ms)
        ✓ should return 401 for GET /api/admin/scraper/history with wrong-secret token (7 ms)
      disabled user (401)
        ✓ should return 401 for POST /api/admin/scraper/run when user is disabled (11 ms)
        ✓ should return 401 for GET /api/admin/scraper/status when user is disabled (7 ms)
        ✓ should return 401 for GET /api/admin/scraper/history when user is disabled (8 ms)
      non-existent user (401)
        ✓ should return 401 for POST /api/admin/scraper/run when user does not exist (7 ms)
        ✓ should return 401 for GET /api/admin/scraper/status when user does not exist (6 ms)
        ✓ should return 401 for GET /api/admin/scraper/history when user does not exist (6 ms)
      non-admin user (403)
        ✓ should return 403 for POST /api/admin/scraper/run for non-admin user (7 ms)
        ✓ should return 403 for GET /api/admin/scraper/status for non-admin user (7 ms)
        ✓ should return 403 for GET /api/admin/scraper/history for non-admin user (7 ms)
      valid admin user (200/202)
        ✓ should return 202 for POST /api/admin/scraper/run for admin user (9 ms)
        ✓ should return 200 for GET /api/admin/scraper/status for admin user (9 ms)
        ✓ should return 200 for GET /api/admin/scraper/history for admin user (9 ms)
      rate limiting
        ✓ should allow the first 5 requests within the rate limit window (286 ms)
        ✓ should return 429 when the 6th request exceeds the rate limit (302 ms)
        ✓ should include Retry-After header in 429 response (300 ms)
        ✓ should track rate limits per user (different admins have separate limits) (309 ms)
        ✓ should reset rate limit after window expires (305 ms)
        ✓ should check rate limit before mutex (rate limit takes precedence) (300 ms)
      middleware ordering
        ✓ should return 401 (not 403) when token is missing, even for non-admin scenario (6 ms)
        ✓ should return 401 (not 403) for disabled admin user (10 ms)

PASS __tests__/scraper/scraperJob.test.js
  scraperJob - mutex lock
    isScraperRunning()
      ✓ should return false when not running (2 ms)
      ✓ should return true after acquireLock()
      ✓ should return false after releaseLock() (1 ms)
    getCurrentJobId()
      ✓ should return null when not running (1 ms)
      ✓ should return jobId when running
      ✓ should return null after releaseLock() (1 ms)
    acquireLock()
      ✓ should return true when lock is available (1 ms)
      ✓ should return false when lock is already held
      ✓ should set isScraperRunning to true (1 ms)
      ✓ should set currentJobId to the provided jobId
      ✓ should not overwrite existing lock when already held (1 ms)
    releaseLock()
      ✓ should clear running state (1 ms)
      ✓ should clear currentJobId
      ✓ should allow a new lock to be acquired after release (1 ms)
      ✓ should be safe to call when no lock is held (1 ms)
    module exports
      ✓ should export isScraperRunning as a function
      ✓ should export getCurrentJobId as a function (1 ms)
      ✓ should export acquireLock as a function
      ✓ should export releaseLock as a function (1 ms)
      ✓ should export initializeScheduler as a function
      ✓ should export stopScheduler as a function (1 ms)
      ✓ should export getScheduleExpression as a function
      ✓ should export getNextScheduledRun as a function
  scraperJob - scheduler
    initializeScheduler()
      ✓ should not initialize when SCRAPER_ENABLED is false (2 ms)
      ✓ should validate cron expression with cron.validate() (39 ms)
      ✓ should fall back to default schedule 0 6 * * * if cron expression is invalid (18 ms)
      ✓ should enforce minimum 1-hour interval - adjust too-frequent schedules (17 ms)
      ✓ should enforce minimum 1-hour interval for every-minute schedule (17 ms)
      ✓ should create cron j
## ❌ Pipeline Failed ### Test: ❌ Failed ``` > apartment-api@1.0.0 test > jest --runInBand console.log {"timestamp":"2026-02-07T17:40:10.060Z","level":"error","message":"Async scrape failed","jobId":"scraper-admin","context":{"errorType":"Error","errorMessage":"Scrape failed"}} at log (services/scraperLogger.js:34:13) console.log {"timestamp":"2026-02-07T17:40:11.048Z","level":"error","message":"Error fetching scraper status","jobId":"scraper-admin","context":{"errorType":"Error","errorMessage":"Database connection lost"}} at log (services/scraperLogger.js:34:13) console.log {"timestamp":"2026-02-07T17:40:11.283Z","level":"error","message":"Error fetching scraper history","jobId":"scraper-admin","context":{"errorType":"Error","errorMessage":"Database connection lost"}} at log (services/scraperLogger.js:34:13) PASS __tests__/scraper/scraperRoutes.test.js Scraper Routes POST /api/admin/scraper/run authentication and authorization ✓ should return 401 without auth token (195 ms) ✓ should return 403 for non-admin user (30 ms) successful scrape trigger ✓ should return 202 with jobId for admin user (17 ms) ✓ should return response with correct structure (jobId, status, message, dryRun) (14 ms) ✓ should return a UUID-formatted jobId (12 ms) scraper already running ✓ should return 409 when scraper is already running (12 ms) ✓ should check isScraperRunning before starting (11 ms) mutex lock behavior ✓ should acquire lock before starting scrape (11 ms) ✓ should release lock after scrape completes (116 ms) ✓ should release lock even when scrape fails (157 ms) activity logging ✓ should log ADMIN_TRIGGER_SCRAPE activity (15 ms) ✓ should log dryRun flag in activity metadata (15 ms) dryRun option ✓ should handle dryRun option in request body (15 ms) ✓ should pass dryRun option to runScrape (112 ms) ✓ should default dryRun to false when not provided (12 ms) htmlContent option ✓ should pass htmlContent option to runScrape (115 ms) ✓ should log usingProvidedHtml in activity metadata when htmlContent provided (15 ms) scrape execution ✓ should call runScrape with trigger "manual" (115 ms) ✓ should call runScrape with the generated jobId (115 ms) ✓ should return 202 immediately without waiting for scrape to finish (116 ms) GET /api/admin/scraper/status Authentication and Authorization ✓ should return 401 without authentication (58 ms) ✓ should return 403 for non-admin user (9 ms) Successful status response (admin) ✓ should return 200 with status for admin (12 ms) ✓ should return currentStatus "idle" when scraper is not running (10 ms) ✓ should return currentStatus "running" with runningJobId when scraper is running (10 ms) ✓ should return lastRun as null when no runs exist (10 ms) ✓ should return lastRun object with correct fields when a run exists (21 ms) ✓ should return lastRun with errors array when last run had errors (11 ms) ✓ should return the most recent run when multiple runs exist (11 ms) ✓ should return nextScheduledRun as ISO timestamp (10 ms) ✓ should return nextScheduledRun as null when scheduler is disabled (10 ms) ✓ should return schedule cron expression (10 ms) Error handling ✓ should return 503 on database error (10 ms) GET /api/admin/scraper/history Authentication and Authorization ✓ should return 401 without authentication (9 ms) ✓ should return 403 for non-admin user (8 ms) Successful history response (admin) ✓ should return 200 with data array for admin (11 ms) ✓ should return empty array when no runs exist (10 ms) ✓ should return run records with correct fields (12 ms) ✓ should sort results by startedAt descending (newest first) (13 ms) ✓ should default limit to 30 records (16 ms) ✓ should respect custom limit query parameter (13 ms) ✓ should cap limit at 100 even if higher value requested (10 ms) ✓ should apply offset to skip records (11 ms) ✓ should return empty array when offset exceeds total records (10 ms) ✓ should accept offset of 0 as valid (10 ms) ✓ should accept limit of 1 as valid (11 ms) ✓ should accept limit of 100 as valid (9 ms) Validation errors ✓ should return 400 for limit less than 1 (13 ms) ✓ should return 400 for negative limit (9 ms) ✓ should return 400 for non-numeric limit (9 ms) ✓ should return 400 for negative offset (9 ms) ✓ should return 400 for non-numeric offset (10 ms) ✓ should return specific error message for invalid limit (9 ms) ✓ should return specific error message for invalid offset (9 ms) Error handling ✓ should return 503 on database error (9 ms) Authentication & Authorization Middleware missing JWT token (401) ✓ should return 401 for POST /api/admin/scraper/run without auth token (4 ms) ✓ should return 401 for GET /api/admin/scraper/status without auth token (5 ms) ✓ should return 401 for GET /api/admin/scraper/history without auth token (5 ms) invalid JWT token (401) ✓ should return 401 for POST /api/admin/scraper/run with invalid token (5 ms) ✓ should return 401 for GET /api/admin/scraper/status with invalid token (6 ms) ✓ should return 401 for GET /api/admin/scraper/history with invalid token (4 ms) expired JWT token (401) ✓ should return 401 for POST /api/admin/scraper/run with expired token (8 ms) ✓ should return 401 for GET /api/admin/scraper/status with expired token (6 ms) ✓ should return 401 for GET /api/admin/scraper/history with expired token (6 ms) token signed with wrong secret (401) ✓ should return 401 for POST /api/admin/scraper/run with wrong-secret token (7 ms) ✓ should return 401 for GET /api/admin/scraper/status with wrong-secret token (6 ms) ✓ should return 401 for GET /api/admin/scraper/history with wrong-secret token (7 ms) disabled user (401) ✓ should return 401 for POST /api/admin/scraper/run when user is disabled (11 ms) ✓ should return 401 for GET /api/admin/scraper/status when user is disabled (7 ms) ✓ should return 401 for GET /api/admin/scraper/history when user is disabled (8 ms) non-existent user (401) ✓ should return 401 for POST /api/admin/scraper/run when user does not exist (7 ms) ✓ should return 401 for GET /api/admin/scraper/status when user does not exist (6 ms) ✓ should return 401 for GET /api/admin/scraper/history when user does not exist (6 ms) non-admin user (403) ✓ should return 403 for POST /api/admin/scraper/run for non-admin user (7 ms) ✓ should return 403 for GET /api/admin/scraper/status for non-admin user (7 ms) ✓ should return 403 for GET /api/admin/scraper/history for non-admin user (7 ms) valid admin user (200/202) ✓ should return 202 for POST /api/admin/scraper/run for admin user (9 ms) ✓ should return 200 for GET /api/admin/scraper/status for admin user (9 ms) ✓ should return 200 for GET /api/admin/scraper/history for admin user (9 ms) rate limiting ✓ should allow the first 5 requests within the rate limit window (286 ms) ✓ should return 429 when the 6th request exceeds the rate limit (302 ms) ✓ should include Retry-After header in 429 response (300 ms) ✓ should track rate limits per user (different admins have separate limits) (309 ms) ✓ should reset rate limit after window expires (305 ms) ✓ should check rate limit before mutex (rate limit takes precedence) (300 ms) middleware ordering ✓ should return 401 (not 403) when token is missing, even for non-admin scenario (6 ms) ✓ should return 401 (not 403) for disabled admin user (10 ms) PASS __tests__/scraper/scraperJob.test.js scraperJob - mutex lock isScraperRunning() ✓ should return false when not running (2 ms) ✓ should return true after acquireLock() ✓ should return false after releaseLock() (1 ms) getCurrentJobId() ✓ should return null when not running (1 ms) ✓ should return jobId when running ✓ should return null after releaseLock() (1 ms) acquireLock() ✓ should return true when lock is available (1 ms) ✓ should return false when lock is already held ✓ should set isScraperRunning to true (1 ms) ✓ should set currentJobId to the provided jobId ✓ should not overwrite existing lock when already held (1 ms) releaseLock() ✓ should clear running state (1 ms) ✓ should clear currentJobId ✓ should allow a new lock to be acquired after release (1 ms) ✓ should be safe to call when no lock is held (1 ms) module exports ✓ should export isScraperRunning as a function ✓ should export getCurrentJobId as a function (1 ms) ✓ should export acquireLock as a function ✓ should export releaseLock as a function (1 ms) ✓ should export initializeScheduler as a function ✓ should export stopScheduler as a function (1 ms) ✓ should export getScheduleExpression as a function ✓ should export getNextScheduledRun as a function scraperJob - scheduler initializeScheduler() ✓ should not initialize when SCRAPER_ENABLED is false (2 ms) ✓ should validate cron expression with cron.validate() (39 ms) ✓ should fall back to default schedule 0 6 * * * if cron expression is invalid (18 ms) ✓ should enforce minimum 1-hour interval - adjust too-frequent schedules (17 ms) ✓ should enforce minimum 1-hour interval for every-minute schedule (17 ms) ✓ should create cron j ```
stephen force-pushed scraper/fixtures from 1080121386 to fcfd35aa3f 2026-02-07 10:56:42 -07:00 Compare
stephen merged commit 1f7faab1c7 into dev 2026-02-07 10:58:14 -07:00
stephen deleted branch scraper/fixtures 2026-02-07 10:58:14 -07:00
Sign in to join this conversation.
No Reviewers
No Label
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: stephen/apartment-dashboard-api#28
No description provided.