Accept logger as second parameter to match the pattern used by all other
functions in scraperService.js (recordScraperRun, upsertUnits, etc.).
Replaces console.log with logger.info for production code consistency.
Add createScraperIndexes() to create indexes on the scraper_runs
collection: a compound index on status+startedAt for active job queries
and a descending index on startedAt for recent run lookups.
Change default collection names from production tables
(units_migration_test, unit_prices_migration_test) to dedicated
validation collections (units_scraper, unit_prices_scraper). This
enables the Node.js scraper to run in parallel with the existing Python
scraper during validation without interfering with production data.
Update all scraper test files to reference the new default collection
names.
Truncate lint/test output to 10000 chars at the source (CI job) before
writing to GITHUB_OUTPUT, instead of in the downstream notify job.
Previously the full output was passed as an env var between jobs, which
could exceed Linux's ARG_MAX limit and prevent bash from launching.
Phase tests share a single MongoMemoryServer database and conflict
when run in parallel. Sequential execution is needed until tests
use isolated databases per file.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Combine lint and test into a single 'ci' job (eliminates duplicate
checkout + npm ci, saving ~60-90s)
- Remove --runInBand flag so Jest parallelizes across worker pools
- Remove unused mongo:7 service container (tests use MongoMemoryServer)
- Fix failure detection: check step outcomes instead of job result,
which was always 'success' due to continue-on-error
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
## Summary
Implements the top-level runScrape() orchestration function that coordinates the entire scraper pipeline end-to-end.
### What it does
- Full pipeline orchestration: Calls fetchPage, parseUnits, convertDataTypes, upsertUnits, insertPrices, markStaleUnits, updateDailySummary in sequence
- dryRun mode: When enabled, parses and validates HTML but skips all database writes
- htmlContent injection: Accepts raw HTML directly, bypassing the fetch step
- New/rented unit calculation: Diffs currently scraped units against previously active units to determine newUnitsCount and rentedUnitsCount for the daily summary
- Run history recording: Every scrape (success or failure) is recorded to the scraper_runs collection via recordScraperRun()
- Structured logging: All pipeline stages log with jobId correlation for traceability
- Error resilience: Catches and handles errors at each stage, ensuring partial failures are logged and recorded
### Test coverage (15 tests)
- Full workflow with mocked dependencies
- Result structure validation and jobId generation
- dryRun mode skips DB writes
- htmlContent bypasses fetch
- Success and failure history recording
- Fetch error handling with retry exhaustion
- Database operation error handling
- New/rented unit count calculation
- Default and scheduled trigger types
- Empty HTML (no units) edge case
Reviewed-on: #16
Co-authored-by: Stephen Minakian <stephenminakian@gmail.com>
Co-committed-by: Stephen Minakian <stephenminakian@gmail.com>
- Add ESLint 9 with flat config for Node.js linting
- Add lint and lint:fix npm scripts
- Add lint job to CI/CD pipeline
- Add notify job to send test/lint results to n8n webhook
- Webhook reports pass/fail status with failure details
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
- Update to node:20-alpine base image
- Add apk upgrade to fix OS-level vulnerabilities
- Update npm to latest to fix bundled package vulnerabilities
- Add scan-deps job that runs in parallel with tests
- Replace deprecated --only=production with --omit=dev
Add yesterdayPrice and priceChangeToday fields to show price delta
between today and yesterday instead of high-low spread.
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
Use pre-configured Docker credentials on server instead of passing
Harbor credentials through SSH script, avoiding shell interpolation
issues with special characters in robot account username.
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
Gitea Actions doesn't support type=gha caching, causing timeout errors.
Removing cache config to get builds working. Can add registry-based
caching later if needed.
- Use SHA-based image tags instead of 'latest' for deployments
- Pass exact image tag from build job to deploy job
- Add default for SSH_PORT
- Health check now fails deploy if not passing
- Added script_stop for proper error handling
- Add image variable substitution to docker-compose.yml
- Build image in CI, push to Harbor registry
- Deploy pulls from Harbor instead of rebuilding
- Supports both local dev (build) and prod (pull from registry)
- Separate test and deploy jobs
- Tests run first and block deployment on failure
- Uses MongoDB service container for tests
- SSH-based deployment for security and flexibility
- Health check verification after deployment
- Requires secrets: SSH_HOST, SSH_USER, SSH_PRIVATE_KEY, DEPLOY_PATH
- Update all test files to use response.body.data.* instead of
response.body.* to match the API response convention
- Skip SEC-4.3 rate limiting tests (moved to Phase 5)
- All 202 tests now pass
- Document completed Google OAuth authentication system
- Include implementation details, security measures, and troubleshooting
- Track future enhancements (admin dashboard, Apple Sign-In, partial public access)