## Summary
Implements the top-level runScrape() orchestration function that coordinates the entire scraper pipeline end-to-end.
### What it does
- Full pipeline orchestration: Calls fetchPage, parseUnits, convertDataTypes, upsertUnits, insertPrices, markStaleUnits, updateDailySummary in sequence
- dryRun mode: When enabled, parses and validates HTML but skips all database writes
- htmlContent injection: Accepts raw HTML directly, bypassing the fetch step
- New/rented unit calculation: Diffs currently scraped units against previously active units to determine newUnitsCount and rentedUnitsCount for the daily summary
- Run history recording: Every scrape (success or failure) is recorded to the scraper_runs collection via recordScraperRun()
- Structured logging: All pipeline stages log with jobId correlation for traceability
- Error resilience: Catches and handles errors at each stage, ensuring partial failures are logged and recorded
### Test coverage (15 tests)
- Full workflow with mocked dependencies
- Result structure validation and jobId generation
- dryRun mode skips DB writes
- htmlContent bypasses fetch
- Success and failure history recording
- Fetch error handling with retry exhaustion
- Database operation error handling
- New/rented unit count calculation
- Default and scheduled trigger types
- Empty HTML (no units) edge case
Reviewed-on: #16
Co-authored-by: Stephen Minakian <stephenminakian@gmail.com>
Co-committed-by: Stephen Minakian <stephenminakian@gmail.com>
- Add ESLint 9 with flat config for Node.js linting
- Add lint and lint:fix npm scripts
- Add lint job to CI/CD pipeline
- Add notify job to send test/lint results to n8n webhook
- Webhook reports pass/fail status with failure details
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
- Update to node:20-alpine base image
- Add apk upgrade to fix OS-level vulnerabilities
- Update npm to latest to fix bundled package vulnerabilities
- Add scan-deps job that runs in parallel with tests
- Replace deprecated --only=production with --omit=dev
Use pre-configured Docker credentials on server instead of passing
Harbor credentials through SSH script, avoiding shell interpolation
issues with special characters in robot account username.
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
Gitea Actions doesn't support type=gha caching, causing timeout errors.
Removing cache config to get builds working. Can add registry-based
caching later if needed.
- Use SHA-based image tags instead of 'latest' for deployments
- Pass exact image tag from build job to deploy job
- Add default for SSH_PORT
- Health check now fails deploy if not passing
- Added script_stop for proper error handling
- Add image variable substitution to docker-compose.yml
- Build image in CI, push to Harbor registry
- Deploy pulls from Harbor instead of rebuilding
- Supports both local dev (build) and prod (pull from registry)
- Separate test and deploy jobs
- Tests run first and block deployment on failure
- Uses MongoDB service container for tests
- SSH-based deployment for security and flexibility
- Health check verification after deployment
- Requires secrets: SSH_HOST, SSH_USER, SSH_PRIVATE_KEY, DEPLOY_PATH