Add fetchPage() with axios, timeout, and retry logic
Some checks failed
CI/CD Pipeline - Apartment API / Run Linting (pull_request) Successful in 9m39s
CI/CD Pipeline - Apartment API / Run Tests (pull_request) Successful in 9m50s
CI/CD Pipeline - Apartment API / Scan Dependencies (pull_request) Successful in 15s
CI/CD Pipeline - Apartment API / Send Webhook Notification (pull_request) Failing after 2s
CI/CD Pipeline - Apartment API / Build & Push Image (pull_request) Has been skipped
CI/CD Pipeline - Apartment API / Deploy to Production (pull_request) Has been skipped

Implement fetchPage() function for HTTP scraping:
- Uses axios for HTTP GET requests
- Configurable timeout (default 30s) and User-Agent header
- Retry with exponential backoff (1s, 2s, 4s) on 5xx and network errors
- Does not retry on 4xx client errors
- Logs each attempt with attempt number and error details
- Returns HTML string on success, throws after retries exhausted
This commit is contained in:
2026-01-30 11:22:32 -07:00
parent 2993d019c5
commit d09aa1d179
3 changed files with 651 additions and 2 deletions

View File

@ -17,14 +17,14 @@ module.exports = {
SCRAPER_ENABLED: process.env.SCRAPER_ENABLED !== 'false',
// HTTP settings
SCRAPER_TIMEOUT: parseInt(process.env.SCRAPER_TIMEOUT) || 30000,
SCRAPER_TIMEOUT: parseInt(process.env.SCRAPER_TIMEOUT, 10) || 30000,
USER_AGENT: 'Mozilla/5.0 (compatible; ApartmentScraper/1.0)',
// Retry configuration for HTTP requests
RETRY_CONFIG: {
maxRetries: 3,
baseDelay: 1000, // 1 second, exponential backoff: 1s, 2s, 4s
timeout: parseInt(process.env.SCRAPER_TIMEOUT) || 30000
timeout: parseInt(process.env.SCRAPER_TIMEOUT, 10) || 30000
},
// MongoDB collection names (environment variable overrides for development isolation)