Compare commits

29 Commits

Author SHA1 Message Date
fcfd35aa3f Update parseUnits tests to match real fixture data
All checks were successful
CI/CD Pipeline - Apartment API / Scan Dependencies (pull_request) Successful in 13s
CI/CD Pipeline - Apartment API / Lint & Test (pull_request) Successful in 43s
CI/CD Pipeline - Apartment API / Send Webhook Notification (pull_request) Successful in 2s
CI/CD Pipeline - Apartment API / Build & Push Image (pull_request) Has been skipped
CI/CD Pipeline - Apartment API / Deploy to Production (pull_request) Has been skipped
The fixture file was updated with real website data (10 units) but
the tests still referenced old synthetic data (5 units with CCT-*
codes). Updated tests to use real unit codes from the fixture (W2707,
E3205, W2603) and moved edge case tests (image extraction, missing
attributes, unavailable units) to use inline HTML for isolation.
2026-02-07 10:56:37 -07:00
724115f2b3 Add HTML fixture files for scraper unit tests
Create three HTML fixture files in __tests__/scraper/fixtures/ to support
deterministic testing of the scraper's HTML parsing logic:

- sample-listing.html: Contains 10 real unit articles extracted from the
  live listings page, covering studios, 1BR, and 2BR floor plans with
  varied pricing and availability dates.

- sample-listing-empty.html: Minimal page structure with an empty units
  section, for testing graceful handling of pages with no listings.

- sample-listing-call.html: Contains units with "Call for pricing" instead
  of numeric rent values, for testing the parser's handling of non-numeric
  price fields.
2026-02-07 10:53:04 -07:00
6f1651436d Fix webhook notification failure when CI tests produce large output
Truncate lint/test output to 10000 chars at the source (CI job) before
writing to GITHUB_OUTPUT, instead of in the downstream notify job.
Previously the full output was passed as an env var between jobs, which
could exceed Linux's ARG_MAX limit and prevent bash from launching.
2026-02-07 10:51:28 -07:00
e70f7429ea SCRAPE-23: Sanitize error messages in scraper_runs (#27)
Co-authored-by: Stephen Minakian <stephenminakian@gmail.com>
Co-committed-by: Stephen Minakian <stephenminakian@gmail.com>
2026-02-06 23:44:02 -07:00
e654d59fea SCRAPE-22: Implement rate limiting for trigger endpoint (#26)
Co-authored-by: Stephen Minakian <stephenminakian@gmail.com>
Co-committed-by: Stephen Minakian <stephenminakian@gmail.com>
2026-02-06 23:17:56 -07:00
a1ee25ef17 SCRAPE-21: Add requireAuth + requireAdmin middleware verification (#25)
Co-authored-by: Stephen Minakian <stephenminakian@gmail.com>
Co-committed-by: Stephen Minakian <stephenminakian@gmail.com>
2026-02-06 23:08:50 -07:00
f389970d16 SCRAPE-20: Implement credential redaction in scraper log calls (#24)
Co-authored-by: Stephen Minakian <stephenminakian@gmail.com>
Co-committed-by: Stephen Minakian <stephenminakian@gmail.com>
2026-02-06 23:00:08 -07:00
d0538334a5 SCRAPE-18: Implement GET /admin/scraper/history endpoint (#23)
Co-authored-by: Stephen Minakian <stephenminakian@gmail.com>
Co-committed-by: Stephen Minakian <stephenminakian@gmail.com>
2026-02-06 22:34:46 -07:00
9d0b9debbb SCRAPE-17: Implement GET /admin/scraper/status endpoint (#22)
Co-authored-by: Stephen Minakian <stephenminakian@gmail.com>
Co-committed-by: Stephen Minakian <stephenminakian@gmail.com>
2026-02-06 22:20:06 -07:00
daa234428e SCRAPE-16: Implement POST /admin/scraper/run endpoint (#21)
Co-authored-by: Stephen Minakian <stephenminakian@gmail.com>
Co-committed-by: Stephen Minakian <stephenminakian@gmail.com>
2026-02-06 22:07:10 -07:00
c6d480a870 SCRAPE-15: Add getNextScheduledRun() test coverage (#20)
Co-authored-by: Stephen Minakian <stephenminakian@gmail.com>
Co-committed-by: Stephen Minakian <stephenminakian@gmail.com>
2026-02-06 18:21:29 -07:00
0c07baa977 SCRAPE-14: Add graceful shutdown handling (#19)
Co-authored-by: Stephen Minakian <stephenminakian@gmail.com>
Co-committed-by: Stephen Minakian <stephenminakian@gmail.com>
2026-02-06 18:15:24 -07:00
9d789d38fd SCRAPE-13: Create node-cron job initialization (#18)
Co-authored-by: Stephen Minakian <stephenminakian@gmail.com>
Co-committed-by: Stephen Minakian <stephenminakian@gmail.com>
2026-02-06 17:28:53 -07:00
dd0269eb30 SCRAPE-12: Implement in-process mutex (isScraperRunning flag) (#17)
Co-authored-by: Stephen Minakian <stephenminakian@gmail.com>
Co-committed-by: Stephen Minakian <stephenminakian@gmail.com>
2026-02-06 16:27:29 -07:00
d818296eeb Restore --runInBand for test stability in merged CI job
Phase tests share a single MongoMemoryServer database and conflict
when run in parallel. Sequential execution is needed until tests
use isolated databases per file.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-06 10:11:34 -07:00
3509da8185 Speed up CI pipeline: merge lint+test, remove runInBand, drop unused mongo service
- Combine lint and test into a single 'ci' job (eliminates duplicate
  checkout + npm ci, saving ~60-90s)
- Remove --runInBand flag so Jest parallelizes across worker pools
- Remove unused mongo:7 service container (tests use MongoMemoryServer)
- Fix failure detection: check step outcomes instead of job result,
  which was always 'success' due to continue-on-error

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-06 09:54:02 -07:00
0ad4c98abe SCRAPE-11: Create main runScraper() orchestration function (#16)
## Summary

Implements the top-level runScrape() orchestration function that coordinates the entire scraper pipeline end-to-end.

### What it does

- Full pipeline orchestration: Calls fetchPage, parseUnits, convertDataTypes, upsertUnits, insertPrices, markStaleUnits, updateDailySummary in sequence
- dryRun mode: When enabled, parses and validates HTML but skips all database writes
- htmlContent injection: Accepts raw HTML directly, bypassing the fetch step
- New/rented unit calculation: Diffs currently scraped units against previously active units to determine newUnitsCount and rentedUnitsCount for the daily summary
- Run history recording: Every scrape (success or failure) is recorded to the scraper_runs collection via recordScraperRun()
- Structured logging: All pipeline stages log with jobId correlation for traceability
- Error resilience: Catches and handles errors at each stage, ensuring partial failures are logged and recorded

### Test coverage (15 tests)

- Full workflow with mocked dependencies
- Result structure validation and jobId generation
- dryRun mode skips DB writes
- htmlContent bypasses fetch
- Success and failure history recording
- Fetch error handling with retry exhaustion
- Database operation error handling
- New/rented unit count calculation
- Default and scheduled trigger types
- Empty HTML (no units) edge case

Reviewed-on: #16
Co-authored-by: Stephen Minakian <stephenminakian@gmail.com>
Co-committed-by: Stephen Minakian <stephenminakian@gmail.com>
2026-02-06 09:45:32 -07:00
d1f717891a SCRAPE-10: Implement recordScraperRun() for scraper_runs (#15)
Co-authored-by: Stephen Minakian <stephenminakian@gmail.com>
Co-committed-by: Stephen Minakian <stephenminakian@gmail.com>
2026-02-06 02:00:19 -07:00
b4978caf31 SCRAPE-9: Implement updateDailySummary() aggregation (#14)
Co-authored-by: Stephen Minakian <stephenminakian@gmail.com>
Co-committed-by: Stephen Minakian <stephenminakian@gmail.com>
2026-02-06 00:15:30 -07:00
d8adfbf90c SCRAPE-8: Implement markStaleUnits() (#13)
Co-authored-by: Stephen Minakian <stephenminakian@gmail.com>
Co-committed-by: Stephen Minakian <stephenminakian@gmail.com>
2026-02-05 23:48:57 -07:00
2b53288fbe SCRAPE-7: Implement insertPrices() bulk operation (#12)
Co-authored-by: Stephen Minakian <stephenminakian@gmail.com>
Co-committed-by: Stephen Minakian <stephenminakian@gmail.com>
2026-02-05 22:32:02 -07:00
4863dc1824 SCRAPE-6: Implement upsertUnits() bulk operation (#11)
Co-authored-by: Stephen Minakian <stephenminakian@gmail.com>
Co-committed-by: Stephen Minakian <stephenminakian@gmail.com>
2026-02-05 22:10:54 -07:00
0cbeaaf9d5 SCRAPE-5: Implement convertDataTypes() with type parsing helpers (#10)
Co-authored-by: Stephen Minakian <stephenminakian@gmail.com>
Co-committed-by: Stephen Minakian <stephenminakian@gmail.com>
2026-02-05 21:33:58 -07:00
072e80d069 SCRAPE-4: Implement parseUnits() with cheerio selectors (#9)
Co-authored-by: Stephen Minakian <stephenminakian@gmail.com>
Co-committed-by: Stephen Minakian <stephenminakian@gmail.com>
2026-02-05 17:14:01 -07:00
c6beaf8333 SCRAPE-3: Implement fetchPage() with axios, timeout, retry (#8)
Co-authored-by: Stephen Minakian <stephenminakian@gmail.com>
Co-committed-by: Stephen Minakian <stephenminakian@gmail.com>
2026-02-05 15:58:07 -07:00
2993d019c5 SCRAPE-2: Add scraper config constants (#7) 2026-01-31 20:54:43 -07:00
3af5a90c93 Add PR number to webhook payload
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2026-01-28 10:54:12 -07:00
9dba679221 Add scraperLogger.js with credential redaction (#6)
Co-authored-by: Stephen Minakian <stephenminakian@gmail.com>
Co-committed-by: Stephen Minakian <stephenminakian@gmail.com>
2026-01-28 10:52:00 -07:00
6daacf4d12 Add ESLint and n8n webhook notifications to CI/CD
- Add ESLint 9 with flat config for Node.js linting
- Add lint and lint:fix npm scripts
- Add lint job to CI/CD pipeline
- Add notify job to send test/lint results to n8n webhook
- Webhook reports pass/fail status with failure details

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2026-01-28 09:39:55 -07:00
17 changed files with 88 additions and 515 deletions

View File

@ -44,7 +44,6 @@ jobs:
set +e set +e
OUTPUT=$(npm run lint 2>&1) OUTPUT=$(npm run lint 2>&1)
EXIT_CODE=$? EXIT_CODE=$?
echo "$OUTPUT"
# Truncate before writing to GITHUB_OUTPUT to prevent # Truncate before writing to GITHUB_OUTPUT to prevent
# "argument list too long" in downstream jobs # "argument list too long" in downstream jobs
echo "lint_output<<EOF" >> $GITHUB_OUTPUT echo "lint_output<<EOF" >> $GITHUB_OUTPUT
@ -59,7 +58,6 @@ jobs:
set +e set +e
OUTPUT=$(npm test -- --runInBand 2>&1) OUTPUT=$(npm test -- --runInBand 2>&1)
EXIT_CODE=$? EXIT_CODE=$?
echo "$OUTPUT"
# Truncate before writing to GITHUB_OUTPUT to prevent # Truncate before writing to GITHUB_OUTPUT to prevent
# "argument list too long" in downstream jobs # "argument list too long" in downstream jobs
echo "test_output<<EOF" >> $GITHUB_OUTPUT echo "test_output<<EOF" >> $GITHUB_OUTPUT
@ -77,7 +75,7 @@ jobs:
name: Send Webhook Notification name: Send Webhook Notification
runs-on: ubuntu-latest runs-on: ubuntu-latest
needs: [ci] needs: [ci]
if: always() && github.ref != 'refs/heads/main' if: always()
steps: steps:
- name: Send results to n8n webhook - name: Send results to n8n webhook
@ -200,12 +198,8 @@ jobs:
build: build:
name: Build & Push Image name: Build & Push Image
runs-on: ubuntu-latest runs-on: ubuntu-latest
needs: [ci, scan-deps] needs: [ci, scan-deps, notify]
if: >- if: github.ref == 'refs/heads/main' && github.event_name != 'pull_request'
github.ref == 'refs/heads/main' &&
github.event_name != 'pull_request' &&
needs.ci.outputs.lint_status == 'success' &&
needs.ci.outputs.test_status == 'success'
outputs: outputs:
image_tag: ${{ steps.set-tag.outputs.tag }} image_tag: ${{ steps.set-tag.outputs.tag }}
@ -270,13 +264,12 @@ jobs:
# ============================================================ # ============================================================
# SBOM Generation with Syft # SBOM Generation with Syft
# ============================================================ # ============================================================
- name: Install Syft
run: |
curl -sSfL https://raw.githubusercontent.com/anchore/syft/main/install.sh | sh -s -- -b /usr/local/bin
- name: Generate SBOM with Syft - name: Generate SBOM with Syft
run: | uses: anchore/sbom-action@v0
syft ${{ steps.set-tag.outputs.full_image }} -o spdx-json=sbom.spdx.json with:
image: ${{ steps.set-tag.outputs.full_image }}
format: spdx-json
output-file: sbom.spdx.json
# ============================================================ # ============================================================
# Image Signing with Cosign # Image Signing with Cosign
@ -317,21 +310,8 @@ jobs:
if: github.ref == 'refs/heads/main' && github.event_name != 'pull_request' if: github.ref == 'refs/heads/main' && github.event_name != 'pull_request'
steps: steps:
- name: Checkout code
uses: actions/checkout@v4
- name: Sync docker-compose.yml to server
uses: appleboy/scp-action@v0.1.7
with:
host: ${{ secrets.SSH_HOST }}
username: ${{ secrets.SSH_USER }}
key: ${{ secrets.SSH_PRIVATE_KEY }}
port: ${{ secrets.SSH_PORT || 22 }}
source: "docker-compose.yml"
target: ${{ secrets.DEPLOY_PATH }}
- name: Deploy via SSH - name: Deploy via SSH
uses: appleboy/ssh-action@v1.2.4 uses: appleboy/ssh-action@v1.0.3
with: with:
host: ${{ secrets.SSH_HOST }} host: ${{ secrets.SSH_HOST }}
username: ${{ secrets.SSH_USER }} username: ${{ secrets.SSH_USER }}
@ -373,7 +353,7 @@ jobs:
fi fi
- name: Verify deployment - name: Verify deployment
uses: appleboy/ssh-action@v1.2.4 uses: appleboy/ssh-action@v1.0.3
with: with:
host: ${{ secrets.SSH_HOST }} host: ${{ secrets.SSH_HOST }}
username: ${{ secrets.SSH_USER }} username: ${{ secrets.SSH_USER }}

View File

@ -15,15 +15,14 @@
const { MongoClient } = require('mongodb'); const { MongoClient } = require('mongodb');
const { MongoMemoryServer } = require('mongodb-memory-server'); const { MongoMemoryServer } = require('mongodb-memory-server');
// Will require insertPrices and getNow after implementation // Will require insertPrices after implementation
let insertPrices; let insertPrices;
let getNow;
let mongoServer; let mongoServer;
let client; let client;
let db; let db;
const PRICES_COLLECTION = 'unit_prices_scraper'; const PRICES_COLLECTION = 'unit_prices_migration_test';
// Mock logger for capturing log calls // Mock logger for capturing log calls
function createMockLogger() { function createMockLogger() {
@ -42,7 +41,7 @@ beforeAll(async () => {
db = client.db('test_apartments'); db = client.db('test_apartments');
// Dynamically require to pick up implementation // Dynamically require to pick up implementation
({ insertPrices, getNow } = require('../../services/scraperService')); ({ insertPrices } = require('../../services/scraperService'));
}); });
afterAll(async () => { afterAll(async () => {
@ -242,12 +241,12 @@ describe('insertPrices', () => {
it('should include last_updated as an ISO timestamp string', async () => { it('should include last_updated as an ISO timestamp string', async () => {
const logger = createMockLogger(); const logger = createMockLogger();
const beforeTime = getNow(); const beforeTime = new Date().toISOString();
const units = [{ unit_code: 'F401', price: 1600 }]; const units = [{ unit_code: 'F401', price: 1600 }];
await insertPrices(db, units, TODAY, logger); await insertPrices(db, units, TODAY, logger);
const afterTime = getNow(); const afterTime = new Date().toISOString();
const doc = await db.collection(PRICES_COLLECTION).findOne({}); const doc = await db.collection(PRICES_COLLECTION).findOne({});
expect(doc.last_updated).toBeDefined(); expect(doc.last_updated).toBeDefined();

View File

@ -21,7 +21,7 @@ let mongoServer;
let client; let client;
let db; let db;
const UNITS_COLLECTION = 'units_scraper'; const UNITS_COLLECTION = 'units_migration_test';
// Mock logger for capturing log calls // Mock logger for capturing log calls
function createMockLogger() { function createMockLogger() {

View File

@ -384,93 +384,7 @@ describe('parseUnits', () => {
}); });
// --------------------------------------------------------------- // ---------------------------------------------------------------
// 7. Fixture-based tests (sample-listing-empty.html, sample-listing-call.html) // 7. Multiple units with varied data
// ---------------------------------------------------------------
describe('fixture: sample-listing-empty.html', () => {
let emptyHtml;
beforeAll(() => {
emptyHtml = fs.readFileSync(
path.join(fixturesDir, 'sample-listing-empty.html'),
'utf-8'
);
});
it('should return empty array when fixture has container but no articles', () => {
const units = parseUnits(emptyHtml, mockLogger);
expect(units).toEqual([]);
// Container exists so no error about missing container
expect(mockLogger.error).not.toHaveBeenCalled();
});
it('should log zero parsed units for empty fixture', () => {
parseUnits(emptyHtml, mockLogger);
expect(mockLogger.info).toHaveBeenCalledWith(
'Units parsed',
expect.objectContaining({ count: 0 })
);
});
});
describe('fixture: sample-listing-call.html', () => {
let callHtml;
beforeAll(() => {
callHtml = fs.readFileSync(
path.join(fixturesDir, 'sample-listing-call.html'),
'utf-8'
);
});
it('should extract both units from the call-for-pricing fixture', () => {
const units = parseUnits(callHtml, mockLogger);
expect(units).toHaveLength(2);
expect(units.map(u => u.unit_code)).toEqual(['P-101', 'P-102']);
});
it('should preserve "Call for pricing" as raw string in the price field', () => {
const units = parseUnits(callHtml, mockLogger);
expect(units[0].price).toBe('Call for pricing');
expect(units[1].price).toBe('Call for pricing');
});
it('should extract all attributes correctly from call-for-pricing units', () => {
const units = parseUnits(callHtml, mockLogger);
const unit = units.find(u => u.unit_code === 'P-101');
expect(unit.id).toBe('9001');
expect(unit.unit_id).toBe('9001');
expect(unit.floor).toBe('1');
expect(unit.area).toBe('750');
expect(unit.bed_count).toBe('1');
expect(unit.bath_count).toBe('1');
expect(unit.available).toBe('true');
expect(unit.unavailable).toBe('false');
expect(unit.soonest).toBe('Now');
expect(unit.date_available).toBe('1706745600');
expect(unit.plan_id).toBe('500');
expect(unit.plan_name).toBe('Penthouse A');
expect(unit.community).toBe('Country Club Towers');
expect(unit.asset).toBe('420');
expect(unit.href).toBe('?spaces_tab=unit-detail&detail=9001');
expect(unit.inventory_href).toBe('?spaces_tab=unit-detail&detail=9001');
expect(unit.specials_content).toBe('');
});
it('should not have image_url when fixture units lack img elements', () => {
const units = parseUnits(callHtml, mockLogger);
expect(units[0].image_url).toBeUndefined();
expect(units[1].image_url).toBeUndefined();
});
});
// ---------------------------------------------------------------
// 8. Multiple units with varied data
// --------------------------------------------------------------- // ---------------------------------------------------------------
describe('varied unit data', () => { describe('varied unit data', () => {
it('should handle unavailable units correctly', () => { it('should handle unavailable units correctly', () => {

View File

@ -15,7 +15,6 @@ const { MongoMemoryServer } = require('mongodb-memory-server');
// Will require after implementation // Will require after implementation
let recordScraperRun; let recordScraperRun;
let getNow;
let mongoServer; let mongoServer;
let client; let client;
@ -30,7 +29,7 @@ beforeAll(async () => {
db = client.db('test_apartments'); db = client.db('test_apartments');
// Dynamically require to pick up implementation // Dynamically require to pick up implementation
({ recordScraperRun, getNow } = require('../../services/scraperService')); ({ recordScraperRun } = require('../../services/scraperService'));
}); });
afterAll(async () => { afterAll(async () => {
@ -270,26 +269,26 @@ describe('recordScraperRun', () => {
// 6. recordedAt is a valid Date object // 6. recordedAt is a valid Date object
// --------------------------------------------------------------- // ---------------------------------------------------------------
describe('recordedAt field', () => { describe('recordedAt field', () => {
it('should set recordedAt as a string timestamp', async () => { it('should set recordedAt as a Date instance', async () => {
const runData = createSampleRunData(); const runData = createSampleRunData();
await recordScraperRun(db, runData, logger); await recordScraperRun(db, runData, logger);
const doc = await db.collection(SCRAPER_RUNS_COLLECTION).findOne({}); const doc = await db.collection(SCRAPER_RUNS_COLLECTION).findOne({});
expect(typeof doc.recordedAt).toBe('string'); expect(doc.recordedAt).toBeInstanceOf(Date);
}); });
it('should set recordedAt close to current time', async () => { it('should set recordedAt close to current time', async () => {
const beforeTime = getNow(); const beforeTime = new Date();
const runData = createSampleRunData(); const runData = createSampleRunData();
await recordScraperRun(db, runData, logger); await recordScraperRun(db, runData, logger);
const afterTime = getNow(); const afterTime = new Date();
const doc = await db.collection(SCRAPER_RUNS_COLLECTION).findOne({}); const doc = await db.collection(SCRAPER_RUNS_COLLECTION).findOne({});
expect(doc.recordedAt >= beforeTime).toBe(true); expect(doc.recordedAt.getTime()).toBeGreaterThanOrEqual(beforeTime.getTime());
expect(doc.recordedAt <= afterTime).toBe(true); expect(doc.recordedAt.getTime()).toBeLessThanOrEqual(afterTime.getTime());
}); });
it('should not overwrite any existing runData fields with recordedAt', async () => { it('should not overwrite any existing runData fields with recordedAt', async () => {
@ -302,7 +301,7 @@ describe('recordScraperRun', () => {
// Verify the original fields still exist alongside recordedAt // Verify the original fields still exist alongside recordedAt
expect(doc.jobId).toBe(runData.jobId); expect(doc.jobId).toBe(runData.jobId);
expect(doc.status).toBe(runData.status); expect(doc.status).toBe(runData.status);
expect(typeof doc.recordedAt).toBe('string'); expect(doc.recordedAt).toBeInstanceOf(Date);
}); });
}); });
}); });

View File

@ -50,15 +50,15 @@ describe('config/scraper', () => {
}); });
describe('SCRAPER_TIMEZONE', () => { describe('SCRAPER_TIMEZONE', () => {
test('should default to "America/Denver" when env var not set', () => { test('should default to "UTC" when env var not set', () => {
const config = require('../../config/scraper'); const config = require('../../config/scraper');
expect(config.SCRAPER_TIMEZONE).toBe('America/Denver'); expect(config.SCRAPER_TIMEZONE).toBe('UTC');
}); });
test('should use env var override when set', () => { test('should use env var override when set', () => {
process.env.SCRAPER_TIMEZONE = 'Europe/London'; process.env.SCRAPER_TIMEZONE = 'America/Denver';
const config = require('../../config/scraper'); const config = require('../../config/scraper');
expect(config.SCRAPER_TIMEZONE).toBe('Europe/London'); expect(config.SCRAPER_TIMEZONE).toBe('America/Denver');
}); });
}); });
@ -142,14 +142,14 @@ describe('config/scraper', () => {
describe('COLLECTIONS', () => { describe('COLLECTIONS', () => {
describe('default values', () => { describe('default values', () => {
test('UNITS should default to "units_scraper"', () => { test('UNITS should default to "units_migration_test"', () => {
const config = require('../../config/scraper'); const config = require('../../config/scraper');
expect(config.COLLECTIONS.UNITS).toBe('units_scraper'); expect(config.COLLECTIONS.UNITS).toBe('units_migration_test');
}); });
test('PRICES should default to "unit_prices_scraper"', () => { test('PRICES should default to "unit_prices_migration_test"', () => {
const config = require('../../config/scraper'); const config = require('../../config/scraper');
expect(config.COLLECTIONS.PRICES).toBe('unit_prices_scraper'); expect(config.COLLECTIONS.PRICES).toBe('unit_prices_migration_test');
}); });
test('DAILY_SUMMARIES should default to "daily_summaries"', () => { test('DAILY_SUMMARIES should default to "daily_summaries"', () => {

View File

@ -118,8 +118,8 @@ jest.mock('../../config/scraper', () => ({
RETRY_CONFIG: { maxRetries: 3, baseDelay: 1000, timeout: 30000 }, RETRY_CONFIG: { maxRetries: 3, baseDelay: 1000, timeout: 30000 },
SHUTDOWN_TIMEOUT: 30000, SHUTDOWN_TIMEOUT: 30000,
COLLECTIONS: { COLLECTIONS: {
UNITS: 'units_scraper', UNITS: 'units_migration_test',
PRICES: 'unit_prices_scraper', PRICES: 'unit_prices_migration_test',
DAILY_SUMMARIES: 'daily_summaries', DAILY_SUMMARIES: 'daily_summaries',
SCRAPER_RUNS: 'scraper_runs' SCRAPER_RUNS: 'scraper_runs'
} }

View File

@ -20,10 +20,9 @@ let mongoServer;
let client; let client;
let db; let db;
// We will require runScrape, recordScraperRun, and createScraperIndexes after implementation // We will require runScrape and recordScraperRun after implementation
let runScrape; let runScrape;
let recordScraperRun; let recordScraperRun;
let createScraperIndexes;
// Store original module references so we can mock individual functions // Store original module references so we can mock individual functions
let scraperService; let scraperService;
@ -37,7 +36,7 @@ beforeAll(async () => {
// Dynamically require to pick up implementation // Dynamically require to pick up implementation
scraperService = require('../../services/scraperService'); scraperService = require('../../services/scraperService');
({ runScrape, recordScraperRun, createScraperIndexes } = scraperService); ({ runScrape, recordScraperRun } = scraperService);
}); });
afterAll(async () => { afterAll(async () => {
@ -121,7 +120,7 @@ describe('recordScraperRun', () => {
const saved = await db.collection('scraper_runs').findOne({ jobId: 'test-job-001' }); const saved = await db.collection('scraper_runs').findOne({ jobId: 'test-job-001' });
expect(saved).toBeTruthy(); expect(saved).toBeTruthy();
expect(saved.status).toBe('success'); expect(saved.status).toBe('success');
expect(typeof saved.recordedAt).toBe('string'); expect(saved.recordedAt).toBeInstanceOf(Date);
}); });
it('should not throw when insert fails', async () => { it('should not throw when insert fails', async () => {
@ -220,11 +219,11 @@ describe('runScrape', () => {
expect(result.staleUnitsCount).toBe(0); expect(result.staleUnitsCount).toBe(0);
// Verify no units were written to the units collection // Verify no units were written to the units collection
const unitsCount = await db.collection('units_scraper').countDocuments(); const unitsCount = await db.collection('units_migration_test').countDocuments();
expect(unitsCount).toBe(0); expect(unitsCount).toBe(0);
// Verify no prices were written // Verify no prices were written
const pricesCount = await db.collection('unit_prices_scraper').countDocuments(); const pricesCount = await db.collection('unit_prices_migration_test').countDocuments();
expect(pricesCount).toBe(0); expect(pricesCount).toBe(0);
}); });
@ -263,7 +262,7 @@ describe('runScrape', () => {
const runRecord = await db.collection('scraper_runs').findOne({ jobId: 'test-record-success' }); const runRecord = await db.collection('scraper_runs').findOne({ jobId: 'test-record-success' });
expect(runRecord).toBeTruthy(); expect(runRecord).toBeTruthy();
expect(runRecord.status).toBe('success'); expect(runRecord.status).toBe('success');
expect(typeof runRecord.recordedAt).toBe('string'); expect(runRecord.recordedAt).toBeInstanceOf(Date);
expect(runRecord.unitsProcessed).toBe(1); expect(runRecord.unitsProcessed).toBe(1);
}); });
@ -278,7 +277,7 @@ describe('runScrape', () => {
const faultyDb = { const faultyDb = {
collection: (name) => { collection: (name) => {
const realCollection = db.collection(name); const realCollection = db.collection(name);
if (name === 'units_scraper') { if (name === 'units_migration_test') {
return new Proxy(realCollection, { return new Proxy(realCollection, {
get(target, prop) { get(target, prop) {
if (prop === 'bulkWrite') { if (prop === 'bulkWrite') {
@ -344,7 +343,7 @@ describe('runScrape', () => {
const faultyDb = { const faultyDb = {
collection: (name) => { collection: (name) => {
const realCollection = db.collection(name); const realCollection = db.collection(name);
if (name === 'unit_prices_scraper') { if (name === 'unit_prices_migration_test') {
return new Proxy(realCollection, { return new Proxy(realCollection, {
get(target, prop) { get(target, prop) {
if (prop === 'bulkWrite') { if (prop === 'bulkWrite') {
@ -389,7 +388,7 @@ describe('runScrape', () => {
const yesterdayStr = yesterday.toISOString().split('T')[0]; const yesterdayStr = yesterday.toISOString().split('T')[0];
// Yesterday had units: OLD-A, OLD-B, OLD-C // Yesterday had units: OLD-A, OLD-B, OLD-C
await db.collection('unit_prices_scraper').insertMany([ await db.collection('unit_prices_migration_test').insertMany([
{ unit_code: 'OLD-A', date_checked: yesterdayStr, price: 1000 }, { unit_code: 'OLD-A', date_checked: yesterdayStr, price: 1000 },
{ unit_code: 'OLD-B', date_checked: yesterdayStr, price: 1100 }, { unit_code: 'OLD-B', date_checked: yesterdayStr, price: 1100 },
{ unit_code: 'OLD-C', date_checked: yesterdayStr, price: 1200 } { unit_code: 'OLD-C', date_checked: yesterdayStr, price: 1200 }
@ -632,7 +631,7 @@ describe('sanitizeError', () => {
const faultyDb = { const faultyDb = {
collection: (name) => { collection: (name) => {
const realCollection = db.collection(name); const realCollection = db.collection(name);
if (name === 'units_scraper') { if (name === 'units_migration_test') {
return new Proxy(realCollection, { return new Proxy(realCollection, {
get(target, prop) { get(target, prop) {
if (prop === 'bulkWrite') { if (prop === 'bulkWrite') {
@ -675,65 +674,3 @@ describe('sanitizeError', () => {
}); });
}); });
}); });
// ============================================================
// Test: createScraperIndexes
// ============================================================
describe('createScraperIndexes', () => {
const mockLogger = { info: jest.fn(), warn: jest.fn(), error: jest.fn() };
beforeEach(() => {
mockLogger.info.mockClear();
mockLogger.warn.mockClear();
mockLogger.error.mockClear();
});
it('should create the status_startedAt compound index', async () => {
await createScraperIndexes(db, mockLogger);
const indexes = await db.collection('scraper_runs').indexes();
const statusIndex = indexes.find(idx => idx.name === 'status_startedAt');
expect(statusIndex).toBeTruthy();
expect(statusIndex.key).toEqual({ status: 1, startedAt: -1 });
});
it('should create the startedAt_desc index', async () => {
await createScraperIndexes(db, mockLogger);
const indexes = await db.collection('scraper_runs').indexes();
const startedAtIndex = indexes.find(idx => idx.name === 'startedAt_desc');
expect(startedAtIndex).toBeTruthy();
expect(startedAtIndex.key).toEqual({ startedAt: -1 });
});
it('should log success after creating indexes', async () => {
await createScraperIndexes(db, mockLogger);
expect(mockLogger.info).toHaveBeenCalledWith('Scraper indexes created successfully');
});
it('should be idempotent (no error on second call)', async () => {
await createScraperIndexes(db, mockLogger);
// Calling again should not throw
await expect(createScraperIndexes(db, mockLogger)).resolves.not.toThrow();
// Indexes should still exist
const indexes = await db.collection('scraper_runs').indexes();
const statusIndex = indexes.find(idx => idx.name === 'status_startedAt');
const startedAtIndex = indexes.find(idx => idx.name === 'startedAt_desc');
expect(statusIndex).toBeTruthy();
expect(startedAtIndex).toBeTruthy();
});
it('should handle database errors gracefully', async () => {
// Pass an object that will throw when collection() is called
const faultyDb = {
collection: () => {
throw new Error('Connection lost');
}
};
await expect(createScraperIndexes(faultyDb, mockLogger)).rejects.toThrow('Connection lost');
});
});

View File

@ -1,177 +0,0 @@
/**
* Tests for scraper integration in server.js startup
*
* Covers:
* - Server initializes scraper scheduler on startup
* - Server does not crash if scheduler init fails
* - Indexes are created before scheduler starts
* - Scheduler not initialized when SCRAPER_ENABLED=false
*
* Strategy: Instead of importing server.js directly (which starts Express and
* connects to MongoDB), we replicate the scraper initialization logic from
* connectToMongoDB() and test it with mocked dependencies. This verifies the
* integration contract without needing a full server boot.
*/
const { MongoClient } = require('mongodb');
describe('Server Scraper Integration', () => {
let connection;
let db;
beforeAll(async () => {
const uri = process.env.MONGO_URI;
connection = await MongoClient.connect(uri);
db = connection.db('apartments_server_integration_test');
});
afterAll(async () => {
if (connection) {
await connection.close();
}
});
describe('scraper initialization during startup', () => {
let mockCreateScraperIndexes;
let mockInitializeScheduler;
let mockCreateLogger;
let mockLogger;
beforeEach(() => {
mockLogger = {
info: jest.fn(),
warn: jest.fn(),
error: jest.fn(),
debug: jest.fn()
};
mockCreateLogger = jest.fn(() => mockLogger);
mockCreateScraperIndexes = jest.fn().mockResolvedValue(undefined);
mockInitializeScheduler = jest.fn();
});
/**
* Replicates the scraper initialization logic from server.js connectToMongoDB().
* This is the exact pattern used in production:
*
* try {
* const scraperLogger = createLogger('server-init');
* await createScraperIndexes(db, scraperLogger);
* if (scraperConfig.SCRAPER_ENABLED) {
* initializeScheduler(db);
* }
* } catch (error) {
* console.error('Scraper initialization failed (server continues):', error.message);
* }
*/
async function runScraperInit(testDb, config) {
try {
const scraperLogger = mockCreateLogger('server-init');
await mockCreateScraperIndexes(testDb, scraperLogger);
if (config.SCRAPER_ENABLED) {
mockInitializeScheduler(testDb);
}
return { success: true };
} catch (error) {
return { success: false, error: error.message };
}
}
it('should initialize scraper scheduler on startup when enabled', async () => {
const config = { SCRAPER_ENABLED: true };
const result = await runScraperInit(db, config);
expect(result.success).toBe(true);
expect(mockInitializeScheduler).toHaveBeenCalledTimes(1);
expect(mockInitializeScheduler).toHaveBeenCalledWith(db);
});
it('should not crash if createScraperIndexes throws an error', async () => {
mockCreateScraperIndexes.mockRejectedValue(new Error('Index creation failed'));
const config = { SCRAPER_ENABLED: true };
const result = await runScraperInit(db, config);
// The server should continue (error is caught, not thrown)
expect(result.success).toBe(false);
expect(result.error).toBe('Index creation failed');
// initializeScheduler should NOT have been called since the error occurred before it
expect(mockInitializeScheduler).not.toHaveBeenCalled();
});
it('should create indexes before initializing the scheduler', async () => {
const callOrder = [];
mockCreateScraperIndexes.mockImplementation(async () => {
callOrder.push('createScraperIndexes');
});
mockInitializeScheduler.mockImplementation(() => {
callOrder.push('initializeScheduler');
});
const config = { SCRAPER_ENABLED: true };
await runScraperInit(db, config);
expect(callOrder).toEqual(['createScraperIndexes', 'initializeScheduler']);
});
it('should not initialize scheduler when SCRAPER_ENABLED is false', async () => {
const config = { SCRAPER_ENABLED: false };
const result = await runScraperInit(db, config);
expect(result.success).toBe(true);
expect(mockCreateScraperIndexes).toHaveBeenCalledTimes(1);
expect(mockInitializeScheduler).not.toHaveBeenCalled();
});
});
describe('server.js imports and exports verification', () => {
it('should be able to import createScraperIndexes from scraperService', () => {
const { createScraperIndexes } = require('../../services/scraperService');
expect(typeof createScraperIndexes).toBe('function');
});
it('should be able to import initializeScheduler from scraperJob', () => {
const { initializeScheduler } = require('../../jobs/scraperJob');
expect(typeof initializeScheduler).toBe('function');
});
it('should be able to import registerSignalHandlers from scraperJob', () => {
const { registerSignalHandlers } = require('../../jobs/scraperJob');
expect(typeof registerSignalHandlers).toBe('function');
});
it('should be able to import SCRAPER_ENABLED from scraper config', () => {
const scraperConfig = require('../../config/scraper');
expect(typeof scraperConfig.SCRAPER_ENABLED).toBe('boolean');
});
});
describe('createScraperIndexes error handling with real db', () => {
it('should handle index creation gracefully with a valid db', async () => {
const { createScraperIndexes } = require('../../services/scraperService');
const { createLogger } = require('../../services/scraperLogger');
const logger = createLogger('test-server-init');
// Should not throw - creates indexes on the test db
await expect(createScraperIndexes(db, logger)).resolves.not.toThrow();
});
it('should throw when given an invalid db object', async () => {
const { createScraperIndexes } = require('../../services/scraperService');
const mockLogger = { info: jest.fn(), error: jest.fn(), warn: jest.fn(), debug: jest.fn() };
// Passing null should throw (simulates what the try/catch in server.js protects against)
await expect(createScraperIndexes(null, mockLogger)).rejects.toThrow();
});
});
describe('registerSignalHandlers', () => {
it('should register SIGTERM and SIGINT handlers without throwing', () => {
const { registerSignalHandlers } = require('../../jobs/scraperJob');
// Should not throw when called
expect(() => registerSignalHandlers()).not.toThrow();
});
});
});

View File

@ -554,8 +554,8 @@ describe('updateDailySummary', () => {
const doc = await db.collection(DAILY_SUMMARIES_COLLECTION).findOne({ date: '2026-02-05' }); const doc = await db.collection(DAILY_SUMMARIES_COLLECTION).findOne({ date: '2026-02-05' });
expect(doc.timestamp).toBeDefined(); expect(doc.timestamp).toBeDefined();
expect(typeof doc.timestamp).toBe('string'); expect(typeof doc.timestamp).toBe('string');
// Verify it matches the Denver-local timestamp format (YYYY-MM-DDTHH:mm:ss.mmm) // Verify it is a valid ISO timestamp
expect(doc.timestamp).toMatch(/^\d{4}-\d{2}-\d{2}T\d{2}:\d{2}:\d{2}\.\d{3}$/); expect(new Date(doc.timestamp).toISOString()).toBe(doc.timestamp);
}); });
}); });
}); });

View File

@ -15,9 +15,8 @@
const { MongoClient } = require('mongodb'); const { MongoClient } = require('mongodb');
const { MongoMemoryServer } = require('mongodb-memory-server'); const { MongoMemoryServer } = require('mongodb-memory-server');
// We will require upsertUnits and getNow after implementation // We will require upsertUnits after implementation
let upsertUnits; let upsertUnits;
let getNow;
let mongoServer; let mongoServer;
let client; let client;
@ -40,7 +39,7 @@ beforeAll(async () => {
db = client.db('test_apartments'); db = client.db('test_apartments');
// Dynamically require to pick up implementation // Dynamically require to pick up implementation
({ upsertUnits, getNow } = require('../../services/scraperService')); ({ upsertUnits } = require('../../services/scraperService'));
}); });
afterAll(async () => { afterAll(async () => {
@ -57,7 +56,7 @@ beforeEach(async () => {
}); });
describe('upsertUnits', () => { describe('upsertUnits', () => {
const UNITS_COLLECTION = 'units_scraper'; const UNITS_COLLECTION = 'units_migration_test';
// --------------------------------------------------------------- // ---------------------------------------------------------------
// 1. bulkWrite is called with correct updateOne operations // 1. bulkWrite is called with correct updateOne operations
@ -146,12 +145,12 @@ describe('upsertUnits', () => {
describe('$set fields', () => { describe('$set fields', () => {
it('should set last_scraped timestamp on every upsert', async () => { it('should set last_scraped timestamp on every upsert', async () => {
const logger = createMockLogger(); const logger = createMockLogger();
const beforeTime = getNow(); const beforeTime = new Date().toISOString();
const units = [{ unit_code: 'T100', price: 1000 }]; const units = [{ unit_code: 'T100', price: 1000 }];
await upsertUnits(db, units, logger); await upsertUnits(db, units, logger);
const afterTime = getNow(); const afterTime = new Date().toISOString();
const doc = await db.collection(UNITS_COLLECTION).findOne({ unit_code: 'T100' }); const doc = await db.collection(UNITS_COLLECTION).findOne({ unit_code: 'T100' });
expect(doc.last_scraped).toBeDefined(); expect(doc.last_scraped).toBeDefined();
@ -222,7 +221,7 @@ describe('upsertUnits', () => {
describe('$setOnInsert for first_seen', () => { describe('$setOnInsert for first_seen', () => {
it('should set first_seen on initial insert', async () => { it('should set first_seen on initial insert', async () => {
const logger = createMockLogger(); const logger = createMockLogger();
const beforeTime = getNow(); const beforeTime = new Date().toISOString();
await upsertUnits(db, [{ unit_code: 'F100', price: 1000 }], logger); await upsertUnits(db, [{ unit_code: 'F100', price: 1000 }], logger);

View File

@ -13,7 +13,7 @@ module.exports = {
// Scheduling configuration // Scheduling configuration
SCRAPER_SCHEDULE: process.env.SCRAPER_SCHEDULE || '0 6 * * *', SCRAPER_SCHEDULE: process.env.SCRAPER_SCHEDULE || '0 6 * * *',
SCRAPER_TIMEZONE: process.env.SCRAPER_TIMEZONE || 'America/Denver', SCRAPER_TIMEZONE: process.env.SCRAPER_TIMEZONE || 'UTC',
SCRAPER_ENABLED: process.env.SCRAPER_ENABLED !== 'false', SCRAPER_ENABLED: process.env.SCRAPER_ENABLED !== 'false',
// HTTP settings // HTTP settings
@ -30,11 +30,10 @@ module.exports = {
// Graceful shutdown timeout (how long to wait for running job before force-stopping) // Graceful shutdown timeout (how long to wait for running job before force-stopping)
SHUTDOWN_TIMEOUT: parseInt(process.env.SCRAPER_SHUTDOWN_TIMEOUT, 10) || 30000, SHUTDOWN_TIMEOUT: parseInt(process.env.SCRAPER_SHUTDOWN_TIMEOUT, 10) || 30000,
// MongoDB collection names - defaults to validation collections for scraper validation; // MongoDB collection names (environment variable overrides for development isolation)
// switch to production collections via env vars (SCRAPER_UNITS_COLLECTION, SCRAPER_PRICES_COLLECTION) when ready
COLLECTIONS: { COLLECTIONS: {
UNITS: process.env.SCRAPER_UNITS_COLLECTION || 'units_scraper', UNITS: process.env.SCRAPER_UNITS_COLLECTION || 'units_migration_test',
PRICES: process.env.SCRAPER_PRICES_COLLECTION || 'unit_prices_scraper', PRICES: process.env.SCRAPER_PRICES_COLLECTION || 'unit_prices_migration_test',
DAILY_SUMMARIES: process.env.SCRAPER_SUMMARIES_COLLECTION || 'daily_summaries', DAILY_SUMMARIES: process.env.SCRAPER_SUMMARIES_COLLECTION || 'daily_summaries',
SCRAPER_RUNS: process.env.SCRAPER_RUNS_COLLECTION || 'scraper_runs' SCRAPER_RUNS: process.env.SCRAPER_RUNS_COLLECTION || 'scraper_runs'
} }

View File

@ -1,5 +1,3 @@
name: apartment-api
services: services:
apartment-api: apartment-api:
image: ${IMAGE:-apartment-api:latest} image: ${IMAGE:-apartment-api:latest}
@ -22,7 +20,9 @@ services:
- "traefik.enable=true" - "traefik.enable=true"
- "traefik.docker.network=traefik" - "traefik.docker.network=traefik"
- "traefik.http.routers.apartment-api.rule=Host(`apartments.maverickapplications.com`) && PathPrefix(`/api`)" - "traefik.http.routers.apartment-api.rule=Host(`apartments.maverickapplications.com`) && PathPrefix(`/api`)"
- "traefik.http.routers.apartment-api.entrypoints=web" - "traefik.http.routers.apartment-api.entrypoints=websecure"
- "traefik.http.routers.apartment-api.tls=true"
- "traefik.http.routers.apartment-api.tls.certresolver=letsencrypt"
- "traefik.http.routers.apartment-api.priority=100" - "traefik.http.routers.apartment-api.priority=100"
- "traefik.http.services.apartment-api.loadbalancer.server.port=8080" # Updated to match - "traefik.http.services.apartment-api.loadbalancer.server.port=8080" # Updated to match
- "traefik.http.middlewares.apartment-api-stripprefix.stripprefix.prefixes=/api" - "traefik.http.middlewares.apartment-api-stripprefix.stripprefix.prefixes=/api"

16
package-lock.json generated
View File

@ -1,15 +1,15 @@
{ {
"name": "apartment-api", "name": "apartment-api",
"version": "1.0.1", "version": "1.0.0",
"lockfileVersion": 3, "lockfileVersion": 3,
"requires": true, "requires": true,
"packages": { "packages": {
"": { "": {
"name": "apartment-api", "name": "apartment-api",
"version": "1.0.1", "version": "1.0.0",
"license": "MIT", "license": "MIT",
"dependencies": { "dependencies": {
"axios": "^1.13.5", "axios": "^1.13.4",
"cheerio": "^1.2.0", "cheerio": "^1.2.0",
"cookie-parser": "^1.4.7", "cookie-parser": "^1.4.7",
"cors": "^2.8.5", "cors": "^2.8.5",
@ -2032,13 +2032,13 @@
"license": "MIT" "license": "MIT"
}, },
"node_modules/axios": { "node_modules/axios": {
"version": "1.13.5", "version": "1.13.4",
"resolved": "https://registry.npmjs.org/axios/-/axios-1.13.5.tgz", "resolved": "https://registry.npmjs.org/axios/-/axios-1.13.4.tgz",
"integrity": "sha512-cz4ur7Vb0xS4/KUN0tPWe44eqxrIu31me+fbang3ijiNscE129POzipJJA6zniq2C/Z6sJCjMimjS8Lc/GAs8Q==", "integrity": "sha512-1wVkUaAO6WyaYtCkcYCOx12ZgpGf9Zif+qXa4n+oYzK558YryKqiL6UWwd5DqiH3VRW0GYhTZQ/vlgJrCoNQlg==",
"license": "MIT", "license": "MIT",
"dependencies": { "dependencies": {
"follow-redirects": "^1.15.11", "follow-redirects": "^1.15.6",
"form-data": "^4.0.5", "form-data": "^4.0.4",
"proxy-from-env": "^1.1.0" "proxy-from-env": "^1.1.0"
} }
}, },

View File

@ -1,6 +1,6 @@
{ {
"name": "apartment-api", "name": "apartment-api",
"version": "1.0.1", "version": "1.0.0",
"description": "API backend for Country Club Towers & Gardens apartment dashboard", "description": "API backend for Country Club Towers & Gardens apartment dashboard",
"main": "server.js", "main": "server.js",
"scripts": { "scripts": {
@ -21,7 +21,7 @@
"author": "Stephen", "author": "Stephen",
"license": "MIT", "license": "MIT",
"dependencies": { "dependencies": {
"axios": "^1.13.5", "axios": "^1.13.4",
"cheerio": "^1.2.0", "cheerio": "^1.2.0",
"cookie-parser": "^1.4.7", "cookie-parser": "^1.4.7",
"cors": "^2.8.5", "cors": "^2.8.5",

View File

@ -11,11 +11,6 @@ const activityRoutes = require('./routes/activity');
const adminRoutes = require('./routes/admin'); const adminRoutes = require('./routes/admin');
const { createIndexes } = require('./models/user'); const { createIndexes } = require('./models/user');
const { createActivityIndexes } = require('./services/activityLogger'); const { createActivityIndexes } = require('./services/activityLogger');
const { createScraperIndexes } = require('./services/scraperService');
const { initializeScheduler, registerSignalHandlers } = require('./jobs/scraperJob');
const scraperConfig = require('./config/scraper');
const { version: API_VERSION } = require('./package.json');
const app = express(); const app = express();
const PORT = process.env.PORT || 3000; const PORT = process.env.PORT || 3000;
@ -85,32 +80,14 @@ async function connectToMongoDB() {
await createIndexes(db); await createIndexes(db);
await createActivityIndexes(db); await createActivityIndexes(db);
console.log('✅ Passport configured and indexes created'); console.log('✅ Passport configured and indexes created');
// Initialize scraper (indexes + scheduler) - errors logged but don't crash server
try {
const { createLogger } = require('./services/scraperLogger');
const scraperLogger = createLogger('server-init');
await createScraperIndexes(db, scraperLogger);
if (scraperConfig.SCRAPER_ENABLED) {
initializeScheduler(db);
console.log('✅ Scraper scheduler initialized');
} else {
console.log('ℹ️ Scraper scheduling disabled');
}
} catch (error) {
console.error('⚠️ Scraper initialization failed (server continues):', error.message);
}
} catch (error) { } catch (error) {
console.error('❌ Failed to connect to MongoDB:', error); console.error('❌ Failed to connect to MongoDB:', error);
process.exit(1); process.exit(1);
} }
} }
// Register scraper signal handlers for graceful shutdown // Helper function to get today's date in UTC
registerSignalHandlers(); const getTodayDateUTC = () => new Date().toISOString().split('T')[0];
// Helper function to get today's date in configured timezone
const getTodayDate = () => new Date().toLocaleDateString('en-CA', { timeZone: scraperConfig.SCRAPER_TIMEZONE });
// Cache for the most recent date with data // Cache for the most recent date with data
let cachedLatestDate = null; let cachedLatestDate = null;
@ -145,7 +122,7 @@ const getLatestDateWithData = async () => {
} }
// Fallback to UTC today if no data found // Fallback to UTC today if no data found
return getTodayDate(); return getTodayDateUTC();
}; };
// Helper function to get yesterday relative to the latest date with data // Helper function to get yesterday relative to the latest date with data
@ -177,7 +154,6 @@ app.use('/admin', adminRoutes);
app.get('/health', (req, res) => { app.get('/health', (req, res) => {
res.json({ res.json({
status: 'ok', status: 'ok',
version: API_VERSION,
timestamp: new Date().toISOString(), timestamp: new Date().toISOString(),
service: 'apartment-api' service: 'apartment-api'
}); });
@ -644,10 +620,10 @@ app.get('/analytics', requireAuth, async (req, res) => {
]).toArray(); ]).toArray();
// 4. Historical Comparison - current vs 30/90/365 day averages // 4. Historical Comparison - current vs 30/90/365 day averages
const tz = scraperConfig.SCRAPER_TIMEZONE; const now = new Date();
const date30 = new Date(Date.now() - 30 * 24 * 60 * 60 * 1000).toLocaleDateString('en-CA', { timeZone: tz }); const date30 = new Date(now.getTime() - 30 * 24 * 60 * 60 * 1000).toISOString().split('T')[0];
const date90 = new Date(Date.now() - 90 * 24 * 60 * 60 * 1000).toLocaleDateString('en-CA', { timeZone: tz }); const date90 = new Date(now.getTime() - 90 * 24 * 60 * 60 * 1000).toISOString().split('T')[0];
const date365 = new Date(Date.now() - 365 * 24 * 60 * 60 * 1000).toLocaleDateString('en-CA', { timeZone: tz }); const date365 = new Date(now.getTime() - 365 * 24 * 60 * 60 * 1000).toISOString().split('T')[0];
const [currentAvg, avg30, avg90, avg365] = await Promise.all([ const [currentAvg, avg30, avg90, avg365] = await Promise.all([
db.collection(PRICES_COLLECTION).aggregate([ db.collection(PRICES_COLLECTION).aggregate([

View File

@ -403,7 +403,7 @@ function getYesterday(dateStr) {
*/ */
async function upsertUnits(db, units, logger) { async function upsertUnits(db, units, logger) {
const collection = db.collection(config.COLLECTIONS.UNITS); const collection = db.collection(config.COLLECTIONS.UNITS);
const now = getNow(); const now = new Date().toISOString();
const operations = units.map(unit => ({ const operations = units.map(unit => ({
updateOne: { updateOne: {
@ -461,7 +461,7 @@ async function upsertUnits(db, units, logger) {
*/ */
async function insertPrices(db, units, date, logger) { async function insertPrices(db, units, date, logger) {
const collection = db.collection(config.COLLECTIONS.PRICES); const collection = db.collection(config.COLLECTIONS.PRICES);
const now = getNow(); const now = new Date().toISOString();
// Filter out units without prices // Filter out units without prices
const unitsWithPrices = units.filter(u => u.price !== null); const unitsWithPrices = units.filter(u => u.price !== null);
@ -581,7 +581,7 @@ async function updateDailySummary(db, summaryData, logger) {
const summary = { const summary = {
date, date,
timestamp: getNow(), timestamp: new Date().toISOString(),
new_units: newUnits, new_units: newUnits,
rented_units: rentedUnits, rented_units: rentedUnits,
stale_units: [], // Stale units list is not tracked per PRD stale_units: [], // Stale units list is not tracked per PRD
@ -626,31 +626,11 @@ async function updateDailySummary(db, summaryData, logger) {
// ============================================================ // ============================================================
/** /**
* Get today's date in YYYY-MM-DD format using the configured timezone. * Get today's date in YYYY-MM-DD format (UTC).
* @returns {string} Today's date string * @returns {string} Today's date string
*/ */
function getToday() { function getTodayUTC() {
return new Date().toLocaleDateString('en-CA', { timeZone: config.SCRAPER_TIMEZONE }); return new Date().toISOString().split('T')[0];
}
/**
* Get current timestamp in the configured timezone as an ISO-like string.
* Returns format: YYYY-MM-DDTHH:mm:ss.mmm (no Z suffix, local time).
* @returns {string} Current timestamp in configured timezone
*/
function getNow() {
const now = new Date();
const formatter = new Intl.DateTimeFormat('en-CA', {
timeZone: config.SCRAPER_TIMEZONE,
year: 'numeric', month: '2-digit', day: '2-digit',
hour: '2-digit', minute: '2-digit', second: '2-digit',
hour12: false
});
const parts = Object.fromEntries(
formatter.formatToParts(now).map(({ type, value }) => [type, value])
);
const ms = String(now.getMilliseconds()).padStart(3, '0');
return `${parts.year}-${parts.month}-${parts.day}T${parts.hour}:${parts.minute}:${parts.second}.${ms}`;
} }
/** /**
@ -738,7 +718,7 @@ async function recordScraperRun(db, runData, logger) {
const collection = db.collection(config.COLLECTIONS.SCRAPER_RUNS); const collection = db.collection(config.COLLECTIONS.SCRAPER_RUNS);
const result = await collection.insertOne({ const result = await collection.insertOne({
...runData, ...runData,
recordedAt: getNow() recordedAt: new Date()
}); });
return result; return result;
@ -750,37 +730,6 @@ async function recordScraperRun(db, runData, logger) {
} }
} }
// ============================================================
// Index Management
// ============================================================
/**
* Create indexes on the scraper_runs collection.
* Should be called once during application startup.
* Idempotent - safe to call multiple times.
*
* @param {Db} db - MongoDB database instance
* @param {Object} logger - Logger instance
* @returns {Promise<void>}
*/
async function createScraperIndexes(db, logger) {
const collection = db.collection(config.COLLECTIONS.SCRAPER_RUNS);
// Compound index for querying runs by status sorted by most recent
await collection.createIndex(
{ status: 1, startedAt: -1 },
{ name: 'status_startedAt' }
);
// Index for sorting all runs by start time (most recent first)
await collection.createIndex(
{ startedAt: -1 },
{ name: 'startedAt_desc' }
);
logger.info('Scraper indexes created successfully');
}
// ============================================================ // ============================================================
// Main Orchestration Function // Main Orchestration Function
// ============================================================ // ============================================================
@ -810,7 +759,7 @@ async function runScrape(db, options = {}) {
trigger, trigger,
dryRun, dryRun,
status: 'running', status: 'running',
startedAt: getNow(), startedAt: new Date().toISOString(),
completedAt: null, completedAt: null,
duration: null, duration: null,
unitsProcessed: 0, unitsProcessed: 0,
@ -840,7 +789,7 @@ async function runScrape(db, options = {}) {
result.unitsProcessed = units.length; result.unitsProcessed = units.length;
// Step 4: Database operations // Step 4: Database operations
const today = getToday(); const today = getTodayUTC();
// Get yesterday's unit codes for comparison // Get yesterday's unit codes for comparison
const yesterdayUnits = await getYesterdayUnitCodes(db, today); const yesterdayUnits = await getYesterdayUnitCodes(db, today);
@ -895,7 +844,7 @@ async function runScrape(db, options = {}) {
result.errors.push(cleanError.message); result.errors.push(cleanError.message);
} finally { } finally {
result.completedAt = getNow(); result.completedAt = new Date().toISOString();
result.duration = Date.now() - startTime; result.duration = Date.now() - startTime;
// Record run to history (always runs, even on failure) // Record run to history (always runs, even on failure)
@ -922,10 +871,8 @@ module.exports = {
markStaleUnits, markStaleUnits,
updateDailySummary, updateDailySummary,
recordScraperRun, recordScraperRun,
createScraperIndexes,
// Export helpers for testing // Export helpers for testing
getToday, getTodayUTC,
getNow,
getYesterdayUnitCodes, getYesterdayUnitCodes,
getYesterday, getYesterday,
parseInteger, parseInteger,