Files
tessera-ctl/.planning/phases/10-ausschreibungs-radar-foundation-d-e-ingestion/10-04-PLAN.md
T

174 lines
14 KiB
Markdown

---
phase: 10-ausschreibungs-radar-foundation-d-e-ingestion
plan: 04
type: execute
wave: 4
depends_on: ["10-01", "10-02", "10-03"]
files_modified:
- apps/api/src/tenders/tender-ingestion.service.ts
- apps/api/src/tenders/tender-ingestion.service.spec.ts
- apps/api/src/tenders/tender-scheduler.service.ts
- apps/api/src/tenders/tender-scheduler.service.spec.ts
- apps/api/src/tenders/tenders.module.ts
autonomous: true
requirements: [SCHEMA-02, INGEST-06]
must_haves:
truths:
- "A changed DÖE notice (same ocid/dedupKey, new contentHash) updates the existing Tender row instead of creating a duplicate (SCHEMA-02)"
- "DÖE is polled once on a shared global schedule regardless of tenant count — poll-once-fan-out-many, never per-tenant (INGEST-06)"
- "Activating the module for a 2nd tenant triggers zero additional DÖE HTTP calls, zero additional cron jobs, and zero additional Tender rows (Success Criteria 4 & 5)"
- "The poll tick is day-cursor gated: no upstream fetch when dayCursor >= today Europe/Berlin (D-01 from-now, no historical backfill)"
- "Tenders past their deadline are marked expired and pruned after 90 days; deadline-less rows are never auto-expired (D-05)"
artifacts:
- apps/api/src/tenders/tender-ingestion.service.ts
- apps/api/src/tenders/tender-scheduler.service.ts
key_links:
- "Scheduler cron tick → TenderIngestionService.pollDueSources(); day-cursor gate lives in the ingestion service, not the scheduler"
- "prisma.tender.upsert({ where: { dedupKey } }) is the SCHEMA-02 change-detection seam; plain PrismaService (no forTenant) on the global table"
---
<objective>
Wire the ingestion orchestration and the shared scheduler — the poll-once-fan-out-many heart of the phase (INGEST-06) and the content-hash change detection (SCHEMA-02) — and prove the two-tenant safety property that is the phase's headline acceptance criterion.
## Phase Goal (user story)
**As a** Tessera-Administrator, **I want to** dass DÖE genau einmal plattformweit auf einem admin-konfigurierbaren Intervall abgefragt wird — egal wie viele Mandanten das Modul aktiviert haben, **so that** kein redundanter Poll und keine doppelte Ingestion pro Mandant entsteht (INGEST-06, Success Criteria 4 & 5).
Purpose: This plan turns "parse a fixture" (Plan 03) into "the platform continuously ingests real DÖE data safely." The multi-tenant scheduler is where the DKV `findFirst()` single-tenant anti-pattern must NOT be copied; the DÖE config is a genuine platform-wide singleton (RESEARCH Pitfall D).
Output: `TenderIngestionService`, `TenderSchedulerService`, and the two-tenant integration test.
</objective>
<execution_context>
@$HOME/.claude/gsd-core/workflows/execute-plan.md
@$HOME/.claude/gsd-core/templates/summary.md
</execution_context>
<context>
@.planning/PROJECT.md
@.planning/ROADMAP.md
@.planning/phases/10-ausschreibungs-radar-foundation-d-e-ingestion/10-RESEARCH.md
@.planning/phases/10-ausschreibungs-radar-foundation-d-e-ingestion/10-PATTERNS.md
@apps/api/src/dkv/dkv-scheduler.service.ts
@apps/api/src/prisma/prisma-tenant.extension.ts
</context>
<tasks>
<task type="auto" tdd="true">
<name>Task 1: TenderIngestionService — day-cursor gate + upsert + change-detect + retention</name>
<read_first>
- .planning/phases/10-ausschreibungs-radar-foundation-d-e-ingestion/10-PATTERNS.md (tender-ingestion.service section: pollDueSources shape, findUnique on sourceType, NO forTenant)
- .planning/phases/10-ausschreibungs-radar-foundation-d-e-ingestion/10-RESEARCH.md (Pattern 1 day-cursor gate + nextDayToFetch; Pitfall A check-vs-fetch; Pattern 4 D-05 null-deadline handling)
- apps/api/src/dkv/dkv.service.ts (orchestration shape: fetch → normalize → persist → log; catch-and-log)
- apps/api/src/prisma/prisma-tenant.extension.ts (forTenant — the extension to AVOID here)
</read_first>
<behavior>
- Test: when dayCursor >= today (Europe/Berlin), pollDueSources() makes NO adapter call and returns (no-op tick, expected — Pitfall A).
- Test: a fresh notice is inserted; re-ingesting the identical notice (same dedupKey, same contentHash) does NOT create a duplicate and does NOT change field data.
- Test: re-ingesting a changed notice (same dedupKey, new contentHash — e.g. extended deadline) UPDATES the existing row (SCHEMA-02).
- Test: after a successful fetch, lastIngestedDay advances by one day; catch-up loops from lastIngestedDay+1 up to today-1.
- Test: pruneExpiredTenders() marks status='expired' for rows with deadlineAt < now, deletes rows expired+deadlineAt older than 90 days, and NEVER touches rows with deadlineAt = null (D-05).
</behavior>
<files>apps/api/src/tenders/tender-ingestion.service.ts, apps/api/src/tenders/tender-ingestion.service.spec.ts, apps/api/src/tenders/tenders.module.ts</files>
<action>
Write RED spec first, then implement `TenderIngestionService` (Injectable) using the plain global `PrismaService` — do NOT call `forTenant()` on `Tender`/`TenderSourcePollConfig` (global RLS-exempt tables, D-03; wrapping them in the tenant RLS extension would make a 2nd tenant's session silently filter out platform data).
`pollDueSources()`: load the singleton config via `prisma.tenderSourcePollConfig.findUnique({ where: { sourceType: 'doe-opendata' } })` (fixed-slug lookup makes the singleton intent explicit — NOT a per-tenant `findFirst()`). If not `isActive`, return. Compute `nextDay = nextDayToFetch(config.lastIngestedDay)` (RESEARCH day-cursor snippet: first eligible day is activation day per D-01 "from now"; returns null if nextDay is not strictly before Berlin-today). If null, log at debug level ("no new DÖE day yet") and return — this no-op is expected, NOT an error (Pitfall A). Otherwise loop the day-cursor from nextDay forward while still `< today`: for each day call `doeAdapter.fetchTenders(day)`, `normalizer.normalize()` each record, and `prisma.tender.upsert({ where: { dedupKey }, update: {...fields, contentHash}, create: {...} })`. Track newly-created / contentHash-changed ids (the delta — consumed by Phase 11 matching, not this phase). After each day, `update` config.lastIngestedDay. Between successive day fetches insert a ~1-2s polite delay (RESEARCH Open Question 1). Never throw out of the tick — catch-and-log per RESEARCH/DKV pattern.
`pruneExpiredTenders()` (D-05 retention): mark `status='expired'` where `deadlineAt < now` AND `status='active'`; delete where `status='expired'` AND `deadlineAt < now-90d`. Rows with `deadlineAt IS NULL` are excluded from both operations (Pattern 4 recommendation (a) — a wrongly-deleted no-deadline tender is unrecoverable). Call `pruneExpiredTenders()` once per successful tick.
</action>
<verify>
<automated>cd apps/api && pnpm test -- tender-ingestion 2>&1 | tail -20</automated>
</verify>
<acceptance_criteria>
- Day-cursor no-op test green (no adapter call when nothing new).
- SCHEMA-02: changed notice updates in place, identical notice does not duplicate — both green.
- D-05: expired-marking + 90-day prune skips null-deadline rows — green.
- Service uses plain PrismaService (assert no `forTenant` call in the file).
</acceptance_criteria>
<done>Ingestion orchestration with change detection + retention; spec green.</done>
</task>
<task type="auto">
<name>Task 2: TenderSchedulerService — single global cron (poll-once-fan-out-many)</name>
<read_first>
- apps/api/src/dkv/dkv-scheduler.service.ts (REUSE the CronJob require()-resolution + SchedulerRegistry addCronJob/deleteCronJob mechanics; do NOT copy activeTenantId / per-tenant framing)
- .planning/phases/10-ausschreibungs-radar-foundation-d-e-ingestion/10-PATTERNS.md (tender-scheduler.service section + ANTI-PATTERN note)
- .planning/phases/10-ausschreibungs-radar-foundation-d-e-ingestion/10-RESEARCH.md (Pitfall D: the one legitimate singleton findUnique; two-tenant verification framing)
</read_first>
<files>apps/api/src/tenders/tender-scheduler.service.ts, apps/api/src/tenders/tenders.module.ts</files>
<action>
Implement `TenderSchedulerService` (Injectable, OnModuleInit) reusing DKV's cron mechanics: the `require('cron').CronJob` resolution workaround, and a `setInterval(intervalMin)` that removes any existing job then `schedulerRegistry.addCronJob(JOB_NAME, job)` with `JOB_NAME = 'tender-doe-poll'`. Cron expression computed exactly as DKV (`*/${min} * * * *` under 60, else `0 */${hours} * * *`). The tick calls `tenderIngestionService.pollDueSources().catch(logErr)`.
CRITICAL deviations from DKV (poll-once-fan-out-many): the scheduler has NO `activeTenantId` field and `setInterval` takes NO tenant argument — there is exactly ONE global cron job for the whole platform (D-04 default hourly interval). `onModuleInit()` loads the singleton config via `findUnique({ where: { sourceType: 'doe-opendata' } })` (NOT `findFirst()`); if `isActive`, call `setInterval(config.pollIntervalMin)`, else register nothing. Provide `stopJob()` (stop+delete the single job). The day-cursor gate is NOT in the scheduler — the scheduler only sets cron-tick frequency; whether an HTTP call happens is decided inside `pollDueSources()` (Pitfall A separation). Register `TenderSchedulerService` in `TendersModule.providers`.
</action>
<verify>
<automated>cd apps/api && pnpm exec tsc --noEmit -p tsconfig.json 2>&1 | tail -5 && grep -c "activeTenantId" src/tenders/tender-scheduler.service.ts</automated>
</verify>
<acceptance_criteria>
- Single named cron job `tender-doe-poll`; `setInterval` has no tenant parameter.
- No `activeTenantId` field (`grep -c "activeTenantId"` returns 0 — proves the per-tenant anti-pattern was not copied).
- `onModuleInit` uses `findUnique` on `sourceType`, not `findFirst`.
- `tsc --noEmit` passes.
</acceptance_criteria>
<done>One global cron drives the shared DÖE poll; no per-tenant scheduling dimension.</done>
</task>
<task type="auto" tdd="true">
<name>Task 3: Two-tenant safety integration test (Success Criteria 4 & 5)</name>
<read_first>
- .planning/phases/10-ausschreibungs-radar-foundation-d-e-ingestion/10-RESEARCH.md (Pitfall D "Verification for this phase's two-tenant acceptance criterion" — assert ABSENCE of tenant-scaled behavior)
- apps/api/src/tenders/tender-scheduler.service.spec.ts (create)
- apps/api/src/module-registry/module-registry.service.ts (isModuleActive / activation mechanism the test drives)
</read_first>
<files>apps/api/src/tenders/tender-scheduler.service.spec.ts</files>
<action>
Write an integration test proving the poll-once-fan-out-many property. Spy on the DÖE adapter's `fetchTenders` (or on the underlying `fetch`) and on `schedulerRegistry.addCronJob`. Simulate activating the `tender-radar` module for tenant A, run/inspect scheduler init, capture counts. Then activate the module for a SECOND tenant B and assert: zero ADDITIONAL DÖE adapter/HTTP calls, zero additional cron jobs registered (still exactly one `tender-doe-poll`), and zero additional `Tender` rows created purely as a result of the 2nd activation. The assertion is the ABSENCE of tenant-count-scaled behavior (there are no per-tenant DÖE configs to iterate) — not correct per-tenant iteration. This is the phase's headline acceptance criterion.
</action>
<verify>
<automated>cd apps/api && pnpm test -- tender-scheduler 2>&1 | tail -20</automated>
</verify>
<acceptance_criteria>
- Test asserts: after a 2nd tenant activation → additional DÖE fetch calls == 0, additional cron jobs == 0, additional Tender rows == 0.
- `pnpm --filter @tessera/api test -- tender-scheduler` green.
</acceptance_criteria>
<done>Two-tenant safety proven by an automated integration test.</done>
</task>
</tasks>
<threat_model>
## Trust Boundaries
| Boundary | Description |
|----------|-------------|
| background job → global DB | Scheduled ingestion writes to the platform-global Tender table with no tenant context |
| tenant activation → scheduler | A tenant activating the module must not alter platform-wide poll behavior |
## STRIDE Threat Register
| Threat ID | Category | Component | Severity | Disposition | Mitigation Plan |
|-----------|----------|-----------|----------|-------------|-----------------|
| T-10-09 | Information Disclosure | `TenderIngestionService` DB access | high | mitigate | Uses plain `PrismaService` (no `forTenant()`) on the global Tender/config tables by design; wrapping global tables in tenant RLS would silently hide platform data from a 2nd tenant. Verified by the no-forTenant assertion + two-tenant integration test |
| T-10-10 | Elevation of Privilege | scheduler per-tenant scaling | high | mitigate | Exactly one global cron job; `setInterval` takes no tenant arg; singleton config via `findUnique` on a fixed slug. Two-tenant integration test asserts zero additional jobs/calls/rows on a 2nd activation (Success Criteria 4 & 5) |
| T-10-11 | Tampering | outbound `pubDay` URL | medium | mitigate | The day-cursor is internally computed from `lastIngestedDay` (never user-supplied); `nextDayToFetch` bounds it to strictly-past days; no admin "fetch date X" feature added in this phase |
| T-10-12 | Denial of Service | catch-up loop hammering DÖE | low | mitigate | ~1-2s polite delay between successive day fetches during catch-up (RESEARCH Open Question 1); day-cursor gate prevents re-fetching the same day repeatedly |
</threat_model>
<verification>
- `pnpm --filter @tessera/api test -- tender-ingestion tender-scheduler` all green.
- SCHEMA-02 change detection proven (update-not-duplicate on contentHash change).
- INGEST-06 poll-once proven (single global cron, two-tenant test).
- D-01 day-cursor no-op, D-05 retention with null-deadline skip proven.
- No `forTenant` / no `activeTenantId` in the tender services (grep gates green).
</verification>
<success_criteria>
- SCHEMA-02: a changed DÖE notice updates its existing record instead of duplicating.
- INGEST-06: DÖE polled once on a shared admin-configurable interval regardless of tenant count.
- Success Criteria 4 & 5: 2nd-tenant activation triggers no redundant poll / no duplicate ingestion.
</success_criteria>
<output>
Create `.planning/phases/10-ausschreibungs-radar-foundation-d-e-ingestion/10-04-SUMMARY.md` when done.
</output>