Files
tessera-ctl/.planning/research/SUMMARY.md
T
schalli d2ac1997d6 docs: complete project research
Ausschreibungs-Radar (v1.1) STACK/FEATURES/ARCHITECTURE/PITFALLS research plus SUMMARY.md synthesis.
2026-07-17 10:12:59 +02:00

21 KiB

Project Research Summary

Project: Tessera — Ausschreibungs-Radar Module (v1.1 Milestone) Domain: German public-procurement (Vergabe) tender aggregation, filtering, and notification, built as a new module on an existing multi-tenant NestJS/Prisma platform Researched: 2026-07-17 Confidence: MEDIUM-HIGH overall (HIGH on architecture/platform integration and domain-feature patterns; MEDIUM on new-library version currency and unverified DÖE pagination details)

Executive Summary

The Ausschreibungs-Radar module aggregates German public tenders from multiple incompatible sources — a central auth-free API (DÖE OpenData, ~75% of market value by €), two structurally different scraping targets (AI AG NetServer and cosinex Vergabemarktplatz), RSS feeds, and email-alert ingestion for the long tail — into one normalized, searchable, filterable, notifiable catalog. Every mature tender-monitoring product (Stotles, Tendium, native German portals) follows the same shape: ingest -> normalize -> dedupe -> saved-search filter -> results UI -> per-user triage state -> digest/instant notification. Tessera's module maps onto this shape directly and, crucially, can reuse substantial existing platform infrastructure: the DKV module's inbox-provider abstraction (ImapProvider/ExchangeInboxProvider), its SmtpConfig fresh-transport mail pattern, and @nestjs/schedule dynamic cron. No new frontend dependencies are needed at all — the results UI is built entirely on the existing Next.js/shadcn/TanStack Query stack.

The recommended approach is a DÖE-only MVP first: build the full pipeline (schema, one adapter, normalizer, filter engine, saved searches, results UI, read/favourite state, digest + instant notification) against the single easiest, legally unambiguous, highest-value source before touching any scraping. This validates the entire module concept on live data with zero ToS risk, and every subsequent source (AI-AG adapter, cosinex adapter, RSS, email-alert) slots into the same pipeline via a TenderSourceAdapter interface without touching ingestion/matching/notification code. The single most important architectural decision — and the biggest deviation from the DKV module template — is that tender data is platform-global, not tenant-scoped: a DÖE notice is a public fact relevant to every tenant, so ingestion/dedup run once for the whole platform, while only saved searches, match results, and notification preferences are tenant-scoped.

Key risks: (1) legal exposure from scraping vergabe24/aumass, both of which have explicit AGB bans on automated access — this must be a hard-coded denylist, not just documentation; (2) copying the DKV scheduler's known single-tenant findFirst() pattern, which would silently break polling for every tenant but the first; (3) notification storms on first saved-search activation (backfill) or on tender updates re-triggering "new" alerts, both requiring an explicit matched-vs-notified state split; (4) cross-source deduplication complexity — correctly deferred until a second source exists, since dedup against a single source is meaningless. All four are addressed by explicit phase-level design decisions documented in PITFALLS.md and ARCHITECTURE.md.

Key Findings

The v1.1 delta adds five focused libraries to apps/api, no new apps/web dependencies. All additions were checked against the existing CommonJS build target (ruling out ESM-only packages like p-queue) and the established native-fetch-only HTTP pattern (ruling out axios).

Core technologies (new):

  • fast-xml-parser (5.10.1) — parses both eForms-DE XML (DÖE) and RSS 2.0 feeds with one library; zero-dependency, pure-TS, avoids adding a separate stale rss-parser (last released 2023).
  • cheerio (1.2.0) — HTML parsing for the AI-AG and cosinex scraping adapters; 10-50x cheaper than a headless browser for server-rendered HTML.
  • csv-parse (7.0.1) — DÖE CSV export as a secondary/verification format alongside eForms-XML/OCDS-JSON.
  • tough-cookie + fetch-cookie (6.0.2 / 3.2.0) — session-cookie handling for scraping adapters, wrapping native fetch rather than introducing axios.
  • playwright (1.61.1, conditional) — only added if a phase-1 spike confirms a target portal's public search requires client-side JS rendering; not added speculatively (~300MB+ Docker image cost).

Explicitly rejected in favor of reuse: a dedicated DÖE/eForms SDK (none exists; plain REST+parsers suffice), p-queue/bottleneck (ESM-only or unmaintained; a ~20-line internal delay helper suffices at this scale), @nestjs-modules/mailer (documented DKV pitfall — can't change SMTP transport at runtime), a second cron mechanism (reuse @nestjs/schedule + SchedulerRegistry), and a new IMAP/EWS client (reuse InboxProvider/ImapProvider/ExchangeInboxProvider verbatim, extended for email body/HTML access rather than PDF attachments).

Expected Features

Must have (table stakes, DÖE-only MVP):

  • DÖE OpenData ingestion (eForms/OCDS/CSV, auth-free) as the sole MVP data source
  • Normalized OCDS-oriented tender schema (title, buyer, CPV codes, region/PLZ, deadline, value, procedure type, source URL, raw payload)
  • Filter engine: full-text keyword (Postgres tsvector/GIN), region/PLZ/Bundesland, CPV code, deadline range, value range
  • Saved-search/filter profiles, per-user (not just per-tenant) and tenant-aware
  • Searchable/sortable results list + detail view linking to source (never mirror bid documents)
  • Read/unread and favourite/shortlist marking, both per-user state
  • Periodic email digest (configurable interval) and instant alert on new match, both reusing MailModule/SMTP infra

Should have (add after MVP validation, v1.x):

  • AI-AG NetServer adapter (covers lhs-vpbw, tender24, vergabe.landbw + other AI-AG portals via one config-driven adapter)
  • cosinex VMP adapter (DTVP + other Länder marketplaces)
  • RSS ingestion (subreport-elvis, service.bund.de) — lowest-effort source expansion
  • Email-alert ingestion for remaining Unterschwellen portals, reusing DKV inbox infra
  • Cross-source deduplication — mandatory once a 2nd source exists, meaningless before
  • "Manual watch" indicator for vergabe24/aumass (excluded portals, trust-building)

Defer (v2+): TED API v3 (redundant to DÖE for DE-only scope), relevance/ranking scoring, dashboard "upcoming deadlines" widget (high-leverage but depends on Phase-8 widget-SDK reuse), CSV/Excel export, team/collaboration features (scope creep toward bid-management).

Anti-features to actively avoid: scraping vergabe24/aumass directly (ToS-prohibited), a generic "scrape any portal URL" framework (fragile across 3+ incompatible portal families), sub-hourly polling (tender lifecycles run days-to-weeks), full bid-management/CRM scope, AI-generated summaries/scoring (premature), and mirroring/hosting bid documents locally (legal ambiguity).

Architecture Approach

The module is a new, self-contained TendersModule built entirely on existing Tessera infrastructure (module-registry self-seed, dynamic cron via SchedulerRegistry, extracted shared InboxModule, tenant-scoped SMTP). The one deliberate break from the DKV template: Tender rows carry no tenantId — they are platform-global reference data, deduplicated once and shared across all tenants. Tenant scoping lives one layer up, in TenderSavedSearch and TenderMatch. This avoids multiplying scraping requests and storage by tenant count and keeps cross-tenant dedup coherent.

Major components:

  1. TenderSourceAdapter implementations (DÖE, AI-NetServer, cosinex, RSS, email-alert) — one interface, N implementations, directly analogous to the existing InboxProvider pattern; each source is a config-driven adapter, not one adapter per portal instance.
  2. TenderNormalizerService — maps each source's raw shape into the unified OCDS-oriented Tender schema, computes a priority-ordered dedup key (OCID -> sourcePortal:noticeId -> fuzzy fingerprint hash) and a separate contentHash for change detection.
  3. TenderIngestionService + TenderSchedulerService — poll -> normalize -> upsert -> change-detect, with one global cron per source (not per tenant) and per-tenant cron only for the genuinely per-tenant email-alert mailbox path.
  4. TenderMatchingService — evaluates active saved searches as DB-scoped Prisma queries against only the newly-changed delta batch per poll, not full-table in-memory scans; keyword matching via Postgres GIN full-text index, not ILIKE.
  5. TenderMailService — instant (on match creation) and digest (batched cron) notification, reusing the DkvMailService fresh-transport-per-send pattern against tenant-specific SmtpConfig.

Build order ships DÖE-first as a complete, demoable slice (schema -> adapter -> normalizer -> ingestion+scheduler -> matching+controller -> web UI -> mail), then adds scraping adapters (Phase B), then RSS + email-alert ingestion (Phase C, lowest ROI, requires the inbox/ extraction refactor as a prerequisite).

Critical Pitfalls

  1. Copying the DKV scheduler's single-tenant findFirst() pattern — would silently break polling for every tenant except the first. Must design as findMany({ where: { isActive: true } }) with one job per tenant/source from day one; add an explicit two-tenant scheduler test as an acceptance criterion.
  2. Poll-per-tenant instead of poll-once-fan-out-many — naive per-tenant-per-search polling multiplies load against shared upstream portals (throttling/ban risk, looks like abuse). Decouple: poll each upstream source once on a shared schedule, fan results out to all matching tenant saved searches from the ingested data.
  3. Legal/ToS risk scraping vergabe24 and aumass — both have explicit AGB bans on automated/scripted access. Enforce as a hard denylist in code (adapter registry refuses to register a denylisted portal id), not just documentation; their Oberschwellen data is already covered via DÖE, so there's no completeness argument to scrape them.
  4. Notification storms (backfill + duplicate-alert flooding) — first saved-search activation can surface months of historical DÖE matches; a tender update (deadline extension) can re-trigger a "new" alert if only "matched" is tracked, not "notified." Requires a separate matched-vs-notified state (TenderMatch.notifiedAt) with explicit backfill-suppression on saved-search creation.
  5. Cross-source deduplication attempted too early or done as exact-ID match — meaningless with one source; once >=2 sources exist, requires fuzzy fingerprint matching (buyer+title+CPV+deadline+value), not exact ID equality, since no shared cross-portal identifier exists below the eForms/TED notice number.
  6. Tenant-scoping the Tender table like DKV data — the single highest-leverage anti-pattern to avoid; multiplies storage and scraping load N-times for identical public data and makes dedup incoherent.

Implications for Roadmap

Based on research, suggested phase structure:

Phase 1: DÖE-Only Tender Radar (End-to-End MVP)

Rationale: DÖE is auth-free, legally unambiguous, covers ~75% of market value by €, and is the only source needed to validate the entire module concept — schema, filter engine, notifications — against real live data before any scraping work begins. Everything else is additive on top of this pipeline. Delivers: Prisma schema (Tender, TenderSourcePollConfig, TenderSavedSearch, TenderMatch + GIN full-text index), TendersModule skeleton + self-seed, DoeOpenDataAdapter, TenderNormalizerService (OCID-based dedup key + contentHash), TenderIngestionService + global scheduler, TenderMatchingService (delta-scoped DB query filtering), TendersController (list/detail/saved-search CRUD), full Next.js results UI (trefferliste, filters, detail view, saved-search management), read/unread + favourite/shortlist per-user state, TenderMailService (instant + digest, reusing SmtpConfig/DKV mail pattern). Addresses: All P1 table-stakes features from FEATURES.md — ingestion, normalized schema, filter engine, saved searches, results UI, triage state, both notification modes. Avoids: Pitfall 24 (single-tenant scheduler copy — build multi-tenant from day one), Pitfall 23 (notification storms — matched/notified split + backfill suppression), Pitfall 25 (tenant isolation gaps — clear schema split between global Tender and tenant-scoped TenderSavedSearch/TenderMatch).

Phase 2: Scraping Adapters — AI-AG NetServer + cosinex VMP

Rationale: Second-best ROI per feasibility research: one AI-AG adapter covers 3+ portal instances (lhs-vpbw, tender24, vergabe.landbw) plus many unlisted AI-AG-family portals; one cosinex adapter covers DTVP and is reusable for other Länder marketplaces. Both are public, unauthenticated search pages with "Niedrig-Mittel" ToS risk. Building this after Phase 1 means the ingestion/matching/notification pipeline is already proven and untouched — only new adapters plug in. Delivers: AiNetServerAdapter and CosinexAdapter behind the existing TenderSourceAdapter interface, config-driven per portal instance (not one class per portal), with the required cross-source dedup logic (fuzzy fingerprint match) now meaningful for the first time. Uses: cheerio, tough-cookie/fetch-cookie from STACK.md; conditionally playwright only if a phase-start spike confirms JS-rendered search results on either portal. Implements: Pattern 1 (Source-Adapter Abstraction), Pattern 2's cross-source dedup key fallback (sourcePortal:noticeId). Avoids: Pitfall 15 (scraper fragility — session tokens, layout drift; build structural health checks per adapter run from day one), Pitfall 21 (cross-source dedup — fuzzy fingerprint, not exact-ID), Pitfall 16 (legal denylist — vergabe24/aumass must be structurally impossible to add, enforced in the adapter registry, not just documented).

Phase 3: RSS + Email-Alert Long Tail

Rationale: Lowest ROI per feasibility research (do last) but lowest implementation cost — RSS ingestion is a small field-mapper over the already-added fast-xml-parser, and email-alert ingestion directly reuses DKV's inbox infrastructure once extracted into a shared module. This closes coverage on the remaining Unterschwellen long tail (portals 2, 6 via alert-email; 7, 10 via RSS) without adding any new scraping ToS exposure. Delivers: inbox/ shared-module extraction (prerequisite refactor moving ImapProvider/ExchangeInboxProvider out of dkv/), RssAdapter (subreport-elvis, service.bund.de), TenderInboxConfig (per-tenant, mirrors DkvModuleConfig credential pattern) + EmailAlertAdapter + per-tenant scheduler jobs for the genuinely per-tenant mailbox-polling path. Delivers (feature-level): "Manual watch" indicator for vergabe24/aumass, source-coverage transparency in the UI (Pitfall 20 — communicate the Oberschwelle-vs-Unterschwelle completeness gap explicitly rather than presenting a falsely-unified result list). Avoids: Pitfall 20 (false completeness — must ship coverage-transparency metadata alongside this phase, not reactively later).

Phase Ordering Rationale

  • DÖE-first is a hard dependency, not a preference: the filter engine, saved searches, and results UI all require normalized data to exist before they can be built and validated — nothing else can be usefully tested without Phase 1's pipeline in place.
  • Scraping adapters (Phase 2) come before RSS/email-alert (Phase 3) despite RSS/email being individually cheaper, because cross-source deduplication only becomes meaningful and testable once a second scraped source exists at meaningful volume — RSS/email sources are comparatively low-volume long-tail additions that exercise the same dedup logic but don't independently justify building it.
  • The inbox/ extraction refactor is scoped into Phase 3, not Phase 1, because it's only a hard prerequisite for the email-alert adapter — doing it earlier would be premature refactoring of code Phase 1/2 never touch.
  • Global-vs-tenant data split (Pattern 3) and the multi-tenant scheduler (avoiding Pitfall 24) must both be correct in Phase 1, even though only one tenant may be active during initial testing — retrofitting either after data/schedules exist is expensive and risks silent data-model migration bugs.

Research Flags

Phases likely needing deeper research during planning:

  • Phase 1: DÖE OpenData API pagination and rate-limit parameters are explicitly unverified (Swagger UI is JS-rendered; feasibility doc flags this as an open point) — confirm via a live API call early in phase planning, not assumed from docs.
  • Phase 1: OCDS field-mapping edge cases (conditional/optional fields, parties[].name resolution, release-vs-record package choice) warrant a focused read of the OCDS eForms profile during normalizer design, per Pitfalls 17-18.
  • Phase 2: Whether AI-AG NetServer or cosinex VMP search pages require JS rendering is unverified — needs a phase-start spike (confirm cheerio-only feasibility per-portal) before committing to playwright.
  • Phase 2: Exact AGB anti-scraping clause wording for RIB, subreport, deutsche-evergabe, tender24 remains an open verification point from the feasibility doc — recommend a quick manual AGB check before scraping any additional portal beyond AI-AG/cosinex.

Phases with standard patterns (skip research-phase):

  • Phase 1's notification/mail integration: directly copies the proven DkvMailService pattern — no new research needed.
  • Phase 3's inbox-module extraction: mechanical refactor of already-working, already-understood code (ImapProvider/ExchangeInboxProvider).

Confidence Assessment

Area Confidence Notes
Stack MEDIUM Versions cross-verified via WebSearch + direct npm registry lookup (authoritative for version/publish-date facts); no Context7/docs-MCP available this session for qualitative maintenance-health claims. All five new packages are well-known, actively-maintained (Playwright is Microsoft's own project) despite an automated SUS flag the researcher identified as a false positive.
Features MEDIUM-HIGH Domain patterns cross-checked against live tender-monitoring SaaS products (Stotles, Tendium, TenderAlerts) plus the OCDS standard; German-specific integration details inherit HIGH confidence from the project's own prior ausschreibungs-portale-feasibility.md research.
Architecture HIGH for platform integration (module registration, tenant scoping, inbox/mail reuse — verified against live dkv/ / module-registry/ source); MEDIUM for OCDS field mapping and DÖE API pagination (verified via spec/Swagger listing, not a live API call).
Pitfalls HIGH Grounded in the project's own feasibility research, official OCDS-for-eForms docs, DÖE OpenData API docs, existing codebase patterns (dkv-scheduler.service.ts, tenant.guard.ts), and German scraping case law (BGH I ZR 224/12). MEDIUM specifically on exact DÖE pagination/rate-limit behavior, an explicitly flagged open verification point.

Overall confidence: MEDIUM-HIGH

Gaps to Address

  • DÖE OpenData API pagination/rate-limit parameters: Swagger UI is JS-rendered and wasn't live-queried this session. Resolve with a direct API call during Phase 1 planning before finalizing poll cadence.
  • Whether AI-AG NetServer / cosinex VMP search pages require JS rendering: unverified; drives the playwright add/skip decision for Phase 2. Resolve with a phase-start technical spike, not assumed.
  • Exact AGB anti-scraping clause wording for RIB, subreport, deutsche-evergabe, tender24: flagged as open in the feasibility doc; low urgency since none of these are in the Phase 1-2 build order, but should be checked before any Phase 3+ expansion beyond the currently-scoped RSS/email-alert sources.
  • Whether vergabe.landbw.de offers an on-page email alert: open verification point from feasibility doc; relevant only if/when its email-alert ingestion is scoped into Phase 3.
  • Package "SUS" flags on fast-xml-parser/csv-parse/playwright/p-queue-alternatives: researcher assessed as false positives (recent-publish velocity, not actual supply-chain risk) but recommends a final npm view/Socket.dev spot-check immediately before pnpm add in Phase 1 implementation, per the repo's existing dependency-hygiene bar.

Sources

Primary (HIGH confidence)

  • Existing codebase (direct read): apps/api/src/dkv/, apps/api/src/module-registry/, apps/api/src/mail/, apps/api/src/prisma/prisma-tenant.extension.ts, apps/api/src/tenant/tenant.middleware.ts, apps/api/prisma/schema.prisma, apps/web/src/lib/module-loader.ts, apps/web/src/app/(portal)/modules/dkv-fleet/
  • .planning/research/ausschreibungs-portale-feasibility.md (project-internal, 2026-07-16) — portal-by-portal capability matrix, ToS risk ratings, DÖE/TED coverage statistics
  • registry.npmjs.org — direct version/publish-date lookup for all new libraries (fetched 2026-07-17)
  • BGH, 30.04.2014 — I ZR 224/12 — German case law on scraping legality and "virtuelles Hausrecht"

Secondary (MEDIUM confidence)

  • Open Contracting Data Standard — Schema/Codelists Reference (standard.open-contracting.org) — OCDS field structure, CPV usage, release-vs-record model
  • oeffentlichevergabe.de OpenData Swagger UI — confirms ocds-mnwr74 German federal OCDS prefix, CC-Zero licensing (page requires JS for full pagination detail)
  • Stotles, Tendium, TenderAlerts.eu — competitive feature-pattern confirmation
  • WebSearch cross-checked against npm registry — ecosystem status for cheerio/fast-xml-parser/csv-parse/playwright and rejected alternatives

Tertiary (LOW confidence)

  • None flagged — all findings traced to at least a MEDIUM-confidence primary or secondary source.

Research completed: 2026-07-17 Ready for roadmap: yes