docs: complete project research
Ausschreibungs-Radar (v1.1) STACK/FEATURES/ARCHITECTURE/PITFALLS research plus SUMMARY.md synthesis.
This commit is contained in:
@@ -2,7 +2,9 @@
|
||||
|
||||
**Domain:** Modular portal platform with marketplace, multi-tenancy, Docker deployment
|
||||
**Project:** Tessera
|
||||
**Researched:** 2026-06-18
|
||||
**Researched:** 2026-06-18 (v1.0 platform pitfalls below); 2026-07-17 (v1.1 Ausschreibungs-Radar module pitfalls in the dedicated section further down)
|
||||
|
||||
> This file accumulates pitfalls research across milestones. The v1.0 section below covers platform-wide architecture pitfalls (multi-tenancy, module system, Docker, i18n, etc.) and remains valid for all subsequent modules built on Tessera. The **v1.1 Ausschreibungs-Radar** section adds pitfalls specific to building a multi-source tender-aggregation module (scraping, eForms/OCDS ingestion, dedup, notifications) on top of that platform.
|
||||
|
||||
## Critical Pitfalls
|
||||
|
||||
@@ -267,7 +269,7 @@ Mistakes that cause rewrites, data breaches, or architectural dead-ends.
|
||||
| Deployment | Volume permissions break non-root containers | Named volumes, entrypoint permission scripts |
|
||||
| VCS integration | Gitea-specific tight coupling | Abstract behind interface, use standard Git ops |
|
||||
|
||||
## Sources
|
||||
## Sources (v1.0 platform pitfalls)
|
||||
|
||||
- [Multi-Tenant SaaS Architecture: What Nobody Tells You Before You Build](https://dev.to/actinode/multi-tenant-saas-architecture-what-nobody-tells-you-before-you-build-a4h) - Confidence: HIGH
|
||||
- [Designing Multi-Tenant SaaS Architecture: Mistakes to Avoid](https://www.saasadviser.co/blog/multi-tenant-saas-architecture-mistakes-best-practices) - Confidence: MEDIUM
|
||||
@@ -280,3 +282,383 @@ Mistakes that cause rewrites, data breaches, or architectural dead-ends.
|
||||
- [Electron vs. Tauri](https://www.dolthub.com/blog/2025-11-13-electron-vs-tauri/) - Confidence: MEDIUM
|
||||
- [Building Interactive Dashboards with React Grid Layout](https://www.ilert.com/blog/building-interactive-dashboards-why-react-grid-layout-was-our-best-choice) - Confidence: MEDIUM
|
||||
- [AWS: Multi-tenant data isolation with PostgreSQL Row Level Security](https://aws.amazon.com/blogs/database/multi-tenant-data-isolation-with-postgresql-row-level-security/) - Confidence: HIGH
|
||||
|
||||
---
|
||||
|
||||
# v1.1 Milestone: Ausschreibungs-Radar — Multi-Source Tender Aggregation Pitfalls
|
||||
|
||||
**Domain:** Multi-source tender/procurement aggregation module (scraping + eForms/OCDS ingestion + email-alert ingestion + normalization + dedup + notification), built as a new NestJS/Prisma/PostgreSQL-RLS module on the existing multi-tenant Tessera platform.
|
||||
**Researched:** 2026-07-17
|
||||
**Confidence:** HIGH (grounded in `.planning/research/ausschreibungs-portale-feasibility.md`, official OCDS-for-eForms docs, DÖE OpenData API docs, existing Tessera codebase patterns — `dkv-scheduler.service.ts`, `tenant.guard.ts`, `DkvModuleConfig` schema — and German scraping case law). MEDIUM on exact DÖE pagination/rate-limit behavior (Swagger is JS-rendered, not yet live-verified per the feasibility doc's open verification points).
|
||||
|
||||
## Critical Pitfalls
|
||||
|
||||
### Pitfall 15: HTML scraper fragility — session tokens, jsessionid, layout drift
|
||||
|
||||
**What goes wrong:**
|
||||
AI-AG NetServer (lhs-vpbw, tender24, vergabe.landbw) and cosinex VMP (DTVP) both gate the "public search" behind server-side session state — `jsessionid` in the URL or a hidden CSRF/viewstate token in the search form that must be replayed on every paginated request. A scraper that treats these as static query params breaks the moment the portal rotates the session, adds a token, or reflows the HTML (even a CSS-only redesign can shift selectors). Because both platforms are proprietary, undocumented, and outside Tessera's control, there is no changelog or deprecation notice — the adapter just silently starts returning zero results or garbage.
|
||||
|
||||
**Why it happens:**
|
||||
Scrapers are built once against a snapshot of the DOM/session flow and treated as "done." Portal vendors change markup, add bot-detection (rate-based session invalidation), or migrate frontend frameworks without any notice to third parties, because scraping was never a supported integration path.
|
||||
|
||||
**How to avoid:**
|
||||
- Build a thin `PortalAdapter` interface (fetch session → search → paginate → parse detail) per platform (one AI-AG adapter, one cosinex adapter — per the feasibility doc's Effort/Value ranking), not per portal instance, so a fix in one place covers 4/8/9 (AI-AG) or all cosinex marketplaces.
|
||||
- Never hardcode a session token's lifetime — always re-derive it from the search page response on each scrape run, don't cache it across runs.
|
||||
- Store raw HTML of the search + detail pages for the last N successful runs (short retention, not permanent) so a layout-change diagnosis doesn't require reproducing the failure live.
|
||||
- Add a structural health check per adapter: assert on stable anchors (e.g., "did we get >0 rows AND did known static fields — DE, umlauts, CPV-looking codes — parse") before treating a scrape as successful; a scraper that "succeeds" with zero rows for days is a silent failure, not a quiet success.
|
||||
|
||||
**Warning signs:**
|
||||
Result count drops to zero (or spikes to an implausible number) for a portal that previously returned steady volume; parse errors on fields that were previously stable; HTTP 200 responses with unexpected redirect chains (session expiry manifests as a redirect to a login/error page, not a 4xx).
|
||||
|
||||
**Phase to address:**
|
||||
Portal Adapter phase (AI-AG + cosinex scrapers) — build the adapter interface with health-check-on-every-run baked in from the first adapter, not bolted on later.
|
||||
|
||||
---
|
||||
|
||||
### Pitfall 16: Legal/ToS risk — scraping portals with explicit automated-access bans
|
||||
|
||||
**What goes wrong:**
|
||||
vergabe24.de and plattform.aumass.de have AGB clauses that explicitly forbid automated extraction/scripted access (vergabe24 even names rate limits in its ToS). German case law (BGH, 30.04.2014 — I ZR 224/12) shows scraping itself is not per se illegal and a "virtuelles Hausrecht" has no independent legal basis — but an *effectively incorporated* AGB prohibition, combined with any technical protection measure (bot detection, CAPTCHA, rate limiting) the portal has in place, shifts the analysis toward Wettbewerbsrecht (§ 3a UWG — Rechtsbruch) and potential Datenbankherstellerrecht (§ 87a UrhG) claims, since both portals invest in curating/aggregating tender data as their core product. Building against these two anyway (or extending a generic scraper framework to "just try" them later) creates real legal exposure that has nothing to do with code quality.
|
||||
|
||||
**Why it happens:**
|
||||
A generic `PortalAdapter` abstraction makes it *technically* trivial to add a new source once the interface exists — the temptation to "just add vergabe24, we already have the crawler" bypasses the legal review that should gate it.
|
||||
|
||||
**How to avoid:**
|
||||
- Hard exclusion, enforced in code, not just documentation: maintain an explicit denylist/allowlist of source identifiers in config (not just a comment), and have the adapter registry refuse to register an adapter for a denylisted portal id even if someone writes the code.
|
||||
- Their Oberschwelle notices are already covered via DÖE (per the feasibility doc), so there is no data-completeness reason to scrape them — document this rationale next to the denylist so a future contributor doesn't "rediscover" the idea without the legal context.
|
||||
- If Unterschwelle coverage from these two ever becomes a business requirement, the only acceptable path is the same one already used for portals 2/3/4/6/7/8/9/10: register a native saved-search + ingest the resulting alert email via the existing DKV inbox infrastructure — not HTML scraping.
|
||||
|
||||
**Warning signs:**
|
||||
A PR adds a new adapter or config entry referencing `vergabe24` or `aumass` in any capacity beyond "denylisted"; someone proposes a "generic connector" that accepts an arbitrary portal URL from tenant admins (this would let a customer point the scraper at a banned portal without Tessera's own code ever naming it).
|
||||
|
||||
**Phase to address:**
|
||||
Portal Adapter phase — encode the denylist as a first-class artifact (config + registry guard + test) before any adapter ships, so it's structurally impossible to "just add" a banned source. Revisit only via explicit product decision, not incidentally.
|
||||
|
||||
---
|
||||
|
||||
### Pitfall 17: eForms-DE / SDK schema version drift breaks XPath and field mapping
|
||||
|
||||
**What goes wrong:**
|
||||
eForms-DE is based on the KoSIT eForms SDK, which the EU Publications Office revises periodically (new SDK majors/minors introduce renamed elements, new mandatory fields, or restructured repeatable groups). A parser hardcoded against one SDK version's XPath will silently drop fields — or throw on notices published under a newer SDK — the moment DÖE starts forwarding notices tagged with the new version. Because DÖE re-publishes notices in the *exact* schema version the buyer's e-Sender submitted, a single ingestion run can contain a mix of SDK versions.
|
||||
|
||||
**Why it happens:**
|
||||
Developers build the XML parser against a handful of sample notices at build time and never revisit it; the SDK version is often only visible in a namespace/version attribute deep in the document, easy to ignore until it breaks something.
|
||||
|
||||
**How to avoid:**
|
||||
- Prefer OCDS (`ocds-mnwr74`) over raw eForms-DE XML wherever DÖE offers both — the OCDS profile's official field mappings (open-contracting-extensions/eforms on GitHub) are versioned and maintained upstream, absorbing most SDK churn for you.
|
||||
- If eForms-DE XML is parsed directly (e.g., for fields OCDS doesn't map), read and log the SDK version attribute on every notice and fail loud (not silently drop) on an unrecognized version rather than best-effort parsing it.
|
||||
- Keep a small fixture library of real notices per encountered SDK version for regression tests — this is a case where "test against production data snapshots" beats synthetic fixtures, because the actual failure mode is structural drift you can't anticipate.
|
||||
|
||||
**Warning signs:**
|
||||
A sudden spike in notices with empty/null fields that used to populate; parser exceptions correlating with a specific publication date range (SDK version cutovers happen on fixed EU-mandated dates); OCDS `previouslyWithheldInformation` releases not being picked up (per official docs, this is a self-managed process — TED handles scheduled release, but OCDS consumers must poll for the redaction-lift date themselves).
|
||||
|
||||
**Phase to address:**
|
||||
DÖE/OCDS Ingestion phase — build the SDK-version-aware ingestion pipeline (log + fail loud on unknown versions) as part of the first central-feed integration, since this is the highest-ROI source per the feasibility doc and will run continuously from day one.
|
||||
|
||||
---
|
||||
|
||||
### Pitfall 18: OCDS optional/conditional fields treated as "always present"
|
||||
|
||||
**What goes wrong:**
|
||||
The OCDS eForms profile has fields that are conditionally present based on procedure type, threshold, or procurement stage (e.g., award fields don't exist until an award release; framework-agreement cascade awards look like multiple suppliers on one award but must be distinguished from joint awards by checking `lot.techniques.frameworkAgreement` + bid-ranking presence, per the official "how to use" guide). A normalizer written against a handful of "happy path" above-threshold notices will crash or silently null out data for the many below-threshold / early-stage notices that don't yet have those fields.
|
||||
|
||||
Additionally, `OrganizationReference` objects (buyer, tenderer, supplier) in OCDS only contain an `id` by default — the human-readable `.name` must be resolved by cross-referencing the `parties` array. Skipping this step produces a UI full of opaque org IDs instead of buyer names, which will look broken to end users even though ingestion "succeeded."
|
||||
|
||||
**Why it happens:**
|
||||
OCDS is release-based (multiple releases per contracting process, merged into a "record") — a normalizer that only looks at the latest release, or treats every field as always-populated, misses this structure.
|
||||
|
||||
**How to avoid:**
|
||||
- Treat every OCDS field beyond `id`/`title`/`tender.status` as optional in the normalized schema; write the normalizer defensively (missing ≠ error, just means "not yet known at this stage").
|
||||
- Explicitly implement the `parties[].name` resolution step for every `OrganizationReference` location (buyer, tenderers, suppliers, procuringEntity) before display — this is a documented, mandatory post-processing step, not an edge case.
|
||||
- Decide up front whether Tessera consumes the OCDS "release" or "record" package (record = pre-merged current state, generally simpler for a read-only aggregator than reconciling releases yourself).
|
||||
|
||||
**Warning signs:**
|
||||
UI shows organization IDs (UUID-looking strings) instead of names; award/value fields blank for tenders still in "planning" or "tender" stage that should legitimately have no award data yet vs. tenders where the data genuinely failed to parse — these two cases must be distinguishable in the schema (e.g., a `parseWarnings` field), not conflated.
|
||||
|
||||
**Phase to address:**
|
||||
DÖE/OCDS Ingestion phase and Normalization phase — build the normalized internal schema to make "field not yet applicable at this stage" and "field failed to parse" distinct states from the start; retrofitting this distinction after the UI already treats blanks as identical is expensive.
|
||||
|
||||
---
|
||||
|
||||
### Pitfall 19: CPV code format inconsistencies break category filtering
|
||||
|
||||
**What goes wrong:**
|
||||
CPV (Common Procurement Vocabulary) codes appear in different shapes across sources: full 8-digit + check-digit form (`45000000-7`) in eForms/TED, sometimes truncated or zero-padded differently in scraped HTML (portals often display only the human label, e.g. "Bauleistungen", without the code at all), and CPV is hierarchical (a filter on "45xxxxxx — Bauarbeiten" should match all more-specific codes underneath it). A filter engine built against exact-string CPV matches will miss the majority of relevant results because most notices carry a specific leaf code, not the parent category a user searched for.
|
||||
|
||||
**Why it happens:**
|
||||
CPV's hierarchy (division → group → class → category → subcategory, encoded positionally in the 8 digits) isn't obvious from a single sample notice, and portals that don't expose CPV at all (many of the AI-AG/cosinex HTML listings show only free-text category labels) force a fallback keyword-matching path that behaves completely differently from the CPV-code path.
|
||||
|
||||
**How to avoid:**
|
||||
- Normalize CPV to the canonical 8-digit+check-digit string (strip formatting, validate against the official CPV code list) in the ingestion layer, never in the filter/UI layer.
|
||||
- Implement hierarchical CPV matching (prefix match on the first N significant digits) as the actual filter semantics, not exact match — expose this in the filter UI as "category" (broad) vs. exact code.
|
||||
- For sources without native CPV (HTML-only portals), map their free-text category taxonomy to CPV divisions in a small lookup table rather than pretending they're equivalent to code-based filtering — flag these results as "category: approximate" in the schema so users understand the precision difference.
|
||||
|
||||
**Warning signs:**
|
||||
A CPV-based saved search returns near-zero results despite users reporting matching tenders exist; filter results differ wildly in volume between DÖE-sourced (CPV-tagged) and portal-scraped (label-only) records for the same category.
|
||||
|
||||
**Phase to address:**
|
||||
Normalization phase (CPV canonicalization + hierarchy) and Filter Engine phase (hierarchical matching semantics) — the canonical CPV table should be seeded once, early, since it's static reference data all sources map into.
|
||||
|
||||
---
|
||||
|
||||
### Pitfall 20: Treating below-threshold coverage as complete ("false completeness")
|
||||
|
||||
**What goes wrong:**
|
||||
Per the feasibility research, DÖE/TED guarantee ~100% of *above-threshold* (Oberschwelle) notices (~12% of procedures by count, ~75% by value) but only an estimated 20–35% of *below-threshold* (Unterschwelle) notices by count today, since Unterschwelle eForms publication is only mandatory for Bund and voluntary elsewhere. If the product doesn't clearly communicate this gap, a tenant configuring a saved search sees "0 results" or a thin result set for their region/CPV and reasonably concludes either "nothing matches" or "the module is broken" — when the real answer is "this data isn't centrally available yet, only reachable via the individual portal (which may not even be in Tessera's covered set)."
|
||||
|
||||
**Why it happens:**
|
||||
The DÖE API looks and feels complete (clean OCDS/eForms structure, no visible "gaps" in the data itself) — there's no signal in the API response that tells you what's *missing*, only what's present.
|
||||
|
||||
**How to avoid:**
|
||||
- Explicitly model source coverage per source in the schema/UI: each normalized tender record carries its source(s) and, more importantly, each saved search result set is annotated with which sources were queried and their known coverage tier (central-guaranteed vs. portal-partial vs. not-covered).
|
||||
- Surface this in the UI ("Diese Suche deckt DÖE (Oberschwelle, vollständig) + 3 Portale (Unterschwelle, teilweise) ab — für vollständige Unterschwellen-Abdeckung: X weitere Portale nicht angebunden") rather than presenting a unified, seemingly-authoritative result list.
|
||||
- Track this as a living fact, not a one-time note — the "20–35%" figure is explicitly stated to be *increasing* as the Unterschwellen-Pflicht rolls out, so hardcoding "this data is incomplete" copy without a mechanism to update it will itself become stale/misleading.
|
||||
|
||||
**Warning signs:**
|
||||
Support requests along the lines of "why didn't I get an alert for a tender I found manually on a portal you claim to cover" — this is a coverage-transparency failure, not necessarily a bug.
|
||||
|
||||
**Phase to address:**
|
||||
Normalization phase (source/coverage metadata on every record) and UI/Filter Engine phase (surfacing coverage transparency) — should ship with the MVP result list, not be added reactively after user confusion.
|
||||
|
||||
---
|
||||
|
||||
### Pitfall 21: Cross-source deduplication — same tender, different IDs
|
||||
|
||||
**What goes wrong:**
|
||||
A single tender can legitimately appear via multiple paths: the buyer's e-Sender pushes eForms-DE to DÖE (→ also mirrored to TED), the same tender is *also* listed natively on the buyer's chosen portal (DTVP, an AI-AG portal, etc.) with a portal-internal ID, and if the tenant also has an email-alert saved search on that portal, the same tender arrives a third time via inbox ingestion. Each path uses a different identifier scheme (DÖE/TED notice number vs. portal-internal Vergabenummer vs. whatever the buyer typed in the alert email subject), different field completeness, and often slightly different publication timestamps (portal listing can precede or lag the central eForms push by hours to days). A naive "unique by ID" dedup produces 2–3 duplicate cards for the same real-world procurement; a naive "unique by title" produces false merges of genuinely different tenders with similar names (common for framework/rebid procedures).
|
||||
|
||||
**Why it happens:**
|
||||
There is no shared cross-portal identifier for German procurement below the eForms/TED notice number, and even that number isn't always echoed back verbatim on the originating portal's HTML page.
|
||||
|
||||
**How to avoid:**
|
||||
- Build a fingerprint-based dedup, not ID-based: normalize (buyer name, CPV/category, region/PLZ, deadline date, and a fuzzy-matched title) into a composite key; use fuzzy string matching (e.g., trigram similarity) with a confidence threshold, not exact equality, on the title component.
|
||||
- When the eForms/TED notice number *is* present in a scraped or emailed record (often referenced as "Bekanntmachungs-ID" or similar), treat it as a strong signal but not the sole key — validate it against the fuzzy fingerprint before merging, since transcription errors happen.
|
||||
- Prefer the DÖE/eForms record as the canonical source of truth when a duplicate is detected (richest structured data); merge portal- and email-sourced duplicates into it as "also seen on X" rather than discarding them — this also gives you a natural cross-check for the coverage-transparency feature (Pitfall 20).
|
||||
- Store the dedup decision (merged-from IDs, confidence score) so a human can review/undo false merges — this is an area where silent auto-merge will eventually be visibly wrong to a user who tracked a tender manually.
|
||||
|
||||
**Warning signs:**
|
||||
Tenants report seeing "the same tender twice" or, conversely, ask why a tender they know overlaps with another wasn't flagged as related; dedup confidence scores clustering near the threshold boundary (indicates the fingerprint algorithm needs tuning, not that the data is ambiguous).
|
||||
|
||||
**Phase to address:**
|
||||
Normalization & Dedup phase — this deserves to be its own phase (not folded into ingestion), since it's the piece most likely to need iteration after real multi-source data is flowing and edge cases surface.
|
||||
|
||||
---
|
||||
|
||||
### Pitfall 22: Deadline/timezone handling for "only still-open" filtering
|
||||
|
||||
**What goes wrong:**
|
||||
eForms/OCDS deadline fields are typically ISO 8601 with explicit timezone (UTC or CEST/CET offset), but scraped portal HTML frequently shows local dates/times without explicit timezone, in German format (`DD.MM.YYYY, HH:MM Uhr`), and email-alert bodies vary by portal template. A filter that does naive string-to-Date parsing without normalizing to a single timezone will misjudge "still open" near midnight boundaries and across DST transitions (CEST↔CET) — either showing an already-closed tender as open (embarrassing if a tenant relies on it) or hiding a still-open one. Deadlines expressed only as a date (no time) also need an explicit convention (end-of-day in which timezone?) since German procurement deadlines are almost always "Datum, HH:MM Uhr" — treating a date-only field as "midnight UTC" silently shortens the effective window by hours.
|
||||
|
||||
**Why it happens:**
|
||||
Timezone bugs are invisible in testing unless tests specifically straddle DST transitions or midnight-local boundaries; developers default to `new Date(string)` parsing without checking what timezone the source actually meant.
|
||||
|
||||
**How to avoid:**
|
||||
- Normalize all deadlines to UTC at ingestion time, with the source timezone made explicit per source (DÖE/eForms: use the timezone offset in the ISO string as-is; scraped portals: assume Europe/Berlin unless the portal states otherwise, and record which assumption was applied).
|
||||
- "Still open" filtering must compare against `now()` in UTC, not local server time — verify the deployment container's timezone doesn't leak into date math (Node's `Date` object is UTC-internal but naive string parsing of ambiguous local strings is where bugs live).
|
||||
- Test explicitly across a DST transition date and around a deadline that falls exactly at day-boundary local time.
|
||||
|
||||
**Warning signs:**
|
||||
A tender flips between "open"/"closed" status on page refresh near its deadline (indicates inconsistent timezone handling between where filtering happens vs. where display happens); users report a "still open" tender they click into is actually already closed on the source portal.
|
||||
|
||||
**Phase to address:**
|
||||
Normalization phase (UTC normalization at ingestion) and Filter Engine phase (open/closed comparison logic) — cover with explicit DST/midnight-boundary test cases before the filter engine phase is marked done.
|
||||
|
||||
---
|
||||
|
||||
### Pitfall 23: Notification storms — first-run backfill and duplicate/immediate-alert flooding
|
||||
|
||||
**What goes wrong:**
|
||||
Two related failure modes: (a) on a tenant's first saved search, DÖE alone can return months of matching historical notices — if the notification pipeline treats "newly matched by this search" the same on day one as it does on day two, the tenant's first experience is an inbox flooded with hundreds of "new tender" emails instead of a clean digest; (b) once running, a tender that gets *updated* (deadline extension, correction notice) re-appears in the source feed and, if the notification logic keys only on "is this a new match" rather than "have I already notified this tenant about this specific tender," triggers a second immediate alert for something they already saw — training users to ignore or unsubscribe from alerts entirely.
|
||||
|
||||
**Why it happens:**
|
||||
"New" is ambiguous — new to the source feed vs. new to this tenant's search vs. never-notified-before are three different conditions that get conflated when the notification trigger is implemented as "insert into results table → fire notification" without a separate notified-state tracking table.
|
||||
|
||||
**How to avoid:**
|
||||
- Separate "matched" from "notified": every (tenant, saved-search, tender) match is recorded, but notification firing is a distinct step gated by explicit backfill handling — on first activation of a saved search, mark all currently-matching historical results as seen/backfilled *without* notifying, and only notify going forward for genuinely new matches (or start the digest from "activation time," clearly communicated to the user).
|
||||
- Track per-tender "last notified version/hash" so an update to an already-notified tender triggers, at most, a distinct "updated" notification (clearly differentiated from "new"), not a duplicate "new tender" alert — and make this configurable (some tenants want deadline-extension alerts, most don't want to see the same tender twice).
|
||||
- Default to digest (periodic batch) rather than instant alerts for saved searches with a strong deadline signal that they'll otherwise return many historical hits (e.g., broad CPV + wide region); reserve instant "Sofort-Alert" for narrow, precise searches where volume is naturally low — this should be a UX default/nudge, not just a raw toggle.
|
||||
- Rate-limit/batch outbound email regardless — even a legitimately large first-run result set should render as one digest email with N results, never N individual emails.
|
||||
|
||||
**Warning signs:**
|
||||
A tenant's first day includes an abnormal spike in outbound emails compared to steady-state; support complaints about "getting the same tender email again"; email provider (existing DKV SMTP infra) flags or throttles the sending account for burst volume.
|
||||
|
||||
**Phase to address:**
|
||||
Notification phase — the matched/notified separation and backfill-suppression logic must be designed before any saved search goes live, since retrofitting it after tenants have already been flooded once is a trust problem, not just a technical one.
|
||||
|
||||
---
|
||||
|
||||
### Pitfall 24: Copying the DKV scheduler's single-tenant pattern for the tender module
|
||||
|
||||
**What goes wrong:**
|
||||
The existing `DkvSchedulerService` (`apps/api/src/dkv/dkv-scheduler.service.ts`) is explicitly documented as v1 single-tenant: it loads config via `findFirst()` and runs one cron job for `activeTenantId`, with a code comment stating "multi-tenant scheduling (one cron job per active tenant) is deferred to a future plan." If the tender module's polling scheduler (for portal adapters, email-alert ingestion, DÖE polling) is built by copy-pasting this pattern, only one tenant's saved searches will ever actually poll in a multi-tenant deployment — every other tenant's saved searches will silently never run, with no error, because the scheduler never even looks for their config.
|
||||
|
||||
**Why it happens:**
|
||||
The DKV scheduler is the most recent, most similar in-repo reference implementation for "polling infra + cron job management" — it's the natural template to copy, and its single-tenant limitation is documented in a code comment that's easy to miss when skimming for the cron-registration pattern, not the architecture note.
|
||||
|
||||
**How to avoid:**
|
||||
- Design the tender module's scheduler as N cron jobs (or a single dispatcher cron that iterates all active tenant configs) from the start — Tessera's multi-tenancy is an explicit from-day-one architectural constraint (per PROJECT.md), and a new module regressing to single-tenant scheduling would be a step backward, not a shortcut.
|
||||
- Reuse the *mechanics* of `SchedulerRegistry.addCronJob()` / dynamic interval updates from the DKV pattern (these are sound), but replace `findFirst()` with `findMany({ where: { isActive: true } })` and register one job per tenant (or per saved-search, depending on granularity chosen), keyed by a job name that includes the tenant/search ID.
|
||||
- Add an explicit multi-tenant scheduler test (two tenants, two configs, both must poll independently) as an acceptance criterion for this phase — don't rely on manual review to catch a `findFirst()` regression.
|
||||
|
||||
**Warning signs:**
|
||||
A second tenant activates a saved search and never receives results/notifications despite valid config; scheduler logs only ever mention one tenant ID across a multi-tenant deployment.
|
||||
|
||||
**Phase to address:**
|
||||
Scheduler/Ingestion Orchestration phase — explicitly call out "multi-tenant, not single-tenant like DKV v1" in the phase's acceptance criteria, since this is a known, named trap already present in the codebase.
|
||||
|
||||
---
|
||||
|
||||
### Pitfall 25: Multi-tenant isolation gaps for saved searches, results, and stored portal credentials
|
||||
|
||||
**What goes wrong:**
|
||||
Three distinct isolation surfaces need to hold: (a) saved-search *configuration* (keywords, CPV, region filters) must be tenant-scoped — a leak here exposes one tenant's business intelligence (what they're bidding on) to another; (b) *results* (which tenders matched which tenant's search, and any tenant-specific annotations/status like "we're bidding on this") must be tenant-scoped even though the underlying tender data itself (from DÖE/portals) is shared, public, non-tenant-specific reference data — the natural bug here is applying tenant filtering to the shared tender catalog (wrong) instead of to the join table between tenant-saved-searches and tenders (right); (c) any stored credentials for portal saved-search registration (if the module ever needs a portal login to register a saved search on a tenant's behalf) must never be shared across tenants even if two tenants use the same portal.
|
||||
|
||||
**Why it happens:**
|
||||
The existing `TenantGuard` pattern scopes `req.tenantPrisma` correctly for request-scoped controller calls, but background jobs (scheduler-triggered ingestion, notification dispatch) run outside any HTTP request context — there is no `req` to derive `tenantId` from, so it's easy to accidentally use the raw (non-tenant-scoped) `PrismaService` for a background operation and forget to add the `tenantId` filter by hand, especially on the *shared reference data* tables where "no tenant filter" is often correct (the tender catalog itself) and it's easy to reflexively skip tenant filtering on the *adjacent* tables that do need it (saved searches, match results, credentials).
|
||||
|
||||
**How to avoid:**
|
||||
- Model the schema with a clear split: tender/notice records (shared, no `tenantId`) vs. saved-search, match-result, notification-log, and portal-credential records (all `tenantId`-scoped, indexed, following the `DkvModuleConfig` precedent of `tenantId String @unique` per-config or `tenantId String` + `@@index([tenantId])` for multi-row tables).
|
||||
- In background/cron contexts, explicitly call `forTenant(prisma, tenantId)` (the existing extension used by `TenantGuard`) per-tenant iteration — never fall back to the raw `PrismaService` for tenant-scoped tables just because there's no request object.
|
||||
- If portal credentials are ever stored (for saved-search registration on AI-AG/cosinex portals), reuse the exact `encryptedInboxCreds` AES-256-GCM pattern from `DkvModuleConfig` — don't invent a new credential-storage mechanism for this module.
|
||||
- Add a cross-tenant isolation test as a standard phase gate: two tenants, overlapping saved-search keywords, verify tenant A never sees tenant B's saved-search config, match annotations, or receives tenant B's notifications.
|
||||
|
||||
**Warning signs:**
|
||||
A query joins the shared tender table directly to a tenant-scoped table without an explicit `tenantId` predicate on the join; background job code imports `PrismaService` directly instead of iterating tenants and using `forTenant()`.
|
||||
|
||||
**Phase to address:**
|
||||
Multi-Tenant Saved Searches phase (schema + isolation tests) — should be verified with an explicit two-tenant UAT before the Notification phase ships, since notification is the surface where a leak becomes user-visible (wrong tenant's alert email).
|
||||
|
||||
---
|
||||
|
||||
### Pitfall 26: Scheduler/rate-limit — polling too aggressively and getting throttled or blocked
|
||||
|
||||
**What goes wrong:**
|
||||
With potentially many tenants each running saved searches against the same underlying portal adapters (AI-AG, cosinex) or the same DÖE API, a naive per-tenant-per-search polling design multiplies request volume against a small number of actual upstream endpoints — e.g., 50 tenants each polling the same AI-AG adapter every 15 minutes doesn't mean 50x the useful data, it means 50x the load on one portal for redundant queries, risking IP-based rate limiting, temporary bans, or (worse) inviting the exact "AGB verbietet Skripte, nennt Raten" scrutiny the feasibility research flagged for vergabe24 — even on portals without an explicit ban, aggressive undifferentiated polling looks like abuse.
|
||||
|
||||
**Why it happens:**
|
||||
Scheduling is naturally designed per-tenant (each tenant's saved search has its own poll interval preference), but the underlying data source is shared infrastructure — without a dedup/coalescing layer, the scheduler design conflates "how often does tenant X want fresh results" with "how often should we actually hit the upstream portal."
|
||||
|
||||
**How to avoid:**
|
||||
- Decouple polling from tenant preference: poll each upstream source (DÖE, each AI-AG instance, cosinex, each RSS feed) on a single shared schedule per source (not per tenant), then fan the results out to all matching tenant saved searches from the ingested/normalized data — this also directly fixes the redundant-load problem and is a natural extension of the "one AI-AG adapter, one cosinex adapter" architecture already recommended.
|
||||
- Respect explicit rate signals: HTTP `Retry-After` headers, documented portal rate limits (vergabe24's AGB explicitly names a rate — even though vergabe24 itself is excluded, treat this as a signal that similar limits likely apply industry-wide), and add jitter/backoff on repeated failures rather than fixed-interval retry that can synchronize into a thundering-herd against a recovering endpoint.
|
||||
- For DÖE/TED (the auth-free, ToS-friendly central APIs), still poll conservatively (e.g., hourly, not per-minute) — there's no completeness benefit to sub-hourly polling of a feed that itself batches publications, and it needlessly increases the chance of being deprioritized or rate-limited by a public-good API relied on by many consumers.
|
||||
- Verify DÖE's actual pagination/rate-limit behavior against the live Swagger UI before finalizing the poll cadence — the feasibility doc flags this as an open verification point, not yet confirmed.
|
||||
|
||||
**Warning signs:**
|
||||
HTTP 429s or connection resets from a portal correlating with polling frequency increases; DÖE/TED response times degrading specifically during Tessera's poll windows; multiple tenants' saved searches against the same portal triggering independent, uncoordinated scrape runs within the same minute.
|
||||
|
||||
**Phase to address:**
|
||||
Scheduler/Ingestion Orchestration phase — the "poll source once, fan out to tenants" architecture is a foundational design decision for this phase, not an optimization to add later; retrofitting it after per-tenant polling ships means migrating live saved searches without disrupting notifications.
|
||||
|
||||
## Technical Debt Patterns (v1.1)
|
||||
|
||||
| Shortcut | Immediate Benefit | Long-term Cost | When Acceptable |
|
||||
|----------|-------------------|-----------------|------------------|
|
||||
| Hardcode portal HTML selectors per portal instance instead of a shared AI-AG/cosinex adapter interface | Faster first working scraper | Every one of 4/8/9 AI-AG portals needs its own fix when the platform changes markup | Never — the feasibility doc's whole ROI case for scraping rests on one adapter covering many portal instances |
|
||||
| Store only parsed/normalized fields, discard raw scraped HTML | Less storage | No way to diagnose a broken scraper without reproducing live against a moving target | Only if a short-retention raw-response cache is added later before it's actually needed for a live incident |
|
||||
| Per-tenant polling of shared upstream sources (copy DKV's per-tenant cron mental model) | Simpler initial scheduler code | Redundant load, throttling risk, exactly the pattern flagged in Pitfall 26 | Never for shared sources (DÖE/portals); acceptable only for genuinely tenant-specific sources (a tenant's own email inbox) |
|
||||
| Exact-string CPV/title matching instead of hierarchical/fuzzy matching | Simpler filter/dedup code | Filter results miss most relevant tenders (Pitfall 19); dedup produces visible duplicates (Pitfall 21) | Acceptable only as an explicitly-labeled MVP limitation with a visible roadmap item, not silently shipped as "the filter" |
|
||||
| Skip the matched/notified separation, fire notification directly on new DB row insert | Faster to build first notification | Backfill flood on every new saved search (Pitfall 23) | Never — this is cheap to build correctly from the start and expensive to fix after users have already been flooded once |
|
||||
| Use raw `PrismaService` in scheduler/cron code instead of `forTenant()` per tenant iteration | Slightly less boilerplate | Silent cross-tenant data leakage in background jobs (Pitfall 25) | Never for any tenant-scoped table |
|
||||
|
||||
## Integration Gotchas (v1.1)
|
||||
|
||||
| Integration | Common Mistake | Correct Approach |
|
||||
|-------------|-----------------|-------------------|
|
||||
| DÖE OpenData API (eForms/OCDS/CSV) | Treating it as 100% complete tender coverage | Explicitly model coverage tier (Oberschwelle guaranteed, Unterschwelle partial ~20-35%) in schema/UI (Pitfall 20) |
|
||||
| DÖE OCDS (`ocds-mnwr74`) | Reading fields without resolving `OrganizationReference.name` against `parties[]` | Implement the documented post-processing step for all 9 org-reference locations before display |
|
||||
| eForms-DE raw XML | Hardcoding one SDK version's XPath | Read/log SDK version per notice, fail loud on unrecognized versions, keep per-version fixtures |
|
||||
| AI-AG NetServer (lhs-vpbw/tender24/vergabe.landbw) | Per-portal scraper, static session token caching | One shared adapter class, re-derive session/CSRF token every run |
|
||||
| cosinex VMP (DTVP) | Assuming cosinex markup is compatible with AI-AG markup | Separate adapter — the feasibility research confirms the two platforms are HTML-incompatible |
|
||||
| subreport-elvis / service.bund.de | Scraping HTML instead of using their native RSS feeds | Use RSS — lower ToS risk, cheaper to parse, already the recommended approach |
|
||||
| vergabe24 / aumass | Building a "generic" adapter that could technically reach them | Hard denylist enforced in the adapter registry, not just documentation (Pitfall 16) |
|
||||
| DKV inbox infra (`ImapProvider`/`ExchangeInboxProvider`) reused for portal alert emails | Assuming alert email format/parsing is identical to DKV's existing sender filters | Build portal-specific alert-email parsers (subject/body formats differ per portal vendor) reusing only the inbox *transport*, not the DKV parsing logic |
|
||||
| TED API v3 | Polling it as a primary source for DE-only coverage | Treat as optional EU redundancy per the feasibility doc — largely redundant to DÖE for Germany |
|
||||
|
||||
## Performance Traps (v1.1)
|
||||
|
||||
| Trap | Symptoms | Prevention | When It Breaks |
|
||||
|------|----------|------------|-----------------|
|
||||
| Storing full raw HTML/XML indefinitely per scrape run | Database/disk growth outpaces useful data | Short retention window (e.g., last N runs or M days) on raw payloads, keep only normalized rows long-term | Noticeable within weeks at hourly polling across ~10 sources |
|
||||
| Re-parsing full eForms XML on every filter query instead of caching normalized rows | Slow search/filter UI as tender volume grows | Normalize once at ingestion, query only the normalized table for filtering/UI | Becomes visible once historical backlog (months of DÖE data) is loaded |
|
||||
| Per-tenant polling of shared sources (see Pitfall 26) | Upstream request volume scales with tenant count, not data freshness need | Poll-once-fan-out-many architecture | Breaks (throttling) well before thousands of tenants — even a few dozen tenants on the same AI-AG portal is enough |
|
||||
| No archival/expiry of past-deadline tenders in the active result set | Filter/search UI slows as the table grows unbounded | Partition or flag expired tenders out of the default "open" query path; archive rather than delete for dedup/audit history | Gradual, but compounds — plan before the table becomes large enough to require a migration under load |
|
||||
| Fuzzy-match dedup (Pitfall 21) run as O(n²) comparison across the full tender set | Ingestion job runtime grows non-linearly with catalog size | Bucket candidates first (by CPV + region + rough deadline window) before fuzzy title comparison within buckets | Noticeable once total tender volume reaches the thousands (a few months of DÖE + portal data) |
|
||||
|
||||
## Security Mistakes (v1.1)
|
||||
|
||||
| Mistake | Risk | Prevention |
|
||||
|---------|------|------------|
|
||||
| Storing portal login credentials (for saved-search registration) in plaintext or a custom encryption scheme | Credential leak grants access to a tenant's competitive procurement activity on a live portal account | Reuse the existing `encryptedInboxCreds` AES-256-GCM pattern from `DkvModuleConfig` verbatim |
|
||||
| Logging scraper session tokens or full request/response on error (for debugging layout breaks) | Session/CSRF tokens or embedded credentials leak into log aggregation | Redact tokens/credentials before logging; log structural diagnostics (selector not found, row count) instead of raw payloads with secrets |
|
||||
| Background/cron ingestion jobs using raw `PrismaService` instead of tenant-scoped access for tenant tables | Cross-tenant data leakage in saved searches, results, notifications (Pitfall 25) | Enforce `forTenant()` usage per tenant iteration in all scheduler code; add an isolation test as a phase gate |
|
||||
| A future "custom portal URL" feature accepting tenant-supplied URLs for the scraper to fetch | SSRF — a malicious or compromised tenant admin points the scraper at internal infrastructure | If ever built, restrict to an explicit allowlist of known portal domains, never arbitrary tenant-supplied URLs |
|
||||
| Notification emails including full source-portal deep links without validating they're outbound-safe | Low risk here, but email content built from scraped/parsed HTML fields (buyer name, title) without sanitization risks HTML injection in HTML-format alert emails | Sanitize/escape all scraped and eForms-derived text fields before interpolating into HTML email templates |
|
||||
|
||||
## UX Pitfalls (v1.1)
|
||||
|
||||
| Pitfall | User Impact | Better Approach |
|
||||
|---------|-------------|-------------------|
|
||||
| Presenting a unified result list without coverage transparency (Pitfall 20) | Users assume completeness, miss tenders on uncovered portals, lose trust when they find something manually that wasn't surfaced | Show per-search coverage summary (sources queried + tier) alongside results |
|
||||
| Showing already-past-deadline tenders in the default "open" view due to timezone bugs (Pitfall 22) | Users waste time on tenders they can no longer bid on; erodes trust in the "still open" filter specifically | Rigorous UTC-normalized deadline filtering with explicit tests around DST/midnight |
|
||||
| Instant-alert flood on first saved-search activation (Pitfall 23) | Users immediately mute/unsubscribe from a feature that could otherwise be valuable | Default to digest for broad searches; explicit backfill-suppression on activation |
|
||||
| Displaying visible duplicate tender cards from multiple sources (Pitfall 21 failure mode) | Looks unpolished/broken, undermines confidence in the module's data quality | Merge with "also seen on: [sources]" badge instead of separate cards |
|
||||
| No visibility into *why* a tender was filtered out or not matched | Users can't tune their saved search, assume the module is missing things arbitrarily | Provide a "why not matched" explainer or at least document filter semantics (hierarchical CPV, region matching) clearly in the UI |
|
||||
| Raw German procurement jargon (Bekanntmachungsart, Vergabeart, Losaufteilung) with no glossary for non-specialist users | Users unfamiliar with procurement terminology can't interpret results | Add inline tooltips/glossary for domain terms in the detail view |
|
||||
|
||||
## "Looks Done But Isn't" Checklist (v1.1)
|
||||
|
||||
- [ ] **Portal adapter (AI-AG/cosinex):** Often missing a structural health check — verify it distinguishes "zero real results" from "scraper broke and returned zero" (Pitfall 15), and that it fails loud rather than silently.
|
||||
- [ ] **eForms/OCDS ingestion:** Often missing SDK-version handling and `OrganizationReference.name` resolution — verify organization names render correctly and unknown SDK versions are logged, not silently mis-parsed (Pitfalls 17, 18).
|
||||
- [ ] **CPV filtering:** Often only exact-matches CPV codes — verify a filter on a parent category (e.g., "45" — Bauarbeiten) returns child-code matches, not just exact 8-digit hits (Pitfall 19).
|
||||
- [ ] **Dedup:** Often only tested against clean synthetic fixtures — verify against real cross-source pairs (same tender via DÖE + a scraped portal + an ingested alert email) with realistic ID/title/date variance (Pitfall 21).
|
||||
- [ ] **Deadline filtering:** Often only tested at "normal" times — verify "still open" behavior across a DST transition and at a deadline that falls at local midnight (Pitfall 22).
|
||||
- [ ] **Notification pipeline:** Often only tested with a handful of seed records — verify behavior when a saved search is activated against months of existing DÖE backlog (no flood) and when an already-notified tender is updated (no duplicate "new" alert) (Pitfall 23).
|
||||
- [ ] **Scheduler:** Often copied from the DKV single-tenant `findFirst()` pattern — verify two tenants with independent active configs both actually poll and receive results (Pitfall 24).
|
||||
- [ ] **Multi-tenant isolation:** Often only tested for the "happy path" single-tenant flow during development — verify with two tenants and overlapping search criteria that no cross-tenant data (saved searches, results, notifications, credentials) leaks (Pitfall 25).
|
||||
- [ ] **Rate limiting:** Often untested until a portal actually throttles in production — verify the poll-once-fan-out-many architecture is actually in place before scaling tenant count, not just planned (Pitfall 26).
|
||||
- [ ] **Excluded sources:** Often only documented, not enforced — verify the adapter registry structurally refuses to register vergabe24/aumass adapters, not just that no one has written one yet (Pitfall 16).
|
||||
|
||||
## Recovery Strategies (v1.1)
|
||||
|
||||
| Pitfall | Recovery Cost | Recovery Steps |
|
||||
|---------|---------------|-----------------|
|
||||
| Scraper broken by portal layout change (Pitfall 15) | MEDIUM | Diagnose via cached raw HTML from last successful runs; patch selectors; add the new structural variant to the adapter's health-check assertions so future drift of the same kind is caught faster |
|
||||
| Accidental scraping attempt against a denylisted portal (Pitfall 16) | HIGH | Immediately disable the adapter, audit logs for request volume/duration against that portal, assess legal exposure, do not silently "fix" and continue — this needs explicit sign-off before any retry |
|
||||
| eForms SDK version broke parsing (Pitfall 17) | LOW-MEDIUM | Add the new SDK version's field mapping/fixture, backfill-reparse affected date range from cached raw XML if retained, otherwise accept the gap and document it |
|
||||
| Cross-tenant data leak discovered in background job (Pitfall 25) | HIGH | Treat as a security incident: identify affected tenants/records, patch the missing `forTenant()` scoping, audit all other scheduler code paths for the same pattern, notify affected tenants per data-protection obligations |
|
||||
| Notification flood already sent to tenants (Pitfall 23) | MEDIUM | Send a brief clarifying follow-up (not another flood), fix the backfill-suppression logic, offer an easy re-subscribe/digest-preference change for anyone who unsubscribed in reaction |
|
||||
| Portal starts throttling/blocking Tessera's scraper IP (Pitfall 26) | MEDIUM | Back off immediately (pause the adapter), fall back to that portal's native email-alert ingestion path if available (per the feasibility doc, most non-excluded portals do offer this), re-evaluate poll cadence before resuming |
|
||||
|
||||
## Pitfall-to-Phase Mapping (v1.1)
|
||||
|
||||
| Pitfall | Prevention Phase | Verification |
|
||||
|---------|-------------------|----------------|
|
||||
| 15. Scraper fragility (session/CSRF/layout drift) | Portal Adapter phase (AI-AG + cosinex) | Adapter has a structural health check; simulate a selector failure in tests and confirm it fails loud, not silent-zero |
|
||||
| 16. Legal/ToS exclusion (vergabe24, aumass) | Portal Adapter phase | Adapter registry test asserts registration attempt for denylisted portal IDs is rejected |
|
||||
| 17. eForms SDK version drift | DÖE/OCDS Ingestion phase | Parser logs/handles unknown SDK version explicitly; fixture tests cover ≥2 real SDK versions |
|
||||
| 18. OCDS optional-field / OrganizationReference handling | DÖE/OCDS Ingestion + Normalization phase | Org names resolve correctly in UI for buyer/tenderer/supplier; "not applicable yet" vs "failed to parse" are distinguishable in schema |
|
||||
| 19. CPV format/hierarchy | Normalization + Filter Engine phase | Filter on a parent CPV category returns child-code matches in a test fixture |
|
||||
| 20. Below/above-threshold false completeness | Normalization + Filter/UI phase | UI displays coverage-tier annotation on every result set; documented and testable against the known ~20-35% Unterschwelle figure |
|
||||
| 21. Cross-source dedup | Normalization & Dedup phase (own phase) | Fuzzy fingerprint dedup test against real cross-source sample pairs (DÖE + scraped portal + alert email for the same tender) |
|
||||
| 22. Deadline/timezone handling | Normalization + Filter Engine phase | Explicit DST-transition and local-midnight test cases pass for "still open" filtering |
|
||||
| 23. Notification storms / backfill flooding | Notification phase | Activation of a saved search against historical backlog produces zero immediate individual alerts; update-vs-new distinction tested |
|
||||
| 24. Multi-tenant scheduler (DKV single-tenant regression) | Scheduler/Ingestion Orchestration phase | Two-tenant test: both tenants' active configs poll and produce independent results |
|
||||
| 25. Multi-tenant isolation (saved searches, results, credentials) | Multi-Tenant Saved Searches phase | Two-tenant cross-isolation UAT: no leakage of config, results, or notifications across tenants |
|
||||
| 26. Scheduler rate-limiting / aggressive polling | Scheduler/Ingestion Orchestration phase | Poll-once-fan-out-many architecture verified (single upstream request serves all matching tenant searches); backoff/jitter tested against simulated 429/Retry-After |
|
||||
|
||||
## Sources (v1.1 Ausschreibungs-Radar)
|
||||
|
||||
- `.planning/research/ausschreibungs-portale-feasibility.md` — portal-by-portal ToS/anti-bot classification, coverage percentages, DÖE/TED architecture, adapter consolidation recommendation (2026-07-16 research)
|
||||
- [OCDS for eForms — How to use this profile](https://standard.open-contracting.org/profiles/eforms/latest/en/how/) — OrganizationReference resolution requirement, withheld-information handling, framework-agreement cascade caveat
|
||||
- [OCDS for eForms — Field mappings / Schema / Codelists](https://standard.open-contracting.org/profiles/eforms/latest/en/) — official field mapping reference
|
||||
- [open-contracting-extensions/eforms (GitHub)](https://github.com/open-contracting-extensions/eforms) — profile source, versioning
|
||||
- [Beschaffungsamt — Datenservice Öffentlicher Einkauf](https://www.bescha.bund.de/DE/ElektronischerEinkauf/Datenservice_Oeffentlicher_Einkauf/Datenservice-Oeffentlicher-Einkauf_node.html) — DÖE service components, eForms-DE/OCDS/CSV export formats
|
||||
- [DÖE OpenData Swagger UI](https://oeffentlichevergabe.de/documentation/swagger-ui/opendata/index.html) — API surface (pagination/rate-limit specifics not yet live-verified — open item per feasibility doc)
|
||||
- [BGH, 30.04.2014 — I ZR 224/12 (screen scraping)](https://dejure.org/dienste/vernetzung/rechtsprechung?Gericht=BGH&Datum=30.04.2014&Aktenzeichen=I+ZR+224/12) — German case law on scraping/AGB/"virtuelles Hausrecht"
|
||||
- [VOELKER & Partner — Screen-Scraping / Web-Crawler rechtliche Zulässigkeit](https://www.voelker-gruppe.com/kompetenzen/ip-it-stuttgart/beitraege/screen-scraping-web-crawler) — AGB incorporation requirements, § 3a UWG framing
|
||||
- Existing Tessera codebase: `apps/api/src/dkv/dkv-scheduler.service.ts` (single-tenant scheduler precedent to avoid repeating), `apps/api/src/tenant/tenant.guard.ts` (tenant-scoping pattern via `forTenant()`), `apps/api/prisma/schema.prisma` `DkvModuleConfig` (credential encryption precedent to reuse)
|
||||
|
||||
---
|
||||
*Pitfalls research for: Tessera (v1.0 platform) + Ausschreibungs-Radar (v1.1 multi-source German tender aggregation module)*
|
||||
*Researched: 2026-06-18 (v1.0); 2026-07-17 (v1.1)*
|
||||
|
||||
Reference in New Issue
Block a user