docs: complete project research

Ausschreibungs-Radar (v1.1) STACK/FEATURES/ARCHITECTURE/PITFALLS research plus SUMMARY.md synthesis.
This commit is contained in:
2026-07-17 10:12:59 +02:00
parent ae841205df
commit d2ac1997d6
6 changed files with 1215 additions and 470 deletions
+409 -271
View File
@@ -1,320 +1,458 @@
# Architecture Patterns
# Architecture Research — Ausschreibungs-Radar Module Integration
**Domain:** Modular portal platform with marketplace, multi-tenancy, configurable dashboard, and desktop wrapper
**Researched:** 2026-06-18
**Domain:** Multi-source tender/procurement-notice ingestion module for an existing NestJS 11 modular-monolith platform (Tessera)
**Researched:** 2026-07-17
**Confidence:** HIGH (module registration, tenant scoping, inbox/mail reuse — verified against existing `dkv/` and `module-registry/` source) / MEDIUM (OCDS field mapping, DÖE API pagination — verified via OCDS spec + DÖE Swagger listing, not a live API call)
## Recommended Architecture
> **Note:** This supersedes the v1.0 platform-level `ARCHITECTURE.md` (2026-06-18, described a generic Fastify/Traefik shell that predates the stack decisions now recorded in `CLAUDE.md`). This file is scoped to the v1.1 Ausschreibungs-Radar milestone and describes how the new module integrates into the actual, current NestJS 11 + Next.js codebase — not a from-scratch platform design.
Tessera follows a **modular monolith** pattern: a single deployable backend that internally separates concerns into distinct modules, fronted by a micro-frontend-capable shell. This avoids premature microservice complexity while maintaining clean module boundaries for future extraction.
## Summary
The Ausschreibungs-Radar module ("tenders" module) is a **new, self-contained NestJS module** built strictly on top of existing Tessera infrastructure — module-registry self-seeding, `@nestjs/schedule` dynamic cron jobs, the `ImapProvider`/`ExchangeInboxProvider` inbox abstraction, tenant-scoped SMTP via `SmtpConfig` + fresh-transport-per-send, and the module-loader-driven Next.js portal route. The one architectural decision that **breaks from the DKV template** is deliberate and important: **tender data is platform-global, not tenant-owned.** DKV invoices belong to one tenant; a DÖE/AI-NetServer/cosinex tender notice is a public fact relevant to *every* tenant. Ingestion, normalization, and dedup therefore run once for the whole platform; only **saved searches, matches, and notification preferences** are tenant-scoped. Getting this split right is the single highest-leverage decision in this design — getting it wrong means N-times redundant scraping/storage and N-times the anti-bot exposure per portal.
## Standard Architecture
### System Overview
```
+-------------------------------------------------------------------+
| Desktop Wrapper (Tauri) |
+-------------------------------------------------------------------+
| Frontend Shell (React SPA) |
| +------------+ +------------+ +------------+ +-------------+ |
| | Sidebar | | Dashboard | | Marketplace| | Module UI | |
| | Navigation | | (Widgets) | | Browser | | (lazy-load) | |
| +------------+ +------------+ +------------+ +-------------+ |
+-------------------------------------------------------------------+
| API Gateway (Traefik / Nginx) |
+-------------------------------------------------------------------+
| Backend (Node.js / Fastify) |
| +--------+ +--------+ +--------+ +---------+ +----------+ |
| | Auth | | Tenant | | Module | | Market- | | Dashboard| |
| | Module | | Module | | Loader | | place | | Service | |
| +--------+ +--------+ +--------+ +---------+ +----------+ |
+-------------------------------------------------------------------+
| PostgreSQL (shared, RLS) |
+-------------------------------------------------------------------+
| Docker Compose (orchestration) |
+-------------------------------------------------------------------+
+---------------------------------------------------------------------------+
| SOURCE ADAPTERS (new) |
| +-----------+ +--------------+ +-----------+ +--------+ +--------------+|
| |DoeOpenData| |AiNetServer | |Cosinex | |Rss | |EmailAlert ||
| |Adapter | |Adapter | |Adapter | |Adapter | |Adapter ||
| |(API,auth- | |(HTML scrape, | |(HTML | |(feed | |(reuses Imap/ ||
| | free) | | public | | scrape, | | parse) | | Exchange ||
| | | | search) | | public) | | | | InboxProvider||
| +-----+-----+ +------+-------+ +-----+-----+ +---+----+ +------+-------+|
| | all implement TenderSourceAdapter -> RawTenderRecord[] | |
+--------+--------------+---------------+-----------+-----------------+---+
`--------------`-------+-------`-----------`-------------'
v
+---------------------------------------------------------------------------+
| TenderIngestionService (new, global -- no tenantId) |
| normalize(raw, sourceType) -> Tender (OCDS-oriented) |
| computeDedupKey() -> upsert by dedupKey -> diff contentHash -> mark changed|
+-------------------------------+-------------------------------------------+
v (only NEW / CHANGED tenders this poll)
+---------------------------------------------------------------------------+
| TenderMatchingService (new) -- DB-query filter evaluation |
| for each active TenderSavedSearch: Prisma `where` over the delta batch |
| (structured filters) + Postgres full-text search (keywords) -> TenderMatch|
+-------------------------------+-------------------------------------------+
v
+---------------------------------------------------------------------------+
| TenderMailService (new) -- reuses SmtpConfig + fresh-transport |
| instant: send on TenderMatch create digest: cron batch per SavedSearch |
+---------------------------------------------------------------------------+
|
+------------------------+---------------------------------------------------+
| TendersController (new) -- GET /tenders (search+filter+paginate), |
| /tenders/:id, /tenders/saved-searches (CRUD, tenant-scoped), |
| /tenders/source-config (admin, global), /tenders/check-now |
+------------------------+---------------------------------------------------+
v
apps/web/.../modules/tender-radar/ (new Next.js module UI)
```
### Component Boundaries
### Component Responsibilities
| Component | Responsibility | Communicates With |
|-----------|---------------|-------------------|
| **Frontend Shell** | Application frame (header, sidebar, routing), theme, i18n | Backend API via REST/WebSocket |
| **Dashboard Engine** | Widget grid, drag-and-drop, layout persistence | Backend Dashboard Service for saving layouts |
| **Marketplace UI** | Browse modules, view details, request activation | Backend Marketplace Service |
| **Module UI Slots** | Lazy-loaded UI for activated modules | Module-specific backend endpoints |
| **API Gateway** | Reverse proxy, rate limiting, tenant header injection | All backend services |
| **Auth Module** | Login, session/JWT, LDAP integration, user management | PostgreSQL, LDAP server |
| **Tenant Module** | Tenant CRUD, tenant context resolution, tenant-specific config | PostgreSQL, injected into every request |
| **Module Loader** | Plugin lifecycle (discover, validate, activate, deactivate) | Filesystem/registry, PostgreSQL |
| **Marketplace Service** | Module catalog, licensing, activation per tenant | PostgreSQL, Module Loader |
| **Dashboard Service** | Widget registry, layout CRUD per user per tenant | PostgreSQL |
| **PostgreSQL** | Persistent storage, row-level security for tenant isolation | All backend modules |
| **Desktop Wrapper (Tauri)** | Native window, system tray, local shortcuts | Frontend Shell (wraps the web app) |
| Component | Responsibility | Scope | New/Modified |
|-----------|----------------|-------|---------------|
| `TenderSourceAdapter` implementations | Fetch raw records from one source, no normalization | Global | New |
| `TenderNormalizerService` | Map each source's raw shape → unified `Tender` fields, compute `dedupKey`/`contentHash` | Global | New |
| `TenderIngestionService` | Orchestrate poll → normalize → upsert → change-detect per source | Global | New |
| `TenderSchedulerService` | Dynamic cron per global source + per-tenant cron for email-alert ingestion | Mixed | New |
| `TenderMatchingService` | Evaluate active saved searches against the new/changed delta | Per-tenant read, global data | New |
| `TenderMailService` | Send instant/digest notification emails via tenant SMTP | Per-tenant | New |
| `TendersController` | REST endpoints for list/detail/saved-search CRUD/admin source-config | Mixed | New |
| Shared `InboxModule` (relocated) | `InboxProvider` interface + `ImapProvider`/`ExchangeInboxProvider` | Platform-shared | New (extracted from `dkv/`) |
| `ModuleRegistryService` | Self-seed `tender-radar` module row | Platform | Reused unmodified |
| `SettingsService` / `SmtpConfig` | Decrypted per-tenant SMTP for notification sends | Per-tenant | Reused unmodified |
| `CalendarCryptoService` | AES-256-GCM encryption for any stored credentials | Platform | Reused unmodified |
| `module-loader.ts` `MODULE_REGISTRY` | Whitelist entry mapping slug → lazy component | Platform | Modified (one entry) |
| `AppModule` | Import `TendersModule` | Platform | Modified (one import) |
### Data Flow
**Request flow (authenticated):**
## Recommended Project Structure
```
User Action
-> Desktop Wrapper / Browser
-> Frontend Shell (React Router)
-> HTTP Request with JWT + Tenant-ID header
-> API Gateway (validates JWT, injects tenant context)
-> Backend Route Handler
-> Service Layer (business logic)
-> PostgreSQL (RLS enforces tenant isolation)
<- Response
<- JSON Response
<- Frontend renders
apps/api/src/
├── inbox/ # NEW -- extracted shared module (was dkv/providers/)
│ ├── inbox.module.ts # exports ImapProvider, ExchangeInboxProvider
│ ├── inbox-provider.interface.ts # InboxProvider contract (moved from dkv.types.ts)
│ ├── imap.provider.ts # moved verbatim from dkv/providers/
│ ├── exchange-inbox.provider.ts # moved verbatim from dkv/providers/
│ └── inbox.types.ts # InboxConfig / InboxEmail / InboxAttachment
│
├── dkv/ # MODIFIED -- imports InboxModule instead of local providers/
│ └── ... # (providers/ folder removed, dkv.types.ts trimmed)
│
├── tenders/ # NEW -- Ausschreibungs-Radar module
│ ├── tenders.module.ts
│ ├── tenders.controller.ts # public list/detail + tenant saved-search CRUD + admin source-config
│ ├── tenders.seed.ts # seeds 'tender-radar' into ModuleRegistry
│ ├── tender-ingestion.service.ts # poll → normalize → upsert → change-detect (per source)
│ ├── tender-matching.service.ts # SavedSearch → Prisma where-clause → TenderMatch
│ ├── tender-mail.service.ts # instant + digest notification sends
│ ├── tender-scheduler.service.ts # SchedulerRegistry cron: 1 per global source + 1 per tenant (email-alert)
│ ├── tender.types.ts # RawTenderRecord, NormalizedTenderFields, SourceType
│ ├── dto/
│ │ ├── saved-search.dto.ts
│ │ ├── source-config.dto.ts
│ │ └── tender-query.dto.ts # pagination + filter query params for GET /tenders
│ └── adapters/
│ ├── tender-source-adapter.interface.ts # fetchTenders(config, since) → RawTenderRecord[]
│ ├── doe-opendata.adapter.ts # Build order Phase A
│ ├── ai-netserver.adapter.ts # Build order Phase B
│ ├── cosinex.adapter.ts # Build order Phase B
│ ├── rss.adapter.ts # Build order Phase C
│ └── email-alert.adapter.ts # Build order Phase C (uses InboxModule)
│
apps/web/src/app/(portal)/modules/
├── tender-radar/ # follow the existing static per-module convention (dkv-fleet, cert-manager)
│ ├── page.tsx # searchable trefferliste + filter sidebar
│ ├── [id]/page.tsx # tender detail view
│ ├── saved-searches/page.tsx # saved search CRUD UI
│ └── settings/page.tsx # admin: source poll config, email-alert inbox config
```
**Module activation flow:**
### Structure Rationale
```
Admin browses Marketplace
-> Selects module, clicks "Activate"
-> POST /api/marketplace/modules/:id/activate
-> Marketplace Service checks license entitlement
-> Module Loader registers module for tenant
-> INSERT module_activations (tenant_id, module_id, status)
-> Frontend sidebar updates (module appears)
<- Success response
```
- **`inbox/` extraction is a prerequisite, not optional.** `ImapProvider`/`ExchangeInboxProvider` are already generic over `InboxConfig`/`InboxEmail` — nothing in them is DKV-specific. Today they live in `dkv/providers/` and their types live in `dkv.types.ts`, so `tenders/` would otherwise have to import from inside another feature module's internals (`../dkv/providers/imap.provider`), which couples two unrelated features and breaks if DKV is ever restructured. Moving them to a shared `inbox/` module once, and updating `dkv.module.ts` to import `InboxModule` instead, costs one small refactor now and pays for every future module that needs inbox polling (already two: DKV, Tenders).
- **`tenders/adapters/` mirrors `dkv/providers/`** — same rationale as DKV: consumers (`TenderIngestionService`) depend only on the `TenderSourceAdapter` interface, never on a concrete adapter. This is what makes "ship DÖE first, add AI-NetServer/cosinex/RSS/email later" possible without touching the ingestion/matching/notification pipeline.
- **Normalizer is separate from adapters**, unlike DKV where `DkvParserService` is a single PDF parser. Here there are 5 structurally incompatible raw shapes (eForms/OCDS JSON, two flavors of scraped HTML, RSS/Atom XML, free-text alert emails). Each adapter can either normalize inline or delegate to a per-source mapping function inside `TenderNormalizerService` — either way, the *interface boundary* is `RawTenderRecord[] → Tender[]`, so the ingestion orchestrator never branches on source type.
- **Web module UI follows the existing static `modules/<slug>/` folder pattern** seen in `dkv-fleet/` and `cert-manager/` (not the dynamic `[category]/[moduleSlug]/` route also present in the codebase for vehicle/settings sub-pages) — match whichever of the two conventions the team is actively converging on at execution time; both are already present, so this is a phase-planning decision, not an open architectural question.
**Dashboard widget flow:**
## Architectural Patterns
```
User opens Dashboard
-> GET /api/dashboard/layout
-> Returns user's widget layout (positions, sizes)
-> Frontend renders react-grid-layout with widget components
-> User drags/resizes widget
-> PUT /api/dashboard/layout (debounced save)
-> Persists to PostgreSQL
```
### Pattern 1: Source-Adapter Abstraction (`TenderSourceAdapter`)
## Core Architecture Decisions
### 1. Modular Monolith over Microservices
**Why:** Tessera is built by a single developer (with Claude). Microservices add deployment, debugging, and network complexity that provides zero benefit at this scale. A modular monolith gives clean separation with a single deployment unit.
**Structure:** Each domain (auth, tenant, marketplace, dashboard, modules) lives in its own directory with its own routes, services, and repository files. They communicate through in-process function calls, not HTTP.
**Future path:** If a module becomes a bottleneck, extract it to a separate service behind the API gateway. The clean boundaries make this straightforward.
### 2. Shared Database with Row-Level Security (RLS)
**Why:** Separate databases per tenant adds massive operational overhead. PostgreSQL RLS enforces tenant isolation at the database level, meaning even application bugs cannot leak data across tenants.
**Implementation:**
```sql
-- Every tenant-scoped table has a tenant_id column
ALTER TABLE modules ENABLE ROW LEVEL SECURITY;
CREATE POLICY tenant_isolation ON modules
USING (tenant_id = current_setting('app.current_tenant')::uuid);
```
The backend sets `app.current_tenant` on each database connection based on the authenticated user's tenant. RLS handles the rest transparently.
### 3. Plugin/Module System as Data-Driven Registry
**Why:** Modules should not require restarting the server to be discovered. A database-driven registry with filesystem-based module code gives hot-activation without runtime code loading risks.
**How it works:**
- Module metadata (name, version, category, routes, permissions) stored in `modules` table
- Module code lives in `src/modules/<module-name>/` with a standard interface
- Activation is per-tenant: `module_activations` table links tenant to module
- Frontend lazy-loads module UI bundles only when the module is active for the current tenant
- Backend routes for a module are only registered/accessible when the module is active
### 4. Frontend Shell with Lazy Module Loading
**Why:** Loading all module UIs upfront wastes bandwidth and exposes code for modules the tenant has not licensed. Lazy loading (React.lazy + dynamic import) loads module UIs on demand.
**Pattern:**
```typescript
// Module registry maps module_id -> lazy component
const moduleRegistry: Record<string, () => Promise<{ default: ComponentType }>> = {
'domaincheck': () => import('./modules/domaincheck/DomaincheckPage'),
'email-tools': () => import('./modules/email-tools/EmailToolsPage'),
};
```
The shell only renders module routes that are in the tenant's active module list (fetched from the backend on login).
### 5. Tauri over Electron for Desktop Wrapper
**Why:** Tauri produces ~10MB binaries vs Electron's 100MB+. RAM usage is 20-40MB vs 200-400MB. Tauri uses the system WebView (no bundled Chromium), has a Rust backend for native features, and has a smaller attack surface. The non-programmer maintainer benefits from the simpler, lighter deployment.
**Architecture:** The Tauri wrapper is a thin shell. It loads the same web application served locally or from the server. Native features (system tray, auto-update, window management) are exposed through Tauri commands.
## Patterns to Follow
### Pattern 1: Tenant Context Middleware
**What:** A middleware that extracts tenant identity from the JWT/session, sets it on the request context, and configures the database connection with RLS.
**When:** Every authenticated request.
**What:** One interface, N implementations — directly analogous to `InboxProvider` (`fetchPdfAttachments` → `fetchTenders`).
```typescript
// middleware/tenantContext.ts
async function tenantContext(req: FastifyRequest, reply: FastifyReply) {
const tenantId = req.user.tenantId;
if (!tenantId) return reply.code(403).send({ error: 'No tenant context' });
// tenders/adapters/tender-source-adapter.interface.ts
export interface RawTenderRecord {
sourceType: SourceType; // 'doe-opendata' | 'ai-netserver' | 'cosinex' | 'rss' | 'email-alert'
sourcePortal: string; // 'doe' | 'lhs-vpbw' | 'tender24' | 'vergabe.landbw' | 'dtvp' | 'subreport-elvis' | 'service.bund.de'
sourceRawId: string; // portal-native id/notice number, pre-normalization
sourceUrl: string;
fetchedAt: Date;
payload: unknown; // raw JSON/HTML-extract/RSS-item/email-body — kept for rawPayload + reprocessing
}
// Set RLS context on the database connection
await req.db.query(`SET app.current_tenant = '${tenantId}'`);
req.tenantId = tenantId;
export interface TenderSourceAdapter {
readonly sourceType: SourceType;
/** since: only fetch records new/changed after this timestamp (cursor from TenderSourcePollConfig.lastPolledAt) */
fetchTenders(config: TenderSourceConfig, since?: Date): Promise<RawTenderRecord[]>;
testConnection?(config: TenderSourceConfig): Promise<{ success: boolean; message?: string }>;
}
```
### Pattern 2: Module Interface Contract
**When to use:** Any time a pipeline must ingest structurally different sources into one output shape without the orchestrator knowing about each source. Same pattern Tessera already uses for `ImapProvider`/`ExchangeInboxProvider`.
**What:** Every module (backend) exports a standard interface so the platform can discover, mount, and manage it uniformly.
**Trade-offs:** Each adapter owns its own retry/rate-limit/anti-bot logic (AI-NetServer and cosinex adapters need polite scraping delays; DÖE/RSS don't). The interface intentionally does *not* prescribe HTTP client or scraping library — `TenderSourceConfig` is adapter-specific (a discriminated union or `Json` blob per `sourceType`), same as `InboxConfig` covers both IMAP and Exchange with one shape only because both providers happen to share fields; here they mostly won't, so `TenderSourceConfig` should be `Json` on the Prisma side with adapter-specific Zod/DTO validation, not a single flat interface.
**When:** Building any new module.
### Pattern 2: Normalized OCDS-Oriented Schema + Cross-Source Dedup Key
**What:** One `Tender` Prisma model absorbing eForms/OCDS (DÖE), scraped HTML (AI-NetServer, cosinex), RSS, and email-alert text — with a **stable dedup key** so the same real-world procurement notice appearing on multiple sources (e.g. an AI-NetServer notice that later also appears on DÖE once it crosses the EU threshold, or the same DÖE OCID reappearing on a poll re-run) collapses to one row.
```prisma
model Tender {
id String @id @default(uuid())
// OCDS-oriented core (see standard.open-contracting.org/latest/en/schema/reference/)
ocid String? // Open Contracting ID, e.g. "ocds-mnwr74-XXXXXXXX" — present when sourced via DÖE/TED
noticeId String? // portal-native notice/procedure number (AI-NetServer, cosinex, RSS items)
title String
description String? @db.Text
buyerName String?
buyerId String? // e.g. Vergabestelle-ID if the source exposes it
procedureType String? // OCDS tender.procurementMethod / procurementMethodDetails
status String @default("active") // 'active' | 'awarded' | 'cancelled' | 'expired'
cpvCodes String[] @default([])
region String?
plz String?
bundesland String?
estimatedValue Decimal? @db.Decimal(14, 2)
currency String? @default("EUR")
publishedAt DateTime?
deadlineAt DateTime?
// Source + dedup
sourceType String // 'doe-opendata' | 'ai-netserver' | 'cosinex' | 'rss' | 'email-alert'
sourcePortal String // 'doe' | 'lhs-vpbw' | 'dtvp' | 'subreport-elvis' | ...
sourceUrl String?
dedupKey String @unique // see dedup strategy below
contentHash String // hash of normalized fields — detects "changed" vs "identical re-poll"
rawPayload Json // original adapter payload — debugging + future re-normalization
firstSeenAt DateTime @default(now())
lastSeenAt DateTime @updatedAt
createdAt DateTime @default(now())
matches TenderMatch[]
@@index([sourceType])
@@index([deadlineAt])
@@index([publishedAt])
@@index([bundesland])
}
```
**Dedup key strategy (priority order, computed by `TenderNormalizerService`):**
1. `ocid` — when the source provides an OCDS Open Contracting ID (always true for DÖE/TED; the platform's registered OCDS prefix is `ocds-mnwr74`). OCID is designed exactly for this — joining the same contracting process across publishers.
2. `${sourcePortal}:${noticeId}` — when the portal exposes a stable native notice/procedure number (AI-NetServer Bietercockpit ID, cosinex Vergabenummer, RSS item guid, email-alert reference number). This is the *primary* key for scraped/RSS/email sources, since they never carry an OCID.
3. `sha256(sourcePortal + normalizedTitle + buyerName + deadlineAt)` — last-resort fallback only when a source gives neither an OCID nor a stable ID (should be rare; flag these rows for manual review via a `dedupConfidence: 'low'` marker if this path is hit).
`dedupKey` is the `@unique` upsert target: `prisma.tender.upsert({ where: { dedupKey }, ... })`. `contentHash` (hash of title+deadline+value+status) is separate from `dedupKey` — it answers "did anything about this same notice change since we last saw it" (deadline extension, cancellation), which is what should trigger re-matching and potentially a "notice updated" notification, whereas an unchanged re-poll should just bump `lastSeenAt` and stop.
**Trade-offs:** A single wide table is simpler to query/filter/index than per-source tables + a union view, and matches how the UI wants to browse ("one trefferliste across all sources"). The cost is that source-specific fields that don't map cleanly (e.g. cosinex-specific metadata) live only in `rawPayload` (Json, unindexed) — acceptable, since the feasibility research shows the cross-source overlap (title, buyer, deadline, value, CPV, region) covers what filtering/notification actually need.
### Pattern 3: Global Data, Per-Tenant Filtering (the key deviation from the DKV template)
**What:** `Tender` rows carry **no `tenantId`** — they are platform-wide. Per-tenant scoping happens one layer up, in `TenderSavedSearch` and `TenderMatch`.
```prisma
model TenderSavedSearch {
id String @id @default(uuid())
tenantId String
userId String? // null = tenant-wide search, set = personal search
name String
keywords String[] @default([]) // full-text match against title+description
bundeslaender String[] @default([])
plzPrefixes String[] @default([])
cpvCodes String[] @default([])
minValue Decimal? @db.Decimal(14, 2)
maxValue Decimal? @db.Decimal(14, 2)
deadlineWithinDays Int?
notifyMode String @default("digest") // 'none' | 'digest' | 'instant'
digestHour Int? @default(7) // for digest mode: hour-of-day to send
isActive Boolean @default(true)
createdAt DateTime @default(now())
updatedAt DateTime @updatedAt
matches TenderMatch[]
@@index([tenantId])
}
model TenderMatch {
id String @id @default(uuid())
tenderId String
tender Tender @relation(fields: [tenderId], references: [id], onDelete: Cascade)
savedSearchId String
savedSearch TenderSavedSearch @relation(fields: [savedSearchId], references: [id], onDelete: Cascade)
tenantId String // denormalized for fast tenant-scoped queries/RLS
matchedAt DateTime @default(now())
notifiedAt DateTime? // null = not yet sent (instant) or not yet in a digest
@@unique([tenderId, savedSearchId])
@@index([tenantId])
@@index([notifiedAt])
}
```
**Why this beats a `tenantId` on `Tender`:** DÖE alone publishes thousands of Oberschwelle notices; duplicating that table N times (once per tenant) multiplies storage for zero benefit — every tenant sees the same underlying notice, just filtered differently. It also means the DÖE/AI-NetServer/cosinex/RSS pollers run **once for the whole platform**, not once per active tenant — critical for the scraping sources, where running the same scrape N times per tenant multiplies anti-bot/ToS exposure on portals that already sit at "Niedrig-Mittel" risk per the feasibility research. `TenantModuleActivation` still gates whether a tenant sees the module at all (standard Tessera marketplace pattern) — but activation controls *visibility*, not a second data copy.
**When this pattern does NOT apply:** email-alert ingestion. A tenant's alert emails arrive in *that tenant's own mailbox* (their own registered "gespeicherte Suche" on a portal) — so `TenderInboxConfig` (credentials, mirroring `DkvModuleConfig.encryptedInboxCreds`) is legitimately per-tenant, even though the `Tender` rows it produces still land in the same global table (deduped against whatever DÖE/AI-NetServer/cosinex already ingested for the same notice).
### Pattern 4: Filter Evaluation — DB Query Against the Delta, Not In-Memory Full-Table Scan
**What:** Filtering happens as a **Postgres query scoped to the just-ingested batch**, not (a) a full in-memory scan of all tenders per saved search, nor (b) a full re-scan of the entire `Tender` table on every poll.
```typescript
// modules/<name>/index.ts
export interface TesseraModule {
id: string;
version: string;
category: string;
routes: (app: FastifyInstance) => void;
widgets?: WidgetDefinition[]; // Optional dashboard widgets
permissions?: string[]; // Required permissions
onActivate?: (tenantId: string) => Promise<void>;
onDeactivate?: (tenantId: string) => Promise<void>;
// tender-matching.service.ts (sketch)
async matchDelta(newOrChangedTenderIds: string[]): Promise<void> {
const savedSearches = await this.prisma.tenderSavedSearch.findMany({ where: { isActive: true } });
for (const search of savedSearches) {
const where = this._buildWhereClause(search, newOrChangedTenderIds); // structured filters
const matches = await this.prisma.tender.findMany({ where }); // DB does the heavy lifting
// full-text keyword refinement, if keywords present, folded into the same query via
// a raw `to_tsvector('german', title || ' ' || description) @@ plainto_tsquery(...)`
for (const tender of matches) {
await this.prisma.tenderMatch.upsert({
where: { tenderId_savedSearchId: { tenderId: tender.id, savedSearchId: search.id } },
create: { tenderId: tender.id, savedSearchId: search.id, tenantId: search.tenantId },
update: {}, // matchedAt stays as first-seen; re-match is idempotent
});
}
}
}
```
### Pattern 3: Widget Component Contract
**Why DB, not in-memory:** structured filters (CPV array overlap, numeric value range, date range, region) are exactly what Postgres indexes/arrays/range queries are built for, and the corpus (all German public tenders) will reach tens of thousands of rows within the first year — loading that into Node to filter per saved search per tenant does not scale and duplicates work Postgres already does better. Keyword matching specifically should use a **Postgres full-text `tsvector` GIN index** on `title || description` (German text-search config) rather than `ILIKE '%term%'` scans — `ILIKE` on a growing table degrades linearly, GIN full-text does not.
**What:** Dashboard widgets follow a standard interface for the grid layout system.
**Why "scoped to the delta" and not the whole table every poll:** re-evaluating every saved search against the *entire* `Tender` table on every poll cycle is O(searches × table-size) repeated hourly — wasteful and re-creates matches that already exist (idempotent upsert hides the waste but not the cost). Instead, `TenderIngestionService` passes the list of tender IDs that were newly created or had a `contentHash` change in *this* poll run to `TenderMatchingService.matchDelta()`, so the query is `WHERE id IN (delta) AND <search filters>` — bounded by poll batch size (dozens to low hundreds), not table size.
**When:** Building any dashboard widget (core or module-provided).
**Trade-off:** if a saved search is created or edited *after* a tender was ingested, that tender won't retroactively appear in the search's matches until the next time it's re-touched (re-poll sees no change → no delta → not re-evaluated). Mitigate with an explicit "backfill" action: when a `TenderSavedSearch` is created/edited, run `matchDelta()` once against *all* tenders from the last N days (bounded, on-demand, not a recurring cost) rather than the full historical table.
```typescript
// widgets/types.ts
export interface WidgetDefinition {
id: string;
name: string; // i18n key
defaultSize: { w: number; h: number };
minSize?: { w: number; h: number };
component: () => Promise<{ default: ComponentType<WidgetProps> }>;
configSchema?: JSONSchema; // Optional widget settings
}
### Pattern 5: Scheduled Polling — Global Cron Per Source + Per-Tenant Cron Only for Email-Alerts
export interface WidgetProps {
config: Record<string, unknown>;
tenantId: string;
userId: string;
}
```
**What:** Extends `DkvSchedulerService`'s `SchedulerRegistry.addCronJob()` pattern, but with two distinct job populations:
### Pattern 4: Feature Flag via Module Activation
- **Global source jobs** (DÖE, RSS feeds, AI-NetServer, cosinex): one cron job per row in `TenderSourcePollConfig` (admin-managed, not tenant-scoped), named `tender-poll-${sourceConfigId}`. Interval is source-appropriate — DÖE can poll frequently (auth-free API, cheap), scraping sources should poll less often and with jitter to stay polite.
- **Per-tenant email-alert jobs**: one cron job per tenant with an active `TenderInboxConfig`, named `tender-inbox-poll-${tenantId}` — this is the one place Tessera's existing per-tenant scheduling gap (DKV's scheduler is documented as "v1 single-tenant, `findFirst()`") must actually be solved properly, since two tenants could each have their own mailbox subscribed to different portal alerts. Loop `TenderInboxConfig.findMany({ where: { isActive: true } })` on `onModuleInit` and register one job per row, mirroring `setInterval(intervalMin, tenantId)` but keyed by tenant instead of a single global slot.
- **Digest cron**: one platform-wide cron (e.g. hourly) that queries `TenderMatch` rows where `notifiedAt IS NULL` and the owning `TenderSavedSearch.notifyMode = 'digest'` and `digestHour` matches the current hour, groups by tenant+savedSearch, and hands off to `TenderMailService`.
**What:** Instead of traditional feature flags, Tessera uses module activation as the feature gating mechanism. If a module is not activated for a tenant, its routes return 403, its UI does not load, and its sidebar entry is hidden.
**Change detection:** `contentHash` (see Pattern 2) is the mechanism — `TenderIngestionService.upsert()` compares the freshly computed hash against the stored one; identical → touch `lastSeenAt` only, skip matching; different → update fields, recompute `contentHash`, add to this poll's "changed" delta so it flows through matching/notification again (a deadline extension should re-surface in a saved search, a brand-new notice obviously should).
**When:** Controlling feature access per tenant.
### Pattern 6: Notification Dispatch Reusing `SmtpConfig` + Fresh-Transport, Not the Global `MailerService`
## Anti-Patterns to Avoid
**What:** `TenderMailService` is built exactly like `DkvMailService` — `nodemailer.createTransport()` freshly per send, using `SettingsService.getDecryptedSmtpConfig(tenantId)` — **not** the platform's global `MailerService`/`@nestjs-modules/mailer` (which is reserved for system emails: password reset, welcome mail, configured once at bootstrap with a static transport). This distinction already exists in the codebase (`DkvMailService` vs `MailService`) and should hold here too: tender notifications are tenant-directed business content, and the tenant may have configured their own outbound SMTP relay that differs from the platform's.
### Anti-Pattern 1: Module-to-Module Direct Dependencies
- **Instant:** triggered synchronously (or via a lightweight in-process queue, matching DKV's direct-call style — no message broker in this stack) right after `TenderMatchingService` creates a `TenderMatch` with `savedSearch.notifyMode === 'instant'`.
- **Digest:** triggered by the digest cron (Pattern 5), batching all unnotified matches for a tenant+savedSearch into one summary email, then setting `notifiedAt` on each included `TenderMatch` — same "mark as sent" idempotency DKV uses for `DkvInvoiceHistory`.
**What:** Module A directly imports and calls Module B's internal functions.
**Trade-off:** reusing per-send transport creation means every notification email opens/closes its own SMTP connection (as DKV already accepts) — fine at Tessera's realistic tender-volume/tenant-count, and it guarantees an admin's SMTP config change takes effect on the very next send without a service restart (same Pitfall-3 mitigation DKV already documents).
**Why bad:** Creates coupling that makes modules impossible to activate independently. If Module B is deactivated, Module A breaks.
## Data Flow
**Instead:** Use an event bus or shared service layer. Modules publish events; other modules subscribe. If the publisher is missing, subscribers simply receive no events.
### Anti-Pattern 2: Tenant ID in Application Logic
**What:** Manually filtering by `tenant_id` in every query throughout the codebase.
**Why bad:** A single missed filter leaks data across tenants. Hundreds of places to maintain.
**Instead:** Use PostgreSQL RLS. Set tenant context once per request at the middleware level. All queries are automatically filtered. Defense in depth: the application layer still passes tenant_id, but RLS is the safety net.
### Anti-Pattern 3: Monolithic Frontend Bundle
**What:** Bundling all module UIs into a single JavaScript bundle.
**Why bad:** Users download code for modules they cannot access. Bundle size grows linearly with module count. Exposes unlicensed module code.
**Instead:** Code-split per module. Use dynamic imports. Only load module bundles when the user navigates to an active module.
### Anti-Pattern 4: Per-Tenant Database/Schema
**What:** Creating a separate PostgreSQL database or schema for each tenant.
**Why bad:** Operational nightmare at scale — migrations must run N times, connection pooling is per-tenant, monitoring multiplies. Overkill for a portal where tenants share the same data model.
**Instead:** Shared schema with RLS. Single migration path. Single connection pool. Isolation enforced at row level.
## Scalability Considerations
| Concern | At 10 users (internal) | At 100 tenants | At 1000+ tenants |
|---------|------------------------|----------------|------------------|
| **Database** | Single PostgreSQL, no pooler needed | PgBouncer for connection pooling | Read replicas, consider partitioning large tables by tenant_id |
| **Backend** | Single container | Horizontal scaling behind gateway (2-4 replicas) | Auto-scaling, consider extracting hot modules to separate services |
| **Frontend** | Single static bundle | CDN for static assets | CDN + edge caching, consider module federation for team-developed modules |
| **Module isolation** | Shared process, trust all modules | Same, but add resource limits per module route | Consider containerized module backends for untrusted/third-party modules |
| **File storage** | Local volume | S3-compatible object storage | Same + lifecycle policies |
## Suggested Build Order
Based on component dependencies, the recommended build order is:
### Ingestion → Notification Flow
```
Phase 1: Foundation
├── Docker Compose setup (PostgreSQL + backend + frontend containers)
├── PostgreSQL schema with tenant_id columns + RLS policies
├── Backend skeleton (Fastify + request lifecycle)
└── Frontend shell (React + routing + layout frame)
Phase 2: Authentication & Tenancy
├── Auth module (login, JWT, session management)
├── Tenant middleware (context injection, RLS activation)
├── User management (CRUD, roles)
└── LDAP integration
Phase 3: Module System
├── Module interface contract
├── Module registry (database-driven)
├── Module loader (route mounting, activation/deactivation)
└── First example module (Domaincheck)
Phase 4: Marketplace & Dashboard
├── Marketplace service (catalog, categories, licensing)
├── Marketplace UI (browse, activate, manage)
├── Dashboard service (layout persistence)
└── Dashboard UI (react-grid-layout, core widgets)
Phase 5: Polish & Desktop
├── i18n (DE + EN)
├── Light/Dark theme
├── Tauri desktop wrapper
└── Gitea CI/CD integration
[Cron tick: TenderSchedulerService]
v
[Adapter.fetchTenders(config, since)] -> RawTenderRecord[]
v
[TenderNormalizerService.normalize()] -> { fields, dedupKey, contentHash }
v
[TenderIngestionService.upsert()] -> prisma.tender.upsert({ where: { dedupKey } })
v (only rows that were newly created OR whose contentHash changed)
[TenderMatchingService.matchDelta(deltaIds)]
v (per active TenderSavedSearch, DB-scoped query)
[prisma.tenderMatch.upsert()]
v
+--------------------------+---------------------------+
| notifyMode='instant' | notifyMode='digest' |
v v
[TenderMailService.sendInstant()] [Digest cron batches unnotified matches -> sendDigest()]
v v
[nodemailer via tenant SmtpConfig, fresh transport per send]
```
**Dependency rationale:**
- Phase 1 first because everything depends on the database, backend framework, and frontend shell
- Phase 2 before modules because module activation requires knowing WHO is asking and WHICH tenant they belong to
- Phase 3 before marketplace because the marketplace manages modules — the module system must exist first
- Phase 4 can partially parallelize (dashboard is independent of marketplace) but both need the module system
- Phase 5 is pure enhancement — i18n/theme are cross-cutting but easier to retrofit than to block on
### Read Flow (Portal UI)
```
[User opens Ausschreibungs-Radar]
v
GET /tenders?keywords=&bundesland=&cpv=&minValue=&deadlineBefore= (TendersController)
v
prisma.tender.findMany({ where: <same structured-filter builder as TenderMatchingService> })
v
[Trefferliste UI] -> click -> GET /tenders/:id -> [Detail view, incl. rawPayload debug panel for admins]
[User manages saved searches]
v
POST/PUT/DELETE /tenders/saved-searches (tenant-scoped, req.tenantId)
v
prisma.tenderSavedSearch.upsert({ ..., tenantId })
v (on create/edit) -> one-off matchDelta() backfill against recent tenders (Pattern 4 mitigation)
```
## Scaling Considerations
| Scale | Architecture Adjustments |
|-------|---------------------------|
| Single tenant (current, internal test phase) | Exactly as designed above — global `Tender` table, one poller per source, no extra work needed even though only one tenant exists yet, because the schema is already tenant-agnostic at the data layer. |
| Multiple tenants, few saved searches each | `matchDelta()` cost scales with (poll batch size × active saved searches), both small — no changes needed. |
| Many tenants, many saved searches, high tender volume | Add a GIN full-text index on `title`/`description` (Pattern 4) before this becomes necessary, not after. If `matchDelta()` ever becomes a bottleneck, batch saved-search evaluation into a single SQL query per poll (`Tender × SavedSearch` cross-join filtered in one statement) instead of one query per search — straightforward migration since the where-clause builder is already centralized. |
### Scaling Priorities
1. **First bottleneck:** keyword filtering via `ILIKE` if the full-text index is skipped in the first slice — fix before it matters, it's a single migration (`CREATE INDEX ... USING GIN (to_tsvector('german', title || ' ' || coalesce(description,'')))`).
2. **Second bottleneck:** AI-NetServer/cosinex scraping adapters getting rate-limited or blocked as tender volume/poll frequency grows — mitigate with per-adapter jittered intervals and respecting any `Retry-After`, not a platform-wide fix.
## Anti-Patterns
### Anti-Pattern 1: Tenant-Scoping the `Tender` Table Like DKV Data
**What people do:** Copy the DKV template literally — add `tenantId` to `Tender`, run every adapter poll once per active tenant (matching `DkvSchedulerService`'s per-tenant cron intent).
**Why it's wrong:** Multiplies scraping requests against AI-NetServer/cosinex by tenant count (worse ToS exposure on portals already flagged "Niedrig-Mittel" risk), multiplies storage for identical public data, and makes cross-tenant dedup impossible (the same DÖE notice would need deduping *and* tenant-duplicating, which is incoherent).
**Do this instead:** Global `Tender` table (Pattern 3); tenant scoping lives one layer up in `TenderSavedSearch`/`TenderMatch`. Only the email-alert path (genuinely per-tenant mailbox) needs per-tenant scheduling.
### Anti-Pattern 2: One Adapter Per Portal Instead of Per Platform
**What people do:** Build a `LhsVpbwAdapter`, `Tender24Adapter`, `VergabeLandbwAdapter` as three separate classes because they're three separate portal URLs.
**Why it's wrong:** The feasibility research already established these three (plus many unlisted others) share the same AI AG NetServer fingerprint (`/NetServer/…ControllerServlet`) — one HTML/DOM shape. Three adapter classes triple the maintenance burden for zero behavioral difference; only the base URL and possibly a search-form parameter differ, which belongs in `TenderSourceConfig`, not in three code paths.
**Do this instead:** One `AiNetServerAdapter`, config-driven per portal instance (base URL + optional search params), same for a future `CosinexAdapter` covering DTVP and other cosinex Vergabemarktplatz instances.
### Anti-Pattern 3: In-Memory Filtering Across the Whole Table
**What people do:** `prisma.tender.findMany()` with no `where`, then `.filter()` in TypeScript per saved search.
**Why it's wrong:** Works fine in a demo with 50 rows, degrades badly once DÖE's Oberschwelle backbone plus scraped Unterschwelle notices accumulate over months, and re-does full-table work on every poll instead of scoping to the delta.
**Do this instead:** Pattern 4 — structured Prisma `where` + Postgres full-text index, scoped to the newly-changed batch.
## Integration Points
### External Services
| Service | Integration Pattern | Notes |
|---------|----------------------|-------|
| DÖE OpenData API (`oeffentlichevergabe.de`) | Auth-free HTTP GET against the OpenData/Swagger-documented endpoints; paginate; filter by `publishedAt`/last-poll cursor | eForms-DE / **OCDS `ocds-mnwr74`** / CSV formats available — prefer the OCDS export for direct field alignment with the `Tender` schema; ~75% of market value by € per feasibility doc, zero scraping/ToS risk |
| AI AG NetServer portals (lhs-vpbw, tender24, vergabe.landbw, + others) | HTML scrape of the public search result pages (no login required for search) | One adapter, config-driven per instance; feasibility doc rates ToS risk "Niedrig-Mittel" — implement politely (rate limit, honest UA string, cache ETags if offered) |
| cosinex Vergabemarktplatz (DTVP + other Länder instances) | HTML scrape of public search, structurally incompatible with AI-NetServer — separate adapter | Reusable across NRW/BB/NI/RLP cosinex instances per feasibility doc |
| RSS feeds (subreport-elvis, service.bund.de) | Standard RSS/Atom parse (e.g. `rss-parser` or `fast-xml-parser`), config-driven feed URL list | Lowest ToS risk of the scraped sources; service.bund.de rated "Niedrig" |
| Portal email alerts (Unterschwelle long tail, 8 of 10 portals) | Reuses `InboxProvider` (extracted `ImapProvider`/`ExchangeInboxProvider`) — per-tenant mailbox subscribed to each portal's native "gespeicherte Suche" alert | Requires manual one-time setup per portal (register a saved search on the portal itself); the module only ingests+parses the resulting alert emails, does not create the portal-side saved search |
| TED API v3 | Optional, deprioritized — keyless, EU-wide, largely redundant to DÖE for DE-only coverage | Build order: only if EU-wide coverage becomes a requirement later |
### Internal Boundaries
| Boundary | Communication | Notes |
|----------|----------------|-------|
| `TendersModule` ↔ `ModuleRegistryModule` | Direct DI import, `OnModuleInit` self-seed (`tenders.seed.ts`) — identical to `dkv.seed.ts` | New module import, no registry changes needed beyond the self-seed call |
| `TendersModule` ↔ `InboxModule` (new, extracted) | Direct DI import; `EmailAlertAdapter` depends on `ImapProvider`/`ExchangeInboxProvider` | Requires the one-time extraction of `inbox/` out of `dkv/` (see Structure Rationale) |
| `TendersModule` ↔ `SettingsModule` | Direct DI import, reused unmodified — `SettingsService.getDecryptedSmtpConfig(tenantId)` | Same pattern as `DkvModule` |
| `TendersModule` ↔ `TenantMiddleware`/`TenantGuard` | `req.tenantId` extraction for saved-search/notification endpoints only — **not** for `GET /tenders` list/detail, which is platform-global read access gated only by `TenantModuleActivation` (module licensing), not by tenant-owned data | This is the one controller where "tenant-scoped" and "tenant-gated" genuinely differ — worth flagging explicitly in the phase plan so it isn't implemented as a blanket `where: { tenantId }` by habit |
| `TendersController` ↔ `module-loader.ts` (web) | New `MODULE_REGISTRY['tender-radar']` entry, same as `cert-manager`/`dkv-fleet` | One-line addition, whitelist pattern (`T-03-09`) — must not be skipped or the module page 404s even if activated |
| `AppModule` ↔ `TendersModule` | New import in `app.module.ts` | One line |
## Build Order — Ships DÖE-First as a Usable Slice
Ordered by dependency; each step after step 8 is additive and doesn't touch the pipeline built before it (the point of the adapter abstraction).
**Phase A — DÖE-only usable slice (schema + one source + filter + UI + notification, end-to-end):**
1. Prisma migration: `Tender`, `TenderSourcePollConfig`, `TenderSavedSearch`, `TenderMatch` (+ GIN full-text index on `title`/`description`).
2. `TendersModule` skeleton + `tenders.seed.ts` (module-registry self-seed, category e.g. `procurement`).
3. `TenderSourceAdapter` interface + `DoeOpenDataAdapter` (auth-free, OCDS-formatted fetch).
4. `TenderNormalizerService` (OCDS release → `Tender` fields, `ocid`-based dedup key, `contentHash`).
5. `TenderIngestionService` (poll → normalize → upsert → change-detect) + `TenderSchedulerService` (single global cron for DÖE).
6. `TenderMatchingService` (DB-query filter evaluation, Pattern 4) + `TendersController` (`GET /tenders`, `GET /tenders/:id`, saved-search CRUD).
7. `apps/web/.../modules/tender-radar/` — trefferliste + filter UI + saved-search management + `MODULE_REGISTRY` entry. **This alone is already a usable, demoable slice** — DÖE covers ~75% of market value by € per feasibility doc.
8. `TenderMailService` (instant + digest, reusing `SmtpConfig`) + digest cron. Closes the loop on the milestone's "optional E-Mail-Versand" requirement using only the DÖE source.
**Phase B — Scraping adapters for the Unterschwelle long tail (pipeline unchanged, adapters only):**
9. `AiNetServerAdapter` (covers lhs-vpbw, tender24, vergabe.landbw + any future AI AG portal via config).
10. `CosinexAdapter` (covers DTVP; reusable for other cosinex Länder marketplaces later).
**Phase C — RSS + email-alert (lowest ROI per feasibility doc, do last):**
11. `RssAdapter` (subreport-elvis, service.bund.de).
12. Extract `inbox/` shared module out of `dkv/` (prerequisite refactor).
13. `TenderInboxConfig` (per-tenant, mirrors `DkvModuleConfig` credential pattern) + `EmailAlertAdapter` + per-tenant scheduler jobs.
**Explicitly out of this build order:** TED API v3 (redundant to DÖE for DE-only), vergabe24/aumass (AGB-prohibited scraping — feasibility doc flags these as avoid).
## New vs Modified — Explicit Inventory
**New:**
- `apps/api/src/tenders/` — entire module (controller, services, scheduler, seed, dto/, adapters/, types)
- `apps/api/src/inbox/` — extracted shared inbox module
- Prisma models: `Tender`, `TenderSourcePollConfig`, `TenderSavedSearch`, `TenderMatch`, `TenderInboxConfig`
- `apps/web/src/app/(portal)/modules/tender-radar/` — list, detail, saved-search, settings pages
**Modified:**
- `apps/api/prisma/schema.prisma` — add the 5 new models + migration
- `apps/api/src/dkv/` — `providers/` folder removed, imports `InboxModule` instead; `dkv.types.ts` trimmed of `InboxConfig`/`InboxEmail`/`InboxAttachment` (moved to `inbox/inbox.types.ts`)
- `apps/api/src/app.module.ts` — import `TendersModule` (and `InboxModule` if not auto-imported via `TendersModule`'s own imports)
- `apps/web/src/lib/module-loader.ts` — add `'tender-radar'` entry to `MODULE_REGISTRY`
**Explicitly NOT modified:** `module-registry.service.ts`, `prisma-tenant.extension.ts` (`forTenant`), `mail.module.ts`/`mail.service.ts` (global system mailer stays untouched — tender notifications use the DKV-style per-tenant transport pattern instead), `settings.service.ts`.
## Sources
- [Multi-Tenant Databases with Postgres Row-Level Security](https://www.midnytecity.com.au/blogs/multi-tenant-databases-with-postgres-row-level-security)
- [AWS: Multi-tenant data isolation with PostgreSQL Row Level Security](https://aws.amazon.com/blogs/database/multi-tenant-data-isolation-with-postgresql-row-level-security/)
- [Approaches to implementing multi-tenancy in SaaS applications - Red Hat](https://developers.redhat.com/articles/2022/05/09/approaches-implementing-multi-tenancy-saas-applications)
- [Node.js Plugin Architecture: Build Your Own Plugin System](https://medium.com/codeelevation/node-js-plugin-architecture-build-your-own-plugin-system-with-es-modules-5b9a5df19884)
- [How to Build Plugin Architecture in Node.js](https://oneuptime.com/blog/post/2026-01-26-nodejs-plugin-architecture/view)
- [Building Customizable Dashboard Widgets Using React Grid Layout](https://www.antstack.com/blog/building-customizable-dashboard-widgets-using-react-grid-layout/)
- [react-grid-layout - GitHub](https://github.com/react-grid-layout/react-grid-layout)
- [Micro-Frontend Architecture with Module Federation](https://module-federation.io/)
- [Tauri vs Electron: The Complete Developer's Guide (2026)](https://blog.nishikanta.in/tauri-vs-electron-the-complete-developers-guide-2026)
- [Tauri in 2026: Build Cross-Platform Desktop Apps](https://dev.to/ottoaria/tauri-in-2026-build-cross-platform-desktop-apps-with-web-technologies-better-than-electron-11mo)
- [Using a Reverse Proxy to Expose Multiple Microservices Through a Single Port in Docker Compose](https://dev.to/syed_omair/using-a-reverse-proxy-to-expose-multiple-microservices-through-a-single-port-in-docker-compose-4h9e)
- [Developing a Multi-Tenant SaaS Application: The 2026 Architecture Guide](https://apipilot.com/developing-a-multi-tenant-saas-application-the-2026-architecture-guide/)
- Existing codebase (verified by direct read): `apps/api/src/dkv/*`, `apps/api/src/module-registry/*`, `apps/api/src/mail/*`, `apps/api/src/prisma/prisma-tenant.extension.ts`, `apps/api/src/tenant/tenant.middleware.ts`, `apps/api/prisma/schema.prisma`, `apps/web/src/lib/module-loader.ts`, `apps/web/src/app/(portal)/modules/dkv-fleet/*` — HIGH confidence, ground truth.
- `.planning/research/ausschreibungs-portale-feasibility.md` (2026-07-16) — portal platform fingerprints, ToS risk ratings, DÖE/TED coverage estimates — HIGH confidence (project's own prior research).
- [OCDS Release Reference — Open Contracting Data Standard 1.1.5](https://standard.open-contracting.org/latest/en/schema/reference/) — core schema fields (ocid, release, tender, parties, buyer) — MEDIUM confidence (public spec, not project-specific).
- [OCDS Building Blocks](https://standard.open-contracting.org/latest/en/getting_started/building_blocks/) — OCID composition (registered prefix + publisher-chosen process id) — MEDIUM confidence.
- [oeffentlichevergabe.de OpenData Swagger UI](https://oeffentlichevergabe.de/documentation/swagger-ui/opendata/index.html) — confirms `ocds-mnwr74` as the registered German federal OCDS prefix (registered 2023-02-06) and CC-Zero licensing — MEDIUM confidence (page requires JS to render full endpoint/pagination detail; exact pagination parameters remain an open verification point, already flagged in the feasibility doc).
---
*Architecture research for: Tessera Ausschreibungs-Radar module (v1.1 milestone)*
*Researched: 2026-07-17*
+170 -86
View File
@@ -1,112 +1,196 @@
# Feature Landscape
# Feature Research
**Domain:** Modular portal platform with marketplace, multi-tenancy, and workflow tool integration
**Researched:** 2026-06-18
**Domain:** Tender / public-procurement monitoring & aggregation (German "Ausschreibungs-Radar" module, Tessera v1.1)
**Researched:** 2026-07-17
**Confidence:** MEDIUM-HIGH (domain patterns cross-checked against live tender-monitoring SaaS products — Stotles, Tendium, TenderAlerts, Jorpex — plus OCDS standard docs and the project's own portal feasibility research; German-specific integration details inherit the HIGH confidence already established in `ausschreibungs-portale-feasibility.md`)
## Table Stakes
> **Scope note:** This supersedes the v1.0 platform-level `FEATURES.md` (portal shell, auth, multi-tenancy, marketplace, dashboard widgets) for roadmap purposes — those features already shipped and are not part of this milestone. This document covers ONLY the new Ausschreibungs-Radar tender module.
Features users expect. Missing = product feels incomplete or unprofessional.
## How This Category Works
Every mature tender-monitoring product (Stotles, Tendium, TenderAlerts, deutsche-eVergabe, vergabe24) follows the same shape, regardless of market:
1. **Ingest** from N sources (APIs, scrapers, RSS, email alerts) into a **normalized schema**.
2. **Deduplicate** across sources — the same tender is frequently published on 2-3 portals plus a central feed (DÖE/TED equivalent).
3. Let each user/tenant define **saved-search profiles** (keyword + geography + classification code + deadline + value filters).
4. Show a **searchable results list + detail view** linking back to the source portal/documents.
5. Track **per-user state** (read/unread, shortlisted/favourite) on top of the shared tender catalog.
6. **Notify**: a configurable periodic digest email, plus an instant alert the moment a new tender matches a saved search.
Tessera's module maps directly onto this shape. The only domain-specific complexity is #1-2 (many incompatible German portal formats, see feasibility doc) — #3-6 are standard SaaS patterns Tessera already has adjacent infrastructure for (DKV inbox ingestion, MailModule/SMTP, NestJS `@nestjs/schedule`, multi-tenant config).
## Feature Landscape
### Table Stakes (Users Expect These)
Features users assume exist. Missing these = product feels incomplete.
| Feature | Why Expected | Complexity | Notes |
|---------|--------------|------------|-------|
| User authentication (email/password + admin-created accounts) | Every portal needs login; without it nothing works | Medium | Foundation for all access control |
| Role-Based Access Control (RBAC) | Users expect permission boundaries; admins expect control | Medium | Tenant-scoped roles are critical for multi-tenancy |
| Multi-tenancy with data isolation | Core requirement per PROJECT.md; customers expect their data is isolated | High | Row-level security (tenant_id on every table) or schema-per-tenant |
| Sidebar navigation with categories | Standard portal pattern; users expect hierarchical navigation | Low | Collapsible, categorized by module type |
| Module activation/deactivation per tenant | Marketplace without on/off is just a list; tenants expect control | Medium | Admin toggles module visibility and access |
| Module licensing (admin-managed) | Core business model; tenants expect clear "you have access to X" | Medium | License = permission to activate; no payment integration yet |
| Responsive layout | Web apps must work on different screen sizes | Medium | Desktop-first, but must not break on tablet |
| Light/Dark theme | Expected in modern apps; users notice its absence | Low | CSS custom properties + toggle; store preference per user |
| Internationalization (i18n) - DE + EN | Core requirement; German users expect German UI | Medium | Must be baked in from day one; retrofitting i18n is painful |
| Basic dashboard with widgets | Central landing area gives users a "home" | Medium | Clock, search, notes, calendar as starting widgets |
| Session management | Users expect to stay logged in, and to be logged out on timeout | Low | Token-based with configurable expiry |
| Error handling and user feedback | Users expect clear feedback on actions (success/error toasts) | Low | Global notification/toast system |
| Loading states and skeleton screens | Without them, users think the app is broken | Low | Standard UX pattern |
| Search / filter within module lists | Even with 10 modules, users expect to type-to-find | Low | Client-side filter on marketplace and sidebar |
| User profile and settings | Users expect to change language, theme, password | Low | Per-user preferences storage |
| Central tender ingestion (DÖE OpenData API) | This is the reason the module exists — without it there's no data | MEDIUM | Auth-free eForms-DE/OCDS/CSV, single API, ~75% market value by € — see feasibility doc. No scraping risk. **MVP foundation.** |
| Normalized tender schema (OCDS-oriented) | Every source has a different shape; UI/filters need one consistent model | MEDIUM | Store title, buyer, CPV code(s), region/PLZ, deadline, estimated value, procedure type, source URL(s), raw payload |
| Full-text keyword search | Users think in free text ("Straßenbau Ludwigsburg"), not structured codes alone | LOW-MEDIUM | Postgres `tsvector`/GIN index on title+description is sufficient at this scale; no need for Elasticsearch |
| Region / PLZ / Bundesland filter | German public buyers are geographically scoped; users only want their operating radius | LOW-MEDIUM | PLZ → Bundesland/Landkreis mapping table (static reference data); PLZ range or radius filter |
| CPV code / branch filter | CPV (Common Procurement Vocabulary) is the industry-standard classification for tenders — OCDS confirms this is the canonical item classification | MEDIUM | Needs a static CPV code list (~9,400 codes, hierarchical, EU-published) with autocomplete/browse, not just free-text code entry |
| Submission deadline filter | Deadline is the single most decision-critical field — "can we even respond in time" | LOW | Simple date range filter; must also drive UI sort/highlighting (see below) |
| Estimated contract value min/max filter | Filters out tenders too small/large to be worth pursuing | LOW | Optional field — many Unterschwelle notices omit value; filter must handle NULL gracefully |
| Saved-search / filter profiles | Users don't want to re-enter the same 5 filters every day — every competitor product treats "saved search" as the core object, not the raw search | MEDIUM | Per-user (not just per-tenant) — colleagues in the same tenant often watch different branches/regions |
| Searchable/sortable results list | Baseline UX for any list-based tool; sort by deadline (most common), value, publish date | LOW-MEDIUM | Standard data-table pattern, consistent with rest of Tessera portal UI |
| Detail view with source link + document access | Users must reach the actual Vergabeunterlagen to bid — the module is a radar, not a bidding platform | LOW-MEDIUM | For DÖE/eForms: link to source portal + any document URLs present in the notice; do not attempt to mirror/host bid documents |
| Read/unread marking | Universal "inbox" pattern — users triage a daily list of new matches | LOW | Per-user state table, not a tenant-shared flag |
| Favourite/shortlist marking | Lets users bookmark tenders they intend to pursue, separate from the read/unread triage state | LOW | Per-user state table; simple boolean + optional "shortlist view" filter |
| Periodic email digest | Every native portal alert (see feasibility doc, portals 2,3,4,6,7,8,9,10) already trains users to expect an inbox digest, not "log in and check" | MEDIUM | **Reuses existing `MailModule`/SMTP infra.** Configurable interval (daily/weekly) per saved search or per user, sent via existing scheduler pattern (`dkv-scheduler.service.ts` precedent) |
| Instant alert on new match | Deadline-sensitive tenders lose value if a digest arrives after a same-day cutoff was missed; competitors (Stotles, Tendium, native portal Suchprofile) all offer immediate notification as an alternative to digest | MEDIUM | Event-driven: at ingestion time, run new tenders against active saved searches, queue+send immediately. **Reuses `MailModule`.** |
## Differentiators
### Differentiators (Competitive Advantage)
Features that set Tessera apart from generic portals. Not expected, but valued.
Features that set the product apart. Not required, but valuable — align with Tessera's core value ("one platform instead of switching between tools").
| Feature | Value Proposition | Complexity | Notes |
|---------|-------------------|------------|-------|
| Drag-and-drop dashboard with resizable widgets | Personal workspace feel; makes dashboard actually useful vs static page | High | Use react-grid-layout or gridstack.js; persist layout per user per tenant |
| Widget marketplace/gallery | Users can add widgets to their dashboard from a catalog | Medium | Distinct from module marketplace; lightweight UI components |
| Module hot-activation without restart | Modules appear instantly after license grant; no deploy needed | High | Requires dynamic route loading or micro-frontend approach |
| LDAP/AD directory sync | Enterprise differentiator; auto-provision users from corporate directory | High | JIT provisioning at login + periodic group sync |
| Per-tenant branding/customization | Each tenant gets their own logo, color accent | Medium | Stored in tenant config; CSS variable override |
| Module-provided dashboard widgets | Modules can contribute widgets to the dashboard system | Medium | Module manifest declares available widgets; loose coupling |
| Activity feed / audit log | Transparency on who did what; compliance-ready | Medium | Event sourcing pattern; filterable by user, action, module |
| Admin impersonation ("view as user") | Support tool; admin can see exactly what a user sees | Medium | Scoped session with clear visual indicator |
| Keyboard shortcuts and command palette | Power-user acceleration; "Ctrl+K" to jump anywhere | Low | Global listener + fuzzy search over routes and actions |
| Onboarding wizard for new tenants | Guides new tenant admins through setup; reduces support burden | Medium | Multi-step flow: branding, users, module selection |
| Module dependency declaration | Module A requires Module B; platform enforces this | Low | Manifest-level declaration; block activation if dependency missing |
| Notification center | Unified inbox for system events, module alerts, admin messages | Medium | WebSocket or SSE for real-time; persisted read/unread state |
| Desktop wrapper (Electron/Tauri) | Installable app feel; taskbar presence, native notifications | Medium | Tauri preferred (smaller binary, Rust-backed) |
| Cross-source deduplication with merged source links | The single most-cited "must-have but hard" feature in every tender-aggregator review (Stotles, EU Tenders Monitor) — same Vergabe published on DÖE + AI-AG portal + RSS should appear once, not 3 times | HIGH | No shared ID across sources — needs fuzzy match on buyer+title+CPV+deadline+value within a time window. Only relevant once ≥2 sources are active — **defer past DÖE-only MVP** |
| Dashboard widget: upcoming deadlines | Tessera already has a configurable drag-and-drop dashboard with a Calendar widget — an "Ausschreibungen mit nahender Frist" widget reuses that framework directly instead of forcing users into the module | LOW-MEDIUM (once dashboard widget SDK exists) | Direct synergy with the already-built Phase 8 dashboard-widgets work; high leverage, low net-new complexity |
| Per-tenant + per-user saved searches with team visibility | Multiple colleagues at one tenant can watch overlapping-but-different regions/branches without duplicating ingestion work | LOW (data-model addition on top of table-stakes saved search) | Natural fit given Tessera's existing multi-tenant architecture |
| CSV/Excel export of filtered results | Procurement/BD teams routinely need to hand a tender list to someone outside the tool (management, partner companies) | LOW | Reuse existing export patterns if any exist elsewhere in Tessera; otherwise trivial with the normalized schema |
| Relevance ranking within a saved search | Once keyword+CPV+region filters return more than a handful of hits, a simple recency sort undersells partial matches; competitors (Stotles "Signal Score") rank rather than just filter | MEDIUM-HIGH | Defer — only valuable once ingestion volume is large enough that plain filtering feels noisy; not needed for DÖE-only MVP |
| "Manual watch" flag for anti-scraping portals (vergabe24, aumass) | Rather than silently having no coverage, the UI can list these portals as "nicht automatisch überwacht — manuell prüfen" with a direct bookmark link, so users aren't surprised by a coverage gap | LOW | Cheap trust-building feature; prevents users assuming full coverage when 2 portals are deliberately excluded per ToS |
## Anti-Features
### Anti-Features (Commonly Requested, Often Problematic)
Features to explicitly NOT build. These add complexity without proportional value for Tessera's use case.
Features that seem good but create problems.
| Anti-Feature | Why Avoid | What to Do Instead |
|--------------|-----------|-------------------|
| Third-party module SDK / developer portal | Massive complexity (sandboxing, review pipeline, versioning); only own modules planned | Build a clean internal module contract; open up later if demand arises |
| Payment/billing integration | Out of scope per PROJECT.md; adds regulatory and UX burden | Admin-managed license grants; add Stripe/payment only when selling externally |
| Real-time collaboration (multiplayer editing) | Workflow tools are typically single-user operations; CRDT/OT is enormously complex | Each module handles its own data; no shared editing state |
| AI/ML-powered recommendations | "Suggested modules" adds little value with a small catalog; ML overhead is huge | Manual curation and categories; maybe simple "popular modules" counter later |
| Native mobile app | Web-first + desktop wrapper covers the use case; mobile adds two platforms to maintain | Responsive web design; PWA if truly needed later |
| Complex workflow orchestration engine (BPMN) | Tessera modules ARE the tools; building a meta-workflow layer is a product in itself | Each module handles its own workflow; cross-module orchestration is future scope |
| White-label/full-rebrand per tenant | Different from "per-tenant branding"; full white-label means separate builds, domains, assets | Offer logo + accent color customization; not full theme overhaul per tenant |
| Plugin sandboxing (iframe/WASM isolation) | Only own modules are deployed; sandboxing is for untrusted third-party code | Modules are trusted first-party Docker containers; share the same runtime |
| Social features (comments, reactions, @mentions) | Not a collaboration tool; adds social complexity without clear workflow value | Keep modules focused on their task; add module-specific notes if needed |
| Granular per-field permissions | RBAC at role/module level is sufficient; field-level ACL is enterprise overkill | Role -> Module access mapping; maybe per-module "read/write/admin" tiers |
| Feature | Why Requested | Why Problematic | Alternative |
|---------|---------------|------------------|-------------|
| Scraping vergabe24.de / aumass.de directly | "We already pay for these, just pull the data automatically" | Both portals' AGB explicitly prohibit automated/scripted access; aumass additionally paywalls its alert feature. Legal/ToS risk for a product Tessera intends to resell to customers | Their Oberschwellen notices already flow through DÖE (both are confirmed DÖE data suppliers per feasibility doc); for Unterschwellen coverage, register a native portal Suchprofil manually and ingest the resulting alert emails like any other Unterschwellen source |
| Generic "scrape any portal the user pastes a URL for" framework | Feels maximally flexible — "just point it at any Vergabeportal" | German portal landscape has ≥3 incompatible platform families (cosinex, AI-AG NetServer, subreport, bespoke); a generic scraper is either fragile (breaks on every portal redesign) or a permanent maintenance sink. Also raises ToS risk per-portal at scale | Ship a small, fixed set of well-understood adapters (DÖE API, one AI-AG adapter, one cosinex adapter, 2 RSS feeds) as scoped in PROJECT.md; treat any further portal as a deliberate, individually-evaluated addition |
| Real-time (sub-hourly) polling of every source | "Instant alert" sounds like it needs constant polling | Tender lifecycles run days-to-weeks; native portal alert emails themselves are typically daily-batch. Sub-hourly polling multiplies scraping load/ToS exposure for near-zero user value | Poll DÖE/RSS on an hourly-to-daily schedule (source-dependent); "instant alert" means instant relative to the last poll/ingestion, not sub-minute real-time |
| Full bid-management / CRM (proposal drafting, submission tracking, win/loss pipeline) | "While we're in the tool, why not manage the whole bid lifecycle" | That is a different product category (bid-management software, e.g. Loopio-class tools) — much larger scope, and out of step with the module's stated goal ("durchsuchen, filtern, anzeigen, versenden") | Keep the module a radar: discovery + filtering + notification + shortlisting. Document-heavy bid workflow can be a future, separate module if ever justified |
| AI-generated tender summaries / bid-fit scoring | Sounds like a strong differentiator, "let AI tell us if this tender is worth pursuing" | Real value requires per-tenant scoring criteria/training data that doesn't exist yet; premature AI feature adds cost/complexity/hallucination risk before the basic radar has been validated in production | Defer to v2+; if pursued later, treat as its own AI-SPEC-gated phase per Tessera's dev workflow, not part of the v1.1 module |
| Mirroring/hosting original Vergabeunterlagen (PDFs) locally | "One-stop shop, don't make users leave the app" | Legal ambiguity around redistributing third-party procurement documents; storage/sync burden keeping mirrors current as portals update documents | Link out to the source portal's document URL; only cache what's needed for parsing (e.g. eForms XML), not full bid document sets |
## Feature Dependencies
```
Authentication -> RBAC -> Multi-Tenancy (each layer builds on the previous)
Multi-Tenancy -> Module Licensing (licenses are tenant-scoped)
Module Licensing -> Module Activation (can't activate without license)
Module Activation -> Marketplace UI (marketplace displays activation state)
Dashboard Framework -> Widget System -> Drag-and-Drop Layout
Dashboard Framework -> Module-Provided Widgets (modules contribute to dashboard)
i18n Framework -> All UI Components (must be in place before building UI)
Theme System -> All UI Components (CSS variables must exist before components)
LDAP Integration -> User Management (extends, does not replace manual management)
Notification Center -> Module Events (modules emit events to notification system)
Sidebar Navigation -> Module Registry (sidebar reflects activated modules)
Normalized tender schema (OCDS-oriented)
└──requires──> DÖE OpenData ingestion (MVP source)
└──enables──> Filter engine (keyword, region, CPV, deadline, value)
└──requires──> Saved-search / filter profiles
├──requires──> Instant alert (needs Mail infra)
└──requires──> Periodic digest (needs Mail infra + Scheduler)
Read/unread marking ──requires──> per-user tender state table (independent of shared tender catalog)
Favourite/shortlist marking ──requires──> per-user tender state table (same table as read/unread)
Cross-source deduplication ──requires──> ≥2 active ingestion sources
(portal adapters or RSS) └──blocks──> "Portal-Adapter (AI-AG/cosinex)" phase
└──blocks──> "RSS-Quellen" phase
└──blocks──> "E-Mail-Alert-Ingestion" phase
E-Mail-Alert-Ingestion (Unterschwellen) ──reuses──> DKV inbox infra (ImapProvider / ExchangeInboxProvider)
Dashboard "upcoming deadlines" widget ──enhances──> Saved-search / filter profiles
(reuses existing dashboard widget framework, Phase 8)
Instant alert ──conflicts-with-if-unthrottled──> Periodic digest
(same match sent twice — needs a "already alerted" flag on the per-user tender state
so a tender doesn't also appear in the next digest as if new)
```
## MVP Recommendation
### Dependency Notes
Prioritize in this order:
- **Filter engine requires the normalized schema, which requires at least the DÖE source to exist:** there is nothing to filter until data is flowing. This is why DÖE-only ingestion is the correct MVP anchor — everything else (saved search, results UI, notifications) can be built and validated against real live data from day one, without waiting on portal-adapter scraping work.
- **Cross-source dedup requires ≥2 sources:** building dedup logic against a single source is meaningless (nothing to deduplicate against). This confirms dedup belongs strictly *after* the first portal adapter or RSS source ships, not in the MVP.
- **Instant alert and periodic digest both require Mail infra, but need a shared "already notified" state** to avoid double-notifying the same user about the same tender (once instantly, once again in the next digest). This is a small but easy-to-miss data-model requirement — flag for the phase that builds notifications.
- **Read/unread and favourite/shortlist state must be per-user, not per-tenant:** the tender catalog itself is shared reference data (a public tender exists independent of which tenant is looking at it), but triage state is personal. This is architecturally different from the DKV module, where nearly everything is tenant-scoped by nature. Worth flagging explicitly for the data-model design in this phase.
- **E-Mail-Alert-Ingestion reuses DKV's inbox provider interface** (`apps/api/src/dkv/providers/inbox-provider.interface.ts`, with `ImapProvider`/`ExchangeInboxProvider` implementations) rather than building new inbox-polling code — the DKV module already solved "watch a mailbox, parse structured content out of incoming mail" for a different structured format (fleet PDFs). The tender module needs the same mailbox-watching mechanics, different parser.
- **Periodic digest and instant alert reuse the existing `MailModule`/SMTP** (`apps/api/src/mail/mail.module.ts`, `mail.service.ts`) and the scheduler pattern already established by `dkv-scheduler.service.ts` and `ldap-sync.scheduler.ts` — no new mail-sending or cron infrastructure needed, only new templates and trigger logic.
1. **Authentication + RBAC + Multi-Tenancy** - Foundation; nothing works without it
2. **Sidebar navigation + Module registry** - Portal shell; gives the app structure
3. **i18n framework (DE + EN)** - Must be first, before building UI text
4. **Theme system (light/dark)** - Must be first, before building styled components
5. **Marketplace UI with licensing/activation** - Core business logic
6. **Basic dashboard with static widgets** - User home; clock, search, notes
7. **Domaincheck module** - First real module; validates the entire module architecture
8. **Drag-and-drop dashboard** - Differentiator; upgrade from static layout
## MVP Definition
Defer:
- **LDAP integration**: High complexity, not needed for initial internal use (manual user creation suffices)
- **Desktop wrapper**: Adds build pipeline complexity; browser works fine initially
- **Notification center**: Useful but not critical until multiple modules exist
- **Admin impersonation**: Support tool; not needed until external customers arrive
- **Onboarding wizard**: Only valuable with external tenants; internal users get manual setup
### Launch With (v1 / MVP — DÖE-API-only)
Minimum viable product — validates the whole module concept against real, legally unambiguous, single-source data before any scraping work begins.
- [ ] DÖE OpenData API ingestion (eForms/OCDS/CSV, auth-free) — the only data source needed to prove the concept
- [ ] Normalized tender schema (OCDS-oriented) — required so later sources slot in without a schema rewrite
- [ ] Filter engine: keyword full-text, region/PLZ/Bundesland, CPV code, deadline range, value range
- [ ] Saved-search / filter profiles (per user, tenant-aware)
- [ ] Searchable/sortable UI results list, sortable by deadline/value/publish date
- [ ] Detail view with source link + any document URLs present in the DÖE notice
- [ ] Read/unread marking (per-user)
- [ ] Favourite/shortlist marking (per-user)
- [ ] Periodic email digest, interval configurable in the webinterface — reuses `MailModule`
- [ ] Instant email alert on new match against an active saved search — reuses `MailModule`
### Add After Validation (v1.x)
Features to add once the DÖE-only MVP is live and validated in production use.
- [ ] AI-AG NetServer portal adapter (covers portals lhs-vpbw, tender24, vergabe.landbw + other AI-AG-hosted portals) — trigger: MVP validated, Unterschwellen coverage gap confirmed as a real pain point
- [ ] cosinex VMP adapter (DTVP, reusable across other cosinex-hosted Länder-Marktplätze) — trigger: same as above
- [ ] RSS ingestion (subreport-elvis, service.bund.de) — trigger: lowest-effort source expansion, can land alongside or before the scraping adapters
- [ ] E-Mail-Alert-Ingestion for remaining Unterschwellen portals, reusing DKV inbox infra — trigger: once ≥1 scraping adapter or RSS source exists, so there's a reason to worry about cross-source overlap
- [ ] Cross-source deduplication (fuzzy match on buyer+title+CPV+deadline+value) — trigger: mandatory as soon as a second source goes live, otherwise duplicate tenders will visibly degrade the results list
- [ ] "Manual watch" indicator for vergabe24/aumass (excluded portals) — trigger: first user question about "why don't I see X"
### Future Consideration (v2+)
Features to defer until the core radar has proven its value.
- [ ] TED API v3 (EU-wide redundancy) — defer: largely redundant to DÖE for DE-only scope, only relevant if cross-border tenders become a requirement
- [ ] Relevance/ranking scoring within saved searches — defer: only needed once match volume per saved search is large enough that plain filtering feels noisy
- [ ] Dashboard "upcoming deadlines" widget — defer: high-leverage but depends on prioritizing dashboard-widget-SDK reuse work; not core to validating the radar itself
- [ ] CSV/Excel export — defer: cheap to add later, not needed to validate core value
- [ ] Team/collaboration features (assign tender to colleague, internal notes) — defer: turns the module toward bid-management scope creep; only pursue if users explicitly ask post-launch
## Feature Prioritization Matrix
| Feature | User Value | Implementation Cost | Priority |
|---------|------------|----------------------|----------|
| DÖE OpenData ingestion + normalized schema | HIGH | MEDIUM | P1 |
| Filter engine (keyword/region/CPV/deadline/value) | HIGH | MEDIUM | P1 |
| Saved-search profiles | HIGH | MEDIUM | P1 |
| Results list + detail view | HIGH | LOW-MEDIUM | P1 |
| Read/unread + favourite marking | MEDIUM | LOW | P1 |
| Periodic digest + instant alert (via MailModule) | HIGH | MEDIUM | P1 |
| AI-AG NetServer adapter | MEDIUM-HIGH | HIGH | P2 |
| cosinex DTVP adapter | MEDIUM | MEDIUM-HIGH | P2 |
| RSS ingestion (subreport, service.bund.de) | MEDIUM | LOW-MEDIUM | P2 |
| E-Mail-Alert-Ingestion (Unterschwellen) | MEDIUM | MEDIUM (reuses DKV infra) | P2 |
| Cross-source deduplication | HIGH (once ≥2 sources) | HIGH | P2 |
| Dashboard "upcoming deadlines" widget | MEDIUM | LOW-MEDIUM | P3 |
| Relevance ranking | LOW-MEDIUM | MEDIUM-HIGH | P3 |
| CSV/Excel export | LOW-MEDIUM | LOW | P3 |
| Team/collaboration (notes, assignment) | LOW | MEDIUM-HIGH | P3 |
**Priority key:**
- P1: Must have for launch (DÖE-only MVP)
- P2: Should have, add when possible (portal-adapter / multi-source phases)
- P3: Nice to have, future consideration
## Competitor Feature Analysis
| Feature | Native German portals (DTVP, subreport, deutsche-eVergabe) | Modern SaaS aggregators (Stotles, Tendium, TenderAlerts) | Tessera's approach |
|---------|--------------------------------------------------------------|-----------------------------------------------------------|---------------------|
| Saved search + alert | Yes — "Suchprofil" per portal, daily email | Yes — core object of the product | Yes — single unified saved-search model across all sources, not one per portal |
| Multi-source coverage | No — each portal only shows its own listings | Yes — 50-1,000+ sources aggregated centrally | Yes, phased: DÖE (central feed) → 2 portal adapters → RSS → email-ingestion for the long tail |
| Deduplication | N/A (single source) | Yes — explicitly marketed as a core feature (Stotles, EU Tenders Monitor) | Deferred to v1.x, once ≥2 sources are live (see dependency notes) |
| Read/unread + shortlist | Rare — most native portals only offer a flat list | Yes — standard inbox-style triage pattern | Yes, in MVP — per-user state table |
| Deadline-aware notification | Digest only, no distinction between digest and instant | Some distinguish instant vs. digest | Both instant alert and periodic digest, user-configurable, in MVP |
| Bid-management/CRM add-ons | No | Some upsell into full bid-management (Stotles) | Explicitly out of scope (anti-feature) — module stays a radar |
## Sources
- [WorkOS: Multi-tenant RBAC design](https://workos.com/blog/how-to-design-multi-tenant-rbac-saas)
- [Logto: Build a multi-tenant SaaS application](https://logto.medium.com/build-a-multi-tenant-saas-application-a-complete-guide-from-design-to-implementation-d109d041f253)
- [Cloudscape Design System: Configurable Dashboard](https://cloudscape.design/patterns/general/service-dashboard/configurable-dashboard/)
- [FreeCodeCamp: Type-safe plugin architecture in React](https://www.freecodecamp.org/news/how-to-design-a-type-safe-lazy-and-secure-plugin-architecture-in-react/)
- [Backstage.io: Plugin-based developer portal](https://backstage.io/)
- [AppMaster: Audit logging for internal tools](https://appmaster.io/blog/audit-logging-internal-tools-activity-feed)
- [Frontegg: SaaS Multitenancy components](https://frontegg.com/blog/saas-multitenancy)
- [Medium: Node.js Plugin Architecture with ES Modules](https://medium.com/codeelevation/node-js-plugin-architecture-build-your-own-plugin-system-with-es-modules-5b9a5df19884)
- [DevelopersVoice: Plugin-ready modular monolith](https://developersvoice.com/blog/dotnet/building_plugin_ready_modular_monolith/)
- [Gridstack.js: Interactive dashboards](https://gridstackjs.com/)
- `.planning/research/ausschreibungs-portale-feasibility.md` — project-internal, HIGH confidence (already-verified portal-by-portal capability matrix, native alert features, ToS constraints, DÖE/TED coverage stats)
- [Open Contracting Data Standard — Codelists & Schema Reference](https://standard.open-contracting.org/latest/en/schema/codelists/) — CPV classification usage, deadline/timezone handling, currency-attached values (MEDIUM-HIGH, official standard docs)
- [Open Contracting Data Standard — official site](https://www.open-contracting.org/data-standard/) — schema/building-blocks overview (MEDIUM-HIGH)
- [Stotles — Tender Alerts & Tracker](https://www.stotles.com/platform/track-tenders) — unified feed from 1,000+ portals, Signal Score relevance ranking, saved searches (MEDIUM, vendor marketing but consistent with independent sources)
- [Tendium — Tender Monitoring](https://tendium.ai/en/tender-monitoring/) — monitoring/alert pattern confirmation (MEDIUM)
- [TenderAlerts.eu](https://tenderalerts.eu/) — EU tender aggregation, notification dashboard pattern (MEDIUM)
- [EU Tenders Monitor (Apify)](https://apify.com/nicolas_izquierdo/eu-tenders-monitor) — explicit confirmation that "relevance scoring, deduplication and new-only alerts" are treated as a standard feature set for this category (MEDIUM)
- [deepbloo — What Is a Tender Monitoring Platform](https://deepbloo.com/blog-posts/what-is-a-tender-monitoring-platform-a-complete-guide-to-tools-for-public-procurement-intelligence) — category definition, deduplication described as critical (MEDIUM)
- Codebase inspection: `apps/api/src/dkv/providers/inbox-provider.interface.ts`, `dkv-scheduler.service.ts`, `apps/api/src/mail/mail.service.ts`, `apps/api/src/mail/mail.module.ts` — confirms reusable infra for email-ingestion, scheduling, and outbound mail (HIGH — direct codebase read)
---
*Feature research for: German public-procurement tender monitoring/aggregation module (Ausschreibungs-Radar)*
*Researched: 2026-07-17*
+384 -2
View File
@@ -2,7 +2,9 @@
**Domain:** Modular portal platform with marketplace, multi-tenancy, Docker deployment
**Project:** Tessera
**Researched:** 2026-06-18
**Researched:** 2026-06-18 (v1.0 platform pitfalls below); 2026-07-17 (v1.1 Ausschreibungs-Radar module pitfalls in the dedicated section further down)
> This file accumulates pitfalls research across milestones. The v1.0 section below covers platform-wide architecture pitfalls (multi-tenancy, module system, Docker, i18n, etc.) and remains valid for all subsequent modules built on Tessera. The **v1.1 Ausschreibungs-Radar** section adds pitfalls specific to building a multi-source tender-aggregation module (scraping, eForms/OCDS ingestion, dedup, notifications) on top of that platform.
## Critical Pitfalls
@@ -267,7 +269,7 @@ Mistakes that cause rewrites, data breaches, or architectural dead-ends.
| Deployment | Volume permissions break non-root containers | Named volumes, entrypoint permission scripts |
| VCS integration | Gitea-specific tight coupling | Abstract behind interface, use standard Git ops |
## Sources
## Sources (v1.0 platform pitfalls)
- [Multi-Tenant SaaS Architecture: What Nobody Tells You Before You Build](https://dev.to/actinode/multi-tenant-saas-architecture-what-nobody-tells-you-before-you-build-a4h) - Confidence: HIGH
- [Designing Multi-Tenant SaaS Architecture: Mistakes to Avoid](https://www.saasadviser.co/blog/multi-tenant-saas-architecture-mistakes-best-practices) - Confidence: MEDIUM
@@ -280,3 +282,383 @@ Mistakes that cause rewrites, data breaches, or architectural dead-ends.
- [Electron vs. Tauri](https://www.dolthub.com/blog/2025-11-13-electron-vs-tauri/) - Confidence: MEDIUM
- [Building Interactive Dashboards with React Grid Layout](https://www.ilert.com/blog/building-interactive-dashboards-why-react-grid-layout-was-our-best-choice) - Confidence: MEDIUM
- [AWS: Multi-tenant data isolation with PostgreSQL Row Level Security](https://aws.amazon.com/blogs/database/multi-tenant-data-isolation-with-postgresql-row-level-security/) - Confidence: HIGH
---
# v1.1 Milestone: Ausschreibungs-Radar — Multi-Source Tender Aggregation Pitfalls
**Domain:** Multi-source tender/procurement aggregation module (scraping + eForms/OCDS ingestion + email-alert ingestion + normalization + dedup + notification), built as a new NestJS/Prisma/PostgreSQL-RLS module on the existing multi-tenant Tessera platform.
**Researched:** 2026-07-17
**Confidence:** HIGH (grounded in `.planning/research/ausschreibungs-portale-feasibility.md`, official OCDS-for-eForms docs, DÖE OpenData API docs, existing Tessera codebase patterns — `dkv-scheduler.service.ts`, `tenant.guard.ts`, `DkvModuleConfig` schema — and German scraping case law). MEDIUM on exact DÖE pagination/rate-limit behavior (Swagger is JS-rendered, not yet live-verified per the feasibility doc's open verification points).
## Critical Pitfalls
### Pitfall 15: HTML scraper fragility — session tokens, jsessionid, layout drift
**What goes wrong:**
AI-AG NetServer (lhs-vpbw, tender24, vergabe.landbw) and cosinex VMP (DTVP) both gate the "public search" behind server-side session state — `jsessionid` in the URL or a hidden CSRF/viewstate token in the search form that must be replayed on every paginated request. A scraper that treats these as static query params breaks the moment the portal rotates the session, adds a token, or reflows the HTML (even a CSS-only redesign can shift selectors). Because both platforms are proprietary, undocumented, and outside Tessera's control, there is no changelog or deprecation notice — the adapter just silently starts returning zero results or garbage.
**Why it happens:**
Scrapers are built once against a snapshot of the DOM/session flow and treated as "done." Portal vendors change markup, add bot-detection (rate-based session invalidation), or migrate frontend frameworks without any notice to third parties, because scraping was never a supported integration path.
**How to avoid:**
- Build a thin `PortalAdapter` interface (fetch session → search → paginate → parse detail) per platform (one AI-AG adapter, one cosinex adapter — per the feasibility doc's Effort/Value ranking), not per portal instance, so a fix in one place covers 4/8/9 (AI-AG) or all cosinex marketplaces.
- Never hardcode a session token's lifetime — always re-derive it from the search page response on each scrape run, don't cache it across runs.
- Store raw HTML of the search + detail pages for the last N successful runs (short retention, not permanent) so a layout-change diagnosis doesn't require reproducing the failure live.
- Add a structural health check per adapter: assert on stable anchors (e.g., "did we get >0 rows AND did known static fields — DE, umlauts, CPV-looking codes — parse") before treating a scrape as successful; a scraper that "succeeds" with zero rows for days is a silent failure, not a quiet success.
**Warning signs:**
Result count drops to zero (or spikes to an implausible number) for a portal that previously returned steady volume; parse errors on fields that were previously stable; HTTP 200 responses with unexpected redirect chains (session expiry manifests as a redirect to a login/error page, not a 4xx).
**Phase to address:**
Portal Adapter phase (AI-AG + cosinex scrapers) — build the adapter interface with health-check-on-every-run baked in from the first adapter, not bolted on later.
---
### Pitfall 16: Legal/ToS risk — scraping portals with explicit automated-access bans
**What goes wrong:**
vergabe24.de and plattform.aumass.de have AGB clauses that explicitly forbid automated extraction/scripted access (vergabe24 even names rate limits in its ToS). German case law (BGH, 30.04.2014 — I ZR 224/12) shows scraping itself is not per se illegal and a "virtuelles Hausrecht" has no independent legal basis — but an *effectively incorporated* AGB prohibition, combined with any technical protection measure (bot detection, CAPTCHA, rate limiting) the portal has in place, shifts the analysis toward Wettbewerbsrecht (§ 3a UWG — Rechtsbruch) and potential Datenbankherstellerrecht (§ 87a UrhG) claims, since both portals invest in curating/aggregating tender data as their core product. Building against these two anyway (or extending a generic scraper framework to "just try" them later) creates real legal exposure that has nothing to do with code quality.
**Why it happens:**
A generic `PortalAdapter` abstraction makes it *technically* trivial to add a new source once the interface exists — the temptation to "just add vergabe24, we already have the crawler" bypasses the legal review that should gate it.
**How to avoid:**
- Hard exclusion, enforced in code, not just documentation: maintain an explicit denylist/allowlist of source identifiers in config (not just a comment), and have the adapter registry refuse to register an adapter for a denylisted portal id even if someone writes the code.
- Their Oberschwelle notices are already covered via DÖE (per the feasibility doc), so there is no data-completeness reason to scrape them — document this rationale next to the denylist so a future contributor doesn't "rediscover" the idea without the legal context.
- If Unterschwelle coverage from these two ever becomes a business requirement, the only acceptable path is the same one already used for portals 2/3/4/6/7/8/9/10: register a native saved-search + ingest the resulting alert email via the existing DKV inbox infrastructure — not HTML scraping.
**Warning signs:**
A PR adds a new adapter or config entry referencing `vergabe24` or `aumass` in any capacity beyond "denylisted"; someone proposes a "generic connector" that accepts an arbitrary portal URL from tenant admins (this would let a customer point the scraper at a banned portal without Tessera's own code ever naming it).
**Phase to address:**
Portal Adapter phase — encode the denylist as a first-class artifact (config + registry guard + test) before any adapter ships, so it's structurally impossible to "just add" a banned source. Revisit only via explicit product decision, not incidentally.
---
### Pitfall 17: eForms-DE / SDK schema version drift breaks XPath and field mapping
**What goes wrong:**
eForms-DE is based on the KoSIT eForms SDK, which the EU Publications Office revises periodically (new SDK majors/minors introduce renamed elements, new mandatory fields, or restructured repeatable groups). A parser hardcoded against one SDK version's XPath will silently drop fields — or throw on notices published under a newer SDK — the moment DÖE starts forwarding notices tagged with the new version. Because DÖE re-publishes notices in the *exact* schema version the buyer's e-Sender submitted, a single ingestion run can contain a mix of SDK versions.
**Why it happens:**
Developers build the XML parser against a handful of sample notices at build time and never revisit it; the SDK version is often only visible in a namespace/version attribute deep in the document, easy to ignore until it breaks something.
**How to avoid:**
- Prefer OCDS (`ocds-mnwr74`) over raw eForms-DE XML wherever DÖE offers both — the OCDS profile's official field mappings (open-contracting-extensions/eforms on GitHub) are versioned and maintained upstream, absorbing most SDK churn for you.
- If eForms-DE XML is parsed directly (e.g., for fields OCDS doesn't map), read and log the SDK version attribute on every notice and fail loud (not silently drop) on an unrecognized version rather than best-effort parsing it.
- Keep a small fixture library of real notices per encountered SDK version for regression tests — this is a case where "test against production data snapshots" beats synthetic fixtures, because the actual failure mode is structural drift you can't anticipate.
**Warning signs:**
A sudden spike in notices with empty/null fields that used to populate; parser exceptions correlating with a specific publication date range (SDK version cutovers happen on fixed EU-mandated dates); OCDS `previouslyWithheldInformation` releases not being picked up (per official docs, this is a self-managed process — TED handles scheduled release, but OCDS consumers must poll for the redaction-lift date themselves).
**Phase to address:**
DÖE/OCDS Ingestion phase — build the SDK-version-aware ingestion pipeline (log + fail loud on unknown versions) as part of the first central-feed integration, since this is the highest-ROI source per the feasibility doc and will run continuously from day one.
---
### Pitfall 18: OCDS optional/conditional fields treated as "always present"
**What goes wrong:**
The OCDS eForms profile has fields that are conditionally present based on procedure type, threshold, or procurement stage (e.g., award fields don't exist until an award release; framework-agreement cascade awards look like multiple suppliers on one award but must be distinguished from joint awards by checking `lot.techniques.frameworkAgreement` + bid-ranking presence, per the official "how to use" guide). A normalizer written against a handful of "happy path" above-threshold notices will crash or silently null out data for the many below-threshold / early-stage notices that don't yet have those fields.
Additionally, `OrganizationReference` objects (buyer, tenderer, supplier) in OCDS only contain an `id` by default — the human-readable `.name` must be resolved by cross-referencing the `parties` array. Skipping this step produces a UI full of opaque org IDs instead of buyer names, which will look broken to end users even though ingestion "succeeded."
**Why it happens:**
OCDS is release-based (multiple releases per contracting process, merged into a "record") — a normalizer that only looks at the latest release, or treats every field as always-populated, misses this structure.
**How to avoid:**
- Treat every OCDS field beyond `id`/`title`/`tender.status` as optional in the normalized schema; write the normalizer defensively (missing ≠ error, just means "not yet known at this stage").
- Explicitly implement the `parties[].name` resolution step for every `OrganizationReference` location (buyer, tenderers, suppliers, procuringEntity) before display — this is a documented, mandatory post-processing step, not an edge case.
- Decide up front whether Tessera consumes the OCDS "release" or "record" package (record = pre-merged current state, generally simpler for a read-only aggregator than reconciling releases yourself).
**Warning signs:**
UI shows organization IDs (UUID-looking strings) instead of names; award/value fields blank for tenders still in "planning" or "tender" stage that should legitimately have no award data yet vs. tenders where the data genuinely failed to parse — these two cases must be distinguishable in the schema (e.g., a `parseWarnings` field), not conflated.
**Phase to address:**
DÖE/OCDS Ingestion phase and Normalization phase — build the normalized internal schema to make "field not yet applicable at this stage" and "field failed to parse" distinct states from the start; retrofitting this distinction after the UI already treats blanks as identical is expensive.
---
### Pitfall 19: CPV code format inconsistencies break category filtering
**What goes wrong:**
CPV (Common Procurement Vocabulary) codes appear in different shapes across sources: full 8-digit + check-digit form (`45000000-7`) in eForms/TED, sometimes truncated or zero-padded differently in scraped HTML (portals often display only the human label, e.g. "Bauleistungen", without the code at all), and CPV is hierarchical (a filter on "45xxxxxx — Bauarbeiten" should match all more-specific codes underneath it). A filter engine built against exact-string CPV matches will miss the majority of relevant results because most notices carry a specific leaf code, not the parent category a user searched for.
**Why it happens:**
CPV's hierarchy (division → group → class → category → subcategory, encoded positionally in the 8 digits) isn't obvious from a single sample notice, and portals that don't expose CPV at all (many of the AI-AG/cosinex HTML listings show only free-text category labels) force a fallback keyword-matching path that behaves completely differently from the CPV-code path.
**How to avoid:**
- Normalize CPV to the canonical 8-digit+check-digit string (strip formatting, validate against the official CPV code list) in the ingestion layer, never in the filter/UI layer.
- Implement hierarchical CPV matching (prefix match on the first N significant digits) as the actual filter semantics, not exact match — expose this in the filter UI as "category" (broad) vs. exact code.
- For sources without native CPV (HTML-only portals), map their free-text category taxonomy to CPV divisions in a small lookup table rather than pretending they're equivalent to code-based filtering — flag these results as "category: approximate" in the schema so users understand the precision difference.
**Warning signs:**
A CPV-based saved search returns near-zero results despite users reporting matching tenders exist; filter results differ wildly in volume between DÖE-sourced (CPV-tagged) and portal-scraped (label-only) records for the same category.
**Phase to address:**
Normalization phase (CPV canonicalization + hierarchy) and Filter Engine phase (hierarchical matching semantics) — the canonical CPV table should be seeded once, early, since it's static reference data all sources map into.
---
### Pitfall 20: Treating below-threshold coverage as complete ("false completeness")
**What goes wrong:**
Per the feasibility research, DÖE/TED guarantee ~100% of *above-threshold* (Oberschwelle) notices (~12% of procedures by count, ~75% by value) but only an estimated 20–35% of *below-threshold* (Unterschwelle) notices by count today, since Unterschwelle eForms publication is only mandatory for Bund and voluntary elsewhere. If the product doesn't clearly communicate this gap, a tenant configuring a saved search sees "0 results" or a thin result set for their region/CPV and reasonably concludes either "nothing matches" or "the module is broken" — when the real answer is "this data isn't centrally available yet, only reachable via the individual portal (which may not even be in Tessera's covered set)."
**Why it happens:**
The DÖE API looks and feels complete (clean OCDS/eForms structure, no visible "gaps" in the data itself) — there's no signal in the API response that tells you what's *missing*, only what's present.
**How to avoid:**
- Explicitly model source coverage per source in the schema/UI: each normalized tender record carries its source(s) and, more importantly, each saved search result set is annotated with which sources were queried and their known coverage tier (central-guaranteed vs. portal-partial vs. not-covered).
- Surface this in the UI ("Diese Suche deckt DÖE (Oberschwelle, vollständig) + 3 Portale (Unterschwelle, teilweise) ab — für vollständige Unterschwellen-Abdeckung: X weitere Portale nicht angebunden") rather than presenting a unified, seemingly-authoritative result list.
- Track this as a living fact, not a one-time note — the "20–35%" figure is explicitly stated to be *increasing* as the Unterschwellen-Pflicht rolls out, so hardcoding "this data is incomplete" copy without a mechanism to update it will itself become stale/misleading.
**Warning signs:**
Support requests along the lines of "why didn't I get an alert for a tender I found manually on a portal you claim to cover" — this is a coverage-transparency failure, not necessarily a bug.
**Phase to address:**
Normalization phase (source/coverage metadata on every record) and UI/Filter Engine phase (surfacing coverage transparency) — should ship with the MVP result list, not be added reactively after user confusion.
---
### Pitfall 21: Cross-source deduplication — same tender, different IDs
**What goes wrong:**
A single tender can legitimately appear via multiple paths: the buyer's e-Sender pushes eForms-DE to DÖE (→ also mirrored to TED), the same tender is *also* listed natively on the buyer's chosen portal (DTVP, an AI-AG portal, etc.) with a portal-internal ID, and if the tenant also has an email-alert saved search on that portal, the same tender arrives a third time via inbox ingestion. Each path uses a different identifier scheme (DÖE/TED notice number vs. portal-internal Vergabenummer vs. whatever the buyer typed in the alert email subject), different field completeness, and often slightly different publication timestamps (portal listing can precede or lag the central eForms push by hours to days). A naive "unique by ID" dedup produces 2–3 duplicate cards for the same real-world procurement; a naive "unique by title" produces false merges of genuinely different tenders with similar names (common for framework/rebid procedures).
**Why it happens:**
There is no shared cross-portal identifier for German procurement below the eForms/TED notice number, and even that number isn't always echoed back verbatim on the originating portal's HTML page.
**How to avoid:**
- Build a fingerprint-based dedup, not ID-based: normalize (buyer name, CPV/category, region/PLZ, deadline date, and a fuzzy-matched title) into a composite key; use fuzzy string matching (e.g., trigram similarity) with a confidence threshold, not exact equality, on the title component.
- When the eForms/TED notice number *is* present in a scraped or emailed record (often referenced as "Bekanntmachungs-ID" or similar), treat it as a strong signal but not the sole key — validate it against the fuzzy fingerprint before merging, since transcription errors happen.
- Prefer the DÖE/eForms record as the canonical source of truth when a duplicate is detected (richest structured data); merge portal- and email-sourced duplicates into it as "also seen on X" rather than discarding them — this also gives you a natural cross-check for the coverage-transparency feature (Pitfall 20).
- Store the dedup decision (merged-from IDs, confidence score) so a human can review/undo false merges — this is an area where silent auto-merge will eventually be visibly wrong to a user who tracked a tender manually.
**Warning signs:**
Tenants report seeing "the same tender twice" or, conversely, ask why a tender they know overlaps with another wasn't flagged as related; dedup confidence scores clustering near the threshold boundary (indicates the fingerprint algorithm needs tuning, not that the data is ambiguous).
**Phase to address:**
Normalization & Dedup phase — this deserves to be its own phase (not folded into ingestion), since it's the piece most likely to need iteration after real multi-source data is flowing and edge cases surface.
---
### Pitfall 22: Deadline/timezone handling for "only still-open" filtering
**What goes wrong:**
eForms/OCDS deadline fields are typically ISO 8601 with explicit timezone (UTC or CEST/CET offset), but scraped portal HTML frequently shows local dates/times without explicit timezone, in German format (`DD.MM.YYYY, HH:MM Uhr`), and email-alert bodies vary by portal template. A filter that does naive string-to-Date parsing without normalizing to a single timezone will misjudge "still open" near midnight boundaries and across DST transitions (CEST↔CET) — either showing an already-closed tender as open (embarrassing if a tenant relies on it) or hiding a still-open one. Deadlines expressed only as a date (no time) also need an explicit convention (end-of-day in which timezone?) since German procurement deadlines are almost always "Datum, HH:MM Uhr" — treating a date-only field as "midnight UTC" silently shortens the effective window by hours.
**Why it happens:**
Timezone bugs are invisible in testing unless tests specifically straddle DST transitions or midnight-local boundaries; developers default to `new Date(string)` parsing without checking what timezone the source actually meant.
**How to avoid:**
- Normalize all deadlines to UTC at ingestion time, with the source timezone made explicit per source (DÖE/eForms: use the timezone offset in the ISO string as-is; scraped portals: assume Europe/Berlin unless the portal states otherwise, and record which assumption was applied).
- "Still open" filtering must compare against `now()` in UTC, not local server time — verify the deployment container's timezone doesn't leak into date math (Node's `Date` object is UTC-internal but naive string parsing of ambiguous local strings is where bugs live).
- Test explicitly across a DST transition date and around a deadline that falls exactly at day-boundary local time.
**Warning signs:**
A tender flips between "open"/"closed" status on page refresh near its deadline (indicates inconsistent timezone handling between where filtering happens vs. where display happens); users report a "still open" tender they click into is actually already closed on the source portal.
**Phase to address:**
Normalization phase (UTC normalization at ingestion) and Filter Engine phase (open/closed comparison logic) — cover with explicit DST/midnight-boundary test cases before the filter engine phase is marked done.
---
### Pitfall 23: Notification storms — first-run backfill and duplicate/immediate-alert flooding
**What goes wrong:**
Two related failure modes: (a) on a tenant's first saved search, DÖE alone can return months of matching historical notices — if the notification pipeline treats "newly matched by this search" the same on day one as it does on day two, the tenant's first experience is an inbox flooded with hundreds of "new tender" emails instead of a clean digest; (b) once running, a tender that gets *updated* (deadline extension, correction notice) re-appears in the source feed and, if the notification logic keys only on "is this a new match" rather than "have I already notified this tenant about this specific tender," triggers a second immediate alert for something they already saw — training users to ignore or unsubscribe from alerts entirely.
**Why it happens:**
"New" is ambiguous — new to the source feed vs. new to this tenant's search vs. never-notified-before are three different conditions that get conflated when the notification trigger is implemented as "insert into results table → fire notification" without a separate notified-state tracking table.
**How to avoid:**
- Separate "matched" from "notified": every (tenant, saved-search, tender) match is recorded, but notification firing is a distinct step gated by explicit backfill handling — on first activation of a saved search, mark all currently-matching historical results as seen/backfilled *without* notifying, and only notify going forward for genuinely new matches (or start the digest from "activation time," clearly communicated to the user).
- Track per-tender "last notified version/hash" so an update to an already-notified tender triggers, at most, a distinct "updated" notification (clearly differentiated from "new"), not a duplicate "new tender" alert — and make this configurable (some tenants want deadline-extension alerts, most don't want to see the same tender twice).
- Default to digest (periodic batch) rather than instant alerts for saved searches with a strong deadline signal that they'll otherwise return many historical hits (e.g., broad CPV + wide region); reserve instant "Sofort-Alert" for narrow, precise searches where volume is naturally low — this should be a UX default/nudge, not just a raw toggle.
- Rate-limit/batch outbound email regardless — even a legitimately large first-run result set should render as one digest email with N results, never N individual emails.
**Warning signs:**
A tenant's first day includes an abnormal spike in outbound emails compared to steady-state; support complaints about "getting the same tender email again"; email provider (existing DKV SMTP infra) flags or throttles the sending account for burst volume.
**Phase to address:**
Notification phase — the matched/notified separation and backfill-suppression logic must be designed before any saved search goes live, since retrofitting it after tenants have already been flooded once is a trust problem, not just a technical one.
---
### Pitfall 24: Copying the DKV scheduler's single-tenant pattern for the tender module
**What goes wrong:**
The existing `DkvSchedulerService` (`apps/api/src/dkv/dkv-scheduler.service.ts`) is explicitly documented as v1 single-tenant: it loads config via `findFirst()` and runs one cron job for `activeTenantId`, with a code comment stating "multi-tenant scheduling (one cron job per active tenant) is deferred to a future plan." If the tender module's polling scheduler (for portal adapters, email-alert ingestion, DÖE polling) is built by copy-pasting this pattern, only one tenant's saved searches will ever actually poll in a multi-tenant deployment — every other tenant's saved searches will silently never run, with no error, because the scheduler never even looks for their config.
**Why it happens:**
The DKV scheduler is the most recent, most similar in-repo reference implementation for "polling infra + cron job management" — it's the natural template to copy, and its single-tenant limitation is documented in a code comment that's easy to miss when skimming for the cron-registration pattern, not the architecture note.
**How to avoid:**
- Design the tender module's scheduler as N cron jobs (or a single dispatcher cron that iterates all active tenant configs) from the start — Tessera's multi-tenancy is an explicit from-day-one architectural constraint (per PROJECT.md), and a new module regressing to single-tenant scheduling would be a step backward, not a shortcut.
- Reuse the *mechanics* of `SchedulerRegistry.addCronJob()` / dynamic interval updates from the DKV pattern (these are sound), but replace `findFirst()` with `findMany({ where: { isActive: true } })` and register one job per tenant (or per saved-search, depending on granularity chosen), keyed by a job name that includes the tenant/search ID.
- Add an explicit multi-tenant scheduler test (two tenants, two configs, both must poll independently) as an acceptance criterion for this phase — don't rely on manual review to catch a `findFirst()` regression.
**Warning signs:**
A second tenant activates a saved search and never receives results/notifications despite valid config; scheduler logs only ever mention one tenant ID across a multi-tenant deployment.
**Phase to address:**
Scheduler/Ingestion Orchestration phase — explicitly call out "multi-tenant, not single-tenant like DKV v1" in the phase's acceptance criteria, since this is a known, named trap already present in the codebase.
---
### Pitfall 25: Multi-tenant isolation gaps for saved searches, results, and stored portal credentials
**What goes wrong:**
Three distinct isolation surfaces need to hold: (a) saved-search *configuration* (keywords, CPV, region filters) must be tenant-scoped — a leak here exposes one tenant's business intelligence (what they're bidding on) to another; (b) *results* (which tenders matched which tenant's search, and any tenant-specific annotations/status like "we're bidding on this") must be tenant-scoped even though the underlying tender data itself (from DÖE/portals) is shared, public, non-tenant-specific reference data — the natural bug here is applying tenant filtering to the shared tender catalog (wrong) instead of to the join table between tenant-saved-searches and tenders (right); (c) any stored credentials for portal saved-search registration (if the module ever needs a portal login to register a saved search on a tenant's behalf) must never be shared across tenants even if two tenants use the same portal.
**Why it happens:**
The existing `TenantGuard` pattern scopes `req.tenantPrisma` correctly for request-scoped controller calls, but background jobs (scheduler-triggered ingestion, notification dispatch) run outside any HTTP request context — there is no `req` to derive `tenantId` from, so it's easy to accidentally use the raw (non-tenant-scoped) `PrismaService` for a background operation and forget to add the `tenantId` filter by hand, especially on the *shared reference data* tables where "no tenant filter" is often correct (the tender catalog itself) and it's easy to reflexively skip tenant filtering on the *adjacent* tables that do need it (saved searches, match results, credentials).
**How to avoid:**
- Model the schema with a clear split: tender/notice records (shared, no `tenantId`) vs. saved-search, match-result, notification-log, and portal-credential records (all `tenantId`-scoped, indexed, following the `DkvModuleConfig` precedent of `tenantId String @unique` per-config or `tenantId String` + `@@index([tenantId])` for multi-row tables).
- In background/cron contexts, explicitly call `forTenant(prisma, tenantId)` (the existing extension used by `TenantGuard`) per-tenant iteration — never fall back to the raw `PrismaService` for tenant-scoped tables just because there's no request object.
- If portal credentials are ever stored (for saved-search registration on AI-AG/cosinex portals), reuse the exact `encryptedInboxCreds` AES-256-GCM pattern from `DkvModuleConfig` — don't invent a new credential-storage mechanism for this module.
- Add a cross-tenant isolation test as a standard phase gate: two tenants, overlapping saved-search keywords, verify tenant A never sees tenant B's saved-search config, match annotations, or receives tenant B's notifications.
**Warning signs:**
A query joins the shared tender table directly to a tenant-scoped table without an explicit `tenantId` predicate on the join; background job code imports `PrismaService` directly instead of iterating tenants and using `forTenant()`.
**Phase to address:**
Multi-Tenant Saved Searches phase (schema + isolation tests) — should be verified with an explicit two-tenant UAT before the Notification phase ships, since notification is the surface where a leak becomes user-visible (wrong tenant's alert email).
---
### Pitfall 26: Scheduler/rate-limit — polling too aggressively and getting throttled or blocked
**What goes wrong:**
With potentially many tenants each running saved searches against the same underlying portal adapters (AI-AG, cosinex) or the same DÖE API, a naive per-tenant-per-search polling design multiplies request volume against a small number of actual upstream endpoints — e.g., 50 tenants each polling the same AI-AG adapter every 15 minutes doesn't mean 50x the useful data, it means 50x the load on one portal for redundant queries, risking IP-based rate limiting, temporary bans, or (worse) inviting the exact "AGB verbietet Skripte, nennt Raten" scrutiny the feasibility research flagged for vergabe24 — even on portals without an explicit ban, aggressive undifferentiated polling looks like abuse.
**Why it happens:**
Scheduling is naturally designed per-tenant (each tenant's saved search has its own poll interval preference), but the underlying data source is shared infrastructure — without a dedup/coalescing layer, the scheduler design conflates "how often does tenant X want fresh results" with "how often should we actually hit the upstream portal."
**How to avoid:**
- Decouple polling from tenant preference: poll each upstream source (DÖE, each AI-AG instance, cosinex, each RSS feed) on a single shared schedule per source (not per tenant), then fan the results out to all matching tenant saved searches from the ingested/normalized data — this also directly fixes the redundant-load problem and is a natural extension of the "one AI-AG adapter, one cosinex adapter" architecture already recommended.
- Respect explicit rate signals: HTTP `Retry-After` headers, documented portal rate limits (vergabe24's AGB explicitly names a rate — even though vergabe24 itself is excluded, treat this as a signal that similar limits likely apply industry-wide), and add jitter/backoff on repeated failures rather than fixed-interval retry that can synchronize into a thundering-herd against a recovering endpoint.
- For DÖE/TED (the auth-free, ToS-friendly central APIs), still poll conservatively (e.g., hourly, not per-minute) — there's no completeness benefit to sub-hourly polling of a feed that itself batches publications, and it needlessly increases the chance of being deprioritized or rate-limited by a public-good API relied on by many consumers.
- Verify DÖE's actual pagination/rate-limit behavior against the live Swagger UI before finalizing the poll cadence — the feasibility doc flags this as an open verification point, not yet confirmed.
**Warning signs:**
HTTP 429s or connection resets from a portal correlating with polling frequency increases; DÖE/TED response times degrading specifically during Tessera's poll windows; multiple tenants' saved searches against the same portal triggering independent, uncoordinated scrape runs within the same minute.
**Phase to address:**
Scheduler/Ingestion Orchestration phase — the "poll source once, fan out to tenants" architecture is a foundational design decision for this phase, not an optimization to add later; retrofitting it after per-tenant polling ships means migrating live saved searches without disrupting notifications.
## Technical Debt Patterns (v1.1)
| Shortcut | Immediate Benefit | Long-term Cost | When Acceptable |
|----------|-------------------|-----------------|------------------|
| Hardcode portal HTML selectors per portal instance instead of a shared AI-AG/cosinex adapter interface | Faster first working scraper | Every one of 4/8/9 AI-AG portals needs its own fix when the platform changes markup | Never — the feasibility doc's whole ROI case for scraping rests on one adapter covering many portal instances |
| Store only parsed/normalized fields, discard raw scraped HTML | Less storage | No way to diagnose a broken scraper without reproducing live against a moving target | Only if a short-retention raw-response cache is added later before it's actually needed for a live incident |
| Per-tenant polling of shared upstream sources (copy DKV's per-tenant cron mental model) | Simpler initial scheduler code | Redundant load, throttling risk, exactly the pattern flagged in Pitfall 26 | Never for shared sources (DÖE/portals); acceptable only for genuinely tenant-specific sources (a tenant's own email inbox) |
| Exact-string CPV/title matching instead of hierarchical/fuzzy matching | Simpler filter/dedup code | Filter results miss most relevant tenders (Pitfall 19); dedup produces visible duplicates (Pitfall 21) | Acceptable only as an explicitly-labeled MVP limitation with a visible roadmap item, not silently shipped as "the filter" |
| Skip the matched/notified separation, fire notification directly on new DB row insert | Faster to build first notification | Backfill flood on every new saved search (Pitfall 23) | Never — this is cheap to build correctly from the start and expensive to fix after users have already been flooded once |
| Use raw `PrismaService` in scheduler/cron code instead of `forTenant()` per tenant iteration | Slightly less boilerplate | Silent cross-tenant data leakage in background jobs (Pitfall 25) | Never for any tenant-scoped table |
## Integration Gotchas (v1.1)
| Integration | Common Mistake | Correct Approach |
|-------------|-----------------|-------------------|
| DÖE OpenData API (eForms/OCDS/CSV) | Treating it as 100% complete tender coverage | Explicitly model coverage tier (Oberschwelle guaranteed, Unterschwelle partial ~20-35%) in schema/UI (Pitfall 20) |
| DÖE OCDS (`ocds-mnwr74`) | Reading fields without resolving `OrganizationReference.name` against `parties[]` | Implement the documented post-processing step for all 9 org-reference locations before display |
| eForms-DE raw XML | Hardcoding one SDK version's XPath | Read/log SDK version per notice, fail loud on unrecognized versions, keep per-version fixtures |
| AI-AG NetServer (lhs-vpbw/tender24/vergabe.landbw) | Per-portal scraper, static session token caching | One shared adapter class, re-derive session/CSRF token every run |
| cosinex VMP (DTVP) | Assuming cosinex markup is compatible with AI-AG markup | Separate adapter — the feasibility research confirms the two platforms are HTML-incompatible |
| subreport-elvis / service.bund.de | Scraping HTML instead of using their native RSS feeds | Use RSS — lower ToS risk, cheaper to parse, already the recommended approach |
| vergabe24 / aumass | Building a "generic" adapter that could technically reach them | Hard denylist enforced in the adapter registry, not just documentation (Pitfall 16) |
| DKV inbox infra (`ImapProvider`/`ExchangeInboxProvider`) reused for portal alert emails | Assuming alert email format/parsing is identical to DKV's existing sender filters | Build portal-specific alert-email parsers (subject/body formats differ per portal vendor) reusing only the inbox *transport*, not the DKV parsing logic |
| TED API v3 | Polling it as a primary source for DE-only coverage | Treat as optional EU redundancy per the feasibility doc — largely redundant to DÖE for Germany |
## Performance Traps (v1.1)
| Trap | Symptoms | Prevention | When It Breaks |
|------|----------|------------|-----------------|
| Storing full raw HTML/XML indefinitely per scrape run | Database/disk growth outpaces useful data | Short retention window (e.g., last N runs or M days) on raw payloads, keep only normalized rows long-term | Noticeable within weeks at hourly polling across ~10 sources |
| Re-parsing full eForms XML on every filter query instead of caching normalized rows | Slow search/filter UI as tender volume grows | Normalize once at ingestion, query only the normalized table for filtering/UI | Becomes visible once historical backlog (months of DÖE data) is loaded |
| Per-tenant polling of shared sources (see Pitfall 26) | Upstream request volume scales with tenant count, not data freshness need | Poll-once-fan-out-many architecture | Breaks (throttling) well before thousands of tenants — even a few dozen tenants on the same AI-AG portal is enough |
| No archival/expiry of past-deadline tenders in the active result set | Filter/search UI slows as the table grows unbounded | Partition or flag expired tenders out of the default "open" query path; archive rather than delete for dedup/audit history | Gradual, but compounds — plan before the table becomes large enough to require a migration under load |
| Fuzzy-match dedup (Pitfall 21) run as O(n²) comparison across the full tender set | Ingestion job runtime grows non-linearly with catalog size | Bucket candidates first (by CPV + region + rough deadline window) before fuzzy title comparison within buckets | Noticeable once total tender volume reaches the thousands (a few months of DÖE + portal data) |
## Security Mistakes (v1.1)
| Mistake | Risk | Prevention |
|---------|------|------------|
| Storing portal login credentials (for saved-search registration) in plaintext or a custom encryption scheme | Credential leak grants access to a tenant's competitive procurement activity on a live portal account | Reuse the existing `encryptedInboxCreds` AES-256-GCM pattern from `DkvModuleConfig` verbatim |
| Logging scraper session tokens or full request/response on error (for debugging layout breaks) | Session/CSRF tokens or embedded credentials leak into log aggregation | Redact tokens/credentials before logging; log structural diagnostics (selector not found, row count) instead of raw payloads with secrets |
| Background/cron ingestion jobs using raw `PrismaService` instead of tenant-scoped access for tenant tables | Cross-tenant data leakage in saved searches, results, notifications (Pitfall 25) | Enforce `forTenant()` usage per tenant iteration in all scheduler code; add an isolation test as a phase gate |
| A future "custom portal URL" feature accepting tenant-supplied URLs for the scraper to fetch | SSRF — a malicious or compromised tenant admin points the scraper at internal infrastructure | If ever built, restrict to an explicit allowlist of known portal domains, never arbitrary tenant-supplied URLs |
| Notification emails including full source-portal deep links without validating they're outbound-safe | Low risk here, but email content built from scraped/parsed HTML fields (buyer name, title) without sanitization risks HTML injection in HTML-format alert emails | Sanitize/escape all scraped and eForms-derived text fields before interpolating into HTML email templates |
## UX Pitfalls (v1.1)
| Pitfall | User Impact | Better Approach |
|---------|-------------|-------------------|
| Presenting a unified result list without coverage transparency (Pitfall 20) | Users assume completeness, miss tenders on uncovered portals, lose trust when they find something manually that wasn't surfaced | Show per-search coverage summary (sources queried + tier) alongside results |
| Showing already-past-deadline tenders in the default "open" view due to timezone bugs (Pitfall 22) | Users waste time on tenders they can no longer bid on; erodes trust in the "still open" filter specifically | Rigorous UTC-normalized deadline filtering with explicit tests around DST/midnight |
| Instant-alert flood on first saved-search activation (Pitfall 23) | Users immediately mute/unsubscribe from a feature that could otherwise be valuable | Default to digest for broad searches; explicit backfill-suppression on activation |
| Displaying visible duplicate tender cards from multiple sources (Pitfall 21 failure mode) | Looks unpolished/broken, undermines confidence in the module's data quality | Merge with "also seen on: [sources]" badge instead of separate cards |
| No visibility into *why* a tender was filtered out or not matched | Users can't tune their saved search, assume the module is missing things arbitrarily | Provide a "why not matched" explainer or at least document filter semantics (hierarchical CPV, region matching) clearly in the UI |
| Raw German procurement jargon (Bekanntmachungsart, Vergabeart, Losaufteilung) with no glossary for non-specialist users | Users unfamiliar with procurement terminology can't interpret results | Add inline tooltips/glossary for domain terms in the detail view |
## "Looks Done But Isn't" Checklist (v1.1)
- [ ] **Portal adapter (AI-AG/cosinex):** Often missing a structural health check — verify it distinguishes "zero real results" from "scraper broke and returned zero" (Pitfall 15), and that it fails loud rather than silently.
- [ ] **eForms/OCDS ingestion:** Often missing SDK-version handling and `OrganizationReference.name` resolution — verify organization names render correctly and unknown SDK versions are logged, not silently mis-parsed (Pitfalls 17, 18).
- [ ] **CPV filtering:** Often only exact-matches CPV codes — verify a filter on a parent category (e.g., "45" — Bauarbeiten) returns child-code matches, not just exact 8-digit hits (Pitfall 19).
- [ ] **Dedup:** Often only tested against clean synthetic fixtures — verify against real cross-source pairs (same tender via DÖE + a scraped portal + an ingested alert email) with realistic ID/title/date variance (Pitfall 21).
- [ ] **Deadline filtering:** Often only tested at "normal" times — verify "still open" behavior across a DST transition and at a deadline that falls at local midnight (Pitfall 22).
- [ ] **Notification pipeline:** Often only tested with a handful of seed records — verify behavior when a saved search is activated against months of existing DÖE backlog (no flood) and when an already-notified tender is updated (no duplicate "new" alert) (Pitfall 23).
- [ ] **Scheduler:** Often copied from the DKV single-tenant `findFirst()` pattern — verify two tenants with independent active configs both actually poll and receive results (Pitfall 24).
- [ ] **Multi-tenant isolation:** Often only tested for the "happy path" single-tenant flow during development — verify with two tenants and overlapping search criteria that no cross-tenant data (saved searches, results, notifications, credentials) leaks (Pitfall 25).
- [ ] **Rate limiting:** Often untested until a portal actually throttles in production — verify the poll-once-fan-out-many architecture is actually in place before scaling tenant count, not just planned (Pitfall 26).
- [ ] **Excluded sources:** Often only documented, not enforced — verify the adapter registry structurally refuses to register vergabe24/aumass adapters, not just that no one has written one yet (Pitfall 16).
## Recovery Strategies (v1.1)
| Pitfall | Recovery Cost | Recovery Steps |
|---------|---------------|-----------------|
| Scraper broken by portal layout change (Pitfall 15) | MEDIUM | Diagnose via cached raw HTML from last successful runs; patch selectors; add the new structural variant to the adapter's health-check assertions so future drift of the same kind is caught faster |
| Accidental scraping attempt against a denylisted portal (Pitfall 16) | HIGH | Immediately disable the adapter, audit logs for request volume/duration against that portal, assess legal exposure, do not silently "fix" and continue — this needs explicit sign-off before any retry |
| eForms SDK version broke parsing (Pitfall 17) | LOW-MEDIUM | Add the new SDK version's field mapping/fixture, backfill-reparse affected date range from cached raw XML if retained, otherwise accept the gap and document it |
| Cross-tenant data leak discovered in background job (Pitfall 25) | HIGH | Treat as a security incident: identify affected tenants/records, patch the missing `forTenant()` scoping, audit all other scheduler code paths for the same pattern, notify affected tenants per data-protection obligations |
| Notification flood already sent to tenants (Pitfall 23) | MEDIUM | Send a brief clarifying follow-up (not another flood), fix the backfill-suppression logic, offer an easy re-subscribe/digest-preference change for anyone who unsubscribed in reaction |
| Portal starts throttling/blocking Tessera's scraper IP (Pitfall 26) | MEDIUM | Back off immediately (pause the adapter), fall back to that portal's native email-alert ingestion path if available (per the feasibility doc, most non-excluded portals do offer this), re-evaluate poll cadence before resuming |
## Pitfall-to-Phase Mapping (v1.1)
| Pitfall | Prevention Phase | Verification |
|---------|-------------------|----------------|
| 15. Scraper fragility (session/CSRF/layout drift) | Portal Adapter phase (AI-AG + cosinex) | Adapter has a structural health check; simulate a selector failure in tests and confirm it fails loud, not silent-zero |
| 16. Legal/ToS exclusion (vergabe24, aumass) | Portal Adapter phase | Adapter registry test asserts registration attempt for denylisted portal IDs is rejected |
| 17. eForms SDK version drift | DÖE/OCDS Ingestion phase | Parser logs/handles unknown SDK version explicitly; fixture tests cover ≥2 real SDK versions |
| 18. OCDS optional-field / OrganizationReference handling | DÖE/OCDS Ingestion + Normalization phase | Org names resolve correctly in UI for buyer/tenderer/supplier; "not applicable yet" vs "failed to parse" are distinguishable in schema |
| 19. CPV format/hierarchy | Normalization + Filter Engine phase | Filter on a parent CPV category returns child-code matches in a test fixture |
| 20. Below/above-threshold false completeness | Normalization + Filter/UI phase | UI displays coverage-tier annotation on every result set; documented and testable against the known ~20-35% Unterschwelle figure |
| 21. Cross-source dedup | Normalization & Dedup phase (own phase) | Fuzzy fingerprint dedup test against real cross-source sample pairs (DÖE + scraped portal + alert email for the same tender) |
| 22. Deadline/timezone handling | Normalization + Filter Engine phase | Explicit DST-transition and local-midnight test cases pass for "still open" filtering |
| 23. Notification storms / backfill flooding | Notification phase | Activation of a saved search against historical backlog produces zero immediate individual alerts; update-vs-new distinction tested |
| 24. Multi-tenant scheduler (DKV single-tenant regression) | Scheduler/Ingestion Orchestration phase | Two-tenant test: both tenants' active configs poll and produce independent results |
| 25. Multi-tenant isolation (saved searches, results, credentials) | Multi-Tenant Saved Searches phase | Two-tenant cross-isolation UAT: no leakage of config, results, or notifications across tenants |
| 26. Scheduler rate-limiting / aggressive polling | Scheduler/Ingestion Orchestration phase | Poll-once-fan-out-many architecture verified (single upstream request serves all matching tenant searches); backoff/jitter tested against simulated 429/Retry-After |
## Sources (v1.1 Ausschreibungs-Radar)
- `.planning/research/ausschreibungs-portale-feasibility.md` — portal-by-portal ToS/anti-bot classification, coverage percentages, DÖE/TED architecture, adapter consolidation recommendation (2026-07-16 research)
- [OCDS for eForms — How to use this profile](https://standard.open-contracting.org/profiles/eforms/latest/en/how/) — OrganizationReference resolution requirement, withheld-information handling, framework-agreement cascade caveat
- [OCDS for eForms — Field mappings / Schema / Codelists](https://standard.open-contracting.org/profiles/eforms/latest/en/) — official field mapping reference
- [open-contracting-extensions/eforms (GitHub)](https://github.com/open-contracting-extensions/eforms) — profile source, versioning
- [Beschaffungsamt — Datenservice Öffentlicher Einkauf](https://www.bescha.bund.de/DE/ElektronischerEinkauf/Datenservice_Oeffentlicher_Einkauf/Datenservice-Oeffentlicher-Einkauf_node.html) — DÖE service components, eForms-DE/OCDS/CSV export formats
- [DÖE OpenData Swagger UI](https://oeffentlichevergabe.de/documentation/swagger-ui/opendata/index.html) — API surface (pagination/rate-limit specifics not yet live-verified — open item per feasibility doc)
- [BGH, 30.04.2014 — I ZR 224/12 (screen scraping)](https://dejure.org/dienste/vernetzung/rechtsprechung?Gericht=BGH&Datum=30.04.2014&Aktenzeichen=I+ZR+224/12) — German case law on scraping/AGB/"virtuelles Hausrecht"
- [VOELKER & Partner — Screen-Scraping / Web-Crawler rechtliche Zulässigkeit](https://www.voelker-gruppe.com/kompetenzen/ip-it-stuttgart/beitraege/screen-scraping-web-crawler) — AGB incorporation requirements, § 3a UWG framing
- Existing Tessera codebase: `apps/api/src/dkv/dkv-scheduler.service.ts` (single-tenant scheduler precedent to avoid repeating), `apps/api/src/tenant/tenant.guard.ts` (tenant-scoping pattern via `forTenant()`), `apps/api/prisma/schema.prisma` `DkvModuleConfig` (credential encryption precedent to reuse)
---
*Pitfalls research for: Tessera (v1.0 platform) + Ausschreibungs-Radar (v1.1 multi-source German tender aggregation module)*
*Researched: 2026-06-18 (v1.0); 2026-07-17 (v1.1)*
+103 -4
View File
@@ -1,6 +1,101 @@
# Technology Stack
**Project:** Tessera - Modular Portal Platform with Marketplace
**Project:** Tessera — Modular Portal Platform with Marketplace
**Latest research:** 2026-07-17 (v1.1 Ausschreibungs-Radar delta) — base stack researched 2026-06-18
**Overall Confidence:** MEDIUM for the v1.1 delta below (versions cross-verified via WebSearch + direct npm registry lookup; no Context7/docs-MCP available this session), HIGH for the v1.0 base stack (unchanged, see bottom section)
---
# v1.1 Delta — Ausschreibungs-Radar Module
**Domain:** German public-procurement (Vergabe) tender aggregation — API client + HTML scraping adapters + RSS + XML/JSON normalization, on top of the existing NestJS/Prisma monorepo and the DKV module's inbox/mail/scheduler infrastructure.
This section is a **delta only**. It assumes the validated v1.0 stack below (pnpm/Turborepo, NestJS 11, Prisma 7/PostgreSQL 16, Next.js 16, Vitest) and covers only **new** libraries needed for this module. No new `apps/web` dependencies are required — the results UI (searchable list, filters, detail view) is built entirely with the existing Next.js/shadcn/TanStack Query/Zustand stack already in place.
## Recommended Additions
### Core Technologies (new)
| Technology | Version | Purpose | Why Recommended |
|------------|---------|---------|-----------------|
| **fast-xml-parser** | 5.10.1 | Parse eForms-DE XML (DÖE) + RSS 2.0 feeds (subreport-elvis, service.bund.de) | Zero-dependency, pure-TS, handles both jobs with one library (`XMLParser`/`XMLBuilder`/`XMLValidator`). Used by Microsoft, NASA, VMware. Faster and simpler than `xml2js` (callback/SAX-based, heavier). One parser covers eForms XML *and* RSS XML — see "What NOT to Add" for why a dedicated `rss-parser` is skipped. |
| **cheerio** | 1.2.0 | HTML parsing/traversal for the AI AG NetServer and cosinex Vergabemarktplatz scraping adapters | De-facto jQuery-style static HTML parser for Node — ~21k dependents, actively maintained (last release 2026-01-23). Wraps `parse5`/`htmlparser2`. 10–50× faster and far cheaper than a headless browser for server-rendered HTML, which both target portals produce for their public search-result pages. |
| **csv-parse** | 7.0.1 | Parse the DÖE OpenData CSV export as a fallback/cross-check format | `stream.Transform`-based, part of the `adaltas/node-csv` toolkit, ~3100 dependents, actively released (14 days old at research time). Ships dual CJS/ESM — safe for `apps/api`'s CommonJS build. Primary DÖE ingestion should use eForms-XML or OCDS-JSON (richer, structured); CSV is secondary/verification. |
| **tough-cookie** + **fetch-cookie** | 6.0.2 / 3.2.0 | Cookie-jar session handling for the scraping adapters (search-form POST → paginated results) | `apps/api` already standardizes on native `fetch` (see `favorites/icon-discovery.service.ts`, `calendar/providers/ics.provider.ts`) — no `axios` anywhere in the app. `fetch-cookie` wraps global `fetch` with a `tough-cookie` `CookieJar` so session cookies (e.g. `JSESSIONID` on the AI AG servlet portal) persist across a scrape run without introducing a second HTTP client family. |
| **playwright** | 1.61.1 (conditional — see note) | Headless-browser fallback adapter, only if a specific portal's search UI turns out to require client-side JS rendering | AI AG NetServer's `PublicationControllerServlet` URL pattern is a classic Java servlet MVC front-controller (JSP/form-based) — server-rendered HTML with a session cookie, **not** an SPA, and (unlike ASP.NET WebForms) has no ViewState/EventValidation mechanic to fight. cosinex markets Vergabemarktplatz as "fully browser-based" (marketing language, not confirmed SPA), making it the more likely candidate to need JS rendering. **Build both adapters cheerio-first; only pull in `playwright` for whichever adapter's public search page is confirmed (phase-1 spike) to require JS execution.** Don't add it speculatively — it downloads browser binaries, growing the Docker image ~300MB+, which the "wartbar/verständlich" constraint argues against unless proven necessary. |
### Supporting / Dev Tools (optional)
| Tool | Purpose | Notes |
|------|---------|-------|
| **undici** (devDependency only) | Mock outbound `fetch` calls in Vitest tests for the DÖE client and scraping adapters (`MockAgent` + `setGlobalDispatcher`) | Node's global `fetch` is already powered by undici under the hood; adding it as a devDependency gives official, zero-extra-runtime-cost HTTP mocking without introducing `nock`/`msw` (neither currently used anywhere in the repo). Add only when writing the adapter test suite. |
## Installation
```bash
# apps/api — core additions
pnpm --filter @tessera/api add fast-xml-parser cheerio csv-parse tough-cookie fetch-cookie
# apps/api — conditional, only if phase-1 spike confirms a JS-rendered portal
pnpm --filter @tessera/api add playwright
pnpm --filter @tessera/api exec playwright install chromium # only the one engine needed
# apps/api — dev/test only, optional
pnpm --filter @tessera/api add -D undici
```
## What NOT to Add — Reuse Existing Instead
The DKV module already solved most of the infrastructure this feature needs. This is the most important section for scope control.
| Don't add | Why not | Reuse instead |
|-----------|---------|----------------|
| **A dedicated DÖE / eForms / OCDS SDK** | None exists as a maintained npm package for the DÖE OpenData API specifically. It's a plain REST/Swagger endpoint returning XML/JSON/CSV — a heavy SDK would be unmaintained-dependency risk for zero benefit. | Native `fetch` (already the established pattern) + `fast-xml-parser` for eForms-XML + plain `JSON.parse` for OCDS (it's just JSON, no schema library needed on day 1) + `csv-parse` for the CSV fallback. |
| **`rss-parser`** | Latest release is 3.13.0 from **2023-04-11** — 3+ years stale (no security concern since RSS 2.0 is a frozen spec, but it's an extra dependency for a job `fast-xml-parser` already does). Only 2 feeds to consume (subreport-elvis, service.bund.de), both plain RSS 2.0 — a ~20-line field mapper (`title`/`link`/`pubDate`/`guid`/`description`) over `fast-xml-parser`'s output is simpler than a second library's API. | `fast-xml-parser` (already added for eForms) + a small internal `parseRssFeed()` helper. |
| **`axios` / `axios-cookiejar-support`** | Would introduce a second HTTP client family. `apps/api` has zero `axios` usage today — both existing HTTP call sites use native `fetch`. | Native `fetch` + `fetch-cookie` (wraps `fetch`, not `axios`). |
| **`node-fetch`** | Obsolete since Node 18 ships `fetch` natively (powered by undici); only relevant for Node ≤16 or legacy stream-handling quirks, neither applies here. | Native global `fetch`. |
| **`p-queue`** (v7+) | **ESM-only** — `apps/api`'s `tsconfig.json` is `"module": "commonjs"` (confirmed). A native-ESM package in a CJS NestJS build means either a dynamic `import()` workaround (the codebase already has one messy precedent for `cron` in `dkv-scheduler.service.ts`) or build breakage. Overkill anyway: the module polls a handful of sources (DÖE, 1 AI-AG adapter, 1 cosinex adapter, 2 RSS feeds, 1 inbox) on independent cron ticks — no need for a promise-concurrency-queue library. | A ~20-line internal `politeDelay(ms)` / sequential-`for`-loop helper between requests within one adapter run. Simpler, CJS-native, matches the "wartbar für Nicht-Programmierer" constraint. |
| **`bottleneck`** | Latest release 2.19.5 is from **2019** — unmaintained for 7 years. A community fork (`@rutter/bottleneck`) exists but adds a supply-chain trust question for zero real benefit given the low concurrency needs above. | Same internal delay helper as above. |
| **A separate portal-login-credential store for scraping** | The feasibility research's recommended tactic explicitly avoids automating authenticated scraping: for portals 2,3,4,6,7,8,9,10 (all except vergabe24/aumass, avoided entirely per their ToS), the admin registers **one saved search per portal manually**, and Tessera ingests the resulting **alert emails** — not the authenticated portal UI. Only the *public, unauthenticated* search-result pages (AI AG NetServer, cosinex DTVP) are scraped. | Reuse the existing tenant-scoped encrypted-credential pattern (`CalendarCryptoService`, AES-256-GCM, via `SettingsModule`) *only* for the inbox connection used to ingest alert emails — same shape of secret the DKV module already stores, not a new credential type. |
| **`@nestjs-modules/mailer`** for digest/alert sending | Already an established pitfall in this codebase: it cannot change its SMTP transport after startup, so a runtime SMTP-config change in the admin UI wouldn't take effect without a service restart (documented in `dkv-mail.service.ts`). | Copy the `DkvMailService` pattern: a dedicated `AusschreibungMailService` that calls `nodemailer.createTransport()` fresh on every send, reading `SettingsService.getDecryptedSmtpConfig(tenantId)` each time. |
| **A second cron/scheduling mechanism** | `@nestjs/schedule` + `SchedulerRegistry.addCronJob()` (dynamic, runtime-updatable) is already wired into `AppModule` and proven in `DkvSchedulerService`. | Reuse the same dynamic-cron pattern — one job per source (DÖE poll, AI-AG adapter poll, cosinex adapter poll, 2× RSS poll, inbox alert poll), each independently configurable in minutes via the admin UI, same as DKV's `pollIntervalMin`. |
| **A new IMAP/EWS client** | `ImapProvider` and `ExchangeInboxProvider` (both implementing `InboxProvider`) already handle IMAP (imapflow) and Exchange/EWS (raw SOAP over `httpntlm`, NTLM auth) inbox polling, including attachment size limits (PDF-bomb mitigation) and credential-safe logging. | Reuse `InboxProvider`/`ImapProvider`/`ExchangeInboxProvider` directly for the Unterschwellen alert-email ingestion. Note: the current attachment filter is PDF-specific (`collectPdfParts`/`looksLikePdf`); the new module needs the **email body/links**, not attachments, for most portal alert mails (search-agent notifications are usually HTML emails linking to the tender, occasionally with a PDF). This means extending `InboxProvider` with a body-text/HTML-fetching capability, or adding a sibling interface — a phase-plan design decision, not a new dependency. |
## Alternatives Considered
| Recommended | Alternative | When to Use Alternative |
|-------------|-------------|--------------------------|
| `fetch` + `cheerio` (adapter-first) | `playwright` (adapter-first) | If the phase-1 spike shows a target portal's public search results are rendered/paginated via client-side JS (React/Angular) rather than server HTML — confirm per-portal before committing, don't assume both need it. |
| `fetch-cookie` + `tough-cookie` | `got` (built-in cookie support via `got.extend({ cookieJar })`) | If the adapters later need retry/backoff/HTTP2 features beyond what a thin `fetch` wrapper offers — `got` bundles those, but it's another HTTP client family; only justified if manual retry logic (already needed per DKV's `D-16` exponential-backoff pattern) becomes unwieldy. |
| `fast-xml-parser` for both eForms-XML and RSS | `xml2js` | Only if a specific eForms XML document needs `xml2js`'s SAX-based tolerance for malformed/legacy XML — unlikely for a DÖE-generated, schema-validated eForms-DE feed. |
| No dedicated OCDS validation library | `ajv` + the published OCDS JSON Schema | If/when the module needs to strictly validate incoming OCDS JSON against the official schema (e.g. to catch DÖE-side data quality issues) rather than just mapping known fields defensively. Defer until proven necessary — DÖE's OCDS output is itself schema-validated upstream. |
| Internal `politeDelay()` helper | `p-queue` (dynamic `import()`) or `bottleneck` (unmaintained fork) | Only if the module later needs true concurrent multi-portal fan-out with a shared global rate budget (e.g. dozens of AI-AG-family portals at once, per the feasibility doc's "deckt viele weitere AI-Portale" note) — at that scale a real queue library becomes worth the complexity. Not needed for the 2-adapter v1.1 scope. |
## Version Compatibility
| Package A | Compatible With | Notes |
|-----------|------------------|-------|
| `fast-xml-parser@5.x` | Node 18+ (project already on Node 22-class runtime per `@types/node@^22`) | Pure TS/JS, no native bindings — no compatibility risk. |
| `cheerio@1.2.x` | Node ≥18.17 | Matches project's Node baseline. |
| `fetch-cookie@3.x` | Native global `fetch` (Node 18+) or any WHATWG-`fetch`-compatible function | Wraps whatever `fetch` implementation is passed in — works with Node's built-in `fetch` (undici) with zero extra config. |
| `tough-cookie@6.x` | `fetch-cookie@3.x` | `fetch-cookie` accepts any `tough-cookie`-compatible jar per its own docs; pin both current majors together. |
| `playwright@1.61.x` (if added) | Requires downloading Chromium binary at install time (`playwright install chromium`) | Docker implication: the `apps/api` production image must run `playwright install --with-deps chromium` in its build stage — another reason to add it only if actually needed, not speculatively. |
| `csv-parse@7.x` | Node 18+, ESM **and** CJS builds published | Unlike `p-queue`/`bottleneck`, `csv-parse` ships dual CJS/ESM — safe for the CommonJS `apps/api` build. |
## Sources (v1.1 delta)
- npm registry (`registry.npmjs.org`) — direct authoritative version/publish-date lookup for `cheerio`, `fast-xml-parser`, `csv-parse`, `rss-parser`, `playwright`, `p-queue`, `tough-cookie`, `fetch-cookie`, `bottleneck` — fetched 2026-07-17. Confirms: cheerio 1.2.0 (2026-01-23), fast-xml-parser 5.10.1 (2026-07-16), csv-parse 7.0.1 (2026-07-02), rss-parser 3.13.0 (2023-04-11, stale), playwright 1.61.1 (2026-06-23), p-queue 9.3.1 (2026-07-03, ESM-only since v7), tough-cookie 6.0.2 (2026-07-07), fetch-cookie 3.2.0 (2025-12-15), bottleneck 2.19.5 (2019-08-03, unmaintained).
- WebSearch (MEDIUM confidence, cross-checked against npm registry above where version-critical) — cheerio/fast-xml-parser/rss-parser/csv-parse/playwright/p-queue ecosystem status; undici-vs-node-fetch guidance; tough-cookie/fetch-cookie/axios-cookiejar-support session-handling patterns; Playwright-vs-cheerio scraping tradeoff for stateful ASP.NET-style vs static HTML portals; OCDS-for-eForms mapping (no dedicated npm library found — confirmed via `standard.open-contracting.org` documentation search, not a package).
- `https://www.oeffentlichevergabe.de/documentation/swagger-ui/opendata/` and `bescha.bund.de` DÖE pages — confirms eForms-DE/OCDS/CSV export formats, no-auth access.
- `.planning/research/ausschreibungs-portale-feasibility.md` (this repo, 2026-07-16) — portal platform identification (AI AG NetServer = Java `ControllerServlet`, cosinex = separate incompatible HTML), which portals to scrape vs. email-alert-ingest vs. avoid entirely.
- Codebase inspection (`apps/api/src/dkv/`, `apps/api/src/settings/`, `apps/api/tsconfig.json`, `apps/api/package.json`) — confirmed: CommonJS build target (rules out ESM-only libs), native `fetch` already the established HTTP client (rules out adding `axios`), existing `InboxProvider`/`ImapProvider`/`ExchangeInboxProvider`/`DkvMailService`/`DkvSchedulerService`/`CalendarCryptoService` patterns to reuse verbatim.
**Gap / low-confidence area:** No Context7 or other docs-MCP server was available in this session (`.mcp.json` only configures `playwright` for browser automation, not doc lookup) — all version numbers here come from WebSearch cross-checked directly against `registry.npmjs.org`, which is authoritative for version/publish-date facts but not for qualitative maintenance-health claims. This repo's automated confidence classifier flagged `fast-xml-parser`, `csv-parse`, `playwright`, `p-queue`, and `tough-cookie` as `SUS` — manual review indicates this is a false positive tripped by recent-publish velocity, not an actual supply-chain concern: all five are widely-adopted, well-known-maintainer packages (Playwright is Microsoft's own project). Recommend a final `npm view <pkg>` / Socket.dev spot-check immediately before `pnpm add` in the implementation phase, per this repo's existing dependency-hygiene bar.
---
# v1.0 Base Stack (reference — unchanged)
**Researched:** 2026-06-18
**Overall Confidence:** HIGH
@@ -64,6 +159,8 @@ tessera/
**Why Keycloak over custom auth:** The project requires LDAP integration, multi-tenancy, admin user management, and token-based auth. Building this from scratch would take weeks and introduce security vulnerabilities. Keycloak provides all of this as a Docker container with zero custom code.
> **Note (v1.1):** Live implementation diverged from this section — see `apps/api/src/settings` and the LDAP work already shipped directly against `ldapts` rather than via Keycloak federation. This base-stack section is kept as originally researched; treat the "Authentication & Authorization" row above as historical context, not current fact, when planning new auth-adjacent work.
### Desktop Wrapper
| Technology | Version | Purpose | Why |
@@ -83,6 +180,8 @@ tessera/
| Traefik | 3.x | Reverse proxy | Automatic SSL, Docker-native service discovery, routing rules via labels. Simpler than nginx for Docker-compose setups |
| Gitea | existing | Version control | Already in place. Automate via webhooks and Gitea API |
> **Note (v1.1):** Reverse proxy in production is **Nginx Proxy Manager (NPM)**, external to the Tessera Docker stack — not Traefik. Tessera containers do not include a reverse proxy; NPM handles SSL termination and routing on the host. Treat the "Traefik" row above as historical/superseded.
### Testing
| Technology | Version | Purpose | Why |
@@ -117,7 +216,7 @@ tessera/
| Linting | Biome | ESLint + Prettier | Single tool, 100x faster, less config. ESLint is being replaced in NestJS 12 roadmap anyway |
| Component Library | shadcn/ui | Material UI / Ant Design | shadcn gives ownership of components (no dep lock-in), built on Radix primitives, Tailwind-native |
| Dashboard Grid | react-grid-layout | Gridstack.js | React-native, TypeScript rewrite in v2, hooks API, responsive breakpoints |
| Reverse Proxy | Traefik | nginx | Docker-native service discovery, auto-SSL, config via labels not files |
| Reverse Proxy | Traefik | nginx | Docker-native service discovery, auto-SSL, config via labels not files — **superseded in practice by external Nginx Proxy Manager, see note above** |
## Multi-Tenancy Strategy
@@ -177,7 +276,7 @@ services:
db: # PostgreSQL 16 (port 5432)
redis: # Redis 7 (port 6379)
keycloak: # Keycloak 26.6.x (port 8080)
traefik: # Reverse proxy (port 80/443)
traefik: # Reverse proxy (port 80/443) — superseded by external NPM, see note above
```
## Version Pinning Strategy
@@ -187,7 +286,7 @@ services:
- **Tailwind/shadcn:** Follow latest within major — utility additions are non-breaking
- **Keycloak Docker image:** Pin to minor (e.g., `quay.io/keycloak/keycloak:26.6`)
## Sources
## Sources (v1.0 base)
- [Next.js 16 Docs](https://nextjs.org/docs/app/guides/upgrading/version-16) — Version 16.2.7+ stable
- [NestJS Documentation](https://docs.nestjs.com/) — Version 11.1.x
+97 -106
View File
@@ -1,165 +1,156 @@
# Project Research Summary
**Project:** Tessera - Modular Portal Platform with Marketplace
**Domain:** Multi-tenant SaaS portal with module marketplace, configurable dashboard, desktop wrapper
**Researched:** 2026-06-18
**Confidence:** HIGH
**Project:** Tessera — Ausschreibungs-Radar Module (v1.1 Milestone)
**Domain:** German public-procurement (Vergabe) tender aggregation, filtering, and notification, built as a new module on an existing multi-tenant NestJS/Prisma platform
**Researched:** 2026-07-17
**Confidence:** MEDIUM-HIGH overall (HIGH on architecture/platform integration and domain-feature patterns; MEDIUM on new-library version currency and unverified DÖE pagination details)
## Executive Summary
Tessera is a modular portal platform where tenants activate licensed workflow modules from a marketplace, interact through a configurable drag-and-drop dashboard, and optionally use a lightweight desktop wrapper. The expert consensus for this type of product is a modular monolith backend (NestJS) with PostgreSQL Row-Level Security for tenant isolation, a Next.js frontend shell that lazy-loads module UIs on demand, and Keycloak for identity management. This stack avoids premature microservice complexity while maintaining clean extraction boundaries.
The Ausschreibungs-Radar module aggregates German public tenders from multiple incompatible sources — a central auth-free API (DÖE OpenData, ~75% of market value by €), two structurally different scraping targets (AI AG NetServer and cosinex Vergabemarktplatz), RSS feeds, and email-alert ingestion for the long tail — into one normalized, searchable, filterable, notifiable catalog. Every mature tender-monitoring product (Stotles, Tendium, native German portals) follows the same shape: ingest -> normalize -> dedupe -> saved-search filter -> results UI -> per-user triage state -> digest/instant notification. Tessera's module maps onto this shape directly and, crucially, can reuse substantial existing platform infrastructure: the DKV module's inbox-provider abstraction (ImapProvider/ExchangeInboxProvider), its SmtpConfig fresh-transport mail pattern, and @nestjs/schedule dynamic cron. No new frontend dependencies are needed at all — the results UI is built entirely on the existing Next.js/shadcn/TanStack Query stack.
The recommended approach is to build foundation-first: Docker infrastructure with network segmentation, database schema with RLS policies, and the frontend shell with i18n and theming baked in from day one. Authentication and tenant context middleware come next, followed by the module system with a versioned API contract, then marketplace and dashboard features. The desktop wrapper (Tauri) is a late-stage concern that wraps the already-working web app.
The recommended approach is a DÖE-only MVP first: build the full pipeline (schema, one adapter, normalizer, filter engine, saved searches, results UI, read/favourite state, digest + instant notification) against the single easiest, legally unambiguous, highest-value source before touching any scraping. This validates the entire module concept on live data with zero ToS risk, and every subsequent source (AI-AG adapter, cosinex adapter, RSS, email-alert) slots into the same pipeline via a TenderSourceAdapter interface without touching ingestion/matching/notification code. The single most important architectural decision — and the biggest deviation from the DKV module template — is that tender data is platform-global, not tenant-scoped: a DÖE notice is a public fact relevant to every tenant, so ingestion/dedup run once for the whole platform, while only saved searches, match results, and notification preferences are tenant-scoped.
The primary risks are tenant data leakage (mitigated by RLS at the database level, not application-level filtering), module API instability (mitigated by a versioned SDK contract from the start), and i18n/theme retrofitting pain (mitigated by establishing both before the first UI component). Docker network segmentation is a day-one infrastructure requirement to prevent lateral movement between containers.
Key risks: (1) legal exposure from scraping vergabe24/aumass, both of which have explicit AGB bans on automated access — this must be a hard-coded denylist, not just documentation; (2) copying the DKV scheduler's known single-tenant findFirst() pattern, which would silently break polling for every tenant but the first; (3) notification storms on first saved-search activation (backfill) or on tender updates re-triggering "new" alerts, both requiring an explicit matched-vs-notified state split; (4) cross-source deduplication complexity — correctly deferred until a second source exists, since dedup against a single source is meaningless. All four are addressed by explicit phase-level design decisions documented in PITFALLS.md and ARCHITECTURE.md.
## Key Findings
### Recommended Stack
A pnpm + Turborepo monorepo with three apps (Next.js 16 frontend, NestJS 11 backend, Tauri 2 desktop) and shared packages for types, UI components, config, and database schema.
The v1.1 delta adds five focused libraries to apps/api, no new apps/web dependencies. All additions were checked against the existing CommonJS build target (ruling out ESM-only packages like p-queue) and the established native-fetch-only HTTP pattern (ruling out axios).
**Core technologies:**
- **Next.js 16 + React 19**: App Router, Server Components, Turbopack, multi-tenant middleware support
- **NestJS 11**: Modular architecture maps directly to Tessera's module system; DI, guards, interceptors
- **PostgreSQL 16 + Prisma 7**: RLS for multi-tenancy, Prisma Client Extensions for tenant context, pure TS runtime in v7
- **Keycloak 26.6**: Native multi-tenancy via realms, LDAP federation, eliminates custom auth
- **Tauri 2**: 5MB installer vs Electron's 150MB; thin wrapper around the web app
- **Tailwind 4 + shadcn/ui**: Utility-first styling with owned component library, trivial dark mode
- **react-grid-layout 2.2**: TypeScript rewrite with hooks API for drag-and-drop dashboard
- **Traefik 3**: Docker-native reverse proxy with auto-SSL and label-based config
**Core technologies (new):**
- **fast-xml-parser** (5.10.1) — parses both eForms-DE XML (DÖE) and RSS 2.0 feeds with one library; zero-dependency, pure-TS, avoids adding a separate stale rss-parser (last released 2023).
- **cheerio** (1.2.0) — HTML parsing for the AI-AG and cosinex scraping adapters; 10-50x cheaper than a headless browser for server-rendered HTML.
- **csv-parse** (7.0.1) — DÖE CSV export as a secondary/verification format alongside eForms-XML/OCDS-JSON.
- **tough-cookie + fetch-cookie** (6.0.2 / 3.2.0) — session-cookie handling for scraping adapters, wrapping native fetch rather than introducing axios.
- **playwright** (1.61.1, conditional) — only added if a phase-1 spike confirms a target portal's public search requires client-side JS rendering; not added speculatively (~300MB+ Docker image cost).
**Explicitly rejected in favor of reuse:** a dedicated DÖE/eForms SDK (none exists; plain REST+parsers suffice), p-queue/bottleneck (ESM-only or unmaintained; a ~20-line internal delay helper suffices at this scale), @nestjs-modules/mailer (documented DKV pitfall — can't change SMTP transport at runtime), a second cron mechanism (reuse @nestjs/schedule + SchedulerRegistry), and a new IMAP/EWS client (reuse InboxProvider/ImapProvider/ExchangeInboxProvider verbatim, extended for email body/HTML access rather than PDF attachments).
### Expected Features
**Must have (table stakes):**
- Authentication with RBAC and multi-tenancy with data isolation
- Sidebar navigation with module categories
- Module activation/deactivation and admin-managed licensing per tenant
- i18n (DE + EN), light/dark theme, responsive layout
- Basic dashboard with widgets, session management, error handling/loading states
- User profile and settings, search/filter in module lists
**Must have (table stakes, DÖE-only MVP):**
- DÖE OpenData ingestion (eForms/OCDS/CSV, auth-free) as the sole MVP data source
- Normalized OCDS-oriented tender schema (title, buyer, CPV codes, region/PLZ, deadline, value, procedure type, source URL, raw payload)
- Filter engine: full-text keyword (Postgres tsvector/GIN), region/PLZ/Bundesland, CPV code, deadline range, value range
- Saved-search/filter profiles, per-user (not just per-tenant) and tenant-aware
- Searchable/sortable results list + detail view linking to source (never mirror bid documents)
- Read/unread and favourite/shortlist marking, both per-user state
- Periodic email digest (configurable interval) and instant alert on new match, both reusing MailModule/SMTP infra
**Should have (differentiators):**
- Drag-and-drop dashboard with resizable widgets
- Widget marketplace/gallery
- Module hot-activation without restart
- Per-tenant branding (logo + accent color)
- Module-provided dashboard widgets
- Keyboard shortcuts and command palette
- LDAP/AD directory sync
**Should have (add after MVP validation, v1.x):**
- AI-AG NetServer adapter (covers lhs-vpbw, tender24, vergabe.landbw + other AI-AG portals via one config-driven adapter)
- cosinex VMP adapter (DTVP + other Länder marketplaces)
- RSS ingestion (subreport-elvis, service.bund.de) — lowest-effort source expansion
- Email-alert ingestion for remaining Unterschwellen portals, reusing DKV inbox infra
- Cross-source deduplication — mandatory once a 2nd source exists, meaningless before
- "Manual watch" indicator for vergabe24/aumass (excluded portals, trust-building)
**Defer (v2+):**
- LDAP integration (manual user creation suffices initially)
- Desktop wrapper (browser works fine first)
- Notification center, admin impersonation, onboarding wizard
**Defer (v2+):** TED API v3 (redundant to DÖE for DE-only scope), relevance/ranking scoring, dashboard "upcoming deadlines" widget (high-leverage but depends on Phase-8 widget-SDK reuse), CSV/Excel export, team/collaboration features (scope creep toward bid-management).
**Anti-features to actively avoid:** scraping vergabe24/aumass directly (ToS-prohibited), a generic "scrape any portal URL" framework (fragile across 3+ incompatible portal families), sub-hourly polling (tender lifecycles run days-to-weeks), full bid-management/CRM scope, AI-generated summaries/scoring (premature), and mirroring/hosting bid documents locally (legal ambiguity).
### Architecture Approach
Modular monolith: single deployable NestJS backend with domain-separated modules (auth, tenant, marketplace, dashboard, module loader) communicating via in-process calls. Frontend is a React shell that lazy-loads module UIs via dynamic imports based on tenant activation state. Shared PostgreSQL database with RLS enforces tenant isolation at the database level. Modules follow a standard interface contract (TesseraModule) for uniform discovery and lifecycle management.
The module is a new, self-contained TendersModule built entirely on existing Tessera infrastructure (module-registry self-seed, dynamic cron via SchedulerRegistry, extracted shared InboxModule, tenant-scoped SMTP). The one deliberate break from the DKV template: Tender rows carry no tenantId — they are platform-global reference data, deduplicated once and shared across all tenants. Tenant scoping lives one layer up, in TenderSavedSearch and TenderMatch. This avoids multiplying scraping requests and storage by tenant count and keeps cross-tenant dedup coherent.
**Major components:**
1. **Frontend Shell** -- routing, sidebar, theme, i18n; lazy-loads module UIs
2. **API Gateway (Traefik)** -- reverse proxy, rate limiting, tenant header injection
3. **Auth Module** -- Keycloak integration, JWT, session management
4. **Tenant Module** -- context resolution, RLS activation per request
5. **Module Loader** -- plugin lifecycle: discover, validate, activate, deactivate
6. **Marketplace Service** -- module catalog, licensing, activation per tenant
7. **Dashboard Engine** -- widget grid, layout persistence, module-contributed widgets
1. **TenderSourceAdapter implementations** (DÖE, AI-NetServer, cosinex, RSS, email-alert) — one interface, N implementations, directly analogous to the existing InboxProvider pattern; each source is a config-driven adapter, not one adapter per portal instance.
2. **TenderNormalizerService** — maps each source's raw shape into the unified OCDS-oriented Tender schema, computes a priority-ordered dedup key (OCID -> sourcePortal:noticeId -> fuzzy fingerprint hash) and a separate contentHash for change detection.
3. **TenderIngestionService + TenderSchedulerService** — poll -> normalize -> upsert -> change-detect, with one global cron per source (not per tenant) and per-tenant cron only for the genuinely per-tenant email-alert mailbox path.
4. **TenderMatchingService** — evaluates active saved searches as DB-scoped Prisma queries against only the newly-changed delta batch per poll, not full-table in-memory scans; keyword matching via Postgres GIN full-text index, not ILIKE.
5. **TenderMailService** — instant (on match creation) and digest (batched cron) notification, reusing the DkvMailService fresh-transport-per-send pattern against tenant-specific SmtpConfig.
Build order ships DÖE-first as a complete, demoable slice (schema -> adapter -> normalizer -> ingestion+scheduler -> matching+controller -> web UI -> mail), then adds scraping adapters (Phase B), then RSS + email-alert ingestion (Phase C, lowest ROI, requires the inbox/ extraction refactor as a prerequisite).
### Critical Pitfalls
1. **Tenant data leaks** -- Enforce PostgreSQL RLS from day one; never rely on application-level WHERE clauses alone
2. **Module API contract instability** -- Define a versioned @tessera/sdk with semver before shipping the first module
3. **Flat Docker network** -- Segment into frontend-net, backend-net, data-net from the initial docker-compose.yml
4. **i18n retrofitting** -- Every string through t('key') from the first component; design for German text length
5. **LDAP as direct auth** -- Use Keycloak as identity layer; LDAP for directory sync only, never raw bind for auth
1. **Copying the DKV scheduler's single-tenant findFirst() pattern** — would silently break polling for every tenant except the first. Must design as findMany({ where: { isActive: true } }) with one job per tenant/source from day one; add an explicit two-tenant scheduler test as an acceptance criterion.
2. **Poll-per-tenant instead of poll-once-fan-out-many** — naive per-tenant-per-search polling multiplies load against shared upstream portals (throttling/ban risk, looks like abuse). Decouple: poll each upstream source once on a shared schedule, fan results out to all matching tenant saved searches from the ingested data.
3. **Legal/ToS risk scraping vergabe24 and aumass** — both have explicit AGB bans on automated/scripted access. Enforce as a hard denylist in code (adapter registry refuses to register a denylisted portal id), not just documentation; their Oberschwellen data is already covered via DÖE, so there's no completeness argument to scrape them.
4. **Notification storms (backfill + duplicate-alert flooding)** — first saved-search activation can surface months of historical DÖE matches; a tender update (deadline extension) can re-trigger a "new" alert if only "matched" is tracked, not "notified." Requires a separate matched-vs-notified state (TenderMatch.notifiedAt) with explicit backfill-suppression on saved-search creation.
5. **Cross-source deduplication attempted too early or done as exact-ID match** — meaningless with one source; once >=2 sources exist, requires fuzzy fingerprint matching (buyer+title+CPV+deadline+value), not exact ID equality, since no shared cross-portal identifier exists below the eForms/TED notice number.
6. **Tenant-scoping the Tender table like DKV data** — the single highest-leverage anti-pattern to avoid; multiplies storage and scraping load N-times for identical public data and makes dedup incoherent.
## Implications for Roadmap
### Phase 1: Foundation and Infrastructure
**Rationale:** Everything depends on the database, Docker setup, backend skeleton, and frontend shell. RLS, network segmentation, i18n, and theming must exist before any feature code.
**Delivers:** Running Docker Compose stack with PostgreSQL (RLS enabled), NestJS skeleton, Next.js shell with routing/layout/sidebar frame, i18n infrastructure, theme/design tokens, Traefik reverse proxy, segmented Docker networks.
**Addresses:** Responsive layout, theme system, i18n framework, error handling/loading states (table stakes infrastructure)
**Avoids:** Pitfalls 1 (tenant isolation), 4 (flat Docker network), 6 (i18n retrofitting), 14 (theme without design tokens), 11 (volume permissions)
Based on research, suggested phase structure:
### Phase 2: Authentication, Tenancy, and User Management
**Rationale:** Module activation requires knowing who the user is and which tenant they belong to. Auth is the prerequisite for everything else.
**Delivers:** Keycloak integration, login/logout, JWT-based sessions, RBAC, tenant context middleware setting RLS per request, user management CRUD, user profile/settings.
**Addresses:** Authentication, RBAC, multi-tenancy, session management, user profile (table stakes)
**Avoids:** Pitfalls 5 (LDAP as auth), 10 (tenant context lost in async)
### Phase 1: DÖE-Only Tender Radar (End-to-End MVP)
**Rationale:** DÖE is auth-free, legally unambiguous, covers ~75% of market value by €, and is the only source needed to validate the entire module concept — schema, filter engine, notifications — against real live data before any scraping work begins. Everything else is additive on top of this pipeline.
**Delivers:** Prisma schema (Tender, TenderSourcePollConfig, TenderSavedSearch, TenderMatch + GIN full-text index), TendersModule skeleton + self-seed, DoeOpenDataAdapter, TenderNormalizerService (OCID-based dedup key + contentHash), TenderIngestionService + global scheduler, TenderMatchingService (delta-scoped DB query filtering), TendersController (list/detail/saved-search CRUD), full Next.js results UI (trefferliste, filters, detail view, saved-search management), read/unread + favourite/shortlist per-user state, TenderMailService (instant + digest, reusing SmtpConfig/DKV mail pattern).
**Addresses:** All P1 table-stakes features from FEATURES.md — ingestion, normalized schema, filter engine, saved searches, results UI, triage state, both notification modes.
**Avoids:** Pitfall 24 (single-tenant scheduler copy — build multi-tenant from day one), Pitfall 23 (notification storms — matched/notified split + backfill suppression), Pitfall 25 (tenant isolation gaps — clear schema split between global Tender and tenant-scoped TenderSavedSearch/TenderMatch).
### Phase 3: Module System and First Module
**Rationale:** The module system must exist before the marketplace can manage it. Building the first real module (Domaincheck) validates the entire architecture.
**Delivers:** Module interface contract (@tessera/sdk), database-driven module registry, module loader with activation/deactivation, lazy-loading frontend integration, Domaincheck module as proof of concept.
**Addresses:** Module activation/deactivation, sidebar navigation reflecting active modules (table stakes); module dependency declaration (differentiator)
**Avoids:** Pitfalls 2 (API contract instability), 3 (monolithic module loading), 8 (scattered licensing), 12 (over-engineered module communication)
### Phase 2: Scraping Adapters — AI-AG NetServer + cosinex VMP
**Rationale:** Second-best ROI per feasibility research: one AI-AG adapter covers 3+ portal instances (lhs-vpbw, tender24, vergabe.landbw) plus many unlisted AI-AG-family portals; one cosinex adapter covers DTVP and is reusable for other Länder marketplaces. Both are public, unauthenticated search pages with "Niedrig-Mittel" ToS risk. Building this after Phase 1 means the ingestion/matching/notification pipeline is already proven and untouched — only new adapters plug in.
**Delivers:** AiNetServerAdapter and CosinexAdapter behind the existing TenderSourceAdapter interface, config-driven per portal instance (not one class per portal), with the required cross-source dedup logic (fuzzy fingerprint match) now meaningful for the first time.
**Uses:** cheerio, tough-cookie/fetch-cookie from STACK.md; conditionally playwright only if a phase-start spike confirms JS-rendered search results on either portal.
**Implements:** Pattern 1 (Source-Adapter Abstraction), Pattern 2's cross-source dedup key fallback (sourcePortal:noticeId).
**Avoids:** Pitfall 15 (scraper fragility — session tokens, layout drift; build structural health checks per adapter run from day one), Pitfall 21 (cross-source dedup — fuzzy fingerprint, not exact-ID), Pitfall 16 (legal denylist — vergabe24/aumass must be structurally impossible to add, enforced in the adapter registry, not just documented).
### Phase 4: Marketplace and Dashboard
**Rationale:** Both depend on the module system but are largely independent of each other. Marketplace manages module catalog and licensing; dashboard provides the user home with widgets.
**Delivers:** Marketplace UI (browse, activate, manage licenses), licensing service, dashboard with react-grid-layout, core widgets (clock, search, notes, calendar), drag-and-drop with layout persistence, module-provided widgets.
**Addresses:** Marketplace UI with licensing (table stakes); drag-and-drop dashboard, widget gallery, module-provided widgets (differentiators)
**Avoids:** Pitfall 7 (widget state explosion -- debounced persistence, JSON column)
### Phase 5: Polish, Desktop, and Advanced Features
**Rationale:** Enhancement layer on a stable web platform. Desktop wrapper requires stable web app. LDAP and advanced features add value but are not launch-blocking.
**Delivers:** Tauri desktop wrapper, LDAP/AD directory sync via Keycloak, per-tenant branding, keyboard shortcuts/command palette, Gitea CI/CD integration.
**Addresses:** Desktop wrapper, LDAP integration, per-tenant branding, command palette (differentiators)
**Avoids:** Pitfall 9 (desktop divergence -- keep wrapper thin), 13 (Gitea tight coupling)
### Phase 3: RSS + Email-Alert Long Tail
**Rationale:** Lowest ROI per feasibility research (do last) but lowest implementation cost — RSS ingestion is a small field-mapper over the already-added fast-xml-parser, and email-alert ingestion directly reuses DKV's inbox infrastructure once extracted into a shared module. This closes coverage on the remaining Unterschwellen long tail (portals 2, 6 via alert-email; 7, 10 via RSS) without adding any new scraping ToS exposure.
**Delivers:** inbox/ shared-module extraction (prerequisite refactor moving ImapProvider/ExchangeInboxProvider out of dkv/), RssAdapter (subreport-elvis, service.bund.de), TenderInboxConfig (per-tenant, mirrors DkvModuleConfig credential pattern) + EmailAlertAdapter + per-tenant scheduler jobs for the genuinely per-tenant mailbox-polling path.
**Delivers (feature-level):** "Manual watch" indicator for vergabe24/aumass, source-coverage transparency in the UI (Pitfall 20 — communicate the Oberschwelle-vs-Unterschwelle completeness gap explicitly rather than presenting a falsely-unified result list).
**Avoids:** Pitfall 20 (false completeness — must ship coverage-transparency metadata alongside this phase, not reactively later).
### Phase Ordering Rationale
- Phases follow strict dependency chains: infrastructure -> auth/tenancy -> modules -> marketplace/dashboard -> polish
- i18n and theming are in Phase 1 (not Phase 5) because retrofitting is a known moderate pitfall
- Module system before marketplace because the marketplace manages modules -- the registry must exist first
- Dashboard is grouped with marketplace (Phase 4) because both depend on the module system and can partially parallelize
- Desktop wrapper is last because it wraps an already-working web app and building it earlier risks divergence
- **DÖE-first is a hard dependency, not a preference:** the filter engine, saved searches, and results UI all require normalized data to exist before they can be built and validated — nothing else can be usefully tested without Phase 1's pipeline in place.
- **Scraping adapters (Phase 2) come before RSS/email-alert (Phase 3)** despite RSS/email being individually cheaper, because cross-source deduplication only becomes meaningful and testable once a second scraped source exists at meaningful volume — RSS/email sources are comparatively low-volume long-tail additions that exercise the same dedup logic but don't independently justify building it.
- **The inbox/ extraction refactor is scoped into Phase 3**, not Phase 1, because it's only a hard prerequisite for the email-alert adapter — doing it earlier would be premature refactoring of code Phase 1/2 never touch.
- **Global-vs-tenant data split (Pattern 3) and the multi-tenant scheduler (avoiding Pitfall 24) must both be correct in Phase 1**, even though only one tenant may be active during initial testing — retrofitting either after data/schedules exist is expensive and risks silent data-model migration bugs.
### Research Flags
Phases likely needing deeper research during planning:
- **Phase 2:** Keycloak realm-per-tenant configuration and NestJS integration patterns; RLS session variable management with Prisma Client Extensions
- **Phase 3:** Module interface contract design; lazy-loading strategy for module UIs with Next.js App Router
- **Phase 4:** react-grid-layout integration with module-provided widgets; layout persistence schema design
- **Phase 1:** DÖE OpenData API pagination and rate-limit parameters are explicitly unverified (Swagger UI is JS-rendered; feasibility doc flags this as an open point) — confirm via a live API call early in phase planning, not assumed from docs.
- **Phase 1:** OCDS field-mapping edge cases (conditional/optional fields, parties[].name resolution, release-vs-record package choice) warrant a focused read of the OCDS eForms profile during normalizer design, per Pitfalls 17-18.
- **Phase 2:** Whether AI-AG NetServer or cosinex VMP search pages require JS rendering is unverified — needs a phase-start spike (confirm cheerio-only feasibility per-portal) before committing to playwright.
- **Phase 2:** Exact AGB anti-scraping clause wording for RIB, subreport, deutsche-evergabe, tender24 remains an open verification point from the feasibility doc — recommend a quick manual AGB check before scraping any additional portal beyond AI-AG/cosinex.
Phases with standard patterns (skip research-phase):
- **Phase 1:** Docker Compose, NestJS/Next.js scaffolding, Tailwind/shadcn setup -- well-documented, established patterns
- **Phase 5:** Tauri wrapper is a thin shell; LDAP sync via Keycloak is documented
- **Phase 1's notification/mail integration:** directly copies the proven DkvMailService pattern — no new research needed.
- **Phase 3's inbox-module extraction:** mechanical refactor of already-working, already-understood code (ImapProvider/ExchangeInboxProvider).
## Confidence Assessment
| Area | Confidence | Notes |
|------|------------|-------|
| Stack | HIGH | All technologies are mature, well-documented, and version-verified. Clear rationale for each choice with alternatives considered. |
| Features | HIGH | Feature landscape is well-defined with clear table stakes vs differentiators. Dependency chain is explicit. |
| Architecture | HIGH | Modular monolith with RLS is a proven pattern for multi-tenant SaaS. Multiple authoritative sources confirm the approach. |
| Pitfalls | HIGH | Critical pitfalls sourced from production incident reports and authoritative guides. Prevention strategies are concrete. |
| Stack | MEDIUM | Versions cross-verified via WebSearch + direct npm registry lookup (authoritative for version/publish-date facts); no Context7/docs-MCP available this session for qualitative maintenance-health claims. All five new packages are well-known, actively-maintained (Playwright is Microsoft's own project) despite an automated SUS flag the researcher identified as a false positive. |
| Features | MEDIUM-HIGH | Domain patterns cross-checked against live tender-monitoring SaaS products (Stotles, Tendium, TenderAlerts) plus the OCDS standard; German-specific integration details inherit HIGH confidence from the project's own prior ausschreibungs-portale-feasibility.md research. |
| Architecture | HIGH for platform integration (module registration, tenant scoping, inbox/mail reuse — verified against live dkv/ / module-registry/ source); MEDIUM for OCDS field mapping and DÖE API pagination (verified via spec/Swagger listing, not a live API call). |
| Pitfalls | HIGH | Grounded in the project's own feasibility research, official OCDS-for-eForms docs, DÖE OpenData API docs, existing codebase patterns (dkv-scheduler.service.ts, tenant.guard.ts), and German scraping case law (BGH I ZR 224/12). MEDIUM specifically on exact DÖE pagination/rate-limit behavior, an explicitly flagged open verification point. |
**Overall confidence:** HIGH
**Overall confidence:** MEDIUM-HIGH
### Gaps to Address
- **Prisma + RLS integration specifics:** Prisma Client Extensions for setting session variables per request need validation during Phase 2 implementation. The pattern is documented but not battle-tested at scale with Prisma 7.
- **Module hot-activation with Next.js App Router:** Lazy-loading module UIs via dynamic imports is standard React, but integrating with Next.js App Router file-based routing may require a custom module route convention. Needs Phase 3 research.
- **Keycloak realm-per-tenant vs single-realm with groups:** Research assumes realm-per-tenant mapping but single-realm with tenant claims may be simpler for the initial scale. Decide during Phase 2 planning.
- **react-grid-layout v2 maturity:** The TypeScript rewrite (v2.2) is recent. Fallback to v1 /legacy API if v2 hooks have issues.
- **DÖE OpenData API pagination/rate-limit parameters:** Swagger UI is JS-rendered and wasn't live-queried this session. Resolve with a direct API call during Phase 1 planning before finalizing poll cadence.
- **Whether AI-AG NetServer / cosinex VMP search pages require JS rendering:** unverified; drives the playwright add/skip decision for Phase 2. Resolve with a phase-start technical spike, not assumed.
- **Exact AGB anti-scraping clause wording for RIB, subreport, deutsche-evergabe, tender24:** flagged as open in the feasibility doc; low urgency since none of these are in the Phase 1-2 build order, but should be checked before any Phase 3+ expansion beyond the currently-scoped RSS/email-alert sources.
- **Whether vergabe.landbw.de offers an on-page email alert:** open verification point from feasibility doc; relevant only if/when its email-alert ingestion is scoped into Phase 3.
- **Package "SUS" flags on fast-xml-parser/csv-parse/playwright/p-queue-alternatives:** researcher assessed as false positives (recent-publish velocity, not actual supply-chain risk) but recommends a final npm view/Socket.dev spot-check immediately before pnpm add in Phase 1 implementation, per the repo's existing dependency-hygiene bar.
## Sources
### Primary (HIGH confidence)
- [Next.js 16 Docs](https://nextjs.org/docs) -- App Router, Server Components, middleware
- [NestJS Documentation](https://docs.nestjs.com/) -- Modules, guards, interceptors
- [Prisma ORM 7](https://www.prisma.io/blog/announcing-prisma-orm-7-0-0) -- Pure TS runtime, Client Extensions
- [PostgreSQL RLS](https://aws.amazon.com/blogs/database/multi-tenant-data-isolation-with-postgresql-row-level-security/) -- Multi-tenant isolation
- [Keycloak 26.6](https://www.keycloak.org/) -- Identity provider, realm multi-tenancy
- [Tauri 2.0](https://v2.tauri.app/) -- Desktop wrapper
- [Martin Fowler: Feature Toggles](https://martinfowler.com/articles/feature-toggles.html) -- Module activation as feature gating
- Existing codebase (direct read): apps/api/src/dkv/*, apps/api/src/module-registry/*, apps/api/src/mail/*, apps/api/src/prisma/prisma-tenant.extension.ts, apps/api/src/tenant/tenant.middleware.ts, apps/api/prisma/schema.prisma, apps/web/src/lib/module-loader.ts, apps/web/src/app/(portal)/modules/dkv-fleet/*
- .planning/research/ausschreibungs-portale-feasibility.md (project-internal, 2026-07-16) — portal-by-portal capability matrix, ToS risk ratings, DÖE/TED coverage statistics
- registry.npmjs.org — direct version/publish-date lookup for all new libraries (fetched 2026-07-17)
- BGH, 30.04.2014 — I ZR 224/12 — German case law on scraping legality and "virtuelles Hausrecht"
### Secondary (MEDIUM confidence)
- [Multi-Tenant SaaS Architecture guides](https://apipilot.com/developing-a-multi-tenant-saas-application-the-2026-architecture-guide/) -- Architecture patterns
- [Node.js Plugin Architecture](https://medium.com/codeelevation/node-js-plugin-architecture-build-your-own-plugin-system-with-es-modules-5b9a5df19884) -- Module system design
- [Docker Network Isolation](https://hexshift.medium.com/docker-network-isolation-pitfalls-that-put-your-applications-at-risk-b60356a14033) -- Network segmentation
- [LDAP Authentication Anti-Pattern](https://blog.lithnet.io/2018/03/the-ldap-authentication-anti-pattern.html) -- Auth layer design
- Open Contracting Data Standard — Schema/Codelists Reference (standard.open-contracting.org) — OCDS field structure, CPV usage, release-vs-record model
- oeffentlichevergabe.de OpenData Swagger UI — confirms ocds-mnwr74 German federal OCDS prefix, CC-Zero licensing (page requires JS for full pagination detail)
- Stotles, Tendium, TenderAlerts.eu — competitive feature-pattern confirmation
- WebSearch cross-checked against npm registry — ecosystem status for cheerio/fast-xml-parser/csv-parse/playwright and rejected alternatives
### Tertiary (LOW confidence)
- None flagged — all findings traced to at least a MEDIUM-confidence primary or secondary source.
---
*Research completed: 2026-06-18*
*Research completed: 2026-07-17*
*Ready for roadmap: yes*
@@ -0,0 +1,51 @@
# Feasibility — Ausschreibungs-Modul (10 Vergabeportale)
_Research 2026-07-16. Source portals: `user-files/portale.txt`._
## Kernkorrektur zur Annahme
`/NetServer/…ControllerServlet` = Fingerprint von **Administration Intelligence AG (AI AG)**, NICHT cosinex.
- **DTVP (3)** = cosinex Vergabemarktplatz (einziges cosinex-Portal der Liste).
- **lhs-vpbw (4), tender24 (8), vergabe.landbw (9)** = AI AG NetServer → **ein AI-Adapter deckt 4/8/9** (plus viele weitere AI-Portale).
- cosinex- und AI-HTML sind inkompatibel — kein gemeinsamer Adapter.
## Portal-Tabelle
| # | Portal | Plattform | Öffentl. Suche | API/Feed | Login | Native E-Mail-Alerts | Anti-Bot/ToS | Zentral verfügbar |
|---|--------|-----------|:--:|--------|-------|:--:|------|------|
| 1 | vergabe24.de | AI AG (Staatsanzeiger) | Nein (paywall) | eForms only | OIDC IdentityServer | Ja (Suchprofile) | **Hoch** — AGB verbietet Extraktion, nennt Raten | eForms→DÖE/TED |
| 2 | deutsche-evergabe.de | Healy Hudson DEVA | Ja | eForms only | Form | Ja (Suchagent ~5€/mo) | Mittel | eForms→DÖE/TED |
| 3 | dtvp.de | **cosinex** VMP | Ja | – | cosinex Konto | Ja (Suchprofil, daily) | Mittel | eForms→DÖE/TED; service.bund.de |
| 4 | lhs-vpbw.vmstart.de | AI AG NetServer | Ja | – | Form/Bietercockpit | Ja | Niedrig-Mittel | eForms→DÖE/TED |
| 5 | plattform.aumass.de | aumass | Ja | nur XML pro Verfahren | Form | Ja (paywall) | **Hoch** — AGB verbietet Skripte | Ja (DÖE-Lieferant) |
| 6 | rib.de | RIB iTWO / meinauftrag | Ja | – | RIB SSO | Ja (~5€/mo) | Mittel | Ja (DÖE-Lieferant) |
| 7 | subreport-elvis.de | subreport ELViS | Ja | **RSS** (veröff./vergeben) | Form | Ja (Suchprofil daily) | Mittel | eForms→TED |
| 8 | tender24.de | AI AG NetServer | Ja | – | Form/Bietercockpit | Ja | Niedrig-Mittel | eForms→DÖE/TED |
| 9 | vergabe.landbw.de | AI AG NetServer | Ja | – | Form/Bietercockpit | wahrsch. | Niedrig-Mittel | eForms→DÖE/TED; service.bund.de |
| 10 | service.bund.de | Beschaffungsamt BMI | Ja (voll öffentl.) | **RSS pro Suche** | keiner | Ja (Newsletter/RSS) | **Niedrig** | zentraler Aggregator (→DÖE) |
## Zentraler Feed (eForms / TED / DÖE)
Seit 25.10.2023: alle **EU-Oberschwellen**-Bekanntmachungen als eForms-DE → **Datenservice Öffentlicher Einkauf (DÖE)** → TED.
- **DÖE OpenData API** (`oeffentlichevergabe.de`, Swagger `/documentation/swagger-ui/opendata/`): eForms-DE XML, **OCDS `ocds-mnwr74`**, CSV. Keine Registrierung. Deutschland-Backbone.
- **TED API v3** (`api.ted.europa.eu/v3/notices/search`): keyless read, ITERATION-Mode uncapped. EU-weit, für DE-only weitgehend redundant zu DÖE.
**Abdeckung (nach Anzahl, Vergabestatistik 2023):**
- ~12% aller Verfahren = Oberschwelle → **garantiert** zentral auf DÖE+TED (≈75% nach €-Wert).
- Unterschwelle eForms noch freiwillig (Pflicht nur Bund). Heute zentral geschätzt ~20-35% nach Anzahl, steigend.
- Restliche ~65-80% nach Anzahl (Länder/kommunal Unterschwelle, Verhandlungs-/Direktvergaben) nur auf Einzelportalen. Genau hier verdienen die 10 Portale ihren Wert; Anteil schrumpft mit kommender Unterschwellen-Pflicht.
DÖE-Datenlieferanten der Liste: aumass, RIB, cosinex, AI-AG-Plattformen → deren Oberschwellen-Daten zentral ohne Scraping.
## Empfehlung (Effort/Value)
1. **DÖE OpenData API zuerst** — eine auth-freie Integration (eForms/OCDS/CSV), ~75% Marktwert, null Scraping/ToS-Risiko. Höchster ROI.
2. **Ein AI-AG NetServer-Adapter** — Portale 4/8/9 + viele weitere AI-Portale. Bester Scraping-ROI für Unterschwellen-Long-Tail; öffentliche Suche, geringes ToS-Risiko.
3. **Ein cosinex VMP-Adapter** — DTVP (3), wiederverwendbar für viele Länder-Marktplätze (NRW, BB, NI, RLP).
4. **TED API v3** — optionale EU-Redundanz, für DE-only fast redundant zu #1.
**Meiden/depriorisieren als Scraping-Ziel:** vergabe24 (1) + aumass (5) — explizite AGB-Verbote automatisierter Zugriffe; Oberschwellen-Daten eh via #1. subreport (7): **RSS statt Scraping**.
**Querschnitts-Taktik:** Portale 2,3,4,6,7,8,9,10 haben **native gespeicherte-Suche + E-Mail-Alert**. Für den Unterschwellen-Rest: pro Portal eine gespeicherte Suche registrieren und die Benachrichtigungs-Mails ingesten — nutzt vorhandene DKV-Inbox-Infra (`ImapProvider`/`ExchangeInboxProvider`), geringeres ToS-Risiko als HTML-Scraping.
## Offene Verifikationspunkte vor Build
- Exakte Anti-Scraping-Klauseln in AGB von RIB, subreport, deutsche-evergabe, tender24.
- Ob vergabe.landbw.de den E-Mail-Alert on-page anbietet.
- Live DÖE OpenData Pagination-Parameter (Swagger JS-rendered).