docs: complete project research
Ausschreibungs-Radar (v1.1) STACK/FEATURES/ARCHITECTURE/PITFALLS research plus SUMMARY.md synthesis.
This commit is contained in:
+103
-4
@@ -1,6 +1,101 @@
|
||||
# Technology Stack
|
||||
|
||||
**Project:** Tessera - Modular Portal Platform with Marketplace
|
||||
**Project:** Tessera — Modular Portal Platform with Marketplace
|
||||
**Latest research:** 2026-07-17 (v1.1 Ausschreibungs-Radar delta) — base stack researched 2026-06-18
|
||||
**Overall Confidence:** MEDIUM for the v1.1 delta below (versions cross-verified via WebSearch + direct npm registry lookup; no Context7/docs-MCP available this session), HIGH for the v1.0 base stack (unchanged, see bottom section)
|
||||
|
||||
---
|
||||
|
||||
# v1.1 Delta — Ausschreibungs-Radar Module
|
||||
|
||||
**Domain:** German public-procurement (Vergabe) tender aggregation — API client + HTML scraping adapters + RSS + XML/JSON normalization, on top of the existing NestJS/Prisma monorepo and the DKV module's inbox/mail/scheduler infrastructure.
|
||||
|
||||
This section is a **delta only**. It assumes the validated v1.0 stack below (pnpm/Turborepo, NestJS 11, Prisma 7/PostgreSQL 16, Next.js 16, Vitest) and covers only **new** libraries needed for this module. No new `apps/web` dependencies are required — the results UI (searchable list, filters, detail view) is built entirely with the existing Next.js/shadcn/TanStack Query/Zustand stack already in place.
|
||||
|
||||
## Recommended Additions
|
||||
|
||||
### Core Technologies (new)
|
||||
|
||||
| Technology | Version | Purpose | Why Recommended |
|
||||
|------------|---------|---------|-----------------|
|
||||
| **fast-xml-parser** | 5.10.1 | Parse eForms-DE XML (DÖE) + RSS 2.0 feeds (subreport-elvis, service.bund.de) | Zero-dependency, pure-TS, handles both jobs with one library (`XMLParser`/`XMLBuilder`/`XMLValidator`). Used by Microsoft, NASA, VMware. Faster and simpler than `xml2js` (callback/SAX-based, heavier). One parser covers eForms XML *and* RSS XML — see "What NOT to Add" for why a dedicated `rss-parser` is skipped. |
|
||||
| **cheerio** | 1.2.0 | HTML parsing/traversal for the AI AG NetServer and cosinex Vergabemarktplatz scraping adapters | De-facto jQuery-style static HTML parser for Node — ~21k dependents, actively maintained (last release 2026-01-23). Wraps `parse5`/`htmlparser2`. 10–50× faster and far cheaper than a headless browser for server-rendered HTML, which both target portals produce for their public search-result pages. |
|
||||
| **csv-parse** | 7.0.1 | Parse the DÖE OpenData CSV export as a fallback/cross-check format | `stream.Transform`-based, part of the `adaltas/node-csv` toolkit, ~3100 dependents, actively released (14 days old at research time). Ships dual CJS/ESM — safe for `apps/api`'s CommonJS build. Primary DÖE ingestion should use eForms-XML or OCDS-JSON (richer, structured); CSV is secondary/verification. |
|
||||
| **tough-cookie** + **fetch-cookie** | 6.0.2 / 3.2.0 | Cookie-jar session handling for the scraping adapters (search-form POST → paginated results) | `apps/api` already standardizes on native `fetch` (see `favorites/icon-discovery.service.ts`, `calendar/providers/ics.provider.ts`) — no `axios` anywhere in the app. `fetch-cookie` wraps global `fetch` with a `tough-cookie` `CookieJar` so session cookies (e.g. `JSESSIONID` on the AI AG servlet portal) persist across a scrape run without introducing a second HTTP client family. |
|
||||
| **playwright** | 1.61.1 (conditional — see note) | Headless-browser fallback adapter, only if a specific portal's search UI turns out to require client-side JS rendering | AI AG NetServer's `PublicationControllerServlet` URL pattern is a classic Java servlet MVC front-controller (JSP/form-based) — server-rendered HTML with a session cookie, **not** an SPA, and (unlike ASP.NET WebForms) has no ViewState/EventValidation mechanic to fight. cosinex markets Vergabemarktplatz as "fully browser-based" (marketing language, not confirmed SPA), making it the more likely candidate to need JS rendering. **Build both adapters cheerio-first; only pull in `playwright` for whichever adapter's public search page is confirmed (phase-1 spike) to require JS execution.** Don't add it speculatively — it downloads browser binaries, growing the Docker image ~300MB+, which the "wartbar/verständlich" constraint argues against unless proven necessary. |
|
||||
|
||||
### Supporting / Dev Tools (optional)
|
||||
|
||||
| Tool | Purpose | Notes |
|
||||
|------|---------|-------|
|
||||
| **undici** (devDependency only) | Mock outbound `fetch` calls in Vitest tests for the DÖE client and scraping adapters (`MockAgent` + `setGlobalDispatcher`) | Node's global `fetch` is already powered by undici under the hood; adding it as a devDependency gives official, zero-extra-runtime-cost HTTP mocking without introducing `nock`/`msw` (neither currently used anywhere in the repo). Add only when writing the adapter test suite. |
|
||||
|
||||
## Installation
|
||||
|
||||
```bash
|
||||
# apps/api — core additions
|
||||
pnpm --filter @tessera/api add fast-xml-parser cheerio csv-parse tough-cookie fetch-cookie
|
||||
|
||||
# apps/api — conditional, only if phase-1 spike confirms a JS-rendered portal
|
||||
pnpm --filter @tessera/api add playwright
|
||||
pnpm --filter @tessera/api exec playwright install chromium # only the one engine needed
|
||||
|
||||
# apps/api — dev/test only, optional
|
||||
pnpm --filter @tessera/api add -D undici
|
||||
```
|
||||
|
||||
## What NOT to Add — Reuse Existing Instead
|
||||
|
||||
The DKV module already solved most of the infrastructure this feature needs. This is the most important section for scope control.
|
||||
|
||||
| Don't add | Why not | Reuse instead |
|
||||
|-----------|---------|----------------|
|
||||
| **A dedicated DÖE / eForms / OCDS SDK** | None exists as a maintained npm package for the DÖE OpenData API specifically. It's a plain REST/Swagger endpoint returning XML/JSON/CSV — a heavy SDK would be unmaintained-dependency risk for zero benefit. | Native `fetch` (already the established pattern) + `fast-xml-parser` for eForms-XML + plain `JSON.parse` for OCDS (it's just JSON, no schema library needed on day 1) + `csv-parse` for the CSV fallback. |
|
||||
| **`rss-parser`** | Latest release is 3.13.0 from **2023-04-11** — 3+ years stale (no security concern since RSS 2.0 is a frozen spec, but it's an extra dependency for a job `fast-xml-parser` already does). Only 2 feeds to consume (subreport-elvis, service.bund.de), both plain RSS 2.0 — a ~20-line field mapper (`title`/`link`/`pubDate`/`guid`/`description`) over `fast-xml-parser`'s output is simpler than a second library's API. | `fast-xml-parser` (already added for eForms) + a small internal `parseRssFeed()` helper. |
|
||||
| **`axios` / `axios-cookiejar-support`** | Would introduce a second HTTP client family. `apps/api` has zero `axios` usage today — both existing HTTP call sites use native `fetch`. | Native `fetch` + `fetch-cookie` (wraps `fetch`, not `axios`). |
|
||||
| **`node-fetch`** | Obsolete since Node 18 ships `fetch` natively (powered by undici); only relevant for Node ≤16 or legacy stream-handling quirks, neither applies here. | Native global `fetch`. |
|
||||
| **`p-queue`** (v7+) | **ESM-only** — `apps/api`'s `tsconfig.json` is `"module": "commonjs"` (confirmed). A native-ESM package in a CJS NestJS build means either a dynamic `import()` workaround (the codebase already has one messy precedent for `cron` in `dkv-scheduler.service.ts`) or build breakage. Overkill anyway: the module polls a handful of sources (DÖE, 1 AI-AG adapter, 1 cosinex adapter, 2 RSS feeds, 1 inbox) on independent cron ticks — no need for a promise-concurrency-queue library. | A ~20-line internal `politeDelay(ms)` / sequential-`for`-loop helper between requests within one adapter run. Simpler, CJS-native, matches the "wartbar für Nicht-Programmierer" constraint. |
|
||||
| **`bottleneck`** | Latest release 2.19.5 is from **2019** — unmaintained for 7 years. A community fork (`@rutter/bottleneck`) exists but adds a supply-chain trust question for zero real benefit given the low concurrency needs above. | Same internal delay helper as above. |
|
||||
| **A separate portal-login-credential store for scraping** | The feasibility research's recommended tactic explicitly avoids automating authenticated scraping: for portals 2,3,4,6,7,8,9,10 (all except vergabe24/aumass, avoided entirely per their ToS), the admin registers **one saved search per portal manually**, and Tessera ingests the resulting **alert emails** — not the authenticated portal UI. Only the *public, unauthenticated* search-result pages (AI AG NetServer, cosinex DTVP) are scraped. | Reuse the existing tenant-scoped encrypted-credential pattern (`CalendarCryptoService`, AES-256-GCM, via `SettingsModule`) *only* for the inbox connection used to ingest alert emails — same shape of secret the DKV module already stores, not a new credential type. |
|
||||
| **`@nestjs-modules/mailer`** for digest/alert sending | Already an established pitfall in this codebase: it cannot change its SMTP transport after startup, so a runtime SMTP-config change in the admin UI wouldn't take effect without a service restart (documented in `dkv-mail.service.ts`). | Copy the `DkvMailService` pattern: a dedicated `AusschreibungMailService` that calls `nodemailer.createTransport()` fresh on every send, reading `SettingsService.getDecryptedSmtpConfig(tenantId)` each time. |
|
||||
| **A second cron/scheduling mechanism** | `@nestjs/schedule` + `SchedulerRegistry.addCronJob()` (dynamic, runtime-updatable) is already wired into `AppModule` and proven in `DkvSchedulerService`. | Reuse the same dynamic-cron pattern — one job per source (DÖE poll, AI-AG adapter poll, cosinex adapter poll, 2× RSS poll, inbox alert poll), each independently configurable in minutes via the admin UI, same as DKV's `pollIntervalMin`. |
|
||||
| **A new IMAP/EWS client** | `ImapProvider` and `ExchangeInboxProvider` (both implementing `InboxProvider`) already handle IMAP (imapflow) and Exchange/EWS (raw SOAP over `httpntlm`, NTLM auth) inbox polling, including attachment size limits (PDF-bomb mitigation) and credential-safe logging. | Reuse `InboxProvider`/`ImapProvider`/`ExchangeInboxProvider` directly for the Unterschwellen alert-email ingestion. Note: the current attachment filter is PDF-specific (`collectPdfParts`/`looksLikePdf`); the new module needs the **email body/links**, not attachments, for most portal alert mails (search-agent notifications are usually HTML emails linking to the tender, occasionally with a PDF). This means extending `InboxProvider` with a body-text/HTML-fetching capability, or adding a sibling interface — a phase-plan design decision, not a new dependency. |
|
||||
|
||||
## Alternatives Considered
|
||||
|
||||
| Recommended | Alternative | When to Use Alternative |
|
||||
|-------------|-------------|--------------------------|
|
||||
| `fetch` + `cheerio` (adapter-first) | `playwright` (adapter-first) | If the phase-1 spike shows a target portal's public search results are rendered/paginated via client-side JS (React/Angular) rather than server HTML — confirm per-portal before committing, don't assume both need it. |
|
||||
| `fetch-cookie` + `tough-cookie` | `got` (built-in cookie support via `got.extend({ cookieJar })`) | If the adapters later need retry/backoff/HTTP2 features beyond what a thin `fetch` wrapper offers — `got` bundles those, but it's another HTTP client family; only justified if manual retry logic (already needed per DKV's `D-16` exponential-backoff pattern) becomes unwieldy. |
|
||||
| `fast-xml-parser` for both eForms-XML and RSS | `xml2js` | Only if a specific eForms XML document needs `xml2js`'s SAX-based tolerance for malformed/legacy XML — unlikely for a DÖE-generated, schema-validated eForms-DE feed. |
|
||||
| No dedicated OCDS validation library | `ajv` + the published OCDS JSON Schema | If/when the module needs to strictly validate incoming OCDS JSON against the official schema (e.g. to catch DÖE-side data quality issues) rather than just mapping known fields defensively. Defer until proven necessary — DÖE's OCDS output is itself schema-validated upstream. |
|
||||
| Internal `politeDelay()` helper | `p-queue` (dynamic `import()`) or `bottleneck` (unmaintained fork) | Only if the module later needs true concurrent multi-portal fan-out with a shared global rate budget (e.g. dozens of AI-AG-family portals at once, per the feasibility doc's "deckt viele weitere AI-Portale" note) — at that scale a real queue library becomes worth the complexity. Not needed for the 2-adapter v1.1 scope. |
|
||||
|
||||
## Version Compatibility
|
||||
|
||||
| Package A | Compatible With | Notes |
|
||||
|-----------|------------------|-------|
|
||||
| `fast-xml-parser@5.x` | Node 18+ (project already on Node 22-class runtime per `@types/node@^22`) | Pure TS/JS, no native bindings — no compatibility risk. |
|
||||
| `cheerio@1.2.x` | Node ≥18.17 | Matches project's Node baseline. |
|
||||
| `fetch-cookie@3.x` | Native global `fetch` (Node 18+) or any WHATWG-`fetch`-compatible function | Wraps whatever `fetch` implementation is passed in — works with Node's built-in `fetch` (undici) with zero extra config. |
|
||||
| `tough-cookie@6.x` | `fetch-cookie@3.x` | `fetch-cookie` accepts any `tough-cookie`-compatible jar per its own docs; pin both current majors together. |
|
||||
| `playwright@1.61.x` (if added) | Requires downloading Chromium binary at install time (`playwright install chromium`) | Docker implication: the `apps/api` production image must run `playwright install --with-deps chromium` in its build stage — another reason to add it only if actually needed, not speculatively. |
|
||||
| `csv-parse@7.x` | Node 18+, ESM **and** CJS builds published | Unlike `p-queue`/`bottleneck`, `csv-parse` ships dual CJS/ESM — safe for the CommonJS `apps/api` build. |
|
||||
|
||||
## Sources (v1.1 delta)
|
||||
|
||||
- npm registry (`registry.npmjs.org`) — direct authoritative version/publish-date lookup for `cheerio`, `fast-xml-parser`, `csv-parse`, `rss-parser`, `playwright`, `p-queue`, `tough-cookie`, `fetch-cookie`, `bottleneck` — fetched 2026-07-17. Confirms: cheerio 1.2.0 (2026-01-23), fast-xml-parser 5.10.1 (2026-07-16), csv-parse 7.0.1 (2026-07-02), rss-parser 3.13.0 (2023-04-11, stale), playwright 1.61.1 (2026-06-23), p-queue 9.3.1 (2026-07-03, ESM-only since v7), tough-cookie 6.0.2 (2026-07-07), fetch-cookie 3.2.0 (2025-12-15), bottleneck 2.19.5 (2019-08-03, unmaintained).
|
||||
- WebSearch (MEDIUM confidence, cross-checked against npm registry above where version-critical) — cheerio/fast-xml-parser/rss-parser/csv-parse/playwright/p-queue ecosystem status; undici-vs-node-fetch guidance; tough-cookie/fetch-cookie/axios-cookiejar-support session-handling patterns; Playwright-vs-cheerio scraping tradeoff for stateful ASP.NET-style vs static HTML portals; OCDS-for-eForms mapping (no dedicated npm library found — confirmed via `standard.open-contracting.org` documentation search, not a package).
|
||||
- `https://www.oeffentlichevergabe.de/documentation/swagger-ui/opendata/` and `bescha.bund.de` DÖE pages — confirms eForms-DE/OCDS/CSV export formats, no-auth access.
|
||||
- `.planning/research/ausschreibungs-portale-feasibility.md` (this repo, 2026-07-16) — portal platform identification (AI AG NetServer = Java `ControllerServlet`, cosinex = separate incompatible HTML), which portals to scrape vs. email-alert-ingest vs. avoid entirely.
|
||||
- Codebase inspection (`apps/api/src/dkv/`, `apps/api/src/settings/`, `apps/api/tsconfig.json`, `apps/api/package.json`) — confirmed: CommonJS build target (rules out ESM-only libs), native `fetch` already the established HTTP client (rules out adding `axios`), existing `InboxProvider`/`ImapProvider`/`ExchangeInboxProvider`/`DkvMailService`/`DkvSchedulerService`/`CalendarCryptoService` patterns to reuse verbatim.
|
||||
|
||||
**Gap / low-confidence area:** No Context7 or other docs-MCP server was available in this session (`.mcp.json` only configures `playwright` for browser automation, not doc lookup) — all version numbers here come from WebSearch cross-checked directly against `registry.npmjs.org`, which is authoritative for version/publish-date facts but not for qualitative maintenance-health claims. This repo's automated confidence classifier flagged `fast-xml-parser`, `csv-parse`, `playwright`, `p-queue`, and `tough-cookie` as `SUS` — manual review indicates this is a false positive tripped by recent-publish velocity, not an actual supply-chain concern: all five are widely-adopted, well-known-maintainer packages (Playwright is Microsoft's own project). Recommend a final `npm view <pkg>` / Socket.dev spot-check immediately before `pnpm add` in the implementation phase, per this repo's existing dependency-hygiene bar.
|
||||
|
||||
---
|
||||
|
||||
# v1.0 Base Stack (reference — unchanged)
|
||||
|
||||
**Researched:** 2026-06-18
|
||||
**Overall Confidence:** HIGH
|
||||
|
||||
@@ -64,6 +159,8 @@ tessera/
|
||||
|
||||
**Why Keycloak over custom auth:** The project requires LDAP integration, multi-tenancy, admin user management, and token-based auth. Building this from scratch would take weeks and introduce security vulnerabilities. Keycloak provides all of this as a Docker container with zero custom code.
|
||||
|
||||
> **Note (v1.1):** Live implementation diverged from this section — see `apps/api/src/settings` and the LDAP work already shipped directly against `ldapts` rather than via Keycloak federation. This base-stack section is kept as originally researched; treat the "Authentication & Authorization" row above as historical context, not current fact, when planning new auth-adjacent work.
|
||||
|
||||
### Desktop Wrapper
|
||||
|
||||
| Technology | Version | Purpose | Why |
|
||||
@@ -83,6 +180,8 @@ tessera/
|
||||
| Traefik | 3.x | Reverse proxy | Automatic SSL, Docker-native service discovery, routing rules via labels. Simpler than nginx for Docker-compose setups |
|
||||
| Gitea | existing | Version control | Already in place. Automate via webhooks and Gitea API |
|
||||
|
||||
> **Note (v1.1):** Reverse proxy in production is **Nginx Proxy Manager (NPM)**, external to the Tessera Docker stack — not Traefik. Tessera containers do not include a reverse proxy; NPM handles SSL termination and routing on the host. Treat the "Traefik" row above as historical/superseded.
|
||||
|
||||
### Testing
|
||||
|
||||
| Technology | Version | Purpose | Why |
|
||||
@@ -117,7 +216,7 @@ tessera/
|
||||
| Linting | Biome | ESLint + Prettier | Single tool, 100x faster, less config. ESLint is being replaced in NestJS 12 roadmap anyway |
|
||||
| Component Library | shadcn/ui | Material UI / Ant Design | shadcn gives ownership of components (no dep lock-in), built on Radix primitives, Tailwind-native |
|
||||
| Dashboard Grid | react-grid-layout | Gridstack.js | React-native, TypeScript rewrite in v2, hooks API, responsive breakpoints |
|
||||
| Reverse Proxy | Traefik | nginx | Docker-native service discovery, auto-SSL, config via labels not files |
|
||||
| Reverse Proxy | Traefik | nginx | Docker-native service discovery, auto-SSL, config via labels not files — **superseded in practice by external Nginx Proxy Manager, see note above** |
|
||||
|
||||
## Multi-Tenancy Strategy
|
||||
|
||||
@@ -177,7 +276,7 @@ services:
|
||||
db: # PostgreSQL 16 (port 5432)
|
||||
redis: # Redis 7 (port 6379)
|
||||
keycloak: # Keycloak 26.6.x (port 8080)
|
||||
traefik: # Reverse proxy (port 80/443)
|
||||
traefik: # Reverse proxy (port 80/443) — superseded by external NPM, see note above
|
||||
```
|
||||
|
||||
## Version Pinning Strategy
|
||||
@@ -187,7 +286,7 @@ services:
|
||||
- **Tailwind/shadcn:** Follow latest within major — utility additions are non-breaking
|
||||
- **Keycloak Docker image:** Pin to minor (e.g., `quay.io/keycloak/keycloak:26.6`)
|
||||
|
||||
## Sources
|
||||
## Sources (v1.0 base)
|
||||
|
||||
- [Next.js 16 Docs](https://nextjs.org/docs/app/guides/upgrading/version-16) — Version 16.2.7+ stable
|
||||
- [NestJS Documentation](https://docs.nestjs.com/) — Version 11.1.x
|
||||
|
||||
Reference in New Issue
Block a user