Commit Graph

5 Commits

Author SHA1 Message Date
schalli a44e40f101 fix(dkv): correct invoice date extraction; add Exchange IsRead filter + mark-as-read
Tessera CI/CD / Lint & Type Check (push) Successful in 44s
Tessera CI/CD / Tests (push) Successful in 44s
Tessera CI/CD / Build & Publish Images (push) Successful in 21s
Date fix: previous regex matched payment-due date ("10 Tage nach Rechnungsdatum...
10.04.2026") instead of actual Rechnungsdatum. New approach anchors on the
invoice number line (DD/DDDDDDDDD/DDD) and takes the date on the next line,
which is always the actual Rechnungsdatum in DKV PDFs.

Exchange dedup: FindItem now filters IsRead=false (combined with sender filter
via <t:And>), so already-processed emails are skipped automatically.
After downloading attachments, UpdateItem marks the message as read
(using ItemId + ChangeKey from GetItem response), mirroring IMAP \Seen behavior.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-30 09:47:17 +02:00
schalli 66ffad149e fix(dkv): extract invoice number/date from PDF, rename export files to RG-DKV format
Tessera CI/CD / Lint & Type Check (push) Successful in 41s
Tessera CI/CD / Tests (push) Successful in 38s
Tessera CI/CD / Build & Publish Images (push) Successful in 23s
- Parser now extracts Rechnungsnummer (DD/DDDDDDDDD/DDD) and Rechnungsdatum
  from PDF text, so filename doesn't rely on email subject
- Export filename changed from DKV_YYYY-MM_... to RG-DKV-{nr}-{YYMMDD}.xlsx
  e.g. RG-DKV-26-650869002-002-260331.xlsx
- Subject fallback now also matches slash-separated invoice numbers (26/NNN/NNN)
- writeAndPrune simplified to accept baseName instead of separate fields
- Validation regex and prune prefix updated to match new RG-DKV- pattern

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-30 09:39:51 +02:00
schalli 5c1aa03270 fix(dkv): handle EV charging rows and service rows in single-tx tab parser
Tessera CI/CD / Lint & Type Check (push) Successful in 40s
Tessera CI/CD / Tests (push) Successful in 37s
Tessera CI/CD / Build & Publish Images (push) Successful in 21s
Analyzed Invoice-4302486921-26_650869002_000.pdf text structure. Three cases:

1. FUEL (fields[4] = numeric tx-nr): km+product merged in fields[5], unit in
   fields[6]. Already working; no change.

2. EV CHARGING (fields contains "DDDD KWH" or "DDDD MIN" unit): column layout
   shifts — no km field, station+ort sometimes merged in fields[1]. Detected by
   regex on unit field; kwhIdx drives relative offset for menge/netto/brutto.
   Ort extracted from fields[2] (kwhIdx>=5) or fields[1] (kwhIdx=4, compact).
   Kilometerstand = 0 (EV chargers don't record odometer).

3. SERVICE ROWS (e.g. "DKV Analytics Premiu"): appear inside a VEHICLE: block but
   fields[4] is non-numeric (product description, not a transaction number). These
   were being parsed as fake vehicle transactions producing wrong ort/km values.
   Now filtered out (return null).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-30 09:20:12 +02:00
schalli a759e816a0 fix(dkv): fix Kennzeichen matching and NaN/invalid km values in Excel export
Tessera CI/CD / Lint & Type Check (push) Successful in 43s
Tessera CI/CD / Tests (push) Successful in 38s
Tessera CI/CD / Build & Publish Images (push) Successful in 23s
Three issues fixed:

1. Kennzeichen normalization: DKV PDF extracts plates without hyphens
   ("GP JL 740E" vs CSV-imported "GP-JL 740E"). Added _normalizeKennzeichen()
   which strips hyphens, spaces, and dots before lookup — resolves vehicle
   master match failure that caused Marke/Modell/Fahrer to appear empty.

2. Empty-string NaN: parser used ?? '0' which doesn't catch empty strings,
   causing parseDE('') = NaN. Changed to || '0' for km, menge, and totals.

3. Invalid km values: EV charging rows from DKV have misaligned columns —
   km position contains a decimal price (e.g. 18.64 EUR or kWh). Added
   sanity check: non-integer km values are written as null (empty cell)
   instead of a misleading decimal. ExportRow.kilometerstand is now number|null.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-30 09:12:13 +02:00
schalli 6235aaf31c feat(07-01): DKV PDF parser validated against real invoice + injectable service
Tessera CI/CD / Build & Deploy (push) Blocked by required conditions
Tessera CI/CD / Lint & Type Check (push) Successful in 41s
Tessera CI/CD / Tests (push) Waiting to run
- dkv-parser.validate.ts: empirical validation script against user-files/invoice.pdf
  - 27 vehicle blocks, 66 transactions extracted (matches expected count)
  - Handles two PDF extraction formats: single-tx (tab-separated) + multi-tx (columnar)
  - German number parsing: replace(/\./g,'').replace(',','.') applied to km and menge
  - Exits 1 with full raw text dump if zero vehicle blocks parsed (assertion guard)
- dkv-parser.service.ts: @Injectable() NestJS service wrapping validated logic
  - parsePdf(buffer: Buffer): Promise<DkvVehicleBlock[]>
  - Uses pdf-parse v2 class API: new PDFParse({data:buffer}) — NOT v1 pdfParse()
  - Calls destroy() after extraction (T-07-01 memory safety)
  - Generic error messages only on parse failure (T-07-02 info disclosure)

Regex adjustment vs Research Pattern 4: space-based regex replaced with
tab-split (single-tx) + columnar transpose (multi-tx) after empirical analysis
of actual invoice.pdf text extraction output.
2026-06-26 19:36:45 +02:00