16c2d5b6af
Synthesize stack, features, architecture, and pitfalls research into unified summary with roadmap implications and phase suggestions. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
283 lines
19 KiB
Markdown
283 lines
19 KiB
Markdown
# Domain Pitfalls
|
|
|
|
**Domain:** Modular portal platform with marketplace, multi-tenancy, Docker deployment
|
|
**Project:** Tessera
|
|
**Researched:** 2026-06-18
|
|
|
|
## Critical Pitfalls
|
|
|
|
Mistakes that cause rewrites, data breaches, or architectural dead-ends.
|
|
|
|
### Pitfall 1: Tenant Data Leaks via Missing tenant_id Filtering
|
|
|
|
**What goes wrong:** In a shared-database multi-tenant model, every database query must be scoped to the current tenant. A single missed WHERE clause exposes one tenant's data to another. This is not theoretical — it has caused real security incidents in production SaaS platforms.
|
|
|
|
**Why it happens:** Developers rely on remembering to add tenant filtering per query rather than enforcing it structurally. As the codebase grows, a new endpoint or background job skips the filter, and cross-tenant data leaks silently.
|
|
|
|
**Consequences:** Complete breach of tenant isolation. Legal liability. Loss of customer trust. Forces emergency patching and potential data breach notifications.
|
|
|
|
**Prevention:**
|
|
- Use PostgreSQL Row-Level Security (RLS) as a database-level safety net. Application bugs cannot bypass RLS policies.
|
|
- Set `current_setting('app.current_tenant')` on every database session/connection via middleware, then let RLS policies enforce filtering automatically.
|
|
- Create a base repository/query class that injects tenant scope — never let raw queries bypass it.
|
|
- Integration tests must explicitly verify that Tenant A cannot access Tenant B's data at the database layer.
|
|
|
|
**Detection:** Code review flag: any raw SQL or ORM query on a tenant-scoped table that does not include tenant filtering. Automated test suite with cross-tenant access attempts.
|
|
|
|
**Phase relevance:** Must be addressed in Phase 1 (core architecture). Retrofitting RLS after data exists is painful.
|
|
|
|
---
|
|
|
|
### Pitfall 2: Module API Contract Instability (Breaking Changes)
|
|
|
|
**What goes wrong:** The platform core exposes APIs to modules. As the platform evolves, these APIs change, breaking existing modules. Without versioning, every core update risks breaking the marketplace.
|
|
|
|
**Why it happens:** The module API is treated as internal code rather than a public contract. Tight coupling between core and modules means core refactoring cascades into module breakage.
|
|
|
|
**Consequences:** Module developers (even if initially just Claude) cannot trust the platform. Upgrading the core requires simultaneously updating all modules. At scale, this becomes a coordination nightmare that blocks all releases.
|
|
|
|
**Prevention:**
|
|
- Define a versioned Module SDK/API from day one (e.g., `@tessera/sdk@1.x`).
|
|
- Use semantic versioning: breaking changes require a new major version with a migration path.
|
|
- Modules declare which SDK version they target. Platform maintains backward compatibility for at least one prior major version.
|
|
- Keep the API surface small — expose only what modules actually need.
|
|
|
|
**Detection:** Any PR that modifies the module-facing interface without a version bump. Module integration tests failing after core changes.
|
|
|
|
**Phase relevance:** Phase 2 (module system design). The API contract must be defined before the first marketplace module ships.
|
|
|
|
---
|
|
|
|
### Pitfall 3: Monolithic Module Loading Destroys Isolation
|
|
|
|
**What goes wrong:** Modules run in the same process as the platform core, sharing memory, dependencies, and crash domains. A buggy module takes down the entire platform. A module with a conflicting dependency version breaks other modules.
|
|
|
|
**Why it happens:** In-process loading is simpler to implement initially. Teams plan to "isolate later" but the architecture cements coupling.
|
|
|
|
**Consequences:** One module crash = entire platform down for all tenants. Dependency conflicts between modules. Security: a malicious module can access other modules' data in-process.
|
|
|
|
**Prevention:**
|
|
- Each module runs as its own Docker container (or at minimum, its own isolated process).
|
|
- Communication between core and modules via well-defined API (REST/gRPC), not shared memory.
|
|
- Module containers have their own dependency trees — no shared node_modules or library conflicts.
|
|
- Resource limits (CPU/memory) per module container prevent a runaway module from starving the platform.
|
|
|
|
**Detection:** Architecture review: can you stop/restart one module without affecting others? If not, isolation is insufficient.
|
|
|
|
**Phase relevance:** Phase 1 (architecture decisions). This is a foundational constraint that shapes everything.
|
|
|
|
---
|
|
|
|
### Pitfall 4: Docker Network Flat Topology Enables Lateral Movement
|
|
|
|
**What goes wrong:** All containers (frontend, backend, database, modules, Redis) placed on a single Docker network. A compromised module container can directly access the database, other modules, and internal services.
|
|
|
|
**Why it happens:** Single-network setups are simpler during development. Teams use docker-compose with one default network and never segment.
|
|
|
|
**Consequences:** A vulnerability in any module gives attackers access to the database, cache, and all other modules. The blast radius of any compromise is the entire system.
|
|
|
|
**Prevention:**
|
|
- Segment Docker networks by trust boundary:
|
|
- `frontend-net`: reverse proxy + frontend containers
|
|
- `backend-net`: API server + module containers
|
|
- `data-net`: database + cache (only accessible from API server)
|
|
- Module containers CANNOT reach the database directly — they communicate only with the platform API.
|
|
- Never publish internal ports (PostgreSQL 5432, Redis 6379) to the host unless explicitly needed for development.
|
|
- Use Docker Compose network aliases to control service discovery.
|
|
|
|
**Detection:** Run `docker network inspect` — if all containers appear on one network, isolation is broken. Penetration test: from a module container, can you reach the database port?
|
|
|
|
**Phase relevance:** Phase 1 (infrastructure setup). Docker Compose network topology must be designed before anything else runs.
|
|
|
|
---
|
|
|
|
### Pitfall 5: LDAP as Primary Authentication Instead of Identity Layer
|
|
|
|
**What goes wrong:** The application directly binds to LDAP with user credentials for authentication (the "simple bind" pattern). Users' passwords pass through the application, which must handle them in cleartext. The application becomes a credential-harvesting target.
|
|
|
|
**Why it happens:** LDAP bind is the most "obvious" integration — send username/password to LDAP, check if bind succeeds. It works, so teams ship it without considering the security implications.
|
|
|
|
**Consequences:** If the application server is compromised, all credentials used during the compromise window are exposed. Cannot add MFA later (LDAP bind is password-only). Cannot support SSO without rearchitecting.
|
|
|
|
**Prevention:**
|
|
- Use LDAP for user directory sync (group membership, attributes) but NOT as the authentication mechanism.
|
|
- Implement authentication via a proper identity layer: local password hashing (bcrypt/argon2) for the initial admin flow, LDAP sync for user provisioning, and optionally OIDC/SAML for enterprise SSO later.
|
|
- If LDAP bind is required for compatibility: always use LDAPS (port 636), validate certificates, use a dedicated service account for directory lookups, and rate-limit bind attempts.
|
|
- Architecture should allow swapping the auth backend without rewriting the application.
|
|
|
|
**Detection:** Code that passes raw user passwords to an LDAP bind operation without TLS. No abstraction layer between "authenticate user" and "LDAP bind."
|
|
|
|
**Phase relevance:** Phase 2 (authentication system). Design the auth abstraction layer first, then implement LDAP as one provider behind it.
|
|
|
|
---
|
|
|
|
## Moderate Pitfalls
|
|
|
|
### Pitfall 6: i18n Retrofitting After UI Is Built
|
|
|
|
**What goes wrong:** UI components use hardcoded strings. When i18n is added later, every component must be touched. German text is 30-40% longer than English, breaking layouts that were designed for English string lengths.
|
|
|
|
**Why it happens:** "We'll add translations later" feels reasonable because it seems like a search-and-replace task. In reality, it requires layout redesign, context-aware pluralization, date/number formatting changes, and string extraction tooling.
|
|
|
|
**Prevention:**
|
|
- Use i18n from the very first component. Every user-visible string goes through a translation function (`t('key')`).
|
|
- Design layouts with flexible widths — avoid fixed-width containers for text.
|
|
- Use ICU MessageFormat for pluralization rules (German has different plural forms than English).
|
|
- Store translations in separate JSON files per locale, never inline.
|
|
- Test the UI with the longest locale (German) as the default during development to catch overflow early.
|
|
|
|
**Detection:** Any user-visible string literal in component code that is not wrapped in a translation call. Layout breakage when switching to German.
|
|
|
|
**Phase relevance:** Phase 1 (frontend scaffolding). i18n infrastructure must be present from the first rendered component.
|
|
|
|
---
|
|
|
|
### Pitfall 7: Dashboard Widget State Explosion
|
|
|
|
**What goes wrong:** The drag-and-drop dashboard stores layout state per user per tenant, with frequent updates on every drag/resize event. This generates massive write amplification to the database and state management complexity.
|
|
|
|
**Why it happens:** Naive implementations persist layout on every mouse event. Combined with multi-tenancy, the permutations of (user x tenant x widget configuration) grow rapidly.
|
|
|
|
**Prevention:**
|
|
- Debounce layout persistence — save only after drag/resize completes (on mouse-up), not during.
|
|
- Store layout as a single JSON column per user-dashboard, not as individual widget position rows.
|
|
- Use optimistic UI updates with background sync — don't block on database writes.
|
|
- Limit the maximum number of widgets per dashboard (e.g., 20) to bound rendering complexity.
|
|
- Use a proven grid library (react-grid-layout or gridstack.js) rather than building from scratch.
|
|
|
|
**Detection:** Database write frequency during dashboard interaction. UI jank when >10 widgets are rendered simultaneously.
|
|
|
|
**Phase relevance:** Phase 3 (dashboard implementation). Choose the grid library and persistence strategy before building widgets.
|
|
|
|
---
|
|
|
|
### Pitfall 8: Licensing State Scattered Across the System
|
|
|
|
**What goes wrong:** Module activation/licensing checks are scattered throughout the codebase — in route guards, API middleware, UI rendering logic, and module loaders. When the licensing model changes, dozens of locations need updating.
|
|
|
|
**Why it happens:** Each feature gate feels like a simple `if (licensed)` check. Over time, these proliferate into an unmaintainable web.
|
|
|
|
**Prevention:**
|
|
- Centralize licensing in a single service/module: `LicenseService.canAccess(tenant, module) -> boolean`.
|
|
- All other code calls this service — no direct database checks for license status.
|
|
- Cache license state per tenant (licenses change rarely — invalidate on admin action only).
|
|
- Module containers should not even start/be routable if the tenant lacks a license — enforce at the orchestration layer, not inside the module.
|
|
|
|
**Detection:** Grep for license-related conditionals outside the license service. Multiple database tables tracking activation state.
|
|
|
|
**Phase relevance:** Phase 2 (marketplace/module system). The license service must exist before the first licensable module.
|
|
|
|
---
|
|
|
|
### Pitfall 9: Desktop Wrapper (Electron/Tauri) Divergence from Web
|
|
|
|
**What goes wrong:** The desktop wrapper introduces platform-specific behavior (file system access, window management, notifications) that diverges from the web version. Bugs exist only in the wrapper. Features work in-browser but not in the desktop app.
|
|
|
|
**Why it happens:** Teams treat the wrapper as "just a browser window" but it isn't — WebView quirks, different security contexts, and OS integration create a separate test surface.
|
|
|
|
**Prevention:**
|
|
- Keep the desktop wrapper as thin as possible — it should be a shell around the same web app, not a fork.
|
|
- Use Tauri over Electron for smaller binary size and better security defaults (Tessera is already considering this).
|
|
- Run the same URL in the wrapper that the browser uses (point to localhost or the deployed server).
|
|
- Do NOT add desktop-only features to the web codebase — keep OS integration in the wrapper layer only.
|
|
- Accept that the desktop wrapper is a Phase 4+ concern — do not build it until the web platform is stable.
|
|
|
|
**Detection:** Features that work in Chrome but fail in the desktop app. Desktop-specific code paths in the shared frontend codebase.
|
|
|
|
**Phase relevance:** Phase 4+ (desktop client). Do not start this until the web platform is feature-complete and stable.
|
|
|
|
---
|
|
|
|
### Pitfall 10: Tenant Context Lost in Async Operations
|
|
|
|
**What goes wrong:** Background jobs, event handlers, and scheduled tasks lose the tenant context that was present during the original HTTP request. Jobs process data without tenant scoping, or worse, process one tenant's job with another tenant's context.
|
|
|
|
**Why it happens:** HTTP middleware sets tenant context on the request. But when a job is enqueued, the worker that picks it up has no request context. If the job payload doesn't explicitly include tenant_id, the worker either fails or defaults to a wrong/global context.
|
|
|
|
**Consequences:** Data corruption across tenants. Background jobs that silently operate on wrong tenant's data. Difficult to reproduce because it depends on job scheduling order.
|
|
|
|
**Prevention:**
|
|
- Every job payload MUST include `tenant_id` as a required field — enforce via TypeScript types or schema validation.
|
|
- Worker initialization must set the tenant context (including database RLS session variable) before processing any job.
|
|
- Never rely on ambient/global state for tenant identification in async contexts.
|
|
- Add assertion checks: if a job's tenant_id doesn't match the database session's tenant, abort with an error.
|
|
|
|
**Detection:** Jobs that succeed without a tenant_id in their payload. Log analysis showing jobs processing data from multiple tenants in a single execution.
|
|
|
|
**Phase relevance:** Phase 2 (background processing). Must be enforced from the first background job.
|
|
|
|
---
|
|
|
|
## Minor Pitfalls
|
|
|
|
### Pitfall 11: Docker Volume Permission Mismatches
|
|
|
|
**What goes wrong:** Containers run as non-root (correctly) but mounted volumes have root ownership from the host. The application cannot write to its data directory, causing silent failures or crashes on startup.
|
|
|
|
**Prevention:** Explicitly set user/group in Dockerfile. Use named volumes instead of bind mounts in production. Set permissions in an entrypoint script. Test with `docker-compose up` from a fresh state.
|
|
|
|
**Phase relevance:** Phase 1 (Docker setup).
|
|
|
|
---
|
|
|
|
### Pitfall 12: Over-Engineering the Module Communication Protocol
|
|
|
|
**What goes wrong:** Teams design complex message buses, event sourcing, or gRPC streaming for module communication when simple REST calls would suffice. The protocol becomes harder to debug than the business logic.
|
|
|
|
**Prevention:** Start with synchronous REST between core and modules. Add async messaging only when you have a proven need (e.g., long-running tasks). Keep the protocol debuggable with standard HTTP tools (curl, Postman).
|
|
|
|
**Phase relevance:** Phase 2 (module system). Start simple, evolve based on real needs.
|
|
|
|
---
|
|
|
|
### Pitfall 13: Gitea Automation Tight Coupling
|
|
|
|
**What goes wrong:** The platform's deployment pipeline is so tightly integrated with Gitea that changing version control systems or CI tools requires rewriting the deployment process.
|
|
|
|
**Prevention:** Abstract the VCS integration behind an interface. Use standard Git operations rather than Gitea-specific APIs where possible. Keep CI/CD configuration in standard formats (Dockerfiles, compose files) rather than Gitea-specific workflows exclusively.
|
|
|
|
**Phase relevance:** Phase 3+ (CI/CD integration). Design the abstraction before implementing Gitea-specific hooks.
|
|
|
|
---
|
|
|
|
### Pitfall 14: Theme System Without Design Tokens
|
|
|
|
**What goes wrong:** Light/dark theme is implemented with scattered CSS variables or conditional classes. Adding a third theme or adjusting the color palette requires touching hundreds of files.
|
|
|
|
**Prevention:** Use a design token system from day one — a single source of truth for colors, spacing, typography. Themes are just alternative token sets. Use CSS custom properties at the `:root` level with a theme class toggle.
|
|
|
|
**Phase relevance:** Phase 1 (frontend scaffolding). Design tokens must be established before the first component is styled.
|
|
|
|
---
|
|
|
|
## Phase-Specific Warnings
|
|
|
|
| Phase Topic | Likely Pitfall | Mitigation |
|
|
|-------------|---------------|------------|
|
|
| Core architecture | Tenant isolation not enforced at DB level | Implement PostgreSQL RLS from day one |
|
|
| Core architecture | Flat Docker network | Design network segmentation in initial docker-compose.yml |
|
|
| Core architecture | i18n deferred | Set up translation infrastructure with first component |
|
|
| Authentication | LDAP as auth instead of directory sync | Build auth abstraction layer; LDAP is one provider |
|
|
| Module system | No versioned API contract | Define module SDK with semantic versioning before first module |
|
|
| Module system | In-process module loading | Each module = own container from the start |
|
|
| Module system | License checks scattered | Centralized LicenseService as single gate |
|
|
| Dashboard | Widget state write amplification | Debounced persistence, JSON column, proven grid library |
|
|
| Background jobs | Tenant context lost in workers | Mandatory tenant_id in all job payloads |
|
|
| Desktop client | Built too early, diverges from web | Defer until web is stable; keep wrapper minimal |
|
|
| Deployment | Volume permissions break non-root containers | Named volumes, entrypoint permission scripts |
|
|
| VCS integration | Gitea-specific tight coupling | Abstract behind interface, use standard Git ops |
|
|
|
|
## Sources
|
|
|
|
- [Multi-Tenant SaaS Architecture: What Nobody Tells You Before You Build](https://dev.to/actinode/multi-tenant-saas-architecture-what-nobody-tells-you-before-you-build-a4h) - Confidence: HIGH
|
|
- [Designing Multi-Tenant SaaS Architecture: Mistakes to Avoid](https://www.saasadviser.co/blog/multi-tenant-saas-architecture-mistakes-best-practices) - Confidence: MEDIUM
|
|
- [Multi-Tenant Row-Level Security in PostgreSQL: A Production Pattern](https://dev.to/uaslimcreate/building-multi-tenant-row-level-security-in-postgresql-a-production-pattern-4n2k) - Confidence: HIGH
|
|
- [Docker Network Isolation Pitfalls](https://hexshift.medium.com/docker-network-isolation-pitfalls-that-put-your-applications-at-risk-b60356a14033) - Confidence: MEDIUM
|
|
- [The LDAP Authentication Anti-Pattern](https://blog.lithnet.io/2018/03/the-ldap-authentication-anti-pattern.html) - Confidence: HIGH
|
|
- [Plugin Versioning Best Practices](https://devactivity.com/insights/mastering-plugin-versioning-ensuring-compatibility-for-your-software-measurement-tool/) - Confidence: MEDIUM
|
|
- [Feature Toggles (Martin Fowler)](https://martinfowler.com/articles/feature-toggles.html) - Confidence: HIGH
|
|
- [Common Technical Challenges in i18n](https://activeloc.com/blog/i18n-technical-challenges/) - Confidence: MEDIUM
|
|
- [Electron vs. Tauri](https://www.dolthub.com/blog/2025-11-13-electron-vs-tauri/) - Confidence: MEDIUM
|
|
- [Building Interactive Dashboards with React Grid Layout](https://www.ilert.com/blog/building-interactive-dashboards-why-react-grid-layout-was-our-best-choice) - Confidence: MEDIUM
|
|
- [AWS: Multi-tenant data isolation with PostgreSQL Row Level Security](https://aws.amazon.com/blogs/database/multi-tenant-data-isolation-with-postgresql-row-level-security/) - Confidence: HIGH
|