0010: Enterprise Platform + SaaS + Global Infrastructure (Module 10)
Context
Modules 1-9 built a complete single-tenant-feeling (though Workspace-scoped
since Module 1) bug bounty / pentesting platform: auth, recon, vulnerability
scanning, AI copilot, a full bug bounty operating system, a Desktop Agent, and
a Browser Extension. Module 10's spec asks for the platform to become
something a real company can run as infrastructure: real multi-tenant
organizations above the existing Workspace layer, a distributed job system
that scales scan execution beyond one process, a plugin marketplace, native
third-party integrations, team workspaces, enterprise reporting/analytics, a
public API platform (REST, GraphQL, webhooks, SDKs in three languages), a
visual workflow automation engine, compliance tooling (retention, backup,
GDPR export), and the operational trilogy every real deployment eventually
needs: observability, high availability, and security hardening — plus CI/CD
and deployment templates to actually run it. Fourteen subsystems, the largest
module by scope in this project's history. This ADR covers the decisions that
span multiple subsystems; each subsystem's own file-level doc comments carry
the narrower, local decisions.
Decision
1. Organization sits above Workspace, doesn't replace it
Every prior module's tenant boundary is Workspace. Module 10 adds
Organization as a new, optional parent: a Workspace can belong to a
Team, which belongs to an Organization, or remain organization-less
exactly as before. This is additive in the strongest sense — no existing
Workspace-scoped query changes shape, no existing route's authorization
model changes. Organization exists specifically for the new Module 10
surfaces that only make sense above workspace scope: billing/seats,
organization-wide RBAC roles, the Distributed Job queue, Plugin
installations, Integration connections, Enterprise Reporting/Analytics, and
the GraphQL/Webhook/SDK platform. A user or workspace that never touches any
Module 10 feature never encounters Organization at all.
2. Distributed Job System: claim-based work queue, not a message broker
DistributedJob/WorkerNode implement a pull-based queue directly on
Postgres — workers claim() the next QUEUED job matching their declared
capabilities via a SELECT ... FOR UPDATE SKIP LOCKED-style claim (the
same optimistic-claim pattern ReconJobPollerService established back in
Module 3), rather than introducing a message broker (RabbitMQ, SQS, Redis
Streams). This keeps the "zero new infrastructure dependency" posture every
module since Module 1 has held to, at the cost of not getting a broker's
built-in features (priority queues, dead-letter handling, fanout) for free —
DistributedJobPriority and job-level retry/lastError tracking reimplement
the pieces of that this project actually needs on top of Postgres directly.
3. Plugin Marketplace: capability-scoped, permission-accepted, sandboxed by contract not by runtime
Plugin/PluginInstallation/PluginVersion model a marketplace where each
plugin declares required permissions (network:fetch, specific entity
read/write scopes) that an installing admin must explicitly accept
(acceptedPermissions on PluginInstallation). This is a contractual
sandbox, not a runtime one — there is no separate process/container isolating
a plugin's hook execution from the rest of the API process. A plugin hook
today runs in-process via ExecutePluginHookCommand. This is a disclosed,
deliberate v1 scope boundary: real code-execution isolation (a WASM runtime,
a separate worker process with a locked-down syscall filter, or similar) is
flagged as the concrete follow-up before accepting untrusted third-party
plugin code in a real multi-tenant SaaS deployment — the permission model is
real and enforced, but it's an honor-system contract against a plugin's
declared behavior, not a technical guarantee against malicious code.
4. Integrations framework: capability registry + generic action executor
IntegrationCapability (what each provider supports: authMethod,
actions) is a static registry entry per provider; IntegrationConnection
is a tenant's actual OAuth/token-authenticated connection to one. Every
provider-specific action (create a Jira ticket, post a Slack message, sync a
GitHub issue) goes through one generic ExecuteIntegrationActionCommand
rather than one command class per provider action — this is what lets
Module 10's other subsystems (the Automation Engine's create_ticket/
sync_data step kinds in particular) integrate with any connected
provider through a single CommandBus.execute() call without importing
IntegrationsModule or knowing which providers exist.
5. Enterprise Reporting/Analytics reuse Module 6/9's rendering, don't reimplement it
ReportContentBuilderService (Enterprise Reporting) computes an
organization/workspace-scoped content payload and is the single code path
both POST .../reports/generate (on-demand) and ReportScheduleRunnerService
(scheduled) call — matching the Module 9 pattern where
BugReportGeneratorService's five renderers are the one code path for both
manual and (now, transitively, via the Automation Engine's report step
kind) automated report generation. Enterprise Analytics similarly computes
its aggregates via read-time Prisma aggregation queries scoped by
organization, not a separate pre-computed data warehouse — appropriate at
this project's scale; a real high-volume SaaS deployment would eventually
want a proper OLAP/warehouse layer, noted as a scaling follow-up rather than
built speculatively now.
6. API Platform: three protocols (REST, GraphQL, Webhooks), three SDK languages, one shared DTO source of truth
packages/shared/src/*.ts remains the single source of truth for every
DTO/endpoint-path constant, exactly as established in Module 1 — the GraphQL
resolvers (Module 10, new) return the same DTOs the REST controllers do
rather than defining a parallel GraphQL-specific type layer, and all three
SDKs (TypeScript, Python, Go) were hand-written by cross-referencing those
same shared endpoint/DTO definitions rather than generated from an OpenAPI
spec — see each SDK's own README for its specific verification status. The
webhook event catalog (WEBHOOK_EVENT_TYPES) was corrected during this
module to match real domain-event verb strings after discovering the
original catalog (authored ahead of the events that would eventually
populate it) never actually matched anything real events emit — see
webhook-signature.util.ts's toWebhookEventKey() for the normalization
this required.
7. Webhook delivery: enqueue synchronously, deliver asynchronously
WebhookDispatchHandler (an @EventsHandler over ~25 domain event classes)
only inserts WebhookDelivery rows — fast, synchronous, no network I/O
inline with the request/event that triggered it. WebhookDeliveryQueueService
(a separate minute-tick poller) performs the actual signed HTTP delivery with
exponential backoff ([1, 5, 15, 60, 180] minutes, 6 attempts before
DEAD_LETTER). This same split — fast synchronous enqueue, slow async
delivery on a separate tick — is reused for the Automation Engine's workflow
triggering (TriggerWorkflowCommand calls WorkflowExecutorService.start()
directly since workflow execution, unlike webhook delivery, has no external
network dependency in its common path) and for every other Module 9/10
scheduled-work service.
8. Automation Engine: a small step-graph interpreter, not a general workflow language
WorkflowExecutorService.executeStep() supports nine step kinds (recon,
scan, report, notify, ai_analysis, create_ticket, sync_data,
condition, loop, delay) via a next/whenTrue/whenFalse/onFailure
step graph, not an embedded scripting language or a dependency on an external
workflow engine (Temporal, Airflow, n8n). condition branches control flow
in-process; delay is the one kind that doesn't complete synchronously — it
sets pausedAtStepKey/resumeAt on the WorkflowRun and returns, and
WorkflowRunnerService's tick resumes it once due. recon/scan steps
honestly record the intent to trigger a job ({note: 'Recorded — wire to ReconJob/VulnScanJob creation...'}) rather than fabricating a job from
incomplete parameters, mirroring the exact same honest-recording pattern
ScheduledJobRunnerService established in Module 9 for its own
DAILY_RECON/WEEKLY_NUCLEI kinds — a generic scheduler/workflow engine
structurally lacks the concrete targetId + tool selection a real recon/scan
job needs, and fabricating one would be worse than recording the gap
honestly.
9. Compliance: real enforcement, narrow and disclosed scope
DataRetentionEnforcerService only auto-deletes two safely-scoped
operational tables (API_REQUEST_LOG, WEBHOOK_DELIVERY) — both have a
clear createdAt cutoff and are disposable log data, unlike e.g.
AiMessage or BugBountyFinding rows, which are user content an automated
deletion policy shouldn't touch without much more deliberate scoping. A
retention policy for any other resourceType is stored and returned by the
API (so an operator can document their retention posture even before
automatic enforcement covers it) but logged-and-skipped rather than silently
doing nothing forever unnoticed. ComplianceExportProcessorService's GDPR
export scope (v1: AUDIT_EVENTS, ORGANIZATION_MEMBERS, WORKFLOWS,
INTEGRATION_CONNECTIONS) was caught missing an organization-membership
filter on AuditEvent during implementation — AuditEvent has no
organizationId column (it's user-scoped, not tenant-scoped), so the
original where: scope.subjectUserId ? {...} : {} fallback would have
exported every user's audit trail system-wide to any org admin. Fixed by
resolving organizationMember userIds first and filtering
AuditEvent.userId IN (...) — see that service's file for the corrected
implementation.
10. Observability: hand-rolled Prometheus registry, real OpenTelemetry tracing, structured JSON logging — zero-to-minimal new dependencies
MetricsRegistryService is a from-scratch Counter/Histogram
implementation rendering real Prometheus text-exposition format, not
prom-client — consistent with this project's dependency posture, and
simple enough (two metric types, a documented text format) that a library
adds more surface than it saves. Distributed tracing (tracing.ts) is the
one disclosed exception: real @opentelemetry/sdk-node +
exporter-trace-otlp-http were added as genuine new dependencies (not
verified installed/built in this sandbox — see that file's own disclosure
comment), since correctly implementing the OTLP wire protocol and W3C
trace-context propagation from scratch would be reinventing a
standards-compliance-critical wheel poorly. withSpan() is wired into two
representative call sites (WorkflowExecutorService.executeStep,
WebhookDeliveryQueueService's outbound fetch()) rather than exhaustively
into every service — a disclosed narrower v1 scope, not an oversight.
StructuredLoggerService (JSON-lines, dependency-free, implements Nest's
LoggerService) replaces the default console logger app-wide via
NestFactory.create(AppModule, { logger: new StructuredLoggerService() }).
Two-tier metrics endpoint security: the global /observability/metrics
endpoint (process-level gauges + the two Counters above) is @Public() with
an optional shared-secret check (open by default — correct for local/
sandboxed environments; Prometheus scrapers have no OAuth/JWT identity to
present in the first place, so network-level access control, not application
auth, is the standard way to protect a metrics endpoint in production). The
per-organization queue-metrics endpoint
(/enterprise/organizations/:id/distributed-jobs/metrics/prometheus) makes
the same token check mandatory instead — this route exposes tenant-scoped
data with no other authorization at all. An earlier draft of this endpoint
reused GetQueueMetricsQuery (which authorizes via the requesting user's
org membership) with a fabricated 'system' userId on a @Public() route —
a confused-deputy/enumeration risk an unauthenticated caller could have used
to probe org membership via a ?userId= parameter. Fixed by computing the
same aggregates directly via Prisma with zero user-identity involvement
instead.
11. High Availability: PollerLeaseService, a Postgres row-based distributed lease — not pg_advisory_lock, not a new broker
Every setInterval-based poller added across Modules 9-10
(ScheduledJobRunnerService, ReportScheduleRunnerService,
WebhookDeliveryQueueService, WorkflowRunnerService,
DataRetentionEnforcerService, BackupPolicyRunnerService,
ComplianceExportProcessorService) originally ticked independently on every
replica of the API process — harmless at one replica, but a real correctness
problem at N (duplicate webhook deliveries, duplicate destructive retention
deletions, duplicate workflow-run resumptions). PollerLeaseService
(apps/api/src/common/concurrency/poller-lease.service.ts) fixes this with
one new table (PollerLease: id = lease name, holderId, expiresAt) and
a conditional UPDATE ... WHERE id = ? AND (expiresAt < now() OR holderId = ?) compare-and-swap — Postgres's own row-level locking during that UPDATE
statement is what actually prevents two replicas from both believing they
won the lease, not application-level reasoning. This was chosen over
pg_advisory_lock/pg_try_advisory_xact_lock specifically because those
either require pinning the entire critical section to one Postgres
connection (awkward with Prisma's pooled driver adapter) or to one open
transaction for the section's full duration (undesirable when the guarded
work makes outbound HTTP/AI calls, as webhook delivery and workflow
execution both do) — a plain leased row avoids both, at the cost of being a
lease with a TTL (not an unconditional mutex), which is the correct trade-off
for this use case regardless: a crashed replica's stale lease expires and
fails over automatically rather than deadlocking the tick forever.
Every service's tick() was refactored to tick() { withLease(name, ttl, () => this.runTick()) } — a small, uniform, mechanical change applied
identically across all seven services rather than seven bespoke locking
implementations.
Beyond the poller lease, HA/performance work in this pass: main.ts now
calls app.enableShutdownHooks() so SIGTERM runs every service's
onModuleDestroy (clearing intervals, releasing lease holds, closing the
Prisma pool) before the process exits, instead of the process being killed
mid-request; and DATABASE_POOL_MAX (env.validation.ts, wired into
PrismaService's PrismaPg adapter) makes the per-replica Postgres
connection pool size explicit and tunable rather than left at
node-postgres's library default, since the correct value scales inversely
with replica count.
12. Security Hardening: tuned CSP/HSTS, not a redesign of Module 1's auth model
Module 1 already implemented the bulk of this project's auth hardening
(refresh token rotation, account lockout, OAuth CSRF state, MFA, argon2
hashing) and Module 9 added encryption/RBAC/audit history — this pass's
concrete addition is fixing helmet()'s bare default configuration, which
would have broken SwaggerModule.setup()'s self-hosted Swagger UI page in
any deployment that actually enabled it (the default CSP's script-src 'self'/style-src 'self' blocks Swagger UI's inline bundle) while adding
explicit HSTS (max-age=180d, includeSubDomains, deliberately not
preload — that's a one-way opt-in to browsers' hardcoded HSTS list this
application shouldn't make on a deployment's behalf) and an explicit
frameguard: deny clickjacking policy. See bootstrap.ts's
configureApp() for the full configuration and reasoning. A genuinely new
auth-model redesign (e.g. moving from bearer tokens to httpOnly session
cookies) was considered out of scope — it would be exactly the kind of
broad Module 1 rewrite the standing "don't rewrite previous modules"
instruction guards against, and bearer-token SPA auth is a legitimate,
common architecture, not a defect.
13. CI/CD + Deployment: four independent paths sharing one build
.github/workflows/ci.yml runs lint/typecheck/test/build via Turborepo
across the whole workspace, with a pgvector/pgvector Postgres service
container for apps/api's Jest suite (prisma db push, not migrate deploy — see point 14) plus dedicated jobs for the Python and Go SDKs (the
first environment in this project's history with an actual Go toolchain —
every prior Go SDK claim was "field-audited, not compiler-verified"; this CI
job is where that finally gets checked for real). docker-publish.yml
builds and pushes three images (api, worker, web) to GHCR on main/tags.
security-scan.yml adds CodeQL and pnpm audit. Four deployment paths
(docker-compose.yml, deploy/k8s/, deploy/linux/, deploy/windows/) run
the identical application artifact — the same Docker images or the same
pnpm build output — so behavior doesn't depend on which is chosen; see
deploy/README.md for the comparison table and what's genuinely
production-hardened (the API/worker tier's HA story) versus a documented
starting point (Postgres itself has no built-in HA in any of these
manifests — see deploy/k8s/03-postgres.yaml's closing comment).
14. Migration history gap, disclosed and worked around, not silently ignored
apps/api/prisma/migrations/ only has real migration files through Module
5 (20260501000000_ai_security_copilot) — every schema change from Module 6
onward, including all of Module 10's, was authored directly in
schema.prisma without a corresponding prisma migrate dev run, because
the sandboxed development environment for this session had no live database
connection to generate migrations against. This is a real, disclosed gap:
prisma migrate deploy (the correct production migration command) cannot
be used against this repository's actual migration history past Module 5.
Every deployment path in this module instead documents prisma db push
(applies the current schema state directly, no migration history required)
for first-time setup, and docs/deployment/deploy/README.md both
explicitly recommend generating a real baseline migration
(prisma migrate dev --name module_10_baseline, run once locally against a
disposable database) before any team adopts migrate deploy for ongoing
schema changes going forward.
Consequences
- Genuinely new capability, zero breakage. Every one of the fourteen
subsystems is additive — no Module 1-9 route, DTO shape, or database
column was renamed, dropped, or had its meaning changed. Backward
compatibility holds by construction (additive Prisma schema changes,
new endpoints/modules registered alongside existing ones in
app.module.ts) not by a separate compatibility test suite. - Multiple disclosed, narrower-than-ideal v1 scopes, each documented at
its point of implementation rather than silently shipped as if complete:
Plugin sandboxing is contractual, not runtime-enforced; Compliance
auto-deletion covers two resource types, not all of them; tracing
withSpan()coverage is two call sites, not exhaustive; the three SDKs are hand-verified against shared DTOs but only the Go SDK has ever been compiled (in CI, for the first time, via this module's ownci.yml). Each of these is a legitimate scope boundary for a single module pass, not a defect, but a team building on this platform should treat this ADR's "disclosed scope" language as a literal list of next steps, not boilerplate. - One genuine architectural risk carried forward on purpose: Postgres
itself is a single point of failure in every deployment path documented
here unless the operator supplies a managed HA offering or a Postgres
operator. Making every application-tier poller safe to run at N
replicas (point 11) doesn't help if the one database backing all of them
is down — this is called out explicitly in three places
(
03-postgres.yaml,deploy/README.md, this ADR) specifically so it isn't mistaken for a solved problem. - The migration-history gap (point 14) is the one item here that a real
team must actively resolve, not just be aware of, before this becomes a
long-running production system with a schema that changes over time
through normal
migrate deployoperations rather than one-offdb pushcalls.