ADR 0014 — Cloud Platform, Enterprise SaaS & High Availability
Status
Accepted (source-complete; verification loop constrained by sandbox toolchain limits, of a different shape than prior modules — see §11).
Context
Modules 1-13 built a full pentest/bug-bounty product with three
first-class clients and a Module 10 "Enterprise Platform" pass that
already introduced multi-tenancy, a Distributed Job System, and four
deployment paths. What Module 10 left as a starting point rather than a
finished story — disclosed explicitly in its own release notes — is
everything that turns "the application runs" into "the application
survives a node failure, scales horizontally, backs itself up, updates
itself safely, and ships through a real pipeline instead of a manual
docker build." Module 14 closes that gap: pluggable infrastructure
abstractions (queue/cache/storage), real distributed-worker resilience,
backup/disaster-recovery, an infrastructure Admin Console, Enterprise-only
tenant management, an API Gateway, cross-client auto-update, and
production-grade CI/CD — all while holding a non-negotiable constraint
carried over from every prior module's Community Edition posture: no
mandatory paid service, self-hosted deployment fully functional, external
cloud services always optional.
Decisions
1. Provider abstractions, not a forced migration to managed
infrastructure. Three seams — QUEUE_PROVIDER (memory | redis),
CACHE_PROVIDER (memory | redis | disk), STORAGE_PROVIDER
(local | s3) — each default to the zero-external-dependency binding.
apps/api/src/common/providers/ defines one interface per seam
(QueueService, CacheService, StorageService) with every consumer
(distributed jobs, HTTP response caching, evidence/attachment uploads)
coded against the interface, never a concrete implementation — a Redis
or S3 deployment is a config change, not a code change, and a self-hosted
single-node deployment never needs either.
2. High Availability is a Postgres-row-based lease, extended, not
reinvented. Module 10 already introduced PollerLeaseService for
seven single-replica-bound background pollers. Module 14 (task #303)
adds leader election on top of the same primitive for cluster-wide
singleton responsibilities (schedule enforcement, cross-node
coordination) rather than introducing a second consensus mechanism
(no Raft/etcd/Zookeeper dependency) — consistent with the "no message
broker for the job system either" precedent Module 10 set for the
Distributed Job System.
3. Distributed Workers get heartbeat/crash-recovery/priority/
concurrency-limit extensions on the existing WorkerNode/
DistributedJob tables (task #304), not a new job-runner subsystem.
A dead worker's claimed jobs are detected via heartbeat staleness and
requeued; per-worker concurrency limits and job priority are additive
columns/query changes to Module 10's claim-based dequeue, not a
replacement of it.
4. Database support is pooling, replica routing, and migration
validation — not a second ORM layer. DATABASE_POOL_MAX/
WORKER_DATABASE_POOL_MAX size Prisma's connection pool per-replica;
an optional read-replica DATABASE_URL routes read-only queries when
configured and falls back to the primary when it isn't (task #307).
apps/api/scripts/validate-migrations.mjs is a standalone script (no
new dependency) checking prisma validate (schema-only, no DB) and,
when a DATABASE_URL is available, prisma migrate diff --exit-code
against it — drift is reported, not fatal, since this repository's own
migration history has known gaps disclosed since Module 10 (see
deploy/README.md's "Schema setup on first deploy" section).
5. Backup/DR is file-based and provider-agnostic, matching the storage
abstraction it reuses. Database dumps (pg_dump), attachment/evidence
directories, and config snapshots are written through the same
StorageService seam as everything else — a backup lands on local disk
by default and in S3-compatible storage only if STORAGE_PROVIDER=s3
is already configured for the deployment. Point-in-time recovery is
documented as a Postgres-operator/managed-service responsibility (task
#312), not re-implemented, consistent with deploy/README.md's existing
disclosure that "Postgres itself is not made highly-available by
anything in this repository."
6. The Admin Console is a read/action surface over infrastructure the
platform already exposes, not a new source of truth. GET /admin-console/* (task #314) aggregates WorkerNode/DistributedJob
state, CacheService.stats(), StorageService usage, the observability
health/metrics endpoints, and backup/plugin/update state that Modules 14
subsystems already persist — an operations dashboard, not a second
schema.
7. Tenant management is Enterprise-only and fails closed to
Community. TenantController/TenantService (task #315) sit behind
EditionGuard, which — per EditionService.isEnterpriseFor()'s existing
Module 10 design — treats a deployment with zero License rows as a
trusted, permitted environment for everything except tenant-specific
billing/branding/custom-domain features, so a self-hosted Community
deployment is never blocked from anything it was already doing; it
simply doesn't see tenant admin as an option.
8. The API Gateway is opt-in caching plus existing middleware wired
together, not a new reverse-proxy layer. @CacheResponse(ttlSeconds) +
CacheResponseInterceptor (task #316) cache GET-only responses through
the existing CacheService abstraction, keyed per-user
(http-cache:<userId|anon>:<path>) to prevent cross-tenant replay, with
no invalidation on write (a documented TTL-bounded staleness trade-off,
applied only to the one low-churn route it was added to this pass —
knowledge-base/references). Rate limiting, compression, and API-key
auth were already real (Module 9/10); this pass's job was composing them
correctly with the new interceptor's guard/interceptor nesting order,
not building them from scratch.
9. Auto-update is check-and-notify everywhere except the Desktop
Agent, never silent self-replacement. Per client: the API/worker have
no runtime update concept (deploy-time versioning only, GET /observability/health exposes the running version); the web app polls
/observability/health and shows a dismissible banner; the CLI's
pentesthub update command prints upgrade instructions rather than
overwriting its own binary; mobile does a store-policy-compliant
check-only comparison; the Desktop Agent's existing Tauri updater (from
Module 7) is the one component with real self-update capability, with a
disclosed no-rollback gap. See docs/architecture/auto-update.md.
10. CI/CD splits PR validation from release/deployment into separate
workflows, and every optional distribution channel is secret-gated.
ci.yml (build/lint/typecheck/test/migration-validation, Turborepo-cached)
runs on every PR; release.yml re-verifies at the exact tagged ref
against a fresh ephemeral database (never trusts main's last green run),
generates a CycloneDX SBOM, and creates the GitHub Release; npm/PyPI
publishing are if: vars.PUBLISH_*_TO_* gated so a fork with neither
configured still gets a complete release artifact set.
docker-publish.yml adds BuildKit-native SBOM/provenance attestations,
Trivy scanning (informational, matching this codebase's existing pnpm audit --audit-level high/continue-on-error: true posture), and cosign
keyless signing via GitHub's own OIDC — no account, key, or paid service
required for any of it. See docs/architecture/ci-cd.md.
11. Performance and Community Edition guarantee passes are dedicated,
not folded silently into feature work. Task #319 found and fixed one
real N+1 (GetAvailablePluginUpdatesHandler, one findFirst per
installation → two total queries plus a supporting composite index).
Task #320 audited every Module 14 subsystem against the standing
guarantee and found one real, unrelated bug in the process:
deploy/k8s/01-configmap.yaml shipped QUEUE_PROVIDER: "in-memory",
which does not match env.validation.ts's zod enum ("memory" | "redis") — applying that ConfigMap as shipped would have failed
validateEnv() at boot. Both are documented in
docs/reviews/0006-module-14-performance-review.md and
docs/reviews/0007-module-14-community-edition-guarantee.md.
Consequences
- Three new pluggable-provider seams (queue/cache/storage) mean every
future feature that needs background work, caching, or file storage
must be written against the interface, not
ioredis/aws-sdkdirectly — a constraint, not just a convenience, since bypassing it would silently reintroduce a hard Redis/S3 dependency the Community Edition guarantee forbids. - The Admin Console and Tenant Management surfaces both read from tables/services other Module 14 subsystems own; neither should grow its own persisted state without a specific reason, to avoid two sources of truth for the same infrastructure facts.
- CI/CD is now genuinely load-bearing:
release.yml's from-scratch re-verification at the tagged ref is the first point in this project's history where a release build is confirmed independent of whatever statemain's last CI run happened to be in. - The Community Edition guarantee, reaffirmed by task #320's dedicated
audit, is now cross-checked against every subsystem this module
touched, not just asserted — the one bug that audit found
(
QUEUE_PROVIDERtypo) is exactly the class of drift that guarantee passes exist to catch before it reaches a real deployment.
Verification constraints (sandbox)
Unlike Modules 7-13, where the standing constraint was "no pnpm
binary, no registry network access at all," this pass's sandbox had a
different and more specific set of blockers, diagnosed in detail
during task #323:
pnpmis installed but not onPATHby default per shell invocation — workable oncePATHis corrected per call.- Network to
registry.npmjs.orgis not fully blocked but is severely degraded (intermittentUND_ERR_CONNECT_TIMEOUT, sub-50 KiB/s throughput) — a fullpnpm installcannot complete within this sandbox's per-command time budget. - The
turbonpm package itself was never fully installed in this sandbox —node_modules/.bin/turbois a dangling symlink with nonode_modules/turbobehind it, sopnpm turbo run ...cannot execute here at all, independent of the network issue. - Direct
tsc --noEmitandprisma validateinvocations hang with zero output until the sandbox's own per-call timeout — root-caused to filesystem I/O latency on the mounted host folder (confirmed via targetedfind/readdirtiming tests against the ~1,266-package pnpm store), not a defect in the code being checked. - No Docker daemon/CLI and no
kubectl/kubeval/kubeconformare available in this sandbox, so Docker image builds and live Kubernetes manifest validation could not be executed here either.
What could be verified here: structural YAML validation
(yaml.safe_load_all) of every new/modified GitHub Actions workflow and
every deploy/k8s/*.yaml/docker-compose*.yml file — all passed. Full
build/lint/typecheck/test/Docker/Kubernetes verification requires a
normal development machine; see the final report accompanying this
module's completion for the exact commands.