All documentation

Release Notes

Module 18 Release Notes — Production Hardening, Observability & Reliability

Scope

Module 18 does not add product features. Every one of its 31 phases hardens, observes, tests, or documents the platform built across Modules 1-17 — reusing existing services, repositories, ToolRunner, ScopeEngine, PollerLeaseService, audit logging, and RBAC throughout, per this module's standing "do not rewrite, do not duplicate infrastructure" constraint. Nothing below replaces a Module 1-17 capability; everything either closes a gap in it or documents it honestly, including gaps this module found but deliberately did not fix.

What was hardened

Configuration & startup — a Zod-validated env.validation.ts schema (fails boot immediately, with a clear multi-line error, on any missing or invalid required variable) and a field-for-field .env.example organized into the categories this module defined: Database, Core/ Deployment, Auth, OAuth, Email, AI, Storage, Redis/Queue, Security, Rate Limiting, Observability, Threat Intelligence, Feature Flags, Deployment.

Error handling & correlation — a global exception filter with a closed AppErrorCode union (no leaking raw stack traces to clients), and end-to-end request correlation: every request gets an X-Request-Id (client-supplied or generated), threaded through structured logs, audit events, and the DistributedJob.requestId column added this module, so one ID can trace a request from HTTP entry through to a background job it enqueued.

Structured logging & secret redaction — JSON logs via StructuredLoggerService, with an allowlist-based redaction pass so tokens/passwords/API keys never reach a log line even transitively.

Health, readiness, liveness, version — /observability/health (liveness), /observability/health/ready (readiness, real DB connectivity check), and a version endpoint, all reused by deploy/installer/health-check.sh and this module's own docs/operations/production-deployment.md checklist.

Database & migrations — a schema reliability review (see Phase 6's IDOR sweep below), and docs/operations/database-migrations.md documenting the prisma db push vs. prisma migrate deploy decision this repository's migration-history state actually requires operators to make (disclosed, not glossed over).

Job queue hardening — atomic conditional-updateMany claiming (ClaimNextJobsHandler, now covered by a new regression test — see Phase 25 below) replacing read-then-write races throughout the Distributed Jobs, Recon, and Vuln execution engines; backed-off (not immediate) retry for crash-orphaned jobs (WorkerReaperService, ReconJobPollerService.sweepOrphans() — both now covered by Phase 27's new disaster-recovery tests).

Tool execution, SSRF, rate limiting, file/storage security — consolidated single-authoritative-path tool execution, centralized SSRF guarding (ssrf-guard.util.ts), a rate-limiting/abuse-protection audit, and a file/storage security pass (upload validation, magic-byte checks).

AI safety — prompt-injection regression tests around every AI surface that ingests untrusted content (threat intel, scan output, free-text analyst input), building on the fenceUntrusted() pattern Module 17 established.

API performance — pagination caps, N+1 query elimination, and index review across the higher-traffic query paths.

SSE hardening — workspace-scoped auth and connection cleanup on every SSE stream, plus bounded exponential-backoff auto-reconnect added to the two job-log streaming hooks (useReconJobLogs, useVulnScanJobLogs) — deliberately not applied to the AI chat/agent orchestration streams; see Known Gaps.

Public sharing & audit log integrity — dedicated security passes over ShareLink token handling and the AuditEvent trail.

Metrics & observability abstraction — a Prometheus-format /observability/metrics endpoint via MetricsRegistryService, with alert rules (deploy/prometheus/alerts.yml) and Alertmanager routing/inhibition configured — see docs/operations/alerting-and-observability.md for the full inventory and its own disclosed gaps.

Backup & recovery — docs/operations/backup-recovery.md documents the real backup scopes, scheduling, CLI wizard, and the disclosed PITR scope boundary (WAL archiving is left to the infrastructure layer by design).

Frontend security headers — a real CSP (with connect-src correctly computed from the deployment's actual API origin, not just 'self'), HSTS, and the standard security-header set on every apps/web response.

Dependency audit — a manual, hands-on review of every package.json/Cargo.toml in the monorepo (this sandbox could not run pnpm audit/cargo audit — disclosed, not silently skipped), which surfaced this module's most consequential finding: see Known Gaps.

Client-surface reliability — timeout + bounded retry added to the TypeScript SDK's HttpClient (previously unbounded — a hung connection could await forever) and to the CLI's own npm-registry update check; Desktop Agent and Mobile app reliability layers were verified already solid on inspection (bounded exponential backoff in both, plus an auth-refresh interceptor in Mobile more complete than the SDK's own — see Known Gaps).

Security-critical test hardening — a 16-category coverage audit (docs/reviews/0017-...md) found 13 categories already well-tested and added real regression tests for the two genuine gaps it found: the ApiKeyAuthGuard's auth/IP-allowlist/rate-limit logic, and the Distributed Jobs claim-race protection.

Disaster/failure-mode tests — new regression tests for the two crash-recovery poller loops (WorkerReaperService, ReconJobPollerService.sweepOrphans()) that had never been directly tested despite each having its own documented history of a real race condition fixed in an earlier phase.

Final security grep sweep — a pattern-based pass (command injection, SQL injection, XSS, hardcoded secrets, CORS, JWT verification, insecure randomness, @Public() route audit, path traversal, TODO/FIXME markers) across the whole codebase at once. No new finding — see docs/reviews/0019-...md.

Known gaps — consolidated

Every gap this module found and disclosed rather than fixed, gathered in one place. None of these are silent; each was recorded in its own phase's review doc at the time, cross-referenced here.

Rate limiting is per-instance, not shared across replicas (Phase 10). Both the global ThrottlerGuard and ApiKeyAuthGuard's hand- rolled per-key limiter keep counters in process memory — correct for a single-instance Community Edition deployment, but the effective limit across a multi-replica deployment is configured limit × replica count. Closing this needs a shared store (Redis, matching the existing RedisCacheService pattern) wired into both. See docs/operations/production-deployment.md §7.

No admin/owner-tier endpoint to review a workspace's audit history (Phases 17, 19, 25). AuditService.listRecentForUser is self-service- only (used solely by the GDPR data-export flow). Adding a real admin-facing audit trail is unbuilt product work — a new workspace- scoped query handler + controller route gated by an admin/owner role guard — not a configuration toggle.

AuditEvent has no tamper-evidence mechanism (Phase 19). Rows are plain, mutable database records; nothing hash-chains or signs them. Acceptable for this platform's current threat model (the audit log's purpose today is operational visibility, not court-admissible evidence), but worth stating plainly rather than implying more integrity than exists.

Observability signals are disconnected (Phase 19). Metrics (Prometheus), logs (structured JSON + request-ID correlation), and traces (none — no distributed tracing exists in this codebase) are three separate systems with no unified exemplar/trace-linking. A request-ID lets an operator manually correlate a log line to... nothing else yet, since there's no tracing backend to link to.

No wired Alertmanager receiver by default (Phase 19, restated in Phase 26 as the single highest-priority go-live checklist item). deploy/prometheus/alertmanager.yml ships with routing/inhibition rules but no Slack/PagerDuty/email receiver configured — alerts fire into nothing until an operator adds one.

RealtimeHubService keeps an unbounded in-memory connection map, and Recon/Vuln SSE streams have no heartbeat during quiet polling (Phase 19, prior-window finding, restated here for completeness). The frontend- side fix (bounded reconnect on the two job-log hooks) does not close the backend-side absence of a heartbeat, which is what makes idle-timeout disconnects happen in the first place on some reverse-proxy configurations.

AI chat and agent orchestration SSE streams do not auto-reconnect (Phase 23, deliberate). The two job-log hooks got a bounded exponential- backoff reconnect fix because their state is simple and append-only. The AI chat/agent streams carry partially-streamed, in-progress content — a naive reconnect risks duplicating or corrupting a mid-generation response in a way this sandbox could not verify safely. Deferred with reasoning recorded, not silently left broken.

The TypeScript SDK has no automatic refresh-on-401 (Phase 24). PentestHubClient.auth.refresh() (and the onTokenRefreshed option that persists its result) only run when a consumer calls them explicitly — there is no interceptor that catches a 401 and retries after refreshing, unlike apps/web's own client.ts or the Mobile app's ApiClient, both of which do exactly that.

The CLI never uses the refresh token it stores (Phase 24). pentesthub login saves a refreshToken to ~/.pentesthub/config.json, but no CLI command ever reads it or calls the refresh endpoint — the CLI also bypasses PentestHubClient/ AuthResource entirely (it builds a raw HttpClient for M13-surface coverage reasons predating this module), so even the SDK-level onTokenRefreshed callback it wires up can never fire. Net effect: a CLI session's access token eventually expires and requires a manual pentesthub login — a real UX gap, not silently swallowed.

No true point-in-time recovery (Phase 20, restated in Phase 26). Backups are discrete pg_dump snapshots; PITR requires Postgres-level WAL archiving, deliberately left to the infrastructure layer rather than bundled, since the right tool varies by where Postgres is hosted.

Postgres itself is not made highly available by anything in this repository (Phase 26, inherited from Module 14's own disclosure). Use a managed offering or a Postgres operator.

Crash-recovered jobs are not idempotent against external side effects (Phase 27, from WorkerReaperService's own doc comment). Migrating a job off a dead worker re-runs its handler from scratch on a new worker — safe for the database (handlers upsert on deterministic keys), but if the dead worker had already issued a real external action before it crashed (kicked off a scan against a live target, called a paid AI API, sent a webhook), that action gets re-issued a second time. Closing this fully needs a stable idempotency key threaded through every external call site (ToolRunner, AI provider adapters, webhook dispatch) and honored by each provider — a cross-cutting change out of proportion to any single phase's scope.

CORRECTED post-Module-18: Phase 22's "Python/Go SDKs don't exist" finding was a false negative. A dedicated reconciliation checkpoint (docs/reviews/0021-post-m18-sdk-reconciliation.md) found both SDKs alive, real, and at full resource parity with the TypeScript SDK — at sdks/python and sdks/go, not packages/sdk-python/packages/sdk-go. Phase 22 checked only the packages/sdk-* path convention and, finding nothing there, wrongly concluded the SDKs were absent rather than searching for them elsewhere. Module 16's #402/#429 and Module 17's #491 reported this work accurately; Phase 22's audit was the error, not the task history. Two small, real, previously-undisclosed issues were found and fixed during reconciliation: the Go SDK's HttpClient used http.DefaultClient, which has no timeout at all (same defect class as this module's own TypeScript SDK fix below — corrected to a 30s default); and .github/dependabot.yml genuinely was missing pip/gomod entries for sdks/python/sdks/go (added). Both SDK READMEs' "Scope (v1)" sections were also stale (still describing the pre-Module-16 5-resource list despite the code already having all 14) and have been corrected. See the reconciliation doc for the full audit, including a real, passing Python test run (python3 -m unittest tests.test_client -v — 5/5 tests passed in-sandbox) and the Go toolchain's continued unavailability in this sandbox (consistent with every prior module).

Verification

mcp__workspace__bash was denied for the entirety of this module's session (confirmed repeatedly from Phase 22 onward). Every new test file this module added (Phases 25 and 27, five spec files total) was written to match this codebase's established Jest/hand-rolled-mock conventions and traced field-for-field against the real source it exercises, but could not be executed in this sandbox. This is disclosed explicitly per this module's standing discipline rather than claimed as a passing run — see Phase 30 for the final verification-loop attempt.