All documentation

Reviews & Audits

Module 18 — Disaster / Failure-Mode Tests

Task #529 (Phase 27). Where Phase 25 hardened security-critical correctness paths, this phase targets the two crash-recovery poller loops that are this codebase's actual answer to "what happens when a worker process dies mid-job" — the scenario every one of Module 18's prior phases assumed would be handled correctly (Phase 8's job-queue hardening, Phase 15's SSE cleanup, this module's repeated "idempotent- by-contract" framing) but that had no test proving it.

What was tested

WorkerReaperService (modules/distributed-jobs/services/worker-reaper.service.spec.ts, new) — the Distributed Jobs system's dead-worker detection and job migration. Five cases: stale-heartbeat detection marks a worker UNHEALTHY while a fresh one is left alone; nothing is touched when no worker is stale; an UNHEALTHY worker's ASSIGNED/RUNNING jobs are migrated back to QUEUED via the atomic UPDATE ... RETURNING (not a read-then-write pair — this is the specific fix task #510 made for a disclosed TOCTOU race), with a MIGRATED JobEvent per job and currentLoad reset; the same atomic query returning zero rows (the race it guards against actually happening) results in no event and no currentLoad reset; and the poll loop survives a thrown error on one tick and keeps ticking rather than dying silently.

ReconJobPollerService's sweepOrphans() (modules/recon/execution/recon-job-poller.service.spec.ts, new) — the Recon Engine's equivalent crash-recovery path, predating and structurally mirroring the Distributed Jobs one. Four cases: an orphaned job under its retry budget gets a scheduled retry with real, strictly- future backoff (not immediate re-eligibility — the exact regression task #510 also fixed here, for a target that reliably crashes its scanner); a job that has exhausted its retry budget is marked FAILED instead; a job this same process still has genuinely in flight (tracked in the poller's own running map) is correctly skipped even if findOrphaned's time-based cutoff would otherwise also match it; and the same poll-loop error-resilience property as above.

Why these two and not others

Both are the two places in this codebase where "a process died while holding work" has a dedicated recovery mechanism with previously- undisclosed, previously-untested logic (as opposed to, say, an HTTP request simply timing out and returning an error, which is already covered by the ordinary request-handling test suite). VulnScanJobPollerService almost certainly has the same shape (Module 4's execution engine mirrors Recon's), but was not independently re-verified this phase — noted here rather than silently assumed identical; a future pass following this same file's pattern would be the natural next step if that surface's crash-recovery behavior needs its own regression coverage.

Environment constraint

Same as every other test-writing phase this module: mcp__workspace__bash is denied this session, so neither new spec file could actually be executed here. Both were traced field-for-field against the real WorkerReaperService/ReconJobPollerService/ReconJobsRepository source (including fixing one real mock-sequencing bug found during self-review — an early draft of the "still running, not really orphaned" test had findOrphaned return the job on every call, which would have made it register as an orphan before the job was even claimed) before being left for Phase 30's verification attempt or a real CI run.