All documentation

Architecture Decision Records

0004: Vulnerability Engine module (Module 4)

Status: Accepted Date: 2026-07-13

Context

Module 4 adds automated vulnerability scanning: run external security scanners (template-based vulnerability matching, content discovery/fuzzing, web server misconfiguration scanning) against a Target, normalize their output into a common Finding schema, deduplicate it, auto-attach evidence, and surface it as validated Findings with a full triage lifecycle (New → Confirmed/Duplicate/Accepted Risk/Fixed/False Positive) — while leaving explicit event seams for a future AI Assistant to hook into without this module knowing it exists.

The scope was fixed by the project owner up front: Module 4 only, continuing the existing architecture and not breaking Module 3. No AI Assistant, Report Generator, Browser Extension, Desktop App, Mobile App, Billing, or Marketplace work. This document also formally supersedes this repository's original PROJECT_SPEC.md placeholder for "Module 4 — Vulnerability Tools" (a client-side payload/encoder toolbox), which the project owner redefined into the scanner-orchestration engine described below — see the note added to that section.

Decision

1. Reuse Module 3's job/worker architecture verbatim, extend it where the spec asked for more

VulnScanJobsController / TargetVulnScanJobsController / ProjectVulnScanJobsController only ever create a VulnScanJob row (status PENDING) and its per-scanner VulnScanJobToolRun rows, then return immediately — identical division of responsibility to Recon.

Execution is owned by VulnScanJobPollerService, which runs inside the same worker OS process as ReconJobPollerService (see worker.main.ts / worker.module.ts) rather than a second binary: Module 4 doesn't need its own deployment unit, it needs its own poll loop and claim query against its own table. Both pollers are independently started/stopped and claim from different tables (vuln_scan_jobs vs recon_jobs), so nothing about their concurrency, timeout, or retry budgets is shared or coupled.

Job states extend Recon's PENDING → RUNNING → (COMPLETED | FAILED | CANCELLED) (with RETRYING as an intermediate state) by adding PAUSED — the spec's explicit "Pause/Resume" requirement, which Recon never needed.

2. Distributed-worker-safe claiming, now with Priority

claimNext(claimantId) is the same atomic conditional-UPDATE ... FOR UPDATE SKIP LOCKED pattern as Recon, extended with ORDER BY priority DESC, "createdAt" ASC — the spec's "Priority" requirement. priority is a plain Int (0–100, default 0, validated at the DTO layer) with no separate queue infrastructure; higher-priority PENDING/RETRYING jobs are simply claimed first by any available worker slot.

Worker Heartbeat / Dead Worker Recovery replaces Recon's claim-age-only staleness check with something more accurate for scans that can legitimately run far longer than a claim-age cutoff would tolerate: VulnScanJob. lastHeartbeatAt is updated on a slower cadence than the poll loop while a scanner is running (tracked via an in-memory lastHeartbeatSentAt per running job, to avoid a DB write on every tick), and the orphan sweep (findOrphaned) keys off heartbeat staleness (falling back to claim-age only if no heartbeat was ever recorded) rather than claim-age alone. A dead worker's job is routed through the same retry/fail budget as any other execution failure — no bespoke recovery path, matching Recon precedent.

3. Pause is a second, gentler mechanism alongside Cancel — not a variant of it

Cancel and Timeout share Recon's mechanism exactly: one AbortController per claimed job, .abort() hard-kills an in-flight scanner subprocess immediately via the signal threaded through to ScannerProcessSpawner.

Pause is deliberately different: PauseVulnScanJobCommand only records intent (pauseRequestedAt), never aborts anything. VulnScanJobExecutionService polls pauseRequestedAt from the database between scanner runs (never mid-run) and, if set, stops looping and transitions the job to PAUSED itself — the currently running scanner is always allowed to finish. Resume (ResumeVulnScanJobCommand) transitions PAUSED → PENDING; the poller picks it back up on its next tick like any other pending job, and since VulnScanJobExecutionService only re-runs PENDING tool runs (already- COMPLETED ones are skipped), a resumed job continues from where it left off rather than restarting from scratch.

One asymmetry worth calling out: Stop/Cancel on a PAUSED job transitions it directly (no worker is watching a paused job to observe cancelRequestedAt), publishing VulnScanJobCancelledEvent from the command handler itself — mirroring how Resume also publishes immediately for the same reason. Every other transition (RUNNING/RETRYING → anything) is worker-observed and worker-published.

4. Scanner integration: one interface, one registry, additive extension — Recon's pattern, renamed

text
ScannerRunner.run(ctx): Promise<ScannerRunResult>         // spawns the CLI scanner, returns raw stdout
Normalizer.normalize(rawOutput): VulnFindingInput[]        // parses raw output into typed findings

ScannerRunnerRegistry resolves both by VulnScanner enum value, and is additionally responsible for the target-type compatibility check Recon's registry didn't need in the same place (isScannerCompatibleWithTargetType). Adding a new scanner is the same closed, three-step recipe as Module 3's tools:

  1. A ScannerRunner class (scanner-runners/adapters/<scanner>.scanner-runner.ts) — declares readonly scanner and readonly supportedTargetTypes, spawns the CLI via ScannerProcessSpawner with an argv array.
  2. A Normalizer class (scanner-runners/normalizers/<scanner>.normalizer.ts) — pure function from raw stdout to VulnFindingInput[], additionally responsible for filling the unified Finding schema's severity/CVSS/CWE/ OWASP-category/evidence fields (see §5).
  3. One entry each in ScannerRunnerRegistry's runners/normalizers maps, VulnCoreModule's providers, the VulnScanner Prisma enum + migration, and SCANNER_TARGET_TYPE_COMPATIBILITY in packages/shared/src/vuln.ts.

v1 ships five scanners this way — Nuclei, ffuf, dirsearch, feroxbuster, Nikto — with the registry, the VulnScanner enum, and SCANNER_TARGET_TYPE_COMPATIBILITY already shaped to accept SQLMap, XSStrike, Dalfox, Wapiti, and OWASP ZAP later without touching any of the five existing adapters.

5. Unified Finding model with the exact field list the spec required

Every scanner's output normalizes into one VulnFinding row carrying every field the spec named: Title, Description, Severity (INFO/LOW/MEDIUM/ HIGH/CRITICAL), CVSS (score + vector), CWE, OWASP Category (mapped heuristically from tags/CWE via owasp-mapping.util.ts), Scanner, Scanner Version, Confidence (LOW/MEDIUM/HIGH/CONFIRMED — Nuclei template matches are always CONFIRMED, content-discovery hits from ffuf/dirsearch/ feroxbuster are only ever suggestive), Status (see §6), Target/Project/ Workspace (FKs), Evidence (see §7), References, Tags, Created By, Created At, Updated At.

Unlike Recon's polymorphic-by-type model, every Vuln finding shares the same shape — there's no analog to Recon's twelve different ReconFindingType payload shapes — so this is a normal typed table, not a data: Json-carrying polymorphic one (though data: Json still exists for scanner-specific raw fields worth preserving, e.g. Nuclei's template-id).

6. Finding Status lifecycle and Deduplication

Status: NEW → CONFIRMED | DUPLICATE | ACCEPTED_RISK | FIXED | FALSE_POSITIVE — a flat set of terminal-ish triage outcomes (not a strict state machine; any status is reachable from any other via UpdateVulnFindingStatusCommand or BulkUpdateVulnFindingsCommand, since triage is a human judgment call the system shouldn't gate).

Deduplication reuses Recon's exact philosophy: a pure key function (vulnFindingDedupKey, combining normalized path + vulnerability type + signature) plus a DB unique constraint — @@unique([targetId, scanner, dedupKey]) — with upsertMany() upserting on that constraint. The spec additionally asked for dedup on "Target, Path, Scanner, Vulnerability Type, Signature" specifically; targetId and scanner are real columns in the constraint, while path/vulnType/signature are folded into the single dedupKey string (mirroring how Recon folds multiple identity fields into one key) — Postgres enforces the invariant either way, the key-computation function is the only thing under unit test.

7. Evidence and Activity Timeline are automatic, reusing Module 2's Evidence model — no new storage system

Every v1 scanner's evidence is plain text (HTTP request/response bodies, curl commands, raw scanner output) — none produce images — so VulnFinding gets six nullable named-relation FKs into the existing Evidence model (requestEvidenceId, responseEvidenceId, screenshotEvidenceId, payloadEvidenceId, commandEvidenceId, rawOutputEvidenceId) rather than a parallel evidence system. EvidenceType gained two purely-additive enum values (PAYLOAD, COMMAND) for the two kinds Recon never needed. Because everything is text, this module needs no Attachment/StorageService/blob-storage integration at all for v1 — a genuine simplification versus Recon's GowitnessNormalizer screenshot-file path. screenshotEvidenceId is reserved (nullable, unused by any v1 normalizer) for a future visual-capture scanner like OWASP ZAP.

Domain events (VulnScanJobCreatedEvent, ...StartedEvent, ...CompletedEvent, ...FailedEvent, ...CancelledEvent, ...PausedEvent, ...ResumedEvent, VulnFindingCreatedEvent, VulnFindingUpdatedEvent, VulnFindingConfirmedEvent, VulnFindingResolvedEvent) are published via the same EventBus and picked up by the same generic ActivityRecordingHandler Recon and every other module already uses — Module 4 only needed its events to implement ActivityDomainEvent. VulnFindingUpdatedEvent is always published on any status change; VulnFindingConfirmedEvent/VulnFindingResolvedEvent are published additionally (not instead) when the new status is CONFIRMED or one of {FIXED, FALSE_POSITIVE, ACCEPTED_RISK} respectively — so a future AI module (or any other subscriber) can listen narrowly ("resolved findings only") or broadly ("every status change") without this module changing.

Because upsertMany()'s upsert() doesn't report insert-vs-update, PrismaVulnFindingsRepository does a findUnique existence check before each upsert() to compute a wasCreated flag, which VulnScanJobExecutionService uses to choose FindingCreated vs FindingUpdated — an accepted small race-condition tradeoff that only affects which event fires, never data integrity (the unique constraint still guarantees that).

8. REST API and Swagger

Every endpoint the spec named exists: Start/Stop/Pause/Resume/Retry Scan, Findings (list with severity/status/scanner/search filters + bulk status update), Finding Details, Evidence (via the finding's evidence FKs, reusing GET /evidence/:id), Scanner History, Scan Templates (CRUD, built-in templates protected from deletion), plus a Dashboard summary endpoint (severity counts, status counts, running-job count, recent findings, scanner history — one aggregate query for the frontend dashboard instead of N separate calls). All controllers use @ApiTags('vuln')/@ApiBearerAuth()/ @ApiOperation() — Swagger docs are auto-populated, no manual tag registry to maintain.

9. Frontend: mirrors Recon's integration points, plus the additional pages the spec asked for

  • A Vulnerabilities card on the target detail page (parallel to Recon's card) — job list scoped to that target, "New scan," scanner-compatibility- aware scanner selector.
  • A Vulnerabilities tab on the project detail page, alongside Recon's existing Scan History tab.
  • /vuln — workspace-wide Vulnerability Dashboard: severity chart, scanner status, running-job count, recent findings.
  • /vuln/jobs, /vuln/jobs/[jobId] — scan job list and detail (progress, live SSE log panel, per-job findings), with Pause/Resume/Stop/Retry actions gated by the job's current status.
  • /vuln/findings, /vuln/findings/[findingId] — workspace-wide, server- filtered findings table (severity/status/scanner/search) with row- selection Bulk Actions and a / keyboard shortcut to focus search; finding detail shows classification (CVSS/CWE/OWASP), tags/references, attached evidence (one <pre> block per non-null evidence FK), and a lightweight timeline (first seen/last seen/created/updated) built from the finding's own timestamps rather than a dedicated Activity-by-entity query (which the Activity module doesn't expose yet).
  • /vuln/templates, /vuln/scanners — Scan Templates CRUD and Scanner History, both new resource types Recon has no analog for.

Live Updates reuse Recon's exact two mechanisms: React Query polling (refetchInterval, idle once a job/dashboard query's status is terminal) for job/list/dashboard state, and a manual fetch() + ReadableStream SSE read loop for live logs (the browser's native EventSource can't attach the Authorization: Bearer header this endpoint requires). Dark Mode and responsive layout are inherited from the existing global shell — Module 4 added no new theme or breakpoint logic.

10. Future extension points — hooks only, nothing implemented

  • AI Assistant: every Vuln event above is published whether or not anything is listening. No AI logic exists in this module.
  • Report Generator / Bug Bounty workflow: the Finding/Evidence/ Activity trail this module produces is the seam a future module would build on — no code in this module references them.
  • SQLMap, XSStrike, Dalfox, Wapiti, OWASP ZAP: VulnScanner enum values, registry map entries, and SCANNER_TARGET_TYPE_COMPATIBILITY rows are the only three places any of these would ever touch, per §4 — none are implemented.

Consequences

  • Adding scanner #6 touches exactly three new files plus three additive registration lines — no existing scanner's code is at risk, proven by the five scanners shipped this way without touching each other.
  • Running the Recon and Vuln pollers in one shared worker process means one deployment unit still covers two independently-scaled-by-concurrency-cap job types; splitting them into separate processes later is possible but wasn't necessary for v1's job volume.
  • The heartbeat-based orphan sweep is strictly more accurate than Recon's claim-age-only sweep for long-running scans, at the cost of one extra timestamp column and a periodic (not per-tick) DB write per running job — a worthwhile tradeoff given how much longer a content-discovery fuzz scan can run than a typical Recon tool invocation.
  • No blob storage integration for v1 Evidence is a real simplification, but it's contingent on every v1 scanner's output being plain text; the first visual-capture scanner (most likely OWASP ZAP) will need to revisit this the same way Gowitness did for Recon.
  • The lightweight per-finding timeline (built from the finding's own timestamps) is a stand-in for a proper "Activity filtered by entity" view — revisit if/when the Activity module grows an entityId filter.

Scanner inventory (as of this ADR)

ScannerFinding confidenceTarget types
NucleiCONFIRMED (template match is exact)URL, DOMAIN, SUBDOMAIN, IP_ADDRESS
ffufLOW/MEDIUM (suggestive)URL
dirsearchLOW/MEDIUM (suggestive)URL
feroxbusterLOW/MEDIUM (suggestive)URL
NiktoMEDIUM/HIGH (misconfiguration checks)URL, DOMAIN, SUBDOMAIN, IP_ADDRESS

Reserved, not implemented: SQLMap, XSStrike, Dalfox, Wapiti, OWASP ZAP.