0011: AI Security Copilot + Autonomous Multi-Agent System (Module 11)
Status: Accepted — source-complete Date: 2026-07-26
Context
Module 5 gave the platform a workspace-aware AI chat assistant: one conversation, one model turn at a time, with RAG and tool calling. Module 11's spec asks for something categorically different — a Master AI Orchestrator that takes a free-text goal ("recon this target and draft a findings summary"), decomposes it into a tree of subtasks, delegates each subtask to one of twelve specialist agents (Recon, Web, API, Mobile, Cloud, Active Directory, AI Research, Exploit Analysis, Report Writer, Code Review, Risk Assessment, Workflow Coordinator), runs them with dependency ordering and bounded concurrency, retries failures, merges outputs, and produces a self-evaluated final report — an autonomous multi-agent system, not a single chat turn. Around that core, the spec also asks for: a searchable AI Knowledge Base (OWASP/CWE/CAPEC/MITRE ATT&CK/NIST/CVE), an Explainability layer ("show your work" for every task), per-workspace Privacy Mode controls (cloud/local/hybrid), and a set of narrower enhancements to Module 5's existing AI Chat/Memory/Voice/Automation surfaces. Nineteen subsystems (tasks #219-#237) plus tests and docs. This ADR covers the decisions that span the module; narrower decisions live in the file-level doc comments of the services they describe.
The standing instructions for this module were explicit: do not rewrite previous modules except where integration is required, and stop after Module 11 is fully completed — every decision below was made against that constraint.
Decision
1. OrchestrationRun -> AgentTask -> AgentToolCallLog as a new, additive domain — not a repurposing of AiConversation
Module 5's AiConversation/AiMessage model one human ↔ one assistant,
one message at a time. An orchestration run is structurally different: one
goal produces a tree of tasks (via parentTaskId, enabling Task
Delegation — a WORKFLOW_COORDINATOR task can itself propose further
sub-tasks up to AI_ORCHESTRATOR_MAX_TASK_DEPTH), each with its own
status lifecycle, retry count, reasoning summary, confidence score, and
tool-call log, independent of any single chat conversation. Rather than
bolting tree semantics onto AiConversation (which every Module 5/6/9
call site already assumes is flat), Module 11 adds four new tables
(OrchestrationRun, AgentTask, AgentToolCallLog, plus
KnowledgeBaseReferenceEntry) and a new agent module alongside the
existing ai module — additive in the same sense every prior module's
schema growth has been: no Module 5 table, column, or query shape changed.
2. Twelve specialist agents share one PromptDrivenAgent base, not twelve bespoke implementations
Every specialist agent (ReconAgent, WebAgent, ..., WorkflowCoordinatorAgent)
is a thin subclass supplying its own system prompt and allowed tool
subset; PromptDrivenAgent.execute() owns the actual model-call + tool-call
loop (budgeted at MAX_AGENT_TOOL_CALLS = 3 per task, mirroring
AiChatService's MAX_TOOL_CALLS_PER_TURN convention exactly) and the
defensive-JSON extraction (extractAgentJson/clampConfidence —
see prompt-driven-agent.base.spec.ts) every agent's final response goes
through. This is the same "one shared engine, many thin configurations"
shape Module 4's ScannerRunner and Module 3's ToolRunner established —
adding a 13th agent later means one new subclass, not touching the
execution loop.
3. The Orchestrator reuses Module 5's AiProviderRegistryService/tool-calling framework, doesn't fork them
AgentOrchestratorService.planGoal() and every PromptDrivenAgent
resolve their model through the exact same AiProviderRegistryService
Module 5/BYOK Phase 1 built — same 5-provider abstraction, same
workspace-BYOK-credential resolution, same collectStreamedCompletion
helper. Tool calling reuses Module 5's AiToolRegistryService/
findToolCall parser rather than a parallel agent-specific tool-call
format. This is the reuse the "don't rewrite previous modules" instruction
was written to encourage: Module 11 is a new consumer of Module 5's seams,
not a fork of them.
4. Explainability (task #233) is pure aggregation, not a new persisted concept
AgentExplainabilityService.explainRun() reshapes data every AgentTask/
AgentToolCallLog already persists (reasoningSummary, confidenceScore,
sources, each tool call's input/output) into one top-to-bottom
narrative (plan reasoning → per-task reasoning + evidence → overall
result) — no new table, no new write path. "Show your work" was true by
construction the moment task #223/#224 started recording that data; this
service exists so a caller doesn't have to reconstruct the narrative
themselves from the raw task tree.
5. AI Knowledge Base: exact/token-match ranking over the catalog itself, not semantic search — by design
KnowledgeBaseReferenceService.search() is plain relevance-ranked text
matching (exact code match scores 1; otherwise a normalized token-overlap
score) over the reference catalog, not an embedding-based nearest-neighbor
search. Deliberately: the catalog is a few hundred well-curated entries
with exact, well-known identifiers ("CWE-79", "T1190"), and a security
analyst or an agent citing one almost always knows the code or a close
paraphrase of the title — exact/substring matching is both cheaper (no
embedding API call per lookup) and more precise than approximate
nearest-neighbor search for that specific access pattern.
reindexForSemanticSearch() still exists as an explicit opt-in per
workspace for the case an agent does want the catalog folded into its
normal RAG recall (AiMemoryService) alongside findings/notes/recon data
in one unified semantic search — never run automatically, since embedding
the whole catalog costs one provider call per entry per workspace.
6. Privacy Mode (task #234): a real enforcement point, with an honestly narrower scope than ADR 0007's "local-first"
AiPrivacyModeService/AiProviderRegistryService.resolveProvider()
implement a genuine per-workspace CLOUD/LOCAL/HYBRID restriction at
the one seam every AI/agent call already passes through: LOCAL throws
AiPrivacyModeViolationError rather than silently falling back to a cloud
provider when nothing local is configured; HYBRID reorders the fallback
candidate list to try local providers first (though a configured default
provider still wins via the pre-existing default-provider shortcut — see
that service's own tests for this exact, intentional precedence); CLOUD
is unrestricted (today's Module 5 behavior, unchanged). This is a real,
enforced restriction — not a persisted preference nobody reads — but it is
explicitly narrower than ADR 0007's full "local-first" definition:
it only restricts which backend AI provider a request may use, not
Desktop-Agent-routing execution. A LOCAL-mode workspace still runs
inference against a provider from this server process (typically Ollama
reachable from it), not necessarily the user's own Desktop Agent. This gap
is disclosed here and in the service's own doc comment, not silently
narrowed without a trace.
7. Performance (task #235): fix real N+1s and unnecessary round-trips found by review, not speculative optimization
Four concrete issues were found and fixed by a dedicated research pass,
not by pre-emptively micro-optimizing: (a) the orchestration run's
task_started SSE event was calling the full toTaskDto() (two DB
round-trips: tool calls + sub-tasks) for a task that, being freshly
started, cannot yet have either — fixed with a skipChildren fast path;
(b) AgentExplainabilityService.explainRun() and
AgentOrchestratorService's own run-detail view both independently
N+1'd one listToolCallLogsForTask query per task — fixed by one new
batched listToolCallLogsForTasks(agentTaskIds) repository method (a
single WHERE "agentTaskId" IN (...) round-trip) that both now share; (c)
KnowledgeBaseReferenceService was re-querying the full (small,
rarely-changing) catalog from the database on every single search()/
listEntries() call despite being a hot path (one agent tool call per
orchestration task) — fixed with an in-memory cache invalidated on
write, not time-based expiry, since the only writes are curation
actions/the one-time boot seed; (d) AiMemoryService.upsertChunks() was
issuing one $executeRaw INSERT per chunk in a loop when indexing long
content — fixed with one multi-row VALUES (...), (...), ... INSERT via
Prisma.join.
8. HTTP surface (task #236): CQRS commands/queries + SSE, matching every prior AI-adjacent controller exactly
Before task #236, AgentOrchestratorService/KnowledgeBaseReferenceService/
AgentExplainabilityService were reachable only via direct injection from
AutomationEngineModule's WorkflowExecutorService (the ai_agent_run
workflow step) — no human user or frontend could start/inspect/explain a
run, or browse/search/curate the Knowledge Base, at all. AgentRunsController
and KnowledgeBaseReferencesController close that gap with the exact same
CQRS command/query dispatch (CommandBus/QueryBus) and @Sse()
live-stream-with-persisted-fallback pattern AiConversationsController
established in Module 5 — a run's stream endpoint checks an in-memory
AgentOrchestrationStreamHubService first, falling back to one final
synthetic event from the persisted OrchestrationRun if no live stream is
active. Domain events (OrchestrationRunCreated/Completed/Failed/Cancelled)
publish into the same single global ActivityRecordingHandler every other
module's events already flow through — no new event-handling
infrastructure.
9. Frontend (task #237): a distinct "Agent Copilot" surface from Module 5's "AI Assistant," not a merged UI
/agent (run list + start-a-run + live activity stream + task tree +
explainability tabs) and /knowledge-base (catalog browse/search) are new
top-level nav entries, deliberately separate from Module 5's existing
/ai chat page — a multi-step, multi-agent orchestration run has a
fundamentally different shape (a task tree that grows over minutes, not a
message that streams over seconds) than a chat turn, and conflating the
two UIs would have made both worse. The live activity panel reuses the
exact manual-fetch()-plus-ReadableStream SSE-consumption pattern
useReconJobLogs established in Module 3 (native EventSource can't
attach the required Bearer JWT header), and the run-detail page polls
GetOrchestrationRun the same way useReconJob polls a running scan —
the SSE stream is a live "what's happening right now" feed layered on top
of, not a replacement for, that polling-based authoritative state.
Consequences
- A second, deliberately distinct AI surface, not a Module 5 rewrite.
Every Module 5 table, endpoint, and UI page is unchanged. A workspace
that never starts an orchestration run never encounters
agentmodule code at all. - The Privacy Mode gap (point 6) is a disclosed, intentional scope boundary, not a defect — but it is the one place in this module where a workspace operator's mental model ("LOCAL means nothing leaves my infrastructure") could diverge from actual behavior if the gap to ADR 0007's Desktop-Agent-routing definition isn't read. Flagged here specifically so it isn't mistaken for solved.
- Sandbox verification, disclosed honestly.
packages/sharedandapps/api's fulltsc --noEmit/eslintwere run to completion this pass via a detached-background-process technique (setsid nohup ... & disown, polled via separate tool calls) that worked reliably for the backend and confirmed zero compile/lint errors across every Module 11 backend file. The identical technique did not reproduce forapps/webin this session — background processes were observed to terminate between tool calls rather than persisting, and even a maximally-scopedtscrun (a handful of new files only) exceeded the available foreground command budget, consistent with this environment's mounted-filesystem I/O being the bottleneck rather than compute. The Module 11 frontend was instead verified by direct, systematic manual cross-reference of every new file's types, imports, and prop shapes against the real@pentesthub/sharedDTOs and the actual signatures of every hook/component it calls (seedocs/modules/11-ai-security-copilot-multi-agent-system.md's verification section for the specific issue this caught and fixed: a tuple-inference bug innew Map(array.map(...))on the Knowledge Base page). This is a real, disclosed gap versus an automatedtscpass — not claimed as equivalent, and called out as the first thing a real verification pass (ci.yml, once runnable) should re-check. - Test coverage is real but intentionally scoped to the highest-risk
logic, not exhaustive. The
agentmodule had zero test files before this pass (a fact this pass discovered mid-way, since directory-listing tools were unreliable in this sandbox session and had to be cross-checked via content search instead). New/fixed specs now cover: Privacy Mode enforcement end-to-end (the security-critical path — a regression here is a silent data-exfiltration risk, not just a broken feature), Explainability's tree-reconstruction and evidence summarization, Knowledge Base caching/invalidation/search-ranking, and the Orchestrator's pure/independently-testable logic (mergeOutputs,cancel/retry/getRun, thetoTaskDtoskipChildren fast path). Full end-to-end orchestration-run execution (a live model driving all twelve agents through a real task graph), and the narrower task #228-#232 enhancements (autonomous recon/vuln analysis, incremental report writing, chat branching/bookmarks, voice copilot, workflow automation integration), do not have dedicated new specs in this pass — a disclosed next step, not an oversight.