Module 5 Phase 1 Release Notes — AI Security Copilot (the "brain")
Full decision record: docs/adr/0005-ai-security-copilot-module.md.
Per the agreed phasing, this release covers Phase 1 only — AI Chat, automatic context, RAG, the provider abstraction, tool calling, and streaming. Phase 2 (Voice, code export, artifacts, canvas diagrams, terminal mode) is explicitly out of scope for this release; work stops here pending approval to continue.
Features added
Backend — the AI Security Copilot engine:
- AI Chat: workspace/project-aware, persistent conversations
(
AiConversation→AiMessage), streamed token-by-token over SSE, with Cancel, Regenerate, and Continue all wired throughAiChatStreamHubService. - Automatic Context:
AiContextServicegathers targets, recon findings, vulnerability findings, notes, evidence, activity, and scan templates for the active project on every turn — the user never pastes data in manually. - RAG: pgvector-backed semantic memory (
ai_memory_chunks,vector(1536), raw-SQL cosine-distance search), isolated per workspace and optionally per project; conversation turns are indexed automatically after completion so the assistant can recall its own prior answers. - AI Provider abstraction: 5 adapters — OpenAI, Anthropic, Gemini,
Ollama, OpenRouter — behind one
AiProviderinterface and anAiProviderRegistryServiceregistry (same pattern as Module 4'sScannerRunnerRegistry). No API keys are configured in this environment by explicit agreement — every provider is a valid, visibly "not configured" entry inGET /ai/providers, wired for later. - Prompt Architecture: 6 independently-editable prompt modules (system, security, recon, finding-analysis, report, coding) — never hardcoded in controllers or services.
- AI Tool Calling: a provider-agnostic
<tool_call>text-tag protocol (works even on providers with no native function-calling, e.g. Ollama). 11 tools at launch: search findings/recon data/evidence/notes/projects/ targets/templates, create note, draft report, generate PoC, generate payloads — registered viaAiToolRegistryService, extensible by future modules without touching the chat engine. - Streaming: SSE token deltas, tool-call/tool-result events, mid-stream
cancellation via
AbortController, capped at 3 tool calls per turn. - Security: prompt-injection guardrails (
PromptInjectionGuardServiceneutralizes literal<tool_call>tags in user input, truncates and data-boundary-wraps tool output), the same enumeration-safe workspace- isolation checkpoint every other module uses (resolveAiConversationForCaller/resolveAiMessageForCaller), and a tighter rate limit on message-sending endpoints (20/min). - Event Integration:
AiSummaryHandlersubscribes toReconFindingCreatedEvent/VulnFindingCreatedEvent/VulnFindingResolvedEvent, filtered to Critical/High severity, running in both the API and worker processes (same dual-registration shape asActivityRecordingHandler). - REST API: conversations (CRUD, pin, archive), messages (send, regenerate,
continue, cancel, bookmark, pin, delete), providers status, context
summary + suggested questions, memory search — fully documented in
Swagger (
@ApiTags('ai')).
Frontend — /ai AI Dashboard:
?projectId=+<ProjectPicker>fallback (matching every other project-scoped page's convention), wrapped in<Suspense>.AiConversationSidebar(search, create, pin, delete),AiChatPanel(streaming message list with the Phase 1 action set — copy, bookmark, pin, regenerate, continue, delete — driven byuseAiMessageStream, the same manualfetch()+ReadableStreamSSE pattern as Module 4's scan job logs),AiContextPanel(context indicator badges + suggested questions wired into the chat panel via an imperative handle).- Rich Markdown rendering (tables, code blocks with a hover copy button,
GFM) reusing the existing
react-markdown/remark-gfm/rehype-highlightstack already used by Notes.
Architectural improvements
- Provider-agnostic tool calling via a text-tag protocol is a new pattern for this codebase — avoids needing 5 different native function-calling integrations (one of which, Ollama, doesn't have one), at the cost of parsing model output rather than a structured API.
- pgvector on the existing Postgres instance proves the "no new infrastructure service" approach for vector search — one extension, one raw-SQL-isolated repository, no separate vector database to operate.
- In-process SSE pub/sub hub (
AiChatStreamHubService) is a genuinely new pattern versus Module 3/4's polling-based live updates — true token-by-token push, at the documented cost of being single-process-only for now. AiCoreModule/AiModulesplit extends the Module 3/4 core-module/HTTP-module pattern to a module whose event handler (AiSummaryHandler) must run in both the API and worker OS processes.
Breaking changes
None. Every addition is additive: new enums, new tables (including the new
vector Postgres extension), a new module wired into the existing
AppModule/WorkerModule, one new frontend route (/ai, promoted from
FUTURE_NAV to PRIMARY_NAV). No Module 1–4 endpoint, contract, or schema
column changed.
Known limitations
- No AI provider is configured. Every provider-calling code path
(chat generation, RAG embedding,
draft_report/generate_poc/generate_payloads) will throwAiProviderNotConfiguredErroruntil real API keys (or a reachable local Ollama endpoint) are supplied — by explicit agreement with the project owner, this is expected Phase 1 state, not a bug. AiChatStreamHubServiceis single-process only — no Redis or shared broker backs it yet. Correct for one API instance; would need a shared pub/sub layer before horizontally scaling the API process.- Phase 2 was unbuilt at the time of this release — see
docs/releases/module-5-phase-2-release-notes.md, shipped the same day once Phase 1 verification/commit/push completed and Phase 2 was approved: Voice Interaction, Smart Code Block toolbar, One-Click Code/ Conversation Export, Terminal Mode, Artifact Mode, and Interactive Canvas are now all implemented. - The tool-call text-tag protocol depends on the model reliably emitting well-formed tags — malformed JSON degrades to an empty-input tool call rather than failing the turn, and an unrecognized tool name resolves to an error result fed back to the model, but a model that never learns to emit the tag at all (unlikely with the system prompt's explicit instructions, but possible with a very small local Ollama model) simply never calls tools — there's no fallback native-function-calling path.
- Full
pnpm build/lint/check-types/testverification for this module is run by the project owner outside this environment (the sandbox this work was authored in cannot fetch Prisma engine binaries, run the pgvector-dependent migration against a real Postgres instance, or reliably run a full Next.js typecheck) — see the ADR's Consequences section and the commit history for the verification pass.
Recommended next step
Per explicit instruction, work stopped here pending approval — Phase 2
was subsequently approved and shipped the same day; see
docs/releases/module-5-phase-2-release-notes.md.