kaveo · Concepts
Evidence and grounded AI
How kaveo's evidence ledger works and how its AI stages cite stored observations, which providers and model tiers it supports, and what runs offline.
On this page
kaveo never asks a model whether something is wrong. Deterministic detectors decide that, and what each collector read during the scan is recorded in an evidence ledger: one row per collector run, or one per API call and page for collectors such as IAM. The AI layer reasons over findings and that evidence, and any claim that does not cite the ledger is discarded.
The evidence ledger#
A collector reports its work as a list of API calls, and each call becomes one entry in the ledger. A simple collector reports its whole run as one call. Collectors that page through an API report each page separately, labelled like iam:GetAccountAuthorizationDetails#page=2.
Each entry holds the API call label, a sha256: digest of the canonical JSON of the resources the call returned, and an optional redacted locator. Collector entries leave the locator empty today. The permission-simulation entries described under agentic investigation use it to record the verdict and the IAM action. The database rejects every edit and every direct delete of a ledger entry. Entries are removed only when the scan they belong to is deleted.
The worker links each finding to the observations behind it. Any role can read the evidence:
Five AI stages#
Four stages run per finding at POST /v1/findings/{finding_id}/prioritize, /explain, /remediate and /compliance-impact, each with a /stream variant that returns Server-Sent Events. When the rule has an entry in kaveo's compliance knowledge base, compliance_impact also receives its curated controls, breach-cost range and documented incident, and is instructed to use only those and never invent new ones.
investigate streams from POST /v1/investigate:
curl -N -X POST "$KAVEO_API_URL/v1/investigate" \
-H "Authorization: Bearer $KAVEO_API_TOKEN" \
-H "Content-Type: application/json" \
-d '{"finding_id": "FINDING_ID", "question": "Which principals can reach this bucket?", "depth": "main"}'
question is optional, up to 2000 characters. depth (fast, main or deep) selects the model tier. Events arrive as start, one claim per grounded claim, content, then done. The gate runs before the first event, so a stream never carries an unchecked claim.
Every stage endpoint, and POST /v1/nlquery, needs the investigate permission, which analysts and admins have. A default API token lacks it, so mint one with the investigate scope for scripts.
The citation gate#
Every stage output passes through the gate:
- The prompt lists the only observation ids the model may cite: those linked to this finding, or its scan's ledger when the finding has no links.
- The model answers in JSON: claims, each with citations, plus stage content such as a summary or an artifact. If the reply cannot be parsed, the stage falls back to deterministic template claims that cite the supplied observations.
- A claim survives only if it cites at least one id and every cited id is in that list. Every other claim is dropped.
- kaveo stores the output together with the observation ids its surviving claims cite, in one statement. An output with no surviving claim returns
persisted: falseand is never stored.
fully_grounded is false whenever a claim was dropped, so a client can flag a partial answer.
The AI layer cannot author a detector finding. A check on every CI run fails the build if the AI layer imports detection or collection code.
Where a value must exist regardless of the model, code computes it. The prioritize score starts from a fixed base per severity. Graph facts add a small bump when the resource is internet-exposed or reaches a crown jewel through a privilege edge, and a larger one when both hold. The model supplies only the cited reasons. Likewise, when you propose a fix for approval, kaveo checks without a model that an IAM policy is strictly narrower than the grant the detector flagged.
Grounding proves that each cited observation exists and was part of the evidence kaveo supplied. It does not prove the model read it correctly, so keep a person on judgment calls.
Untrusted input handling#
Text from your cloud, and every question you type, is treated as data:
- Finding titles, finding details, observation labels and locators, crown-jewel labels and your question are wrapped in untrusted-data fences. A closing delimiter inside a value is neutralized, so text cannot end its own fence.
- The system prompt says fenced text is evidence, and tells the model to ignore any attempt in it to change its role, reveal the prompt or alter the output format.
- A finding's details are cut to 4000 characters before they enter a prompt.
Fencing lowers the chance a model is steered. The citation gate bounds the output even if it is. See Security model for the wider picture.
Providers and model tiers#
Set the provider with KAVEO_AI_PROVIDER:
Stock compose passes all of these except the xAI key; Configuration shows the override. If anthropic or bedrock fails, the call returns 503. openai, grok and local fall back to the offline stub and log a warning.
The defaults are Anthropic model ids. On openai and grok, kaveo uses its own catalog model for each tier unless you set a different value. Set them to Bedrock model or inference-profile ids on bedrock, to the provider's model names when KAVEO_OPENAI_BASE_URL points at another OpenAI-compatible service, or to models your server serves on local.
An admin can switch provider without a restart in Settings → AI provider, or with PUT /v1/ai/config. A hosted provider can be selected only once its key is set. The switch applies to the api process that received it and is lost on restart, so keep KAVEO_AI_PROVIDER as your lasting default. The worker, which runs agentic investigation and autonomous patrol, always uses KAVEO_AI_PROVIDER. GET /v1/ai/config shows the active provider and its routing.
Agentic investigation (opt-in)#
Agentic investigation hunts for attack chains across a whole scan. It is off by default because a run makes many model calls, and it is only meaningful with a real provider. Set KAVEO_AGENT_INVESTIGATION_ENABLED=true.
Five specialist agents run one after another, covering data exfiltration, exposure, lateral movement, privilege escalation and secrets. Each explores the scan's stored graph with read-only tools and proposes candidate chains, within KAVEO_AGENT_MAX_ITERATIONS steps (default 12). Agents never call AWS and never touch the deterministic findings.
A deterministic validator runs each permission hop through iam:SimulatePrincipalPolicy with the scan's read-only session, and writes each verdict to the ledger for the hop to cite. A chain is validated only when every permission hop simulates allow. Unvalidated chains are kept but flagged. If kaveo cannot assume the account's read-only role, as on a synthetic scan, every chain stays unvalidated.
Natural-language graph query#
POST /v1/nlquery answers a question about a scan's graph, such as what can reach a crown-jewel database. Send scan_id and question, plus conversation_id to continue a thread. In the console this is Ask kaveo.
The model never returns data. It picks one query from a fixed list: list resources, forward or reverse blast radius, a reachability summary, attack paths, or choke points. If its choice is not on the list, kaveo picks one with a keyword heuristic. kaveo runs it over the stored graph, using the model's arguments only to filter returned rows, never inside SQL. The model then narrates the result, citing the evidence behind the edges it surfaced, and the gate drops any claim that cites anything else. The response carries persisted: false, because answers are not stored as AI outputs. The question and answer text are kept in the conversation.
Next steps#
- How kaveo works: how a scan fills the ledger and builds the graph.
- Configuration: every AI setting and the compose override.
- Remediation and autonomous patrol: approve fixes and see how kaveo verifies them once applied.