Skip to content

kaveo · Concepts

How kaveo works

How kaveo turns read-only cloud API calls into findings, an evidence ledger, a reachability graph and grounded AI output, and what each container does.

On this page

kaveo keeps two jobs apart. Deterministic code collects your cloud configuration and decides what is wrong with it. An AI layer then explains and prioritizes those results, and it can only cite evidence the deterministic side already stored. This page follows a scan through the system and names the boundaries that keep the two sides separate.

The pipeline#

The kaveo pipelineThe kaveo pipeline, left to right. Cloud accounts are read through read-only access by the collectors. The collectors produce two things: entries in the evidence ledger, which holds hashed observations, and resources plus the edges between them. Detectors turn resources and edges into findings, and graph queries walk the edges to build attack paths. The AI layer reads the evidence ledger and the findings and graph results through its stages: prioritize, explain, remediate, investigate and compliance impact. Everything it produces goes through the citation gate, where ungrounded claims are dropped. What passes is stored as AI output with its citations, or as drafted fixes. Drafted fixes start remediation sagas, from the console or the patrol.cloudaccountsREAD-ONLYACCESScollectorsevidence ledgerHASHED OBSERVATIONSresources+ edgesdetectorsfindingsgraph queries,attack pathsAI layerprioritizeexplainremediateinvestigatecompliance impactCITATION GATEungrounded claimsdroppedAI outputwith its citationsdrafted fixesremediation sagasfrom the consoleor the patrolThe kaveo pipelineThe kaveo pipeline, left to right. Cloud accounts are read through read-only access by the collectors. The collectors produce two things: entries in the evidence ledger, which holds hashed observations, and resources plus the edges between them. Detectors turn resources and edges into findings, and graph queries walk the edges to build attack paths. The AI layer reads the evidence ledger and the findings and graph results through its stages: prioritize, explain, remediate, investigate and compliance impact. Everything it produces goes through the citation gate, where ungrounded claims are dropped. What passes is stored as AI output with its citations, or as drafted fixes. Drafted fixes start remediation sagas, from the console or the patrol.cloud accountsREAD-ONLY ACCESScollectorsevidence ledgerHASHEDOBSERVATIONSresources+ edgesdetectorsfindingsgraph queries,attack pathsAI layerprioritize · explain · remediateinvestigate · compliance impactCITATION GATEungroundedclaims droppedAI output withits citationsdrafted fixesremediation sagasfrom the consoleor the patrol

Collectors read your accounts through read-only access and produce resources, plus an evidence ledger entry for the calls behind them. Edges between resources form the IAM and reachability graph. Detectors turn resources and edges into findings. Graph queries walk the edges recursively in Postgres to build attack paths, choke points and blast radius, and findings are grouped into issues.

The AI layer reads findings, graph results and observations. Every stage output goes through a citation gate that checks each cited observation id against the ledger and drops claims whose citations don't resolve. An output is stored, together with its citations, only when cited evidence survives. A fix drafted by the remediate stage is also checked without a model (for an IAM policy: valid JSON that is strictly narrower than the flagged grant) and stored as a draft fix. Every remediation saga starts from one of those drafts, whether you propose it in the console or the patrol drafts it after a scan.

The MCP server sits beside this flow, not after it. It is off unless you set KAVEO_MCP_ENABLED=true. Its 7 read-only tools query scans, findings, attack paths, compliance status and observations directly, scoped to the org of the token that calls them. See MCP and integrations.

Containers#

kaveo runs as one Docker Compose stack of five containers on a private bridge network.

ServiceRole
caddyIngress, on host ports 80 and 443 by default. /v1/* and /mcp go to api, everything else to web. It serves plain HTTP on port 80 until you set KAVEO_SITE_ADDRESS to a real hostname, which turns on automatic TLS.
webThe Next.js console, on port 3000 inside the network.
apiFastAPI service on port 8000. At startup it applies pending migrations and, when no user exists yet, creates the owner account.
workerSame build as api, started as a background worker. Runs scans, remediation sagas, the scan and report schedulers, the patrol and agent investigations.
postgresPostgres 16, the system of record.

The worker waits until api reports healthy, so the api has finished its startup migrations before the worker claims a job. An optional model service runs Ollama for the local AI provider. It is off unless you start the stack with --profile local-model.

Caddy doesn't route /health or /metrics. Both are served by api on port 8000 inside the compose network, so probe and scrape them from there. See Deployment.

A scan, step by step#

The API inserts a scan as queued. The worker claims it with a row lock that skips jobs another worker already holds, marks it running and commits right away, so the long-running scan never holds a row lock. Several worker replicas can pull from the same queue without handing out a job twice.

The scan then runs as a single transaction:

  1. Scope the transaction to the scan's org, so every read and write below runs under that org's row-level security policy.
  2. Load the account and collect. For AWS, the worker assumes the read-only role and resolves the regions: the ones the scan requested, or every region enabled for the account from ec2:DescribeRegions. Global collectors such as IAM, S3 and CloudTrail run once per account, and every other AWS collector runs once per region. A collector that raises is logged, recorded with a count of zero and skipped. KAVEO_SCAN_MODE sets the fallback: aws collects for real only, synthetic never contacts a cloud and scans a labelled fixture, and auto uses that fixture only when kaveo has no base AWS credentials. A denied assume fails the scan. Azure, GCP, Kubernetes and GitHub accounts run their own collectors with the credential you registered, and any error there fails the scan rather than falling back to demo data. See Connect accounts.
  3. Append observations. A collector writes one ledger entry per run by default. Collectors that report per-call provenance, such as IAM, write one entry per API call, and each pagination page counts as a call. Each entry stores a SHA-256 hash of the normalized resources that call produced, not the payload itself.
  4. Persist resources, then derive and persist graph edges. Each edge cites an observation, normally the one behind its source resource.
  5. Run every detector registered for the account's provider over an in-memory view of those facts. A detector that raises is logged and skipped. Findings pass through suppression, which labels them rather than dropping them, and are written with their compliance mappings. A finding about a collected resource is also linked to the observation behind that resource.
  6. Import native GuardDuty, Inspector and Security Hub findings. Each record becomes one finding marked as native, goes through the same suppression pass, and cites the observation of the call that collected it.
  7. Validate attack paths. Toxic-path findings start hidden, because a graph edge shows that a step is possible, not that it is permitted. On a live AWS scan, each permission hop is checked with the read-only iam:SimulatePrincipalPolicy call. Only a chain where every hop passes is shown.
  8. Record the coverage map and mark the scan succeeded.

Everything commits at once, so a crash never leaves a half-written scan. On an error the transaction rolls back and the scan is marked failed. If a worker's lease expires and another worker re-claims the scan, the final write detects it and discards the first worker's results.

After a successful commit, the worker runs follow-up steps. A failure in any of them is logged and never changes the scan's outcome. If the brief fails, the steps after it are skipped for that scan:

  • Brief. New and resolved findings compared with the previous successful scan, plus regressions of verified fixes.
  • Patrol. When KAVEO_PATROL_ENABLED=true, it drafts fixes for the top new risks as proposed sagas. Nothing is applied without approval.
  • Issues and alerts. Findings are grouped into issues per resource, new findings go to your enabled alert channels, and framework scores are stored.
  • Slack. When a Slack webhook is configured, the brief is posted there.

Worker loop order#

Each pass of the worker loop runs in this order:

  1. Write a heartbeat. /metrics exposes its age as kaveo_worker_seconds_since_heartbeat, so you can alert on a stalled worker.
  2. Reap crashed jobs. Scans and sagas whose 900-second lease expired go back to the queue, and agent runs use an 1800-second lease.
  3. Scheduler tick: queue scans for accounts whose schedule is due.
  4. Report tick: email the latest brief to accounts whose report schedule is due.
  5. Run one approved remediation saga, if any.
  6. Otherwise, run one queued scan.
  7. Otherwise, run one agent investigation.
  8. Otherwise, sleep for the poll interval, 2 seconds by default.

After a job, the next pass starts straight away. Approved sagas go first, so an approval never waits behind a queue of scans, and agent investigations go last, so they never delay the deterministic pipeline.

A scan that runs longer than its lease is presumed dead and re-queued. If your largest accounts take more than 15 minutes to scan, raise KAVEO_WORKER_LEASE_SECONDS above your longest scan. See Configuration.

Plugins by file drop#

Detectors and collectors register themselves. Each one is a class in a catalog directory that subclasses a base class, and discovery imports every module in that directory, including subdirectories. There is no central list.

Both registries sort by id, so scans run in a stable order, and both reject a duplicate id with an error instead of letting one class replace another. If discovery fails, the api logs it and keeps serving, so confirm what loaded: /health on the api service reports the detector and collector counts, and the worker logs them at the start of each scan. Adding coverage means adding a file. Today kaveo ships 97 detectors and 48 collectors. See Detection.

What is stored#

Postgres 16 holds everything kaveo keeps, grouped by area:

  • Scan results. Connected accounts, scans, the resources and graph edges each scan collected, findings with their links to evidence, compliance mappings and framework scores.
  • Evidence. The evidence ledger: one hashed entry per collector run or API call.
  • AI output. Each stored stage output together with the observations it cites.
  • Remediation workflow. Drafted fixes, remediation sagas, agent investigation runs, scan briefs, patrol runs and issues.
  • Tenancy and access. Orgs, memberships, users, sessions, and MCP and API tokens.
  • Integrations and audit. Alert channels and the audit log.

The evidence ledger and the audit log are append-only. The database rejects any update and any direct delete. Rows leave only through a cascade from their parent: deleting a scan removes its evidence, and erasing a tenant removes its audit entries. A migration also revokes update and delete on the audit log from the app's database role.

Everything that holds your data, from scan results and evidence to workflow, alert channels and the audit log, is scoped to an org and protected by a row-level security policy. In the stock compose file, api and worker connect as a non-superuser database role, so Postgres enforces those policies, and only the migration runner connects as the superuser. The schema comes from 45 forward-only SQL migrations. api applies pending ones at startup behind a Postgres advisory lock. When migrations are pending, it takes a pg_dump snapshot first, still under the lock. The stock compose file turns this on with KAVEO_BACKUP_BEFORE_MIGRATE.

Boundaries enforced in CI#

The backend is one codebase, but the boundaries between its parts are checked on every CI run, and the build fails if one is crossed:

  • The shared data models that pass between parts depend on nothing else in kaveo.
  • The AI layer cannot import detection or collection code, so it can't run a detector or a collector.
  • Collection code cannot import detection or AI code: collectors report and detectors judge.
  • Product logic cannot import the API layer, the app's startup code or the worker.
  • The kaveo CLI does not reach into server code. It talks to /v1/* over HTTP only.

For how these boundaries protect your data, see Security model. For how citations are checked, see Evidence and grounded AI.