kaveo · Concepts
How kaveo works
How kaveo turns read-only cloud API calls into findings, an evidence ledger, a reachability graph and grounded AI output, and what each container does.
On this page
kaveo keeps two jobs apart. Deterministic code collects your cloud configuration and decides what is wrong with it. An AI layer then explains and prioritizes those results, and it can only cite evidence the deterministic side already stored. This page follows a scan through the system and names the boundaries that keep the two sides separate.
The pipeline#
Collectors read your accounts through read-only access and produce resources, plus an evidence ledger entry for the calls behind them. Edges between resources form the IAM and reachability graph. Detectors turn resources and edges into findings. Graph queries walk the edges recursively in Postgres to build attack paths, choke points and blast radius, and findings are grouped into issues.
The AI layer reads findings, graph results and observations. Every stage output goes through a citation gate that checks each cited observation id against the ledger and drops claims whose citations don't resolve. An output is stored, together with its citations, only when cited evidence survives. A fix drafted by the remediate stage is also checked without a model (for an IAM policy: valid JSON that is strictly narrower than the flagged grant) and stored as a draft fix. Every remediation saga starts from one of those drafts, whether you propose it in the console or the patrol drafts it after a scan.
The MCP server sits beside this flow, not after it. It is off unless you set KAVEO_MCP_ENABLED=true. Its 7 read-only tools query scans, findings, attack paths, compliance status and observations directly, scoped to the org of the token that calls them. See MCP and integrations.
Containers#
kaveo runs as one Docker Compose stack of five containers on a private bridge network.
The worker waits until api reports healthy, so the api has finished its startup migrations before the worker claims a job. An optional model service runs Ollama for the local AI provider. It is off unless you start the stack with --profile local-model.
Caddy doesn't route /health or /metrics. Both are served by api on port 8000 inside the compose network, so probe and scrape them from there. See Deployment.
A scan, step by step#
The API inserts a scan as queued. The worker claims it with a row lock that skips jobs another worker already holds, marks it running and commits right away, so the long-running scan never holds a row lock. Several worker replicas can pull from the same queue without handing out a job twice.
The scan then runs as a single transaction:
- Scope the transaction to the scan's org, so every read and write below runs under that org's row-level security policy.
- Load the account and collect. For AWS, the worker assumes the read-only role and resolves the regions: the ones the scan requested, or every region enabled for the account from
ec2:DescribeRegions. Global collectors such as IAM, S3 and CloudTrail run once per account, and every other AWS collector runs once per region. A collector that raises is logged, recorded with a count of zero and skipped.KAVEO_SCAN_MODEsets the fallback:awscollects for real only,syntheticnever contacts a cloud and scans a labelled fixture, andautouses that fixture only when kaveo has no base AWS credentials. A denied assume fails the scan. Azure, GCP, Kubernetes and GitHub accounts run their own collectors with the credential you registered, and any error there fails the scan rather than falling back to demo data. See Connect accounts. - Append observations. A collector writes one ledger entry per run by default. Collectors that report per-call provenance, such as IAM, write one entry per API call, and each pagination page counts as a call. Each entry stores a SHA-256 hash of the normalized resources that call produced, not the payload itself.
- Persist resources, then derive and persist graph edges. Each edge cites an observation, normally the one behind its source resource.
- Run every detector registered for the account's provider over an in-memory view of those facts. A detector that raises is logged and skipped. Findings pass through suppression, which labels them rather than dropping them, and are written with their compliance mappings. A finding about a collected resource is also linked to the observation behind that resource.
- Import native GuardDuty, Inspector and Security Hub findings. Each record becomes one finding marked as native, goes through the same suppression pass, and cites the observation of the call that collected it.
- Validate attack paths. Toxic-path findings start hidden, because a graph edge shows that a step is possible, not that it is permitted. On a live AWS scan, each permission hop is checked with the read-only
iam:SimulatePrincipalPolicycall. Only a chain where every hop passes is shown. - Record the coverage map and mark the scan
succeeded.
Everything commits at once, so a crash never leaves a half-written scan. On an error the transaction rolls back and the scan is marked failed. If a worker's lease expires and another worker re-claims the scan, the final write detects it and discards the first worker's results.
After a successful commit, the worker runs follow-up steps. A failure in any of them is logged and never changes the scan's outcome. If the brief fails, the steps after it are skipped for that scan:
- Brief. New and resolved findings compared with the previous successful scan, plus regressions of verified fixes.
- Patrol. When
KAVEO_PATROL_ENABLED=true, it drafts fixes for the top new risks asproposedsagas. Nothing is applied without approval. - Issues and alerts. Findings are grouped into issues per resource, new findings go to your enabled alert channels, and framework scores are stored.
- Slack. When a Slack webhook is configured, the brief is posted there.
Worker loop order#
Each pass of the worker loop runs in this order:
- Write a heartbeat.
/metricsexposes its age askaveo_worker_seconds_since_heartbeat, so you can alert on a stalled worker. - Reap crashed jobs. Scans and sagas whose 900-second lease expired go back to the queue, and agent runs use an 1800-second lease.
- Scheduler tick: queue scans for accounts whose schedule is due.
- Report tick: email the latest brief to accounts whose report schedule is due.
- Run one approved remediation saga, if any.
- Otherwise, run one queued scan.
- Otherwise, run one agent investigation.
- Otherwise, sleep for the poll interval, 2 seconds by default.
After a job, the next pass starts straight away. Approved sagas go first, so an approval never waits behind a queue of scans, and agent investigations go last, so they never delay the deterministic pipeline.
A scan that runs longer than its lease is presumed dead and re-queued. If your largest accounts take more than 15 minutes to scan, raise KAVEO_WORKER_LEASE_SECONDS above your longest scan. See Configuration.
Plugins by file drop#
Detectors and collectors register themselves. Each one is a class in a catalog directory that subclasses a base class, and discovery imports every module in that directory, including subdirectories. There is no central list.
Both registries sort by id, so scans run in a stable order, and both reject a duplicate id with an error instead of letting one class replace another. If discovery fails, the api logs it and keeps serving, so confirm what loaded: /health on the api service reports the detector and collector counts, and the worker logs them at the start of each scan. Adding coverage means adding a file. Today kaveo ships 97 detectors and 48 collectors. See Detection.
What is stored#
Postgres 16 holds everything kaveo keeps, grouped by area:
- Scan results. Connected accounts, scans, the resources and graph edges each scan collected, findings with their links to evidence, compliance mappings and framework scores.
- Evidence. The evidence ledger: one hashed entry per collector run or API call.
- AI output. Each stored stage output together with the observations it cites.
- Remediation workflow. Drafted fixes, remediation sagas, agent investigation runs, scan briefs, patrol runs and issues.
- Tenancy and access. Orgs, memberships, users, sessions, and MCP and API tokens.
- Integrations and audit. Alert channels and the audit log.
The evidence ledger and the audit log are append-only. The database rejects any update and any direct delete. Rows leave only through a cascade from their parent: deleting a scan removes its evidence, and erasing a tenant removes its audit entries. A migration also revokes update and delete on the audit log from the app's database role.
Everything that holds your data, from scan results and evidence to workflow, alert channels and the audit log, is scoped to an org and protected by a row-level security policy. In the stock compose file, api and worker connect as a non-superuser database role, so Postgres enforces those policies, and only the migration runner connects as the superuser. The schema comes from 45 forward-only SQL migrations. api applies pending ones at startup behind a Postgres advisory lock. When migrations are pending, it takes a pg_dump snapshot first, still under the lock. The stock compose file turns this on with KAVEO_BACKUP_BEFORE_MIGRATE.
Boundaries enforced in CI#
The backend is one codebase, but the boundaries between its parts are checked on every CI run, and the build fails if one is crossed:
- The shared data models that pass between parts depend on nothing else in kaveo.
- The AI layer cannot import detection or collection code, so it can't run a detector or a collector.
- Collection code cannot import detection or AI code: collectors report and detectors judge.
- Product logic cannot import the API layer, the app's startup code or the worker.
- The
kaveoCLI does not reach into server code. It talks to/v1/*over HTTP only.
For how these boundaries protect your data, see Security model. For how citations are checked, see Evidence and grounded AI.