Skip to content

kaveo · Concepts

Security model

kaveo's trust model: read-only scan roles, a separate write role for approved fixes, no outbound traffic by default, tenant isolation and hashed credentials.

On this page

This page covers what kaveo can read and change in your cloud, what leaves the stack, how tenants are kept apart and how callers are authenticated. It ends with what to set before you expose a deployment.

Read-only by default#

kaveo scans AWS by assuming a cross-account role in the target account, so it never holds long-lived keys for that account. The read-only role template that ships with kaveo, in CloudFormation and Terraform versions, creates the role kaveo-readonly (see Connect accounts):

  • The trust policy allows sts:AssumeRole only from kaveo's own principal (KaveoPrincipalArn) and only with a matching sts:ExternalId. The CloudFormation template requires the ExternalId to be at least 8 characters.
  • It attaches the AWS-managed read-only policies SecurityAudit and job-function/ViewOnlyAccess.
  • MaxSessionDuration is 3600 seconds.

The collector contract forbids calling any mutating API or storing long-lived credentials.

Writes are separate, scoped and approved#

kaveo cannot change a scanned AWS account unless you deploy a second role there and turn the executor on.

  • A separate write role. The write role template, in CloudFormation and Terraform versions (default role name kaveo-remediation, see Remediation), has the same trust shape: kaveo's principal, an ExternalId and a one-hour session. Its inline policy lists only the actions the 7 shipped remediation operations need, including the reads they use to locate and check their target. A drift test fails the build if the template stops matching the actions the executor declares it needs.
  • Assumed only for an approved saga. The write role is used to execute a remediation a human has approved, never during scans.
  • Dry run by default. KAVEO_REMEDIATION_EXECUTOR defaults to offline. Real execution needs KAVEO_REMEDIATION_EXECUTOR=aws, set through a compose override, and a write role registered on the account.
  • Proposing and approving are split by role. Proposing a fix needs investigate (analyst). Approving, rejecting, executing, rolling back or opening a PR needs approve_remediation, which only admins have. Approval takes a typed reason.
  • Checked after the change. Each executed fix is verified. Most operations re-collect the resource and re-run the finding's own detector. The three containment operations (quarantine an instance, revoke role sessions, take forensic snapshots) have no detector to re-run, so they re-read the resource and check the expected end state.

The optional autonomous patrol only drafts sagas in the proposed state and never applies them. See Remediation and autonomous patrol.

Nothing leaves by default#

Out of the box, the stack makes no outbound calls beyond the cloud APIs it scans.

  • AI is offline. Compose defaults KAVEO_AI_PROVIDER to offline, a deterministic stub that calls no model. Hosted AI client libraries load lazily, only when a hosted provider is selected.
  • Telemetry stays local. KAVEO_TELEMETRY defaults to local, which writes usage rows to your own database; off records nothing. Billing defaults to off.
  • Everything else is opt-in. Hosted AI, the Slack brief, GitHub remediation PRs, Jira, webhooks and other alert channels, outbound email, Security Hub push, the MCP server and self-serve signup do nothing until you configure them.
  • Air-gapped. With the air-gapped bundle and the offline or self-hosted local AI provider, the only runtime egress is sts:AssumeRole plus read-only calls to the AWS account being scanned.

If you turn on a hosted AI provider, the fenced cloud-derived text in each prompt goes to that provider.

Evidence and grounded AI#

Findings come only from deterministic code. A check on every CI run fails the build if the AI layer imports detection or collection code. Each call a collector reports becomes an entry in the evidence ledger (one entry per run by default; per API call and page for collectors such as IAM), and the database rejects any edit or direct delete of a ledger entry.

The AI layer treats everything derived from your cloud as data, not instructions:

  • Input fencing. Cloud-derived text and your investigate question are wrapped in untrusted-data fences. A closing delimiter inside a value is neutralized, so content cannot end its own fence. The system prompt tells the model that fenced text is data.
  • Bounded input. A finding's details are truncated at 4000 characters before they enter a prompt.
  • Output citation gate. The gate keeps a claim only if it cites at least one observation id and every cited id resolves to a stored observation. Uncited or unresolved claims are dropped before anything is stored.

See Evidence and grounded AI for the full pipeline.

Tenant isolation#

  • Tenant tables use Postgres row-level security. A request scoped to an org sees only that org's rows.
  • In the stock compose stack, the api and worker connect as a non-superuser database role, so the database itself enforces RLS. Only the migration runner uses the superuser, through MIGRATE_DATABASE_URL.
  • The citation resolver checks ids against observations loaded under the caller's org. An id from another tenant does not resolve, so a claim that cites one is dropped.

Identity and credentials#

  • Local passwords are hashed with argon2id. Five failed logins for an email within 15 minutes lock it for 15 minutes.
  • Session, MCP (kv_mcp_) and API (kv_api_) tokens are stored only as SHA-256 hashes. MCP and API token plaintext is shown once, at creation.
  • An API token can do at most the intersection of its owner's role grants and its scopes. The default scopes are view_findings and run_scan. Tokens stop working when their owner is deactivated.
  • The kaveo CLI writes its config file with mode 0600.
  • An AWS account record holds only the role ARN and ExternalId, never access keys. Non-AWS accounts store an opaque credential reference such as env:NAME, and the secret itself stays in the environment, never in the database.
  • The Slack brief webhook, GitHub token, Jira API token and SCIM token are read from the environment, never the database.
  • Per-org alert channels are configured in the console, so a channel's destination settings (a webhook URL and its signing secret, or a Slack webhook URL) are stored in the database with that org's settings, under row-level security.
  • TOTP secrets are encrypted at rest with KAVEO_MFA_SECRET_KEY. Recovery codes are stored only as SHA-256 hashes.
  • Optional: TOTP MFA, OIDC, SAML, Google, Microsoft or GitHub SSO, and SCIM 2.0 provisioning and de-provisioning.
  • The audit log is append-only: the app's database role cannot update or delete its entries, and the database rejects any update or direct delete. An org's rows are removed only when the org itself is deleted.

Roles#

RoleActions
viewerview_findings
analystview_findings, run_scan, investigate, share_report, triage_findings
adminevery action, including approve_remediation, manage_accounts and manage_users

The role table lives in code and is evaluated in process. Platform superadmins, listed in KAVEO_SUPERADMIN_EMAILS or promoted at runtime by another superadmin, pass every role check and can act in any org, so keep that list short. A scoped token is limited to its scopes even when a superadmin mints it.

Hardening before exposure#

Set these before anyone else can reach the deployment. The stock compose file passes only a fixed list of variables to the api and worker, and KAVEO_RATE_LIMIT_PER_MINUTE, KAVEO_MFA_SECRET_KEY and KAVEO_SCIM_TOKEN are not on it. Add them to both services with a compose override file, as described in Configuration. The full checklist is in Deployment.

  • Required secrets. Set POSTGRES_PASSWORD, POSTGRES_APP_PASSWORD and KAVEO_BOOTSTRAP_PASSWORD before the first start, generating each one separately, for example with openssl rand -hex 24. Compose will not start without them.
  • TLS. Set KAVEO_SITE_ADDRESS to your hostname so Caddy obtains a certificate, and set KAVEO_PUBLIC_ORIGIN to the matching https:// origin.
  • Secure cookie. Compose defaults KAVEO_ENVIRONMENT to prod, and any value other than dev marks the session cookie Secure. Also set KAVEO_SESSION_COOKIE_SECURE=true on a TLS deployment: when it is unset, compose passes an empty value, which the api does not accept at startup. Use false only for an internal bring-up over plain HTTP.
  • Rate limit. Set KAVEO_RATE_LIMIT_PER_MINUTE to a positive value. The default, 0, is off. The limiter counts requests per client address as the api sees it. Behind the stock Caddy that is Caddy's address rather than each user's, so size the value as one budget shared by all users.
  • MFA key. Set KAVEO_MFA_SECRET_KEY to a strong random value before anyone enrolls. MFA enrollment is refused while it is empty.
  • Internal endpoints. The stock Caddyfile routes only /v1/* and /mcp to the api. Keep /health and /metrics on the compose network. If your IdP needs /scim/v2, route only that path and set KAVEO_SCIM_TOKEN.
  • Signup gates. Keep KAVEO_SIGNUP_ENABLED=false and KAVEO_APPROVAL_REQUIRED=true unless you need self-serve signup. Leave KAVEO_DEPLOYMENT_MODE at its default, on_prem: setting it to hosted also opens signup.