kaveo · Reference
FAQ
Answers about kaveo's access model, data residency, AI use, offline operation, supported clouds, remediation safety and how it relates to AWS-native tools.
On this page
- Access and data
- Does kaveo need write access to my cloud?
- Where does my data go?
- How are tenants kept apart?
- AI and evidence
- Does kaveo work without an AI model?
- Can the AI invent a finding?
- Which AI providers are supported?
- Coverage
- Which clouds are supported?
- How is this different from Security Hub or GuardDuty?
- Can it run air-gapped?
- Operations
- Will kaveo change my infrastructure on its own?
- How do I gate CI on kaveo?
- Can my own agents query kaveo?
- How do I reduce noise?
- Why does a setting in my .env have no effect?
- Getting kaveo
- How do I get kaveo?
Short answers to common questions about deploying and running kaveo. Each answer links to the page that covers it in full.
Access and data#
Does kaveo need write access to my cloud?#
No. Scanning assumes a read-only role, kaveo-readonly by default, which carries the SecurityAudit and ViewOnlyAccess managed policies. Its trust policy admits only kaveo's own principal, and only when the call presents your ExternalId. Write access exists only if you deploy the separate remediation role, register it on the account and set KAVEO_REMEDIATION_EXECUTOR=aws. Even then, kaveo assumes that role only to apply or roll back a remediation an admin has approved, never during a scan. See Security model.
Where does my data go?#
When you run kaveo yourself, into its own Postgres database, inside your perimeter. The stock compose file sets the AI provider to offline, telemetry writes usage rows to that local database only, and billing is off, so by default nothing else leaves. Hosted AI, the Slack brief, GitHub PRs, Jira, alert channels, outbound email, Security Hub push, the MCP server and self-serve signup each stay off until you configure them. If you configure a hosted AI provider, the AI stages send it prompts built from your findings and their evidence. See Security model.
How are tenants kept apart?#
With Postgres row-level security. In the stock compose file, the api and worker connect as the non-superuser application role, so the database itself enforces RLS, and only the migration runner uses the superuser. The citation resolver is scoped to your org too, so an AI claim that cites another tenant's observation does not resolve and is dropped. See Security model.
AI and evidence#
Does kaveo work without an AI model?#
Yes. Detection, attack paths, compliance mapping, the Risk Report's findings and counts, the patrol's triage and the MCP server are all deterministic. With KAVEO_AI_PROVIDER=offline, the five AI stages, including the priority and explanation added to the report's top findings, return deterministic stub output and no model is called. See Evidence and grounded AI.
Can the AI invent a finding?#
No. Findings come only from deterministic detectors or from imported native AWS findings. An import-linter contract fails CI if the AI layer imports the detector or collector code. Any AI claim whose citations are missing, or do not resolve to a stored observation, is dropped before it is saved. See Evidence and grounded AI.
Which AI providers are supported?#
Set KAVEO_AI_PROVIDER to one of these:
anthropicopenai, which covers any OpenAI-compatible API throughKAVEO_OPENAI_BASE_URLgrokbedrock, which uses the deployment's AWS credentialslocal, a self-hosted OpenAI-compatible endpoint such as vLLM or Ollamaofflineauto
auto picks the most cost-effective hosted provider whose key is set, and resolves to offline when no key is set. It is the application default when the variable is unset, but the stock compose file sets offline. You can override the model id for each tier with KAVEO_MODEL_FAST, KAVEO_MODEL_MAIN and KAVEO_MODEL_DEEP. See Configuration.
Coverage#
Which clouds are supported?#
AWS is covered in depth, with 44 collectors and 93 detectors. Azure, GCP, Kubernetes and GitHub are early: one collector and one detector each today. Each needs a pip extra (kaveo[azure], kaveo[gcp], kaveo[k8s] or kaveo[saas]) that the default image does not install, so you rebuild the api and worker image with it. If the extra is missing, the credential does not resolve or the provider API errors, the scan fails with the reason. kaveo does not substitute demo data for a real account; only KAVEO_SCAN_MODE=synthetic scans the labelled synthetic dataset. See Connect cloud accounts.
How is this different from Security Hub or GuardDuty?#
kaveo works alongside them. It imports your Security Hub, GuardDuty and Inspector findings, marks them native so you can tell them apart from its own, and runs them through the same suppression rules. On top of those, it adds its own deterministic detectors, an IAM and reachability graph with attack paths, and compliance mapping across 10 frameworks. It also keeps an evidence ledger behind every result. With the AWS executor enabled, it adds remediation that is checked by a targeted re-scan. See Detection and attack paths.
Can it run air-gapped?#
Yes. The air-gapped bundle ships every container image in one archive, with a SHA256 manifest over every file that you verify before you load it. With the provider set to offline, the only runtime egress is sts:AssumeRole and the read-only calls to the accounts you scan. For self-hosted AI, build the bundle with --with-model to include the optional model service, then set KAVEO_AI_PROVIDER=local so inference stays inside the stack. Model weights are never bundled, so you copy them in separately. See Deployment.
Operations#
Will kaveo change my infrastructure on its own?#
No. The optional autonomous patrol drafts up to KAVEO_PATROL_MAX_PROPOSALS fixes per scan (3 by default) as proposed sagas and never applies them. Approving, executing or rolling back a fix needs the approve_remediation permission, which only admins have, and approval requires a typed reason. GitHub remediation PRs go to a separate branch, and kaveo never merges them. See Remediation and autonomous patrol.
How do I gate CI on kaveo?#
Run the CLI with --wait and --fail-on. Point it at your deployment with KAVEO_API_URL and authenticate with an API token in KAVEO_API_TOKEN. kaveo auth token create --name ci mints one scoped to view_findings and run_scan by default.
kaveo scan start --account ACCOUNT_ID --wait --fail-on high
ACCOUNT_ID is kaveo's id for the account, from kaveo accounts list. Exit code 1 means the scan found at least one finding at or above the threshold. Exit code 4 means kaveo did not answer (a 5xx, a timeout or a refused connection), or the scan failed or did not finish within --timeout (900 seconds by default). Exit code 3 means the token is missing, rejected or lacks the scope. That way your pipeline can tell a failed gate from an outage. --fail-on requires --wait. See CLI and API reference.
Can my own agents query kaveo?#
Yes, through the read-only MCP server at POST /mcp. Set KAVEO_MCP_ENABLED=true, then create a token on the console's MCP settings page or through POST /v1/mcp/tokens. The token inherits your role and org. The 7 tools read scans, findings, attack paths, compliance status and evidence, and none of them can change anything. See MCP server and integrations.
How do I reduce noise?#
kaveo gives you several controls, from the report view down to individual rules:
- Relevance mode.
KAVEO_REPORT_RELEVANCE_MODEshapes the Risk Report.real_risk(the default) collapses hygiene findings unless the resource is internet-exposed or on a toxic path.strictnarrows the report further, andallshows everything. The mode changes only the report view;/v1/findingsand compliance are unaffected. - Suppression.
KAVEO_DETECT_SUPPRESS_RULESandKAVEO_DETECT_MIN_SEVERITYmark matching findings as suppressed. They are still stored and counted, but hidden by default, and the settings apply to new scans only. - Thresholds. Settings such as
KAVEO_DETECT_ROTATE_AFTER_DAYStune individual detectors. Each value is read once per process, so restart the worker after a change. - Waivers. A waiver marks a compliance control, or one rule on one resource, as accepted risk, with an approver, a justification and an expiry. See Compliance and reporting.
See Detection and attack paths.
Why does a setting in my .env have no effect?#
The stock compose file has no env_file, so it passes to the api and worker only the variables listed in its x-svc-env block. KAVEO_REPORT_RELEVANCE_MODE, KAVEO_PATROL_ENABLED, KAVEO_MCP_ENABLED and the Google, Microsoft and GitHub sign-in settings are on that list. KAVEO_DETECT_*, KAVEO_REMEDIATION_EXECUTOR, the OIDC, SAML and SCIM settings, the Slack, GitHub and Jira integration settings, and the xAI key for grok are not. Add those to both services in a compose override file. See Configuration.
Getting kaveo#
How do I get kaveo?#
Book a demo to talk to the ZHASK team. For a product-level overview, see the kaveo product page.