kaveo · Get started
Deployment
Run kaveo in production or fully air-gapped: TLS, required secrets, base credentials, backups, health checks and the settings to confirm before you expose it.
On this page
Topology#
kaveo runs as one Docker Compose project, defined in infra/docker-compose.yml. Five services, plus an optional model service, share the bridge network kaveo-net, and only Caddy publishes ports to the host.
The api and worker are built from the same Dockerfile, so detectors and collectors are identical at request time and scan time. The image is based on python:3.12-slim, runs as the non-root user kaveo (uid 10001), and ships postgresql-client-16 for pg_dump. The worker starts after the api passes its healthcheck, which happens once the api's boot sequence, migrations included, has finished. Scans are claimed with SELECT ... FOR UPDATE SKIP LOCKED, so you can run more than one worker with --scale worker=2. The named volumes are pgdata, caddy_data, caddy_config, kaveo_backups and ollama_models.
Caddy sends /v1/* and /mcp to the api and everything else to the console. It adds HSTS, X-Frame-Options: DENY, nosniff, a referrer policy and a content security policy to every response. It does not send /health, /metrics or /scim/v2/* to the api. Those paths fall through to the console, so the api's versions of them stay inside the compose network. If your identity provider provisions users over SCIM, add a handle /scim/v2/* block that proxies to api:8000, and put it above the catch-all. Also pass KAVEO_SCIM_TOKEN, the bearer token your identity provider presents, to the api in an override file.
cp .env.example .env
openssl rand -hex 24 # once for each required secret
docker compose --env-file .env -f infra/docker-compose.yml up -d --build
Production checklist#
Work through this list before anyone else can reach the console.
- Set the required secrets. Compose refuses to start until
POSTGRES_PASSWORD,POSTGRES_APP_PASSWORDandKAVEO_BOOTSTRAP_PASSWORDare set. Generate each one withopenssl rand -hex 24, and setKAVEO_BOOTSTRAP_EMAILto an address you control. The bootstrap owner is an admin account, so treat its password like any other admin credential. Set these before the first start. Postgres applies its passwords only when it initialises an emptypgdatavolume, and the owner is created only while the users table is empty. - Turn on TLS. Set
KAVEO_SITE_ADDRESSto your hostname, for examplekaveo.example.com. Caddy then obtains and renews a certificate automatically. For that to work, the name must resolve to this host and ports 80 and 443 must be reachable. SetKAVEO_PUBLIC_ORIGINto the browser-facing origin, such ashttps://kaveo.example.com, because the Google, Microsoft and GitHub sign-in callback URLs are built from it. - Keep the session cookie Secure. Compose sets
KAVEO_ENVIRONMENT=prod. Any value other thandevmarks the cookieSecure. SetKAVEO_SESSION_COOKIE_SECURE=falseonly for an internal bring-up over plain HTTP. - Keep sign-up closed. Leave
KAVEO_SIGNUP_ENABLED=false(the default) unless you want self-service sign-up. If you turn it on, keepKAVEO_APPROVAL_REQUIRED=true(also the default), so new accounts stay pending until a platform super admin approves them. Super admins are the addresses listed inKAVEO_SUPERADMIN_EMAILS. - Rate-limit the API.
KAVEO_RATE_LIMIT_PER_MINUTEis 0 (off) by default. If the console is reachable beyond a trusted network, set a positive value. Requests over the budget get a429. The limiter counts requests per client address, as the api sees it. Behind the stock Caddy, every request arrives from Caddy's address, so the value is one budget shared by all users on each api replica. Size it for that, or set uvicorn'sFORWARDED_ALLOW_IPSon the api to Caddy's address or the compose network's subnet, so the api trusts theX-Forwarded-Forheader that Caddy sets. - Set an MFA key. Set
KAVEO_MFA_SECRET_KEYto a strong random value before anyone enrolls TOTP. It encrypts TOTP secrets at rest, and enrollment refuses to start without it. Keep it stable afterwards, because secrets already stored can only be decrypted with the same key. - Keep
/healthand/metricsinternal. Both are open by design, for probes and scrapers. Do not add Caddy routes for them, and scrape/metricsfrom inside the compose network. - Ship structured logs. If you forward logs to a collector, set
KAVEO_LOG_FORMAT=json. - Pass other settings with an override file. The compose file has no
env_fileand passes only a fixed list of variables to the api and worker. The settings in items 5, 6 and 8, includingFORWARDED_ALLOW_IPS, fall outside that list, as doesKAVEO_SCIM_TOKEN. See Configuration for which settings need an override.
For example, save this as infra/docker-compose.override.yml:
services:
api:
environment:
KAVEO_MFA_SECRET_KEY: ${KAVEO_MFA_SECRET_KEY:?set KAVEO_MFA_SECRET_KEY in .env}
KAVEO_RATE_LIMIT_PER_MINUTE: ${KAVEO_RATE_LIMIT_PER_MINUTE:-0}
KAVEO_LOG_FORMAT: json
worker:
environment:
KAVEO_LOG_FORMAT: json
When you use -f, Compose does not load override files on its own, so name both files:
docker compose --env-file .env \
-f infra/docker-compose.yml -f infra/docker-compose.override.yml up -d
Base AWS credentials#
The worker assumes each account's read-only role from kaveo's own AWS identity. You can provide that identity in three ways:
- An instance or task role (preferred). On EC2, attach an instance profile; if you run the containers as ECS tasks, give them a task role. Leave
AWS_ACCESS_KEY_ID,AWS_SECRET_ACCESS_KEYandAWS_SESSION_TOKENempty. On EC2 with IMDSv2, set the metadata hop limit to 2 so that containers on the bridge network can reach it. - Static keys. Set
AWS_ACCESS_KEY_IDandAWS_SECRET_ACCESS_KEYin.env, plusAWS_SESSION_TOKENfor temporary credentials. Compose passes them to the api and worker. - A named profile. Compose does not pass
AWS_PROFILE. Mount~/.awsread-only at/home/kaveo/.awsin the api and worker, and setAWS_PROFILEin an override file. The files must be readable by uid 10001.
Check which identity the worker resolves:
docker compose --env-file .env -f infra/docker-compose.yml exec worker \
python -c "import boto3; print(boto3.client('sts').get_caller_identity()['Arn'])"
That identity is the principal that each read-only role must trust (KaveoPrincipalArn in the templates). For an IAM user, use the ARN as printed. For an instance or task role, STS prints a session ARN of the form arn:aws:sts::<account>:assumed-role/<role>/<session>. The session part changes with each instance or task, so trust the role itself: arn:aws:iam::<account>:role/<role>. Set that value as KAVEO_PRINCIPAL_ARN so onboarding pre-fills it. In production, also set KAVEO_SCAN_MODE=aws. Under the default auto, a scan without base credentials falls back to the labeled synthetic dataset. Under aws, it fails instead. See Connect cloud accounts.
Backups and migrations#
The api applies migrations on every boot and holds a Postgres advisory lock while it does, so concurrent starts never apply a file twice. Migrations are 45 forward-only SQL migrations shipped with the release, applied in order. Each runs in its own transaction and is recorded as applied.
Compose sets KAVEO_BACKUP_BEFORE_MIGRATE=true, so when migrations are pending, the api first takes a pg_dump snapshot while holding the lock. Snapshots go to the kaveo_backups volume at /var/lib/kaveo/backups, with names like kaveo-2026-09-30T14-05-00Z.dump. They are in custom format, so you restore them with pg_restore. There are no down migrations, which makes this snapshot your rollback point for an upgrade.
For a snapshot on demand, run the backup command from the api image using the migration connection (MIGRATE_DATABASE_URL), as documented in the release README and AIRGAP.md. The app role is scoped by row-level security, and pg_dump needs a role that can read every row, so make backup, which connects as the app role, is not enough.
The volume lives on the same host as the database, so copy archives off the host on a schedule.
Health and monitoring#
The api serves two unauthenticated operational endpoints on port 8000, reachable from inside the compose network:
GET /healthreturnsstatus(okordegraded),version,ai_provider,database(up,downorunknown), and how many detectors and collectors loaded. It returns 200 even when the database is down, so alert on the body. If the database was unreachable while the api booted,databasereadsunknownandstatusstill readsok, so alert wheneverdatabaseis notup. The compose healthcheck uses this endpoint and checks only the status code.GET /metricsreturns Prometheus text, includingkaveo_detectors_loaded,kaveo_collectors_loadedandkaveo_database_up. After a worker checks in, it also includeskaveo_worker_seconds_since_heartbeat, the age of the most recent check-in from any worker. Workers check in between jobs, so the value climbs during a long scan. Alert when it exceeds your longest expected scan time. By default, the console marks workers as down after 300 seconds.
docker compose --env-file .env -f infra/docker-compose.yml exec api \
python -c "import urllib.request; print(urllib.request.urlopen('http://localhost:8000/health').read().decode())"
API responses carry an X-Request-ID header. If the client sent one, the api reuses it. Log lines include the same id, so use it to match a client error to its log line.
Air-gapped install#
On a connected machine, run the release's air-gap bundle script. It writes dist/kaveo-airgap-<version>.tar.gz and a .sha256 file. The bundle contains:
images.tar, adocker saveof every image in the stack- the compose file, Caddyfile, database init scripts and migrations
.env.example,AIRGAP.md,LICENSE,manifest.jsonandSHA256SUMS
On the target host:
shasum -a 256 -c kaveo-airgap-<version>.tar.gz.sha256
mkdir kaveo-airgap
tar -xzf kaveo-airgap-<version>.tar.gz -C kaveo-airgap && cd kaveo-airgap
shasum -a 256 -c SHA256SUMS
docker load -i images.tar
cp .env.example .env
If a checksum fails, the bundle changed in transit, so stop. In .env, set KAVEO_AI_PROVIDER=offline, KAVEO_SCAN_MODE=aws (or synthetic for a demo without AWS), the required secrets and KAVEO_BOOTSTRAP_EMAIL. Then start the stack. Migrations apply on boot.
docker compose --env-file .env -f infra/docker-compose.yml up -d
With the offline provider and no outbound integrations configured, such as mail, alert channels, Jira, GitHub pull requests or SSO, the only runtime egress is sts:AssumeRole and read-only API calls to the accounts you scan. Automatic TLS needs a public certificate authority, which an isolated network cannot reach. Keep the default :80 behind your own TLS terminator, or add a tls directive with your own certificate to the Caddyfile and mount the certificate and key into the caddy container.
Optional self-hosted model#
To include the Ollama image, build the bundle with --with-model. Model weights are never bundled. On a connected machine, pull them into the model container (for example, ollama pull llama3.1). Then copy the ollama_models volume (named kaveo_ollama_models on the Docker host), which is mounted at /root/.ollama, to the target host. In .env, set the following. The base URL is already the compose default.
KAVEO_AI_PROVIDER=local
KAVEO_LOCAL_BASE_URL=http://model:8000/v1
Set KAVEO_MODEL_FAST, KAVEO_MODEL_MAIN and KAVEO_MODEL_DEEP to model names that your server serves, then start with the profile:
docker compose --env-file .env -f infra/docker-compose.yml --profile local-model up -d
If the model endpoint is unreachable, the local provider falls back to the offline stub instead of failing. See Evidence and grounded AI for how the AI layer uses evidence.