kaveo · Guides
Remediation and autonomous patrol
Approve AI-drafted fixes through audited sagas, apply them with a scoped write role, prove them with a re-scan, open GitHub PRs, and let patrol draft fixes.
On this page
kaveo does not change your cloud on its own. Every fix is a remediation saga: a drafted artifact inside a lifecycle that an admin must approve. By default the saga runs as a dry run, and kaveo holds no write credentials.
The saga lifecycle#
Every state change appends an observation to the evidence ledger. Its hash covers the actor, the transition and the time, plus the typed reason for approvals and rejections. The saga itself keeps the approver and the reason. A finding has at most one saga in flight. If you propose again, you get the active saga back.
After each successful scan, kaveo checks the account's recorded sagas. If the same rule fires again on the same resource, the saga is flagged as a regression. The flag appears in the post-scan brief and in the Remediations queue.
Propose, review, approve#
- Propose. On a finding's Remediation tab, choose Propose for approval. You can also call
POST /v1/findings/{finding_id}/remediations/propose. kaveo drafts the artifact with the grounded remediate stage, from the finding and the observations behind it. This needsinvestigate, which analysts and admins have. - Review. The Remediations page (
/remediations) shows each saga with its finding, severity, rule and artifact.GET /v1/remediationsreturns the same queue and accepts astatefilter. - Decide. In the queue, choose Approve or Reject and type a reason. Or call
POST /v1/remediations/{saga_id}/approveor/rejectwith{"reason": "..."}. A blank reason is refused. Both needapprove_remediation, which only admins have, so an analyst can propose a fix but cannot approve it. - Execute. The worker claims approved sagas itself.
POST /v1/remediations/{saga_id}/executeonly confirms the approval. It returns202withexecution_pending, or409if the saga has no approval. - Roll back. Choose Roll back in the queue, or call
POST /v1/remediations/{saga_id}/rollback. This restores arecordedorfailedsaga from the snapshot taken before execution. Under theofflineexecutor, the rollback is recorded and nothing changes. This needsapprove_remediation.
GET /v1/remediations/operations lists the rules that have a typed operation, and the console offers the propose flow only for those. For any finding, including rules without an operation, POST /v1/findings/{finding_id}/remediate/saga drafts and stores a Terraform, IAM policy or CLI artifact that you apply yourself. GET /v1/findings/{finding_id}/remediations lists the stored artifacts.
Dry run and real apply#
KAVEO_REMEDIATION_EXECUTOR has two modes:
offline(default). A deterministic dry run. It checks that the artifact is well formed and records the result. Nothing in your cloud changes.aws. kaveo assumes a separate write role for that one saga and applies the typed operation for the finding's rule. It applies that operation, not the artifact text, which is for review and pull requests.
Under offline, a saga still moves through executing and verifying to recorded. There, recorded means the dry run passed, not that anything changed. The finding stays open, so the next scan flags that saga as a regression.
To turn on real apply:
-
In the target account, deploy the write role next to the read-only scan role:
bashaws cloudformation deploy \ --template-file infra/aws/kaveo-write-role.cfn.yaml \ --stack-name kaveo-remediation \ --capabilities CAPABILITY_NAMED_IAM \ --parameter-overrides \ KaveoPrincipalArn=<kaveo-principal-arn> \ ExternalId=<external-id>The Terraform version is
kaveo-write-role.tf. The role trusts only kaveo's principal with an ExternalId, and sessions last at most one hour. Its inline policy lists only the actions the operations below need. A drift test keeps that list in line with the executor. -
Register the role with
POST /v1/accounts, which needsmanage_accounts. The console has no field for it yet. Resend the account'sname,aws_account_id,role_arnandexternal_id, and addwrite_role_arnandwrite_external_id. The call updates the account in place, so include every field. If you later reconnect the same account from the console, the write role is cleared, so register it again. -
Set
KAVEO_REMEDIATION_EXECUTOR=awson the api and the worker. You need a compose override for this; see Configuration.
Before execution, the worker saves a snapshot of the resource's stored configuration. If the rule has no operation or the account has no write role, the saga fails and nothing changes. Verification runs under the read-only scan role. For the four operations that fix a detector finding, kaveo re-collects the resource and re-runs that finding's own detector. The saga is recorded only when the finding no longer fires. The three incident-response operations check the resource's end state instead. The result is written to the ledger as a remediation_verify:<rule> observation.
Supported automatic operations#
Block Public Access clears the S3 finding only when disabled Block Public Access caused it. If a public ACL grant or a public bucket policy caused it, verification fails and the saga lands in failed. Fix the grant or the policy yourself.
Quarantine blocks inbound traffic only. The quarantine group keeps AWS's default outbound rule. It does not remove the public IP, so exposure.public_ec2_instance keeps firing, and the next scan flags a recorded quarantine saga as a regression.
response.forensic_snapshot and identity.role_active_sessions are response actions, not detector rules. Nothing in kaveo creates findings with these rule ids today, so no scan or console flow reaches them. Treat them as preview. Rolling back a forensic snapshot saga deletes the snapshots.
Remediation as a pull request#
- Set
KAVEO_GITHUB_REPO(owner/name) andKAVEO_GITHUB_TOKENon the api through a compose override. Use a fine-grained token with onlycontents:writeandpull_requests:write, limited to that one repository. kaveo callsapi.github.com. - Choose Open PR in the queue, which appears once both are set, or call
POST /v1/remediations/{saga_id}/pr. This needsapprove_remediation. Without both settings the call returns409. - kaveo commits
kaveo/remediations/<rule>-<id>.<ext>to the branchkaveo/remediation-<id>and opens a pull request against the default branch. In the file name, the rule's dots become dashes. The id is the first 12 hex characters of the saga id. The extension istf,jsonorsh, by artifact format.
The call is idempotent: a repeat returns the pull request that is already open. kaveo never pushes to the default branch and never merges.
Autonomous patrol#
Set KAVEO_PATROL_ENABLED=true. The stock compose file already passes it through.
After each successful scan, patrol takes the findings that are new since the previous successful scan. It drops the built-in hygiene rules, including dormant access keys and key rotation. It then ranks the rest by severity, with a bonus for exposure. rules and a larger one for exposure.toxic_attack_path. For the top KAVEO_PATROL_MAX_PROPOSALS findings (default 3), it drafts a fix and opens a saga in proposed.
Each draft is keyed on the scan and the finding, so a re-run never duplicates it. Nothing is applied until an admin approves. GET /v1/scans/{scan_id}/patrol returns the run.
Patrol can draft for rules that have no typed operation. With the aws executor on, check GET /v1/remediations/operations before you approve. For any other rule, open a pull request or apply the fix yourself.
Scheduled scans and the post-scan brief#
- Schedule. Call
PUT /v1/accounts/{account_id}/schedulewith{"interval_hours": 24}. The interval can be 1 to 720 hours, andnullturns scheduling off. The first scheduled scan runs one interval after you set it. This call needsmanage_accounts. - Brief. After each successful scan, kaveo lists findings that are new or resolved since the previous successful scan, matched on rule and resource ARN, plus any regressions. Read the brief at
GET /v1/scans/{scan_id}/briefand the full comparison atGET /v1/scans/{scan_id}/diff. - Slack. Set
KAVEO_SLACK_WEBHOOK_URLthrough a compose override, and the worker posts each brief. When patrol drafts fixes, it also posts the patrol summary.
Least-privilege drafts#
GET /v1/scans/{scan_id}/least-privilege proposes a narrower policy for each IAM user and role that has at least one granted service it has not used in 90 days. It skips AWS service-linked roles. It uses the Access Advisor data the scan already collected. A draft drops the unused services. It keeps service:* for the services that have been used, except EC2, IAM, Lambda and S3. For those four, it lists the exact actions that were used, when Access Advisor reports them. kaveo does not store or apply these drafts. They are proposals for you to review.