kaveo · Concepts
Detection and attack paths
kaveo's 97 deterministic detectors, 48 read-only collectors, native AWS findings, the reachability graph, toxic paths, choke points, issues and tuning.
On this page
Everything on this page is deterministic code, with no model involved. Collectors read your accounts with read-only access, detectors turn what they collected into findings, and graph queries turn resource relationships into attack paths. The AI layer works on these results afterwards and never creates a finding. See Evidence and grounded AI.
Collectors#
kaveo ships 48 collectors: 44 for AWS, plus one each for Azure storage accounts, GCP Cloud Storage buckets, Kubernetes RBAC and GitHub organizations. A collector calls only read APIs. What it reads is recorded in the evidence ledger: one row per run by default, or one per API call and pagination page for collectors that report per-call provenance, such as IAM. A finding on a collected resource cites the observation behind that resource.
The AWS collectors cover:
- Identity: IAM, the credential report, the password policy, Access Advisor, Organizations SCPs, Cognito
- Compute and containers: EC2 instances, AMIs, EBS snapshots, instance user data (checked for secret patterns, plaintext dropped), Lambda, ECS, ECR, EKS
- Data: S3, RDS, DynamoDB, Redshift, ElastiCache, EFS, OpenSearch, SageMaker and Bedrock
- Network and edge: VPC topology, security groups, VPC Flow Logs coverage, load balancers, API Gateway, CloudFront, Route 53 zones and records, ACM
- Secrets and keys: Secrets Manager and SSM Parameter Store (metadata only, never values), KMS
- Messaging: SNS, SQS
- Logging and monitoring: CloudTrail trails and events, CloudWatch alarms, AWS Config, and which security services are enabled in each region
- Native findings: GuardDuty, Inspector and Security Hub
The Azure, GCP, Kubernetes and GitHub collectors each need an optional Python package that the default image does not include. See Connect cloud accounts.
Detectors#
A detector is a pure function of one scan's collected resources and edges. It makes no network calls and has no randomness, so the same facts always produce the same findings, each with source=deterministic. Only detectors for the scanned account's provider run.
kaveo ships 97 detectors:
What each rule declares#
- A stable rule id, for example
exposure.s3_public_bucket. - A default severity:
critical,high,medium,loworinfo. - A title and compliance control mappings.
remediation_effort, from 1 (trivial) to 5 (major project).- MITRE ATT&CK technique ids.
confidence:high,mediumorlow. Severity is how bad a finding is if real. Confidence is how likely it is to be real.noise_notes: known false-positive patterns and the compensating controls that justify muting the rule.GET /v1/findings/{finding_id}returns them, so you can see why a rule fires before you suppress it.
Example rules#
How "public" is decided#
Every rule that reads a resource-policy or role trust-policy document shares one definition. That covers S3 bucket, SNS, SQS, Lambda, ECR, Secrets Manager, KMS and OpenSearch policies, and role trust policies. A statement is public when it allows access, its principal is a wildcard ("*" or {"AWS": "*"}) or it uses NotPrincipal, and no condition narrows who may use it.
The narrowing keys are aws:PrincipalOrgID, aws:PrincipalOrgPaths, aws:PrincipalArn, aws:PrincipalAccount, aws:SourceAccount, aws:SourceArn, aws:SourceOwner, aws:SourceVpc, aws:SourceVpce, sts:ExternalId, kms:CallerAccount, and aws:SourceIp when no listed CIDR is 0.0.0.0/0 or ::/0. Keys such as aws:SecureTransport and aws:MultiFactorAuthPresent limit how a call is made, not who makes it, so a wildcard trust guarded only by MFA is still public.
exposure.s3_public_bucket works differently. It flags a bucket when an ACL grants access to the AllUsers or AuthenticatedUsers group, when every Block Public Access setting is off, or when AWS's own GetBucketPolicyStatus reports the bucket policy as public.
identity.admin_policy_attached and identity.access_advisor_unused_services skip AWS service-linked roles, because you cannot change their grants.
Native AWS findings#
GuardDuty, Inspector and Security Hub findings collected during a scan become kaveo findings with source=native. They go through the same suppression pass, and each one cites the observation of the call that collected it. GET /v1/findings?source=native lists only these.
Known exploited vulnerabilities#
vuln.known_exploited_vulnerability raises a critical finding for any Inspector CVE on CISA's Known Exploited Vulnerabilities (KEV) catalog, whatever severity Inspector gave it.
kaveo never fetches the catalog. It ships a curated snapshot of 30 widely exploited CVEs. For full coverage, download CISA's known_exploited_vulnerabilities.json yourself, mount it into the api and worker containers and set KAVEO_KEV_CATALOG_PATH to its path. The file extends the snapshot and wins where both list a CVE. kaveo reads it once at startup, so restart after replacing it. An unreadable file falls back to the snapshot.
Graph and attack paths#
Each scan stores a graph of resource relationships. An internet node connects to each internet-facing resource. The other relationships cover a principal assuming a role, a role or policy attached to a principal, access to a resource, privilege escalation, passing a role to a service, group membership, and trust.
kaveo starts at every internet-facing resource, walks those relationships forward for up to 8 hops without revisiting a node, and keeps paths that end at a crown jewel. A path is toxic when it gets there through at least one hop other than a trust relationship. Toxic paths are listed first. Enumeration stops at KAVEO_ATTACK_PATH_MAX_ROWS rows (default 50000), and kaveo logs when it cuts a result short.
A crown jewel is:
- an RDS instance or cluster, S3 bucket or DynamoDB table
- an IAM user, role or managed policy that carries admin (
AdministratorAccess, or an allow on action*) - any resource with the tag key
crown-jewel,crown_jewelorsensitive, the tag valuecrown-jewelorsensitive, or a data-classification tag such asdata-classification: pci
A name alone never makes a crown jewel.
A choke point is a node inside a path, not its entry or target. kaveo ranks choke points by how many paths pass through them. One path is medium, 2 is high, and 3 or more is critical. Fixing the top choke point cuts every path through it.
Two detectors turn graph facts into findings:
exposure.toxic_attack_path(critical) emits one finding per internet entry and crown jewel pair, with the ordered hops. It starts suppressed. After the scan, the worker simulates each permission hop with the read-onlyiam:SimulatePrincipalPolicycall and records each result as an observation. Two kinds of hop are simulated: assuming a role (assts:AssumeRole) and passing a role to a service (asiam:PassRole). Structural hops (a role or policy attached to a principal, group membership) are accepted from the stored relationship without a call. Only a chain where every simulated hop is allowed is shown and markedvalidated. If a hop is denied, cannot be resolved, or the worker cannot assume the role again, the chain stays hidden. A synthetic scan has no account to simulate against, so it can confirm only chains with no simulated hops.exposure.toxic_combinationscores each internet-facing resource on public exposure, privilege reach to an admin-level crown jewel, and reach to a data-store crown jewel. Two factors is high, all three is critical. It is not suppressed.
Issues and triage#
After each successful scan, kaveo groups the non-suppressed findings by resource ARN into issues. An issue takes the worst severity among its findings. Its status (open, investigating, resolved or accepted_risk), assignee and note carry over to later scans, while its severity and member findings refresh. Findings with no resource are not grouped. This step is best-effort and never fails a scan. POST /v1/scans/{scan_id}/issues runs it again for one scan and is safe to repeat.
List issues with GET /v1/issues, filtered by account_id, status or severity. Triage one with PATCH /v1/issues/{issue_id}, which needs the analyst or admin role.
Other deterministic views#
These views are built on request from what a scan stored. They make no extra cloud calls and store nothing.
- Identity inventory (
GET /v1/scans/{scan_id}/identities): IAM roles and users that hold active access keys, ranked by a risk score so powerful, unused identities come first. Service-linked roles are listed but not scored. - Least-privilege drafts (
GET /v1/scans/{scan_id}/least-privilege): what each principal was granted minus what Access Advisor shows it used in the last 90 days, at service-wildcard granularity (s3:*). For EC2, IAM, Lambda and S3, where Access Advisor tracks individual actions, the draft lists the actions used. Service-linked roles get no draft. Drafts are proposals and are never applied. - Data inventory (
GET /v1/scans/{scan_id}/data-inventory): data stores, likely-sensitive and unprotected first. Sensitivity comes from tags (declared) and names (heuristic). Object contents are never read.
Tuning and noise#
None of these settings deletes a finding.
Suppression. KAVEO_DETECT_SUPPRESS_RULES takes comma-separated rule ids. KAVEO_DETECT_MIN_SEVERITY suppresses anything below that level. Suppressed findings are stored with data.suppressed_reason naming the setting responsible. They are hidden by default behind a visible count, and include_suppressed=true returns them. They do not count against compliance controls and are not grouped into issues. Suppression applies when a scan runs, so a change affects new scans only.
Thresholds. Each value is read once when the detector code loads. A value that is not an integer falls back to the default.
Risk Report relevance. KAVEO_REPORT_RELEVANCE_MODE shapes the Risk Report only. It never changes /v1/findings, the findings export or compliance.
The built-in hygiene list has 10 rules, such as identity.dormant_users and logging.cloudtrail_not_multi_region. KAVEO_REPORT_HYGIENE_RULES replaces it.