Vectasec · Concepts
Identities and attack paths
How Vectasec derives non-human identities, credentials and grants from every connector, observes runtime use, and chains findings into attack paths.
On this page
When the engine runs a scan, it derives an inventory of non-human identities (NHIs) from what the connectors collected, runs identity and credential rules over it, and correlates identities with exposure and data sensitivity into attack paths. This page covers each layer and where it stops.
The identity model#
Before any detector runs, the scan makes two derivation passes over the collected resources: data classification and NHI derivation. Neither makes another call to your system. If either pass fails, the scan continues without it and records the error in its stats.
NHI derivation produces three resource kinds:
Edges link them: credential_of (credential to identity), granted_to (grant to holder) and grants_access (grant to the resource it names, when the same scan collected that resource). Derivation is idempotent by ext_id, so a rescan restates the inventory instead of adding to it.
29 of the 39 connector types have an identity extractor. For example:
- Kafka: ACL binding principals, one grant per binding. Kafka exposes no credential material, and ACLs are readable only when the cluster has an authorizer.
- RabbitMQ: users, password metadata and per-vhost permissions. RabbitMQ serves users and permissions only to an administrator-tagged account. The connector still issues only GET requests.
- Kong: consumers, credential counts per auth plugin and ACL group memberships.
- Redis: ACL users, a password credential unless the user is
nopass, and the key and command scope as a grant. - PgBouncer: configured users, a password credential whose hashing follows
auth_typeunless it istrustorany, admin or stats console access as a grant, and a backend identity for each database pinned to aforce_userorauth_user. - AWS: principals in SQS, SNS and EventBridge resource policies.
- Kubernetes: ServiceAccounts, their long-lived token secrets and RBAC role bindings.
Discovery, Flink, GraphQL, gRPC, integration-as-code, Memcached, OpenAPI, Redpanda, Schema Registry and Solr contribute no identities today.
A grant is marked wildcard when it covers everything in scope, for example a Kafka binding on * or ALL, a RabbitMQ permission of .*, Redis ~* or +@all, or a Kubernetes binding to cluster-admin.
View the inventory#
The console's Identities page lists every NHI across your systems, with wildcard, external and default flags. Each identity opens a drawer with its credential metadata, grants, open findings and a Blast radius view inferred from grant scope. From the CLI:
From the engine directory of the release:
uv run vectasec identities
It prints identity, credential, grant and wildcard-grant totals, then one line per identity. Pass --tenant to read a tenant other than the demo tenant.
Identity and credential rules#
These two packs run on the derived kinds, so one rule covers every system. Their findings attach to the derived resources, so they stay separate from a connector's own rules, which can describe the same issue in that system's terms.
A not-assessable finding is a coverage gap, never a pass. Findings, evidence and compliance explains the three outcomes.
Runtime evidence#
Configuration shows what a system permits. Five connectors also read, without an agent, who is connected when the scan runs:
Each live client becomes a runtime_client resource, and five rules read it:
RUNTIME.CREDENTIAL.CLEARTEXT_ON_WIRE(critical): a client authenticated withPLAIN,AMQPLAIN,LOGINorCRAM-MD5outside TLS.RUNTIME.IDENTITY.SHARED_DEFAULT_IN_USE(high): a default account such asguest,adminorpostgreshas live connections.RUNTIME.TRANSPORT.PLAINTEXT_IN_USE(high): a live session is not TLS, whatever its auth mechanism.RUNTIME.IDENTITY.UNATTRIBUTED_CLIENT(medium): the server cannot name the client's principal.RUNTIME.ACCESS.UNEXPECTED_PEER(medium): a client connects from a public IP address.
Where a source doesn't report transport, RUNTIME.TRANSPORT.PLAINTEXT_IN_USE files a not-assessable finding instead of guessing, and so does RUNTIME.CREDENTIAL.CLEARTEXT_ON_WIRE when it sees a cleartext mechanism. Redis CLIENT LIST never reports transport, and Redis ACL LOG entries, which record refused logins and commands, are kept as separate access-denial resources rather than live clients. Kafka's consumer-group API shows consumers only, not producers, and doesn't report transport. Its client.id is self-asserted, so Vectasec never treats it as a principal. Evidence is sampled at scan time, and there is no in-cluster sensor.
Never-observed identities#
After each scan, Vectasec adds the observed clients to a running evidence record for the connection and compares them with the identities the connection declares. RUNTIME.GRANT.NEVER_OBSERVED names an unused identity only when all of these hold:
- The runtime source names authenticated principals, which Kafka's doesn't.
- At least one principal has been observed on the connection.
- At least 14 days have passed since observation began on that connection.
- At least 20 completed scans ran in that window.
When it fires, each finding is medium and states the window and scan count, because a job that runs less often than the window can look unused. Until then, the rule files one info-severity, not-assessable finding per connection that says why no claim is made yet, such as a window that is still too short. On Kafka that gap never closes, and the finding says so. Connectors with no runtime source stay silent.
To fill the window, give the connection a rescan cadence on the console's Connections page, or set scan_interval_minutes with PATCH /connections/{id}. Scheduled rescans need pg_cron, as Configuration describes. Both thresholds apply, so a 24-hour cadence needs about 20 days to reach 20 scans, while a 6-hour cadence meets both in 14 days. The console's Runtime page shows which connections have a runtime source and how long the observation window has run.
Data classification#
Vectasec never reads a message, row or payload. It classifies data-carrying resources by the whole words in their name, namespace and label values. A match means the resource is named like that data, not that it contains it.
The classifier accepts topics, queues, streams, exchanges, subjects, routes, services, MCP tools, keyspaces and channels. Today the connectors that emit those kinds are Kafka (topics), Kong (routes and services) and MCP (tools). Connector-specific kinds, such as AWS SQS queues or Pulsar topics, aren't classified yet.
DATA.SENSITIVE.INVENTORY (low) lists resources named like PCI, PHI or secrets data. DATA.EXPOSURE.SENSITIVE_PLAINTEXT applies to Kafka only. It flags a classified topic whose broker has a plaintext client listener or no authorizer. The finding is high, or critical when the topic is named like PCI, PHI or secrets data and the broker has both.
Attack paths#
The attack-path engine correlates three signals across the resources of a single scan:
- Exposure: a plaintext client listener, a broker with no authorizer, or a surface that is unauthenticated or publicly reachable.
- Over-privilege: an identity that holds a wildcard allow grant. If the identity is also a shared or default account, the hop says so.
- Sensitivity: a resource with any data class from the table above.
Each path is an ordered list of hops, and each hop cites a resource ext_id and the signal read there.
A sensitive resource gets at most one data-anchored path, because the three-factor rule replaces the two-factor one. If a surface exists but its exposure can't be read, ATTACKPATH.SENSITIVE_ON_EXPOSED files a not-assessable finding. On the quick-start Kafka broker, which has plaintext listeners and no authorizer, the payments topic gets this observation (wrapped here):
toxic combination (sensitive data on an exposed surface):
[broker (broker-1): exposed — client listener INTERNAL maps to PLAINTEXT; authorizer.class.name is unset (no ACL enforcement)]
-> [topic 'payments' (topic-payments): named like pci data (matched pci:payments)]
Keep in mind what a path claims:
- It is a correlation over the scan's current graph, not proof of exploitation. "Reachable" means an authorization or hosting relationship that the scan recorded, not a proven network path.
- Paths stay inside one scan. They don't trace data lineage across systems.
- When a hop can't be tied to a specific surface, the engine uses the most exposed surface in the same scan.
- Attack-path findings map to no compliance control. The findings underneath them, such as
KAFKA.LISTENER.PLAINTEXT, carry the mappings.
The console's Attack paths page draws each path as a chain of hops.
Blast radius#
GET /graph/blast-radius/{resource_id} walks the stored graph from one resource. It returns everything reachable, the shortest path to each resource, and each resource's open findings by severity:
curl -s "http://127.0.0.1:8000/graph/blast-radius/$RESOURCE_ID?depth=4" \
-H "Authorization: Bearer $VECTASEC_ENGINE_TOKEN" \
-H "X-VectaSec-Tenant: $TENANT_ID"
resource_idis the resource's UUID, not itsext_id.depthaccepts 1 to 10 and defaults to 4. The walk never goes past 8 hops, so a larger value acts as 8.- The walk is undirected by default. Pass
directed=trueto follow edges from source to destination only. - A per-path cycle guard keeps the root from showing up in its own blast radius.
depth_bounded: truemeans the walk reached resources at the depth limit, so more may lie beyond it. Read the result as a minimum.- Resolved and suppressed findings aren't counted.
Next steps#
- Findings, evidence and compliance: the three outcomes, the evidence ledger and how controls are evaluated.
- Connector reference: what each connector collects and the rules it feeds.
- CLI and API reference: every command and route, including how to authenticate API calls.