Vectasec · Concepts
How Vectasec works
The Vectasec architecture: Postgres control plane, engine, worker, console and collector, and the scan pipeline from read-only collection to cited findings.
On this page
Vectasec reads your integration layer through read-only connectors and routes none of its traffic. This page follows one scan from request to committed result.
The four pieces#
- Console. A Next.js 16 app that reads Supabase with your session, so row-level security decides what you see. Product changes go to the engine API, never straight to the database.
- Engine. The product: the Python engine, holding the connectors, detectors, scan loop, queue worker and API. Connectors use list, get and describe calls and never write to the systems they scan; the few extra probes some connectors support, such as the MCP connector's
active_probes, run only when you turn them on. The engine's Postgres connection is privileged and not subject to RLS, so the engine scopes every statement to your tenant itself. - Collector. Runs the engine's connectors and detectors, unmodified, inside your network. Target credentials stay in its local configuration. It connects outbound only, and every result payload is redacted before it leaves.
- Schema. The 16 database migrations define behavior, not only tables. pg_cron schedules scans, a trigger records posture history when a scan finishes, triggers hash-chain the evidence ledger per finding and reject edits to evidence, so the ledger is tamper-evident and
vectasec doctor --verify-ledgerfinds breaks, and a deferred constraint trigger enforces the citation gate on AI output.
The scan pipeline#
One scan routine runs every scan the engine performs, whether it starts from the CLI, the API or the queue worker:
- Ensure the connection row exists. Fields the connector declares secret go to Supabase Vault in the same transaction.
- Resolve the connection's Vault secret and merge it into the config, in memory only.
- Insert the scan as
running, or adopt the queued row, and commit. - Call the connector, which returns resources and edges. Collection runs outside any write transaction. If the connector raises, the scan is marked
error. - Enrich the collection: derive non-human identities, credentials and grants, and tag data-carrying resources with the data classes their names and labels suggest. A failure here is recorded in the scan stats and the scan continues.
- Upsert resources on
(connection_id, ext_id). A resource that is absent now is marked stale, not deleted, so its history survives. - Replace this connection's edges. An edge to something the scan did not collect is dropped.
- Run every detector against resources of the kinds it declares. A detector that raises is recorded and skipped.
- Upsert each finding on
(tenant_id, fingerprint)and insert its observations, the evidence it cites, in the same transaction. When a detector returns several hits for one resource, the finding takes the worst severity and every hit is kept as an observation. Active suppressions apply here. - Re-evaluate compliance in the same transaction, so control state never lags the findings. An evaluation error is recorded and does not void the scan.
- Mark the scan
done, write thescan.completedaudit row, and commit.
Steps 6 to 11 are one transaction. You see all of a scan's results or none of them.
After the scan#
The engine then runs these follow-up steps, each in its own transaction:
- Estate correlation. When the scan was a discovery sweep, it compares the services found with your connections and flags any that nobody is scanning.
- Runtime evidence. On connectors that can observe live clients, it records who was seen connected, and flags declared identities never seen in use once there is enough history to say so.
- OSV and NVD lookups. Each is a separate step. They send component identifiers (package URLs to OSV, CPEs to NVD) to those public feeds and file published advisories that match the running version. Set
VECTASEC_OSV_DISABLEDorVECTASEC_NVD_DISABLEDto turn either off on an install with no outbound access. - Notifications. Findings first seen in this scan go to your routing rules.
A failure in any of them never changes the scan's outcome. Correlation, runtime and feed errors are recorded in the scan's stats, and every notification attempt, sent or failed, is recorded as a delivery.
Scans run by a collector#
A collector runs steps 4 and 8 inside your network and pushes resource metadata, edges, and findings with their observations. It never sends the raw configuration the detectors read. The engine writes resources, edges, findings, observations, compliance and the audit row in one transaction, as in steps 6, 7 and 9 to 11.
Queue and scheduling#
Scans started through the API are queued by default, on pgmq inside Postgres.
POST /scans, or Scan now in the console (POST /connections/{id}/scan), inserts aqueuedscan row and a message on the scan queue in one transaction. The message carries only the public config; secrets stay in Vault.vectasec workerclaims up to 5 messages at a time with a 600-second visibility timeout and runs the pipeline. If a worker dies mid-scan, the message is redelivered; a rerun is safe because findings dedupe on their fingerprint.- A failing message is archived after 5 reads. Failures that a retry cannot fix, such as rejected credentials or a configuration error, are archived at once. Terminal failures go to your scan-failed routing rules.
- The worker marks any scan still
runningan hour after it started aserror.--reap-onlyruns only that check;--oncedrains the queue and exits. - For recurring scans, set
scan_interval_minuteswithPATCH /connections/{id}: 5 to 10080, or null for none. The console's schedule setting calls the same route and offers 30 minutes to 7 days. A scheduled database job runs every minute and queues each active connection that is due, skipping any with a scan already queued or running. - For a collector-mode connection, Scan now and the scheduler create only the queued row, never a pgmq message. The collector claims it through
GET /collector/jobs.
vectasec scan in the CLI, and POST /scans with "inline": true, run the pipeline in-process instead of queueing.
Connectors fan out#
Each connector and each detector registers itself when the engine starts, so there is no central list to edit. The engine loads 39 connectors and 448 detectors today.
The wide connectors share one pattern:
- One credential, many surfaces. One AWS connection reaches SQS, SNS, EventBridge, MSK and API Gateway in every region you list. One Kubernetes connection reads core resources, RBAC, network policies and admission webhooks, plus Istio, Linkerd and Gateway API where installed.
- Errors are isolated per surface. Each AWS service and region, and each Kubernetes resource list, is collected separately. A denied read is recorded on the
aws_accountork8s_clusterresource instead of failing the scan. For AWS, only the openingsts:GetCallerIdentitycheck can fail the scan, because it proves the credentials work. A Kubernetes API that returns 404 is recorded as not installed, not as an error. - Denied reads become
not_assessablefindings. A not-assessable finding can only leave a controlunknown; it never counts as a pass or a fail. - Resource kinds are distinct, such as
aws_sqs_queueormesh_gateway, so each detector runs only against the kinds it was written for.
Fingerprints and deduplication#
A finding's identity is a fingerprint of tenant, connection, rule and resource id: which rule fired, on which resource, through which connection. The scan id is not part of it.
- The same issue seen across 100 scans is one finding. Each scan that sees it adds observations to its evidence trail.
- The status you set (triaged, in progress, resolved, accepted) is kept across rescans, and a rescan never reopens a finding you resolved. Suppressions are re-evaluated on every scan, so an open finding becomes suppressed when a suppression starts and reopens when it expires.
- The same topic name in production and staging stays two findings, because the connection is part of the key.
- The collector computes the same fingerprint, so a finding has the same identity whether the engine or a collector scanned the connection.
Console reads, engine writes#
The console reads Supabase as the authenticated role, which has SELECT policies only, and RLS decides which tenant's rows you see. A few pages, such as integration settings, also call engine read routes from the server. The only database writes the console makes itself go through two narrow database functions: one creates your organization at sign-up, and one records when you were last active.
Every other product change is a Server Action. It re-verifies your session and your role; for a change to an existing row, it reads that row under RLS to confirm you can see it. It then calls the engine with the service token plus the tenant and actor it derived on the server. The engine checks that actor's role, then writes the change and its audit row in one transaction. The browser never sees the engine URL or token.