Vectasec · Guides
Run a collector in your network
Deploy the outbound-only Vectasec Collector next to private brokers so credentials and raw configuration never leave your network.
On this page
A collector scans systems that only answer from inside your network. It runs the engine's own connectors and detectors next to your brokers and gateways, keeps their credentials local and connects outbound only. It pushes findings, their observations and resource metadata to the engine, never the raw configuration it read.
When to use a collector#
Use one for systems no cloud API can read, such as self-hosted Kafka ACLs, the Kong admin API or the RabbitMQ management API, or when a target's credentials must not leave your network. If a worker can already reach the target and its credential may live in Supabase Vault, a cloud-mode connection is simpler.
The collector is version 0.0.1 and is an early release. The engine serves the four /collector/* routes it calls, and the full loop has been run end to end against a development Redis. Read the limitations below before you rely on it.
Before you start#
- An engine reachable over HTTPS. The collector refuses a plain
http://URL, and the engine itself serves plain HTTP, on127.0.0.1:8000by default (ENGINE_HOST,ENGINE_PORT). Put a TLS-terminating proxy in front of it and expose only the/collector/*routes to the collector's network. VECTASEC_COLLECTOR_TOKEN_SECRETset on the engine to an independent random value before you expose those routes. It signs collector session tokens. When it is unset, the engine derives a key fromDATABASE_URL, which is meant for single-host development only.- Docker and a checkout of the Vectasec source (book a demo to get it).
- A read-only credential for each system you want to scan.
1. Create an enrollment token#
Generate a token into a file on the collector host, print its SHA-256 hash, then hand the file to the container's user:
openssl rand -hex 32 > enroll.token
chmod 600 enroll.token
tr -d '\n' < enroll.token | openssl dgst -sha256 -r | cut -d' ' -f1
sudo chown 10001:10001 enroll.token
The container runs as uid 10001, so on a Linux host it cannot read a mode-600 file you own; the last line fixes that without widening the mode. The collector strips the trailing newline before it sends the token, so hash the token without it, as above. The engine hashes what it receives the same way and looks the result up in collectors.enroll_token_hash. Insert a row holding the hash, so the token itself never reaches the database:
insert into collectors (tenant_id, name, enroll_token_hash)
values ('<tenant-uuid>', 'dc1-collector', '<sha256-hex>');
Use one row and one token per collector. The row starts pending and turns active at the first enrollment, when the collector also replaces name with its VECTASEC_COLLECTOR_NAME (the hostname by default). The token is not single-use: the collector enrolls again on every restart and whenever its session token nears expiry, and the engine maps the token to the same row each time, so keep the file in place. To rotate it, write a new token file the same way, set enroll_token_hash to its hash and restart the collector.
2. Describe your targets#
targets.json lists the systems this collector may scan. Write each credential as a ${ENV_VAR} reference, so the file holds no secrets and can live in git or a ConfigMap. An unset reference stops the collector at boot.
{
"targets": [
{ "name": "prod-kafka", "type": "kafka",
"config": { "bootstrap": "kafka-1.internal:9092,kafka-2.internal:9092" } },
{ "name": "prod-kong", "type": "kong",
"config": { "url": "http://kong.internal:8001", "admin_token": "${KONG_ADMIN_TOKEN}" } },
{ "name": "prod-rabbit", "type": "rabbitmq",
"config": { "url": "https://rabbit.internal:15672", "username": "vectasec-ro",
"password": "${RABBIT_PASSWORD}" } },
{ "name": "internal-mcp", "type": "mcp",
"config": { "url": "https://mcp.internal/mcp", "token": "${MCP_TOKEN}" } }
]
}
Target names must be unique. A scan job names a target and carries no host or credential, so the engine cannot point the collector at a system you did not list. Each type must be a registered connector, which the collector checks at boot; the connector reference lists each one's fields.
3. Build the image#
Build it with the collector Dockerfile from the release; the build context must include the engine, because the image installs the engine's connectors and detectors unmodified. Tag it, for example, vectasec/collector:0.0.1.
The image runs as uid 10001, has no shell or package manager, and exposes no port, so use check, preview and the logs rather than docker exec.
4. Check, then run#
docker run --rm --read-only --cap-drop=ALL --security-opt no-new-privileges \
-e VECTASEC_CONTROL_PLANE_URL=https://<your-engine-host> \
-e VECTASEC_ENROLL_TOKEN_FILE=/run/secrets/enroll \
-e VECTASEC_COLLECTOR_NAME=dc1-collector \
-e KONG_ADMIN_TOKEN -e RABBIT_PASSWORD -e MCP_TOKEN \
-v /path/to/enroll.token:/run/secrets/enroll:ro \
-v /path/to/targets.json:/etc/vectasec/targets.json:ro \
-v vectasec-state:/var/lib/vectasec \
vectasec/collector:0.0.1 preview
Neither review command contacts the engine, but both load the full configuration, so the engine URL and enrollment token must still be set:
checkvalidates the configuration and connector types and prints each target's redacted config, without connecting to any target.previewscans every target with your real credentials and prints exactly what would be pushed, as JSON.
Read the preview output first. Then run the same command with run, the default, in place of preview. For a long-running collector, replace --rm with -d --restart unless-stopped.
Allow egress only to the engine and your targets; the collector needs no inbound rule. If a proxy inspects outbound TLS, mount its CA bundle and point VECTASEC_CA_BUNDLE at that path rather than turning verification off.
5. Bind connections to the collector#
Scans reach a collector through collector-mode connections. Create one per target with the same type and a name identical to the target name, because the job carries only that name.
curl -s -X POST "$VECTASEC_ENGINE_URL/connections" \
-H "Authorization: Bearer $VECTASEC_ENGINE_TOKEN" \
-H "X-VectaSec-Tenant: <tenant-uuid>" \
-H "X-VectaSec-Actor: <your-user-uuid>" \
-H "Content-Type: application/json" \
-d '{"type":"kafka","name":"prod-kafka","mode":"collector","config":{"bootstrap":"kafka-1.internal:9092"}}'
The actor must be a tenant member with the analyst role or higher; a Supabase user JWT also works as the bearer token. With exactly one collector registered to your tenant, the engine binds the connection to it. With several, add collector_id from GET /collectors, which lists each collector's id, status and whether it sent a heartbeat in the last two minutes.
The engine checks a connector's required fields in collector mode too, but the collector never reads this config. Send non-secret values only, and where a credential field is required, such as the RabbitMQ password, send a placeholder.
Use check and preview on the collector to test reachability, not the connection's Test action. Test runs the connector from the engine against the stored config. Unless the engine can reach the target itself, the test fails and marks the connection error, and scheduled scans only queue for active connections.
Queue a scan with Scan now in the console, POST /connections/{id}/scan, or a schedule set through scan_interval_minutes on PATCH /connections/{id}. The engine writes a queued scan row, and the bound collector claims it on its next poll.
How the collector talks to the engine#
- Enroll.
POST /collector/enrollexchanges the enrollment token for an HMAC-signed session token bound to your tenant and this collector. It expires after 3600 seconds, lives in memory only, and is renewed a minute early. - Heartbeat every 60 seconds with its version, status, and target names and types.
- Poll
GET /collector/jobs. The engine hands over up to 5 queued scans and marks them running. The collector sends a 25-second wait hint, but the engine answers immediately today, so an idle collector polls every 5 seconds. - Scan locally. The connector reads the target with your credential, and every detector runs against what it read.
- Redact and push. Redaction runs over the whole payload, then
POST /collector/resultsdelivers it. The engine confirms the scan belongs to this collector and writes resources, edges, findings, observations and compliance in one transaction.
A failed cycle backs off exponentially: the delay starts at 1 second and doubles up to 300 seconds (VECTASEC_MAX_BACKOFF_S), and each wait is jittered by up to 50 percent either way. Redaction scrubs in three layers: every held credential of 8 or more characters, every credential-shaped key, and URL userinfo or PEM private keys. Log records get the same scrub.
What leaves your network#
An observation such as security.inter.broker.protocol=PLAINTEXT is the evidence for its finding, so the values a rule cites do leave. The configuration as a whole does not.
On disk, the collector writes only state.json in its state directory, mode 0600: the collector id, the enrollment token's SHA-256 hash, the enrollment time and the version. The session token is held in memory only.
Current limitations#
- No mTLS. The collector authenticates with server-side TLS plus a bearer token.
- Unsigned images. No signature, SBOM or provenance attestation. Build from source and pin the tag.
- No result spool. If a push fails, that scan's results are dropped. The scan stays
runninguntilvectasec worker, which ages out scans left running for more than an hour, marks it as an error, and scheduled scans for that connection wait until then. Queue a new scan with Scan now to replace it sooner. - One request per scan. Results are not chunked, so allow large request bodies on the proxy in front of the engine.
- No per-target rate limiting. A scan calls your target's admin API as fast as the connector does.
- No state resume. The state file is a record only; every restart enrolls again.
- Kafka with SASL or TLS. The Kafka connector passes only the bootstrap address, with no SASL or TLS settings, so clusters that require either on the listener it connects to cannot be scanned yet.
- Fewer engine-side steps. Collector scans skip enrichment and the follow-up steps: no derived non-human identities or data classes, estate correlation, OSV or NVD lookups, runtime evidence or notifications. See How Vectasec works.
Next steps#
- Configuration: every collector setting and its default
- Deployment: exposing the engine and the production checklist
- Security model: authentication paths and roles
- CLI and API reference: the connection and collector endpoints