aegis · Concepts
Audit and observability
How aegis records every decision in a tamper-evident ClickHouse hash chain, verifies it, and exposes Prometheus metrics, correlation IDs and SIEM export.
On this page
Every audited decision, whether a tools/call outcome or an admin change, goes through one recording step. That step writes the tamper-evident audit log in ClickHouse and feeds the optional SIEM and alert sinks. Tool-call outcomes are also counted in Prometheus.
Tamper-evident audit log#
When CLICKHOUSE_URL is set, aegis appends every tools/call decision and every admin mutation to the audit_log table in CLICKHOUSE_DB (default aegis). At startup the gateway creates the database and table if they are missing, retrying for about 30 seconds while ClickHouse comes up. The log line Audit log enabled (HMAC-keyed chain continues from seq N) confirms the chain is live. If ClickHouse stays unreachable, or the chain cannot be read, the gateway keeps serving with the audit log disabled and logs Audit log DISABLED at error level. Check for the enabled line after every deploy.
Each row carries an HMAC-SHA256 tag keyed by AUDIT_HMAC_KEY. The tag binds the previous row's tag to every field of the row except correlation_id. The key never reaches ClickHouse, so someone with write access to the database cannot rebuild a valid chain. Editing a field, or moving a row to another tenant, changes its tag.
What gets recorded#
Admin mutations are rows with event_type set to admin. They cover role, policy, role-grant and server changes, quarantine releases, approval decisions and kill-switch changes. Their decision is create, update, delete, approve, deny, kill or unkill, and code is 0. server holds an identifier for what changed, such as the server, role, subject, approval id or kill-switch key. Deleting a server that had policies writes a second delete row for NAME:policies.
Rate-limited (-32029) and quota-refused (-32060) calls stop before dispatch, so they appear in Prometheus but not in the audit log. initialize, ping and tools/list are not audited.
Row fields#
The writer#
Each gateway process runs one writer task that owns the chain and continues it from the table's last row at boot. The request path hands events to it over a bounded queue of 4096 and never waits. When the queue is full, each dropped event is logged as audit queue full with a running total. If an insert fails, the writer re-reads the chain tip from the table, drops that event and logs audit insert failed. Alert on both lines, and size ClickHouse so the queue keeps draining.
Verify the chain#
curl -s http://localhost:8080/admin/v1/audit/verify -H "Authorization: Bearer $TOKEN"
$TOKEN is an admin token, such as the one from the quick start. The endpoint recomputes every tag in seq order, then checks that the table still reaches the highest seq this process knows was written. It returns {ok, rows, first_break_seq}:
It catches edited, inserted, deleted and reordered rows, and a truncated tail while the gateway runs. The chain is one global sequence across tenants, so the caller needs catalog:read on aegis-admin in the default tenant, and other tenants get 403. The endpoint returns 503 when RBAC is off (no DATABASE_URL) or the audit log is disabled, and 502 if ClickHouse cannot be read. It reads the whole table on every call, so run it on a schedule, not in a tight loop.
Query events#
curl -s "http://localhost:8080/admin/v1/audit/events?limit=50&decision=deny" -H "Authorization: Bearer $TOKEN"
The response is {"events": [...]}, newest first, with the row fields above minus the chain tags. limit defaults to 100 and is clamped to 1–1000. decision matches the column exactly, so deny also returns denied approvals logged as admin events. event_type tells them apart.
Results are always scoped to the caller's tenant. The endpoint needs catalog:read on aegis-admin. It returns 503 when RBAC is off or the audit log is disabled, and 502 if the query fails. GET /admin/v1/overview also reports your tenant's row count.
Correlation IDs#
On /mcp and /v1/invoke, aegis adopts the trace id from a valid W3C traceparent header, or mints a new 128-bit id. The id is returned in X-Correlation-ID and written to the call's audit row and SIEM event. A forwarded tools/call also sends it upstream in a new traceparent with a fresh span id, so the server joins your trace. For example, traceparent: 00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01 comes back as X-Correlation-ID: 4bf92f3577b34da6a3ce929d0e0e4736. Admin events have an empty correlation_id.
Prometheus metrics#
/metrics is closed by default and returns 503. Set METRICS_TOKEN to require a matching bearer token (preferred; a missing or wrong token gets 401), or METRICS_PUBLIC=1 for unauthenticated scraping when the endpoint is isolated at the network layer.
curl -s http://localhost:8080/metrics -H "Authorization: Bearer $METRICS_TOKEN" | grep '^aegis_'
The label sets are fixed, and no series carries a subject, tool or server label. Counters are per replica and reset on restart, so use rate() and sum across replicas. indeterminate is not a decision label, so find timeouts in the audit log. An egress DLP refusal (-32053) counts as decision="block" with no kind.
A failed upstream call returns HTTP 200 with a JSON-RPC error body, so it never appears as a 5xx. Alert on the dispatch error rate instead, the same ratio the Helm chart's canary analysis uses:
100 * sum(rate(aegis_tool_calls_total{decision="error"}[5m]))
/ sum(rate(aegis_tool_calls_total[5m]))
In the Helm chart, podMonitor.enabled: true creates a Prometheus Operator PodMonitor. Unless config.metricsPublic is "1", it reads the bearer token from the metricsToken key of the gateway's Secret (the chart's own, or secrets.existingSecret).
Operator console#
/dashboard is a single page compiled into the binary, served with a strict Content-Security-Policy, X-Frame-Options: DENY and Referrer-Policy: no-referrer. It shows a demo dataset until you connect it.
Paste a bearer token with catalog:read on aegis-admin to switch to live data. The console then polls every 3 seconds for the overview, the latest 100 audit events, servers, roles, policies, users, shadow findings and kill-switches. It verifies the chain once on connect, so chain status appears only for an operator token in the default tenant. The token stays in that browser's localStorage until you choose Disconnect, so use a dedicated read-only subject.
The console only sends GET requests. Make changes through the admin API. It does not read approvals, circuit breakers, SIEM sinks, alert channels or tenants yet, so those panes stay empty in live mode.
SIEM and alerts (commercial)#
SIEM export ships every tool-call and admin event, with or without ClickHouse.
Each sink has its own bounded queue and sends batches of up to 64 events, so a slow sink never adds latency to a call. Delivery has a 5-second timeout and no retry: a batch the sink rejects or cannot receive is logged and dropped. Keep the ClickHouse audit log as your record of truth. Events carry the tenant and correlation id, but not seq or the chain tags.
Alerts are the urgent subset: injection, PII, integrity and credential blocks, plus rug-pull and shadow findings from discovery and the background sweeps.
Repeats of the same kind, tenant, server and tool are deduplicated for ALERT_COOLDOWN_SECS (default 300). Each tenant is also capped at 60 alerts a minute, and low-severity findings such as shadow collisions may use at most 20 of those. Set ALERT_COOLDOWN_SECS well above SHADOW_SWEEP_SECS. A Slack incoming-webhook URL is itself a credential, so supply it as a secret, as the Helm chart does with secrets.slackWebhookUrl.
Related pages#
- Configuration: every audit, metrics, SIEM and alert variable.
- Error codes: what each
codein an audit row means. - Hold, approve and stop tool calls: the actions behind the approval and kill-switch rows.
- Deploy aegis: supplying
AUDIT_HMAC_KEYandMETRICS_TOKENthrough the Helm chart.