Skip to content

aegis · Concepts

Audit and observability

How aegis records every decision in a tamper-evident ClickHouse hash chain, verifies it, and exposes Prometheus metrics, correlation IDs and SIEM export.

On this page

Every audited decision, whether a tools/call outcome or an admin change, goes through one recording step. That step writes the tamper-evident audit log in ClickHouse and feeds the optional SIEM and alert sinks. Tool-call outcomes are also counted in Prometheus.

Tamper-evident audit log#

When CLICKHOUSE_URL is set, aegis appends every tools/call decision and every admin mutation to the audit_log table in CLICKHOUSE_DB (default aegis). At startup the gateway creates the database and table if they are missing, retrying for about 30 seconds while ClickHouse comes up. The log line Audit log enabled (HMAC-keyed chain continues from seq N) confirms the chain is live. If ClickHouse stays unreachable, or the chain cannot be read, the gateway keeps serving with the audit log disabled and logs Audit log DISABLED at error level. Check for the enabled line after every deploy.

Each row carries an HMAC-SHA256 tag keyed by AUDIT_HMAC_KEY. The tag binds the previous row's tag to every field of the row except correlation_id. The key never reaches ClickHouse, so someone with write access to the database cannot rebuild a valid chain. Editing a field, or moving a row to another tenant, changes its tag.

What gets recorded#

decisionRecorded whenCodes
allowThe call was forwarded and the response scan did not block it. PII and secrets in the response may have been redacted.0
denyUnknown tool, kill-switch, RBAC denial, approval held, denied or unavailable, or an input-schema violation-32601, -32062, -32003, -32061, -32602
blockQuarantined tool, injection, egress DLP, a response blocked for PII, or a brokered credential that could not be resolved-32050, -32051, -32053, -32052, -32054
errorUpstream at capacity or its circuit open, or the upstream call failed-32063, -32000
indeterminateThe upstream deadline expired after the request was sent, so the tool may have run-32064

Admin mutations are rows with event_type set to admin. They cover role, policy, role-grant and server changes, quarantine releases, approval decisions and kill-switch changes. Their decision is create, update, delete, approve, deny, kill or unkill, and code is 0. server holds an identifier for what changed, such as the server, role, subject, approval id or kill-switch key. Deleting a server that had policies writes a second delete row for NAME:policies.

Rate-limited (-32029) and quota-refused (-32060) calls stop before dispatch, so they appear in Prometheus but not in the audit log. initialize, ping and tools/list are not audited.

Row fields#

FieldContent
seq, ts_msPosition in the chain, starting at 1, and the write time in Unix milliseconds
event_typetool_call or admin
tenant, subjectThe caller's tenant and JWT subject (anonymous when authentication is off)
server, toolThe routed server and the tool name
decision, code, reasonThe outcome, its JSON-RPC code and a short reason
session_id, request_idThe session (from X-Session-ID, or the one minted for the call) and the JSON-RPC request id. Empty on admin rows.
correlation_idThe trace id. It is stored but is not covered by the tag.
prev_hash, row_hashThe chain tags. The events API does not return them.

The writer#

Each gateway process runs one writer task that owns the chain and continues it from the table's last row at boot. The request path hands events to it over a bounded queue of 4096 and never waits. When the queue is full, each dropped event is logged as audit queue full with a running total. If an insert fails, the writer re-reads the chain tip from the table, drops that event and logs audit insert failed. Alert on both lines, and size ClickHouse so the queue keeps draining.

Verify the chain#

bash
curl -s http://localhost:8080/admin/v1/audit/verify -H "Authorization: Bearer $TOKEN"

$TOKEN is an admin token, such as the one from the quick start. The endpoint recomputes every tag in seq order, then checks that the table still reaches the highest seq this process knows was written. It returns {ok, rows, first_break_seq}:

FieldMeaning
oktrue when every tag matches and no rows are missing from the tail
rowsRows read, across all tenants
first_break_seqThe first row that fails, or the seq after the last surviving row when the tail was cut. null when ok is true.

It catches edited, inserted, deleted and reordered rows, and a truncated tail while the gateway runs. The chain is one global sequence across tenants, so the caller needs catalog:read on aegis-admin in the default tenant, and other tenants get 403. The endpoint returns 503 when RBAC is off (no DATABASE_URL) or the audit log is disabled, and 502 if ClickHouse cannot be read. It reads the whole table on every call, so run it on a schedule, not in a tight loop.

Query events#

bash
curl -s "http://localhost:8080/admin/v1/audit/events?limit=50&decision=deny" -H "Authorization: Bearer $TOKEN"

The response is {"events": [...]}, newest first, with the row fields above minus the chain tags. limit defaults to 100 and is clamped to 1–1000. decision matches the column exactly, so deny also returns denied approvals logged as admin events. event_type tells them apart.

Results are always scoped to the caller's tenant. The endpoint needs catalog:read on aegis-admin. It returns 503 when RBAC is off or the audit log is disabled, and 502 if the query fails. GET /admin/v1/overview also reports your tenant's row count.

Correlation IDs#

On /mcp and /v1/invoke, aegis adopts the trace id from a valid W3C traceparent header, or mints a new 128-bit id. The id is returned in X-Correlation-ID and written to the call's audit row and SIEM event. A forwarded tools/call also sends it upstream in a new traceparent with a fresh span id, so the server joins your trace. For example, traceparent: 00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01 comes back as X-Correlation-ID: 4bf92f3577b34da6a3ce929d0e0e4736. Admin events have an empty correlation_id.

Prometheus metrics#

/metrics is closed by default and returns 503. Set METRICS_TOKEN to require a matching bearer token (preferred; a missing or wrong token gets 401), or METRICS_PUBLIC=1 for unauthenticated scraping when the endpoint is isolated at the network layer.

bash
curl -s http://localhost:8080/metrics -H "Authorization: Bearer $METRICS_TOKEN" | grep '^aegis_'
SeriesLabelsCounts
aegis_tool_calls_totaldecision: allow, deny, block, error, limitedtools/call outcomes. limited is a rate-limited call.
aegis_threat_blocks_totalkind: injection, pii, integrity, credentialBlocks by kind
aegis_http_requests_totalstatus: 1xx, 2xx, 3xx, 4xx, 5xxEvery HTTP response on every route, including probes and scrapes
aegis_quota_exceeded_totalmode: observe, enforceOver-quota calls (commercial billing)
aegis_approvals_totalstate: required, granted, deniedApproval-gated calls by outcome
aegis_input_rejected_totalmode: monitor, enforceInput-schema violations
aegis_killswitch_blocked_totalscope: tenant, server, toolCalls refused by a kill-switch
aegis_upstream_rejected_totalreason: circuit_open, busyCalls fast-failed by the resilience layer
aegis_circuit_opened_totalnoneCircuit breaker trips
aegis_build_infoversionGauge, always 1

The label sets are fixed, and no series carries a subject, tool or server label. Counters are per replica and reset on restart, so use rate() and sum across replicas. indeterminate is not a decision label, so find timeouts in the audit log. An egress DLP refusal (-32053) counts as decision="block" with no kind.

A failed upstream call returns HTTP 200 with a JSON-RPC error body, so it never appears as a 5xx. Alert on the dispatch error rate instead, the same ratio the Helm chart's canary analysis uses:

text
100 * sum(rate(aegis_tool_calls_total{decision="error"}[5m]))
  / sum(rate(aegis_tool_calls_total[5m]))

In the Helm chart, podMonitor.enabled: true creates a Prometheus Operator PodMonitor. Unless config.metricsPublic is "1", it reads the bearer token from the metricsToken key of the gateway's Secret (the chart's own, or secrets.existingSecret).

Operator console#

/dashboard is a single page compiled into the binary, served with a strict Content-Security-Policy, X-Frame-Options: DENY and Referrer-Policy: no-referrer. It shows a demo dataset until you connect it.

Paste a bearer token with catalog:read on aegis-admin to switch to live data. The console then polls every 3 seconds for the overview, the latest 100 audit events, servers, roles, policies, users, shadow findings and kill-switches. It verifies the chain once on connect, so chain status appears only for an operator token in the default tenant. The token stays in that browser's localStorage until you choose Disconnect, so use a dedicated read-only subject.

The console only sends GET requests. Make changes through the admin API. It does not read approvals, circuit breakers, SIEM sinks, alert channels or tenants yet, so those panes stay empty in live mode.

SIEM and alerts (commercial)#

SIEM export ships every tool-call and admin event, with or without ClickHouse.

SinkVariables
Splunk HECSPLUNK_HEC_URL, SPLUNK_HEC_TOKEN
Datadog LogsDATADOG_LOGS_URL, DATADOG_API_KEY
OTLP/HTTP logsOTLP_LOGS_URL
Generic webhookSIEM_WEBHOOK_URL

Each sink has its own bounded queue and sends batches of up to 64 events, so a slow sink never adds latency to a call. Delivery has a 5-second timeout and no retry: a batch the sink rejects or cannot receive is logged and dropped. Keep the ClickHouse audit log as your record of truth. Events carry the tenant and correlation id, but not seq or the chain tags.

Alerts are the urgent subset: injection, PII, integrity and credential blocks, plus rug-pull and shadow findings from discovery and the background sweeps.

SinkVariables
SlackSLACK_WEBHOOK_URL
PagerDuty Events v2PAGERDUTY_ROUTING_KEY, optional PAGERDUTY_EVENTS_URL
Generic webhookALERT_WEBHOOK_URL

Repeats of the same kind, tenant, server and tool are deduplicated for ALERT_COOLDOWN_SECS (default 300). Each tenant is also capped at 60 alerts a minute, and low-severity findings such as shadow collisions may use at most 20 of those. Set ALERT_COOLDOWN_SECS well above SHADOW_SWEEP_SECS. A Slack incoming-webhook URL is itself a credential, so supply it as a secret, as the Helm chart does with secrets.slackWebhookUrl.