aegis · Concepts
How aegis works
The ordered request pipeline behind every aegis tool call, from authentication to audit, and the Redis, Postgres and ClickHouse stores behind it.
On this page
The request pipeline#
Every tools/call runs the same ordered sequence of checks. Cheap checks run first and the upstream forward runs last, so a refused call never reaches your MCP server. Any stage can refuse the call, and each refusal carries its own JSON-RPC code (see Error codes).
- Authenticate. Middleware on the protected routes verifies the bearer JWT: HS256 with a shared secret, RS256 through a JWKS endpoint, or a token from aegis's own OAuth issuer (commercial). A missing or invalid token gets HTTP 401 with a plain JSON body, not a JSON-RPC error. So does a valid token whose tenant is not listed in
AUTH_ALLOWED_TENANTS, when that variable is set. - Correlation ID. aegis adopts the trace ID from the caller's W3C
traceparentheader, or mints one. It is returned inX-Correlation-IDand sent upstream as a childtraceparent. - Parse. A body that is not JSON gets HTTP 400 with
-32700. JSON that is not a valid JSON-RPC 2.0 request gets HTTP 400 with-32600. A notification (noid) gets HTTP 202 with an empty body and goes no further. - Session. Without an
X-Session-IDheader, aegis mints a session. A session you supply must exist (otherwise HTTP 401) and belong to you (otherwise HTTP 403). - Rate limit. Per session (
RATE_LIMIT_PER_SESSION, default 120 a minute) and per session and tool (RATE_LIMIT_PER_TOOL, default 60 a minute).0turns a limit off. Over the limit gets HTTP 429 with-32029. - Monthly quota (commercial). With billing enabled and
BILLING_ENFORCE=1, a tenant over its monthly quota gets HTTP 402 with-32060. With billing enabled but not enforced, the overage is counted in metrics and the call continues. - Route. aegis resolves the owning server once and uses that answer for both the RBAC check and the forward, so nothing can change between check and use. An unknown tool gets
-32601after a throttled rediscovery. - Kill-switch. A switch on the tenant, server or tool refuses the call with
-32062. It overrides RBAC. - RBAC. With
DATABASE_URLset, a call with no matchingallow, or with any matchingdeny, gets HTTP 403 with-32003. Adenywins overapprove, andapprovewins overallow. - Approval hold. When an
approvepolicy matches, the call is held and refused with-32061and anapproval_id. After an operator approves it, the same caller can run the identical call (same arguments) once, withinAPPROVAL_TTL_SECS(default 3600) of the original request. - Input-schema validation. When the tool declares an
inputSchema, the arguments are checked against it. A violation gets-32602.INPUT_VALIDATIONdefaults to enforce.monitorlogs the violation instead, andoffskips the check. - Threat gate. The whole
paramssubtree is scanned. A quarantined tool gets-32050and an injection pattern gets-32051. A credential, card number, SSN or private key in the outbound arguments gets-32053(EGRESS_ACTION=monitorlogs it instead, andoffskips the scan).THREAT_MODE=monitorlogs all of these without blocking. - Credential broker (commercial). If the server declares a
secret_ref, aegis resolves it from AWS Secrets Manager and injects it asAuthorization: Bearer. If it cannot, or the broker is turned off, the call is refused with-32054. - Forward. Backpressure and the circuit breaker, both keyed by tenant and server, fast-fail with
-32063. The breaker opens afterCIRCUIT_FAILURE_THRESHOLD(default 5) consecutive failures. Connection failures are retried up toUPSTREAM_MAX_RETRIES(default 2), andhttpsupstreams can use mTLS. A timeout (UPSTREAM_TIMEOUT_SECS, default 10) gets-32064with"safe_to_retry": false, because the tool may have run. Any other upstream failure gets-32000. - Response scan. Any echo of a brokered credential becomes
[REDACTED:UPSTREAM_CREDENTIAL]. The response is then scanned for PII and secrets. By default each hit is redacted as[REDACTED:KIND]. WithPII_ACTION=block, a response carrying a credential, card number, SSN or private key is refused with-32052.PII_ACTION=monitoronly logs. - Record and return. The decision is recorded, the call is metered if billing is on, and the response is fanned out to SSE subscribers on the session before it returns with
X-Session-IDandX-Correlation-ID.
From stage 7 on, every refusal except an RBAC deny comes back as HTTP 200 with a JSON-RPC error body. Check error.code, not only the HTTP status.
One recording path#
From stage 7 on, every outcome goes through one recording step. It increments the Prometheus decision counters and writes one audit event, which goes to the HMAC-SHA256 hash-chained audit log in ClickHouse and to the SIEM sinks (commercial). Integrity, injection, PII and credential blocks (-32050, -32051, -32052, -32054) also go to the alert channels (commercial). Admin changes to servers, roles, policies, role grants, approvals, kill-switches and quarantines go through the same recording step. Because one event feeds every destination, you do not wire the audit log, SIEM and alerts separately.
Rate-limit and quota refusals stop before stage 7. They are counted in Prometheus but not written to the audit log. See Audit and observability.
Background work#
- Integrity re-sweep. Every
INTEGRITY_RESWEEP_SECS(default 15,0disables), aegis re-lists each upstream's tools and compares every definition with its SHA-256 pin. A changed definition is logged and alerted (commercial). UnlessTHREAT_MODE=monitor, it is also quarantined: it disappears fromtools/listand calls get-32050until an operator re-pins it withPOST /admin/v1/quarantine. Quarantine is sticky, so reverting the upstream does not lift it. - Shadow sweep. At startup, after catalog changes and every
SHADOW_SWEEP_SECS(default 60), aegis compares tool names across servers for exact shadows and look-alike names. Findings appear atGET /admin/v1/shadows. The sweep only reports. It never blocks. - Catalog reconciler. With
DATABASE_URLset, everyCATALOG_RECONCILE_SECS(default 30) aegis compares the enabled servers in Postgres with its routing table and rebuilds the table if they differ. - Kill-switch cache. Every
KILLSWITCH_REFRESH_SECS(default 2), each replica reloads the switch set from Redis. The replica that flips a switch applies it at once, and the others pick it up on their next refresh. - Policy snapshot. This refresh runs on demand, not on a timer. When the RBAC snapshot is older than
POLICY_REFRESH_SECS(default 30), a single request reloads it while the others keep using the current snapshot. Policy changes take effect within about that window.
MCP surface#
Point MCP clients at POST /mcp. POST /v1/invoke is the same handler at its original path.
aegis does not advertise or proxy resources or prompts. It tracks sessions in X-Session-ID (1-hour TTL, refreshed on use) and does not issue an Mcp-Session-Id to clients. To see results live, open GET /v1/subscribe?session_id=... as an SSE stream, as the session's owner. Each response an upstream returns on that session goes out to every subscriber, up to 64 per session. aegis's own refusals and errors are not streamed, and a slow subscriber drops events instead of slowing the call.
Toward your servers, aegis is an MCP client over Streamable HTTP. It runs initialize once per upstream (and again if the upstream expires the session with HTTP 404), echoes any Mcp-Session-Id the upstream issues, and accepts a JSON or SSE reply. Each upstream request is built from scratch, so your client's headers are never passed through.
Backing services#
Shared state lives in these stores, so the gateway scales horizontally. The Helm chart runs 2 replicas by default (see Deploy aegis). Some state stays in each replica's memory:
- SSE subscriptions. A subscriber only receives responses for calls that its own replica handles. If you rely on
/v1/subscribe, keep each session on one replica. - Integrity pins, quarantines and shadow findings. Each replica builds these from its own sweeps, and a restarted replica pins every tool again on first sight.
GETandPOST /admin/v1/quarantineact on the replica that serves the request. - Circuit-breaker and backpressure state.
UPSTREAM_MAX_CONCURRENT(default 64,0disables) caps in-flight calls to each upstream on each replica, so the cluster-wide ceiling grows with the replica count.