Skip to content

aegis · Reference

FAQ

Answers to common questions about aegis: MCP support, licensing, performance, what runs where, and how it behaves when a dependency fails.

On this page

These are short answers to common questions about evaluating and running aegis. Most answers link to the page that covers the topic in full.

Product and licensing#

Do I need to change my MCP servers or clients?#

Not if your servers speak MCP over Streamable HTTP. Point your client at the gateway's POST /mcp endpoint, or at /v1/invoke, which runs the same handler, and have it send a bearer token once you have configured a JWT verifier. Your MCP servers stay where they are, behind the gateway. aegis reaches them over HTTP or HTTPS and doesn't launch stdio servers. You register them in the server catalog, or, when you run without Postgres, you name a single upstream with UPSTREAM_URL. Both options are covered in Connect and secure MCP servers.

Which MCP features are supported?#

aegis covers the tools surface of MCP. It answers initialize and ping itself, serves tools/list from the tools it discovered on your servers, and forwards tools/call to the server that owns the tool. Notifications get HTTP 202 with an empty body. Any other method returns -32601, which means MCP resources and prompts are not proxied today. aegis declares MCP protocol version 2025-06-18 by default, both to its upstreams and in its own initialize reply, and MCP_PROTOCOL_VERSION overrides it. For the full request path, see How aegis works.

Is aegis open source?#

aegis is open-core. The gateway core is MIT-licensed and runs on its own. It includes the proxy, authentication, RBAC, threat detection, the audit log, approvals, the kill-switch, metrics, the dashboard, the discovery CLI and the Compose and Helm deployments. Billing, OAuth issuance and SCIM, the SIEM connectors, alerting, the credential broker and per-tenant namespace isolation are commercial, and every commercial file carries the SPDX identifier LicenseRef-AEGIS-Commercial. The code is not in a public repository today. To get access to it or to a commercial license, book a demo.

Is there a release version?#

Not yet. The project is pre-1.0 and has no tagged releases. The gateway, the Helm chart and the chart's appVersion are all at 0.1.0. You'll find the chart in Deploy aegis.

Running aegis#

What does it need to run?#

Redis is required. It holds sessions, rate-limit counters and kill-switches, and the gateway exits at startup if it can't connect to REDIS_URL. Postgres (DATABASE_URL) turns on the server catalog, RBAC, approvals and the admin API. Without it, aegis forwards every call to a single UPSTREAM_URL with RBAC disabled, and the admin API returns HTTP 503. ClickHouse (CLICKHOUSE_URL) turns on the audit log.

Before you expose the gateway, configure a JWT verifier (AUTH_HS256_SECRET or AUTH_JWKS_URL), set DATABASE_URL, and set CLICKHOUSE_URL together with a strong AUDIT_HMAC_KEY. With no verifier, every caller runs as anonymous. Without DATABASE_URL, no RBAC is applied. Without AUDIT_HMAC_KEY, the audit chain is keyed with a built-in development key. All the variables are listed in Configuration.

How much latency does it add?#

We do not publish a latency or throughput figure. The aegis team has run internal load tests against mock upstreams, but they are not an independent benchmark, they do not run in CI, and the host dominates tail latency, so measure on your own hardware during an evaluation. The checks each call passes through are listed in How aegis works.

What happens if Redis, Postgres or ClickHouse goes down?#

Each control has a defined failure behavior:

  • Redis. The gateway exits at startup if it can't reach Redis. While Redis is unreachable, calls to /mcp and /v1/invoke fail with HTTP 500, because each one creates or looks up a session in Redis. If only the rate-limit check fails, the call is allowed. Kill-switches already in force stay in force, and an outage never adds new ones.
  • Postgres. With DATABASE_URL set, the gateway tries the catalog load 10 times at startup and exits if Postgres is still unreachable. Once it is running, approve-gated calls fail closed, so they never run unapproved. RBAC keeps enforcing the last policy snapshot it loaded and retries the refresh. If the first snapshot load fails, RBAC denies every call until a load succeeds.
  • ClickHouse. Tool traffic never waits on the audit log. If ClickHouse is unreachable when the gateway starts, the gateway gives up after 30 attempts, about 30 seconds, logs Audit log DISABLED and runs without an audit log until it restarts. If a later write fails, or the write queue is full, that event is dropped and the gateway logs audit insert failed or audit queue full. Alert on these log lines.
  • Billing (commercial). Usage metering fails open, so a billing outage never blocks traffic.
  • Health checks. /healthz returns a static 200 and checks none of these stores, so monitor each store directly.

For every fail-closed default, see Security model.

Security and policy#

Can it run multi-tenant?#

Yes. By default, one deployment serves every tenant. aegis reads the tenant from a JWT claim (AUTH_TENANT_CLAIM, default tenant), and a token without that claim maps to default. Tenants are kept apart in routing, RBAC, circuit breakers and the audit log. Serving more than one tenant needs the Postgres catalog, because the single UPSTREAM_URL mode serves only the default tenant. To restrict an instance to specific tenants, set AUTH_ALLOWED_TENANTS. If the variable is set but empty, aegis refuses every tenant. The Helm chart's isolated tenancy mode sets the variable for you and can also render a Namespace, ResourceQuota and LimitRange for the release. Those per-tenant namespace guardrails are commercial. For details, see Identity and access.

Does it block tools that impersonate other tools?#

Not on name alone. Shadow and typosquat detection compares tool names across the servers in a tenant, including Unicode look-alikes and near matches, and reports what it finds at GET /admin/v1/shadows. It never blocks, because identical names across servers are often legitimate. A tool whose definition changes after aegis pinned it is a different case. In the default enforcing mode, aegis quarantines it: it removes the tool from tools/list and refuses calls with -32050 until an operator releases it with POST /admin/v1/quarantine, which pins the current definition.

Pinning is trust-on-first-use. A tool that is malicious the first time aegis sees it becomes its own baseline, so review a server before you add it to the catalog. Pins and quarantines live in each gateway process's memory, so resolve open quarantines before you restart the gateway. For more, see Threat detection.

Can I try policies without blocking traffic?#

For threat checks, yes. THREAT_MODE=monitor makes the integrity, injection, egress and PII checks log their findings without blocking, and response PII is not redacted in that mode. You can instead relax one check at a time with PII_ACTION=monitor, EGRESS_ACTION=monitor or INPUT_VALIDATION=monitor. Input-schema validation has its own switch and ignores THREAT_MODE. RBAC has no monitor mode: a deny rule, or a call with no matching allow, is always refused, so try new policies on a test subject or tenant first. The commercial monthly quota only observes by default and starts refusing calls when you set BILLING_ENFORCE=1. For details, see Threat detection and Identity and access.

Operating aegis#

Is the dashboard live?#

Only partly. /dashboard shows a built-in demo dataset until you paste an operator bearer token. It then reads live data from /admin/v1 on the same origin and keeps the token in that browser's localStorage. The console only reads: changes you make in it are not sent to the gateway. In live mode, the approvals, circuit breakers, SIEM sinks, alert channels and tenants panes stay empty because the console doesn't load them yet. Decide approvals through the admin API with GET /admin/v1/approvals and POST /admin/v1/approvals/:id. See Audit and observability and the API reference.