Skip to content

aegis · Get started

Deploy aegis

Deploy aegis with Docker Compose or the hardened Helm chart: external stores, secrets handling, network policy, tenancy modes and canary releases.

On this page

Deployment options#

You can run aegis in two ways. Docker Compose is for local work and evaluation: it starts the gateway with every backing store and a set of mock servers. The Helm chart is for Kubernetes and is where you run production traffic. Both run the same gateway image, which you build from the gateway directory of the source tree:

bash
cd gateway
docker build -t aegis-gateway:local .

The image runs as a non-root user (uid 1000) and listens on port 8080. aegis is pre-1.0 and has no tagged releases yet, so build the image from source.

For Kubernetes, push the image to a registry your cluster can pull from, then set image.repository and image.tag. The chart's default image.repository may not be reachable from your cluster. An empty image.tag falls back to the chart's appVersion (0.1.0). In production, pin a digest or an immutable tag.

Docker Compose#

From the repository root:

bash
docker compose -f deploy/docker-compose.yml up --build

This starts 19 services: the gateway, an Envoy TLS front end, Redis, Postgres, ClickHouse, LocalStack (standing in for AWS Secrets Manager), a test certificate generator, several mock MCP servers plus an MCP conformance server, and mock OIDC, SIEM and Stripe services for the end-to-end tests. A 20th service, the discovery scanner, sits behind the tools profile and runs only on demand. The gateway is published on :8080 and Envoy on :8443.

bash
curl -k https://localhost:8443/healthz

A healthy gateway returns {"status":"ok","service":"aegis-gateway"}. Envoy generates a self-signed localhost certificate the first time its container starts, which is why the command uses -k. In production, put a real certificate or load balancer in front of the gateway.

Kubernetes (Helm)#

The chart is at deploy/helm/aegis (chart version 0.1.0). A minimal install reads credentials from a Secret you manage. Create it first. It must contain at least auditHmacKey, or the pod does not start:

bash
kubectl create secret generic aegis-credentials \
  --from-literal=auditHmacKey="$(openssl rand -hex 32)"

Then install the chart:

bash
helm install aegis deploy/helm/aegis \
  --set config.clickhouseUrl='http://clickhouse:8123' \
  --set secrets.existingSecret=aegis-credentials

config.redisUrl defaults to redis://redis:6379. To turn on the server catalog and RBAC, add databaseUrl to the Secret.

Backing stores are external#

Redis, Postgres and ClickHouse are not subcharts. You point the chart at stores you already run. Each missing or blocked store behaves differently:

  • No databaseUrl: the gateway runs a single upstream (config.upstreamUrl) with RBAC off, and logs a warning at startup.
  • No clickhouseUrl: there is no audit log, and the gateway logs Audit log DISABLED at startup.
  • Redis blocked: the gateway never binds :8080, and the pod restarts in a loop.
  • Postgres blocked (with databaseUrl set): the gateway retries the catalog load 10 times, two seconds apart, then exits.
  • ClickHouse blocked: the gateway logs Audit log DISABLED at startup and keeps serving traffic without an audit log. Nothing fails, so watch for that line.

Secrets#

Put every DSN that contains a password under secrets.*. Values under config.* render into a plaintext ConfigMap. The preferred approach is secrets.existingSecret, a Secret you manage outside the chart, for example with External Secrets or SOPS. When it is set, the other secrets.* values are ignored and the chart renders no Secret of its own. The keys in that Secret match the names below.

KeyEnvironment variableUsed for
auditHmacKeyAUDIT_HMAC_KEYKeys the audit hash chain. Required in an existing Secret.
authHs256SecretAUTH_HS256_SECRETHS256 shared secret. In production, use your IdP's JWKS instead.
scimBearerTokenSCIM_BEARER_TOKENSCIM provisioning (commercial)
metricsTokenMETRICS_TOKENBearer token for /metrics
oauthSigningKeyPemOAUTH_SIGNING_KEY_PEMStable OAuth signing key (commercial)
splunkHecTokenSPLUNK_HEC_TOKENSplunk sink (commercial)
datadogApiKeyDATADOG_API_KEYDatadog sink (commercial)
slackWebhookUrlSLACK_WEBHOOK_URLSlack alerts. The URL is the credential. (commercial)
pagerdutyRoutingKeyPAGERDUTY_ROUTING_KEYPagerDuty alerts (commercial)
awsAccessKeyIdAWS_ACCESS_KEY_IDCredential broker (commercial)
awsSecretAccessKeyAWS_SECRET_ACCESS_KEYCredential broker (commercial)
databaseUrlDATABASE_URLCatalog and RBAC. Overrides config.databaseUrl.
redisUrlREDIS_URLSessions, rate limits, kill-switch. Overrides config.redisUrl.
clickhouseUrlCLICKHOUSE_URLAudit log. Overrides config.clickhouseUrl.

If a secret is unset, the chart leaves its variable out rather than setting it to an empty string, so the gateway falls back to its own default. The exception is auditHmacKey: with an existing Secret, the pod does not start until that key is present. If you pass secrets inline instead, always set secrets.auditHmacKey alongside a clickhouseUrl. The gateway logs an error at startup when it is missing.

Hardening defaults#

The chart ships with these defaults. CI renders the chart and fails the build if the security context, the probes, the resource requests and limits or the PodDisruptionBudget go missing.

  • The pod runs as non-root uid 1000 with a read-only root filesystem, no privilege escalation, all capabilities dropped and the RuntimeDefault seccomp profile. A writable emptyDir is mounted at /tmp.
  • No ServiceAccount token is mounted, because the gateway calls no Kubernetes API.
  • A PodDisruptionBudget is on, with minAvailable: 1.
  • CPU and memory requests and limits are set.
  • Liveness and readiness probes are on /healthz.
  • replicaCount is 2.

/healthz always returns a static 200 and checks no dependencies. A pod can report ready while ClickHouse is unreachable, so alert on the startup logs as well as the probes.

Network policy#

networkPolicy.enabled is false by default. Egress rules have a separate switch, networkPolicy.egress.enabled, which is also off.

  • Admit your ingress controller. With ingress.enabled, the chart refuses to render until you set networkPolicy.ingress.fromIngressController selectors or networkPolicy.ingress.fromCidrs.
  • Kubelet probes come from the node. On CNIs that police host-to-pod traffic, a default deny fails the readiness probes. Add your node or pod CIDRs to networkPolicy.ingress.fromCidrs.
  • Everything shares port 8080. A NetworkPolicy cannot scope traffic by path, so METRICS_TOKEN, the SCIM token and RBAC remain the real access controls.
  • Egress must list every destination. DNS is allowed by default. If you turn egress on, add rules in networkPolicy.egress.rules for Redis, Postgres, ClickHouse (8123), your upstream MCP servers, your IdP and every SIEM or alert sink. Splunk HEC defaults to 8088 and OTLP/HTTP to 4318. A missing ClickHouse or sink rule does not fail requests, so the only signal is a log line.
  • Allow the AWS link-local credential endpoints if the credential broker uses IMDS, EKS Pod Identity or ECS. Without them, every brokered upstream fails closed. IRSA avoids the dependency.

Autoscaling#

autoscaling.enabled is false by default. When you enable it, the HPA scales between minReplicas: 2 and maxReplicas: 10 on a 75% CPU target.

Tenancy modes#

  • tenancy.mode: shared (the default): one deployment serves every tenant. Tenants are isolated by their JWT claim at the routing, RBAC and audit layers.
  • tenancy.mode: isolated: you install one release per tenant.
bash
helm install aegis-acme deploy/helm/aegis -n tenant-acme --create-namespace \
  --set tenancy.mode=isolated \
  --set 'tenancy.tenants={acme}'

Isolated mode sets AUTH_ALLOWED_TENANTS, so the release refuses any token for another tenant, including SCIM tokens. The chart refuses to render if the tenant list is empty. The backing stores are still yours to isolate, so pair each release with its own database or schema. The optional Namespace, ResourceQuota and LimitRange objects for isolated mode come from tenant-namespace.yaml, which is a commercial module. The tenant allow-list itself is part of the MIT core.

Canary releases#

canary.enabled is false by default. When you enable it, the chart uses Flagger with provider: kubernetes, which is a blue/green rollout with no traffic split. With the loadtester enabled, a pre-rollout acceptance webhook runs before any traffic reaches the new build and confirms that unauthenticated requests to /v1/invoke, /admin/v1/* and /scim/v2/* are still refused. If any of them succeeds, the rollout is aborted.

The chart does not install or check these prerequisites: the Flagger controller, a Prometheus at canary.prometheus.address, the Prometheus Operator (for the PodMonitor) and flagger-loadtester. Setting canary.enabled=true alone fails to render, because metric gating is on by default. You also need canary.prometheus.address, podMonitor.enabled, the loadtester settings (canary.loadtester.enabled, canary.loadtester.url and canary.loadtester.load.enabled) and a scrapeable /metrics. Alternatively, set canary.metrics.enabled=false and gate on the acceptance webhook alone. Because flagger-loadtester runs the commands in the Canary spec, restrict who can patch canaries.flagger.app.

SSE streams on /v1/subscribe do not survive a traffic split, because the stream registry is per process. This does not affect the default provider, but resolve it before you switch to a traffic-shifting provider.

Before you go live#

  • Configure a JWT verifier (config.authJwksUrl, with config.authIssuer and config.authAudience). Without one, every caller runs as anonymous.
  • Set databaseUrl so that RBAC is enforced.
  • Set a strong auditHmacKey and a clickhouseUrl.
  • Set metricsToken. Without it, /metrics returns 503 unless you opt into unauthenticated scraping.
  • Terminate TLS at your ingress controller and pin the image.

See Configuration for every variable and Security model for what each control covers.