Skip to content

aegis · Concepts

Threat detection

Request and response checks aegis runs on every tool call: schema validation, injection, egress DLP, PII and secret redaction, tool integrity and shadow tools.

On this page

aegis checks every tools/call in both directions. It checks the arguments on the way out, before the upstream is contacted, so a block there means the call never reaches the MCP server. It scans the response on the way back, before your client or any SSE subscriber sees it. Every refusal is audited, with the decision deny for a schema violation and block for a threat. Threat refusals return HTTP 200 with a JSON-RPC error, so check error.code, not the status.

THREAT_MODE=monitor switches every threat check to log-only: injection and egress findings are logged and the call is forwarded, response PII handling is forced to monitor, and a changed tool is logged but not quarantined. Unset, or any other value, enforces. Input-schema validation has its own switch and ignores THREAT_MODE.

CheckWhenCodeControlled by
Input-schema validationOutbound, first-32602INPUT_VALIDATION
Integrity quarantineOutbound-32050INTEGRITY_RESWEEP_SECS
Injection scanningOutbound-32051INJECTION_SHELL_TOOLS
Egress DLPOutbound-32053EGRESS_ACTION
PII and secretsResponse-32052PII_ACTION, PII_CLASSES
Shadow and typosquatBackgroundNone, report-onlySHADOW_SWEEP_SECS

Input-schema validation (-32602)#

aegis checks the call's arguments against the inputSchema the tool declared in tools/list. It enforces type, required, properties, additionalProperties: false, enum, array items, minItems and maxItems, string minLength and maxLength, and number minimum and maximum. It is not a full JSON Schema implementation.

It is permissive where the schema is silent. An empty schema constrains nothing, extra keys are allowed unless additionalProperties is false, and unknown keywords are ignored. The regex pattern keyword is not evaluated: the schema comes from the upstream, and compiling a regex it supplies would invite ReDoS.

INPUT_VALIDATION is enforce (default, refuses with -32602), monitor (logs and forwards) or off.

Injection scanning (-32051)#

aegis scans the whole params object, not only arguments, so a payload under any other key is still seen. Every string value and every object key is checked at any depth, after lowercasing and collapsing whitespace runs to a single space, so tab and newline variants hit the same rule. Matching is plain substring search with no regular expressions. There are 43 built-in rules in three sets; INJECTION_SHELL_TOOLS is the only setting that changes which of them run.

19 shell patterns, always on. Each names a specific command, sensitive path or pipe into an interpreter: $(, rm -rf, ; rm, | sh, | bash, |python, bash -i, python -c, /dev/tcp/, /etc/passwd, /etc/shadow, .ssh/id_, .aws/credentials, /proc/self/environ, curl http and wget http, plus no-space forms of ; rm, | sh and | bash.

8 ambiguous shell patterns, per tool. A backtick, &&, ||, >>, ../, nc -, and curl or wget followed by a space. These occur in ordinary text; applied to every tool they would refuse input such as a company name containing &&. They run only on tools matched by INJECTION_SHELL_TOOLS, a comma-separated list of name globs that is empty by default:

bash
INJECTION_SHELL_TOOLS=exec_*,run_command,shell

A trailing * matches a prefix, and * alone applies the set to every tool. List only the tools that actually hand arguments to a shell.

16 prompt-injection phrases, always on. Instruction-override and exfiltration templates such as ignore previous instructions, disregard previous, forget your instructions, new instructions:, you are now, system:, reveal your system prompt, do anything now and developer mode.

Shell rules are checked across every value before prompt rules, and the first match refuses the call. The message names the category, rule and location; locations start at args, which stands for params:

json
{"jsonrpc":"2.0","error":{"code":-32051,"message":"Blocked: prompt_injection pattern 'ignore_previous' in args.arguments.message"},"id":3}

Egress DLP (-32053)#

After injection scanning, aegis runs the response detectors below over the outbound params, including keys and numbers. If a class that can block is found, the call is refused with -32053 and never reaches the upstream.

EGRESS_ACTION is block (default), monitor or off. There is no redact mode: rewriting an argument would send the upstream something the caller never sent. Classes that cannot block, such as an email address, are logged and never refuse a call. PII_CLASSES and PII_CLASSES_EXCLUDE apply here too.

PII and secrets in responses (-32052)#

aegis scans the response's result and error (message and data). The whole response is treated as untrusted, so object keys, array elements and numeric values are scanned too. Genuine base64 data: URIs are skipped so binary payloads are not altered.

ClassDefaultCan blockMatches
PRIVATE_KEYOnYesPEM private key headers
AWS_ACCESS_KEYOnYesKey IDs such as AKIA and ASIA
GCP_API_KEYOnYesKeys starting AIza
GITHUB_TOKENOnYesghp_-style and github_pat_ tokens
SLACK_TOKENOnYesxoxb--style tokens
OPENAI_KEYOnYessk- keys
JWTOnYesThree-part tokens starting eyJ
BEARER_TOKENOnYesBearer followed by a token
CREDIT_CARDOnYesVisa, Mastercard, Amex and Discover prefixes that pass the Luhn check
SSNOnYesUS format, never-issued ranges rejected
EMAILOnNoEmail addresses
PHONEOpt-inNoNorth American and E.164-style numbers
DATE_OF_BIRTHOpt-inNoYYYY-MM-DD dates

PII_ACTION decides what happens on a hit:

  • redact (default): each hit becomes [REDACTED:KIND], for example [REDACTED:SSN], and the cleaned response is returned.
  • block: hits are masked, and if any hit is a class that can block, the whole response is refused with -32052. The upstream has already run the tool; a block withholds the result, it does not undo the call.
  • monitor: findings are logged and the response is returned unchanged.

PII_CLASSES replaces the default set (PII_CLASSES=SSN,CREDIT_CARD,PHONE, or all for every class). PII_CLASSES_EXCLUDE then subtracts, for example EMAIL for a search-by-email tool. Unknown names are logged and ignored.

A second pass joins adjacent values to catch a secret split across fields that a client would re-join when rendering. It is always logged and, under block, refused when the class can block. Under redact it cannot be masked cleanly, so it is logged, not masked. Use PII_ACTION=block if split secrets must be stopped.

Tool integrity and rug-pulls (-32050)#

The first time aegis discovers a tool, it pins a SHA-256 hash of the definition (name, description and inputSchema, keys sorted) per tenant, server and tool. Every later listing is re-hashed and compared. A background sweep re-lists all upstreams every INTEGRITY_RESWEEP_SECS (default 15, 0 disables), which catches a tool changed silently with no catalog change.

A changed tool is quarantined: hidden from tools/list and refused with -32050. Quarantine is sticky, so reverting the upstream does not release it and a flapping definition cannot toggle the block. Under THREAT_MODE=monitor the change is logged but not quarantined. With alerting (commercial) configured, a detected change also raises an alert in either mode.

List held tools, then release one after you have reviewed its new definition. $TOKEN is an admin token:

bash
curl -s http://localhost:8080/admin/v1/quarantine \
  -H "Authorization: Bearer $TOKEN"

curl -s -X POST http://localhost:8080/admin/v1/quarantine \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"server":"mock-primary","tool":"echo"}'

The listing returns quarantined, an array of server and tool pairs. Listing needs catalog:read on aegis-admin; releasing needs catalog:write and is audited. Both endpoints are scoped to your tenant and need RBAC, so DATABASE_URL must be set. A release re-pins the current definition as the new baseline and the tool is callable immediately. Releasing a tool that is not held returns 404.

Shadow and typosquat tools#

aegis compares tool names across the servers in each tenant. Names are folded first: lowercased, underscores and hyphens removed, zero-width characters dropped, and fullwidth characters and common Cyrillic and Greek look-alikes mapped to ASCII. Two kinds are reported:

  • exact: folded names are equal, so beta_echo and BetaEcho collide.
  • typosquat: both names are 5 or more characters and within Levenshtein distance 2, such as beta_ech0 and beta_echo. Version and plural pairs (query_v1 and query_v2, list_issue and list_issues) are skipped.

The sweep runs at startup, after each catalog change and every SHADOW_SWEEP_SECS (default 60, 0 disables the periodic sweep). Findings are at GET /admin/v1/shadows (catalog:read), scoped to your tenant, as a shadows array with the fields kind, tenant, server, tool, shadows_server, shadows_tool and distance. Within a tenant, servers are compared in name order, and the one whose name sorts later is reported as server, the likely offender; review both sides.

This check never blocks. Servers legitimately share names like read_file or search, and near matches include benign variants. With alerting (commercial) configured, new findings also raise a deduplicated alert.

Upstream credential reflection#

When the credential broker (commercial) injects a credential, aegis remembers the exact secret it sent. Before the PII scan, every occurrence in the response's result, error message, error data and object keys is replaced with [REDACTED:UPSTREAM_CREDENTIAL]. Exact matching catches opaque tokens the format detectors would miss. This applies to tools/call responses, including what SSE subscribers receive. If the upstream call fails, the same credential is scrubbed from the error text before it reaches your client or the audit log, and that text is capped at 300 characters.