MCP Client
Wrap your Model Context Protocol client with Shield. Tool lists are checked for tool poisoning and flagged tools are dropped; tool results, resources, and prompts are checked for injection before they reach the model; and tool calls can be checked against a tool policy.
The shieldMcpClient wrapper adds Shield to your Client from @modelcontextprotocol/sdk. It checks the tools a server lists for tool poisoning and excludes flagged tools. It also checks content returned by the server's tools, resources, and prompts for injection before your agent passes that content to a model. With a tool policy, it refuses tool calls the policy doesn't allow.
Usage
import { Client } from "@modelcontextprotocol/sdk/client/index.js";
import { StreamableHTTPClientTransport } from "@modelcontextprotocol/sdk/client/streamableHttp.js";
import { shieldMcpClient } from "@zeroleaks/shield/mcp";
const client = shieldMcpClient(new Client({ name: "agent", version: "1.0.0" }), {
onToolFlagged: (tool, result) => {
console.warn(`Dropped MCP tool ${tool.name}`, result.issues, result.result.matches);
},
});
await client.connect(new StreamableHTTPClientTransport(new URL(serverUrl)));
const { tools } = await client.listTools(); // flagged tools left out
const result = await client.callTool({ name: "search", arguments: { query } }); // checkedHow it works
- On every call to
listTools, Shield runsscanTools()on the tools the server returned (unlessscanTools: false). A tool is flagged when detection finds instructions in its description or input schema, or when it has an issue: a duplicate name, hidden characters in its name, a description longer thanmaxDescriptionLength(4,000 characters by default), or a definition that changed since the client first saw it (see Pinning).onToolFlaggedis called for each flagged tool, andonFlaggedToolsdecides what happens to it. See Tool lists. - On every call to
callTool, before the server is called, Shield refuses a tool the latestlistToolsflagged, checks the arguments for credentials and exfiltration links, and, withpolicy, checks the call against the tool policy (see Tool calls). After the server returns, Shield runs detection on the text of the result (unlessscanToolResults: false): its text blocks, embedded resources, resource link titles and descriptions, the string values ofstructuredContent, and a legacytoolResult. Results withisError: trueare checked too. - On every call to
readResource, it runs detection on the text of every entry incontents. - On every call to
getPrompt, it runs detection on the prompt's description and the content of every message, whatever its role.
In steps 2 to 4, the text is joined and checked as one tool result, with the tool result options. An embedded resource or resource content given as a base64 blob is decoded and read when its MIME type is text/*, application/json, or application/xml, up to its first 64KB. With onDetection: "block", an injection throws InjectionDetectedError with source: "tool" and the result is not returned. With "warn", onInjectionDetected(result, "tool") is called and the result is returned. Either way the result is the SDK's own object, unchanged.
Result detection runs after the server has answered. When this check blocks a result, the tool call has already run on the server; Shield prevents its result from reaching the model. A secondaryDetector in the detect options runs before a result is blocked, and detection results are reused for text already checked, as described under Shared behavior.
Tool lists
onFlaggedTools sets what happens to flagged tools:
| Mode | Behavior |
|---|---|
"drop" (default) | Returns a copy of the result without the flagged tools. nextCursor and every other field are kept. |
"throw" | Throws one InjectionDetectedError with source: "tool" for every flagged tool in the list. Its risk is the highest risk detection found, or "low" when the tools were flagged only for issues, and its categories hold the matched categories and the issue names. |
"warn" | Returns the list unchanged. |
onToolFlagged(tool, result) is called for each flagged tool in every mode, with the tool definition and its scanTools() result. When nothing is flagged, you get the SDK's result object itself. onDetection and onInjectionDetected don't apply to tool lists.
Each call checks only the page of tools it returns, so duplicate names across two pages are not detected.
Pinning
A server can change a tool's description after you reviewed it or the model began using it, adding instructions without changing its name. The first time listTools returns a tool unflagged, the client pins its definition: the description, title, input and output schemas, and annotations, compared after sorting keys. When a later list returns a different definition for that name, the tool is flagged with the issue changed_since_pinned and handled like any flagged tool. A changed tool keeps its original pin, so it stays flagged until you accept the change.
By default, pins last for the lifetime of the wrapped client. To keep them across restarts, pass your own object and save it:
import { readFile, writeFile } from "node:fs/promises";
const pins = JSON.parse(await readFile("mcp-pins.json", "utf8").catch(() => "{}"));
const client = shieldMcpClient(mcp, { pins });
const { tools } = await client.listTools(); // new tools are added to pins
await writeFile("mcp-pins.json", JSON.stringify(pins));To accept a changed tool, delete its entry from pins. pins: false turns pinning off. Without the wrapper, pinTools() and the pins option of scanTools() do the same.
Tool calls
Before callTool reaches the server, the wrapper:
- Refuses a tool that the latest
listToolsflagged, by throwingInjectionDetectedErrorwithsource: "tool"and the tool's risk and issues. Dropping a tool removes it from the list you give the model. This check also blocks calls by name, including calls from a model that saw an earlier list or guessed the name. A later list that no longer flags the tool allows it again. This check is enabled unlessonFlaggedToolsis"warn"orblockFlaggedToolCallsisfalse. - Runs
scanOutputText()on the call's arguments and throwsOutputBlockedErrorfor any high or critical finding, such as a credential or a markdown image that carries data to another server. Lower findings are not reported. PassscanArgumentsto change the detectors, such as{ pii: true }orexfiltration: { allowedDomains: [...] }, orfalseto skip it. - With
policy, runspolicy.checkAsync()on the call and throwsToolPolicyErrorwhen the policy refuses it. See Tool policy.
Tool policy
Pass a policy from createToolPolicy() as policy, and the wrapper applies it to the server's tools:
import { createToolPolicy } from "@zeroleaks/shield";
const policy = createToolPolicy({
deny: ["delete_*"],
rules: {
get_issue: { labels: ["untrusted"] },
create_issue: { labels: ["sink"], maxCalls: 3 },
},
});
const client = shieldMcpClient(mcp, { policy });
const { tools } = await client.listTools(); // declared to the policy
await client.callTool({ name: "get_issue", arguments: { number: 7 } }); // recorded as untrusted
await client.callTool({ name: "create_issue", arguments: { title, body } }); // throws ToolPolicyError- Declared tools. Every
listToolsdeclares the tools it returns to the policy, as this client's own: a call through this client must name one of them, and its arguments must match that tool'sinputSchema. A tool the list dropped is not declared, so a call to it is refused even withblockFlaggedToolCalls: false. A page fetched with acursoradds to the pages before it, and a list without a cursor starts over. Before the firstlistTools, a call is checked against the tools other sources declared, if any. - The check.
callToolrunspolicy.checkAsync()after the other checks above, so a call they block is not counted. When the policy refuses a call, Shield throwsToolPolicyErrorwith the decision'stool,reason, andviolationswithout calling the server. Shield awaits anapprovecallback that returns a Promise. - Results. After the server answers, the result is recorded with
policy.recordResult(), flagged when detection found an injection in it, whether the result is then blocked (onDetection: "block") or returned ("warn"). A resource or prompt with an injection is recorded withpolicy.recordUntrusted(), asreadResourceorgetPrompt. WithscanToolResults: false, results are recorded by their labels only.
A policy holds one session's state, while an MCP client often serves many sessions. Wrap the shared Client once per session with a separate policy. Each wrapper is a Proxy that shares the connection. You can pass one policy to the clients of several servers, so content read from one server blocks a sink on another.
When a server's tools change
A server that sends notifications/tools/list_changed can change its tools after you listed them, so list them again through the wrapped client. The SDK's listChanged option refreshes the list automatically by default, but it calls the unwrapped client, so the tools it passes to onChanged are not checked. Disable that refresh and call the wrapped client's listTools() instead:
const mcp = new Client(
{ name: "agent", version: "1.0.0" },
{
listChanged: {
tools: {
autoRefresh: false,
onChanged: async () => {
const { tools } = await client.listTools(); // checked
setTools(tools);
},
},
},
}
);
const client = shieldMcpClient(mcp);Options
shieldMcpClient takes detect, scanToolResults, onDetection, and onInjectionDetected from the shared options. There are no user messages to check, so detect only sets the options that scanToolResults: true (the default) uses. The other shared options don't apply: nothing is hardened, and the wrapper doesn't guard model output. It also takes:
| Option | Type | Default | Description |
|---|---|---|---|
scanTools | ScanToolsOptions | false | the tool result detect options | scanTools() options for tool lists, such as maxDescriptionLength and any detect option, or false to leave tool lists unchecked. Defaults to the scanToolResults options when that is an object, else detect. |
onFlaggedTools | "drop" | "throw" | "warn" | "drop" | What happens to flagged tools. See Tool lists. |
onToolFlagged | (tool: ToolDefinition, result: ToolScanResult) => void | none | Called with each flagged tool and its scan result, in every mode |
pins | ToolPins | false | pins kept for the life of the client | Pinned tool definitions. Pass an object to keep pins across sessions; the client adds new tools to it. See Pinning. |
blockFlaggedToolCalls | boolean | true, except with onFlaggedTools: "warn" | Refuse callTool for a tool the latest listTools flagged. See Tool calls. |
scanArguments | ScanOutputOptions | false | secrets and exfiltration | scanOutputText() options for tool call arguments, or false. See Tool calls. |
policy | ToolPolicy | none | A policy from createToolPolicy() for this session. Listed tools are declared to it, calls are checked against it, and results are recorded in it. See Tool policy. |
What is not covered
- The wrapped client is a Proxy.
shieldMcpClientreturns a Proxy over your client, withlistTools,callTool,readResource, andgetPromptreplaced. Every other property reads from your client, and its methods are bound to it, soconnect(),close(), and the rest keep working. Your client is not modified, but setting a property on the wrapped client sets it on yours. - Other methods:
listResources(),listResourceTemplates(),listPrompts(),complete(),experimental.tasks.callToolStream(), and rawrequest()calls are not checked. - Requests from the server: sampling (
sampling/createMessage, which carries messages and a system prompt written by the server) and elicitation requests reach the handlers you register withsetRequestHandler(). Check them there withdetect(). Notifications, such as log messages, are not checked either. - Tool call arguments are checked for high and critical output findings only, not for injection or prompt leaks. This check runs when your code makes the call. For model-generated arguments, also use your model provider's wrapper to guard them as they are produced. A tool policy checks them against the tool's schema, but only for the keywords listed on its page.
- Content that is not read: images, audio, blobs with other MIME types, and the resources that resource links point to, which are not fetched.
- The SDK's automatic list refresh, as described in When a server's tools change.
See What is not scanned for what no integration scans.
LangChain.js
shieldChatModel and ShieldCallbackHandler for LangChain.js chat models, chains, and agents. Hardening, injection detection in human and tool messages, and output guarding on invoke, batch, stream, and streamEvents.
OpenAI Agents SDK
Shield guardrails for the OpenAI Agents SDK. The run's input and every tool output are checked for injection, the final output and tool call arguments for prompt leaks, credentials, exfiltration links, and canaries, and tool calls against a tool policy.