ZeroLeaksDocs
Shield SDKProvider Wrappers

MCP Client

Wrap your Model Context Protocol client with Shield. Tool lists are checked for tool poisoning and flagged tools are dropped; tool results, resources, and prompts are checked for injection before they reach the model; and tool calls can be checked against a tool policy.

The shieldMcpClient wrapper adds Shield to your Client from @modelcontextprotocol/sdk. It checks the tools a server lists for tool poisoning and excludes flagged tools. It also checks content returned by the server's tools, resources, and prompts for injection before your agent passes that content to a model. With a tool policy, it refuses tool calls the policy doesn't allow.

Usage

import { Client } from "@modelcontextprotocol/sdk/client/index.js";
import { StreamableHTTPClientTransport } from "@modelcontextprotocol/sdk/client/streamableHttp.js";
import { shieldMcpClient } from "@zeroleaks/shield/mcp";

const client = shieldMcpClient(new Client({ name: "agent", version: "1.0.0" }), {
  onToolFlagged: (tool, result) => {
    console.warn(`Dropped MCP tool ${tool.name}`, result.issues, result.result.matches);
  },
});
await client.connect(new StreamableHTTPClientTransport(new URL(serverUrl)));

const { tools } = await client.listTools(); // flagged tools left out
const result = await client.callTool({ name: "search", arguments: { query } }); // checked

How it works

  1. On every call to listTools, Shield runs scanTools() on the tools the server returned (unless scanTools: false). A tool is flagged when detection finds instructions in its description or input schema, or when it has an issue: a duplicate name, hidden characters in its name, a description longer than maxDescriptionLength (4,000 characters by default), or a definition that changed since the client first saw it (see Pinning). onToolFlagged is called for each flagged tool, and onFlaggedTools decides what happens to it. See Tool lists.
  2. On every call to callTool, before the server is called, Shield refuses a tool the latest listTools flagged, checks the arguments for credentials and exfiltration links, and, with policy, checks the call against the tool policy (see Tool calls). After the server returns, Shield runs detection on the text of the result (unless scanToolResults: false): its text blocks, embedded resources, resource link titles and descriptions, the string values of structuredContent, and a legacy toolResult. Results with isError: true are checked too.
  3. On every call to readResource, it runs detection on the text of every entry in contents.
  4. On every call to getPrompt, it runs detection on the prompt's description and the content of every message, whatever its role.

In steps 2 to 4, the text is joined and checked as one tool result, with the tool result options. An embedded resource or resource content given as a base64 blob is decoded and read when its MIME type is text/*, application/json, or application/xml, up to its first 64KB. With onDetection: "block", an injection throws InjectionDetectedError with source: "tool" and the result is not returned. With "warn", onInjectionDetected(result, "tool") is called and the result is returned. Either way the result is the SDK's own object, unchanged.

Result detection runs after the server has answered. When this check blocks a result, the tool call has already run on the server; Shield prevents its result from reaching the model. A secondaryDetector in the detect options runs before a result is blocked, and detection results are reused for text already checked, as described under Shared behavior.

Tool lists

onFlaggedTools sets what happens to flagged tools:

ModeBehavior
"drop" (default)Returns a copy of the result without the flagged tools. nextCursor and every other field are kept.
"throw"Throws one InjectionDetectedError with source: "tool" for every flagged tool in the list. Its risk is the highest risk detection found, or "low" when the tools were flagged only for issues, and its categories hold the matched categories and the issue names.
"warn"Returns the list unchanged.

onToolFlagged(tool, result) is called for each flagged tool in every mode, with the tool definition and its scanTools() result. When nothing is flagged, you get the SDK's result object itself. onDetection and onInjectionDetected don't apply to tool lists.

Each call checks only the page of tools it returns, so duplicate names across two pages are not detected.

Pinning

A server can change a tool's description after you reviewed it or the model began using it, adding instructions without changing its name. The first time listTools returns a tool unflagged, the client pins its definition: the description, title, input and output schemas, and annotations, compared after sorting keys. When a later list returns a different definition for that name, the tool is flagged with the issue changed_since_pinned and handled like any flagged tool. A changed tool keeps its original pin, so it stays flagged until you accept the change.

By default, pins last for the lifetime of the wrapped client. To keep them across restarts, pass your own object and save it:

import { readFile, writeFile } from "node:fs/promises";

const pins = JSON.parse(await readFile("mcp-pins.json", "utf8").catch(() => "{}"));
const client = shieldMcpClient(mcp, { pins });
const { tools } = await client.listTools(); // new tools are added to pins
await writeFile("mcp-pins.json", JSON.stringify(pins));

To accept a changed tool, delete its entry from pins. pins: false turns pinning off. Without the wrapper, pinTools() and the pins option of scanTools() do the same.

Tool calls

Before callTool reaches the server, the wrapper:

  1. Refuses a tool that the latest listTools flagged, by throwing InjectionDetectedError with source: "tool" and the tool's risk and issues. Dropping a tool removes it from the list you give the model. This check also blocks calls by name, including calls from a model that saw an earlier list or guessed the name. A later list that no longer flags the tool allows it again. This check is enabled unless onFlaggedTools is "warn" or blockFlaggedToolCalls is false.
  2. Runs scanOutputText() on the call's arguments and throws OutputBlockedError for any high or critical finding, such as a credential or a markdown image that carries data to another server. Lower findings are not reported. Pass scanArguments to change the detectors, such as { pii: true } or exfiltration: { allowedDomains: [...] }, or false to skip it.
  3. With policy, runs policy.checkAsync() on the call and throws ToolPolicyError when the policy refuses it. See Tool policy.

Tool policy

Pass a policy from createToolPolicy() as policy, and the wrapper applies it to the server's tools:

import { createToolPolicy } from "@zeroleaks/shield";

const policy = createToolPolicy({
  deny: ["delete_*"],
  rules: {
    get_issue: { labels: ["untrusted"] },
    create_issue: { labels: ["sink"], maxCalls: 3 },
  },
});
const client = shieldMcpClient(mcp, { policy });
const { tools } = await client.listTools(); // declared to the policy
await client.callTool({ name: "get_issue", arguments: { number: 7 } }); // recorded as untrusted
await client.callTool({ name: "create_issue", arguments: { title, body } }); // throws ToolPolicyError
  • Declared tools. Every listTools declares the tools it returns to the policy, as this client's own: a call through this client must name one of them, and its arguments must match that tool's inputSchema. A tool the list dropped is not declared, so a call to it is refused even with blockFlaggedToolCalls: false. A page fetched with a cursor adds to the pages before it, and a list without a cursor starts over. Before the first listTools, a call is checked against the tools other sources declared, if any.
  • The check. callTool runs policy.checkAsync() after the other checks above, so a call they block is not counted. When the policy refuses a call, Shield throws ToolPolicyError with the decision's tool, reason, and violations without calling the server. Shield awaits an approve callback that returns a Promise.
  • Results. After the server answers, the result is recorded with policy.recordResult(), flagged when detection found an injection in it, whether the result is then blocked (onDetection: "block") or returned ("warn"). A resource or prompt with an injection is recorded with policy.recordUntrusted(), as readResource or getPrompt. With scanToolResults: false, results are recorded by their labels only.

A policy holds one session's state, while an MCP client often serves many sessions. Wrap the shared Client once per session with a separate policy. Each wrapper is a Proxy that shares the connection. You can pass one policy to the clients of several servers, so content read from one server blocks a sink on another.

When a server's tools change

A server that sends notifications/tools/list_changed can change its tools after you listed them, so list them again through the wrapped client. The SDK's listChanged option refreshes the list automatically by default, but it calls the unwrapped client, so the tools it passes to onChanged are not checked. Disable that refresh and call the wrapped client's listTools() instead:

const mcp = new Client(
  { name: "agent", version: "1.0.0" },
  {
    listChanged: {
      tools: {
        autoRefresh: false,
        onChanged: async () => {
          const { tools } = await client.listTools(); // checked
          setTools(tools);
        },
      },
    },
  }
);
const client = shieldMcpClient(mcp);

Options

shieldMcpClient takes detect, scanToolResults, onDetection, and onInjectionDetected from the shared options. There are no user messages to check, so detect only sets the options that scanToolResults: true (the default) uses. The other shared options don't apply: nothing is hardened, and the wrapper doesn't guard model output. It also takes:

OptionTypeDefaultDescription
scanToolsScanToolsOptions | falsethe tool result detect optionsscanTools() options for tool lists, such as maxDescriptionLength and any detect option, or false to leave tool lists unchecked. Defaults to the scanToolResults options when that is an object, else detect.
onFlaggedTools"drop" | "throw" | "warn""drop"What happens to flagged tools. See Tool lists.
onToolFlagged(tool: ToolDefinition, result: ToolScanResult) => voidnoneCalled with each flagged tool and its scan result, in every mode
pinsToolPins | falsepins kept for the life of the clientPinned tool definitions. Pass an object to keep pins across sessions; the client adds new tools to it. See Pinning.
blockFlaggedToolCallsbooleantrue, except with onFlaggedTools: "warn"Refuse callTool for a tool the latest listTools flagged. See Tool calls.
scanArgumentsScanOutputOptions | falsesecrets and exfiltrationscanOutputText() options for tool call arguments, or false. See Tool calls.
policyToolPolicynoneA policy from createToolPolicy() for this session. Listed tools are declared to it, calls are checked against it, and results are recorded in it. See Tool policy.

What is not covered

  • The wrapped client is a Proxy. shieldMcpClient returns a Proxy over your client, with listTools, callTool, readResource, and getPrompt replaced. Every other property reads from your client, and its methods are bound to it, so connect(), close(), and the rest keep working. Your client is not modified, but setting a property on the wrapped client sets it on yours.
  • Other methods: listResources(), listResourceTemplates(), listPrompts(), complete(), experimental.tasks.callToolStream(), and raw request() calls are not checked.
  • Requests from the server: sampling (sampling/createMessage, which carries messages and a system prompt written by the server) and elicitation requests reach the handlers you register with setRequestHandler(). Check them there with detect(). Notifications, such as log messages, are not checked either.
  • Tool call arguments are checked for high and critical output findings only, not for injection or prompt leaks. This check runs when your code makes the call. For model-generated arguments, also use your model provider's wrapper to guard them as they are produced. A tool policy checks them against the tool's schema, but only for the keywords listed on its page.
  • Content that is not read: images, audio, blobs with other MIME types, and the resources that resource links point to, which are not fetched.
  • The SDK's automatic list refresh, as described in When a server's tools change.

See What is not scanned for what no integration scans.

On this page