ZeroLeaksDocs
Shield SDKProvider Wrappers

OpenAI Agents SDK

Shield guardrails for the OpenAI Agents SDK. The run's input and every tool output are checked for injection, the final output and tool call arguments for prompt leaks, credentials, exfiltration links, and canaries, and tool calls against a tool policy.

The @zeroleaks/shield/openai-agents module provides five guardrails for the OpenAI Agents SDK (@openai/agents). Register them in the SDK's guardrail slots on an agent, a Runner, a function tool, or an MCP server:

GuardrailSlotChecksOn a finding
shieldInputGuardrail()inputGuardrails of an Agent or RunnerThe run's input for injectionrun() rejects with InputGuardrailTripwireTriggered
shieldOutputGuardrail()outputGuardrails of an Agent or RunnerThe final output for prompt leaks, credentials, exfiltration links, and canariesrun() rejects with OutputGuardrailTripwireTriggered
shieldToolOutputGuardrail()outputGuardrails of a tool(), toolOutputGuardrails of an MCP serverWhat the tool returned, for injection, before the model reads itThe model gets a rejection message instead of the output
shieldToolInputGuardrail()inputGuardrails of a tool(), toolInputGuardrails of an MCP serverThe tool call's arguments, before the tool runs, for credentials, exfiltration links, canaries, and system prompt textThe tool doesn't run, and the model gets a rejection message
shieldToolPolicyGuardrail()inputGuardrails of a tool(), toolInputGuardrails of an MCP serverThe tool call, before the tool runs, against a tool policyThe tool doesn't run, and the model gets the policy's reason

Guardrails can stop a run or a tool call, but they can't change text. Nothing is redacted and instructions are not hardened. Content with a finding either trips the guardrail or passes through unchanged.

Usage

import {
  Agent,
  InputGuardrailTripwireTriggered,
  OutputGuardrailTripwireTriggered,
  run,
  tool,
} from "@openai/agents";
import { z } from "zod";
import {
  shieldInputGuardrail,
  shieldOutputGuardrail,
  shieldToolInputGuardrail,
  shieldToolOutputGuardrail,
} from "@zeroleaks/shield/openai-agents";

const readInbox = tool({
  name: "read_inbox",
  description: "Read new email.",
  parameters: z.object({}),
  execute: async () => fetchInbox(),
  outputGuardrails: [shieldToolOutputGuardrail()],
});

const sendEmail = tool({
  name: "send_email",
  description: "Send an email.",
  parameters: z.object({ to: z.string(), body: z.string() }),
  execute: async ({ to, body }) => send(to, body),
  inputGuardrails: [shieldToolInputGuardrail()],
});

const agent = new Agent({
  name: "Support",
  instructions: "You are a support agent for Acme. Never share internal ticket notes.",
  tools: [readInbox, sendEmail],
  inputGuardrails: [shieldInputGuardrail()],
  outputGuardrails: [shieldOutputGuardrail()],
});

try {
  const result = await run(agent, userInput);
  return result.finalOutput;
} catch (error) {
  if (
    error instanceof InputGuardrailTripwireTriggered ||
    error instanceof OutputGuardrailTripwireTriggered
  ) {
    console.warn(error.result.guardrail.name, error.result.output.outputInfo);
    return "Sorry, I can't help with that.";
  }
  throw error;
}

To guard every tool exposed by an MCP server, pass the tool guardrails to the server (Agents SDK 0.17.1 or later):

import { MCPServerStreamableHttp } from "@openai/agents";

const server = new MCPServerStreamableHttp({
  url: serverUrl,
  toolInputGuardrails: [shieldToolInputGuardrail()],
  toolOutputGuardrails: [shieldToolOutputGuardrail()],
});

To guard every agent a Runner runs, pass the input and output guardrails to it instead: new Runner({ inputGuardrails: [shieldInputGuardrail()], outputGuardrails: [shieldOutputGuardrail()] }).

How it works

Input guardrail

shieldInputGuardrail() runs detection on the run's input:

  • A string input is checked as a user message.
  • In a list of input items, the string content or input_text parts of each user message are checked as a user message, with the detect options. Tool results in the list are checked as tool results, with the tool result options: the output of function_call_result items (read like a tool output, as described below), the stdout and stderr of shell_call_output items, and the output of apply_patch_call_output, program_output, and hosted_tool_call items. A list has tool results when you pass the history of an earlier run, or when a session adds its stored items, which the guardrail sees along with the new input. With onDetection: "warn", an injection that stays in the history is reported on every run.
  • System and assistant messages are not checked.

With onDetection: "block" (the default), the first injection trips the guardrail and run() rejects with the SDK's InputGuardrailTripwireTriggered. Its result.output.outputInfo is { detected: true, source, result }, where source is "user" or "tool" and result is the LocalDetectResult, which holds the risk, the matched categories and patterns, and the classifier score, but not the text. With "warn", onInjectionDetected(result, source) is called for each injection and the run continues; the guardrail's entry in result.inputGuardrailResults still has detected: true.

The SDK runs an input guardrail alongside the agent's first model call unless the guardrail sets runInParallel: false. Shield sets it to false, so the input is checked before the model is called, and a blocked input never reaches the model or a tool. Pass runInParallel: true to reduce latency at the cost of this guarantee, for example with a slow secondaryDetector or escalate detector. With stream: true, run() resolves with the stream, and the error comes from result.completed and the stream.

The SDK runs input guardrails once per run, on the input to the first agent. They don't run again after a handoff or on later turns. Use the tool output guardrail to check tool results produced during the run.

Output guardrail

shieldOutputGuardrail() checks the agent's final output: a string, or every string in a structured output (outputType), keys included. It looks for:

  • A leak of the system prompt: systemPrompt if you pass it, else the agent's instructions when they are a string. With instructions given as a function, or a stored prompt, pass systemPrompt, or the prompt leak check is skipped.
  • Output findings from scanOutputText(): credentials and exfiltration links by default, and personal data with output: { pii: true }.
  • A canary you pass in canary or output.canary.

A prompt leak or a canary trips the guardrail, unless throwOnLeak is false, and so does a high or critical output finding, unless blockOnOutputFindings is false. run() then rejects with the SDK's OutputGuardrailTripwireTriggered. onLeakDetected and onOutputFindings are called whether or not it trips.

outputInfo is { leak, findings }: leak holds the leak's confidence and fragmentCount, and findings holds the type, kind, and severity of every output finding when one of them is high or critical. It never contains the text. The SDK includes it in the error message, so the message is safe to log. The error's result.agentOutput contains the SDK's unredacted copy of the output that tripped the guardrail. Do not log that field.

The SDK runs the output guardrails of the agent that produced the final output, plus those of the Runner. If a handoff occurred, it uses the last agent's guardrails. With stream: true, the guardrail runs after the model finishes and the text has already been streamed. The error is thrown from your for await loop and result.completed after the last chunk.

Tool output guardrail

shieldToolOutputGuardrail() runs after a function tool returns and before its output reaches the model. It checks the output's text with the tool result options:

  • a string, whole;
  • a text part (text, input_text, or output_text);
  • an MCP content block: text, an embedded resource (its text, or a text/*, application/json, or application/xml blob, decoded, up to 64KB), or a resource link's title and description;
  • a list of those;
  • the string values of anything else, not its keys, up to 64K characters.

With onDetection: "block" and the default behavior: "rejectContent", an injection causes the SDK to replace the tool's output with rejectionMessage, and the run continues. With behavior: "throwException", run() rejects with the SDK's ToolCallError, whose error is ToolOutputGuardrailTripwireTriggered. With "warn", onInjectionDetected(result, "tool") is called and the output reaches the model. outputInfo has the same shape as the input guardrail's, and each result is in result.toolOutputGuardrailResults.

The tool has already run when its output is checked. A blocked result does not reach the model.

With policy, each output is also recorded in the tool policy with policy.recordResult(), flagged when detection found an injection in it, whether or not the guardrail rejects it. See Tool policy guardrail.

Tool input guardrail

shieldToolInputGuardrail() checks the arguments generated by the model before a function tool runs. It parses them as JSON and checks every string, keys included, with the same checks as the output guardrail: text of the system prompt, credentials, exfiltration links, and the canary. An agent induced to put a credential, the canary, or its instructions into a tool call is stopped before the tool runs.

Only high and critical findings trip it. A markdown or HTML image that sends data out is critical, but a bare URL that carries data in its query or path, the usual shape of an argument to a fetch tool, is most often a medium url_data finding: it is reported to onOutputFindings, and the tool runs.

A prompt leak or a canary (unless throwOnLeak: false) or a high or critical finding (unless blockOnOutputFindings: false) trips the guardrail. With the default behavior: "rejectContent", the tool doesn't run and the model gets rejectionMessage as its output. With "throwException", run() rejects with the SDK's ToolCallError, whose error is ToolInputGuardrailTripwireTriggered. outputInfo has the same shape as the output guardrail's.

Tool policy guardrail

shieldToolPolicyGuardrail({ policy }) runs policy.checkAsync() on each call before the tool runs, with the call's name and its JSON arguments. A policy from createToolPolicy() remembers what a session has read, so give each run its own: pass a function of the context you pass to run(), and give the same function to shieldToolOutputGuardrail({ policy }), which records each output in the policy:

import { createToolPolicy, type ToolPolicy } from "@zeroleaks/shield";
import {
  shieldToolOutputGuardrail,
  shieldToolPolicyGuardrail,
} from "@zeroleaks/shield/openai-agents";

type Session = { policy: ToolPolicy };
const policyGuardrail = shieldToolPolicyGuardrail<Session>({ policy: (context) => context.policy });
const outputGuardrail = shieldToolOutputGuardrail<Session>({ policy: (context) => context.policy });

const sendEmail = tool({
  name: "send_email",
  description: "Send an email.",
  parameters: z.object({ to: z.string(), body: z.string() }),
  execute: async ({ to, body }) => send(to, body),
  inputGuardrails: [policyGuardrail],
  outputGuardrails: [outputGuardrail],
});

await run(agent, userInput, {
  context: { policy: createToolPolicy({ rules: { send_email: { labels: ["sink"] } } }) },
});

To use one policy throughout a conversation, store it in the context object passed to every run() in that conversation. A policy object passed directly, instead of a function, is shared by every run that uses the guardrail. A function that returns something other than a policy fails the run with a TypeError.

With the default behavior: "rejectContent", a refused call doesn't run, and the model gets the decision's message, or rejectionMessage if you set one, as the tool's output. With "throwException", run() rejects with the SDK's ToolCallError, whose error is ToolInputGuardrailTripwireTriggered. outputInfo is the policy's decision ({ allowed, reason, message, tool, violations? }), which never holds argument values; it is in result.toolInputGuardrailResults.

The policy counts every call it allows. Place this guardrail last in inputGuardrails, after shieldToolInputGuardrail(), so calls rejected by another guardrail are not counted. The SDK already refuses tool names the agent doesn't have. The policy's declared tools therefore matter for tools you leave out of allow, and for argument schemas the SDK doesn't enforce, as with a JSON Schema tool that isn't strict.

Options

shieldInputGuardrail and shieldToolOutputGuardrail take detect, scanToolResults, onDetection, and onInjectionDetected from the shared options. shieldToolOutputGuardrail only checks tool output, so detect only sets the options that scanToolResults: true (the default) uses. Detection results are reused for text already checked, and a secondaryDetector or escalate detector runs before a guardrail trips, as described under Shared behavior.

shieldOutputGuardrail and shieldToolInputGuardrail take systemPrompt, sanitize, output, onLeakDetected, and onOutputFindings from the shared options. shieldToolPolicyGuardrail takes none of the shared options.

The guardrails also take:

OptionGuardrailsTypeDefaultDescription
nameallstring"shield_input", "shield_output", "shield_tool_input", "shield_tool_output", "shield_tool_policy"The name the SDK shows in errors, results, and traces
runInParallelinputbooleanfalseRun alongside the first model call instead of before it. See Input guardrail.
canaryoutput, tool inputstringnoneA canary token you planted in the instructions. Its appearance counts as a leak. It must be a string: these guardrails can't plant one.
throwOnLeakoutput, tool inputbooleantrueTrip on a prompt leak or a canary. The shared option defaults to false, since the wrappers can redact instead.
blockOnOutputFindingsoutput, tool inputbooleantrueTrip on a high or critical output finding. The shared option defaults to false for the same reason.
behaviortool input, tool output, tool policy"rejectContent" | "throwException""rejectContent"What a tripped tool guardrail does: give the model rejectionMessage in place of the output, or fail the run
rejectionMessagetool input, tool output, tool policystringsee belowWhat the model gets in place of the tool's output
policytool policy (required), tool outputToolPolicy | (context) => ToolPolicynoneThe tool policy, or a function that returns it from the context passed to run(). The tool policy guardrail checks calls against it; the tool output guardrail records outputs in it. See Tool policy guardrail.

The default rejection messages are "This tool output was withheld because Shield flagged it as a possible prompt injection." for tool output, and "This tool call was blocked because Shield found a credential, an exfiltration link, a canary, or system prompt text in its arguments." for tool calls. The tool policy guardrail's default is the decision's message, such as Tool "send_email" can send data out, and this session has seen untrusted content from read_inbox.

To use a canary, put it in the instructions yourself and pass the same token to the guardrails:

import { canaryInstruction, createCanary } from "@zeroleaks/shield";

const canary = createCanary();
const agent = new Agent({
  name: "Support",
  instructions: `${SUPPORT_PROMPT}\n${canaryInstruction(canary)}`,
  outputGuardrails: [shieldOutputGuardrail({ canary, systemPrompt: SUPPORT_PROMPT })],
});

What is not covered

  • Redaction and hardening. The guardrails can't rewrite text, so output with a finding is blocked or passed as it was, never redacted, and instructions are not hardened.
  • Streamed output. With stream: true, the output guardrail runs after the text was streamed. It fails the run, but can't take the text back.
  • Tools with an outputSchema. The SDK checks a rejection message against the tool's output schema, and a plain string fails that check, so run() rejects with ToolCallError (its error is the SDK's InvalidToolOutputError) instead of going on. The injected output still doesn't reach the model. For these tools, behavior: "throwException" gives you the tripwire error instead.
  • Tools that run elsewhere. Hosted tools (web search, file search, code interpreter, hosted MCP) run on OpenAI's side and take no tool guardrails. The input guardrail reads the output of a hosted_tool_call item only when one is passed in a later run's input.
  • Tool definitions. The tools an MCP server lists are not checked for tool poisoning. Run scanTools() on what server.listTools() returns.
  • Handoffs. Input guardrails don't run on the input a handoff passes to the next agent, and output guardrails only run for the agent that produced the final output.
  • History stored by OpenAI. With previousResponseId or conversationId, earlier items stay on OpenAI's side, and the input guardrail only sees the new input.
  • Content that is not read: images, files, audio, and screenshots; reasoning; and the arguments of tool calls to tools without shieldToolInputGuardrail().
  • Parallel tool calls and the tool policy. When the model asks for several tools in one turn, the SDK can run them at the same time, so a result can be recorded in the policy before a sink call from the same turn is checked, and that call is refused although the model chose it before it saw the result.
  • Realtime agents (@openai/agents/realtime), which have guardrails of their own.

See What is not scanned for what no integration scans.

On this page