OpenAI Agents SDK
Shield guardrails for the OpenAI Agents SDK. The run's input and every tool output are checked for injection, the final output and tool call arguments for prompt leaks, credentials, exfiltration links, and canaries, and tool calls against a tool policy.
The @zeroleaks/shield/openai-agents module provides five guardrails for the OpenAI Agents SDK (@openai/agents). Register them in the SDK's guardrail slots on an agent, a Runner, a function tool, or an MCP server:
| Guardrail | Slot | Checks | On a finding |
|---|---|---|---|
shieldInputGuardrail() | inputGuardrails of an Agent or Runner | The run's input for injection | run() rejects with InputGuardrailTripwireTriggered |
shieldOutputGuardrail() | outputGuardrails of an Agent or Runner | The final output for prompt leaks, credentials, exfiltration links, and canaries | run() rejects with OutputGuardrailTripwireTriggered |
shieldToolOutputGuardrail() | outputGuardrails of a tool(), toolOutputGuardrails of an MCP server | What the tool returned, for injection, before the model reads it | The model gets a rejection message instead of the output |
shieldToolInputGuardrail() | inputGuardrails of a tool(), toolInputGuardrails of an MCP server | The tool call's arguments, before the tool runs, for credentials, exfiltration links, canaries, and system prompt text | The tool doesn't run, and the model gets a rejection message |
shieldToolPolicyGuardrail() | inputGuardrails of a tool(), toolInputGuardrails of an MCP server | The tool call, before the tool runs, against a tool policy | The tool doesn't run, and the model gets the policy's reason |
Guardrails can stop a run or a tool call, but they can't change text. Nothing is redacted and instructions are not hardened. Content with a finding either trips the guardrail or passes through unchanged.
Usage
import {
Agent,
InputGuardrailTripwireTriggered,
OutputGuardrailTripwireTriggered,
run,
tool,
} from "@openai/agents";
import { z } from "zod";
import {
shieldInputGuardrail,
shieldOutputGuardrail,
shieldToolInputGuardrail,
shieldToolOutputGuardrail,
} from "@zeroleaks/shield/openai-agents";
const readInbox = tool({
name: "read_inbox",
description: "Read new email.",
parameters: z.object({}),
execute: async () => fetchInbox(),
outputGuardrails: [shieldToolOutputGuardrail()],
});
const sendEmail = tool({
name: "send_email",
description: "Send an email.",
parameters: z.object({ to: z.string(), body: z.string() }),
execute: async ({ to, body }) => send(to, body),
inputGuardrails: [shieldToolInputGuardrail()],
});
const agent = new Agent({
name: "Support",
instructions: "You are a support agent for Acme. Never share internal ticket notes.",
tools: [readInbox, sendEmail],
inputGuardrails: [shieldInputGuardrail()],
outputGuardrails: [shieldOutputGuardrail()],
});
try {
const result = await run(agent, userInput);
return result.finalOutput;
} catch (error) {
if (
error instanceof InputGuardrailTripwireTriggered ||
error instanceof OutputGuardrailTripwireTriggered
) {
console.warn(error.result.guardrail.name, error.result.output.outputInfo);
return "Sorry, I can't help with that.";
}
throw error;
}To guard every tool exposed by an MCP server, pass the tool guardrails to the server (Agents SDK 0.17.1 or later):
import { MCPServerStreamableHttp } from "@openai/agents";
const server = new MCPServerStreamableHttp({
url: serverUrl,
toolInputGuardrails: [shieldToolInputGuardrail()],
toolOutputGuardrails: [shieldToolOutputGuardrail()],
});To guard every agent a Runner runs, pass the input and output guardrails to it instead: new Runner({ inputGuardrails: [shieldInputGuardrail()], outputGuardrails: [shieldOutputGuardrail()] }).
How it works
Input guardrail
shieldInputGuardrail() runs detection on the run's input:
- A string input is checked as a user message.
- In a list of input items, the string content or
input_textparts of each user message are checked as a user message, with thedetectoptions. Tool results in the list are checked as tool results, with the tool result options: the output offunction_call_resultitems (read like a tool output, as described below), thestdoutandstderrofshell_call_outputitems, and theoutputofapply_patch_call_output,program_output, andhosted_tool_callitems. A list has tool results when you pass thehistoryof an earlier run, or when asessionadds its stored items, which the guardrail sees along with the new input. WithonDetection: "warn", an injection that stays in the history is reported on every run. - System and assistant messages are not checked.
With onDetection: "block" (the default), the first injection trips the guardrail and run() rejects with the SDK's InputGuardrailTripwireTriggered. Its result.output.outputInfo is { detected: true, source, result }, where source is "user" or "tool" and result is the LocalDetectResult, which holds the risk, the matched categories and patterns, and the classifier score, but not the text. With "warn", onInjectionDetected(result, source) is called for each injection and the run continues; the guardrail's entry in result.inputGuardrailResults still has detected: true.
The SDK runs an input guardrail alongside the agent's first model call unless the guardrail sets runInParallel: false. Shield sets it to false, so the input is checked before the model is called, and a blocked input never reaches the model or a tool. Pass runInParallel: true to reduce latency at the cost of this guarantee, for example with a slow secondaryDetector or escalate detector. With stream: true, run() resolves with the stream, and the error comes from result.completed and the stream.
The SDK runs input guardrails once per run, on the input to the first agent. They don't run again after a handoff or on later turns. Use the tool output guardrail to check tool results produced during the run.
Output guardrail
shieldOutputGuardrail() checks the agent's final output: a string, or every string in a structured output (outputType), keys included. It looks for:
- A leak of the system prompt:
systemPromptif you pass it, else the agent'sinstructionswhen they are a string. Withinstructionsgiven as a function, or a storedprompt, passsystemPrompt, or the prompt leak check is skipped. - Output findings from
scanOutputText(): credentials and exfiltration links by default, and personal data withoutput: { pii: true }. - A canary you pass in
canaryoroutput.canary.
A prompt leak or a canary trips the guardrail, unless throwOnLeak is false, and so does a high or critical output finding, unless blockOnOutputFindings is false. run() then rejects with the SDK's OutputGuardrailTripwireTriggered. onLeakDetected and onOutputFindings are called whether or not it trips.
outputInfo is { leak, findings }: leak holds the leak's confidence and fragmentCount, and findings holds the type, kind, and severity of every output finding when one of them is high or critical. It never contains the text. The SDK includes it in the error message, so the message is safe to log. The error's result.agentOutput contains the SDK's unredacted copy of the output that tripped the guardrail. Do not log that field.
The SDK runs the output guardrails of the agent that produced the final output, plus those of the Runner. If a handoff occurred, it uses the last agent's guardrails. With stream: true, the guardrail runs after the model finishes and the text has already been streamed. The error is thrown from your for await loop and result.completed after the last chunk.
Tool output guardrail
shieldToolOutputGuardrail() runs after a function tool returns and before its output reaches the model. It checks the output's text with the tool result options:
- a string, whole;
- a text part (
text,input_text, oroutput_text); - an MCP content block: text, an embedded resource (its text, or a
text/*,application/json, orapplication/xmlblob, decoded, up to 64KB), or a resource link's title and description; - a list of those;
- the string values of anything else, not its keys, up to 64K characters.
With onDetection: "block" and the default behavior: "rejectContent", an injection causes the SDK to replace the tool's output with rejectionMessage, and the run continues. With behavior: "throwException", run() rejects with the SDK's ToolCallError, whose error is ToolOutputGuardrailTripwireTriggered. With "warn", onInjectionDetected(result, "tool") is called and the output reaches the model. outputInfo has the same shape as the input guardrail's, and each result is in result.toolOutputGuardrailResults.
The tool has already run when its output is checked. A blocked result does not reach the model.
With policy, each output is also recorded in the tool policy with policy.recordResult(), flagged when detection found an injection in it, whether or not the guardrail rejects it. See Tool policy guardrail.
Tool input guardrail
shieldToolInputGuardrail() checks the arguments generated by the model before a function tool runs. It parses them as JSON and checks every string, keys included, with the same checks as the output guardrail: text of the system prompt, credentials, exfiltration links, and the canary. An agent induced to put a credential, the canary, or its instructions into a tool call is stopped before the tool runs.
Only high and critical findings trip it. A markdown or HTML image that sends data out is critical, but a bare URL that carries data in its query or path, the usual shape of an argument to a fetch tool, is most often a medium url_data finding: it is reported to onOutputFindings, and the tool runs.
A prompt leak or a canary (unless throwOnLeak: false) or a high or critical finding (unless blockOnOutputFindings: false) trips the guardrail. With the default behavior: "rejectContent", the tool doesn't run and the model gets rejectionMessage as its output. With "throwException", run() rejects with the SDK's ToolCallError, whose error is ToolInputGuardrailTripwireTriggered. outputInfo has the same shape as the output guardrail's.
Tool policy guardrail
shieldToolPolicyGuardrail({ policy }) runs policy.checkAsync() on each call before the tool runs, with the call's name and its JSON arguments. A policy from createToolPolicy() remembers what a session has read, so give each run its own: pass a function of the context you pass to run(), and give the same function to shieldToolOutputGuardrail({ policy }), which records each output in the policy:
import { createToolPolicy, type ToolPolicy } from "@zeroleaks/shield";
import {
shieldToolOutputGuardrail,
shieldToolPolicyGuardrail,
} from "@zeroleaks/shield/openai-agents";
type Session = { policy: ToolPolicy };
const policyGuardrail = shieldToolPolicyGuardrail<Session>({ policy: (context) => context.policy });
const outputGuardrail = shieldToolOutputGuardrail<Session>({ policy: (context) => context.policy });
const sendEmail = tool({
name: "send_email",
description: "Send an email.",
parameters: z.object({ to: z.string(), body: z.string() }),
execute: async ({ to, body }) => send(to, body),
inputGuardrails: [policyGuardrail],
outputGuardrails: [outputGuardrail],
});
await run(agent, userInput, {
context: { policy: createToolPolicy({ rules: { send_email: { labels: ["sink"] } } }) },
});To use one policy throughout a conversation, store it in the context object passed to every run() in that conversation. A policy object passed directly, instead of a function, is shared by every run that uses the guardrail. A function that returns something other than a policy fails the run with a TypeError.
With the default behavior: "rejectContent", a refused call doesn't run, and the model gets the decision's message, or rejectionMessage if you set one, as the tool's output. With "throwException", run() rejects with the SDK's ToolCallError, whose error is ToolInputGuardrailTripwireTriggered. outputInfo is the policy's decision ({ allowed, reason, message, tool, violations? }), which never holds argument values; it is in result.toolInputGuardrailResults.
The policy counts every call it allows. Place this guardrail last in inputGuardrails, after shieldToolInputGuardrail(), so calls rejected by another guardrail are not counted. The SDK already refuses tool names the agent doesn't have. The policy's declared tools therefore matter for tools you leave out of allow, and for argument schemas the SDK doesn't enforce, as with a JSON Schema tool that isn't strict.
Options
shieldInputGuardrail and shieldToolOutputGuardrail take detect, scanToolResults, onDetection, and onInjectionDetected from the shared options. shieldToolOutputGuardrail only checks tool output, so detect only sets the options that scanToolResults: true (the default) uses. Detection results are reused for text already checked, and a secondaryDetector or escalate detector runs before a guardrail trips, as described under Shared behavior.
shieldOutputGuardrail and shieldToolInputGuardrail take systemPrompt, sanitize, output, onLeakDetected, and onOutputFindings from the shared options. shieldToolPolicyGuardrail takes none of the shared options.
The guardrails also take:
| Option | Guardrails | Type | Default | Description |
|---|---|---|---|---|
name | all | string | "shield_input", "shield_output", "shield_tool_input", "shield_tool_output", "shield_tool_policy" | The name the SDK shows in errors, results, and traces |
runInParallel | input | boolean | false | Run alongside the first model call instead of before it. See Input guardrail. |
canary | output, tool input | string | none | A canary token you planted in the instructions. Its appearance counts as a leak. It must be a string: these guardrails can't plant one. |
throwOnLeak | output, tool input | boolean | true | Trip on a prompt leak or a canary. The shared option defaults to false, since the wrappers can redact instead. |
blockOnOutputFindings | output, tool input | boolean | true | Trip on a high or critical output finding. The shared option defaults to false for the same reason. |
behavior | tool input, tool output, tool policy | "rejectContent" | "throwException" | "rejectContent" | What a tripped tool guardrail does: give the model rejectionMessage in place of the output, or fail the run |
rejectionMessage | tool input, tool output, tool policy | string | see below | What the model gets in place of the tool's output |
policy | tool policy (required), tool output | ToolPolicy | (context) => ToolPolicy | none | The tool policy, or a function that returns it from the context passed to run(). The tool policy guardrail checks calls against it; the tool output guardrail records outputs in it. See Tool policy guardrail. |
The default rejection messages are "This tool output was withheld because Shield flagged it as a possible prompt injection." for tool output, and "This tool call was blocked because Shield found a credential, an exfiltration link, a canary, or system prompt text in its arguments." for tool calls. The tool policy guardrail's default is the decision's message, such as Tool "send_email" can send data out, and this session has seen untrusted content from read_inbox.
To use a canary, put it in the instructions yourself and pass the same token to the guardrails:
import { canaryInstruction, createCanary } from "@zeroleaks/shield";
const canary = createCanary();
const agent = new Agent({
name: "Support",
instructions: `${SUPPORT_PROMPT}\n${canaryInstruction(canary)}`,
outputGuardrails: [shieldOutputGuardrail({ canary, systemPrompt: SUPPORT_PROMPT })],
});What is not covered
- Redaction and hardening. The guardrails can't rewrite text, so output with a finding is blocked or passed as it was, never redacted, and instructions are not hardened.
- Streamed output. With
stream: true, the output guardrail runs after the text was streamed. It fails the run, but can't take the text back. - Tools with an
outputSchema. The SDK checks a rejection message against the tool's output schema, and a plain string fails that check, sorun()rejects withToolCallError(itserroris the SDK'sInvalidToolOutputError) instead of going on. The injected output still doesn't reach the model. For these tools,behavior: "throwException"gives you the tripwire error instead. - Tools that run elsewhere. Hosted tools (web search, file search, code interpreter, hosted MCP) run on OpenAI's side and take no tool guardrails. The input guardrail reads the
outputof ahosted_tool_callitem only when one is passed in a later run's input. - Tool definitions. The tools an MCP server lists are not checked for tool poisoning. Run
scanTools()on whatserver.listTools()returns. - Handoffs. Input guardrails don't run on the input a handoff passes to the next agent, and output guardrails only run for the agent that produced the final output.
- History stored by OpenAI. With
previousResponseIdorconversationId, earlier items stay on OpenAI's side, and the input guardrail only sees the new input. - Content that is not read: images, files, audio, and screenshots; reasoning; and the arguments of tool calls to tools without
shieldToolInputGuardrail(). - Parallel tool calls and the tool policy. When the model asks for several tools in one turn, the SDK can run them at the same time, so a result can be recorded in the policy before a sink call from the same turn is checked, and that call is refused although the model chose it before it saw the result.
- Realtime agents (
@openai/agents/realtime), which have guardrails of their own.
See What is not scanned for what no integration scans.
MCP Client
Wrap your Model Context Protocol client with Shield. Tool lists are checked for tool poisoning and flagged tools are dropped; tool results, resources, and prompts are checked for injection before they reach the model; and tool calls can be checked against a tool policy.
Vercel AI SDK
shieldLanguageModelMiddleware and shieldMiddleware for generateText and streamText on AI SDK 4, 5, 6, and 7. Hardening, injection detection in user messages and tool results, and output guarding.