Coverage
Which prompt injection and agent security checks Shield has, next to open-source libraries and hosted guardrail APIs, and what it doesn't cover.
This page compares documented capabilities as of September 27, 2026. It lists which checks exist, rather than measuring accuracy. Local Shield capabilities require the explicit local exports; the hosted detector makes an API request. The archived benchmark describes an earlier evaluation, while models and plans describes current hosted access.
In the tables, yes means the product documents the capability; for JavaScript packages, it is also present in the shipped code. Partial means the capability has a limit, no means it is absent or documented as not covered, and n/d means the documentation doesn't say.
Open-source libraries and local SDK checks
| Capability | Shield 2.0 | LLM Guard | NeMo Guardrails | LlamaFirewall | Guardrails AI | @openai/guardrails | llm-trust-guard |
|---|---|---|---|---|---|---|---|
| Injection and jailbreak detection | yes | yes | yes | yes | yes | yes (LLM) | yes (regex) |
| Decodes hidden payloads (base64, ROT13, Unicode tags) and scans them | yes | no | n/d | partial (Unicode tags) | n/d | partial (PII check only) | yes |
| Instructions split across turns | partial | no | n/d | partial (LLM audit of the trace) | n/d | yes (LLM) | yes (rules) |
| Injection in tool results and documents | yes | partial | partial | yes | partial | yes (LLM) | yes |
| Poisoned tool definitions | yes | no | no | no | no | no | yes |
| Tool definitions that change after approval | yes | no | no | no | no | no | yes |
| Tool-call policy: declared tools, argument schemas, allow and deny lists | yes | no | partial (experimental) | no | no | no | yes |
| Stop sending data out after untrusted input | yes | no | no | no | no | no | partial (forbidden sequences) |
| System prompt leak in output, checked against the real prompt | yes (including encoded and obfuscated copies) | no | n/d | no | yes (fuzzy match) | no | partial (keywords you supply) |
| Canary tokens | yes | no | no | no | no | no | no |
| Credentials in output | yes (103 kinds) | no | n/d | no | yes | yes | yes |
| Personal data in output | yes (8 types) | yes | yes | partial | yes | yes (42 types) | yes |
| Exfiltration links | yes | partial | n/d | no | partial | yes (URL allowlist) | yes |
| Injection into rendered or executed output (XSS, SQL, shell) | yes (opt-in) | no | yes | no | partial | no | yes |
| Content moderation (toxicity, topics) | no | yes | yes | no | yes | yes (OpenAI API) | no |
| Hallucination and grounding | no | yes | yes | no | yes | yes | no |
| Tool calls checked against the user's intent | no | no | no | yes (LLM) | no | yes (LLM) | partial (rules) |
| Code security scanning | no | no | partial | yes (CodeShield) | no | no | partial |
| Needs a network call or API key | no | model download on first use | depends on the rails | model download; an API key for the LLM checks | depends on the validators | yes | no |
| Language | JavaScript, TypeScript | Python | Python | Python | Python | JavaScript, TypeScript | JavaScript, TypeScript |
| Status | 2.0.0 released October 2, 2026 | archived July 2026 | active | last release May 2025 | active; Hub closed August 2026 | preview | active |
Hosted APIs
Shield’s hosted API returns an injection verdict for supplied text. Tool policies, output sanitization, canaries, and session tracking belong to the local SDK and are not services provided by these endpoints.
| Capability | Shield API | Lakera Guard | Azure Prompt Shields and Foundry guardrails | Google Model Armor | AWS Bedrock Guardrails |
|---|---|---|---|---|---|
| Injection and jailbreak detection | yes | yes | yes | yes | yes |
| Decodes hidden payloads and scans them | yes | n/d | partial | no (documented) | n/d |
| Instructions split across turns | supplied text only; no session state | no | no | no (documented) | no |
| Injection in tool results and documents | yes | yes | yes (listed tool types, not MCP) | partial (Google-hosted MCP servers) | partial (AgentCore Gateway) |
| Poisoned tool definitions | description text only | yes | no | no (documented) | no (documented) |
| Tool definitions that change after approval | no | no | no | no | no |
| Tool-call policy | no | partial (allow and deny lists) | partial (platform setting) | no | n/d |
| Stop sending data out after untrusted input | no | partial (beta, model-based) | no | no | no |
| System prompt leak in output | no | partial (a custom rule) | n/d | n/d | no (input side only) |
| Canary tokens | no | n/d | n/d | n/d | n/d |
| Credentials in output | no | no (custom rules) | partial | partial | partial |
| Personal data in output | no | yes | partial (preview) | yes | yes |
| Exfiltration or malicious links | injection instructions only; no URL reputation check | yes | partial | yes | partial |
| Content moderation | no | yes | yes | yes | yes |
| Hallucination and grounding | no | no | yes (preview) | no | yes |
| Tool calls checked against the user's intent | no | yes (beta) | yes (preview) | no | no |
| Images, documents, or audio | no | partial | partial | yes | yes |
| Runs in your process, with no network call | no | no (self-hosting needs an Enterprise license) | no | no | no |
Where Shield covers more
- Detection accuracy: the archived benchmark reports an earlier SDK evaluation. Its model names and results are not the current hosted release; compare models under the same data, configuration, and hardware.
- Shield and
llm-trust-guardare the only products here that run in your process with no model download, API key, or network call and still cover input injection, tool results, tool definitions, tool calls, and output. In the archived local evaluation, Shield without an optional transformer scored 0.777 on five groups against 0.632 forllm-trust-guard, and it checks output against the real system prompt rather than keywords you supply. - It extracts the text an agent reads from web pages and scans hidden text separately.
- It finds the system prompt in output even when it is encoded, split, or obfuscated. It is also the only product here with canary tokens; Rebuff, now archived, had them too.
- It decodes hidden payloads before scanning. Google documents that Model Armor does not;
llm-trust-guardand@stackone/defenderare the other packages that do. - It is one of two products here that scan tool definitions and catch a definition that changes after approval (the other is
llm-trust-guard), and one of two with spotlighting (the other is Azure). - It wraps nine JavaScript integrations: OpenAI (Chat Completions and Responses), Anthropic, Groq, Google Gen AI, Mistral, LangChain.js, the Vercel AI SDK, the OpenAI Agents SDK, and MCP clients.
What Shield doesn't cover
- Content moderation: toxicity, harm categories, bias, topic restrictions, and profanity. Use a moderation model or API next to Shield.
- Tool calls checked against the user's intent: LlamaFirewall's AlignmentCheck, Azure Task Adherence, Lakera's Dangerous Deviation, and
@openai/guardrailsask a model whether a tool call fits what the user asked for. Shield's tool policy enforces rules you write instead. - Hallucination and grounding checks.
- Code security scanning: LlamaFirewall's CodeShield finds vulnerable code; Shield's injection detectors only look for injected HTML, SQL, shell, and similar shapes.
- Images, PDFs, office documents, and audio.
- Local SDK languages other than JavaScript and TypeScript. Shield has no dedicated Python package; the hosted API works with the OpenAI Python client.
- Attacks that are only attacks in context: a generally benign request that violates a particular system prompt, such as an off-topic question to a support bot. Shield scans text without the system prompt, so on held-out v1, where Meta's CyberSecEval has many such cases, ProtectAI's model and GPT-6 Sol score higher.
Sources
- LLM Guard: repository (archived) and docs
- NeMo Guardrails: repository and guardrail catalog
- LlamaFirewall: repository and docs
- Guardrails AI: repository, validator catalog, and Hub shutdown notice
@openai/guardrails: repositoryllm-trust-guard: version 4.32.7 from npm, README and shipped code- Lakera Guard: screening roles, agent behavior defense, data leakage prevention
- Azure: Prompt Shields, Task Adherence, Foundry intervention points
- Google Model Armor: overview and MCP integration
- AWS Bedrock Guardrails: prompt attacks and AgentCore policy guardrails
Archived benchmarks
Historical SDK detector comparisons, evaluation limitations, and reproduction notes. These results do not describe the current hosted models.
The zeroleaks package
Source-available CLI and TypeScript library that tests a system prompt for extraction and prompt injection. Runs locally with your own OpenRouter or OpenAI key.