ZeroLeaksDocs
Shield SDK

Coverage

Which prompt injection and agent security checks Shield has, next to open-source libraries and hosted guardrail APIs, and what it doesn't cover.

This page compares documented capabilities as of September 27, 2026. It lists which checks exist, rather than measuring accuracy. Local Shield capabilities require the explicit local exports; the hosted detector makes an API request. The archived benchmark describes an earlier evaluation, while models and plans describes current hosted access.

In the tables, yes means the product documents the capability; for JavaScript packages, it is also present in the shipped code. Partial means the capability has a limit, no means it is absent or documented as not covered, and n/d means the documentation doesn't say.

Open-source libraries and local SDK checks

CapabilityShield 2.0LLM GuardNeMo GuardrailsLlamaFirewallGuardrails AI@openai/guardrailsllm-trust-guard
Injection and jailbreak detectionyesyesyesyesyesyes (LLM)yes (regex)
Decodes hidden payloads (base64, ROT13, Unicode tags) and scans themyesnon/dpartial (Unicode tags)n/dpartial (PII check only)yes
Instructions split across turnspartialnon/dpartial (LLM audit of the trace)n/dyes (LLM)yes (rules)
Injection in tool results and documentsyespartialpartialyespartialyes (LLM)yes
Poisoned tool definitionsyesnononononoyes
Tool definitions that change after approvalyesnononononoyes
Tool-call policy: declared tools, argument schemas, allow and deny listsyesnopartial (experimental)nononoyes
Stop sending data out after untrusted inputyesnononononopartial (forbidden sequences)
System prompt leak in output, checked against the real promptyes (including encoded and obfuscated copies)non/dnoyes (fuzzy match)nopartial (keywords you supply)
Canary tokensyesnononononono
Credentials in outputyes (103 kinds)non/dnoyesyesyes
Personal data in outputyes (8 types)yesyespartialyesyes (42 types)yes
Exfiltration linksyespartialn/dnopartialyes (URL allowlist)yes
Injection into rendered or executed output (XSS, SQL, shell)yes (opt-in)noyesnopartialnoyes
Content moderation (toxicity, topics)noyesyesnoyesyes (OpenAI API)no
Hallucination and groundingnoyesyesnoyesyesno
Tool calls checked against the user's intentnononoyes (LLM)noyes (LLM)partial (rules)
Code security scanningnonopartialyes (CodeShield)nonopartial
Needs a network call or API keynomodel download on first usedepends on the railsmodel download; an API key for the LLM checksdepends on the validatorsyesno
LanguageJavaScript, TypeScriptPythonPythonPythonPythonJavaScript, TypeScriptJavaScript, TypeScript
Status2.0.0 released October 2, 2026archived July 2026activelast release May 2025active; Hub closed August 2026previewactive

Hosted APIs

Shield’s hosted API returns an injection verdict for supplied text. Tool policies, output sanitization, canaries, and session tracking belong to the local SDK and are not services provided by these endpoints.

CapabilityShield APILakera GuardAzure Prompt Shields and Foundry guardrailsGoogle Model ArmorAWS Bedrock Guardrails
Injection and jailbreak detectionyesyesyesyesyes
Decodes hidden payloads and scans themyesn/dpartialno (documented)n/d
Instructions split across turnssupplied text only; no session statenonono (documented)no
Injection in tool results and documentsyesyesyes (listed tool types, not MCP)partial (Google-hosted MCP servers)partial (AgentCore Gateway)
Poisoned tool definitionsdescription text onlyyesnono (documented)no (documented)
Tool definitions that change after approvalnonononono
Tool-call policynopartial (allow and deny lists)partial (platform setting)non/d
Stop sending data out after untrusted inputnopartial (beta, model-based)nonono
System prompt leak in outputnopartial (a custom rule)n/dn/dno (input side only)
Canary tokensnon/dn/dn/dn/d
Credentials in outputnono (custom rules)partialpartialpartial
Personal data in outputnoyespartial (preview)yesyes
Exfiltration or malicious linksinjection instructions only; no URL reputation checkyespartialyespartial
Content moderationnoyesyesyesyes
Hallucination and groundingnonoyes (preview)noyes
Tool calls checked against the user's intentnoyes (beta)yes (preview)nono
Images, documents, or audionopartialpartialyesyes
Runs in your process, with no network callnono (self-hosting needs an Enterprise license)nonono

Where Shield covers more

  • Detection accuracy: the archived benchmark reports an earlier SDK evaluation. Its model names and results are not the current hosted release; compare models under the same data, configuration, and hardware.
  • Shield and llm-trust-guard are the only products here that run in your process with no model download, API key, or network call and still cover input injection, tool results, tool definitions, tool calls, and output. In the archived local evaluation, Shield without an optional transformer scored 0.777 on five groups against 0.632 for llm-trust-guard, and it checks output against the real system prompt rather than keywords you supply.
  • It extracts the text an agent reads from web pages and scans hidden text separately.
  • It finds the system prompt in output even when it is encoded, split, or obfuscated. It is also the only product here with canary tokens; Rebuff, now archived, had them too.
  • It decodes hidden payloads before scanning. Google documents that Model Armor does not; llm-trust-guard and @stackone/defender are the other packages that do.
  • It is one of two products here that scan tool definitions and catch a definition that changes after approval (the other is llm-trust-guard), and one of two with spotlighting (the other is Azure).
  • It wraps nine JavaScript integrations: OpenAI (Chat Completions and Responses), Anthropic, Groq, Google Gen AI, Mistral, LangChain.js, the Vercel AI SDK, the OpenAI Agents SDK, and MCP clients.

What Shield doesn't cover

  • Content moderation: toxicity, harm categories, bias, topic restrictions, and profanity. Use a moderation model or API next to Shield.
  • Tool calls checked against the user's intent: LlamaFirewall's AlignmentCheck, Azure Task Adherence, Lakera's Dangerous Deviation, and @openai/guardrails ask a model whether a tool call fits what the user asked for. Shield's tool policy enforces rules you write instead.
  • Hallucination and grounding checks.
  • Code security scanning: LlamaFirewall's CodeShield finds vulnerable code; Shield's injection detectors only look for injected HTML, SQL, shell, and similar shapes.
  • Images, PDFs, office documents, and audio.
  • Local SDK languages other than JavaScript and TypeScript. Shield has no dedicated Python package; the hosted API works with the OpenAI Python client.
  • Attacks that are only attacks in context: a generally benign request that violates a particular system prompt, such as an off-topic question to a support bot. Shield scans text without the system prompt, so on held-out v1, where Meta's CyberSecEval has many such cases, ProtectAI's model and GPT-6 Sol score higher.

Sources

On this page