Shield API
Check untrusted text for prompt injection before it reaches your agent.
Shield classifies user messages, retrieved documents, and tool results for prompt injection. Send text to the API, read the verdict and coverage, then apply your application's policy before that content reaches the agent.
The API has one binary attack category, prompt_injection, which includes jailbreak attempts. It does not generate answers, execute tools, or replace an agent's authorization checks.
Start with a key
Create an API key. The shield model is free with an account and an explicit research acknowledgement. Every paid ZeroLeaks plan also includes shield-base, shield-large, and shield-tiered. Paid accounts can use an existing ZeroLeaks key.
Base URL: https://api.zeroleaks.ai/v1
Authorization: Bearer zl_live_...Keep the key on your server. The same key authenticates ZeroLeaks scan API calls, but access to scanning depends on your plan.
Choose an endpoint
| Endpoint | Use |
|---|---|
POST /moderations | Classify one text or a batch of up to 32 texts. |
POST /chat/completions | Use an OpenAI-compatible chat client to obtain a classification, with JSON and streaming support. |
GET /models | List the model IDs and their availability to your account. |
Both classification endpoints use the same models. Chat completions classifies only the final user or tool message; it does not follow instructions in earlier messages.
Read the whole result
flagged combines the model's decision with deterministic rules. category_scores.prompt_injection is the resulting policy score. The separate shield.model_score preserves the model's score, and shield.rules tells you whether a rule matched.
For long text, Shield scores overlapping windows from the head and tail. Check shield.coverage.truncated: accepting a large input does not mean every token was scored. If complete coverage is required, split content into bounded chunks and check each one, or use the SDK's requireFullCoverage option to reject partial results.
A service failure is an error, never a safe verdict. Decide how your application handles an unavailable classifier before placing it in a critical path.
Data handling
Shield does not retain submitted content from paid accounts, including when they select shield. Unflagged free inputs are not retained. With your free-use acknowledgement, flagged requests may contribute redacted excerpts of up to 2,000 characters to security research for up to 12 months. Shield applies best-effort redaction to detectable personal data and credentials, but cannot guarantee that it removes every sensitive detail. See models, plans, and data handling.