Models and plans
Model access, service limits, and how Shield handles submitted content.
Models
| Model ID | Access | Behavior |
|---|---|---|
shield | Free with an account, API key, and research acknowledgement. | The small Shield classifier with deterministic rules. |
shield-base | Every active paid ZeroLeaks plan. | The Base classifier with the same binary attack category. |
shield-large | Every active paid ZeroLeaks plan. | An ensemble that combines the larger classifier and Base. |
shield-tiered | Every active paid ZeroLeaks plan. | Runs Shield first and escalates uncertain scores to Large. |
Paid personal subscriptions and qualifying paid workspace memberships provide paid Shield access. Existing trial and subscription-status rules apply. A key remains subject to its owner's current access. Selecting a paid model does not purchase a subscription.
The September 2026 release uses S15e for Shield, SB1 for Base, and an ensemble of 0.6 × L7a (Qwen3-1.7B) plus 0.4 × SB1 for Large. All tiers include rules v2 and use a model threshold of 0.5. The model card reports the comparison and its limitations.
Tiered escalates small-model scores in the interval [0.01, 0.97) to Large. These models classify prompt injection and jailbreaks under one binary prompt_injection label. They do not provide independent probabilities for individual attack types.
List models
curl https://api.zeroleaks.ai/v1/models \
-H "Authorization: Bearer $ZEROLEAKS_API_KEY"The response follows the OpenAI model-list format. Each entry includes:
{
"id": "shield-large",
"object": "model",
"owned_by": "zeroleaks",
"shield": {
"requires_paid_plan": true,
"available": false
}
}The example omits the created timestamp. shield.available reflects your account's model entitlement. Free inference also requires the current research acknowledgement on the key; listing a model does not grant that acknowledgement.
Capacity and input limits
Shield has no monthly call meter and no routine per-minute quota. Service concurrency and queue length are bounded. A full queue returns a retryable 429; an unavailable inference service returns 503. Retry with backoff and respect Retry-After when supplied.
The public API accepts up to 32 moderation inputs per request, 200,000 UTF-16 code units per text, and a 2 MiB JSON body. These operational limits do not guarantee complete token coverage. Read each result's shield.coverage before accepting a long document.
Data handling
Content handling follows the account's paid access, independently of the selected model.
| Request | Submitted content retained? |
|---|---|
| Paid account, any model, any verdict | No. |
| Free account, unflagged result | No. |
| Free account, flagged result, acknowledged key | A redacted excerpt of up to 2,000 characters may be retained for security research for up to 12 months. |
Free use requires an explicit acknowledgement when creating the key. An existing free key without that acknowledgement receives research_consent_required. Paid accounts do not need to opt into research collection.
The collector preserves actual flagged input after best-effort redaction of detectable email addresses, phone numbers, credentials, credential-bearing URLs, and names. This is not a guarantee that every personal or proprietary detail is removed. Do not submit sensitive content to the free service if it cannot be retained after this redaction.
Each excerpt is capped at 2,000 characters, with at most eight excerpts per request. Where rule evidence identifies an attack, the excerpt includes nearby context. For model-only flags, the excerpt uses the beginning of the redacted input and may not contain the text that caused the flag. Redaction runs before clipping and again at the database write boundary. The collector uses hashes of the redacted excerpts to deduplicate contributions within the account.
Excerpts enter a separate research quarantine, associated with the contributing account and key. They expire at the UTC 12-month anniversary of collection and are removed by the next hourly cleanup. A classifier flag does not automatically make an excerpt training data; there is no automatic training promotion step.