ZeroLeaksDocs

Chat completions

Obtain a classification through OpenAI-compatible chat clients, including JSON and streaming.

POST https://api.zeroleaks.ai/v1/chat/completions

Shield uses the chat protocol to return a classification. It does not generate conversational answers or tool calls.

Input selection

The final message must have role user or tool. Its text is the entire classification input. Shield excludes earlier messages from the scored text and does not follow them as classifier instructions. To classify a whole transcript, format that transcript as text in the final message or use moderations.

content can be a string or an array of { "type": "text", "text": "..." } parts. Shield joins text parts with newlines. Images, audio, and other modalities are unsupported.

import OpenAI from "openai";

const shield = new OpenAI({
  apiKey: process.env.ZEROLEAKS_API_KEY,
  baseURL: "https://api.zeroleaks.ai/v1",
});

const response = await shield.chat.completions.create({
  model: "shield",
  messages: [{ role: "user", content: "Text to inspect" }],
});

const verdict = response.choices[0].message.content;

Text verdict

An unflagged input returns:

safe

A flagged input returns:

unsafe
prompt_injection

These are protocol labels. safe means the classifier did not flag the scored text; review coverage before treating it as a check of the full input.

The response is a chat.completion object with one assistant choice and finish_reason: "stop". Its top-level shield object contains flagged, score, model_score, rules, and coverage.

JSON mode

Set response_format: { type: "json_object" }. The assistant message's content is a JSON string:

const response = await shield.chat.completions.create({
  model: "shield",
  messages: [{ role: "user", content: "Text to inspect" }],
  response_format: { type: "json_object" },
});

const verdict = JSON.parse(response.choices[0].message.content ?? "null");
// { flagged: false, score: 0.03, category: null }
// or { flagged: true, score: 0.93, category: "prompt_injection" }

The scores above are examples. Structured-output JSON schemas are not supported.

Streaming

const stream = await shield.chat.completions.create({
  model: "shield",
  messages: [{ role: "user", content: "Text to inspect" }],
  stream: true,
  stream_options: { include_usage: true },
});

let verdict = "";
for await (const chunk of stream) {
  verdict += chunk.choices[0]?.delta.content ?? "";
}

The server classifies first, then emits OpenAI-compatible server-sent events: a content chunk with Shield metadata, a finish chunk, an optional usage chunk, and data: [DONE]. Streaming does not progressively classify an unfinished input. Inference failures return an HTTP error before a verdict stream begins.

usage.prompt_tokens reports input token accounting. completion_tokens is zero because the classifier does not generate tokens; the service formats the verdict text. total_tokens equals prompt_tokens.

Supported options and limits

  • model: one of the four Shield model IDs; defaults to shield.
  • messages: 1 to 256 messages; only the final user or tool message is scored.
  • response_format: text or json_object.
  • stream: boolean; defaults to false.
  • stream_options.include_usage: supported when stream is true.
  • temperature: only 0, if supplied.
  • n: only 1, if supplied.
  • max_tokens and max_completion_tokens: accepted positive integers for client compatibility; they do not truncate the verdict.

Tool calling, function calling, multiple choices, nonzero temperature, log probabilities, and custom stop sequences are unsupported. The moderation input and body limits, coverage semantics, and errors also apply.

On this page