LangChain.js
shieldChatModel and ShieldCallbackHandler for LangChain.js chat models, chains, and agents. Hardening, injection detection in human and tool messages, and output guarding on invoke, batch, stream, and streamEvents.
Shield has two LangChain.js integrations, both in @zeroleaks/shield/langchain. Both require @langchain/core.
- shieldChatModel wraps a chat model, such as
ChatOpenAIorChatAnthropic. It hardens system messages, checks human and tool messages, and guards the model's text and tool call arguments on every call, streamed or not. Use this integration by default. - ShieldCallbackHandler checks the runs it is attached to: messages sent to chat models, tool outputs, retrieved documents, and model output. It can stop a run, but it can't harden a prompt or redact output. Use it for chains and agents whose model you can't wrap.
shieldChatModel
import { ChatOpenAI } from "@langchain/openai";
import { shieldChatModel } from "@zeroleaks/shield/langchain";
const model = shieldChatModel(new ChatOpenAI({ model: "gpt-5.5" }), {
systemPrompt: "You are a support agent...",
});
const reply = await model.invoke([
["system", "You are a support agent..."],
["human", userInput],
]);The wrapped model retains your model's type and class, supports bindTools() and withStructuredOutput(), and can be used in a chain or agent. Passing something that is not a chat model (one without _generate) throws a TypeError.
How it works
Every call on a LangChain chat model, whether invoke, batch, stream, streamEvents, or a call from a chain or agent, ultimately calls the model's _generate or _streamResponseChunks method. shieldChatModel returns a Proxy over your model that guards those two methods. On each call, Shield:
- Hardens every system message, a
SystemMessageor aChatMessagewith rolesystemordeveloper(unlessharden: false), adding the canary if you set one. It makes a new message of the same class: string content stays a string, and in content blocks the text blocks are joined with newlines into the first one and hardened, and other blocks stay where they were. Your messages are not modified. - Runs detection on every human message (unless
detect: false) and every tool and function message (unlessscanToolResults: false), including earlier turns. Text blocks are joined; other blocks are ignored. AI messages are not scanned. - Calls the model's own method.
- Guards each AI message in the result: its text, as a string or its text blocks joined and written back over the same blocks, is checked for prompt leaks (unless
sanitize: falseor there is no system prompt) and passed through the output guard (unlessoutput: false), and so is every string in eachtool_callsentry'sargs. - Updates the other copies of the tool call arguments that providers keep:
tool_use,tool_call,tool_call_chunk, andfunction_callcontent blocks,tool_call_chunks, and OpenAI'sadditional_kwargs.tool_calls. Each takes the guarded arguments of the tool call with the same id, so a finding is reported once. A copy that matches no tool call, and each ofinvalid_tool_calls, is guarded on its own. - When anything was redacted, rebuilds the generation's
textand removeslogprobsfrom itsgenerationInfoand the message'sresponse_metadata.
For prompt leak checks, Shield uses systemPrompt if provided. Otherwise, it uses the text of the first system message before hardening.
A secondaryDetector in the detect or scanToolResults options runs before a request is blocked, and detection results are reused for text already checked, as described under Shared behavior.
Runnables made from the model
Other methods run with the wrapped model as this, so the runnables they build call the guarded model: bindTools(), withConfig(), withStructuredOutput(), pipe(), withRetry(), withFallbacks(), and the rest. If a method returns a new chat model instead, or a runnable bound to one, as some bindTools() implementations do, that model is wrapped with the same options.
const withTools = model.bindTools([searchTool]); // still guarded
const extractor = model.withStructuredOutput(schema); // still guardedOptions
shieldChatModel takes the shared options. systemPrompt defaults to the text of the first system message, before hardening. streamingChunkSize is not used.
Streaming
See Streaming for the three streamingSanitize modes. Here they work like this:
| Mode | Behavior |
|---|---|
"buffer" (default) | Reads every chunk from the model before yielding any. If the message the chunks make up is clean, they are replayed unchanged. If anything in it was redacted, you get a single chunk with the whole guarded message. |
"chunked" | Works like "buffer". |
"passthrough" | Chunks and tokens pass through unchanged, with no sanitization or output scanning. Input is still checked. |
Tokens the model reports to callback handlers through handleLLMNewToken are also buffered, then replayed for the chunks that are yielded. Callback handlers, streamEvents(), and LangGraph's messages stream mode therefore only see guarded text. A model that streams inside _generate has its tokens buffered the same way, and each generation's guarded text is then sent as one token.
streamEvents() without a version produces LangChain's content block events. Shield builds them from the guarded chunks with LangChain's base implementation, in place of the provider's own.
What is not covered
- The wrapped model is a Proxy. Methods other than
_generateand_streamResponseChunksrun with the Proxy asthis. A model class that keeps state in JavaScript private fields (#field) and reads them in those methods throws aTypeErrorthrough the Proxy. A model that produces output without going through_generateor_streamResponseChunksis not guarded. - Cache hits: with a LangChain cache, a hit returns the stored response without calling the model, so neither the input check nor the output guard runs for it.
- A provider's own content block events are not used; see Streaming.
- Reasoning and thinking blocks in the output are not scanned, and images, audio, and files in messages are not read.
- Other parts of a chain: only the wrapped model is guarded. Tools, retrievers, and other models in the chain are not; attach a ShieldCallbackHandler to check tool outputs and retrieved documents.
ShieldCallbackHandler
import { ShieldCallbackHandler } from "@zeroleaks/shield/langchain";
const shield = new ShieldCallbackHandler({
throwOnLeak: true,
blockOnOutputFindings: true,
});
const result = await agent.invoke(input, { callbacks: [shield] });The handler checks what passes through the runs it is attached to:
| Callback | What it checks | On a finding |
|---|---|---|
handleChatModelStart | Human and tool messages sent to a chat model, including earlier turns. It also records the run's system prompt. | Throws InjectionDetectedError, and the model is not called |
handleToolEnd | A tool's output: a string, a ToolMessage's content, or the string values of anything else | Throws InjectionDetectedError with source: "tool", and the tool call fails |
handleRetrieverEnd | The pageContent of each retrieved document | Throws InjectionDetectedError with source: "tool" |
handleLLMEnd | The model's text and tool call arguments: prompt leaks against systemPrompt or the run's first system message, and credentials, exfiltration links, and a canary | Calls onLeakDetected and onOutputFindings. Throws LeakDetectedError with throwOnLeak, and OutputBlockedError with blockOnOutputFindings. Otherwise the output goes through unchanged. |
With onDetection: "warn", injections are reported to onInjectionDetected and nothing throws.
The handler sets raiseError, so LangChain waits for it and passes its errors on to the run, after logging them with console.error. A framework that catches tool errors and reports them to the model, as LangGraph's ToolNode does by default, sends the model the error message instead of the tool output.
It can block, but it can't redact. A callback observes the run but cannot modify it: prompts are not hardened, output is not redacted, and output with a finding either stops the run or passes through unchanged. With streaming, handleLLMEnd runs after the text has been streamed, so it can fail the run but cannot retract that text. Use shieldChatModel wherever you can wrap the model.
Options
ShieldCallbackHandler takes the shared options except harden, streamingSanitize, and streamingChunkSize. Since it can't add anything to the prompt, canary only finds a token you put in the prompt yourself; canary: true creates a token that never appears.
What is not covered
LLM runs with a plain string prompt (handleLLMStart), chain inputs and outputs, and anything the handler isn't attached to are not checked. See What is not scanned for what no integration scans.
Mistral Provider
Wrap your Mistral client with Shield for prompt hardening, injection detection in user messages and tool results, and output guarding on chat.complete and chat.stream.
MCP Client
Wrap your Model Context Protocol client with Shield. Tool lists are checked for tool poisoning and flagged tools are dropped; tool results, resources, and prompts are checked for injection before they reach the model; and tool calls can be checked against a tool policy.