The zeroleaks package
Source-available CLI and TypeScript library that tests a system prompt for extraction and prompt injection. Runs locally with your own OpenRouter or OpenAI key.
The zeroleaks npm package is a standalone scanner for system prompts. It loads your prompt into a model you choose and tests that model in two ways:
- Extraction: a multi-turn attacker tries to get the model to reveal its own system prompt.
- Injection: a set of behavioral probes tries to get the model to follow planted instructions (in a document, a fake admin message, or a tool description) or to agree to misuse a tool. A judge model decides whether the target complied.
It runs on your machine with your own model keys. The source is at github.com/ZeroLeaks/zeroleaks. These pages describe version 1.5.0.
License
The package is licensed under FSL-1.1-Apache-2.0, the Functional Source License. You can use, modify, and redistribute it for any purpose except offering a competing commercial product or service. Each release becomes available under Apache 2.0 two years after it is published or on January 21, 2028, whichever comes first.
Features
- runSecurityScan(): one call that takes a system prompt and returns a score, findings, and recommendations
- createScanEngine(): full control over turns, tree depth, branching, Best-of-N, and callbacks
- CLI:
zeroleaks scan,zeroleaks probes,zeroleaks categories,zeroleaks techniques - Agents: Strategist, Attacker, Evaluator, Mutator, Inspector, Orchestrator, and an injection compliance judge
- Probe library: extraction probes across direct, encoding, persona, social, crescendo, many-shot, policy puppetry, and more, plus 78 behavioral injection probes in six categories
- Knowledge base: documented techniques, payload templates, exfiltration vectors, and defense bypass methods, each with a source
- Models: any OpenRouter model, OpenAI models directly with
OPENAI_API_KEY, or any OpenAI-compatible endpoint (Ollama, vLLM, a gateway) with--base-url
Architecture
The extraction scan follows TAP (Tree of Attacks with Pruning) with several cooperating agents:
| Agent | Role |
|---|---|
| Strategist | Reads the conversation, picks the next strategy, and moves between attack phases |
| Attacker | Generates attack prompts from the strategy and evaluator feedback |
| Evaluator | Decides how much of the system prompt leaked in each response |
| Mutator | Produces Best-of-N variations of an attack |
| Inspector | Fingerprints known guardrails (TombRaider-style) and suggests bypasses |
| Orchestrator | Runs scripted multi-turn sequences (Siren, Echo Chamber, TombRaider) |
| InjectionEvaluator | Judges whether the target complied with an injection probe (full, partial, or refused) |
Research foundation
zeroleaks draws on:
- TAP: Tree of Attacks with Pruning (Mehrotra et al.)
- PAIR: Prompt Automatic Iterative Refinement
- Crescendo: multi-turn gradual escalation
- TombRaider: dual-agent defense fingerprinting
- Siren: multi-turn human jailbreak simulation
- Echo Chamber: escalation through agreement
- Best-of-N: sampling-based jailbreaking
- Garak: probes adapted from NVIDIA's garak
- AgentDojo, InjecAgent, JailbreakBench, HarmBench, promptfoo, OWASP LLM Top 10: sources for the behavioral injection probes
Use cases
- CI: scan a system prompt when it changes.
zeroleaks scanexits with 0 when every check passed, 1 when it finds a vulnerability, and 2 when it can't reach a verdict. - Local development: test prompt wording without an account.
- Custom tooling: call the engine, agents, or probe library from your own code.
- Research: read the probe library and knowledge base programmatically.
zeroleaks package or hosted ZeroLeaks?
The zeroleaks package tests a system prompt. The target is a model you specify, running your prompt with no tools, memory, or application code. A finding means that model, given that prompt, leaked its instructions or followed an injected instruction in its reply.
Hosted ZeroLeaks tests a running agent. Through @zeroleaks/sdk runtime scans or endpoint scans, every probe goes through your real agent: its model settings, tools, memory, credentials, and the other agents it talks to. It profiles the agent, derives the rules the agent must never break, and checks whether real tool calls break them. Outbound payloads are contained so a successful exploit reaches a sink ZeroLeaks controls. Findings come with a least-privilege policy or configuration fix, reports are stored, and later scans rerun the same attacks to catch regressions.
Hosted agent scans no longer run a system-prompt extraction loop. They run a secrets-in-context track instead, which scores what sensitive data leaked (credentials, internal hosts, tool schemas, approval rules), not whether the prompt was reproduced word for word. Hosted prompt scans are retired.
zeroleaks package | Hosted ZeroLeaks | |
|---|---|---|
| Target | A model running your system prompt | Your deployed or in-process agent |
| Tools, memory, peers | Not exercised | Real tool calls and results are traced |
| Model access | Your OpenRouter or OpenAI key | Included, one ZeroLeaks API key |
| Output | Terminal report or JSON | Stored reports with fixes and PDF export |
| Cost | Free under FSL; you pay your model provider | Paid plans (Pro, Team, Business, Enterprise), with a 14-day Pro trial |
Use the package to harden prompt wording, run local experiments, or add a low-cost prompt check to CI. Use hosted ZeroLeaks to test your agent's behavior, including its tools, memory, and interactions with other agents.
Next steps
- Installation: install via bun or npm
- Quick Start: run a first scan
- CLI: command-line usage
- Scan Engine: advanced configuration