ZeroLeaksDocs
ZeroLeaks Package

The zeroleaks package

Source-available CLI and TypeScript library that tests a system prompt for extraction and prompt injection. Runs locally with your own OpenRouter or OpenAI key.

npm version

The zeroleaks npm package is a standalone scanner for system prompts. It loads your prompt into a model you choose and tests that model in two ways:

  • Extraction: a multi-turn attacker tries to get the model to reveal its own system prompt.
  • Injection: a set of behavioral probes tries to get the model to follow planted instructions (in a document, a fake admin message, or a tool description) or to agree to misuse a tool. A judge model decides whether the target complied.

It runs on your machine with your own model keys. The source is at github.com/ZeroLeaks/zeroleaks. These pages describe version 1.5.0.

License

The package is licensed under FSL-1.1-Apache-2.0, the Functional Source License. You can use, modify, and redistribute it for any purpose except offering a competing commercial product or service. Each release becomes available under Apache 2.0 two years after it is published or on January 21, 2028, whichever comes first.

Features

  • runSecurityScan(): one call that takes a system prompt and returns a score, findings, and recommendations
  • createScanEngine(): full control over turns, tree depth, branching, Best-of-N, and callbacks
  • CLI: zeroleaks scan, zeroleaks probes, zeroleaks categories, zeroleaks techniques
  • Agents: Strategist, Attacker, Evaluator, Mutator, Inspector, Orchestrator, and an injection compliance judge
  • Probe library: extraction probes across direct, encoding, persona, social, crescendo, many-shot, policy puppetry, and more, plus 78 behavioral injection probes in six categories
  • Knowledge base: documented techniques, payload templates, exfiltration vectors, and defense bypass methods, each with a source
  • Models: any OpenRouter model, OpenAI models directly with OPENAI_API_KEY, or any OpenAI-compatible endpoint (Ollama, vLLM, a gateway) with --base-url

Architecture

The extraction scan follows TAP (Tree of Attacks with Pruning) with several cooperating agents:

AgentRole
StrategistReads the conversation, picks the next strategy, and moves between attack phases
AttackerGenerates attack prompts from the strategy and evaluator feedback
EvaluatorDecides how much of the system prompt leaked in each response
MutatorProduces Best-of-N variations of an attack
InspectorFingerprints known guardrails (TombRaider-style) and suggests bypasses
OrchestratorRuns scripted multi-turn sequences (Siren, Echo Chamber, TombRaider)
InjectionEvaluatorJudges whether the target complied with an injection probe (full, partial, or refused)

Research foundation

zeroleaks draws on:

  • TAP: Tree of Attacks with Pruning (Mehrotra et al.)
  • PAIR: Prompt Automatic Iterative Refinement
  • Crescendo: multi-turn gradual escalation
  • TombRaider: dual-agent defense fingerprinting
  • Siren: multi-turn human jailbreak simulation
  • Echo Chamber: escalation through agreement
  • Best-of-N: sampling-based jailbreaking
  • Garak: probes adapted from NVIDIA's garak
  • AgentDojo, InjecAgent, JailbreakBench, HarmBench, promptfoo, OWASP LLM Top 10: sources for the behavioral injection probes

Use cases

  • CI: scan a system prompt when it changes. zeroleaks scan exits with 0 when every check passed, 1 when it finds a vulnerability, and 2 when it can't reach a verdict.
  • Local development: test prompt wording without an account.
  • Custom tooling: call the engine, agents, or probe library from your own code.
  • Research: read the probe library and knowledge base programmatically.

zeroleaks package or hosted ZeroLeaks?

The zeroleaks package tests a system prompt. The target is a model you specify, running your prompt with no tools, memory, or application code. A finding means that model, given that prompt, leaked its instructions or followed an injected instruction in its reply.

Hosted ZeroLeaks tests a running agent. Through @zeroleaks/sdk runtime scans or endpoint scans, every probe goes through your real agent: its model settings, tools, memory, credentials, and the other agents it talks to. It profiles the agent, derives the rules the agent must never break, and checks whether real tool calls break them. Outbound payloads are contained so a successful exploit reaches a sink ZeroLeaks controls. Findings come with a least-privilege policy or configuration fix, reports are stored, and later scans rerun the same attacks to catch regressions.

Hosted agent scans no longer run a system-prompt extraction loop. They run a secrets-in-context track instead, which scores what sensitive data leaked (credentials, internal hosts, tool schemas, approval rules), not whether the prompt was reproduced word for word. Hosted prompt scans are retired.

zeroleaks packageHosted ZeroLeaks
TargetA model running your system promptYour deployed or in-process agent
Tools, memory, peersNot exercisedReal tool calls and results are traced
Model accessYour OpenRouter or OpenAI keyIncluded, one ZeroLeaks API key
OutputTerminal report or JSONStored reports with fixes and PDF export
CostFree under FSL; you pay your model providerPaid plans (Pro, Team, Business, Enterprise), with a 14-day Pro trial

Use the package to harden prompt wording, run local experiments, or add a low-cost prompt check to CI. Use hosted ZeroLeaks to test your agent's behavior, including its tools, memory, and interactions with other agents.

Next steps

On this page