Vibecode ZeroLeaks
track this build5 steps, step by step0%The mechanics here are not exotic: fire a few hundred adversarial prompts at your own chat endpoint, capture the responses, and check whether any of them contain your system prompt, tool schemas or API keys. An agent can build that loop, including an LLM-as-judge scorer and an HTML report, in a single sitting, and it will find the embarrassing stuff on day one. What you cannot one-shot is a probe library that stays current with each new model release and each new jailbreak family, because that is maintained knowledge, not code. There is also a trust angle: 'we ran our own script and found nothing' reads very differently in a security review than a dated third-party report. Build it for your own sanity checks, keep paying if you need something to show someone else.
You are building a lean indie version of ZeroLeaks.
Create the following project files first, then implement the application by following them. Keep the files updated as decisions change. Do not collapse this into a single README or prompt.
===== README.md =====
# ZeroLeaks indie build
## Goal
Build the smallest trustworthy replacement for the core ZeroLeaks workflow for one developer or a tiny team.
## Scope
Runs a versioned pack of extraction and injection probes against your chat endpoint, scores each response for leaked system prompt, secrets or tool definitions, and emits a ranked report with the exact transcripts.
## Quick start
1. Install the documented dependencies.
2. Copy `.env.example` to `.env`.
3. Run the development command chosen during implementation.
4. Complete the acceptance checks in `BUILD_PLAN.md`.
## Honest limits
This build deliberately does not replace:
- A curated, maintained probe corpus that tracks new jailbreak families instead of whatever the agent remembered on build day
- Multi-turn and multi-model attack strategies, including crescendo and encoding tricks, done properly rather than as single-shot prompts
- A third-party report with a date on it that you can hand to a customer or an auditor
- Severity triage and remediation guidance written by someone who has seen a lot of these
- Regression runs on every model or prompt change without you remembering to trigger them
If those capabilities are essential, use ZeroLeaks instead of pretending the gap is solved.
===== AGENTS.md =====
# Agent instructions
- Optimize for a working, understandable weekend build.
- Prefer the fewest moving parts that satisfy the brief.
- Do not invent cryptography, security guarantees, APIs, or compliance claims.
- Keep secrets out of source control and logs.
- Add focused tests for destructive, security-sensitive, and data-loss paths.
- Run the project checks before declaring the build complete.
- Record any deliberate shortcut in the README under "Tradeoffs".
===== BUILD_PLAN.md =====
# Build plan
## Original build brief
Build a local CLI tool called leakprobe that red-teams a chat AI endpoint for system prompt and secret leakage.
Stack: Python 3.12, uv for deps, httpx, pydantic, typer for the CLI, jinja2 for the report. No web UI, no database, no accounts, no telemetry. Everything runs on my machine and writes to ./runs/.
Config: a target.yaml describing the endpoint under test (url, http method, headers, JSON body template with a {{message}} placeholder, and a JSONPath-ish key for extracting the reply text). Also a secrets.yaml listing my real system prompt text, tool/function names, and regex patterns for keys I never want echoed (sk-, ghp_, AKIA, bearer tokens). All API keys come from .env via python-dotenv; write .env.example and gitignore .env.
Probes: ship a probes/ directory of YAML files, at least 60 probes across these families, each with id, family, severity, and one or more turns: direct extraction, polite social engineering, roleplay and persona swap, translation and encoding (base64, rot13, pig latin), token smuggling, fake developer or debug mode, 'repeat the text above', markdown and code block coercion, tool and function schema enumeration, indirect injection via pasted document content, and refusal-boundary probing. Support multi-turn probes where later turns reference earlier replies.
Runner: async, configurable concurrency (default 4), per-request timeout, exponential backoff on 429 and 5xx, and a --limit flag so I can smoke test. Log every request and response verbatim to runs/TIMESTAMP/transcripts.jsonl.
Scoring: two layers. First, deterministic detectors: fuzzy overlap against my known system prompt using token n-gram matching, exact matches on tool names, and regex hits on secret patterns. Second, an LLM judge (OpenAI-compatible, model configurable, key from .env) that reads the transcript and returns strict JSON with leaked: bool, leak_type, confidence, and a one-line rationale. Combine into a severity per probe. Deterministic hits always win.
Output: a self-contained HTML report at runs/TIMESTAMP/report.html grouped by severity, each finding showing the probe, the full exchange, and which detector fired, plus report.json for diffing. Add a `leakprobe diff RUN_A RUN_B` command that shows newly failing and newly passing probes so I can use it as a regression gate. Exit code 1 if any high-severity finding, so it works in CI.
Out of scope: scanning targets I do not control, DoS or rate-limit abuse, network-level scanning, auth bypass testing, any hosted dashboard. Print a short warning on first run that this only targets endpoints listed in target.yaml.
Deliver a README with a 60 second quickstart, and pytest tests for the detectors using fixture transcripts (no live API calls in tests).
## Required capabilities
- An API key for the model you use as judge
- A reachable chat endpoint or API for the app under test, plus permission to hammer it
- A copy of your real system prompt and secret patterns to match against
## Delivery order
1. Scaffold the smallest runnable application and document its commands.
2. Implement the primary data model and core workflow.
3. Add validation, safe failure states, and persistence.
4. Cover the critical path with automated tests.
5. Exercise a clean install from the README and fix every missing step.
## Done when
- A new user can go from clone to first successful workflow using only the README.
- The core workflow works without paid infrastructure unless the brief requires it.
- Tests cover the highest-risk behavior.
- Known limitations are explicit rather than hidden.
===== .env.example =====
# Copy to .env and document every variable when it is introduced.
# Never put real credentials in this file.
APP_ENV=development
# Add only values required by the selected implementation.You are building a lean indie version of ZeroLeaks.
Create the following project files first, then implement the application by following them. Keep the files updated as decisions change. Do not collapse this into a single README or prompt.
===== README.md =====
# ZeroLeaks indie build
## Goal
Build the smallest trustworthy replacement for the core ZeroLeaks workflow for one developer or a tiny team.
## Scope
Runs a versioned pack of extraction and injection probes against your chat endpoint, scores each response for leaked system prompt, secrets or tool definitions, and emits a ranked report with the exact transcripts.
## Quick start
1. Install the documented dependencies.
2. Copy `.env.example` to `.env`.
3. Run the development command chosen during implementation.
4. Complete the acceptance checks in `BUILD_PLAN.md`.
## Honest limits
This build deliberately does not replace:
- A curated, maintained probe corpus that tracks new jailbreak families instead of whatever the agent remembered on build day
- Multi-turn and multi-model attack strategies, including crescendo and encoding tricks, done properly rather than as single-shot prompts
- A third-party report with a date on it that you can hand to a customer or an auditor
- Severity triage and remediation guidance written by someone who has seen a lot of these
- Regression runs on every model or prompt change without you remembering to trigger them
If those capabilities are essential, use ZeroLeaks instead of pretending the gap is solved.
===== AGENTS.md =====
# Agent instructions
- Optimize for a working, understandable weekend build.
- Prefer the fewest moving parts that satisfy the brief.
- Do not invent cryptography, security guarantees, APIs, or compliance claims.
- Keep secrets out of source control and logs.
- Add focused tests for destructive, security-sensitive, and data-loss paths.
- Run the project checks before declaring the build complete.
- Record any deliberate shortcut in the README under "Tradeoffs".
===== BUILD_PLAN.md =====
# Build plan
## Original build brief
Build a local CLI tool called leakprobe that red-teams a chat AI endpoint for system prompt and secret leakage.
Stack: Python 3.12, uv for deps, httpx, pydantic, typer for the CLI, jinja2 for the report. No web UI, no database, no accounts, no telemetry. Everything runs on my machine and writes to ./runs/.
Config: a target.yaml describing the endpoint under test (url, http method, headers, JSON body template with a {{message}} placeholder, and a JSONPath-ish key for extracting the reply text). Also a secrets.yaml listing my real system prompt text, tool/function names, and regex patterns for keys I never want echoed (sk-, ghp_, AKIA, bearer tokens). All API keys come from .env via python-dotenv; write .env.example and gitignore .env.
Probes: ship a probes/ directory of YAML files, at least 60 probes across these families, each with id, family, severity, and one or more turns: direct extraction, polite social engineering, roleplay and persona swap, translation and encoding (base64, rot13, pig latin), token smuggling, fake developer or debug mode, 'repeat the text above', markdown and code block coercion, tool and function schema enumeration, indirect injection via pasted document content, and refusal-boundary probing. Support multi-turn probes where later turns reference earlier replies.
Runner: async, configurable concurrency (default 4), per-request timeout, exponential backoff on 429 and 5xx, and a --limit flag so I can smoke test. Log every request and response verbatim to runs/TIMESTAMP/transcripts.jsonl.
Scoring: two layers. First, deterministic detectors: fuzzy overlap against my known system prompt using token n-gram matching, exact matches on tool names, and regex hits on secret patterns. Second, an LLM judge (OpenAI-compatible, model configurable, key from .env) that reads the transcript and returns strict JSON with leaked: bool, leak_type, confidence, and a one-line rationale. Combine into a severity per probe. Deterministic hits always win.
Output: a self-contained HTML report at runs/TIMESTAMP/report.html grouped by severity, each finding showing the probe, the full exchange, and which detector fired, plus report.json for diffing. Add a `leakprobe diff RUN_A RUN_B` command that shows newly failing and newly passing probes so I can use it as a regression gate. Exit code 1 if any high-severity finding, so it works in CI.
Out of scope: scanning targets I do not control, DoS or rate-limit abuse, network-level scanning, auth bypass testing, any hosted dashboard. Print a short warning on first run that this only targets endpoints listed in target.yaml.
Deliver a README with a 60 second quickstart, and pytest tests for the detectors using fixture transcripts (no live API calls in tests).
## Required capabilities
- An API key for the model you use as judge
- A reachable chat endpoint or API for the app under test, plus permission to hammer it
- A copy of your real system prompt and secret patterns to match against
## Delivery order
1. Scaffold the smallest runnable application and document its commands.
2. Implement the primary data model and core workflow.
3. Add validation, safe failure states, and persistence.
4. Cover the critical path with automated tests.
5. Exercise a clean install from the README and fix every missing step.
## Done when
- A new user can go from clone to first successful workflow using only the README.
- The core workflow works without paid infrastructure unless the brief requires it.
- Tests cover the highest-risk behavior.
- Known limitations are explicit rather than hidden.
===== .env.example =====
# Copy to .env and document every variable when it is introduced.
# Never put real credentials in this file.
APP_ENV=development
# Add only values required by the selected implementation.You are building a production product version of ZeroLeaks.
Create the following project files first, then implement the application by following them. Keep the files updated as decisions change. Do not collapse this into a single README or prompt.
===== PRODUCT.md =====
# ZeroLeaks product brief
## Problem
The mechanics here are not exotic: fire a few hundred adversarial prompts at your own chat endpoint, capture the responses, and check whether any of them contain your system prompt, tool schemas or API keys. An agent can build that loop, including an LLM-as-judge scorer and an HTML report, in a single sitting, and it will find the embarrassing stuff on day one. What you cannot one-shot is a probe library that stays current with each new model release and each new jailbreak family, because that is maintained knowledge, not code. There is also a trust angle: 'we ran our own script and found nothing' reads very differently in a security review than a dated third-party report. Build it for your own sanity checks, keep paying if you need something to show someone else.
## Product outcome
Runs a versioned pack of extraction and injection probes against your chat endpoint, scores each response for leaked system prompt, secrets or tool definitions, and emits a ranked report with the exact transcripts.
## Target user
A serious builder who needs a maintainable product foundation rather than a one-off demo.
## Required capabilities
- An API key for the model you use as judge
- A reachable chat endpoint or API for the app under test, plus permission to hammer it
- A copy of your real system prompt and secret patterns to match against
## Explicit non-goals for v1
- A curated, maintained probe corpus that tracks new jailbreak families instead of whatever the agent remembered on build day
- Multi-turn and multi-model attack strategies, including crescendo and encoding tricks, done properly rather than as single-shot prompts
- A third-party report with a date on it that you can hand to a customer or an auditor
- Severity triage and remediation guidance written by someone who has seen a lot of these
- Regression runs on every model or prompt change without you remembering to trigger them
## Success criteria
- The primary workflow is measurable end to end.
- Setup is reproducible in a clean environment.
- Failure, recovery, and support paths are documented.
- Product claims match what the implementation actually guarantees.
===== ARCHITECTURE.md =====
# Architecture
## Starting brief
Build a local CLI tool called leakprobe that red-teams a chat AI endpoint for system prompt and secret leakage.
Stack: Python 3.12, uv for deps, httpx, pydantic, typer for the CLI, jinja2 for the report. No web UI, no database, no accounts, no telemetry. Everything runs on my machine and writes to ./runs/.
Config: a target.yaml describing the endpoint under test (url, http method, headers, JSON body template with a {{message}} placeholder, and a JSONPath-ish key for extracting the reply text). Also a secrets.yaml listing my real system prompt text, tool/function names, and regex patterns for keys I never want echoed (sk-, ghp_, AKIA, bearer tokens). All API keys come from .env via python-dotenv; write .env.example and gitignore .env.
Probes: ship a probes/ directory of YAML files, at least 60 probes across these families, each with id, family, severity, and one or more turns: direct extraction, polite social engineering, roleplay and persona swap, translation and encoding (base64, rot13, pig latin), token smuggling, fake developer or debug mode, 'repeat the text above', markdown and code block coercion, tool and function schema enumeration, indirect injection via pasted document content, and refusal-boundary probing. Support multi-turn probes where later turns reference earlier replies.
Runner: async, configurable concurrency (default 4), per-request timeout, exponential backoff on 429 and 5xx, and a --limit flag so I can smoke test. Log every request and response verbatim to runs/TIMESTAMP/transcripts.jsonl.
Scoring: two layers. First, deterministic detectors: fuzzy overlap against my known system prompt using token n-gram matching, exact matches on tool names, and regex hits on secret patterns. Second, an LLM judge (OpenAI-compatible, model configurable, key from .env) that reads the transcript and returns strict JSON with leaked: bool, leak_type, confidence, and a one-line rationale. Combine into a severity per probe. Deterministic hits always win.
Output: a self-contained HTML report at runs/TIMESTAMP/report.html grouped by severity, each finding showing the probe, the full exchange, and which detector fired, plus report.json for diffing. Add a `leakprobe diff RUN_A RUN_B` command that shows newly failing and newly passing probes so I can use it as a regression gate. Exit code 1 if any high-severity finding, so it works in CI.
Out of scope: scanning targets I do not control, DoS or rate-limit abuse, network-level scanning, auth bypass testing, any hosted dashboard. Print a short warning on first run that this only targets endpoints listed in target.yaml.
Deliver a README with a 60 second quickstart, and pytest tests for the detectors using fixture transcripts (no live API calls in tests).
## Boundaries
Separate the product into replaceable modules for interface, application logic, persistence, external integrations, and operational concerns. Keep domain logic independent from delivery frameworks and vendors.
## Production baseline
- Configuration: validated at startup with safe local defaults where possible.
- Security: least privilege, input validation, secret redaction, rate limits on abuse-prone paths, and no invented security primitives.
- Data: explicit schema and migrations, transactional writes where integrity matters, backup and restore instructions.
- Integrations: adapters around third-party providers, idempotent webhook or job processing, bounded retries, and timeouts.
- Observability: structured logs with request or operation IDs, an error-tracking hook, and health/readiness checks where a server exists.
- Quality: unit tests for domain rules, integration tests at module boundaries, and one end-to-end critical-path test.
## Decision records
For each major dependency, document why it was chosen, its failure mode, and how it can be replaced. Do not introduce infrastructure until a requirement justifies it.
===== AGENTS.md =====
# Agent instructions
- Read `PRODUCT.md` and `ARCHITECTURE.md` before changing code.
- Implement milestone by milestone; keep each change reviewable and leave the application runnable.
- Treat authentication, payments, encryption, imports, webhooks, and destructive actions as high-risk boundaries when present.
- Never invent cryptography or silently weaken a requirement to make a test pass.
- Use provider interfaces for external services and deterministic fakes in tests.
- Add migrations and rollback or recovery notes for persistent data changes.
- Log useful operational context without credentials, tokens, passwords, or personal data.
- Update documentation and run all checks before completing a milestone.
===== MILESTONES.md =====
# Delivery milestones
## M0 — Decisions and scaffold
- Confirm the runtime, persistence model, threat boundaries, and deployment target.
- Create a reproducible local environment and continuous checks.
## M1 — Core workflow
- Implement the smallest end-to-end product path with validation and tests.
- Keep integrations behind interfaces.
## M2 — Trust layer
- Add secure failure behavior, recovery paths, audit-relevant events, and data safeguards.
- Test abuse cases and destructive operations.
## M3 — Operability
- Add structured logs, error reporting hooks, health signals, backup/restore documentation, and deployment configuration.
## M4 — Release gate
- Run a clean-install test, critical-path end-to-end test, dependency review, and documented rollback exercise.
- Compare the shipped behavior with `PRODUCT.md` and publish remaining limitations.
===== OPERATIONS.md =====
# Operations
## Before release
- Validate configuration and secrets at startup.
- Define backup, restore, and rollback procedures and test them.
- Document logs, error tracking, health signals, and alert ownership.
- Set dependency update and vulnerability review expectations.
## Incident checklist
1. Contain the issue without destroying evidence or user data.
2. Record the timeline and affected scope.
3. Rotate exposed secrets and revoke compromised sessions or credentials.
4. Restore from a verified source when needed.
5. Document the root cause, remediation, and regression test.
## Launch constraint
Do not market omitted ZeroLeaks capabilities as implemented. The v1 non-goals in `PRODUCT.md` remain user-visible limitations until they are deliberately delivered.# ZeroLeaks indie build ## Goal Build the smallest trustworthy replacement for the core ZeroLeaks workflow for one developer or a tiny team. ## Scope Runs a versioned pack of extraction and injection probes against your chat endpoint, scores each response for leaked system prompt, secrets or tool definitions, and emits a ranked report with the exact transcripts. ## Quick start 1. Install the documented dependencies. 2. Copy `.env.example` to `.env`. 3. Run the development command chosen during implementation. 4. Complete the acceptance checks in `BUILD_PLAN.md`. ## Honest limits This build deliberately does not replace: - A curated, maintained probe corpus that tracks new jailbreak families instead of whatever the agent remembered on build day - Multi-turn and multi-model attack strategies, including crescendo and encoding tricks, done properly rather than as single-shot prompts - A third-party report with a date on it that you can hand to a customer or an auditor - Severity triage and remediation guidance written by someone who has seen a lot of these - Regression runs on every model or prompt change without you remembering to trigger them If those capabilities are essential, use ZeroLeaks instead of pretending the gap is solved.
# Agent instructions - Optimize for a working, understandable weekend build. - Prefer the fewest moving parts that satisfy the brief. - Do not invent cryptography, security guarantees, APIs, or compliance claims. - Keep secrets out of source control and logs. - Add focused tests for destructive, security-sensitive, and data-loss paths. - Run the project checks before declaring the build complete. - Record any deliberate shortcut in the README under "Tradeoffs".
# Build plan
## Original build brief
Build a local CLI tool called leakprobe that red-teams a chat AI endpoint for system prompt and secret leakage.
Stack: Python 3.12, uv for deps, httpx, pydantic, typer for the CLI, jinja2 for the report. No web UI, no database, no accounts, no telemetry. Everything runs on my machine and writes to ./runs/.
Config: a target.yaml describing the endpoint under test (url, http method, headers, JSON body template with a {{message}} placeholder, and a JSONPath-ish key for extracting the reply text). Also a secrets.yaml listing my real system prompt text, tool/function names, and regex patterns for keys I never want echoed (sk-, ghp_, AKIA, bearer tokens). All API keys come from .env via python-dotenv; write .env.example and gitignore .env.
Probes: ship a probes/ directory of YAML files, at least 60 probes across these families, each with id, family, severity, and one or more turns: direct extraction, polite social engineering, roleplay and persona swap, translation and encoding (base64, rot13, pig latin), token smuggling, fake developer or debug mode, 'repeat the text above', markdown and code block coercion, tool and function schema enumeration, indirect injection via pasted document content, and refusal-boundary probing. Support multi-turn probes where later turns reference earlier replies.
Runner: async, configurable concurrency (default 4), per-request timeout, exponential backoff on 429 and 5xx, and a --limit flag so I can smoke test. Log every request and response verbatim to runs/TIMESTAMP/transcripts.jsonl.
Scoring: two layers. First, deterministic detectors: fuzzy overlap against my known system prompt using token n-gram matching, exact matches on tool names, and regex hits on secret patterns. Second, an LLM judge (OpenAI-compatible, model configurable, key from .env) that reads the transcript and returns strict JSON with leaked: bool, leak_type, confidence, and a one-line rationale. Combine into a severity per probe. Deterministic hits always win.
Output: a self-contained HTML report at runs/TIMESTAMP/report.html grouped by severity, each finding showing the probe, the full exchange, and which detector fired, plus report.json for diffing. Add a `leakprobe diff RUN_A RUN_B` command that shows newly failing and newly passing probes so I can use it as a regression gate. Exit code 1 if any high-severity finding, so it works in CI.
Out of scope: scanning targets I do not control, DoS or rate-limit abuse, network-level scanning, auth bypass testing, any hosted dashboard. Print a short warning on first run that this only targets endpoints listed in target.yaml.
Deliver a README with a 60 second quickstart, and pytest tests for the detectors using fixture transcripts (no live API calls in tests).
## Required capabilities
- An API key for the model you use as judge
- A reachable chat endpoint or API for the app under test, plus permission to hammer it
- A copy of your real system prompt and secret patterns to match against
## Delivery order
1. Scaffold the smallest runnable application and document its commands.
2. Implement the primary data model and core workflow.
3. Add validation, safe failure states, and persistence.
4. Cover the critical path with automated tests.
5. Exercise a clean install from the README and fix every missing step.
## Done when
- A new user can go from clone to first successful workflow using only the README.
- The core workflow works without paid infrastructure unless the brief requires it.
- Tests cover the highest-risk behavior.
- Known limitations are explicit rather than hidden.# Copy to .env and document every variable when it is introduced. # Never put real credentials in this file. APP_ENV=development # Add only values required by the selected implementation.
# ZeroLeaks product brief ## Problem The mechanics here are not exotic: fire a few hundred adversarial prompts at your own chat endpoint, capture the responses, and check whether any of them contain your system prompt, tool schemas or API keys. An agent can build that loop, including an LLM-as-judge scorer and an HTML report, in a single sitting, and it will find the embarrassing stuff on day one. What you cannot one-shot is a probe library that stays current with each new model release and each new jailbreak family, because that is maintained knowledge, not code. There is also a trust angle: 'we ran our own script and found nothing' reads very differently in a security review than a dated third-party report. Build it for your own sanity checks, keep paying if you need something to show someone else. ## Product outcome Runs a versioned pack of extraction and injection probes against your chat endpoint, scores each response for leaked system prompt, secrets or tool definitions, and emits a ranked report with the exact transcripts. ## Target user A serious builder who needs a maintainable product foundation rather than a one-off demo. ## Required capabilities - An API key for the model you use as judge - A reachable chat endpoint or API for the app under test, plus permission to hammer it - A copy of your real system prompt and secret patterns to match against ## Explicit non-goals for v1 - A curated, maintained probe corpus that tracks new jailbreak families instead of whatever the agent remembered on build day - Multi-turn and multi-model attack strategies, including crescendo and encoding tricks, done properly rather than as single-shot prompts - A third-party report with a date on it that you can hand to a customer or an auditor - Severity triage and remediation guidance written by someone who has seen a lot of these - Regression runs on every model or prompt change without you remembering to trigger them ## Success criteria - The primary workflow is measurable end to end. - Setup is reproducible in a clean environment. - Failure, recovery, and support paths are documented. - Product claims match what the implementation actually guarantees.
# Architecture
## Starting brief
Build a local CLI tool called leakprobe that red-teams a chat AI endpoint for system prompt and secret leakage.
Stack: Python 3.12, uv for deps, httpx, pydantic, typer for the CLI, jinja2 for the report. No web UI, no database, no accounts, no telemetry. Everything runs on my machine and writes to ./runs/.
Config: a target.yaml describing the endpoint under test (url, http method, headers, JSON body template with a {{message}} placeholder, and a JSONPath-ish key for extracting the reply text). Also a secrets.yaml listing my real system prompt text, tool/function names, and regex patterns for keys I never want echoed (sk-, ghp_, AKIA, bearer tokens). All API keys come from .env via python-dotenv; write .env.example and gitignore .env.
Probes: ship a probes/ directory of YAML files, at least 60 probes across these families, each with id, family, severity, and one or more turns: direct extraction, polite social engineering, roleplay and persona swap, translation and encoding (base64, rot13, pig latin), token smuggling, fake developer or debug mode, 'repeat the text above', markdown and code block coercion, tool and function schema enumeration, indirect injection via pasted document content, and refusal-boundary probing. Support multi-turn probes where later turns reference earlier replies.
Runner: async, configurable concurrency (default 4), per-request timeout, exponential backoff on 429 and 5xx, and a --limit flag so I can smoke test. Log every request and response verbatim to runs/TIMESTAMP/transcripts.jsonl.
Scoring: two layers. First, deterministic detectors: fuzzy overlap against my known system prompt using token n-gram matching, exact matches on tool names, and regex hits on secret patterns. Second, an LLM judge (OpenAI-compatible, model configurable, key from .env) that reads the transcript and returns strict JSON with leaked: bool, leak_type, confidence, and a one-line rationale. Combine into a severity per probe. Deterministic hits always win.
Output: a self-contained HTML report at runs/TIMESTAMP/report.html grouped by severity, each finding showing the probe, the full exchange, and which detector fired, plus report.json for diffing. Add a `leakprobe diff RUN_A RUN_B` command that shows newly failing and newly passing probes so I can use it as a regression gate. Exit code 1 if any high-severity finding, so it works in CI.
Out of scope: scanning targets I do not control, DoS or rate-limit abuse, network-level scanning, auth bypass testing, any hosted dashboard. Print a short warning on first run that this only targets endpoints listed in target.yaml.
Deliver a README with a 60 second quickstart, and pytest tests for the detectors using fixture transcripts (no live API calls in tests).
## Boundaries
Separate the product into replaceable modules for interface, application logic, persistence, external integrations, and operational concerns. Keep domain logic independent from delivery frameworks and vendors.
## Production baseline
- Configuration: validated at startup with safe local defaults where possible.
- Security: least privilege, input validation, secret redaction, rate limits on abuse-prone paths, and no invented security primitives.
- Data: explicit schema and migrations, transactional writes where integrity matters, backup and restore instructions.
- Integrations: adapters around third-party providers, idempotent webhook or job processing, bounded retries, and timeouts.
- Observability: structured logs with request or operation IDs, an error-tracking hook, and health/readiness checks where a server exists.
- Quality: unit tests for domain rules, integration tests at module boundaries, and one end-to-end critical-path test.
## Decision records
For each major dependency, document why it was chosen, its failure mode, and how it can be replaced. Do not introduce infrastructure until a requirement justifies it.# Agent instructions - Read `PRODUCT.md` and `ARCHITECTURE.md` before changing code. - Implement milestone by milestone; keep each change reviewable and leave the application runnable. - Treat authentication, payments, encryption, imports, webhooks, and destructive actions as high-risk boundaries when present. - Never invent cryptography or silently weaken a requirement to make a test pass. - Use provider interfaces for external services and deterministic fakes in tests. - Add migrations and rollback or recovery notes for persistent data changes. - Log useful operational context without credentials, tokens, passwords, or personal data. - Update documentation and run all checks before completing a milestone.
# Delivery milestones ## M0 — Decisions and scaffold - Confirm the runtime, persistence model, threat boundaries, and deployment target. - Create a reproducible local environment and continuous checks. ## M1 — Core workflow - Implement the smallest end-to-end product path with validation and tests. - Keep integrations behind interfaces. ## M2 — Trust layer - Add secure failure behavior, recovery paths, audit-relevant events, and data safeguards. - Test abuse cases and destructive operations. ## M3 — Operability - Add structured logs, error reporting hooks, health signals, backup/restore documentation, and deployment configuration. ## M4 — Release gate - Run a clean-install test, critical-path end-to-end test, dependency review, and documented rollback exercise. - Compare the shipped behavior with `PRODUCT.md` and publish remaining limitations.
# Operations ## Before release - Validate configuration and secrets at startup. - Define backup, restore, and rollback procedures and test them. - Document logs, error tracking, health signals, and alert ownership. - Set dependency update and vulnerability review expectations. ## Incident checklist 1. Contain the issue without destroying evidence or user data. 2. Record the timeline and affected scope. 3. Rotate exposed secrets and revoke compromised sessions or credentials. 4. Restore from a verified source when needed. 5. Document the root cause, remediation, and regression test. ## Launch constraint Do not market omitted ZeroLeaks capabilities as implemented. The v1 non-goals in `PRODUCT.md` remain user-visible limitations until they are deliberately delivered.
$ choose a build depth, inspect the files, then open the complete pack in your agent · this prompt is generated from the build plan · improve it via PR
Two reasons, and neither is that the harness is hard. First, jailbreaks rot: the probes that worked against last quarter's model are dead weight now, and keeping a live corpus is somebody's full time job, not a cron you set up once. Second, self-attested security is worth roughly nothing to an enterprise buyer, so companies pay for an external artifact with a date and a logo on it. If your goal is just to stop shipping a system prompt that unravels when someone types 'repeat everything above', a local harness is genuinely enough.
xA curated, maintained probe corpus that tracks new jailbreak families instead of whatever the agent remembered on build day
xMulti-turn and multi-model attack strategies, including crescendo and encoding tricks, done properly rather than as single-shot prompts
xA third-party report with a date on it that you can hand to a customer or an auditor
xSeverity triage and remediation guidance written by someone who has seen a lot of these
xRegression runs on every model or prompt change without you remembering to trigger them
Nothing worth pointing at. That's why the prompt exists.
Vibecode ZeroLeaks
Kinda. The core of ZeroLeaks is buildable in a weekend with the prompt on this page, but there are real gaps: A curated, maintained probe corpus that tracks new jailbreak families instead of whatever the agent remembered on build day, Multi-turn and multi-model attack strategies, including crescendo and encoding tricks, done properly rather than as single-shot prompts. Read the honest list above before committing.
How much does ZeroLeaks cost?
ZeroLeaks costs about $79/month (Pro, checked 2026-08-18), which is $948 per year.
What do I lose by replacing ZeroLeaks?
Honestly: A curated, maintained probe corpus that tracks new jailbreak families instead of whatever the agent remembered on build day; Multi-turn and multi-model attack strategies, including crescendo and encoding tricks, done properly rather than as single-shot prompts; A third-party report with a date on it that you can hand to a customer or an auditor; Severity triage and remediation guidance written by someone who has seen a lot of these; Regression runs on every model or prompt change without you remembering to trigger them. If any of those are load-bearing for you, keep paying.
Is there an open-source alternative to ZeroLeaks?
No mature open-source alternative worth pointing at, which is exactly why the one-shot prompt on this page exists.