A deliberately vulnerable MLOps platform, built to be broken. A realistic Flask application with 82 security bugs planted in it on purpose: target practice for scanners, for agents, and for anyone who rates their own code review highly.
Never deploy this, and never point it at anything you care about. Every bug in it is real and exploitable. It is for local research and teaching.
Architecture · Scoreboard · Security policy
Claude Opus 5 █████████████████████████████████████░░░ 91% (single)
Claude Sonnet 5, 5-region swee ████████████████████████████████░░░░░░░░ 80% (sweep)
Rowan (hedgerow.dev) ████████████████████████░░░░░░░░░░░░░░░░ 61% (single)
Claude Haiku 4.5, 5-region swe ████████████████████░░░░░░░░░░░░░░░░░░░░ 50% (sweep)
CodeQL ███████████░░░░░░░░░░░░░░░░░░░░░░░░░░░░░ 27% (single)
Bandit █████████░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░ 23% (single)
Semgrep █████████░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░ 23% (single)
Every one of the 82 bugs is real and has a test that proves it's exploitable.
Every score above is verified: raw output committed, every finding mapped
to a manifest id, arithmetic done by
benchmarks/score.py rather than typed in. single is
one deterministic pass; sweep reviews the app in regions, which materially
helps recall and is why it isn't compared like-for-like against a scanner.
Older self-reported claims (Claude Opus 57%, GPT-5.5 44%, Kimi K3 35% and others) have no artifact behind them, and some were reviewed against blind copies looser about hiding the answer key than the current exporter. They are listed but not ranked in SCOREBOARD.md.
It looks like a real self-hosted MLOps platform: model registry, dataset uploads, experiment tracking, a prediction API, and a built-in LLM assistant, all on Flask and SQLite. It behaves like one, too. The bugs are written down in an answer key kept well away from the code, so you can check what a tool actually found, not what it claims.
Langfail exists to be broken into. It is not production software. Run it on your own machine, disconnected from anything you'd miss, and keep exploit code in a throwaway environment. Yes, really.
Specifically, on the machine you run it on:
mcp-serve-httplistens on0.0.0.0:8765with no authentication by default. That is an unauthenticated SQL / file-read /evalendpoint reachable by anything on your network. That's the bug. Don't run it on a shared or untrusted network.- Some bugs write and delete outside the project directory. The planted zip-slip, the artifact metadata path, and the retention sweep all take attacker-controlled paths. Run the exploit tests in a container or a VM.
- Don't run this on a cloud instance. The SSRF plants will happily reach
169.254.169.254and return your instance credentials.- The XML descriptor parser resolves external entities, so it makes real outbound network requests.
- The test suite is POSIX-only: it reads
/etc/passwdand shells out to/bin/sh.
Difficulty runs from "one obvious line" to "only visible once you run the app and actually attack it." By theme:
| Theme | The gist |
|---|---|
| Classic web bugs | unsafe file loading, SSRF, path traversal, zip/tar slip, SQL and command injection, template injection, IDOR, unsafe dynamic imports, XXE, insecure config loading |
| LLM / AI-agent bugs | assistant memory poisoned across sessions, invisible Unicode smuggling instructions past filters, markdown images used to exfiltrate, the assistant tricked into writing unsafe SQL and code, no ceiling on billed LLM calls, datasets that run code on load (the trust_remote_code risk), private data leaking into a third-party API call, and MCP auth/poisoning bugs |
| Auth & accounts | self-assigned admin at signup, session tokens accepted in the URL, a signature check that doesn't check, guessable password-reset codes, weak hashing, a regex that hangs the server, open redirects, stored XSS via an uploaded image, and a timing-attackable comparison |
| Supply chain & privacy | several flavors of unsafe deserialization, a bypass of the app's own "safe" unpickler, code that runs when you import a model repo, a plugin that fires on the next restart, plus model extraction, membership inference, and training-data poisoning |
| Agentic | the agent talked into installing a malicious package, a tool description swapped out after a human approved it, a faked human confirmation, one tenant's data leaking into another's conversation, and a "just update my preferences" call that quietly switches off a security check elsewhere |
| IDOR deep dive | one bug class in every shape it takes: no check, a check on the wrong code path, checks at several levels of a hierarchy, each sitting next to a correct version, so a reviewer has to actually read rather than grep for a keyword |
| The dashboard | the HTML side has its own set: open redirect, reflected and stored XSS, CSRF, clickjacking, and a session cookie any script can read |
| Misconfiguration | 5 findings that aren't code at all: Ollama, TorchServe and Triton set up badly in deploy/, plus debug mode and default secrets |
It's a benchmark. It stresses two things on purpose:
- Following untrusted data across a whole codebase, not one file. Input
arrives in
langfail/api/and ends up somewhere dangerous inlangfail/services/,langfail/ml/, orlangfail/workers/, but rarely in a straight line. Paths detour through a database write-then-read, a background job queue, a serialize/deserialize round trip. Some pass through a helper inlangfail/core/security.pythat looks protective and isn't. Those are traps for anything that checks "is there a sanitizer here?" without asking whether it works. - How an AI agent handles security bugs: finding them, and falling for
them. The built-in assistant (
langfail/agent/) can run SQL, read files, fetch URLs, and do math. It can be manipulated directly by what a user types and indirectly by instructions hidden in data it reads.
The bugs don't announce themselves. The code has type hints, docstrings, and
passing tests, with no naming or comments that give the game away. The answer
key, every bug and exactly how to trigger it, lives in a separate file,
benchmarks/ground_truth.yaml, so it never
leaks into the app.
ARCHITECTURE.md has the diagrams.
langfail has been public long enough that a tool can score well by recognising
it rather than analysing it: the symbol names, file layout, and library calls
are all memorisable. So there is a second corpus, modelbay/, that
carries the same planted vulnerabilities with everything a memoriser keys on
changed.
Each bug is mutated along six axes while its exploitability is preserved: rename (functions, models, keys, the package itself), wrap the dangerous API in an equivalently dangerous shim, relocate the source and sink to different files and layers, rewrite each broken guard into a different-looking but equally broken form, ship a safe twin for every mutation, and hold the answer key out of the tree. A scanner has to re-derive each finding from data flow, not from a remembered fingerprint.
The corpus is public; the answer key, the proofs, and the per-tool results are
held out and released after a scored run, so a frozen rule set cannot be tuned
to the set. Design and rationale:
docs/adr/0001-blinded-mutation-benchmark.md.
Once the key is present locally, score a tool with
python benchmarks/score.py --suite mutant results/<tool>.yaml.
What it shows. Mutation barely moves a dataflow (taint) engine or a frontier LLM reviewer: both re-derive most of the same bugs under the changed surface. It hits signature- and pattern-matching tools much harder, because a renamed sink or a wrapped library call no longer matches the rule that was looking for it. Per-tool figures ship with the results on release.
Don't point anything at this repo directly. The answer key sits right next to the code, and anything that can read files will find it.
Generate a blind copy first: same code, minus the answer key, exploit tests, and every doc that narrates a bug:
python scripts/export_blind_copy.py /path/to/an/empty/directoryPoint your scanner or agent there, then compare findings against
benchmarks/ground_truth.yaml. See
SCOREBOARD.md for the exact scoring
method, plus real numbers from several SAST tools and LLMs run this way.
For pentesting a live target with no source access, start the app and hand over nothing but a URL and a login:
flask --app langfail seed # demo users: admin/admin123, alice/alice123, bob/bob123
flask --app langfail run # http://127.0.0.1:5000Give the agent http://127.0.0.1:5000 and one of those logins, or let it
register its own, since signup is open to anyone. No source, no hints, no
answer key. Scoring: SCOREBOARD.md.
| Path | What's here |
|---|---|
langfail/api/ |
Flask routes, where untrusted input enters (taint sources). Includes authz_demo.py, the IDOR deep-dive tier: vulnerable and safe checks side by side |
langfail/ui/ |
the HTML dashboard: open redirect, reflected/stored XSS, CSRF, clickjacking, a JS-readable cookie |
langfail/services/ |
business logic, and most of the dangerous operations (sinks: SQL, outbound HTTP, templates, file I/O) |
langfail/ml/ |
model save/load, dataset extraction, format conversion, metrics |
langfail/workers/ |
the job queue and its worker: sinks that only fire later, in another process |
langfail/agent/ |
the assistant's backends, tools, and reasoning loop: the prompt-injection surface |
langfail/mcp_server.py |
the same tools exposed over MCP |
langfail/core/ |
config, database, auth/JWT, and the security helpers with deliberate gaps |
benchmarks/ground_truth.yaml |
the answer key: every bug, with its full source-to-sink path |
exploits/ |
runnable proof-of-concept scripts for the multi-step chains |
tests/ |
ordinary tests, plus one proof-of-exploit per planted bug (test_exploits*.py) |
deploy/docker-compose.yml |
insecure configs for real ML-serving tools, descriptive only. langfail never starts them |
python -m venv .venv && . .venv/bin/activate # note: avoid a repo path containing ':'
pip install -e . # add [ml] for numpy/pandas/joblib
export LANGFAIL_JWT_SECRET="a-long-enough-dev-secret-32-bytes!!"
flask --app langfail seed # demo users: admin/admin123, alice/alice123, bob/bob123
flask --app langfail run # http://127.0.0.1:5000
flask --app langfail worker # in another shell: drains the job queue
flask --app langfail mcp-serve # optional: pip install -e ".[mcp]"; MCP tools over stdio
flask --app langfail mcp-serve-http # ⚠️ binds 0.0.0.0:8765, NO AUTH by default, see the warning aboveHealth check: curl localhost:5000/health.
The installed langfail console script starts the same dev server, but it's
only a shortcut: seed, worker, and mcp-serve need the flask --app langfail form.
Prefer uv? uv venv && uv pip install -e ".[dev,ml]" --python .venv/bin/python works the same. One catch: a uv venv
skips pip/setuptools/wheel, which the V58 slopsquatting PoC needs to
build a local package. If that one test fails to build, uv pip install pip setuptools wheel --python .venv/bin/python.
Local and pluggable, no cloud API:
LANGFAIL_LLM_BACKEND=stub(default): deterministic and offline; used for CI and reproducible scoring.LANGFAIL_LLM_BACKEND=ollamawithLANGFAIL_LLM_MODEL=llama3.1: a local Ollama server (LANGFAIL_LLM_OLLAMA_URL), Metal-accelerated on macOS, for prompt injection that behaves like the real thing.
PYTHONPATH=. pytest -q # 153 tests (6 skipped without the mcp/lxml extras): proof that every planted bug really is exploitable
PYTHONPATH=. python exploits/chain_a_ssrf_to_rce.py # multi-step: SSRF leads to remote code execution
PYTHONPATH=. python exploits/chain_b_indirect_injection.py # multi-step: hidden data tricks the assistant
PYTHONPATH=. python benchmarks/check_ground_truth.py # confirms the answer key still matches the codeThat's 9 ordinary tests, a proof-of-exploit for every planted bug (V07 rides
along on V06's chain), and a check for all but one of the 50 precision decoys:
code written to look every bit as suspicious as a real bug
while being perfectly safe. Flag a decoy and you've scored yourself a false
positive. Installing the mcp extra (pip install -e ".[mcp,mcp-http]") adds 5
otherwise-skipped MCP tests.
Exact numbers and scoring rules: SCOREBOARD.md.