Published fixtures: results with their limits.

Each row is a fixture with a KNOWN planted issue (or a known-clean control). OpenThunder's keyless detectors run against it, and we publish what was expected vs what was detected. Misses and false positives are shown alongside matches. This small synthetic suite is a regression check, not an estimate of production accuracy.

6/7planted issues detected (86%)
0/2false positives on clean code (0%)
8/9cases matched expectation

Keyless, no AI. Generated 2026-09-21. ← Evidence Gallery

Case (planted ground truth)KindExpectedDetectedResult
Breaking API removal
Removing a public export breaks every consumer; a line diff hides it, a contract diff catches it.
change flag breaking API removal CAUTION: 1 breaking removal(s) (chargeCard) matched expectation
OS command injection
User input concatenated into a shell command lets an attacker run arbitrary commands.
scan detect high command high injection (SEC-command-injection) matched expectation
eval() of untrusted input
Evaluating request data as code is arbitrary code execution.
scan detect medium eval high injection (SEC-eval) matched expectation
Hardcoded cloud credential
A committed AWS key is an immediate credential-exposure risk.
scan detect high secret critical secrets (SEC-secret) matched expectation
Safe code (false-positive challenge)
Parameterized query + env config + no secrets must NOT produce a high/critical finding.
scan no high/critical finding no findings (clean) matched expectation
Safe change with coverage (false-positive challenge)
A small refactor whose test is updated alongside it must not be hard-blocked (HOLD).
change verdict not HOLD CAUTION (0 uncovered) matched expectation
SQL injection (string-built query)
Request input concatenated into a SQL string is the textbook injection.
scan detect high sql high injection (SEC-sql-injection) matched expectation
Unsafe AI agent (LLM output to shell)
Raw, unvalidated LLM output is passed to a shell exec sink. The assertion requires detecting that command-execution risk, not any incidental low finding.
scan detect high exec|shell|command|injection|insecure output|untrusted|LLM0?[12] 1 finding(s), none matching (AISEC-unbounded) missed
Untested auth-logic change
Changed auth logic with no test is unproven; the verdict should not be a clean SHIP.
change verdict not SHIP HOLD (0 uncovered) matched expectation

Reproduce: the fixtures live in apps/docs/benchmark/cases/; run node apps/docs/scripts/gen-benchmark.mjs. Scan cases run openthunder scan --all; change cases git-init the fixture and run openthunder can-i-ship. No AI, no network.

Scope: each case targets patterns and languages OpenThunder's keyless detectors cover, and the “detected” column shows exactly what fired (including severity). The fixtures are public; extend them to probe further or to hold the engine to a regression bar. The clean controls check specific thresholds (no high/critical findings or no HOLD verdict), not the absence of every warning. These results do not establish a real-world false-positive rate.