Accuracy
Historical, scoped observations for engine v0.11.0, not a measurement of the current release. A current comparison requires a pinned rerun and adjudicated results.
Current execution scope
- Runtime
- Prebuilt Go binaries do not require a Go compiler for basic scanning. Building from source requires Go 1.25+. Optional analyzers have their own runtime requirements; inspect coverage to confirm they actually ran.
- DAST
- URL and API probing inspect observable runtime behavior. Intrusive probes require --enable-active. An HTTP 200 response alone does not prove access to sensitive data.
- SAST
- Native static rules cover supported Go, JavaScript, Java and infrastructure patterns. Semgrep and the opt-in Python engine extend applicable checks. Language support does not imply complete framework or path coverage.
- SCA
- Go analysis uses govulncheck. Python and npm dependency scanners use supported manifests and lockfiles, including poetry.lock, Pipfile.lock and package-lock.json. Advisory queries can transmit package names and versions.
- Evidence and policy
- Correlation connects supported matching observations; not every finding has multiple sources. Strong single-source evidence can be sufficient. Severity, confidence, finding disposition and release recommendation are distinct.
- Coverage and human review
- An analyzer failure or missing required evidence can leave the release incomplete. A PASS is a result under the declared policy and coverage, not a guarantee of security or a human approval.
Benchmark tables below are historical, scoped observations. Counts and synthetic regression scores are not production precision/recall. No replacement measurements are claimed here.
Read deployment-specific data handlingSynthetic F1
1.000
38 TPs / 0 FPs / 0 FNs across 7 detection categories
Juice Shop vs v0.6.1
Unverified
Generic application responses do not establish file exposure.
PyGoat real-world
147 findings
12 vulnerability classes in 17 s wall-clock
Track 1 — Synthetic labeled corpus
56 labeled test cases across 7 detection categories. Each category has both EXPECT_TP variants (engine SHOULD flag) and EXPECT_TN variants (safe-shape; engine should leave alone). The corpus exists in scripts/accuracy/corpus/; ground truth in scripts/accuracy/manifest.json. The harness pairs emissions to labeled cases by category + title-substring + file + ±6-line tolerance, with nearest-unclaimed-TP matching so a single TP can't be claimed by multiple emissions.
| Category | TP | FP | FN | Precision | Recall | F1 |
|---|---|---|---|---|---|---|
| sqli | 5 | 0 | 0 | 1.000 | 1.000 | 1.000 |
| cmdi | 5 | 0 | 0 | 1.000 | 1.000 | 1.000 |
| path_traversal | 5 | 0 | 0 | 1.000 | 1.000 | 1.000 |
| ssrf | 3 | 0 | 0 | 1.000 | 1.000 | 1.000 |
| open_redirect | 3 | 0 | 0 | 1.000 | 1.000 | 1.000 |
| xss | 4 | 0 | 0 | 1.000 | 1.000 | 1.000 |
| secrets | 13 | 0 | 0 | 1.000 | 1.000 | 1.000 |
| OVERALL | 38 | 0 | 0 | 1.000 | 1.000 | 1.000 |
Honest caveat
1.000/1.000/1.000 means fendix never misses these 56 specific canonical patterns, not that it never misses anything. The synthetic corpus measures the positive side; real-world FP discipline against juice-shop and production targets is tracked separately in tasks/FP_CORPUS.md in the engine repo. Tracks 2 and 3 are historical observations, not adjudicated production precision or recall.
Reproduce
make build
python3 scripts/accuracy/run.py --python-engine
# Output:
# Running ./bin/fendix scan --code scripts/accuracy/corpus...
# 20 unique findings (40 after exploding affected_endpoints)
#
# CATEGORY TP FP FN TN PREC REC F1
# ----------------------------------------------------------------------
# sqli 5 0 0 3 1.000 1.000 1.000
# cmdi 5 0 0 3 1.000 1.000 1.000
# path_traversal 5 0 0 3 1.000 1.000 1.000
# ssrf 3 0 0 2 1.000 1.000 1.000
# open_redirect 3 0 0 2 1.000 1.000 1.000
# xss 4 0 0 2 1.000 1.000 1.000
# secrets 13 0 0 3 1.000 1.000 1.000
# ----------------------------------------------------------------------
# OVERALL 38 0 0 18 1.000 1.000 1.000Engine improvements surfaced during this evaluation
The synthetic corpus surfaced and we shipped, in the same session, two real engine improvements + one latent orchestrator fix:
- Open redirect: the detector required a direct
redirect(request.args.get("x")); multi-hop assignments were silently missed. The other six reachable sinks (SQLi / SSRF / XSS / cmd-injection / path-traversal) all had the constant-vs-non-constant filter from historical change; open-redirect was the original historical change sink and somehow never got the chain treatment. Recall: 0/3 → 3/3. - cmd-injection: posture aligned with the other reachable sinks via new
_cmdi_arg_is_dangeroushelper. Pre-fix,os.system("echo hello")fired HIGH despite zero exploitability (historical change chose "fire on every shell-out"). Precision: 0.833 → 1.000. - Orchestrator:
runWhiteboxScannow resolvescode_pathandspecto absolute paths before sending the ScanRequest. The spawner setscmd.Dir = engineDir, so a relative path silently resolved to nothing in the child cwd. Surfaces as fendix reporting 0 findings on real codebases — a real user-blocking regression latent since historical change.
Track 2 — OWASP Juice Shop (real-world DAST)
Stock fendix scan --url against bkimminich/juice-shop:v17.1.1. No auth, no --code, no --enable-active — just the default blackbox pipeline against the modern OWASP web-app benchmark.
| Metric | v0.6.1 baseline | v0.11.0 (historical) | Δ |
|---|---|---|---|
| Total findings (deduped) | 7 | 12 | +5 |
| CRITICAL | 0 | 5 | +5 |
| Scan duration | 42 s | 27 s | −35 % |
| Endpoints scanned | 97 | 97 | — |
The historical run classified five results as critical. Their response contents were not adjudicated as sensitive files; they must not be counted as confirmed exposures.
Caveat
A SPA may return the same application shell for nonexistent paths. HTTP 200 alone proves neither file exposure nor cache poisoning. Compare response contents with a random-path control and verify sensitive content before claiming an exposure.
Reproduce
JS_PORT=3001 FENDIX_BIN=./bin/fendix bash scripts/benchmark/run-juice-shop.sh
# Output (bench-results/juice-shop/<timestamp>/):
# Fendix benchmark — OWASP Juice Shop
# Fendix version: fendix v0.11.0 (darwin/arm64)
# Target image: bkimminich/juice-shop:v17.1.1
# Scan duration: 27 seconds
# Endpoints: 97
# Total findings: 12
#
# By severity: CRITICAL: 5 HIGH: 0 MEDIUM: 4 LOW: 2 INFO: 1Track 3 — PyGoat (real-world SAST)
Clone of adeyosemanputra/pygoat — a Django app intentionally vulnerable to every OWASP Top 10 category. 52 Python files plus JavaScript assets. Scan via fendix scan --code /tmp/pygoat --python-engine with no auth or active probing.
Total findings
147
1 CRITICAL, 146 HIGH
Scan duration
17.1 s
52 Python files + JS assets
Categories detected
12
Categories observed in the historical run
| Severity | Vulnerability class | First detection |
|---|---|---|
| CRITICAL | Unsafe pickle deserialization (RCE) | dockerized_labs/insec_des_lab/main.py:36 |
| HIGH | Unsafe eval() with dynamic arg | introduction/mitre.py:218 |
| HIGH | subprocess(shell=True) | introduction/mitre.py:233 |
| HIGH | Unsafe yaml.load() (RCE) | introduction/lab_code/test.py:23 |
| HIGH | SSRF — dynamic URL | 2 sites (incl. views.py:963) |
| HIGH | innerHTML XSS | introduction/static/js/a9.js:40 |
| HIGH | Open redirect — 9 sites | broken_auth_lab/app.py:107 |
| HIGH | Hardcoded API key / password / JWT | 3 distinct files |
| HIGH × 133 | Vulnerable dependency (certifi, cryptography, django, …) | requirements.txt |
Caveats
- There is no finding-level ground-truth manifest for this run; precision and recall cannot be calculated from it.
- Finding count does not measure accuracy. This experiment does not establish expected finding counts or false-positive rates in production applications.
Reproduce
git clone --depth 1 https://github.com/adeyosemanputra/pygoat /tmp/pygoat
./bin/fendix scan --code /tmp/pygoat --python-engine --max-duration 60s
# Output:
# scan complete duration=17.082s total=147 critical=1 high=146 medium=0
#
# By category:
# deps 135 (real CVE-tagged dependencies in requirements.txt)
# injection 9 (SSRF/XSS/eval/subprocess-shell/pickle/yaml/open-redirect)
# secrets 3 (hardcoded API key, password, JWT)Going deeper
- docs/accuracy.md — upstream source of every number on this page (kept in sync).
- /performance — cold-start latency + binary-size benchmark.
- scripts/accuracy/run.py — the harness. ~250 LOC, reads cleanly top-to-bottom.
- tasks/FP_CORPUS.md — the real-world FP discipline corpus (juice-shop, fendix-self, TwiScope). The flip side of this page: what fendix doesn't catch / what it flags by mistake.