AI code review,
measured on real PRs.
How we test the ScanDrix engine: an 11-domain eval gate with behavioral checks, a weighted composite score, and a fail-closed gate. No competitor runs, no leaderboard — yet.
- Eval domains
- 11
- Score pillars
- 4 · 0.75 bar
- Location tolerance
- ±2 lines
- Gate profiles
- harness · local · CI
Evaluation Design
Large diffs break shallow reviewers
Once a PR grows past a few hundred lines, chunk-based tools lose cross-file context and miss taint flows that span handlers, services, and DB layers. ScanDrix compiles the full AST before reasoning.
Explore AST compilation ↗Static rules plateau early
Pattern linters struggle because they cannot reason about architecture or cross-file data flow. ScanDrix combines tree-sitter AST queries with multi-model contextual reasoning over call graphs.
See Drixy rule engine ↗Deterministic evaluation gate
11 engine eval domains enforce contracts, schemas, and behavioral checks. A fatal failure blocks — it never warns and continues, guaranteeing sub-minute reviews with ±2 lines ground-truth tolerance.
Read benchmark methodology ↗02 — Category breakdown
What the harness checks.
Six capability areas, each exercised by the evaluation harness. We publish the methodology — not leaderboard scores, and no competitor numbers we did not measure.
Eval domains
11 engine areas
Composite score
4 pillars · 0.75 bar
Leaderboard scores
Not published yet
OWASP Top 10 security
Injection, XSS, auth & access flaws
Our approach
AST parsing plus cross-file taint tracking
Evaluation
Severity + secondary evals
Logic bug detection
Broken invariants & edge cases
Our approach
Multi-model reasoning over call graphs
Evaluation
Scorer + investigation evals
Secret & credential leaks
Tokens, keys & passwords in diffs
Our approach
Deterministic secret patterns on every diff
Evaluation
Promotion gate
Cross-file taint tracking
Source-to-sink flows across packages
Our approach
Dataflow analysis across file boundaries
Evaluation
Parser + anchoring evals
Drixy rules enforcement
Custom org & repo policies
Our approach
Declarative rules with Dry Run previews
Evaluation
Behavioral scorer (±2 lines)
PR turnaround
Time from webhook to posted review
Our approach
Async worker queue, streaming comments
Evaluation
Design goal: sub-minute
Evaluation status
Internal evaluation in progress
No public leaderboard yet — and no competitor scores we did not measure. Scores stay with the harness until independent runs exist.
Method: an 11-domain engine eval gate — anchoring, dedup, severity, format, investigation, rules, parser, summary, promotion, scorer, secondary — with a weighted four-pillar composite score and a fail-closed gate. No competitor runs, no leaderboard — yet.
03 — How to read these results
Three things we actually track.
Marketing benchmarks cherry-pick one metric. A review tool lives or dies on all three — miss bugs, slow the team, or cry wolf, and it gets disabled. So these are the three axes of the harness, not three scores.
Location-true scoring
A finding counts when it names the right file within ±2 lines — recall and precision against ground truth, plus F1.
Machine-checkable output
Format and anchor location are scored pillars, so malformed or unplaceable findings fail the same gate as wrong ones.
One bar, fail-closed
Four pillars weighted into a composite with a 0.75 pass bar. Below it the gate blocks — in harness, local, and CI profiles.
04 — Methodology
Reproducible by design.
No tuned demos. One fixed gate, one composite bar — and no rankings we did not earn. The bar is exact file and line, checked automatically.
Eleven engine domains
Anchoring, dedup, severity, format, investigation, rules, parser, summary, promotion, scorer, secondary — each with behavioral checks.
Ground truth with tolerance
Rule violations scored against ground-truth file + line sites within ±2 lines: recall, precision, F1.
One composite bar
Recall, precision, format and anchor location weighted into a single score with a 0.75 pass bar.
Fail-closed gate
Three profiles — harness, local, CI. A fatal failure blocks the gate instead of warning and continuing.
Stop trusting. Start measuring.
Test ScanDrix on your own repositories.
Connect GitHub or GitLab in minutes. Read your first line-precise review before believing any number on this site — free trial, no credit card.
Methodology, not a leaderboard · live service state on /status