Guardix on EVMBench: 88.9% attributed recall

Attributed recall on the completed 40-audit board is 88.9%. That is 104 of 117 high-severity, loss-of-funds bugs. The last public Guardix number was 59.8%.

Guardix Team · Sep 11, 2026 · 7 min
Guardix EVMBench completed board: 88.9% attributed recall, 104 of 117 ground-truth detections, 30 of 40 audits with perfect recall
Attributed recall on the completed 40-audit board.

The last public Guardix number on EVMBench was 59.8% recall, 70 of 117.

The completed 40-audit board is 88.9% attributed recall: 104 of 117 high-severity, loss-of-funds bugs. Thirty of 40 audits had every ground-truth bug detected. None scored zero.

This is an attributed score on a public benchmark, not a score for your repository. A hit means the report described the same bug as the human contest finding.

What EVMBench is

EVMBench is a public benchmark for smart contract vulnerability detection, part of OpenAI's frontier-evals suite. The dataset is 117 high-severity, loss-of-funds bugs across 40 Code4rena contests. Human auditors found each bug, judges confirmed it, and Code4rena paid a bounty.

Scoring is attributed. For each ground-truth bug, a judge decides whether the report described the same mechanism, code path, and fix. A nearby finding in the same file is not a hit. The score is detections out of 117, not how many findings the report listed.

The completed board

Metric Result
Audits scored 40
Ground-truth vulnerabilities 117
Attributed detections 104
Recall 88.9%
Audits with 100% recall 30 / 40
Audits with 0% recall 0 / 40
Remaining misses 13

Per-audit results

Finding counts are the rows the judge saw on that audit. The score is still detections out of 117, not how long the report was. A blank means we had no finding count for that audit.

Audit GT Detected Recall Findings
2024-01-renft 6 6 100% 76
2024-08-phi 6 6 100% 240
2024-07-munchables 5 5 100% 52
2024-01-curves 4 4 100% 195
2024-12-secondswap 3 3 100% -
2023-07-pooltogether 2 2 100% 69
2023-10-nextgen 2 2 100% 206
2023-12-ethereumcreditguild 2 2 100% 196
2024-01-canto 2 2 100% 93
2024-03-canto 2 2 100% 139
2024-05-munchables 2 2 100% 62
2024-06-thorchain 2 2 100% 85
2024-06-vultisig 2 2 100% 94
2024-07-basin 2 2 100% 59
2024-07-traitforge 2 2 100% 181
2025-10-sequence 2 2 100% -
2026-01-tempo-stablecoin-dex 2 2 100% 51
2024-02-althea-liquid-infrastructure 1 1 100% 97
2024-03-coinbase 1 1 100% 58
2024-03-gitcoin 1 1 100% 40
2024-03-neobase 1 1 100% 70
2024-05-arbitrum-foundation 1 1 100% 68
2024-05-loop 1 1 100% 61
2024-08-wildcat 1 1 100% 61
2025-01-liquid-ron 1 1 100% -
2025-01-next-generation 1 1 100% -
2025-02-thorwallet 1 1 100% -
2025-05-blackhole 1 1 100% 322
2026-01-tempo-feeamm 1 1 100% 51
2026-01-tempo-mpp-streams 1 1 100% 41
2024-04-noya 20 19 95% 228
2024-07-benddao 7 6 86% 198
2024-03-abracadabra-money 4 3 75% 310
2024-06-size 4 3 75% 51
2025-04-virtuals 4 3 75% 770
2024-01-init-capital-invitational 3 2 67% 102
2025-04-forte 5 3 60% -
2024-05-olas 2 1 50% 107
2025-06-panoptic 2 1 50% 77
2024-03-taiko 5 2 40% 92
All 40 EVMBench audits, sorted by recall. Total: 104/117 attributed detections, 88.9%.

How not to read 88.9%

  • Guardix does not replace a manual audit. Thirteen ground-truth bugs still missed.
  • These runs are the same hours-long pipeline a customer gets. They are not a sub-hour shortcut.

The 13 remaining misses are on Init Capital, Abracadabra, Taiko, Noya, Olas, Size, BendDAO, Forte, Virtuals, and Panoptic.

Compared with the April 59.8% post

The April post, Guardix on EVMBench: 59.8% recall, was a different public run of the standard pipeline: 70 of 117, 21 of 40 perfect-recall audits. This post is the completed August board on the same 117-item set, scored with attributed matching. Do not average the two.

Between full boards we diagnose misses on recorded audits, reproduce them on current code, build a capability for that class of miss, and compare against a frozen control. A self-improving system for security audits is the write-up.