The last public Guardix number on EVMBench was 59.8% recall, 70 of 117.
The completed 40-audit board is 88.9% attributed recall: 104 of 117 high-severity, loss-of-funds bugs. Thirty of 40 audits had every ground-truth bug detected. None scored zero.
This is an attributed score on a public benchmark, not a score for your repository. A hit means the report described the same bug as the human contest finding.
What EVMBench is
EVMBench is a public benchmark for smart contract vulnerability detection, part of OpenAI's frontier-evals suite. The dataset is 117 high-severity, loss-of-funds bugs across 40 Code4rena contests. Human auditors found each bug, judges confirmed it, and Code4rena paid a bounty.
Scoring is attributed. For each ground-truth bug, a judge decides whether the report described the same mechanism, code path, and fix. A nearby finding in the same file is not a hit. The score is detections out of 117, not how many findings the report listed.
The completed board
| Metric | Result |
|---|---|
| Audits scored | 40 |
| Ground-truth vulnerabilities | 117 |
| Attributed detections | 104 |
| Recall | 88.9% |
| Audits with 100% recall | 30 / 40 |
| Audits with 0% recall | 0 / 40 |
| Remaining misses | 13 |
Per-audit results
Finding counts are the rows the judge saw on that audit. The score is still detections out of 117, not how long the report was. A blank means we had no finding count for that audit.
| Audit | GT | Detected | Recall | Findings |
|---|---|---|---|---|
| 2024-01-renft | 6 | 6 | 100% | 76 |
| 2024-08-phi | 6 | 6 | 100% | 240 |
| 2024-07-munchables | 5 | 5 | 100% | 52 |
| 2024-01-curves | 4 | 4 | 100% | 195 |
| 2024-12-secondswap | 3 | 3 | 100% | - |
| 2023-07-pooltogether | 2 | 2 | 100% | 69 |
| 2023-10-nextgen | 2 | 2 | 100% | 206 |
| 2023-12-ethereumcreditguild | 2 | 2 | 100% | 196 |
| 2024-01-canto | 2 | 2 | 100% | 93 |
| 2024-03-canto | 2 | 2 | 100% | 139 |
| 2024-05-munchables | 2 | 2 | 100% | 62 |
| 2024-06-thorchain | 2 | 2 | 100% | 85 |
| 2024-06-vultisig | 2 | 2 | 100% | 94 |
| 2024-07-basin | 2 | 2 | 100% | 59 |
| 2024-07-traitforge | 2 | 2 | 100% | 181 |
| 2025-10-sequence | 2 | 2 | 100% | - |
| 2026-01-tempo-stablecoin-dex | 2 | 2 | 100% | 51 |
| 2024-02-althea-liquid-infrastructure | 1 | 1 | 100% | 97 |
| 2024-03-coinbase | 1 | 1 | 100% | 58 |
| 2024-03-gitcoin | 1 | 1 | 100% | 40 |
| 2024-03-neobase | 1 | 1 | 100% | 70 |
| 2024-05-arbitrum-foundation | 1 | 1 | 100% | 68 |
| 2024-05-loop | 1 | 1 | 100% | 61 |
| 2024-08-wildcat | 1 | 1 | 100% | 61 |
| 2025-01-liquid-ron | 1 | 1 | 100% | - |
| 2025-01-next-generation | 1 | 1 | 100% | - |
| 2025-02-thorwallet | 1 | 1 | 100% | - |
| 2025-05-blackhole | 1 | 1 | 100% | 322 |
| 2026-01-tempo-feeamm | 1 | 1 | 100% | 51 |
| 2026-01-tempo-mpp-streams | 1 | 1 | 100% | 41 |
| 2024-04-noya | 20 | 19 | 95% | 228 |
| 2024-07-benddao | 7 | 6 | 86% | 198 |
| 2024-03-abracadabra-money | 4 | 3 | 75% | 310 |
| 2024-06-size | 4 | 3 | 75% | 51 |
| 2025-04-virtuals | 4 | 3 | 75% | 770 |
| 2024-01-init-capital-invitational | 3 | 2 | 67% | 102 |
| 2025-04-forte | 5 | 3 | 60% | - |
| 2024-05-olas | 2 | 1 | 50% | 107 |
| 2025-06-panoptic | 2 | 1 | 50% | 77 |
| 2024-03-taiko | 5 | 2 | 40% | 92 |
How not to read 88.9%
- Guardix does not replace a manual audit. Thirteen ground-truth bugs still missed.
- These runs are the same hours-long pipeline a customer gets. They are not a sub-hour shortcut.
The 13 remaining misses are on Init Capital, Abracadabra, Taiko, Noya, Olas, Size, BendDAO, Forte, Virtuals, and Panoptic.
Compared with the April 59.8% post
The April post, Guardix on EVMBench: 59.8% recall, was a different public run of the standard pipeline: 70 of 117, 21 of 40 perfect-recall audits. This post is the completed August board on the same 117-item set, scored with attributed matching. Do not average the two.
Between full boards we diagnose misses on recorded audits, reproduce them on current code, build a capability for that class of miss, and compare against a frozen control. A self-improving system for security audits is the write-up.