← Back to live feed · 1 stories across 1 day
Thursday, Sep 17, 2026
1 story1 Epoch AI Labels 9 of 15 Benchmarks Flawed in First AI Audit AI Sep 17, 5:53 PM EDT 11/8
The Epoch AI research group launched a project to evaluate the quality of AI capability assessments, finding the majority of its initial sample failed to meet standards. The "Benchmark Reviews" initiative audited 15 tools, labeling 9 as flawed, 4 as verified, and 2 as lacking sufficient information for a review. In one instance, the DeepSWE benchmark contained 23 false negatives across 131 tasks, creating an artificial performance ceiling of 79.6% that aligns with the current high score of approximately 74%.
Auditors use a specific rubric to assign verdicts and publish full assessments of limitations for every verified benchmark. The organization excludes its own created tools from the audit to avoid conflicts of interest. Analysis of DeepSWE v1.1 showed that the benchmark's verifier discards agent changes to certain test files without notifying the model, causing failures unrelated to model performance.