← Back to live feed · 1 stories across 1 day

Friday, Sep 18, 2026

1 story
1
Epoch AI Labels 9 of 15 AI Benchmarks Flawed in New Audit
topics 🤖 AI tags AIAI ModelsAI Research keywords Epoch AI Research

Epoch AI Research released its first set of quality reviews for artificial intelligence benchmarks on Sept. 17. The inaugural phase of the Benchmark Reviews project analyzed 15 evaluation sets, labeling nine as flawed, four as verified, and two as insufficient for review. The organization uses a specific rubric to assign these verdicts, focusing on whether errors in the benchmarks fundamentally alter the results of AI capability tests.

Findings from the audit revealed that 45.5% of tasks in Terminal Bench 4.0 were broken, and 46% of sampled questions in the HLE benchmark were faulty. For DeepSWE v1.1, the audit identified a bug that disrupts grading for all tasks by discarding agent changes to test files without informing the model. To prevent conflicts of interest, the group excludes benchmarks it created from the review process.

Image via @epochairesearch on X
Earlier version from Thursday, Sep 17
Epoch AI Labels 9 of 15 AI Benchmarks Flawed in Debut Audit
16 tweets • 12 sources
See all 17 tweets →