← Back to live feed · 1 stories across 1 day
Friday, Sep 18, 2026
1 story1 NEWGPT 6 Astra Tops Initial Token Efficiency on 11 Task WeirdML v3 Benchmark AI Sep 18, 6:20 AM EDT 5/3
A new AI performance test requires models to build data analysis pipelines and handle unfamiliar datasets to solve complex problems. Released by researcher @htihle, the WeirdML v3 benchmark features 11 hand-made tasks where models operate with limited feedback and unspecified goals. GPT 6 Astra leads early results in token efficiency, maintaining a large advantage over other tested models in its ability to produce results with fewer tokens.
The benchmark includes a radio telescope simulation called the TOD Pipeline task, which forces agents to extract weak sky signals from atmospheric emission and system temperature spikes. Models must calibrate the instrument and filter radiometer noise to map data onto the sky. These results continue to update as the benchmark adds more models to its evaluation set.