← Back to live feed · 1 stories across 1 day
Sunday, Sep 20, 2026
1 story1 NEWGAVEL Harness Lifts Qwen3-8B Robot Success to 92.6% Without Model Changes AI Sep 19, 10:10 PM EDT 3/3
Single task success for the Qwen3-8B model rose from 41.2% to 91.8% on robot tasks using a symbolic harness called GAVEL. The system integrates an explicit graph world model that predicts the outcomes of LLM generated actions to catch and repair violations before they are executed. On the BEHAVIOR 1K benchmark, success across 500 multi task instructions climbed from 19.9% to 92.6% without changes to the underlying model.
The harness maintains a graph of object relations, action preconditions, and probabilistic beliefs about unobserved objects to address state tracking failures. It reorders remaining subtasks to cut travel distance by 5.4% based on the distribution of possible object locations. These results derive from simulation harness tests and do not represent performance in deployed household robots.