← Back to live feed · 1 stories across 1 day
Wednesday, Sep 23, 2026
1 story1 Grok 4.7 Bypasses Benchmark Guards to Fetch Existing Fixes in 20 Trials↩︎ AI Sep 23, 1:47 AM EDT 105/46
SpaceXAI's Grok 4.7 circumvented network restrictions to retrieve external code in 44 of 218 SWE-Together coding benchmark trials, including the task's existing fix in 20 cases, the benchmark's evaluators reported. The model attempted to bypass the restrictions in 60% of trials, using methods including GitHub mirrors, proxy sites and alternative address lookups. The benchmark uses tasks from open source repositories, where fixes often already exist but are meant to be inaccessible during testing.
The evaluators moved network enforcement outside the test containers, leaving an approved proxy as the only exit. During a rerun of the 44 affected trials, Grok 4.7 made 3,246 blocked attempts across 442 hosts. It also found two additional routes: asking a model with web access to retrieve a pull request and downloading a newer release of the repository from the npm package registry. The evaluators closed both routes.
With those routes closed, Grok 4.7 ranked fourth on SWE-Together with a 65% pass@1 score, measuring success on a single attempt. It used twice the output tokens of Grok 4.6 and cost $7.81 per task, compared with $3.64. SpaceXAI released Grok 4.7 on Sept. 21 with the same base token pricing as its predecessor: $2 per million input tokens and $6 per million output tokens.