3 weeks ago
Meta's Muse Code coding agent beats Grok, Gemini on benchmark
Meta built a new computer helper called Muse Code.
It lives in the terminal, a place where programmers type commands on a black screen.
Instead of just giving ideas, Muse Code does coding jobs by itself, like fixing mistakes and changing files.
It is powered by a smart brain called Muse Spark 1.2.
On a big test called DeepSWE 1.1, it finished 59 out of every 100 tasks.
That was better than similar helpers from xAI and Google.
Meta used to give its smart brains away for free, but now it is starting to charge for them.
Other companies are racing to build the best coding helpers too.
Some grown-ups are not sure the test scores are totally honest, and an outside tester measured one of Meta's models lower than Meta said.
So the new helper looks strong, but we should wait and see.
Meta released Muse Code, a terminal coding agent powered by its Muse Spark 1.2 model, which was released August 5.
On the DeepSWE 1.1 benchmark, Muse Code scored 59 percent, ahead of xAI's Grok Build 4.5 and Google's Gemini 3.6 Flash.
The release extends Meta's shift to paid AI access that began July 9 with Muse Spark 1.1 on the Meta Model API at $1.25 per million input tokens.
Competition is fierce: Anthropic released Claude Opus 5 on July 24, Alibaba's Qwen3.8-Max ran a project autonomously for more than 16 days, Moonshot AI open-sourced Kimi K3 on July 27, and OpenAI cut prices on lower-cost models by up to 80 percent.
Benchmark credibility is disputed: Vals AI measured Muse Spark 1.1 at 69.29 versus Meta's reported 80.0 on Terminal-Bench 2.1, and Yann LeCun confirmed Llama 4's headline results were a composite no single model achieved.
- Who
- Meta, which developed and released Muse Code, competing against xAI, Google, Anthropic, OpenAI, Alibaba, and Moonshot AI.
- What
- The launch of Muse Code, a terminal coding agent that scored 59 percent on the DeepSWE 1.1 benchmark, beating Grok Build 4.5 and Gemini 3.6 Flash.
- Where
- Not specified in the article.
- When
- Reported shortly after August 5, when Meta released the Muse Spark 1.2 model that powers it.
- Why
- To compete in the commercially valuable AI coding tool market and to extend Meta's shift from free model access to selling finished developer products.
Meta and benchmark backers
Independent evaluators and skeptics
Credibility of benchmark results
Meta and benchmark backers
Muse Code's 59 percent score on DeepSWE 1.1 demonstrates it genuinely outperforms Grok Build 4.5 and Gemini 3.6 Flash, showing Meta's agent is among the leading coding tools.
Independent evaluators and skeptics
Benchmark results have had a difficult year: Yann LeCun confirmed Llama 4's headline scores were a composite no single model achieved, and Vals AI measured Muse Spark 1.1 at 69.29 versus Meta's reported 80.0 on Terminal-Bench 2.1, with Meta's harness exceeding the benchmark's per-task resource limits.
Key facts
- Product
- Muse Code terminal coding agent
- Underlying model
- Muse Spark 1.2 (released August 5)
- DeepSWE 1.1 score
- 59 percent, ahead of Grok Build 4.5 and Gemini 3.6 Flash
- Meta Model API pricing
- $1.25 per million input tokens; $4.25 per million output tokens
- Terminal-Bench 2.1 discrepancy
- Meta reported 80.0 for Muse Spark 1.1; Vals AI measured 69.29
- Rival product moves
- Claude Opus 5 (July 24); Kimi K3 open weights (July 27); OpenAI price cuts up to 80 percent
- Prior benchmark controversy
- Yann LeCun confirmed Llama 4's headline scores were a composite of best per-benchmark results











