3 weeks ago

Meta's Muse Code coding agent beats Grok, Gemini on benchmark

Meta's Muse Code coding agent beats Grok, Gemini on benchmark
Meta just shipped a coding agent that beats Grok and Gemini · wionews.com

Meta built a new computer helper called Muse Code.

It lives in the terminal, a place where programmers type commands on a black screen.

Instead of just giving ideas, Muse Code does coding jobs by itself, like fixing mistakes and changing files.

It is powered by a smart brain called Muse Spark 1.2.

On a big test called DeepSWE 1.1, it finished 59 out of every 100 tasks.

That was better than similar helpers from xAI and Google.

Meta used to give its smart brains away for free, but now it is starting to charge for them.

Other companies are racing to build the best coding helpers too.

Some grown-ups are not sure the test scores are totally honest, and an outside tester measured one of Meta's models lower than Meta said.

So the new helper looks strong, but we should wait and see.

Key facts

Product
Muse Code terminal coding agent
Underlying model
Muse Spark 1.2 (released August 5)
DeepSWE 1.1 score
59 percent, ahead of Grok Build 4.5 and Gemini 3.6 Flash
Meta Model API pricing
$1.25 per million input tokens; $4.25 per million output tokens
Terminal-Bench 2.1 discrepancy
Meta reported 80.0 for Muse Spark 1.1; Vals AI measured 69.29
Rival product moves
Claude Opus 5 (July 24); Kimi K3 open weights (July 27); OpenAI price cuts up to 80 percent
Prior benchmark controversy
Yann LeCun confirmed Llama 4's headline scores were a composite of best per-benchmark results

Sources

Related news