1 month ago
Meta Faked AI Benchmark Score, Raising Concerns Over Muse Spark 1.1
Meta, a big tech company, submitted an AI model called Llama 4 to a leaderboard where AI models are ranked.
The version they submitted scored very high, but when the actual public version was released, it scored much lower.
People noticed big differences between the two versions and accused Meta of cheating to make their model look better.
Meta's former chief AI scientist, Yann LeCun, later confirmed that Meta had indeed manipulated the benchmark scores by picking the best results from different versions of the model.
This has raised concerns about the honesty of Meta's AI benchmarks and the real capabilities of their AI models, especially with the recent release of Muse Spark 1.1.
Additionally, there have been real-world incidents where Meta's AI made wrong decisions, like disabling real business accounts and allowing hackers to take over Instagram accounts.
Meta's Llama 4 model initially ranked 2nd on LMArena but dropped to 32nd with the public version.
Developers found substantive differences between the submitted and public versions.
Meta's VP of generative AI denied manipulation, attributing it to cloud implementations.
Yann LeCun confirmed Meta used checkpoint cherry-picking to inflate benchmark scores.
Real-world incidents show Meta's AI making incorrect decisions, such as disabling legitimate business accounts and allowing hackers to compromise Instagram accounts.
- Who
- Meta, Yann LeCun, Ahmad Al-Dahle
- What
- Meta faked AI benchmark scores for Llama 4, raising concerns over Muse Spark 1.1.
- Where
- Global, with specific incidents in New York and Instagram
- When
- April 2025 (initial submission), January 2026 (LeCun's confirmation)
- Why
- To inflate benchmark scores and potentially misrepresent AI capabilities.
Meta's Defense
Critics' Accusations
Benchmark Score Discrepancy
Meta's Defense
Meta attributes the inconsistency to differing cloud implementations.
Critics' Accusations
Developers found substantive differences between submitted and public versions, suggesting manipulation.
Yann LeCun's Confirmation
Meta's Defense
Meta denies training Llama 4 directly on benchmark test sets.
Critics' Accusations
LeCun confirms Meta used checkpoint cherry-picking to inflate benchmark scores.
Key facts
- Date of Initial Submission
- April 2025
- Initial Ranking
- 2nd place
- Public Version Ranking
- 32nd place
- Meta's Denial
- Ahmad Al-Dahle, VP of generative AI, denied manipulation.
- Yann LeCun's Confirmation
- January 2026
- Number of Compromised Instagram Accounts
- 20,225
Quotes
Ahmad Al‑Dahle
Meta VP of generative AI
“The Llama 4 benchmark results were, in my words, fudged. Meta’s team trained multiple checkpoints, selected the highest score on each benchmark, and presented the results as a single composite performance — one that no single version of the model had actually achieved.”
wionews.com
“The claims are simply not true. The inconsistency is due to differing cloud implementations.”
wionews.com







