1 week ago

AI Model Releases Outpace Independent Testing, Raising Evaluation Concerns

AI Model Releases Outpace Independent Testing, Raising Evaluation Concerns
Eleven AI models in twenty days! Releases have outrun anyone's ability to test them · wionews.com

AI companies are releasing new models very quickly.

Eleven models appeared in one month, including six within just nine days.

Testing an AI model properly can take longer than the time between releases.

This means people may compare a new model's company-reported scores with an older model's company-reported scores.

Companies usually publish the tests where their models perform well, even when the numbers are accurate.

Some models may also have seen test questions during training, which can make their scores look better.

Companies are increasingly competing on speed, price, and the ability to run models themselves.

People choosing an AI model should test it on the work they actually need it to do.

Key facts

Models released
Eleven new AI models shipped during the month.
Providers involved
Seven different providers released the models discussed.
Fastest release period
Six models launched in nine days, from August 6 to August 14.
Gemini release gap
Google released Gemini 3.7 Flash three weeks after Gemini 3.6 Flash.
Benchmark change
Gemini 3.7 Flash reportedly gained 16 points on DeepSWE v1.1 in three weeks.
Independent testing timeline
Independent benchmarking and safety evaluation can take weeks or longer.
Recommended evaluation
The article advises using independent results and testing models on the buyer's actual workload.

Sources

Related news