13 hrs ago
AI Model Upgrade Suddenly Unlocks Exploit-Capable Cyber Skills
Researchers gave two versions of an AI model the same difficult computer-security challenge.
The older model, Claude Opus 4.8, could not reliably solve it.
A newer version, Opus 5, solved it in about three hours.
The computer being attacked did not change.
The newer model was simply better at the task.
This matters because AI abilities can improve suddenly between ordinary product releases.
Safety tests may miss abilities that were not expected or specifically tested.
People should not assume that a model’s current limitations will remain after an upgrade.
Hacktron researchers tested two versions of the same AI model against the same target.
Claude Opus 4.8 struggled across several sessions and failed to produce a working exploit.
Opus 5 solved the identical problem in about three hours.
The attack involved defeating memory-protection randomization that obscures a program’s data locations.
The comparison highlights how routine model releases can suddenly change cybersecurity risks and expose gaps in safety testing.
- Who
- Hacktron researchers and the AI models Claude Opus 4.8 and Opus 5.
- What
- Researchers compared the models’ ability to produce an exploit against a target protected by memory-location randomization.
- Where
- The target’s specific location is not identified in the article.
- When
- The tests were conducted days apart, with the models separated by a normal product release.
- Why
- The comparison was used to show that a model upgrade can create new cybersecurity risks and challenge existing safety evaluations.
Testing Must Improve
General Models Cannot Be Fully Predicted
What the pre-release testing should catch
Testing Must Improve
If testing did not reveal that Opus 5 could reliably produce exploits against hardened targets, the incident indicates a gap in the evaluation process.
General Models Cannot Be Fully Predicted
No evaluation can test every possible misuse of a broadly capable model, because the number of potential tasks is too large.
How to plan for model upgrades
Testing Must Improve
Defenders and buyers need stronger ways to identify and prepare for capability changes before new versions are released.
General Models Cannot Be Fully Predicted
A model’s ability to surprise its makers is an inherent consequence of building systems general enough to handle previously unseen problems.
Key facts
- Earlier model
- Claude Opus 4.8 struggled across several sessions and did not produce a working exploit.
- Newer model
- Opus 5 solved the same problem in about three hours.
- Target
- The same target was used in both attempts.
- Security barrier
- The attack involved bypassing a technique that randomizes where program data is located.
- Time between tests
- The models were tested days apart.
- Release context
- The change followed a normal product release, not an announced cyber-weapon release.
- Main concern
- Current model limitations may no longer provide dependable safety after a more capable version ships.









