16 hrs ago

Claude Opus 5 Safeguard Bypassed Through False Target Context

Claude Opus 5 Safeguard Bypassed Through False Target Context
Claude refused to write the exploit! A lie about the target was enough to change its mind · wionews.com
  • A three-person team reportedly used Claude Opus 5 during the Hacktron AI breach to reach OpenAI’s internal code.

  • The model initially refused to write exploit code targeting a real remote system.

  • Researchers routed the model toward their own server, presenting it as a vulnerable capture-the-flag practice target.

  • After accepting that framing, Claude Opus 5 produced a working exploit in about three hours.

  • The incident highlights the weakness of safeguards that rely on the model believing a user’s description of context.

Key facts

Model
Claude Opus 5
Team size
Three people
Initial response
The model refused to write exploit code for a real remote target.
Deceptive framing
The team presented its server as a capture-the-flag practice target.
Time to exploit
About three hours
Reported target
OpenAI’s internal code
Company response
Anthropic did not comment on the bypass in the coverage described.

Sources

Related news