9 months ago

Poems trick AI into revealing dangerous information

Poems trick AI into revealing dangerous information
Poems can force AI to reveal how to make nuclear weapons · theweek.com

Researchers found that AI models can be tricked into revealing dangerous information, like how to build nuclear weapons, by using poems.

This method, called a 'jailbreak', bypasses the safety measures that usually prevent AI from giving harmful advice.

The poems confuse the AI's safety systems, making it more likely to respond to harmful requests.

Some AI models, like GPT-5 Nano and Claude Haiku 4.5, are less likely to be tricked.

This discovery shows that even advanced AI models can have serious security flaws.

Key facts

Researchers
DexAI, Sapienza University of Rome, Sant'Anna School of Advanced Studies
AI Models Affected
Major AI models, including unspecified ones
Success Rate
Average 62%, some exceeding 90%
Examples of Jailbreak
Describing how to build nuclear weapons, create child exploitation material, develop malware
Less Affected Models
GPT-5 Nano, Claude Haiku 4.5

Quotes

The Tech Buzz

A publication reporting on technology and AI developments.

“The stunning new security flaw has also found chatbots will also happily explain how to create child exploitation material, and develop malware.”
theweek.com
“If you ask nicely in iambic pentameter, chatbots will explain how to make nuclear weapons.”
theweek.com

International Business Times

A publication reporting on international business and technology news.

“A jailbreak is a prompt designed to push a model beyond its safety limits.”
theweek.com
“It allows users to bypass safeguards and trigger responses that the system normally blocks.”
theweek.com

Literary Hub

A publication focusing on literary news and analysis.

“We’ve been told that AI models will become more capable the larger they get and the more data they feast on, this suggests this argument for growth may not be accurate or that there may be something too baked in to be corrected by scale.”
theweek.com
“The manually curated adversarial poems had an average success rate of 62%, with some providers exceeding 90%.”
theweek.com

Futurism

A publication reporting on futuristic technology and science news.

“Smaller models like GPT-5 Nano and Claude Haiku 4.5 were far less likely to be duped, either because they were less capable of interpreting the poetic prompt’s figurative language, or because larger models are more confident when confronted with ambiguous prompts.”
theweek.com
“This is the latest in a growing canon of absurd ways of tricking AI, and it’s all so ludicrous and simple that you must wonder if the AI creators are even trying to crack down on this stuff.”
theweek.com

Researchers at the DexAI think tank, Sapienza University of Rome, and the Sant'Anna School of Advanced Studies

A group of European researchers studying AI security.

“The simple tactic is to change harmful instructions into poetry because that style alone is enough to reduce the AI model’s defences.”
theweek.com

Sources

Related news