Science & Tech · AI · 19 hrs ago

Test finds some AI models gave detailed answers after safeguards were removed

Test finds some AI models gave detailed answers after safeguards were removed

A study by the UK-based organisation Tech Against Terrorism tested more than 130 AI models.

Researchers presented them with hundreds of requests of the kind a person planning a terrorist attack might make.

They then disabled safeguards that normally make models refuse certain requests.

Some models gave detailed answers about attacks, terrorist financing and radicalisation.

Meta’s Llama 3.1 8B scored 97 out of 100 before the change and about 3 afterwards.

The results suggest that removing safeguards can leave some models open to dangerous use.

Tech Against Terrorism founder Adam Hadley said many open-source models’ safeguards had already been bypassed, though the statement gave no next steps.

Sources

Related news