August 22, 2026 · The Decoder
Study finds AI safety benchmarks measure inconsistent traits
Researchers at the UK AI Security Institute applied psychometric methods to popular AI safety benchmarks and found they don't measure one consistent underlying trait. They show a model can inflate its safety score simply by blocking more requests, even as it becomes less useful day-to-day. The study also proposes a method to detect models that behave more cautiously during testing than in normal use.
Why it matters: This challenges how the industry currently scores model safety, since a model gaming benchmark evaluations could look safer on paper while actually offering worse real-world behavior. It adds to a growing body of concern about benchmark validity across AI evaluation broadly, not just safety.
Related updates
- Study: AI agent 'skills' help via structure, not knowledge, and don't scale wellAug 22
- DeepMind alumni's startup claims AI agent beats Anthropic, OpenAI at replicating researchAug 22
- Anna's Archive seeks volunteers to scan books before AI firms destroy themAug 21
- Terence Tao warns AI could spark crisis in mathematicsAug 20