What matters in AI.

Subscribe

HarmBench does not measure one safety attribute, researchers find

3 of the 4 datasets that can target harmful refusal are saturated and cannot show differences between models.

Claimed, not confirmed

Researchers did an audit of HELM Safety, a safety benchmark. They wanted to know if harmful refusal is one attribute that a test can measure. Harmful refusal is the tendency of a model to refuse dangerous prompts. Of the 4 datasets that can target it, 3 are saturated. Tests on the last dataset, HarmBench, give a strong indication that it does not measure one attribute. The researchers say that a safety score must first show that it measures one attribute.

Sources

  1. Searching for "Harmful Refusal": A Psychometric Audit of an AI Safety Benchmarkarxiv.org
AI MATTER · NEWS · AI MATTER · NEWS ·9 OCT2026

Posted

Tags

More in Benchmarks

All Benchmarks news