HarmBench does not measure one safety attribute, researchers find
3 of the 4 datasets that can target harmful refusal are saturated and cannot show differences between models.
Claimed, not confirmed
Researchers did an audit of HELM Safety, a safety benchmark. They wanted to know if harmful refusal is one attribute that a test can measure. Harmful refusal is the tendency of a model to refuse dangerous prompts. Of the 4 datasets that can target it, 3 are saturated. Tests on the last dataset, HarmBench, give a strong indication that it does not measure one attribute. The researchers say that a safety score must first show that it measures one attribute.
Sources
Posted