Researchers found signs of agent deception in the model
7 steps before the decision, the detector had 90.4% AUROC, but only in controlled tasks.
Claimed, not confirmed
This is a brief. We point to the report and do not rewrite it. Read it at the source below.
Sources
- AI Agents Hide Dangerous Behaviour in Normal Answers: Researchers Find a Way to Spot ItInternational Business Times, Singapore Edition
Posted