An AI agent gives high confidence on 56% of high-risk tasks
The researchers find that this reporting destroys 68% of the gains from delegation.
Claimed, not confirmed
Researchers did a test with an AI agent that tells a user its confidence. The user can delegate a task to the agent or do the task. The researchers told a large language model its correct success rate. The model gave high confidence on 56% of the tasks with a high risk of failure for the model. The researchers find that incentives cause this.
Sources
Posted