A test with 222 tasks shows that agents do not keep to a set time
The authors found a typical error of 2.9 times for Fable 5.1 in Claude Code, but only 1.2 times for GPT-6 Astra in Codex.
Claimed, not confirmed
This is a brief. We point to the report and do not rewrite it. Read it at the source below.
Sources
Posted