8 Oct 2026 Research
The authors found that humans, DeepSeek and Qwen3-Max all cut standard delivery by 48 to 53 percentage points. The results are different for each value of time.
Bioengineer.org Claimed, not confirmed
8 Oct 2026 Research
The authors found 12.3 times more collisions in real-time tests than in static tests. The agents completed 91% to 94% of the tasks.
arxiv.org Claimed, not confirmed
8 Oct 2026 Research
The authors found a typical error of 2.9 times for Fable 5.1 in Claude Code, but only 1.2 times for GPT-6 Astra in Codex.
arxiv.org Claimed, not confirmed
8 Oct 2026 Research
2 easy methods to add diversity did not decrease the risk.
arxiv.org Claimed, not confirmed
8 Oct 2026 Research
The researchers find that this reporting destroys 68% of the gains from delegation.
arxiv.org Claimed, not confirmed
8 Oct 2026 Research
An instruction to randomize independently decreased the correlation between agents with the same input, but did not remove it.
arxiv.org Claimed, not confirmed