What matters in AI.

Subscribe

Google adds long tasks to Android Bench 2.0

On the new long tasks, Claude Opus 5.5 is first with a 32% pass rate, then GPT 6 Astra with 28%.

This is a brief. We point to the report and do not rewrite it. Read it at the source below.

Top model passes 32% of long-horizon tasks On Android Bench 2.0, Claude Opus 5.5 leads with a 32% long-horizon task pass rate, shown as 32 of 100 squares filled. GPT 6 Astra follows at 28%, shown as a bar 28% of the grid width. LHT PASS RATE Android Bench 2.0 32% Claude Opus 5.5 Top of leaderboard 100% GPT 6 Astra 28%
On Google's new Android Bench 2.0, the leading model, Claude Opus 5.5, passes 32% of long-horizon tasks; GPT 6 Astra follows at 28%.

Sources

  1. Android Bench 2 Adds Support for Long-Horizon Tasks, Agentic Evaluation, and Continuous Scoringinfoq.com
AI MATTER · NEWS · AI MATTER · NEWS ·9 OCT2026

Posted

Tags

More in Agents

All Agents news