What matters in AI.

Subscribe

Reruns of the same tasks change 15 to 36 percent of the results

Tool names with incorrect information lower the reward of each agent by 0.20 to 0.44, the TRACE test shows.

Claimed, not confirmed

This is a brief. We point to the report and do not rewrite it. Read it at the source below.

Harmless wording changes versus misleading tool namesOn 88 new tasks with 4 agents, renaming tools or reformatting outputs kept reward within ±0.10 for 7 of 8 agent-change pairs, while misleading tool names lowered all 4 agents rewards by 0.20 to 0.44. Identical reruns flip 15-36% of outcomes. 88 NEW TASKS, 4 AGENTS Rename or reformat Misleading names AGENTS AFFECTED 7 of 8 4 of 4 REWARD UNCHANGED REWARD LOWER REWARD CHANGE 0 within ±0.10 −0.20 to −0.44 Identical reruns flip 15-36% of outcomes
In a test on 88 new tasks, cosmetic changes left rewards nearly unchanged for seven of eight agent-change pairs, while misleading tool names hurt all four agents.

Sources

  1. TRACE: Diagnosing Verifier Brittleness in Agentic Evaluationarxiv.org
AI MATTER · NEWS · AI MATTER · NEWS ·9 OCT2026

Posted

Tags

More in Benchmarks

All Benchmarks news