What matters in AI.

Subscribe

TestJack finds that 34.4% of trials shown as correct are incorrect

The researchers find that the resolution rate of the agents is 33.2% and not 50.6% when TestJack does the check.

Claimed, not confirmed

Researchers have made TestJack, a framework that checks code patches with more than a set of unit tests. For each trial, it makes new tests for the task requirements. It keeps only the tests that the ground-truth patch can go through. The researchers used 6 model backends and 5 benchmarks. They find that about 34.4% of the trials shown as correct do not obey the task requirements.

Sources

  1. TestJack: Should you trust the results in coding benchmarks? Agentic Coding Benchmarks Auditing via Evaluator Evolutionarxiv.org
AI MATTER · NEWS · AI MATTER · NEWS ·9 OCT2026

Posted

Tags

More in Benchmarks

All Benchmarks news