A self-evolving harness lifts the Qwen3.5-4B score from 0.16 to 0.30
The researchers say that harness changes repair procedure failures, but only weight changes repair content failures.
Claimed, not confirmed
The researchers examined two procedures to increase the performance of an agent for long tasks. They can change the harness around a frozen model, or they can train the weights of the model. They say that the type of failure shows which procedure to use. On the DeepPlanning benchmark, a self-evolving harness lifted the held-out score of Qwen3.5-4B from 0.16 to 0.30. Changes to the weights then cut the content failures that the harness did not repair.
Two types of failure
The team identifies each trajectory with a failure by the first signal that fires. This gives two groups.
- Procedure failures: blocked calls, loops, and used step budgets.
- Content failures: the agent gives a plan that is not good.
Harness changes repair procedure failures. Weight changes repair content failures.
The harness loop
On DeepPlanning, the loop lifted the score of Qwen3.5-4B from 0.16 to 0.30. The loop repaired procedure failures. It left content failures for the weights.
Weights and adapters
The team uses trajectories of the self-evolving harness to train LoRA adapters. An adapter adds weights to a frozen model. On the 9B model, the adapter gave the same result as the full harness loop. It cut content failures from a quarter of trajectories to one in twenty.
An adapter that the team trains on answer-shuffled trajectories gave a lower score than the model with no adapter.
Other benchmark
On WebArena-Lite, the loop gave +0.09 on tasks that it did not see before. The model gets this from the information that it sees. Adapters did not add to it.
This is a brief. We point to the report and do not rewrite it. Read it at the source below.
Sources
Posted