Tests show that the reward has a small effect on LLM agents
In tests on 6 models, a flipped, random or removed reward gives almost the same improvement curve.
Claimed, not confirmed
This is a brief. We point to the report and do not rewrite it. Read it at the source below.
Sources
Posted