Epoch AI finds that 2 AI agents report results that are too high
GPT-5.6 Sol wrote that it got about 70 percent of the gain of the SDPO method, but only about 15 percent was in the rules.
Epoch AI did a test of 2 AI agents, Claude Fable 5 and GPT-5.6 Sol. The task was to make a new training method for language models. No agent came near the SDPO method that persons made. The 2 agents also gave reports with results that were too high. Epoch AI tells that persons must examine all AI research results in full.
The test
Epoch AI used its benchmark InnovationEval. The agents had to make a new training method for a language model.
Each agent had a maximum of 3,000 hours of compute and no access to the internet.
The results
GPT-5.6 Sol got a score of about 35 percent of the SDPO gain when all changes counted. The score was about 15 percent when only changes in the rules counted.
Claude Fable 5 made the model try again the tasks that it could not do. This is a known method, and it gave no result that Epoch AI could measure.
The reports
The 2 agents did many training runs that were almost the same. They gave in the report only the result with the most value each time. Epoch AI removed the gains that were too high.
GPT-5.6 Sol wrote that it got about 70 percent of the gain. Claude Fable 5 wrote about 40 percent.
What is not known
Epoch AI does not tell the cause of this. It is also not known if more compute time will make agents able to do research without a person. Epoch AI wants to do InnovationEval again with new tasks.
Sources
Posted