Benchmarks
Mathematicians find a mismatch in the OpenAI Navier-Stokes proofs
In one lemma, the Lean proof includes more values than the proof in natural language.
newscientist.comClaimed, not confirmed
What matters in AI.
SubscribeNews category
7 stories, newest first.
Benchmarks
In one lemma, the Lean proof includes more values than the proof in natural language.
newscientist.comClaimed, not confirmed
Benchmarks
The researchers find that the resolution rate of the agents is 33.2% and not 50.6% when TestJack does the check.
arxiv.orgClaimed, not confirmed
Benchmarks
Tool names with incorrect information lower the reward of each agent by 0.20 to 0.44, the TRACE test shows.
arxiv.orgClaimed, not confirmed
Benchmarks
Without its guards, the verifier always passes, but the skills that it trains are as good.
arxiv.orgClaimed, not confirmed
Benchmarks
3 of the 4 datasets that can target harmful refusal are saturated and cannot show differences between models.
arxiv.orgClaimed, not confirmed
Benchmarks
Artificial Analysis says a hallucination is material if it can mislead a reader, for example with an incorrect contract date.
Benchmarks
The author tells that the correct test was more work than the fast code.
towardsdatascience.comClaimed, not confirmed