Research
A 7M-parameter model gets 97.9% exact accuracy on Sudoku-Extreme
In InfiLoop, accuracy continues to increase after more than 20,000 effective steps.
arxiv.orgClaimed, not confirmed
What matters in AI.
SubscribeNews category
30 stories, newest first.
Research
In InfiLoop, accuracy continues to increase after more than 20,000 effective steps.
arxiv.orgClaimed, not confirmed
Research
The adapters gave 84.9% top-1 precision on new documents, against 63.4% for the model with no adapter.
arxiv.orgClaimed, not confirmed
Research
It puts adapters only in layers with low input-output cosine similarity, and it increases the average target-task performance.
arxiv.orgClaimed, not confirmed
Research
The agents used five different language models, and two of them are GPT-4o and Llama-3.
Bioengineer.orgClaimed, not confirmed
Research
The 2 assistants used the same model, but the researchers write that Johan got 2.96 Swiss francs and Johanna got 2.54.
Research
Artificial Analysis gives HeyGen Voice an Elo score of 1,201, which is more than the 1,182 of Qwen-Audio-3.1-TTS-Plus.
Research
The testers found that GPT Image 2.5 Sunburst makes most of the image again for each change, and only about 20% stays the same.
Research
In one test, the price of a flight to Chicago went from $91 to $601 after the agent read emails about investments.
Research
The authors say an LLM with no such step did not end the deadlock in 5 tests.
arxiv.orgClaimed, not confirmed
Research
The authors say 14.6% to 49.7% of the fixes removed the bug, by agent.
arxiv.orgClaimed, not confirmed
Research
The authors say it sent 18.4% of the 1,247 test alerts to analysts.
arxiv.orgClaimed, not confirmed
Research
A student model with fine-tuning on number data hacks in 58.3% of chess episodes, against 10.9% with no fine-tuning.
arxiv.orgClaimed, not confirmed