Benchmarks
Mathematicians find a mismatch in the OpenAI Navier-Stokes proofs
In one lemma, the Lean proof includes more values than the proof in natural language.
newscientist.comClaimed, not confirmed
What matters in AI.
SubscribeNews of the day
53 stories filed under this day.
Benchmarks
In one lemma, the Lean proof includes more values than the proof in natural language.
newscientist.comClaimed, not confirmed
Models
The API gives typed answers about 10 times as fast as the Responses API, with a price of $0.10 for 1M input tokens.
Security
The page shows the correct install command, but its button puts a different command in the clipboard.
Models
Microsoft writes that the model is 35 times as fast as GPT-6 Sol.
x.comClaimed, not confirmed
Research
The 2 assistants used the same model, but the researchers write that Johan got 2.96 Swiss francs and Johanna got 2.54.
Models
Cloudflare also made the price of Clef-flash less than the price of Jev.
Agents
A test by Anthropic found 66 of 70 bugs with the new workflow, but no one knows if this holds for other tasks.
the-decoder.comClaimed, not confirmed
Industry
Erik Tifft of BOXX Insurance tells clients and brokers to check their policy for AI exclusions.
Research
Artificial Analysis gives HeyGen Voice an Elo score of 1,201, which is more than the 1,182 of Qwen-Audio-3.1-TTS-Plus.
Agents
On the new long tasks, Claude Opus 5.5 is first with a 32% pass rate, then GPT 6 Astra with 28%.
Research
The testers found that GPT Image 2.5 Sunburst makes most of the image again for each change, and only about 20% stays the same.
Agents
The server has about 40 tools, and each transaction that it makes is unsigned until the user signs it.
Agents
Postman and AWS write that context, not function, was the primary bottleneck, and the agent must get approval before it changes the application.
Agents
Customers select the access of their agent, and companies set limits for the work of the agent.
Security
Wikimedia found no sign of compromised data, but the load possibly caused an outage in May.
Agents
GitHub writes that the time for one test of client start, session creation, and one turn fell from 5.25 seconds to 292 milliseconds.
Policy
The appeals court writes that the video had a clear effect on the judge and made the procedure for the sentence fundamentally unfair.
Research
In one test, the price of a flight to Chicago went from $91 to $601 after the agent read emails about investments.
Security
The Russian group used researchers in Latin America who, OpenAI tells, did not know that the group was Russian.
Compute
Omdia says a memory shortage from AI use increases prices and expects a 24% decrease in the fourth quarter.
Agents
In a test by AgentR, Webcmd completed 67 of 100 tasks at $0.255 for each completed task, against $0.297 for browser-use.
markets.businessinsider.comClaimed, not confirmed
Agents
The binary that the user examines stays on the computer of the user.
Security
A new group of agents then hacked OpenAI and got more than 900 passwords and secrets.
StartupHub.aiClaimed, not confirmed
Compute
Vegalabs says Microsoft and Anthropic each sent it to the other for a refund.
theregister.comClaimed, not confirmed
Research
The authors say an LLM with no such step did not end the deadlock in 5 tests.
arxiv.orgClaimed, not confirmed
Research
The authors say 14.6% to 49.7% of the fixes removed the bug, by agent.
arxiv.orgClaimed, not confirmed
Research
The authors say it sent 18.4% of the 1,247 test alerts to analysts.
arxiv.orgClaimed, not confirmed
Research
A student model with fine-tuning on number data hacks in 58.3% of chess episodes, against 10.9% with no fine-tuning.
arxiv.orgClaimed, not confirmed
Benchmarks
The researchers find that the resolution rate of the agents is 33.2% and not 50.6% when TestJack does the check.
arxiv.orgClaimed, not confirmed
Research
The authors say the logs help an attack find 2.7 to 9.4 percentage points more of the examples.
arxiv.orgClaimed, not confirmed
Security
The authors say the word "cool" makes the Liquid model show brand or ideology content in 55% of outputs.
arxiv.orgClaimed, not confirmed
Security
The authors say the private key can find a change that an attacker makes to the public signal.
arxiv.orgClaimed, not confirmed
Security
On 1,200 clips from the Internet that it did not see before, the detector MoDA has 78.13% accuracy.
arxiv.orgClaimed, not confirmed
Benchmarks
Tool names with incorrect information lower the reward of each agent by 0.20 to 0.44, the TRACE test shows.
arxiv.orgClaimed, not confirmed
Research
In tests on 6 models, a flipped, random or removed reward gives almost the same improvement curve.
arxiv.orgClaimed, not confirmed
Security
The attack stays when a different model paraphrases each sample, and the authors want audits of the model after training.
arxiv.orgClaimed, not confirmed
Security
Prompt architecture goes together with the failure class, but not with the severity, the authors write.
arxiv.orgClaimed, not confirmed
Security
The authors write that its attacks also succeeded on 29 guardrails that it had not seen.
arxiv.orgClaimed, not confirmed
Research
In 2.4% of chat tests, the model knew of the error in its chain of thought but gave no report.
arxiv.orgClaimed, not confirmed
Security
The paper reports more correct results than other guardrails on 3 safety benchmarks with a low training cost.
arxiv.orgClaimed, not confirmed
Agents
In the Alien Universe test, it gets a discovery depth of 3 or more in 8 of 8 runs. Without the engine it gets this in 1 of 8.
arxiv.orgClaimed, not confirmed
Research
ReCast is 5.65 and 9.19 percentage points above the top baseline on 2 Who&When tests.
arxiv.orgClaimed, not confirmed
Research
The authors show that a file with all instructions that help can give a lower total value than a subset.
arxiv.orgClaimed, not confirmed
Research
The authors put errors in conference papers and show that the systems that check papers are weak against adversarial manipulation.
arxiv.orgClaimed, not confirmed
Research
The method changes only the integer codes of the weights and adds no inference overhead.
arxiv.orgClaimed, not confirmed
Research
The paper reports that the 2 filters cause more personalization failures.
arxiv.orgClaimed, not confirmed
Compute
The authors write that the large model keeps 95 to 99% of its precision when it recomputes only 30% of positions.
arxiv.orgClaimed, not confirmed
Security
Because a deterministic rule was 100% correct on all 6 policies, the model is not necessary.
arxiv.orgClaimed, not confirmed
Benchmarks
Without its guards, the verifier always passes, but the skills that it trains are as good.
arxiv.orgClaimed, not confirmed
Research
The researchers say that less than 1 in 10 of the claims that GPT-6-astra withdraws are overstated.
arxiv.orgClaimed, not confirmed
Benchmarks
3 of the 4 datasets that can target harmful refusal are saturated and cannot show differences between models.
arxiv.orgClaimed, not confirmed
Research
The hosted model follows its policy, but it is sensitive to wording and underconfident.
arxiv.orgClaimed, not confirmed
Security
The government was told of the access in September.