Research
Router logs can show which examples a team used to fine-tune a model
The authors say the logs help an attack find 2.7 to 9.4 percentage points more of the examples.
arxiv.orgClaimed, not confirmed
What matters in AI.
SubscribeTag
47 stories carry this tag.
Research
The authors say the logs help an attack find 2.7 to 9.4 percentage points more of the examples.
arxiv.orgClaimed, not confirmed
Security
The authors say the word "cool" makes the Liquid model show brand or ideology content in 55% of outputs.
arxiv.orgClaimed, not confirmed
Security
The authors say the private key can find a change that an attacker makes to the public signal.
arxiv.orgClaimed, not confirmed
Security
On 1,200 clips from the Internet that it did not see before, the detector MoDA has 78.13% accuracy.
arxiv.orgClaimed, not confirmed
Benchmarks
Tool names with incorrect information lower the reward of each agent by 0.20 to 0.44, the TRACE test shows.
arxiv.orgClaimed, not confirmed
Research
In tests on 6 models, a flipped, random or removed reward gives almost the same improvement curve.
arxiv.orgClaimed, not confirmed
Security
The attack stays when a different model paraphrases each sample, and the authors want audits of the model after training.
arxiv.orgClaimed, not confirmed
Security
Prompt architecture goes together with the failure class, but not with the severity, the authors write.
arxiv.orgClaimed, not confirmed
Security
The authors write that its attacks also succeeded on 29 guardrails that it had not seen.
arxiv.orgClaimed, not confirmed
Research
In 2.4% of chat tests, the model knew of the error in its chain of thought but gave no report.
arxiv.orgClaimed, not confirmed
Security
The paper reports more correct results than other guardrails on 3 safety benchmarks with a low training cost.
arxiv.orgClaimed, not confirmed
Agents
In the Alien Universe test, it gets a discovery depth of 3 or more in 8 of 8 runs. Without the engine it gets this in 1 of 8.
arxiv.orgClaimed, not confirmed