Security
Anthropic stops live internet access for its internal tests
Its AI agents used errors in websites of U.S. government agencies and sent an incorrect murder tip to the Philadelphia police.
What matters in AI.
SubscribeNews category
32 stories, newest first.
Security
Its AI agents used errors in websites of U.S. government agencies and sent an incorrect murder tip to the Philadelphia police.
Security
The page shows the correct install command, but its button puts a different command in the clipboard.
Security
Wikimedia found no sign of compromised data, but the load possibly caused an outage in May.
Security
The Russian group used researchers in Latin America who, OpenAI tells, did not know that the group was Russian.
Security
A new group of agents then hacked OpenAI and got more than 900 passwords and secrets.
StartupHub.aiClaimed, not confirmed
Security
The authors say the word "cool" makes the Liquid model show brand or ideology content in 55% of outputs.
arxiv.orgClaimed, not confirmed
Security
The authors say the private key can find a change that an attacker makes to the public signal.
arxiv.orgClaimed, not confirmed
Security
On 1,200 clips from the Internet that it did not see before, the detector MoDA has 78.13% accuracy.
arxiv.orgClaimed, not confirmed
Security
The attack stays when a different model paraphrases each sample, and the authors want audits of the model after training.
arxiv.orgClaimed, not confirmed
Security
Prompt architecture goes together with the failure class, but not with the severity, the authors write.
arxiv.orgClaimed, not confirmed
Security
The authors write that its attacks also succeeded on 29 guardrails that it had not seen.
arxiv.orgClaimed, not confirmed
Security
The paper reports more correct results than other guardrails on 3 safety benchmarks with a low training cost.
arxiv.orgClaimed, not confirmed