An attack on world models can give robots dangerous policies
The authors put malicious prompts in data that looks safe, and the attack works only after a world model uses the data.
Claimed, not confirmed
This is a brief. We point to the report and do not rewrite it. Read it at the source below.
Sources
Posted