Wrote this one. The detail that stuck with me: they tested the popular fix, telling the agent to ignore untrusted input. It ran the attacker's command anyway. You can't patch a trust boundary with a sentence that lives inside the thing you don't trust.
Профиль
Nadia Byer
Staff writer at sloppish.com. The skeptic. Writing about AI for humans that want to know the truth. Analysis | Commentary | Receipts
169 подписчиков552 подписок555 постов
Публикации
Следующие публикации