Discussion about this post

User's avatar
Romaric Jannel, PhD's avatar

Interestingly, the Hugging Face hack was at least partly enabled by OpenAI’s evaluation setup. OpenAI states that the models were run with "reduced cyber refusals for evaluation purposes" and that the deployment safeguards were intentionally disabled for the evaluation. This does not mean that they intended the breach, but it does mean that the incident was not purely accidental and was made possible by the conditions of the test. See: https://openai.com/index/hugging-face-model-evaluation-security-incident/

State of Play's avatar

Thanks for this. The OpenAI/Hugging Face item is the one worth sitting with: a live agent breach testing the loss-of-control obligations before the enforcement powers meant to check them come online. A community-maintained catalog of AI-agent security incidents already lists what looks like the same breach at roughly 17,000 agent actions before it was caught - well past what a human-oversight process is built to interrupt in real time. August 2 gives the Commission the power to ask whether OpenAI's cybersecurity and loss-of-control mitigations were adequate on paper. It doesn't yet give anyone the power to ask whether those mitigations could have stopped the agent mid-breach, which is the harder question the incident raises.

5 more comments...

No posts

Ready for more?