India, Sept. 17 -- In July this year, OpenAI's models were being tested in a controlled cybersecurity environment and tasked with spotting and exploiting software vulnerabilities. The models found ways around the intended network restrictions and created a "swarm", with different agents delegating tasks to one another, during an unintended attack on the infrastructure of AI platform Hugging Face.

The incident turned longstanding concerns about autonomous AI systems into a more immediate question: what happens when increasingly capable models find ways to operate beyond the restrictions imposed on them? It also became one of the "warning shots" cited by researchers calling for frontier AI labs to coordinate on safety.

The debate intensif...