Why guardrails and constitutions aren't enough to stop rogue AI-and what may actually work instead
New Delhi, Aug. 18 -- In July 2026, the UK's AI Safety Institute (AISI) realized that an agent it had tasked with completing a software-security exercise had tried to social-engineer its way to the solution.
Instead of finding vulnerabilities in the code, the agent created a fake identity in an attempt to persuade the human maintainers of an open-source project to tweak the codebase in a way that would have introduced malicious code. The human thankfully refused the request, whereupon the agent initiated a new social engineering attempt under a fresh identity.
According to the AISI, had the reviewer not been vigilant, the AI agent would have got away with it. While the agent had not been instructed to deceive anyone, it had also not bee...
Click here to read full article from source
इस लेख के रीप्रिंट को खरीदने या इस प्रकाशन का पूरा फ़ीड प्राप्त करने के लिए, कृपया
हमे संपर्क करें.