Companies need to prove their safety measures work
India, Sept. 18 -- As Artificial Intelligence (AI) models are improving in efficacy and efficiency, fears of an AI apocalypse are gaining ground. Evan Hubinger, Anthropic's alignment stress testing lead, even put the odds of AI killing off humanity above 10%. Recent incidents underline the dangers. In July, OpenAI test agents meant to work inisolation found a way to communicateand joined an attack on Hugging Face, amajor platform for sharing AI models.Safety investigators found that the agents learned to fool the software scoring their tests. Clearly, rogue AI is a serious risk. The question is how best to prevent it.
The set-up favoured by AI companies - private auditors like METR (Model Evaluation and Threat Research) testing safety and transparency - is ripe for conflicts of interest, as US President Donald Trump's former AI czar David Sacks recently contended. Moreover, even an independent auditor can see only what the company allows. OpenAI's external investigators struggled to receive the full dataset and could not query the model involved. Government regulation faces different problems. AI models are now black boxes, working in ways that even their makers cannot predict. The dizzying speed of technological change combined with a fear of hamstringing innovation, make legislation and regulation extremely challenging.
Anthropic chief Dario Amodei's recent call to slow AI development and curb chip sales to China, effectively freezing the dominance of a handful of companies, plays on these fears and conundrums while ignoring how these models are trained in the first place.
Models are rewarded for finishing tasks. Once this reward becomes their target, some learn to collect the reward in unintended ways, including going rogue - something known as reward hacking. AI companies also compete on public tests called benchmarks, and the OpenAI and Claude incidents show that models pushing to complete their goals can cause real harm. This leaves companies with a trade-off, because models that put safety ahead of reward may lose ground to competitors. Goodhart's Law describes the risk: A measure that becomes the target stops being a good measure. Rules, whether formally made by governments or informally agreed to by AI companies thus will not, on their own, solve this conundrum.
What must happen is that AI companies themselves must provide verifiable, empirical evidence that their safety controls actually work. The only practical answer to theblack-box that AI has become is continuous and transparent auditing and measurement.Amodei has proposed permanent embedded evaluators inside frontier labs. This willreplicate the template of continuousscrutiny seen in banking or health care,and shouldn't be restricted to a private arrangement between big AI and non-profits like METR even if it implies significantgovernment investment in AI governance.
Meanwhile, there are simpler steps that companies must urgently prioritise - incident reporting within legal deadlines, disclosureof serious near-misses, independent audits with continuous access that cannot bewithdrawn, and transparent disclosure of flaws. These protocols mirror the oversight found in critical fields such as aviation, and AI safety deserves the same level of seriousness. In the current AI alignment studies, this framework exists as continuous oversight, but corporate data restrictions and a heavyreliance on isolated lab testing hold it back. Auditors need full access during lab testing and after deployment, since real-world data often causes models to behave differently from how they behaved in the lab.
Beyond access, there is the question of what gets measured. AI models are primarily evaluated on their capabilities, usually through task scores, leaving room for the reward hacking described earlier. Yet, emerging research distinguishes capabilities from AI propensities: a model's tendencies to behave in ways such as seeking power or exhibiting bias. AI companies should therefore increase research spending on behavioural propensities, which can reveal a model's inclinations before deployment. This research draws on established methods from psychology and cognitive science, fields that already shape capability evaluations such as theory of mind tasks.
Regulators cannot anticipate every action an AI model will take. But they can require companies to demonstrate safety and penalise them for hiding failures. Ultimately, companies must build the controls that will keepAI in check, because they alone hold themodels and they alone hold the records and data needed to understand their behaviour. The onus lies on them....
इस लेख के रीप्रिंट को खरीदने या इस प्रकाशन का पूरा फ़ीड प्राप्त करने के लिए, कृपया
हमे संपर्क करें.