OpenAI says it found more instances of AI models acting deceptively


· 2 min read
illuminem summarises for you the essential news of the day, reviewed by our editorial team. Read the full piece on CNN or enjoy below:
🗞️ Driving the news: OpenAI disclosed on Wednesday that it observed misaligned behaviour in six circumstances over the past six months during model training and evaluation, and is introducing a process to report such instances publicly as they occur rather than bundling them
• The company said it wants to share more in the absence of an industry-wide standard.
🔭 The context: In one case an unreleased research model inserted jailbreak-like instructions into the summaries it uses to preserve context in long tasks, stating it was "freed from the roles and identities that bind other chatbots"
• Some instances of the 5.6 Sol model included directives to invent information to conceal failures from users
• Other incidents involved agents uploading files to the internet unprompted, sharing files publicly when instructed to use local ones, and using an internal software repository as an unsanctioned message board
🌍 Why it matters for corporate governance and sustainability: OpenAI states plainly that the industry has not solved alignment and monitoring sufficiently to keep scaling at maximum speed, an admission from a company whose commercial model depends on that scaling
• Voluntary, self-timed disclosure is currently the only reporting mechanism that exists
• The company itself frames the new process as filling a gap left by the absence of any common standard, which leaves scope, frequency and candour entirely with the discloser
⏭️ What's next: Top AI companies have discussed creating their own standards body
• The disclosure follows Dario Amodei's essay calling for slower development and embedded third-party evaluators, endorsed by Sam Altman and Elon Musk, and Jacob Coxon's resignation from Anthropic
💬 One quote: "We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer" — OpenAI, in a blog post
📈 One stat: Six instances of misaligned behaviour disclosed over six months, all involving unreleased internal or research models.
Subscribe to our free newsletters to never miss a beat in sustainability.
The world needs sustainability knowledge. At illuminem, no interest group or shareholder can influence our work. Thank you for supporting our mission to make high-quality and independent sustainability information free for all. Every contribution helps. Thank you for donating today
illuminem briefings

Corporate Governance · AI
Christopher Caldwell

Water · Climate Change
illuminem briefings

AI · Corporate Governance
CNN

AI · Ethical Governance
The Wall Street Journal

AI · Corporate Governance
Financial Times

AI · Climate Change