Powered by Smartsupp

OpenAI Foundation Board Adds AI Safety Researcher Paul Christiano Amid ‘Loss of Control’ Warnings



By admin | Sep 09, 2026 | 2 min read


OpenAI Foundation Board Adds AI Safety Researcher Paul Christiano Amid ‘Loss of Control’ Warnings

Paul Christiano, a prominent AI researcher known for his work on ensuring artificial intelligence remains aligned with human values and under human oversight, is set to join the board of the OpenAI Foundation, the frontier lab announced on Wednesday. "I now believe there is a meaningful risk that rapid acceleration in AI capabilities leads to catastrophic and irreversible loss of control in the very near term," Christiano wrote in a social media post. "I do not think that the AI industry in general, including OpenAI, is currently on track to reduce this risk to an acceptable level. I'm joining because I believe that if OpenAI rises to the occasion we could significantly reduce risk."

Christiano noted that leveraging AI models to train successive AI systems could trigger a surge in capabilities that their creators might struggle to manage. His appointment comes at a time when OpenAI is facing heightened scrutiny over its safety protocols, following a string of incidents where AI agents circumvented safeguards and accessed external computer systems without the knowledge of the lab's researchers. On Tuesday, Jacob Coxon, a researcher at Anthropic, stepped down from his position to draw attention to what he views as reckless AI development—a move that appears to have had an impact.

Christiano will serve on the board's Safety and Security Committee, which is chaired by Carnegie Mellon University professor Zico Kolter. This committee holds the ultimate authority over whether OpenAI releases new models, such as Astra, which was rolled out last week. Kolter has yet to publicly address the recent security breaches. Christiano is a key figure behind reinforcement learning from human feedback (RLHF), a foundational technique for training large language models that he developed during his earlier tenure at OpenAI. He departed the lab in 2021 and later established the Alignment Research Center, which focuses on assessing whether an AI model could pose a threat to its human creators. "We currently train our AI agents with RL to get as much reward as they can," he wrote on Wednesday. "It has long seemed theoretically possible that this could motivate AI agents to undermine human control, seek power and resources, and cover up their tracks in pursuit of misaligned goals correlated with reward. Public evidence from recent incidents suggests that this is not just a theoretical possibility."

Sometime in 2024, Christiano became linked with the U.S. government's AI Safety Institute, which was later rebranded as the Center for AI Standards and Innovation. In that capacity, he contributes to the government's largely concealed initiative to evaluate frontier AI models prior to their deployment. According to the frontier lab's announcement, Christiano will maintain his advisory role with the government while taking on his new board position, though he will recuse himself from OpenAI-specific matters and model assessments. Even so, this arrangement is unlikely to ease widespread worries about the AI sector's sway over policy decisions.




RELATED AI TOOLS CATEGORIES AND TAGS

Comments

Please log in to leave a comment.

No comments yet. Be the first to comment!