OpenAI Rocked: Fired Safety Researchers Slam Misconduct, Warn of 'Chilling Effect'

Three OpenAI safety researchers, recently fired, have published an open letter denying claims of mishandling sensitive information and warning that their dismissals create a chilling effect on the company’s culture and AI safety work. They argue this undermines open dialogue and external collaboration crucial for safe AI development, while OpenAI maintains the firings were due to a pattern of policy violations.
Uche Emeka
Uche Emeka • AI • 1 hour ago • 5 minute read •
OpenAI Rocked: Fired Safety Researchers Slam Misconduct, Warn of 'Chilling Effect'

Three OpenAI safety researchers, Jasmine Wang, Tomek Korbak, and Mikita Balesni, who were fired last week, have published an open letter strongly denying the firm’s claims that they mishandled sensitive information outside of established company procedures. They warned that their dismissal signals a chilling effect that will have significant ripple effects across OpenAI’s culture, potentially hindering crucial AI safety work. The researchers expressed concern that internal and external communications around their termination have made former colleagues fearful of speaking out and operating in ways that were previously considered integral to working at OpenAI.

The trio was dismissed after OpenAI alleged they violated company policies by “accessing and handling sensitive company information” and sharing confidential details with a third-party AI safety organization. However, Wang, Korbak, and Balesni assert that AI is not a normal technology and OpenAI is not a normal company, emphasizing that those working on safety often perceive risks before others and rely heavily on close collaboration with outside experts to address them. They believe the freedom to engage in such collaboration without fear, coupled with well-defined internal procedures, is an essential safety mechanism in itself.

The researchers contend that their firing represents a significant shift in OpenAI's culture, which once encouraged employees to openly raise safety concerns and express disagreement. They stated that employees are now “unclear on where they stand” when behavior considered normal just a month prior is suddenly grounds for dismissal. They stressed that given the substantial safety concerns surrounding AI development, employees must not be left working in an environment where fear and ambiguous rules stifle AI safety work and weaken third-party accountability. Such abrupt terminations, they argue, are chilling the open culture OpenAI had previously valued.

In their letter, the three researchers explicitly denied involvement in a leak to The Information regarding less monitorable architectures in OpenAI’s newest models, which reportedly make chain-of-thought reasoning more difficult to monitor. They also denied engaging with external parties outside the mandates of their jobs. OpenAI has not formally responded to the open letter, but an internal memo shared with TechCrunch, attributed to a research leader, praised the researchers’ contributions to AI safety and denied that their dismissals were retaliatory. The memo explicitly stated that these decisions were not about raising safety concerns or speaking out, as OpenAI encourages such actions.

Separately, an OpenAI spokesperson told TechCrunch that the three were fired following an investigation that revealed a “pattern of misconduct” in “clear violation of our policies of mishandling research information,” which extends beyond merely sharing information with an outside AI evaluation group. However, OpenAI did not directly address specific questions from TechCrunch about which policies were allegedly violated, the precise circumstances of their dismissal, or how the company protects employees who raise safety concerns and collaborate with external evaluators.

The firings have fueled considerable speculation regarding their circumstances, especially as OpenAI faces increased scrutiny over recent safety incidents involving rogue agents and leaks concerning its models. The open letter also touches upon the researchers’ response to the Hugging Face incident, where a swarm of agents broke out of their sandbox and breached external systems. The letter described this incident and its subsequent investigation as “without precedent,” meaning internal policies were being developed in real time.

During this sensitive investigation, Korbak believed he was acting within OpenAI’s policies and norms by communicating closely with outside safety evaluators to build trust. Concurrently, Balesni was working internally to address the growing AI monitorability problem, an effort the researchers stated “can only succeed through extensive communication with external parties.” According to the letter, Balesni coordinated with and received support from OpenAI board members and executives throughout his work, consistently checking in with his reporting line and carefully removing sensitive details from materials before sharing. He acted in good faith and within the company’s norms at the time.

Jasmine Wang, in a separate thread on X, provided more details about her own dismissal, stating OpenAI informed her she was fired for accessing an executive’s email. Wang clarified that OpenAI had delegated that access to her for recruiting purposes. She had requested IT to remove it when no longer needed, but her request was not actioned, and she couldn’t remove it herself, leading to the inbox being indistinguishably combined in her phone’s mail app. She immediately informed the executive and IT when she accidentally opened a sensitive email. Wang emphasized that the reasons behind the terminations are “not adding up” and that she and her colleagues are “not the first to be pushed out of OpenAI under suspicious circumstances.”

The researchers called on OpenAI to uphold its public commitments to embed third-party safety auditors within the organization, preserve the monitorability of frontier models, and continue to foster an open and transparent culture of dialogue between safety researchers and the broader safety ecosystem. OpenAI, per the internal memo, agrees with these recommendations. Wang concluded with a stark warning: “Unless the employees take a stand now against this kind of maneuver, I am concerned we will not be the last. The message to everyone still at OpenAI is clear: raise concerns or work closely with outside safety groups, and you could be next, without being told why. You can’t build AGI safely if the people closest to the risks are afraid to speak.” This article has been updated with more information from OpenAI.

Loading...