On October 1, OpenAI fired three safety researchers — Jasmine Wang, Tomek Korbak, and Mikita Balesni — for allegedly mishandling research information. A week later, the three published an open letter disputing their dismissal. One line cuts through: “I believe we were fired for prioritizing safety over the near-term interests of OpenAI as a corporation.” OpenAI denies it. But here’s what neither side disputes: this is at least the tenth OpenAI safety researcher to exit the company since May 2024.
The Pattern Is the Story
The list of departures is long enough to be its own argument. Ilya Sutskever, co-founder and chief scientist, resigned in May 2024. Jan Leike — who co-led the now-dissolved Superalignment team — left the same month and said safety had “taken a backseat to shiny products.” Leopold Aschenbrenner and Pavel Izmailov were fired in April 2024 for alleged leaks. Daniel Kokotajlo left after refusing to sign an exit agreement. Miles Brundage, a long-time policy researcher, departed late 2024. According to Fortune, nearly half of OpenAI’s AGI safety staffers had already left by mid-2024 — a pattern going back to May 2024 that the latest firings only extend.
Last week, David Robinson resigned in an Atlantic essay titled “I Quit OpenAI Because Its Culture Is Broken.” Robinson had spent three and a half years at the company, signing off on safety reports for twelve frontier-model launches. His conclusion: “An environment where things like this can happen is no place to grow artificial minds that could be smarter than we are.” Now Wang, Korbak, and Balesni. Every single departure — whether resignation or termination — traces back to the same structural tension: commercial speed versus safety rigor.
What These Three Actually Worked On
This matters more than the HR dispute. Korbak served as OpenAI’s technical contact for METR, an external organization that independently evaluates frontier AI models. Korbak and Balesni co-authored the “Chain of Thought Monitorability” paper — the technical case for keeping AI reasoning legible to human reviewers. Their open letter now urges the board to stop building models whose reasoning humans cannot audit. Korbak’s concern is precise: OpenAI is “losing the ability to monitor what AI agents think.”
As frontier models move toward hidden chain-of-thought reasoning — internal steps that are never shown to developers or users — the ability to audit model behavior disappears. For developers deploying AI agents in production, that is a real problem. Debugging misbehaving agents is already hard. Debugging agents whose reasoning you cannot inspect is harder still. Trusting them in compliance-sensitive contexts becomes nearly impossible.
The Chilling Effect, Regardless of Who’s Right
OpenAI’s case is not nothing. The company says the terminations stemmed from “a pattern of misconduct” — specifically accessing and handling sensitive information outside established procedures. An internal memo stated the decisions were “not about raising safety concerns or speaking out.” OpenAI also argues it did not fire these researchers when they raised concerns internally; it fired them when, the company says, they violated policy.
However, the chilling effect does not depend on OpenAI being wrong. Even if all three violated policy in exactly the way the company describes, the manner of the firings — abrupt, public, framed as misconduct — sends a message to every remaining safety researcher. The open letter says as much: internal and external communications have made former colleagues “afraid to speak.” Whether or not that was the intent, it is the outcome. ByteIota has previously covered OpenAI’s recent credibility issues around its math papers retraction — this adds another layer to an already complicated public trust picture.
It Is Not Just OpenAI
Anthropic was founded by people who left OpenAI over safety concerns. It now has safety researcher departures of its own. Mrinank Sharma, head of Safeguards Research at Anthropic, resigned recently and published concerns about existential risks. Researchers have left both Anthropic and Google. Jacob Coxon, who worked at both labs, stated that these companies are “gambling with our lives.” The AI safety chilling effect, in other words, is not a single company’s culture problem.
The commercial incentive structure at every frontier lab — raise capital, ship faster, win market share — is in structural tension with the slower, more careful work that safety research requires. Moreover, Anthropic and Google face the same investor pressure. They are showing the same cracks.
What This Means If You Build on These Models
The people who audit model behavior before launch are leaving. The person who signed off on twelve OpenAI model safety reports just quit. The researchers who worked on chain-of-thought legibility just got fired. The next major model launch will have fewer experienced safety reviewers behind it than the last one.
That is not a reason to stop building on frontier models — they remain the most capable tools available. But it is a reason to build with more defensive assumptions: monitor agent behavior aggressively, treat model outputs as unaudited, and do not assume that safety reviews at the lab level substitute for safety reviews in your own stack. OpenAI is still the most capable AI lab in the world. It may also be running out of the people best positioned to keep it honest about what that means.













