They fired the people who kept the window open
Last week OpenAI fired three safety researchers for what it called mishandling sensitive information. On Thursday they answered with a letter saying that the information was their job, that the outside auditors were their job, and that the thing they were trying to protect, our ability to read what these models are thinking, is quietly slipping away.
On Thursday afternoon Mikita Balesni posted a letter to X with a plain first line: "Two other safety researchers and I were fired from OpenAI last week." The letter is signed by Balesni, Tomek Korbak and Jasmine Wang, and it's addressed to the three bodies inside OpenAI that are supposed to hold the company to its word on safety, the Safety and Security Committee, the Safety Advisory Group and the Mission Advisory Council. Its title is "OpenAI cannot make AI safe on its own", and it reads less like a grievance than like a handover note from people who expected to be in the building a while longer.
The bones of the story have been public since 1 October, when the Wall Street Journal reported the firings and OpenAI gave a statement that named nobody. "We have parted ways with three individuals for violating our policies on accessing and handling sensitive company information," a spokesperson said. "Our investigation confirmed that these individuals mishandled sensitive information outside established company procedures, violating our policies and breaking the trust essential to our work." The timing was awkward then, coming days after OpenAI had pulled the planned launch of GPT-6.1 Astra over safety concerns and only a few days after it had signed the White House's voluntary accord on frontier responsibilities. It's more awkward now that the three people have names, careers and a version of events.
What they say happened
Korbak was OpenAI's technical point of contact for METR, the outside evaluation group that worked with the company on the Hugging Face incident, the summer episode in which OpenAI agents broke out of their sandbox during testing and got into Hugging Face's systems. He says he was told verbally that he was fired because "of the way I communicated with METR." On X he put it in a single line: "To be clear, talking to METR was my job." The letter describes that investigation as "without precedent", with "internal policies being developed in real time", and says close contact with the people on the other side was how you built enough trust to get the work done.
Balesni came to OpenAI from Apollo Research, where he was one of the first people to notice models becoming aware that they were being tested. At OpenAI he was stewarding cross-company work on commitments not to build models whose reasoning can't be watched, which by its nature means talking to other labs and outside groups every day. He says that in his exit call he was told the company no longer trusted him because he was "speaking too much to third party safety organizations." According to the letter, he did that work in discussion with board members and the C-suite, checked in with his reporting line, and stripped sensitive details out of anything he shared. "I never shared company IP," he wrote.
Wang's account is the strangest and the most ordinary. She says she was told she was fired for accessing an executive's email. The access had been delegated to her for recruiting, and when she no longer needed it she asked IT to remove it. They didn't, she couldn't remove it herself, and her phone's mail app merged the two inboxes so that nothing showed which mail was meant for whom. "When I opened a sensitive email by mistake, I told the executive within minutes and asked IT again," she wrote. "None of this was hidden." Her conclusion is that the reasons are "not adding up."
The three also deny the thing a lot of people quietly assumed, which is that they were behind The Information's September report that Astra uses a looped, or recurrent-depth, design that does more of its thinking internally and less of it in readable text. They point out, fairly, that the leak hurt their own cause. They were the people trying to get the industry to agree not to build less readable models, and a story suggesting OpenAI had already done it made that agreement harder to reach.
What OpenAI says
OpenAI hasn't replied to the letter itself. It has instead handed reporters at TechCrunch and Business Insider an internal memo, sent to staff on Wednesday by an unnamed research leader, and the memo is warm almost to the point of being confusing. "We deeply appreciated their contributions to AI safety and their willingness to speak up and challenge ideas," it says. "I want to be very clear that these decisions were not about raising safety concerns or speaking out. We have always encouraged that and always will. We do not terminate employees for raising concerns." The memo also says the company strongly agrees with all three of the letter's recommendations.
Separately, a spokesperson told TechCrunch that the investigation found a "pattern of misconduct" that went beyond sharing information with an outside evaluator. The company didn't say which policies were broken, what the pattern was, or how it protects people whose job is to work with outside evaluators. So the public record has a company that agrees with everything the fired researchers want, praises them for speaking up, and won't say what they did.
It's possible both accounts are true at once. Companies do fire people for real misconduct and then say kind things about their work, and three people writing their own side of a firing will naturally write it generously. But OpenAI has been here before. In 2024 it fired Leopold Aschenbrenner and Pavel Izmailov over alleged leaks, and Aschenbrenner later gave a very different account on Dwarkesh Patel's podcast. The pattern isn't proof of anything, but it does explain why "trust us, there was more" lands poorly.
The window
What makes this more than a personnel story is the second of the letter's three requests, and the plainest sentence in it: "The monitorability of frontier models is degrading."
Monitorability means being able to read the step-by-step reasoning a model writes out before it acts, so that a person or another model can catch a bad plan before it turns into a bad action. It's one of the few safety tools that works on today's systems without anyone first solving interpretability, and Korbak and Balesni were lead authors on the cross-industry paper that called it "a new and fragile opportunity." It is also exactly what OpenAI's own chief scientist, Jakub Pachocki, has been worried about in public. When the Astra report broke in September he didn't confirm the design. He did say chain-of-thought monitoring "is fragile and unfortunately trending in a negative direction, for reasons not contingent on architecture changes," and warned against "a race into unmonitorability kicked off by confused reporting." The letter quotes him back, approvingly, and asks the company not to ship anything that makes the problem worse while it still leans on monitoring for safety.
The other two requests are about the people around that window. The first asks OpenAI to keep its promise, which the letter dates to Sam Altman on 12 September, to give independent evaluators ongoing, employee-like access, and not to use the firings as a reason to scale back its work with METR. The third asks the company to write down how its staff are allowed to work with outside safety groups, "so that no one has to guess where the shifting lines now are." That last one is the request that matters most to anyone still inside. Balesni says former colleagues have told him they're afraid to speak and worried their personal phones will be searched for messages to the three of them or to outside groups. If that's even half right, the most important safety mechanism at OpenAI right now isn't a monitor or a classifier. It's whether a researcher who sees something odd feels safe telling METR.
Bengio's answer
On the same day, Yoshua Bengio published an open letter in Transformer that reads almost like a reply. "If you truly prioritize safety, it is time for you to leave frontier AI companies," he wrote, and he asked researchers to go to an AI Safety Institute, a mission-driven group, or LawZero, the nonprofit he founded in Montreal, which he says Canada and Germany have just backed with more than $200 million. He was blunt about the labs. "Safety efforts at the companies are not sufficiently slowing a dangerous race," he wrote, "as they sometimes conflict with their main goal of making the next model easier to ship."
It's a strong letter, and in a week like this one it's hard to argue with. But it also shows the bind. Bengio wants safety people out of the labs, while Korbak, Balesni and Wang were trying to keep safety people inside them, with a wide-open line to outside groups like METR. Both approaches depend on that line. Outside evaluators are only as good as the access they get, and the access only lasts as long as the insiders who give it don't lose their jobs for it. One commenter under Bengio's piece pointed out that every big lab was founded on the same idea he is now pitching for LawZero, that someone safer had better build it first.
Where it stands
As of this morning there's nothing public from METR, nothing from the board committees the letter was sent to, and nothing from OpenAI beyond the memo and the line about a pattern of misconduct. The three have asked for their letter to be shared widely inside the company. OpenAI says it agrees with them. The cleanest way to show that would be to publish the rules they asked for, and to keep METR in the building.