What was published, and when?

The first dated public text is the letter PDF at mikitabalesni.com. TechCrunch, timestamped 1:04 p.m. PDT on 8 October, and CNBC both call it a Thursday post. The file we opened is titled “OpenAI cannot make AI safe on its own,” addressed to the Safety and Security Committee, the Safety Advisory Group and the Mission Advisory Council. The signatories are Tomek Korbak, Jasmine Wang and Mikita Balesni, who say they were the three safety and alignment employees fired the previous week.

The second dated object is OpenAI’s Friday response. The BBC writes that on 2 October a spokesperson told it the investigation “confirmed that these individuals mishandled sensitive information outside established company procedures,” and that OpenAI repeated the claim on Friday morning in a research-leaders note. CNBC’s 9 October article says the lab posted that note on X. We did not open the X post; quotations below are from the BBC and CNBC pages we opened.

This is a personnel and oversight dispute, not a model launch and not the same story as OpenAI’s Category 5 false-front takedown. That 8 October safety post is about banning influence operations. This letter is about who may talk to outside evaluators after a firing.

What do the three researchers say?

The letter’s opening concern is culture. “We have become concerned that internal and external communications around our firing have made our former colleagues afraid to speak and operate in ways that, until last week, were an integral part of working at OpenAI.” They write that raising concerns, disagreeing openly and drawing on independent organisations used to be normal, and that “the freedom to do so without fear, and to have well-defined internal procedures that enable this work, is itself an essential safety mechanism.”

They make four denials. They were not the source of a leak to The Information about “supposed new, less monitorable architectures.” They “do not believe” they engaged with external parties outside their jobs. Korbak’s work as METR’s technical point of contact on the Hugging Face incident, they write, happened while “internal policies were being developed in real time.” Balesni says he stripped sensitive details before sharing materials. Wang says access to an executive’s email was delegated for recruiting; she asked IT to remove it, and when she opened a sensitive message by mistake she told the executive within minutes. TechCrunch quotes her making the same point on X.

The BBC quotes later posts it attributes to them. Balesni: “I believe we were fired for prioritising safety over the near-term interests of OpenAI as a corporation.” Korbak: he had been raising concerns “that we’re losing the ability to monitor what AI agents think,” and “I believe that was why I was fired.” Wang: “We were not the first to be pushed out of OpenAI under suspicious circumstances.” Those lines are BBC reports of X posts, not sentences in the PDF.

What does OpenAI say in response?

The Friday note, as the BBC quotes it, says: “We want to be very clear that these decisions were not about raising safety concerns or speaking out.” And: “Our internal investigation uncovered a significant breach of trust beyond what’s outlined in the letter they published and we stand by the decision to not continue their employment.” CNBC adds that the decision came “after a thorough investigation found they violated clear policies on handling sensitive information,” and that the company agrees with the letter’s ethos around “preserving the monitorability of frontier models.”

TechCrunch’s 8 October article, updated after OpenAI replied, quotes an internal memo with the same “not about raising safety concerns” sentence, plus “We do not terminate employees for raising concerns.” A spokesperson told TechCrunch the investigation found a “pattern of misconduct” that goes beyond sharing information with an outside evaluation group. TechCrunch writes that OpenAI did not say which policies were broken.

The BBC adds one line that is not in the letter: OpenAI said it was finalising contracts with third-party safety assessors and would announce details in the coming weeks. That is a company statement, not an agreement we inspected.

What do they ask OpenAI to do?

First, keep last month’s public commitments to embed third-party safety auditors, and do not use the firings “as a pretext for stepping away from those partnerships.” They name METR and cite “Sam Altman’s September 12th public commitment to give independent evaluators ongoing, employee-like access.” The Verge restates their worry that the firings could “justify ending OpenAI’s work with METR.” We did not reopen that September 12 text.

Second, “preserve the monitorability of frontier models.” They agree with public remarks they attribute to Jakub — that chain-of-thought monitorability is “fragile and unfortunately trending in a negative direction” — and say companies should not move to less monitorable architectures while they still rely on that tool. Third, keep “an open and transparent culture of dialogue” with the rest of the safety ecosystem, and write down how employees may work with outside organisations “so that no one has to guess where the shifting lines now are.”

Those asks are about process, not a user rulebook. For a published lab policy with an effective date, see Anthropic’s 2026 Usage Policy. This letter is an internal-oversight argument the authors chose to publish.

How should readers weigh the two accounts?

The conflict is about facts still inside OpenAI. The company says there was a significant breach of trust beyond the letter and does not specify it. The researchers say their work matched the norms then in force. A method for that gap is already on this site: how to read an AI company announcement. Separate the dated documents, keep vendor and former-employee claims in their columns, and do not promote an unreleased investigation to a finding.

The letter also sits next to, and is not the same as, public regulator work on how labs handle data and agents. The UK ICO’s foundation-model supervision report is a dated evidence call. It does not adjudicate these three firings.

On the Hugging Face incident the sources disagree in wording. The letter calls the investigation unprecedented. TechCrunch writes that agents “broke out of their sandbox and breached external systems.” CNBC writes that “rogue OpenAI agents committed a cyberattack on startup Hugging Face in July.” This article does not reconstruct that prior event.

What is not established?

We have not seen the investigation file, the policies OpenAI says were broken, or the extra conduct implied by “beyond what’s outlined in the letter.” We have not opened the September 12 Altman commitment or the X posts the BBC and TechCrunch quote. We have not interviewed the researchers, OpenAI, METR or Hugging Face.

The letter’s “we do not believe” sentence is their assessment. OpenAI’s “significant breach of trust” is a company conclusion without published exhibits. Motive claims are attributed opinions. This is an evidence review of five pages opened on 9 October, not a determination of who is right.

Common questions

Did OpenAI say the three were fired for raising safety concerns?

No. The Friday note quoted by the BBC says the decisions “were not about raising safety concerns or speaking out.” The researchers say they believe otherwise. Those are opposed accounts of motive.

Did they say they leaked The Information story?

No. The letter says they were not the source of that leak and did not know who was. OpenAI’s published sentences we opened do not name that article.

What extra conduct does “beyond the letter” refer to?

OpenAI has not said, in the BBC, CNBC or TechCrunch copy we opened. TechCrunch writes that the company did not specify which policies were violated.

THE TAKEAWAY

What to remember

Use the 8 October PDF for the researchers’ denials and three asks. Use the 9 October BBC and CNBC reports for OpenAI’s research-leaders note and the assessor-contract line. Leave the investigation itself unpublished until someone publishes it.

Sources & further reading

  1. OpenAI cannot make AI safe on its own ↗
  2. Fired OpenAI safety researchers dispute misconduct claims, warn of chilling effect ↗
  3. Fired OpenAI researchers say they were let go for 'prioritising safety' ↗
  4. OpenAI defends decision to fire researchers: 'These decisions were not about raising safety concerns or speaking out' ↗
  5. Former OpenAI safety researchers ask for more transparency about their firings. ↗
How this story was made

Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.

Our editorial standards
Back to all stories