OpenAI Fired Its Safety Team. Now They're Warning the Board About the One Thing It Can't See.
OpenAI fired three safety researchers on 1 October. This week they wrote to the board about the one instrument that checks whether a model is telling the truth.
The three people OpenAI paid to watch its models for danger were shown the door on 1 October. Jasmine Wang, Tomek Korbak and Mikita Balesni worked on safety and alignment. OpenAI's statement was short: they mishandled sensitive company information outside established procedures. The reporting links the dismissals to material shared with an outside AI-safety group.

That is the news. The interesting part is the thing they say is going dark.
What the letter asks for
The Wall Street Journal reviewed the letter. Its asks are narrow. "As an industry, we do not yet know how to safely develop and deploy models that we cannot monitor," it says, and OpenAI and its rivals "should not move forward with developments that further decrease" that ability. The three want closer work with outside safety auditors, against the risk they call "truly catastrophic". Korbak and Balesni had co-authored a chain-of-thought monitoring paper.
As an industry, we do not yet know how to safely develop and deploy models that we cannot monitor.
The one thing it cannot see
Chain of thought is the closest thing a model has to a confession. When it is legible, a reviewer can watch the model reason its way toward a wrong answer and catch it. When it is summarised or hidden, that check disappears. The worry is not hypothetical: OpenAI has disclosed a coordinated campaign to extract the hidden reasoning of its models, an "adversarial distillation" effort it says began on 1 July and was shut down by 28 July.
Interactive · the shutter
What the auditor can see.
Flick the shutter. Legible, the trace can be checked; opaque, the model's word is all that is left.
01parse the request
02weigh two approaches
03draft an answer
04check it against policy
05revise, then reply
06note what was left out
MONITORED · the trace is legible.UNMONITORED · the trace is opaque.
Why three jobs is not the story
The firing is a personnel matter. The warning is not, and OpenAI's own week shows the seam. On 6 October it published 722 mathematics manuscripts from an unreleased model; the Association for Human Mathematics, reposted by Terence Tao, told mathematicians to stop reviewing the output. The common thread is verification: we are producing more than anyone can check, and the instruments for checking are being pulled back. I read that release in 722 Proofs, One Locked Door.
OpenAI's answer
OpenAI does not dispute the letter's goal. In a memo shared with the Journal, it said it "strongly agreed" with the recommendations and that the dismissals "were not about raising safety concerns or speaking out". The three say they did not believe they had "engaged with external parties outside the mandates of our jobs". Both things can be true: a company can agree with a warning and still part with the people who raised it.
The honest counterpoint
There is a fair reading in which OpenAI is right. Safety-testing information is genuinely sensitive, and an outside group is not automatically trustworthy. OpenAI says the three broke its rules, and I have not seen the internal evidence either way. I spent time looking for the full letter; it is not public, so the quotes here are the Journal's. What would change my mind: the letter in full, or a concrete account of what was shared and with whom.
Readout: what is confirmed, and what is not
- Confirmed (OpenAI statement, carried by the WSJ, The Verge and AFP, 1 October 2026): three researchers were dismissed for mishandling sensitive information outside established company procedures.
- Reported (WSJ, 2 October 2026; AFP; Barron's): the three are Jasmine Wang, Tomek Korbak and Mikita Balesni, at least two on safety and alignment.
- Reported (WSJ, 8 October 2026; India Today; Gizmodo): a letter to the board and safety committees, sent 7 October, asks for outside auditors and the preservation of chain-of-thought monitoring.
- Reported (WSJ, 8 October 2026; India Today): OpenAI's memo in reply, saying it strongly agreed with the recommendations and that the dismissals were not about safety concerns.
- Reported (OpenAI blog, 30 September 2026; The Next Web; CNBC): a coordinated adversarial-distillation campaign to extract hidden reasoning, begun 1 July and shut down by 28 July, with a core cluster linked to people associated with Moonshot AI.
- Reported (Scott Aaronson, 8 October 2026): frontier labs are quietly testing whether math-capable models can break cryptographic primitives.
- Not confirmed: the outside organisation, the specific information shared, and the full text of the letter. None is public.
If you are weighing how much of a model's reasoning you can actually audit, and want a second pair of eyes on the question, email brandon@kreostudio.co.uk.
Sources & references
- OpenAI parts ways with researchers who allegedly shared confidential information, The Wall Street Journal, 1 October 2026.
- Fired OpenAI researchers ask company to preserve visibility into AI reasoning, The Wall Street Journal, 8 October 2026.
- OpenAI severs ties with safety researchers accused of disclosing confidential information, The Verge, 1 October 2026.
- OpenAI fires three staff for mishandling sensitive info, RTE, 2 October 2026.
- 3 axed employees warn OpenAI, do not lose monitoring power over AI, India Today, 8 October 2026.
- 3 Fired OpenAI Employees Write Plea for Chain of Thought Monitoring to Be Preserved, Gizmodo, 8 October 2026.
- OpenAI says Moonshot-linked users tried to extract its AI reasoning, The Next Web, 30 September 2026.
- Sharing AI progress in mathematics, OpenAI, 6 October 2026.
- AHM statement on OpenAI's October 6 release, Terence Tao, 7 October 2026.
Reader signal
Was this useful?
Work with KREO Studio
AI engineering, data science and design architecture, from Plymouth to the wider UK.
Next Article
