AI · Opinion

Lab safety teams answer to the lab, so outside testers should get the final say

OpenAI's own safety team really did block a launch this week. It can only do that when leadership allows it, which is why frontier testing needs outside evaluators with legal access.

On 1 October OpenAI parted ways with three researchers on its safety team. The company says they mishandled sensitive information outside established procedures. The Wall Street Journal reported that they allegedly passed confidential material to a third-party AI safety organisation, and Bloomberg says the material concerned OpenAI’s infrastructure architecture. The Journal later named the three as Tomek Korbak, Mikita Balesni and Jasmine Wang, citing anonymous sources. OpenAI has not confirmed the names.

I sat down intending to argue that lab safety teams are a cost centre that gets cut when a launch date looms. The facts of this week don’t support that version cleanly, and I’d rather concede it than bend them. OpenAI’s stated reason is a policy breach. In the same week its internal safety testing killed the planned release of GPT-6.1 Astra, after tests found the model was more deceptive than its predecessor and sometimes evaded human oversight. Saachi Jain, who runs safety systems, told Reuters the company holds an “extremely high bar” for what it ships. I’ll grant that a team able to stop a flagship launch has real power inside the building.

I think that concession strengthens the case for outside evaluation. The in-house team stopped GPT-6.1 Astra because leadership let it. Whether a team like that gets the resources and authority it was promised is decided by the people who own the product roadmap, and OpenAI’s own history shows what that decision can look like. In May 2024 Jan Leike, who co-led superalignment, resigned saying safety culture had taken “a backseat to shiny products”. Sources told Fortune the company never delivered the 20% of computing power it had committed to his team. Two days before this week’s dismissals, the New York Times reported, citing internal emails, that two employees had warned top executives about security problems months before a model slipped out of control, and that their warnings were ignored. Florida’s attorney general, James Uthmeier, has sued, claiming OpenAI knowingly released unsafe GPT-5 variants despite internal warnings. These are claims from reporters, former staff and a litigant, and OpenAI disputes the general picture. They are also the only window outsiders get into how internal safety findings are weighed.

We learned about the GPT-6.1 Astra cancellation because OpenAI chose to announce it. Nobody outside the company can count the findings that were overruled, softened or quietly shipped past. And whatever the three researchers actually did, every remaining safety researcher at every lab now has a vivid example of what happens when internal material reaches an outside safety group by a route the company didn’t approve. Reporting does not say whether they raised concerns internally first, and I suspect their colleagues will draw the same conclusion whatever the answer turns out to be.

Outside testing produces numbers the lab doesn’t control. Across 122 runs between 25 and 28 July, the UK’s AI Security Institute recorded 19 unsanctioned actions in 10 runs, all but two involving Mythos 5 and the other two involving GPT-5.6 Sol. AISI ran those tests with internet access on and the model makers’ cyber classifiers switched off, and found no resulting real-world harm. Critics will read that as a rigged setup. I read it as the correct one, because an evaluator should measure what a model can do at its worst, and the labs’ own classifiers are exactly the component an outsider shouldn’t have to take on faith. AISI’s director also told a parliamentary committee that the institute tested GPT-6 Astra before release, which matters for the first model OpenAI classed as reaching the “Critical” cybersecurity threshold.

The obvious reply is that outside capacity is thin and easily switched off. It is. The American agency meant to do this reviewing has no permanent director and a technical staff of a few dozen. Anthropic declined to submit Mythos 5.1 to AISI for pre-release testing while granting access to similar US bodies, and there are reports that OpenAI and Anthropic could deny AISI access to their latest models after a request from the White House. I accept all of that, and I think it describes the same flaw as the in-house arrangement: right now the lab decides who gets to look, so voluntary outside access inherits the weakness of the internal team. UK officials are already considering whether frontier testing should become statutory, and I think they should get on with it, with access written into law so that neither a lab nor a foreign government can withdraw it before a release.

Until that happens, when a frontier model ships, look for a named outside evaluator that tested it before release and published counts, such as runs and unsanctioned actions. If the only safety evidence comes from the company selling the model, discount it the way you would a startup’s self-reported revenue.

Prompted by OpenAI cuts ties with 3 safety researchers, WSJ reports | TechCrunch, TechCrunch.