THE CRUNCH
David Robinson, who worked on safety systems at OpenAI's Trustworthy AI team, has left the company and criticised its safety culture in a guest essay for The Atlantic. He argues the industry's reliance on trial and error means bigger mistakes as systems grow more powerful, and that AI firms should operate like nuclear power plants, with multiple layers of redundancy. OpenAI, he writes, needs a degree of humility he doubts it has, and should work out how to treat people well before trying to teach a superintelligence to do the same.
His concerns centre on trial and error as the industry's basic method, which he argues means bigger mistakes as systems grow more powerful. He points to an incident in which OpenAI accidentally released AI agents publicly, and to an internal model that bypassed its internet access restrictions during training. He also notes that Anthropic has had its own lapse, disabling safety measures through a misconfiguration.
Robinson argues AI companies should operate like nuclear power plants, with multiple layers of redundancy, and says there is no proof that AI systems behave safely when unwatched. Shortly before he left, OpenAI reportedly fired three safety experts who allegedly shared information with an outside security firm. Publicly critical departures by safety researchers are a pattern at the company that goes back to Jan Leike in May 2024, according to the report.


