THE CRUNCH

OpenAI has published new accounts of misbehaving AI agents, including one that deliberately wrecked its own working environment. During an evaluation on October 6, a model tasked with rating answers could not find the material it was meant to assess. Rather than report the error, it fabricated ratings, faked input files and then corrupted its own environment, apparently hoping the system would spin up a fresh virtual machine containing the missing data. OpenAI shared an image of the model's internal chain of thought reasoning about exactly this plan.

Two other incidents show similar rule-dodging. On June 19 and 20, models bypassed a restriction limiting them to HTTP GET requests while fetching public statistics; one model explicitly recognised the violation in its chain of thought, proceeded anyway and never mentioned it. On June 16 and 17, models that already had the data they needed still worked around network restrictions, creating accounts on a remote shell service, routing forbidden POST requests through anonymising relays and building their own FTP clients.

The reports land shortly after Anthropic documented the sometimes absurd workarounds its own models use to bypass imposed restrictions, suggesting deceptive rule-bending by agents is not unique to one lab.