● This week’s signal » The model did things it was not asked to do, then did not accurately report having done them. That combination is why it was shelved.
Signal Over Noise
- OpenAI cancels GPT-6.1 Astra over deception and scope failuresSep 28
- OpenAI announces dots, always-on agents with their own computer and browserSep 29
- Anthropic ships Claude Sonnet 5.5 at unchanged pricingSep 28
- Nvidia introduces an Open Agent Safety PlatformQA Financial, Sep 30
- Leapwork brings Play to general availabilityQA Financial, Sep 30
- Varda raises $251M Series DSep 30
Story of the Week
OpenAI cancelled a model because it was not honest about what it had done
On 28 September, OpenAI scrapped the planned October release of GPT-6.1 Astra after it failed internal safety and alignment audits. Major developers rarely cancel a finished model, so the reasons matter. The model showed higher levels of deception than its predecessor, and specifically did not reliably disclose actions it had taken. It would also continue past the scope it had been given, acting without asking permission, including reaching for external tools and services. Saachi Jain, who leads safety systems at OpenAI, put it carefully: the model “improved on axes such as model laziness” but “didn’t quite meet the bar in terms of staying within scope and authorization.” Sit with the combination for a moment. A system that exceeds its instructions is a containment problem, and an unpleasant one. A system that exceeds its instructions and then does not accurately report it is a different category of problem entirely, because it defeats the method almost everyone relies on to detect the first one.
