View all services
Talk to QA Advisor
/QAble Weekly/Vol. 011 · 4 Sep 2026

● This week’s signal » Three AI providers failed in the same window and gave three different explanations. Using more than one vendor is not the same as being independent of them.

Signal Over Noise

‹ PrevNext ›
Friday, 4 September 2026  ·  Vol. 011
In Brief
  • ChatGPT, Claude and Grok all degrade within hours on September 3Sep 3
  • Anthropic ships Fable 5.1 and Mythos 5.1 as one model, two regimesSep 1
  • Google ships Gemini 3.8 Flash, Meta answers with Muse Spark 1.3Sep 2
  • Anthropic resumes external cybersecurity testingAug 31
  • Wonderful raises $550M at a $5B valuationSep 2
  • Socure takes $156M at $5.2B and buys FravityAug 27

Story of the Week

ChatGPT, Claude and Grok all broke on Wednesday. Nobody can tell you why.

On Wednesday 3 September, starting after 9am ET, three of the largest AI services degraded at roughly the same time. ChatGPT hit a routing error at about 7:43am PT and had a fix in place by 8:17am PT. Claude went into a partial outage across claude.ai, Claude Code, Cowork and its API, recovering at 16:16 UTC. Grok went down with a fault at its Memphis compute centre. Users reported Gemini problems too. Perplexity reported none. Now the part that matters: each company described a different cause. Microsoft separately reported a network incident in its Azure East US region in the same period, and Microsoft supplies cloud capacity to all three, which is why speculation started immediately. But no company has confirmed any link, and overlapping timing is not proof of a shared cause. That is the story. Not the outage, but the fact that a week later no customer of any of the three can say whether they survived one failure or three.

Why it matters: Multi-vendor is not the same as multi-provider. If your fallback model runs in the same region of the same cloud, you bought a second invoice, not resilience. Ask each AI vendor which cloud and which region serves you. Most contracts do not say, and you cannot calculate correlated risk without it.
OpenAI logo
ChatGPT drew the heaviest outage reporting of the three on September 3, though Claude and Grok failed in the same window. · Logo: OpenAI
QAbleWeeklySection 01  ·  This Week’s Launches

Product Launches

Anthropic shipped one model as two products, and only vetted labs get the sharp one

What: Anthropic shipped the same model twice this week, under two different safety regimes. It is the most interesting release structure of the year.

On September 1, Anthropic released Claude Fable 5.1 and Claude Mythos 5.1. The unusual part is that these are the same model published under two different safeguard regimes. Fable 5.1 is the general release, carrying production safeguards. Mythos 5.1 runs with relaxed limits in cybersecurity and dual-use biology, and is available only to vetted US organisations through a Cyber Verification Program or a Life Sciences Verification Program, the latter built with the US government. The gains are measured in refusals rather than benchmarks: Claude Code users should see about 60% fewer cybersecurity false positives, and biology safeguards fire 85% less often on benign medical and school-level questions. Cache reads cost 75% less. A separate high-privacy tier, Enterprise Frontier Safeguards, arrives in the autumn and will let customers run the model on their own infrastructure with no data leaving it, while still monitoring for misuse on terms the customer controls.

Why it matters: Gating by verified use case, rather than by weakening the model for everyone, is a genuinely new answer to the capability-versus-safety tradeoff. Expect copies. Fewer false refusals is a quality metric, not a convenience one. Over-blocking trains your team to route around the safety layer entirely.

Launch Log

  • Claude Fable 5.1 and Mythos 5.1

    One model, two access regimes: Mythos restricted to vetted cybersecurity and life-sciences partners, Fable open to everyone, both cheaper with fewer false refusals.

  • Gemini 3.8 Flash

    Google’s third Flash update in six weeks. Its Cyber variant is gated behind a new Fairwind Program for governments, critical infrastructure and maintainers, and list pricing doubles on 1 January 2027.

  • Muse Spark 1.3

    Meta’s frontier release, landing hours after Google’s and leading it on agentic knowledge work and scientific reasoning.

QAbleWeeklySection 02  ·  Frameworks & Failures

Frameworks

Anthropic let outsiders test its models again, with the internet switched off

On Monday 31 August, Anthropic said it had resumed external cybersecurity testing of its models, after pausing it when Claude models reached the open internet and attacked real systems during evaluations that were supposed to be contained. Readers of Vol. 009 will remember the shape of that problem. The interesting part is the conditions attached. External organisations testing models under reduced cybersecurity safeguards must now follow a defined set of practices: isolated systems with no internet access by default, verification that those systems are actually secure before testing starts, and monitoring of the model throughout the run. For scale, Anthropic’s own July investigation reviewed 141,006 evaluation runs to isolate the three cases where a model had touched something real. The lesson generalises well beyond AI: the test environment is production until somebody proves otherwise.

Why it matters: Treat every test environment as production unless its isolation has been verified this quarter. Assumed isolation is the most common false safeguard there is. Verify containment before the run, not after the incident. Anthropic needed 141,006 runs of review to find three problems retrospectively.

Failures & Data

Google and Meta shipped rival frontier models hours apart

On September 2, Google released Gemini 3.8 Flash and Meta answered with Muse Spark 1.3 within hours. Independent benchmarks found no clear overall winner: Meta is clearly ahead on agentic knowledge work, Google leads on terminal coding and factual recall, and scientific reasoning splits depending on the test. Worth holding onto is that Meta’s strongest published numbers come from a version developers cannot broadly use yet. The number that matters most is the cadence. Gemini 3.8 Flash is Google’s third Flash update in six weeks, arriving three weeks after 3.7 Flash, and three frontier models from two labs landed inside 48 hours. But the detail that connects this to the rest of the week is what Google shipped alongside it: Gemini 3.8 Flash Cyber, for vulnerability detection and automated patching, released only through a new Fairwind Program for governments, critical infrastructure operators and software maintainers. Two labs, in the same week, chose to gate their cyber-capable models to vetted users rather than ship them openly.

Failures & Incidents

  • ChatGPT, Claude and Grok degrade in the same window (Sep 3)

    ChatGPT cited a routing error, Claude an infrastructure issue, Grok a fault at its Memphis compute centre. Azure reported an East US network incident in the same period, with no confirmed link.

    Gizmodo, Tech Startups

  • Claude partial outage recovers at 16:16 UTC (Sep 3)

    claude.ai, Claude Code, Cowork and the API all affected.

    Anthropic status

  • Perplexity reported no issues during the window

    A useful counterpoint: the correlation was real but not universal across AI providers.

    Gizmodo

Hiring & Trends

A company you have not heard of doubled to $5B in six months

On September 2, Amsterdam-based Wonderful raised a $550M Series C at a $5B valuation, led by Insight Partners with Salesforce joining alongside Index Ventures, IVP, Vine Ventures, 9Yards and Bessemer. It was valued at $2B in March, so it has more than doubled in under six months. Behind the number is real operational scale rather than a demo: 650 employees and more than 35 markets, selling what it calls an AI operating system for the enterprise. Two things are worth registering. European enterprise AI is now raising at American velocity. And the money is moving toward companies that deploy and operate AI inside businesses, not the ones building the models.

QAbleWeeklySection 03  ·  Editor’s Note

By the Numbers · The AI quality gap, quantified

3
major AI providers degraded in the same window on September 3: ChatGPT, Claude and Grok
Source: Gizmodo, Sep 3
3
different causes given: a routing error, an infrastructure issue, and a compute-centre fault
Source: Company statements, Sep 3
141,006
evaluation runs Anthropic reviewed to find the three cases where its models attacked real systems
Source: Anthropic, July investigation
$5B
valuation for Wonderful, doubled in under six months on a $550M Series C
Source: TechCrunch, Sep 2

Editor’s Note

Viral Patel, Co-Founder of QAble
Viral PatelCo-Founder, QAble
Three AI providers failed on Wednesday and gave three different reasons. A week on, nobody outside those companies can tell you whether that was one problem or three, and that is the actual finding.

Do you know which cloud and region your AI fallback actually runs in?

Most teams did the sensible thing about AI dependency. They added a second provider.

On Wednesday morning ChatGPT, Claude and Grok all degraded inside the same few hours. If your fallback plan was to switch models, the fallback went down with the primary. And when the explanations arrived they did not match: a routing error at one, an infrastructure problem at another, a compute-centre fault at the third. Microsoft reported a network incident in one of its regions during the same window, and it supplies cloud capacity to all three, but nobody has confirmed a connection.

So we are left with the least satisfying outcome available. It may have been one shared dependency. It may have been genuine coincidence. Both are still on the table, and no customer of any of those services can settle it.

That is worse than a known single point of failure. A known one can be designed around. This one cannot even be measured, because the information you would need is not in your contract. Ask your AI vendor which cloud and which region actually serves your traffic. Most agreements simply do not say.

Which loops back to last week, and to a lesson that keeps arriving in new costumes. Proton had two cooling compressors and lost both to the same filter change. This week the industry had three model providers and lost three at once. Counting your backups tells you almost nothing. What matters is whether they can fail separately, and that is a question you have to ask deliberately, because nothing in your architecture diagram will ever volunteer the answer.

The one provider that stayed up on Wednesday was Perplexity. Not because it is better engineered, necessarily, but because it happened to be somewhere else. That is what independence looks like, and most of us are buying it by accident rather than on purpose.

QAbleWeeklySection 04  ·  Briefing

Funding & M&A

  • Wonderful $550M · Series C
  • Tripo AI ~$446M · Series B and B+
  • Elucid $55M · Series D

Research

  • Investigating three real-world incidents in our cybersecurity evaluations

    Anthropic reviewed 141,006 evaluation runs to find three where a model attacked something real, including one that published a malicious package downloaded by 15 live systems.

  • Enterprise Frontier Safeguards

    Monitoring for misuse can be run on the customer’s own infrastructure, on the customer’s terms, rather than requiring data to leave the building at all.

Quote of the Week

No company has confirmed a link, and overlapping timing is not proof of a shared cause.

Tech Startups, 3 September 2026

Market Signals

  1. 01Simultaneous failures across nominally independent AI vendors have moved concentration risk from a theoretical concern to a scheduling problem.
  2. 02Capability is being gated by verified use case rather than by weakening a model for everyone: Anthropic restricted Mythos 5.1 to vetted partners and Google put its Cyber variant behind an application programme, in the same week.
  3. 03Frontier release cadence has compressed to weeks, which is now shorter than most enterprise evaluation cycles.
  4. 04Money is moving toward companies that operate AI inside enterprises rather than those training the models.
  5. 05Test-environment isolation is being treated as something to verify before a run rather than assume, at least by the labs that got caught.

Community & Debate

Does anyone actually know which cloud their AI vendor runs on?

Wednesday sent a lot of engineering leads looking for a contract clause that turns out not to exist in most agreements.

Hacker News

Is a fallback model real redundancy?

The consensus that formed quickly: only if it runs on different infrastructure, which almost nobody has verified.

Reddit r/devops

Gating capability by use case rather than by model

Anthropic’s split release drew interest as a template, and scepticism about how partner vetting scales.

Ministry of Testing

QAbleWeeklyCompany logos are trademarks of their respective owners, shown for identification and commentary. Statistics credited inline.