View all services
Talk to QA Advisor
/QAble Weekly/Vol. 016 · 9 Oct 2026

● This week’s signal » Generating results has become cheap and fast. Confirming they are true has not. This week showed the gap at both ends of the scale.

Signal Over Noise

‹ PrevNext ›
Friday, 9 October 2026  ·  Vol. 016
In Brief
  • OpenAI publishes 722 AI-generated mathematical manuscripts in one dayOct 6
  • Mathematicians call the release a demonstration of power, not scholarshipOct 7
  • Anthropic cuts Haiku pricing by about 75% with Claude Haiku 5.5Oct 7
  • Mistral previews Large 4, a one-trillion-parameter modelOct 6
  • 46% of teams have shipped AI code that failed in productionSmartBear, Sep 30
  • Vinci raises $250M at a $1.5B valuationOct 6

Story of the Week

OpenAI published 722 maths papers in a day. Checking them will take years.

On 6 October, OpenAI released a catalogue of 722 mathematical manuscripts, grouped into 372 result families, drawn from evaluations of roughly 4,000 problems. They were produced by an unnamed internal model that OpenAI describes only as significantly more capable than GPT-6 Astra, and which began training on 28 August. The detail that makes this more than a press release is what came before it. In September, after mathematicians circulated an open letter about open problems being used as competitive AI benchmarks, an Advisory Group on Mathematics and Artificial Intelligence was formed, Terence Tao among its members. On 21 September Tao described its task precisely: “advising OpenAI on how to coordinate the release of a large number of significant results in mathematics that they report have been produced by their internal model.” Coordinating a release. Not establishing whether the results are correct. Two weeks later all 722 arrived at once. Tao has separately argued that solutions to open problems are now “being harvested at large scale in an unsustainable fashion, leaving entire fields of mathematics much less fertile”, and that researchers are withholding promising directions for fear of being scooped by an AI. On 7 October his blog carried a guest statement from the Association for Human Mathematics, which was blunter: “Mathematicians did not ask for this work to be done,” and “Releasing over 700 files at once is not a demonstration of scholarship, but a demonstration of power.”

Why it matters: This is the verification gap in its purest form. Producing candidate answers is now effectively free. Establishing which are correct still runs at human speed. Note who absorbs the cost. The producer gets the announcement, and the reviewing community gets years of unpaid checking it did not request.
OpenAI logo
The 722 manuscripts were produced by an unnamed internal model that OpenAI says is still in training. · Logo: OpenAI
QAbleWeeklySection 01  ·  This Week’s Launches

Product Launches

A 75% price cut with a footnote worth reading

What: Anthropic cut its cheapest model’s price by about 75%. Read the footnote on token consumption before you budget against it.

On 7 October, Anthropic released Claude Haiku 5.5 at $0.10 and $0.50 per million input and output tokens for prompts under 100,000 tokens, roughly 75% below Haiku 4.5, with 2.5 times faster agent turns and a one-million-token window. On a desktop-task benchmark it scores 72.4% against 15.7% for its predecessor. The footnote matters: early reporting notes the model consumes considerably more tokens per task, so a 75% cut in unit price is not a 75% cut in your bill. The day before, Mistral previewed Mistral Large 4, a one-trillion-parameter mixture of experts with 49 billion active, available only through a guarded preview with weights promised later this month. Google shipped Gemini Nano Banana 2.1 and Z.AI shipped GLM 5.3 Fast in the same two days.

Why it matters: Price per token and price per task have decoupled. Benchmark your own workload before claiming a saving. Four frontier releases in forty-eight hours is now routine. Pin versions, or your tested system will change without a deployment.

Launch Log

  • Claude Haiku 5.5

    About 75% cheaper per token than Haiku 4.5 at $0.10 and $0.50 per million, with 2.5x faster agent turns, though it consumes considerably more tokens per task.

  • Mistral Large 4

    A one-trillion-parameter mixture of experts with 49 billion active, previewed behind guardrails with open weights promised later this month.

  • UiPath Test Cloud on Oracle Marketplace

    AI-assisted test design and execution across Oracle Fusion, E-Business Suite, PeopleSoft, Siebel and JD Edwards.

QAbleWeeklySection 02  ·  Frameworks & Failures

Frameworks

When generating is free and checking is not

Put this week’s two biggest numbers side by side. 722 mathematical manuscripts produced in a day, requiring years of expert review. And 46% of teams shipping AI-generated code that failed in production, with 47% unable to explain why. These are the same problem at different scales. Generation has become abundant and cheap. Verification has stayed scarce and expensive, because it still depends on someone who understands the domain actually looking. Three things follow. First, verification capacity is now the constraint, so plan it as a budget line rather than an assumption. Second, make failures attributable, because a defect you cannot trace to a cause cannot teach you anything, and half of organisations are currently in that position. Third, be sceptical of volume as evidence. Seven hundred unverified results is not seven hundred results. A thousand generated tests is not a thousand checks. Until something independent has confirmed them, both are a pile of claims.

Why it matters: Budget verification as a first-class cost. It is the only part of the pipeline that has not got cheaper this year. Treat any large volume of AI output as unverified by default, whether it is proofs, tests or code. The burden sits with whoever has to trust it.

Failures & Data

Shipped it, it broke, still confident

SmartBear’s State of Software Quality and Testing report, published 30 September, contains three numbers that belong together. 46% of teams have shipped AI-generated code that subsequently failed in production. Of those teams, 69% still report substantial or complete confidence in AI-generated code. And 47% of organisations could not explain how AI had contributed to a production bug. Read in order, they describe something more specific than overconfidence. Being unable to explain the cause of a failure means there is no mechanism by which the failure can correct the confidence. The belief is not being tested by the evidence, because the evidence never arrives in a legible form.

Failures & Incidents

  • OpenAI publishes 722 AI-generated mathematical manuscripts (Oct 6)

    372 result families drawn from roughly 4,000 problems, produced by an unnamed internal model still in training. Verification is expected to take the community years.

    OpenAI, via Nature

  • Mathematicians object to the scale and manner of the release (Oct 7)

    An advisory group formed in September had been asked to advise on coordinating the release, not on verifying correctness. All 722 arrived at once.

    Association for Human Mathematics

  • 46% have shipped AI code that failed in production

    Of those teams, 69% remain substantially or completely confident in AI-generated code, and 47% of organisations cannot explain how AI contributed to a bug.

    SmartBear, Sep 30

Hiring & Trends

Development is changing faster than the function that checks it

Applause’s 2026 research puts a number on something most engineering leaders have felt. 44.2% reported significant AI-driven change to development, against 35.1% for quality assurance. The two functions are moving at different speeds, and the gap accumulates. The same work found 26.4% seeing reductions in both defect count and severity after introducing AI, 14.7% seeing increases in both, and 20.3% of organisations with no formal documentation of which AI tools are even acceptable to use. Tooling is chasing the gap. In this week’s vendor roundup, UiPath put Test Cloud on the Oracle Marketplace, Anaconda paired agent swarms with autonomous security testing, Katalon added shared rules for testing agents, and Thunders widened agent access to test evidence.

QAbleWeeklySection 03  ·  Editor’s Note

By the Numbers · The AI quality gap, quantified

722
mathematical manuscripts OpenAI published in one day, across 372 result families
Source: OpenAI release, Oct 6
46%
of teams have shipped AI-generated code that later failed in production
Source: SmartBear, State of Software Quality and Testing, Sep 30
69%
of those same teams still report substantial or complete confidence in AI-generated code
Source: SmartBear, Sep 30
47%
of organisations could not explain how AI had contributed to a production bug
Source: SmartBear, Sep 30

Editor’s Note

Viral Patel, Co-Founder of QAble
Viral PatelCo-Founder, QAble
“Seven hundred and twenty-two mathematical papers were published in a single day this week. Nobody knows yet how many of them are correct, and finding out is expected to take years.”

Could you explain how AI contributed to your last production defect?

The objection from mathematicians was not that the results are wrong. Some of them may well be excellent.

The objection was about who pays for checking. A model produced 722 manuscripts in a day. Confirming whether they hold up requires people who understand the field to read them carefully, and there are not many such people, and they did not ask for the work. One statement this week put it as plainly as possible: this is not a demonstration of scholarship, it is a demonstration of power.

Now bring that down to the scale most of us work at. A survey published the same week found that 46% of teams have shipped AI-generated code that later failed in production. Of those teams, 69% remain substantially or completely confident in AI-generated code. And 47% of organisations cannot explain how AI contributed to a bug they have already had.

That last number is the one that should bother you. Confidence surviving a failure is not stubbornness. It is what happens when the failure never becomes legible enough to teach anyone anything. If you cannot attribute the defect, the experience cannot update the belief.

Both stories are the same shape. Generating output got dramatically cheaper this year. Checking output did not get cheaper at all, because it still requires a person who understands the domain to actually look.

So the useful question is not whether AI makes your team faster. It plainly does. It is whether you have funded the part that has not got faster, and whether you would know if it had fallen behind.

QAbleWeeklySection 04  ·  Briefing

Funding & M&A

  • Vinci $250M · Series B
  • Type One Energy $200M · Series B
  • Valon $150M · Series D

Research

  • A human audit of OpenAI’s AI-generated mathematical proofs

    An early attempt to do at small scale what the community now faces at large: establish independently whether the claimed results are correct, which no group had been tasked with.

  • SmartBear, State of Software Quality and Testing 2026

    The three findings that matter sit together: 46% have shipped failing AI code, 69% of those remain confident, and 47% cannot explain how AI contributed to a bug.

Quote of the Week

“Releasing over 700 files at once is not a demonstration of scholarship, but a demonstration of power.”

Association for Human Mathematics, 7 October

Market Signals

  1. 01Verification has become the scarce resource in AI output, whether the output is a mathematical proof or a pull request.
  2. 02Producing results at volume is being treated as evidence of capability, while confirming them remains someone else’s unpaid problem.
  3. 03Model pricing is falling sharply, but token consumption per task is rising, so unit price is no longer a reliable proxy for cost.
  4. 04Development is changing faster than quality assurance by roughly nine percentage points, and the gap compounds.
  5. 05Nearly half of organisations cannot attribute a production defect to an AI contribution, which means their confidence is not being corrected by experience.

Community & Debate

Who is supposed to check 722 papers?

The objection that landed hardest was not about correctness but about consent, and about who absorbs the reviewing cost.

Hacker News

Is a 75% price cut a 75% saving?

Reports that Haiku 5.5 consumes more tokens per task sent teams benchmarking their own workloads rather than trusting the headline.

Reddit r/devops

Can you attribute your last AI-related defect?

The SmartBear finding that 47% cannot explain how AI contributed to a production bug prompted some uncomfortable checking.

Ministry of Testing

QAbleWeeklyCompany logos are trademarks of their respective owners, shown for identification and commentary. Statistics credited inline.