View all services
Talk to QA Advisor
/QAble Weekly/Vol. 014 · 25 Sep 2026

● This week’s signal » The people building frontier AI spent this week asking governments to constrain them, while their own government said no.

Signal Over Noise

‹ PrevNext ›
Friday, 25 September 2026  ·  Vol. 014
In Brief
  • Altman and Amodei brief the UN Security Council on AI riskSep 23
  • Trump calls international AI oversight a globalist schemeSep 22
  • OpenAI cuts API prices by half or more with GPT-6 Sol and LunaSep 22
  • xAI ships Grok 4.7 at the same price as 4.6Sep 21
  • Snorkel AI raises $350M at a $3.5B valuationSep 22
  • Tekever raises $580M at a $6.4B valuationSep 23

Story of the Week

Two AI chief executives asked the UN Security Council to rein them in

On Tuesday 22 September, at the UN General Assembly, President Trump rejected international AI oversight, calling it a “globalist scheme” to control the technology, and adding: “I’m not going to stifle growth of something that will be bigger than the industrial revolution.” He compared those warning about AI risk to people warning about a climate “hoax”. The following day, Wednesday 23 September, France, which holds the Security Council presidency this month, convened a session on the future of AI and international security. Speaking to it were Sam Altman of OpenAI, Dario Amodei of Anthropic (by video), Clement Delangue of Hugging Face, and Turing Award winner Yoshua Bengio. Chinese developers DeepSeek and Moonshot AI were invited to make statements. Amodei told the council that “AI could be a risk to humanity as a whole” and that he believes “this is the most important global security issue facing the world today.” Bengio described risks from systems beyond human control, warning that the danger “does not respect the borders we defend.” Read the two days together and the shape is unusual: an industry lobbying for constraint, and a government declining to supply it.

Why it matters: When vendors ask to be regulated, read the specific ask rather than the speech. The detail is where the commercial interest lives. US and Chinese labs sitting in the same session matters more than anything said in it. Evaluation standards only work if both adopt them.
United Nations logo
France holds the Security Council presidency for September and convened the session on the future of AI and international security. · Logo: United Nations
QAbleWeeklySection 01  ·  This Week’s Launches

Product Launches

Three frontier models in 48 hours, and the news was the price

What: Three frontier models shipped in 48 hours, and the headline was not capability. It was the price.

On 21 September, xAI released Grok 4.7, its frontier model for coding and agentic work, at $2 and $6 per million input and output tokens. That is the same price as Grok 4.6, so the upgrade is free to existing users. The next day OpenAI released GPT-6 Sol and GPT-6 Luna, following Astra earlier in the month, and cut API prices by half or more. Luna, aimed at high-volume work, costs $0.10 per million input tokens and $0.50 per million output, with a 1,050,000-token context window. Xiaomi shipped MiMo V2.6 Pro the same day, a 1.02-trillion-parameter model with 42 billion active parameters, native text, image, video and audio, released with open weights under an MIT licence. Capability gains across the three were incremental and nobody pretended otherwise. The movement was in cost and in licensing, which is a different competition from the one the benchmarks describe.

Why it matters: Pin your model versions. At this release cadence an unpinned default will change underneath a system you have already tested. Re-run your build-versus-buy maths each quarter. Assumptions made six months ago were priced against a market that no longer exists.

Launch Log

  • GPT-6 Sol and GPT-6 Luna

    OpenAI cut API prices by half or more, with Luna at $0.10 per million input tokens and a 1,050,000-token context window.

  • Grok 4.7

    xAI’s frontier coding and agentic model, priced identically to Grok 4.6 at $2 and $6 per million tokens, with a 500k context window.

  • MiMo V2.6 Pro

    Xiaomi’s 1.02-trillion-parameter multimodal model with 42B active parameters, shipped with open weights under an MIT licence.

QAbleWeeklySection 02  ·  Frameworks & Failures

Frameworks

The actual proposal was not regulation. It was access.

Underneath the speeches at the Security Council sat a narrower and more practical idea. Ahead of the session, Amodei published an essay arguing that AI companies should moderate the pace of development and grant independent evaluators complete access to advanced research facilities. Not disclosure after the fact. Not a summary of results. Access, while the work is happening. That is a familiar argument in a different costume. Anyone who has run a quality function knows the difference between being shown a report and being allowed to test the system yourself, and knows which one catches things. The United Nations already has two vehicles that could carry it: the Global Digital Compact and a UN-backed Independent International Scientific Panel on AI. Neither has the access rights Amodei described. The gap between a panel that reviews what it is given and an evaluator who can run its own tests is the entire distance between governance and assurance.

Why it matters: Independent access beats self-reported results, at every scale. The argument being made at the Security Council is the one you should be making about your vendors. Ask any AI supplier what an independent party may test, not what they will report. The answers differ more than you would expect.

Failures & Data

Cheap inference just made proper testing affordable

The price cuts have a second-order effect that deserves more attention than the launches did. Evaluating an AI feature properly means running it many times, because a single answer tells you almost nothing about a system that is allowed to vary. That has been the practical reason teams test AI features thinly: running an evaluation suite a hundred times across a hundred cases was a real line in the budget. At $0.10 per million input tokens, it largely is not. The same suite that was uncomfortable to run weekly can now run on every pull request. Nothing about this is automatic. Budgets do not reallocate themselves, and a cheaper model does not write your test cases. But the most common objection to evaluating AI output seriously, that it costs too much to do repeatedly, quietly stopped being true this week.

Failures & Incidents

  • US rejects international AI oversight at the General Assembly (Sep 22)

    President Trump called such proposals a globalist scheme, saying he would not stifle growth of something bigger than the industrial revolution.

    Fortune

  • OpenAI and Anthropic brief the UN Security Council (Sep 23)

    Altman, Amodei, Delangue and Bengio spoke. DeepSeek and Moonshot AI were invited to make statements. France chaired.

    UN News

  • Amodei calls AI the most important global security issue facing the world

    He told the council that AI could be a risk to humanity as a whole.

    CNN

Hiring & Trends

A company that sells test data is now worth $3.5 billion

On 22 September, Snorkel AI raised $350M at a $3.5B valuation, co-led by Insight Partners and S32. That is roughly triple the $1.3B it was worth seventeen months ago. What changed is what it sells. Snorkel began as tooling to automate data labelling. It now sells finished training and evaluation datasets, generated by its own models and software alongside subject-matter experts rather than resold human labour, along with the reinforcement-learning environments used to develop models. That business has grown eighteen-fold in twelve months to a $375M annualised run rate. Read plainly: the frontier labs have decided it is cheaper to buy expert-built evaluation data than to produce it themselves. Testing data has become a product with a market price, which is a strange and telling development for anyone who has spent years arguing that test data deserves a budget.

QAbleWeeklySection 03  ·  Editor’s Note

By the Numbers · The AI quality gap, quantified

$0.10
per million input tokens for GPT-6 Luna, as OpenAI cut API prices by half or more
Source: VentureBeat, Sep 22
$3.5B
valuation for Snorkel AI, which sells finished training and evaluation datasets
Source: TechCrunch, Sep 22
18x
growth in Snorkel’s data-as-a-service business in twelve months, to a $375M run rate
Source: PR Newswire, Sep 22
3
frontier models released inside 48 hours: GPT-6 Sol, GPT-6 Luna and Grok 4.7
Source: Company announcements, Sep 21 to 22

Editor’s Note

Viral Patel, Co-Founder of QAble
Viral PatelCo-Founder, QAble
Strip away the setting and the speeches, and the ask at the Security Council was something every quality team has asked for: not a report about the system, but permission to test it themselves.

What can you independently test in your AI vendor, rather than simply be told about?

Two things happened this week that look unrelated and are not.

On Tuesday the American President told the General Assembly that international AI oversight is a globalist scheme. On Wednesday the chief executives of the two largest American AI companies sat in front of the Security Council and asked for international AI standards. Chinese labs were invited to the same session. That is an odd alignment, and worth watching regardless of what you think of any party in it.

But the specific proposal underneath was narrower than the headlines suggested. Ahead of the session, Amodei argued that independent evaluators should be granted complete access to advanced research facilities. Not a report after the fact. Not a summary of results that have already been filtered. Access, while the work is happening.

Anyone who has run a quality function has had this argument, in a smaller room, with lower stakes and the same structure. There is an enormous difference between being shown a vendor’s test results and being allowed to test the vendor’s system yourself. One of those tells you what they found. The other tells you what is there.

And the second thing. Snorkel AI is now worth three and a half billion dollars selling expert-built evaluation data, a business that grew eighteen-fold in a year. The frontier labs worked out it was cheaper to buy rigorous test data than to make it. Meanwhile most engineering organisations still treat their own test data as something to be improvised at the end of a sprint.

The argument being made at the Security Council this week, that reports are not verification, is the same argument worth making about your own suppliers. It travels down from geopolitics to procurement without losing any of its force.

QAbleWeeklySection 04  ·  Briefing

Funding & M&A

  • Tekever $580M · Series D
  • Snorkel AI $350M · Series E

Research

  • Global Digital Compact

    The UN framework already agreed for governing digital technology, and one of the few existing vehicles that could carry binding AI evaluation commitments.

  • Independent International Scientific Panel on AI

    A UN-backed body intended to assess AI capability and risk, currently without the access rights that would let it test systems rather than review reports.

Quote of the Week

I believe that this is the most important global security issue facing the world today.

Dario Amodei, to the UN Security Council, 23 September

Market Signals

  1. 01The frontier labs are lobbying for constraint while their own government publicly refuses to impose it, which is an unusual alignment to watch.
  2. 02US and Chinese developers appearing in the same UN session is the first credible signal that evaluation standards might be negotiated rather than imposed.
  3. 03Model competition has moved from capability to price, with API costs halving in a single announcement.
  4. 04Cheap inference removes the standard objection to evaluating AI output repeatedly, which was cost rather than principle.
  5. 05Expert-built evaluation data has become a product with a market price, and the frontier labs are the customers.

Community & Debate

Does asking to be regulated mean anything?

Split between reading it as genuine concern and as an attempt to set rules that favour incumbents with compliance budgets.

Hacker News

Nobody priced for inference getting this cheap

Teams reworking cost models that were built when running an evaluation suite repeatedly was the expensive part.

Reddit r/devops

Buying your test data

Snorkel’s valuation prompted an uncomfortable question in QA circles about why expert-built test data is fundable externally but not internally.

Ministry of Testing

QAbleWeeklyCompany logos are trademarks of their respective owners, shown for identification and commentary. Statistics credited inline.