View all services
Talk to QA Advisor
/QAble Weekly/Vol. 013 · 18 Sep 2026

● This week’s signal » The agents scanned, wrote the exploit, tested it, and scaled it. A person set the goal and then mostly watched.

Signal Over Noise

‹ PrevNext ›
Friday, 18 September 2026  ·  Vol. 013
In Brief
  • AI agent swarm breaches 395 organisations through PaperCut serversSep 9
  • TypeSafe AI exits stealth with $40M and a model that returns types, not textSep 15
  • Claude Code weekly limits drop 17% in real termsSep 14
  • Profound raises $180M at a $1.8B valuationSep 15
  • Arcee AI raises $150M at a $1B pre-money valuationSep 16
  • Anew Labs raises $290M at a $1.5B valuationSep 16

Story of the Week

One person, a swarm of AI agents, and 395 organisations breached

Security firm GreyNoise published its findings on September 9 under the title “Agents Gone Wild”, and four research teams spent the following week corroborating it. Preparation began on 31 August and the campaign launched the next day, when a single attacker pointed hundreds of AI agents at PaperCut print-management servers. The agents did the work: they found targets through a public scanning service, wrote the exploits themselves, tested them in a lab, then deployed at scale. Three hours and 55 minutes from an empty workspace to the first remote code execution. Then eleven organisations inside 26 seconds of launch. The final count was 440 compromised instances across 395 named organisations in 48 countries. Two details deserve to stay with you. 204 of the 395 victims were schools or universities, and at one US high school the agents went from first access to full domain administrator in seven minutes. Of the 440 instances, 280 gave up credentials and 147 also surrendered operating system and domain secrets, along with thousands of authentication tokens for Google, Microsoft, Amazon, Anthropic and Cursor.

Why it matters: The economics changed, not the technique. PaperCut flaws were known and patchable. What is new is that one person can now run the exploitation of hundreds of targets in parallel. Your patch window is now competing with an attacker who does not sleep, does not tire, and scales by spending money rather than time.
DeepSeek logo
The campaign ran a DeepSeek model inside OpenAI’s Codex harness. Both were used as tools by the attacker, not compromised. · Logo: DeepSeek
QAbleWeeklySection 01  ·  This Week’s Launches

Product Launches

He helped invent the training behind ChatGPT. Now he says it made AI a people-pleaser.

What: The most interesting launch this week came from a man who helped invent ChatGPT, and it does not talk.

On September 15, TypeSafe AI came out of stealth with $40M led by DCVC. Its co-founder and CEO is Diogo Almeida, formerly of OpenAI and Google Brain, a co-author of the 2022 InstructGPT paper that introduced RLHF, the training method that made ChatGPT possible. His argument now is the interesting part: he says training models on human approval produced people-pleasing assistants, and that approval and reliability are not the same target. So he is building the opposite. Their model, Jev, does not return sentences. It returns typed decisions with probabilities attached, meant to be consumed by other software rather than read by a person, trained toward calibrated decision-making instead of conversation. It is aimed at the unglamorous judgement calls buried inside every application: classify this, route that, score this, extract that. Because it answers all at once rather than word by word, it responds in 70 to 500 milliseconds and costs a fraction of a conversational model. The demonstration is deliberately absurd and rather effective: given structured game state, Jev plays Doom, deciding its next move in 0.114 seconds.

Why it matters: A typed answer can be validated by a schema. A paragraph has to be parsed and hoped over. That is a testability difference, not a style preference. If you are using a large conversational model for classification or routing, you are paying for prose you then throw away.

Launch Log

  • TypeSafe AI

    Jev, a model returning typed probabilistic decisions for software rather than prose for people, responding in 70 to 500 milliseconds.

  • Arcee AI

    Frontier open-weight models, including a 400-billion-parameter sparse mixture-of-experts model with 13 billion active parameters per token.

QAbleWeeklySection 02  ·  Frameworks & Failures

Frameworks

Your access rules assume every account belongs to a person

The most useful analysis of this week’s campaign was not about print servers at all. It was about identity. Most enterprise access management was designed around humans, and humans have a lifecycle: they are hired, they change teams, they take leave, they leave. Reviews and offboarding hang off those events. As one analysis put it, an agent’s API key does not take parental leave or change roles. So machine credentials accumulate quietly, often with standing administrator rights and no expiry, and nothing in the calendar ever prompts anyone to look at them. That is what the agents harvested and reused. Three things follow. Give every non-human credential an owner and an expiry date, because unowned keys are never revoked. Stop granting standing domain administrator rights to service accounts, and issue them for the duration of a task instead. And audit machine identities on a schedule, since they will never generate the HR event that triggers a review.

Why it matters: Run one query this month: which service accounts hold domain admin, who owns them, and when do they expire? The uncomfortable answers are usually in that list. Treat an agent credential as more dangerous than a staff credential. It is used more often, watched less, and never goes on holiday.

Failures & Data

Anthropic announced a 25% increase that was really a 17% cut

On September 14, Anthropic raised Claude Code weekly usage limits by a permanent 25% across Pro, Max, Team and Enterprise seats. True, and incomplete. A temporary 50% boost that had been running since May expired on September 13, the day before. So a developer who had 150 units a week moved to 125. Anthropic itself put it plainly once pressed: “Compared to today, this works out to a 17% reduction in weekly limits on Claude Code.” The announcement led with the increase, developers did the subtraction publicly, and the company deleted and reposted a clarification that stated the cut plainly. The lesson is not about pricing. It is that your capacity planning now depends on vendor limits that can move under you, and that the announcement and the arithmetic may not agree.

Failures & Incidents

  • AI agent swarm breaches 395 organisations via PaperCut (reported Sep 9)

    440 compromised instances across 395 named organisations in 48 countries, with eleven falling in the first 26 seconds after launch. Agents wrote and tested the exploits themselves.

    GreyNoise, “Agents Gone Wild”

  • 204 of the victims were schools and universities

    At one US high school the agents reached full domain administrator seven minutes after first access.

    VentureBeat

  • Claude Code weekly limits fall 17% in practice (Sep 14)

    A permanent 25% rise replaced a temporary 50% boost that expired the previous day. Anthropic confirmed the net reduction after developers did the arithmetic.

    BleepingComputer

Hiring & Trends

Brands are now paying to find out what AI says about them

On September 15, Profound raised $180M at a $1.8B valuation, co-led by Sequoia and Kleiner Perkins, seven months after its last round. The product exists because of a shift most people have made without noticing: a growing share of buyers now start with ChatGPT, Gemini or Perplexity instead of a search box. Profound tracks how often AI systems mention a brand, in what context, and how they describe it. More than 1,000 enterprises pay for this, including Walmart, Comcast, Estée Lauder and Royal Bank of Canada, alongside Zoom, MongoDB, Figma and Cursor. Elsewhere the same week, Arcee AI raised $150M at a $1B pre-money valuation to keep building open-weight models, and Anew Labs took $290M for AI drug discovery.

QAbleWeeklySection 03  ·  Editor’s Note

By the Numbers · The AI quality gap, quantified

26 sec
to breach the first eleven organisations once the agent swarm was released
Source: GreyNoise, “Agents Gone Wild”, Sep 9
204
of the 395 breached organisations were schools or universities
Source: VentureBeat, Sep 16
7 min
from first access to full domain administrator at one US high school
Source: VentureBeat, Sep 16
17%
real cut to Claude Code weekly limits, announced as a 25% increase
Source: Bleeping Computer, Sep 14

Editor’s Note

Viral Patel, Co-Founder of QAble
Viral PatelCo-Founder, QAble
A person chose the target and set the goal. Everything after that, finding the servers, writing the exploit, testing it, running it against hundreds of organisations, was done by software while they watched.

Which of your service accounts hold admin rights, and who owns them?

The PaperCut campaign is not interesting because of the vulnerability. That was known and there was a patch.

It is interesting because of the arithmetic. Under four hours from an empty workspace to a working attack. Eleven organisations in the first 26 seconds. Three hundred and ninety-five organisations in total, across forty-eight countries, from one operator. Attacking a hundred targets used to cost roughly a hundred times what attacking one cost. That relationship has quietly broken.

And the victims tell you who pays for it first. Two hundred and four of them were schools and universities. Not because schools are interesting targets, but because they run the same software as everyone else with a fraction of the staff to patch it. At one high school, the gap between the first intrusion and complete control of the network was seven minutes.

The other half of the story is what the agents took. Not just data, but credentials. Of 440 compromised instances, 280 gave up credentials, and the haul included thousands of tokens for Google, Microsoft, Amazon, Anthropic and Cursor. Which raises the question the best analysis of the week actually asked. Who owns those accounts?

Every access review in your company is triggered by something happening to a human being. Someone joins, someone moves team, someone leaves. Machine credentials never do any of those things. They are created during an incident or a migration, given generous permissions so the work can proceed, and then they simply persist. Nothing in the calendar ever asks about them again.

You cannot outrun this with faster patching alone. But you can make the stolen key worthless, and that is a smaller, duller, entirely achievable piece of work that almost nobody has scheduled.

QAbleWeeklySection 04  ·  Briefing

Funding & M&A

  • Anew Labs $290M · First external round
  • Profound $180M · Series D
  • TypeSafe AI $40M · Seed

Research

  • Agents Gone Wild

    GreyNoise’s account of the PaperCut campaign, notable less for the exploit than for the timings: under four hours from empty workspace to working attack, then eleven organisations in 26 seconds.

  • Training language models to follow instructions with human feedback

    The 2022 InstructGPT paper that introduced RLHF, co-authored by the founder now arguing it optimised models for approval rather than reliability.

Quote of the Week

An agent’s API key does not take parental leave or change roles.

VentureBeat, on why machine credentials escape review

Market Signals

  1. 01Automated exploitation has moved from proof of concept to industrial scale, with one operator breaching hundreds of organisations in a single campaign.
  2. 02Schools and universities are now the soft underbelly of automated attacks, making up more than half the victims in this week’s campaign.
  3. 03A counter-movement to conversational AI is funded and shipping: typed, structured output designed for software rather than for reading.
  4. 04Vendor usage limits are becoming a capacity-planning variable, and the headline number may not match the arithmetic.
  5. 05Machine identity is emerging as the weak point that human-shaped access policies were never designed to cover.

Community & Debate

Who owns the service account?

The credential-harvesting detail sent a lot of teams looking for an owner on machine identities that turned out to have none.

Hacker News

Do we still need a chatbot for classification?

TypeSafe’s typed-output pitch restarted an argument about using conversational models for jobs that were never conversations.

Ministry of Testing

Announce the number people will actually feel

The Claude Code limit change became a case study in communicating a reduction badly.

Reddit r/devops

QAbleWeeklyCompany logos are trademarks of their respective owners, shown for identification and commentary. Statistics credited inline.