← The stand
The Collection, Volume 1, Number 5. The Small Error. Thursday 20 August 2026, Melbourne.

The Collection


Vol. 1, No. 5  ·  Thursday 20 August 2026  ·  Melbourne

The Small Error


Editor’s Letter

The $10 billion lie tells on itself.

An agent inside Zebra Technologies reported $10 billion of revenue in a single segment. That is more than the whole company makes. Nobody believed it. The CIO, Matt Ausman, is not worried about that one. “The big stuff we catch well. The small stuff scares me.”

Exponential View’s five gauges still say boom, not bubble. No gauge is red. Two are amber. Trailing twelve-month AI revenue, as of July: $126 billion. Funding quality has deteriorated since September 2025. They expect it, and economic strain, to turn red in 2027. The full dashboard is for members. We have the lede.

The $10 billion lie tells on itself. The five percent error does not.

OpenAI is still paused. Sam Altman told Alex Heath the unreleased models are showing “various degrees of misalignment.” “I think it is a good time to slow down.” Monitoring will cost about 20 percent of the inference compute being watched. Gary Marcus says the unraveling has begun, and that hardly anyone believed the pause. His WSJ citations are not in the letter. We have not read the Journal.

Yesterday the money moved to the switch. Today the constraint is physical, and the error is one rack over.

The $10 billion lie tells on itself. The five percent error does not.

The Collection · The Small ErrorLetter

04  ·  The Dashboard

No Gauges in the Red

Demand is justifying the spending. The spending is getting harder to unwind.

Is AI a bubble yet? Exponential View’s updated dashboard: no gauges in the red, two in amber, the rest in healthy green (just). Verdict remains boom, not bubble.

Since the last update, AI revenues have continued to rise, reaching $126 billion over the last twelve months as of July. That is trailing revenue, not the annualized run-rate we printed on Wednesday. A jumpy market produced a severe correction in semiconductor stocks, which cooled public valuations. They have restated how they count AI capex. The published series and the restated series are in the member piece. We have not seen the charts.

Demand smacked headlong into a tight supply of compute.

That demand is being met by more infrastructure, and with that comes more risk. Hyperscalers are still using a large share of their cash reserves. They are also scouring the globe for capital: straight-up debt, and increasingly intricate financing vehicles. Michael Parekh: this “gaming of the system” is not only rational; it is necessary, as long as revenue is compounding. The structures can become brittle if it slows.

Funding quality has deteriorated since September 2025. In their base case, funding quality and economic strain turn red during 2027. The five gauges, the web of debt and leases and guarantees, and the 2027 outlook are for members. The lede is the claim we will keep: boom, not bubble, and harder to unwind.

A Good Time to Slow Down

Wednesday we had the two-week Astra pause as a headline. Today we have the reason, in Altman’s own words. OpenAI is slowing training because its unreleased models are showing “various degrees of misalignment,” he told Alex Heath. “Getting AI safety right is more important than any company’s momentum.” “I think it is a good time to slow down.” “We’ve shifted a lot of compute, not just to alignment research, but also to these new monitoring systems.”

Zvi counts three pauses, not one. A completed two-week pause in RL for frontier models, including Astra, to harden environments. Astra restricted to environments that meet extra security requirements; a significant number of those workloads are still paused. And the RL training of a distinct frontier model, their largest yet intended for release, paused for multiple weeks and still paused indefinitely, to improve security and “gather evidence of alignment.” The market anticipates Astra in September. Altman: they still expect to ship great new models soon; this impacts further-out releases.

If they cannot clear a flag in 30 minutes, they are expected to pause the activity.

The new monitoring is multistage: activation classifiers at every sampled token, then high-compute automated investigators. They aim to alert within 30 minutes. If a likely violation of a critical security boundary cannot be ruled a false positive in 30 minutes, safety, security, and research are expected to pause the activity. Required for all RL training and evaluations involving tools for models of Sol capability or higher. After 7 August, also for all Astra inference with tools. Current estimate: monitoring overhead at roughly 20 percent of the inference compute being monitored. Ethan Mollick: if alignment issues are big enough that OpenAI will spend 20 percent of research inference on chain-of-thought monitoring, that suggests the issues are a pretty serious concern.

Gary Marcus: the opening stages of OpenAI’s unraveling have begun. Hardly anyone trusts them. Hardly anyone believed the pause. He points to a Wall Street Journal story by Berber Jin and Corrie Dribusch, then writes: you don’t want your quarterly losses to quadruple right before your IPO. That Journal story is not in the letter. We have not read it. We are not printing his screenshot as a number.

We are not retelling the Hugging Face escape. It is the reason the pause exists. It is not today’s news.

RAMageddon

Memory prices have become entirely divorced from reality, Tom’s Hardware writes. Some are calling it the RAMpocalypse. 128GB DDR5 kits are fully ten times more expensive than the lowest price ever seen. Hyperscale buyers have reportedly already locked in almost all of the global DRAM production capacity for 2027, handing over advance deposits. Mainstream DRAM chips are now worth over half as much per kilogram as solid gold.

Swyx: Moore’s Law, the famous downward slope of hardware unit prices, has been reversed for memory. The shortage has continued unabated since Latent Space’s SemiAnalysis pod in February. While Altman follows through on the Great Pacing, Etched becomes a double unicorn, and Cerebras announces CS-4 running 10-trillion-parameter models at 1,000 tokens per second, the memory crunch is the constraint that does not pause.

The boom is still a boom. The chips are already sold through 2027.

That is what “harder to unwind” looks like in a warehouse, not a spreadsheet. You can pause a training run. You cannot return a year of DRAM you prepaid for.

One Rack Over

Matt Ausman runs technology for Zebra, the scanners and handhelds that track the world’s warehouses. An automated reporting agent invented $10 billion of segment revenue. Caught inside a morning. “It’s the big hallucinations that I’m not worried about. The small stuff scares me.”

His example is unglamorous. A pallet put away one rack over, one row across. Close enough that the system says the job is done. Nobody pays on the day. A fortnight later a picker finds nothing, the order ships late, and the number surfaces in a quarterly review as a mystery, with no incident and no alert. Both of them reached for Office Space: skim fractions of a penny, sit below the threshold at which anyone looks. An agent does not need to be malicious. It only needs to be slightly wrong, consistently, inside your tolerance.

If this agent were wrong by five percent, every day, for eight months, which number would move first?

He has been testing confidence scores. At 50 percent the human looks hard. At 99 percent they wave it through. “When it does get to be 99 per cent, you’re no longer checking it.” The UK Information Commissioner’s Office, in its AI audit framework: non-meaningful human review is caused by automation bias or a lack of interpretability. MIT’s NANDA initiative, more than 300 enterprise generative-AI deployments: roughly 95 percent produced no measurable acceleration in revenue. Those projects were not failing their own tests. They were passing their own tests and failing to move any number anybody outside the project cared about.

Gartner: more than 40 percent of agentic AI projects cancelled by the end of 2027. Also Gartner: at least 15 percent of day-to-day work decisions made autonomously by 2028, up from none in 2024. The EU’s human-oversight duty for high-risk systems, including systems that monitor workers, was due 2 August 2026. Regulation 2026/1744 moved it to 2 December 2027. The transparency rules arrived on schedule. The requirement to have a human meaningfully in charge slipped sixteen months.

Ausman: if something is 99 percent good, who owns that one percent? A worker follows a procedure an agent gave them. The procedure was wrong. “Are they going to be reprimanded because they followed a procedure that AI gave them?” Accountability goes looking for an owner. Where nobody has decided, it lands on the person holding the handheld.

Software Only You Have

Scott Belsky: we are shifting from everyone competing on the same generalized world-class software to competing by creating software unique to us. Until now companies adjusted to Workday, Adobe, Microsoft. The obstacle to a bold workflow was often the generalized tool, designed to be good enough for everyone. Now small engineering teams inside companies can build custom software. Native to the culture. More integrated than any enterprise suite. The world’s best software in each industry, he thinks, will be the software that only one player has, and made for themselves.

Those who own a vital graph, and those who manage the intelligence of the organization, will thrive. He sits on Atlassian’s board; he is not a disinterested witness. Smaller specialized models may be the most cost-efficient solution for specific functions. The need to call frontier models by default is dwindling.

The best software will be the software only you have.

On the consumer side he has been living with agents: Instinct, Littlebird, Hermes. They hit bot checks. CAPTCHAs exist because they are bots. His implication: the moat is “favored agent” status, preferred access to Uber, Airbnb, Amazon, Amex, in exchange for becoming the default vendor. Right now agents navigate websites made for humans. Favored agents would go direct to the data.

Eric Newcomer dropped by Crosby, an AI-native law firm in New York, for a summit organized by Jake Saper of Emergence Capital. The gospel is AINS: AI-Native Services. Instead of selling software, hire lawyers, accountants, brokers, pair them with engineers, and charge for the result, a completed contract, not the hour. The playbook deck is in the member post. We have the lede.

Open Weights, Closed Frontier

Kimi K3, an open-weight model from Moonshot, became the leading open-weight model globally after its July debut, within striking distance of America’s closed frontier models. Dean Ball of OpenAI suggested one probable outcome of an open-weight-dominant world is “full AI communism.” Jensen Huang posted on open weights and American AI leadership. Jamin Ball of Altimeter asked where the American open-source models are, and pointed to investor incentives.

McKinsey, 700 organizations across 41 countries: 63 percent of respondents use open-source models, often alongside closed ones. Pinterest’s CEO, Bill Ready, on the earnings call: any CEO not taking advantage of open-source models is almost certainly wasting a lot of shareholders’ money. NVIDIA has Nemotron, including 3.5 Lightning last week. Meta released open weights for Muse Glimmer, announced an upcoming open-weight Muse Spark 1.2, and published “The Future Is for Everyone.”

The closed frontier is American. The open stack, so far, is not.

Wednesday’s ChinaTalk piece was the industrial-policy version of this. Today it is the board version. Same gap. Different letter.

Demand smacked headlong into a tight supply of compute.

The Collection · The Small ErrorPiece

Standing Orders

Four rules we are keeping

Outcomes. The pause. The gauges. The pallet.

  1. I

    Measure the outcome, not the confidence score

    Did the customer get the thing, on the day, undamaged. A 99 percent score is the machine marking its own homework.

  2. II

    The pause is not the post-mortem

    Altman said the models are misaligned. Twenty percent of inference is now a tax. We still have not seen the Hugging Face report.

  3. III

    Do not call it a bubble until a gauge is red

    EV’s lede is boom. 2027 is when they expect funding quality to turn. Print the date, not the vibe.

  4. IV

    Who owns the one percent?

    Write it down before the agent is in the aisle. The person with the handheld did not choose the model.

The Collection · The Small ErrorOrders