On June 5, 2026, Anthropic published a report containing a sentence that would have sounded like science fiction three years ago: more than 80 percent of the code merged into the company's production systems is now written by Claude, not by humans. OpenAI has set dates — an automated AI research intern by September 2026, a fully fledged automated AI researcher by March 2028. Google DeepMind has let its AlphaEvolve system improve the algorithms that train Gemini. Recursive self-improvement — AI accelerating the development of AI — has stopped being a thought experiment in philosophical essays and become a stated engineering goal at no fewer than three companies simultaneously.

This site has covered parts of this before. “Every Seven Months” dealt with the doubling rate of what AI systems can accomplish; “Two Documents, Same Week” with the gap between Anthropic's self-improvement report and the company's economic policy thinking. The question that remains is the one increasingly being asked in earnest: not whether the loop can close, but how fast it happens if it does — now, with several actors pulling in the same direction with hundreds of thousands of GPUs each.

What the labs are actually doing

Anthropic's report, authored by Marina Favaro and Jack Clark, is the most detailed window into the daily operations of a frontier lab published to date. The numbers deserve exact citation, with the caveat that they are the company's own and not independently audited. The share of production code written by Claude: over 80 percent in May 2026, versus single digits before February 2025. Code output per engineer: eight times the 2024 level — with the company's own note that lines of code measure quantity, not productivity. In the company's recurring optimization test, where the model is asked to speed up training code, the result went from roughly 3x (May 2025) to roughly 52x (April 2026); a skilled human researcher reaches around 4x in half a working day. And in April 2026, Anthropic published the first example of an open research problem driven from hypothesis to result by AI agents: on a week-long alignment problem, the agents recovered 97 percent of the measurable improvement gap, where two human researchers reached 23 percent — but humans still chose the problem and built the scoring rubric.

OpenAI's announcement is of a different kind: not measured results but stated goals. In a livestream on October 28, 2025, Sam Altman and chief scientist Jakub Pachocki gave internal target dates — an “automated AI research intern” by September 2026, running on hundreds of thousands of GPUs, and a true automated AI researcher by March 2028. Altman explicitly added that they may fail entirely. The goal is informative nonetheless: one of the world's largest AI labs is now organizing its research roadmap around automating its own research.

DeepMind holds the empirically most elegant data point. AlphaEvolve, presented in May 2025, found among other things a method for multiplying 4×4 matrices using 48 multiplications — the first improvement of its kind for complex-valued matrices since Strassen's algorithm in 1969 — and sped up a central computational kernel in Gemini's training by 23 percent, cutting roughly one percent off total training time. One percent sounds small. But it is a closed loop in miniature: a Gemini-powered system making the next Gemini cheaper to train.

All of this should be read with a source-critical spine. The labs have commercial reasons to appear fast, the figures are internal, and Anthropic itself reports that employee estimates of productivity gains tend to be inflated. But the direction is supported by external measurements — and that is where we turn next.

The curve that shortened

The seven-month figure that gave the June essay “Every Seven Months” its name comes from METR's original study in March 2025: the length of tasks AI systems can complete independently doubled roughly every seven months. That figure is already old. METR's updated dataset, Time Horizon 1.1, published on January 29, 2026, shows that the doubling rate for the period from 2023 onward is closer to four months. In METR's pilot measurement from spring 2026, the strongest assessed model sat at an estimated 16–20 hours of task length — at the upper limit of what the measurement suite can handle at all without new, longer tasks. Anthropic uses the same trend in its report: four minutes in March 2024, an hour and a half a year later, twelve-hour tasks this year.

It is worth pausing on what this means methodologically. The confidence intervals are wide, the measurement suite is near its ceiling, and a trend that holds for three years need not hold for five. But the empirical core is uncomfortable for anyone hoping for a slowdown: the best-documented rate has not stood still or flattened — it has shortened, from seven months to roughly four. The development has outrun its own most-cited rule of thumb.

The counterarguments carry weight

Here a worse essay would shift into inevitability. The counterarguments deserve their full weight instead, because they are not straw men — they are the most likely explanation for why the loop is not yet self-sustaining.

The first is the difference between automated engineering and automated research. Anthropic's own report is strikingly honest on this point: Claude is today superhumanly fast at executing well-defined experiments, but the gap persists in what the company calls research taste — choosing which problems are worth attacking, which results to trust, when a line of inquiry is dead. In the agent experiment above, humans chose the problem and wrote the scoring rubric. Automating the perspiration is not the same as automating the judgment, and it is judgment that drives paradigm shifts.

The second is the verification bottleneck. Anthropic itself describes how human code review has already become a chokepoint as Claude produces more than humans can read — Amdahl's law in organizational form: the step that hasn't sped up sets the system's total pace. METR's randomized study from July 2025 cut deeper still: experienced developers who believed AI tools made them faster turned out on average to be 19 percent slower with them. Self-reported acceleration is not acceleration. This is, in my view, the single strongest reason to read the labs' internal figures with caution.

The third is physics and economics. Epoch AI's macro model GATE — the most serious attempt to date to model automated AI development — concludes that a pure software explosion, where AI improves AI without continued expansion of physical compute, does not add up: experiments require compute, and the efficiency gains land near or below the self-sustaining threshold once experiment costs are counted. Compute is built in the physical world, and that world has queues — “The Same Megawatt” showed that the Nordic grid queues alone hold tens of gigawatts and that a transmission line takes seven to eight years. The loop's speed is ultimately a function of concrete, transformers, and permitting processes.

The fourth is that the fastest scenarios have drawn qualified criticism. The AI 2027 scenario — where an automated researcher in 2027 triggers an intelligence explosion — has been picked apart, and the criticism of its timeline models is substantial: assumptions of unbroken exponentiality, sensitivity to parameter choices, absence of negative feedback. The scenario is worth reading as a stress test, not a forecast.

Base assessment and fast scenario

My base assessment, clearly marked as such: the software leg of AI development continues to accelerate, and something resembling OpenAI's “research intern” — systems that independently execute week-long, well-defined research tasks — will be in operation at several labs within one to two years. That is enough for a strong, cumulative speedup of AI research, perhaps a factor of two to four at lab level. But full recursive self-improvement — systems that without humans choose problems, judge results, and build their successors — is braked by the three documented bottlenecks of judgment, verification, and compute, and on the base assessment does not become reality before roughly 2030.

Then there is the fast scenario, and it deserves equally clear marking. The most unsettling partial result in Anthropic's report is not the code share but the judgment measurement: in the company's test of choosing the next step in real research sessions, the models went from beating the human's choice in 51 percent of cases in November 2025 to 64 percent in April 2026. If research taste turns out to be just another capability following the same curve as all the others — and so far every measurable capability has, in the company's own phrasing — the first counterargument falls. Add the race dynamic: three labs with stated goals and hundreds of thousands of GPUs each means it is enough for one of them to succeed. In the fast scenario, the loop starts to bite in earnest as early as 2027–2028, with a doubling rate that keeps shortening rather than leveling out. It is a scenario, not a forecast — but anyone planning around the base assessment should stress-test the plan against the faster course, because the difference between them is only a couple of years.

The loop meets the payroll

What does the pace mean for the labor market? “The Same Megawatt” provided the uncomfortable frame: by that essay's illustrative calculation, every gigawatt of data center capacity in the queues carries revenue requirements equivalent to 100,000–200,000 person-years of work, and the Nordic queues sum to 8–16 million person-years — if the revenues are ultimately drawn from making human labor more efficient. The closing loop is the mechanism that sets the pace of that extraction. The base assessment gives labor markets and education systems perhaps a decade to absorb the shift; the fast scenario gives a couple of years. The difference between the scenarios is thus not abstract — it is the difference between structural transformation and shock.

Early signals exist, but only where someone is measuring. In the United States, AI has topped Challenger, Gray & Christmas's monthly statistics on stated layoff reasons five months in a row, March through July 2026 — roughly 102,000 announced job cuts citing AI in the first half of the year, about 23 percent of all cuts, versus 5 percent for the whole of 2025. In Norway, the unions Nito and Tekna meanwhile see no signs of AI-driven layoffs. Both pictures can be true: diffusion lags, and the Nordics sit later in the chain. But there is a third explanation that should worry us more — the measurement gap. As far as we have been able to find, no Nordic government statistics separate out AI as a stated cause of redundancy; there is no box to tick. Challenger can count because American companies state the reason themselves. With the doubling rate shortened to four months, that gap becomes a policy problem in its own right: you cannot plan for what you do not measure, and the signal will show up in American data long before it shows up in Nordic data.

For the Nordics

The most remarkable thing in Anthropic's report is its conclusion: the company documenting its own acceleration argues that the world needs a verifiable option to jointly slow down — while noting that training runs are easier to conceal than missile silos and that trust infrastructure of that kind has historically taken decades to build. “At the Foot of the Singularity” covered Hassabis's parallel proposal; none of the mechanisms yet exist. The race is thus the default, and the pace is set by the least cautious actor.

For Nordic decision-makers this means three things. The advice from “Every Seven Months” — reversible preparedness, indicators early in the chain — still stands, but the urgency has changed: what doubled every seven months now doubles every four. The indicators to watch are concrete: METR's next data points, whether labs begin crediting models with research ideas rather than execution, and the energy queues — because the Nordics sit, as “The Same Megawatt” showed, on one of the loop's real choke valves and thus hold more bargaining power than the debate pretends. And institutions that plan in budget years now face a process that improves itself in quarters.

The question was never whether the loop can close. The question is whether we manage to decide anything before it does.

Sources: Anthropic, "When AI builds itself", June 5, 2026 · Anthropic, automated weak-to-strong researcher, April 2026 · Sam Altman on X, Oct 28, 2025 · TechCrunch on OpenAI's goals, Oct 28, 2025 · DeepMind, AlphaEvolve, May 2025 · METR, "Measuring AI Ability to Complete Long Tasks", March 2025 · METR, "Time Horizon 1.1", Jan 29, 2026 · METR, time horizons overview · METR, developer study, July 2025 (arXiv 2507.09089) · Epoch AI, GATE · Epoch AI, "AI and explosive growth redux" · AI 2027 · Critique of AI 2027's timeline models (LessWrong) · Challenger, Gray & Christmas, June 2026 report · Challenger, "AI Leads For Fifth Straight Month" · Digi.no on Nito and Tekna · Every Seven Months · Two Documents, Same Week · At the Foot of the Singularity · The Same Megawatt

Assessment note: the base assessment and the fast scenario in the section "Base assessment and fast scenario" are the author's own conclusions; the factual claims elsewhere are documented in the sources above. Anthropic's internal figures are the company's own and not independently audited — this is stated in the text where they are used.

Rolf Skogling writes AI-skiftet from an industry-oriented, practical perspective, grounded in how AI is actually used in organisations and production.