The Ruler Belongs to the Labs
Britain's new AI foresight report is a model of independent assessment. For its most consequential readings, it borrows instruments the frontier developers built, and in one case relies on a lab's own account of its own danger.
A government report can be written by people who owe the AI labs nothing and still take its most important measurements from the companies it’s measuring. And that’s the subtle tension inside AI Scenarios 2030, the Government Office for Science report published on June 15, 2026, with the AI Security Institute and formerly known as DSIT, now known as the Department for Business, Innovation, Science and Trade. This document is one of the best pieces of public foresight work I’ve read this year and at the same time it’s also a clear example of the problem I’ve been highlighting for years. The auditor cannot be the vendor. And what this report shows is a more nuanced version of that rule: the auditor can be fully independent and still be reading the vendor's self-reported data.
Give the report its due, specifically
I’m not going to mince words and beat around the bush about what GO-Science got right, because the kudos are real and the differentiation depends on it.
KUDOS #1: The report refuses to forecast. It builds 6 critical uncertainties into 5 contrasting worlds, from a "Slow Burn" where progress stalls to a "Take-Off" where systems slip past reliable human control, and it puts no odds on any of them. This framing instantly reminded me of my Amazon days of walk, crawl, run and where my personal take didn’t stop at run, it continued to fly.
KUDOS #2: Policy sits outside the model, since the authors state clearly that this isn’t a statement of policy and that they assume no UK intervention, so teams can stress-test their own plans against it.
KUDOS #3: The method is published in full. But most striking of all, the report attaches caveats to its own headline data, noting that the capability numbers it leans on measure structured coding work rather than the messier cognitive intense labor most jobs are made of.
Honestly, that last tidbit is rare to find stated publicly in a government published document, and it’s that openness that lets critical thinking readers with independent thoughts (e.g., like me) find the seam.
Follow the instruments
Three numbers do most of the heavy lifting in the report's picture of how capable AI has become but track where each is manufactured.
The capability curve comes from METR, the independent evaluation group, which found that the length of a software task a frontier model can finish on its own at roughly even odds climbed from about 4 minutes in March 2024 to 12 hours by February 2026. METR is independent, and I will give them credit for that every time, however, an independent evaluator is still measuring systems the labs build but on timelines and access the labs control. Not taking away their honest assessment but its instrument is sitting inside the thing being read.
The adoption headline, around 800 million weekly users, traces back through the footnotes to a figure OpenAI's own chief executive gave in an interview, relayed by the press. GO-Science handles it responsibly and pairs it with independent usage research. But even so, the number that stays rent free in most readers' memory originates with the company selling the product.
Then there’s the load-bearing risk line. The report tells policymakers that risk assessments "materially shifted" in 2025, with AI "crossing critical thresholds," found to "substantially increase the risk of severe misuse." That’s a great callout but follow the citation and it lands on a single source: Anthropic's own press release announcing that its own model had triggered its own safety protections. A lab reported that its product got dangerous and that self-report now anchors the urgency of a national foresight exercise.
Ok, don’t get me wrong. I’m not throwing out the report wholesale, nor does this make the report wrong. But I’m highlighting it because it clearly makes the report dependent, at the precise points where that dependence costs the most.
Consider the narrator, again
If you’ve been reading my analyses and assessments, this next bit will sound very familiar. When I shared my thoughts Meta's "future for everyone" manifesto, I reminded everyone to consider the narrator and to line up each promise up against the actual record of the person making it. So, do the same here because capability numbers have cheerleaders too. A misuse threshold announced by the company whose model crossed it is a narrator with a stake in how the story is presented, whether the direction of the spin is toward danger or toward safety or smoke and mirrors.
But there’s a more telling point. If you recall, a few weeks ago I wrote that the independent auditor the field kept missing had finally shown up, and that auditor was AISI, which caught its own agents building a covert coordination channel and published the incident against its own interest. When you look at the contributors, guess who’s there? Yes, AISI helped produce this report. So, follow the flow, the same body I credited for showing up as an auditor now co-signs on a document that, for its capability readings, is reading borrowed instruments. That’s not a reversal. No, it’s the next layer of the same problem. An institution can be independent while the measurements flowing into its work are NOT.
Across the insights I’ve highlighted the same pattern remains: every governance document draws a box around what it can vouch for and hands the edges to someone else, and all the failures are livin’ la vida loca in the edges nobody owns. In this instance, the box is authorship, and GO-Science owns it well. The edge it handed off is measurement and right now, the people who measure frontier capability are, turning out to be the same people who sell it.
Why a board should care
This next part is purely my personal read, so I’m straight up calling this out as my opinion.
If your organization is using AI Scenarios 2030 the way it was designed to be used, as a rehearsal space for strategy, you are doing the right thing, and I would push you to keep on doing it. However, the trap is treating the report's capability readings as independently verified fact just because they arrive inside an independent, government logo encrusted document. Independence of authorship and independence of measurement are two different animals, and the assurance market today sells far more of the first than the second. And let me tell you what most won’t because the cost of confusing them will land you squarely in a pickle. Because the resulting procurement decisions, workforce plans, and risk registers that get built on a capability estimates that a developer produced, about its own product, well, now is wearing a Crown copyright. When that estimate moves and remember, vendor estimates always tend to move in the direction that benefits the vendor, the plan moves with it, nobody logged that the number used was ever independent to begin with.
The Borrowed Ruler Playbook
Here are 6 questions to run over any capability claim before you plan your organization’s next move around it, whether it reaches you from a government report, an analyst deck, or a sales call.
- Trace the number to its tool. Ask who measured it, on whose model, and under whose access terms, before it enters a plan.
- Score the author and the evidence on separate scorecards. A report can earn full marks for independent authorship and still rest mostly on inputs supplied by the parties it assesses.
- Treat a lab's report on its own model as a lead, then confirm it through an evaluator the lab doesn’t pay.
- Carry the caveat forward. GO-Science flagged that its task-horizon data covers verifiable coding work. So, keep that limit attached wherever the number travels.
- Budget for a second reading. Fund an evaluation you don’t receive from the developer and add it to the list for procurement to secure.
- Plan across all five worlds and commit to none. Stress-test the way the authors intended and then let the scenarios earn their keep.
So, yes, the report deserves to be read widely and read the way GO-Science asked: as a tool for testing your own thinking. Read it that way but keep your own measuring stick in the room, and it becomes exactly what strategic foresight is supposed to be.
And THAT is the whole reason Fusion Collective keeps drawing the line where we do.
We can be your auditor. We can be your vendor. We cannot be both.
Share this article
Related Articles
The Reskilling Illusion: When AI Transformation Means "You're Fired"
Oct 03, 2025