The success rate nobody measured

A completed review of one public document, done by Build n Bloom in September 2026 from public sources. A historical example, not a client engagement and not an installed system.

This is the standalone edition, written to be read without scripts. The same review sits inside the review pack page at /demos/assessment/review-kit#public-proof.

“Only about half of government AI pilots make it into production.”

We wrote that sentence in order to review it. It cites a real survey, and that survey cannot support it.

What the source supports

In autumn 2023, 87 of the 89 UK government bodies the National Audit Office approached answered its survey. Of those, 32 reported at least one fully deployed AI use case and 61 reported piloting or planning one; 29 reported both. That leaves four groups that do not overlap, and no measurement of anything over time.

The 87 responding bodies, split into groups that do not overlap. Derived by Build n Bloom from the counts the report gives: 32 deployed, 61 piloting or planning, 29 reporting both. Shares are rounded to one decimal and total 99.9%.
GroupBodiesShare of 87What the group means
Deployed only 3 3.4% At least one fully deployed use case, and no pilot or planned use case reported alongside it.
Reported both statuses 29 33.3% Deployed, and also piloting or planning. This group is the whole of the overlap between the two headline figures.
Piloting or planning only 32 36.8% A pilot in progress or completed, or a planned use case, with nothing fully deployed.
In neither group 23 26.4% Neither status reported. Not the same as no interest in AI: exploring is a different thing from piloting.
All respondents 87 100% 87 of the 89 bodies surveyed.
Deployed, as the report gives it
37% — 3 + 29 = 32 bodies.
Piloting or planning, as the report gives it
70% — 29 + 32 = 61 bodies.
Why the two cannot be added
29 bodies are in both. 32 + 61 counts them twice and gives 93, against 87 respondents. The two groups together cover 64 bodies.
Respondent count and proportions: summary, paragraph 12, printed page 8. Counts of 32 deployed, 61 piloting or planning, and 29 reporting both: Figure 8, note 2, printed page 30. The four exclusive groups are ours, subtracted from those counts. National Audit Office, Use of artificial intelligence in government, published 15 March 2024, reporting a survey run in autumn 2023.

Source and scope

The National Audit Office surveyed UK government bodies in autumn 2023 and published the result on 15 March 2024, in a 56-page report on the use of artificial intelligence in government. Of 89 bodies approached, 87 answered. It reports two headline figures: 32 bodies (37%) had at least one fully deployed AI use case, and 61 (70%) were piloting or planning one. The groups are not exclusive — a note under Figure 8 gives the overlap as 29 bodies reporting both. “AI” there is broader than generative AI, “piloting” covers pilots in progress and completed, and every answer is self-reported at a single moment. The NAO’s press release describes the same survey rather than a second one. This is a dated historical record: it describes those 87 bodies in autumn 2023. Public records can be revised, so what we reviewed is the version at the links below, read on 8 September 2026.

Page links open the publisher’s own PDF at that page. Nothing substantial is reproduced here; the summary is our own.

Check the reasoning

The finding above stands without any of this. The working is here so it can be argued with.

Why 37 ÷ 70 is not a conversion rate

The report gives the overlap, so the four groups follow from its own counts: 32 − 29 = 3 deployed with nothing else reported, 61 − 29 = 32 piloting or planning with nothing fully deployed, and 87 − 64 = 23 in neither group. No definition moves across those steps, which is what makes the arithmetic legitimate.

Interpreting 37 ÷ 70 as a conversion rate reads two counts of organisations, taken at one moment, as the beginning and the end of one journey. No body was observed changing status, so these figures carry no pilot-to-production conversion rate. Subtracting them for a “stalled” figure fails in the same way.

Figure 5 of the same report sorts each body by its highest status, so those categories do exclude one another. They are not the 70% group. Substituting them into a sentence built on the summary keeps the wording and changes the definition underneath it.

Why we reviewed this source

The subject being AI is incidental, and we know how it looks for an AI firm to choose an AI report. It is public, free to open and dated, so every number we use can be checked. Its two headline figures overlap, which is the shape that invites the mistake in the sentence at the top.

Could we say most responding bodies had fully deployed AI?

No. Full deployment was reported by 32 of the 87. The 70% figure covers piloting and planning, which is broader activity rather than deployment, and the two groups together cover 64 bodies — a majority reporting one status or the other, but not a majority with anything fully deployed.

Could we say the survey shows whether government AI paid off?

No. These two figures record status, not cost or return. Figure 8 asks what impacts bodies expected from AI rather than what they realised, so neither figure supports a statement about value in either direction.

Could we use the report’s Figure 5 percentages instead, to be safer?

No — not in the same sentence. Figure 5 sorts each body by its highest status, so its categories exclude one another. The 70% figure does not. Swapping one in for the other keeps the wording and silently changes the definition.

Could we cite the NAO press release as a second, corroborating source?

No. The press release describes the same autumn 2023 survey. Corroboration means a different instrument or dataset, not one dataset at two URLs.

What would settle the original question

A defined cohort: pilots identified at a stated start date and followed to production, abandonment, or still running at a stated end date. It can be collected going forward or reconstructed from records that already carry dates. What no dataset settles is whether the claim is load-bearing enough to wait for that evidence, soften it, or cut it — that judgement stays with the specialist.

What this shows, and what it does not

It shows that we work from the primary document and say where in it we are, down to the printed page; that we keep what the report gives apart from what we derived, and show the derivation; and that we will publish a finding thinner than the sentence it replaces.

It does not show an AI system running in anyone’s workflow, a measured client outcome, or the state of UK government AI today — the figures describe autumn 2023. It is a prepared review of a public source, not an incident from a live engagement, and it does not establish that your documents behave like this report. A measured outcome needs an engagement with a measurement plan agreed in writing beforehand.

Take the method with you

Free, no email. The editable assessment pack is a separate thing: it is the fictional worked case, sent by email from the form at the top of this page, and does not include these files.

If a claim like this one sits inside work your firm repeats, the free 25-minute workflow call is where to bring one example. You would leave with a clearer view of what evidence a decision about that step would need — not a promise that the step can be lifted out of your week.