Operational notes Observatory

Who answers for the benefit figures you publish?

7 min read

A graduated steel tape measure laid diagonally across embossed card, in black and white
A graduated scale tells you how long something is. It does not tell you why it changed.

On 22 July 2026 the Office for Statistics Regulation — the UK regulator of official statistics — closed with a public statement its casework into the way NHS England presents the benefits of the Federated Data Platform, the English health service’s data platform whose supplier is Palantir. Four days later the Health Foundation published an independent analysis covering the same ground, picked up by the British press on 27 and 28 July. Two separate fronts, one question: did that system actually work?

We have already set out the facts — baseline, measured outcome, what counts as proof. What matters here is something else: what changes, now, for anyone buying technology on a promise of efficiency.

The rebuke did not go to the supplier

Palantir is never named in the OSR statement. The addressee is NHS England, the body that calculated and published those figures, assessed against the Code of Practice for Statistics. That is the point that should sting for buyers, in a company as much as in the public sector: the benefit slides your supplier hands you become yours the moment they end up in a report, a press release or an answer to a parliamentary question. And they do not become evidence because you republished them.

Among the commitments NHS England has given, one says precisely this: to clarify on its website that the platform’s case studies “are written by organisations using FDP themselves and are not authored or verified by NHS England”. As we write, the caveat on the two types of benefit is already online on the main benefits page; on that page and on the methods page, the clarification about case studies has not yet appeared.

The same pattern, on a different document: in the statement it updated on 28 July, the National Data Guardian notes that the data protection impact assessment it had reviewed “stated that access to identifiable patient information would be limited to NHS staff with a legitimate need”, whereas external contractor staff also had such access. NHS England has acknowledged that the assessment “did not accurately reflect the operational arrangements in place”, called it an error and apologised. This is not a penalty: it is the same lesson, namely that a published document belongs to whoever signs it.

Counts of actions and benefit calculations are not the same thing

The first commitment is already visible. The benefits page now states that NHS England publishes two types of data: on one side “a count of benefits deriving from actions taken by local NHS users through products”, on the other “a benefits calculation applied by NHS England”, that is, observational before-and-after comparisons for which “NHS England cannot draw conclusions about cause and effect as other variables have not been controlled for”.

These are quantities of a different nature. How many times the system was used, how many records were reviewed, how many alerts were opened: those are usage data, and they are counted. How much time was saved, how many days of delay were avoided: those are inferences, and they are derived. Confusing the two is the most common error in project reports, and now a regulator has written it down about a system already in service.

Translated into a specification: two separate lists. Usage metrics, with the source of the count. Outcome metrics, with the calculation method, the time window and the comparison group. If an indicator cannot say which of the two lists it belongs to, it is not an indicator.

The independent evaluation comes afterwards, and that is the problem

NHS England has commissioned an independent academic evaluation from Imperial College, prioritising the two products where the causality question is most evident, and has invited Health Foundation analysts to contribute. It is the right decision. But it arrives with the platform in service for years: according to the Financial Times, final results are not expected before 2029.

The most instructive detail sits in another commitment: NHS England “will continue to consider possibilities for incorporating comparisons with controls, noting the issue that there is no common adoption date for FDP at trust-level”. Translated: building a comparison group today is hard because adoption happened piecemeal, without a design. That choice had to be made before installation, while a comparison group could still be assembled.

It has to be demanded during the tender, not after the fact: which indicators, measured how, against which comparison group, from which dated baseline, and who certifies the result. The lessons learned on AI procurement reach the same point from another angle. After signature only the before-and-after is left — and before-and-after does not separate the system’s effect from a process change, from seasonality, from a national programme, or from the fact that the first adopters were organisations already in better shape.

This is not a piece against one supplier

The same dynamic recurs with any technology sold on a promise of efficiency, whatever name is on the invoice. The seller has an interest in producing favourable figures; the buyer has an interest in showing them; neither, alone, has an interest in building the comparison that might contradict them. The contract has to demand it, with verifiable limits rather than confidential undertakings.

And the converse holds, with the same precision. The independent analysis declares itself descriptive: it states that an inferential or causal evaluation “would require additional data and is outside the scope” of the work, and that the tool “may be useful and received well by staff who use it” even without evidence that the benefits translate into measurable reductions. Not having measured a difference does not prove the tool is useless: it proves that, with the available data, effectiveness is not demonstrated. Those are two different statements, and confusing them repeats the very error being criticised.

Where the parties stand

Palantir, quoted by the Financial Times, argues that the analysis “appears to misunderstand what the tool does” and that including short-stay patients in the comparison amounts to “comparing apples and oranges”. NHS England stands by the reduction in long-stay discharge delays, says it made data available to the Health Foundation and discussed the different approaches, and has commissioned the independent evaluation. The OSR did not call the supplier to account: it called to account the body that publishes.

The point for buyers

A system running inside your own perimeter produces, by itself, the data with which to measure it: usage logs, processing times, outcomes, queues. Those data are yours, not the supplier’s. That is the difference between building an evaluation and receiving one already written.

This is why we work in two modes, and only two: on-premise, with the AI installed in the customer’s own environment, or on a dedicated cloud reserved for the single customer, reached over a dedicated VPN, with the data centre in Italy and the premises staffed directly by us. In both, the measurement stays where the data are: how it is built.

About to sign a contract that promises a measurable benefit? Half an hour to write down together how you will measure it.

Sources