How you measure whether a system actually worked
6 min read
On 26 July 2026 the Health Foundation — an independent British health research charity — published a descriptive analysis of the relationship between adoption of OPTICA and delayed discharge from English hospitals, with the letter its Chief Executive, Jennifer Dixon, sent on 22 July to Ed Humpherson, Director General of the Office for Statistics Regulation. That same 22 July the OSR — the UK regulator of official statistics — closed with a public statement its investigation into the performance metrics NHS England derives from the Federated Data Platform.
Two distinct facts: an outcome measure that does not confirm the expected benefit, and a regulator asking for caution in how benefits are claimed. Together they explain how you check whether such a system worked — and it concerns us too, who are in the same trade.
One point the headlines skip: on NHS England’s pages OPTICA (Optimised Patient Tracking & Intelligent Choices Application) is a “nationally commissioned, locally developed” product, built by North of England Care System Support (NECS) for discharge planning; it runs on the Federated Data Platform, supplied by Palantir.
What was measured, and with what design
In the English system a “delayed discharge” is a statistical return, not an impression: hospitals report daily, in the Acute Discharge Situation Reports, how many patients remain in hospital despite being declared medically fit to leave. Different metrics are built on that base: that is where everything is decided.
NHS England publishes the percentage change in average “delay days” by length-of-stay cohort, comparing the twelve months before adoption with the twelve after. As at 31 March 2026 the site reports 25 trusts realising benefits and a fall of 13.58% for stays beyond 7 days and 15.25% beyond 21.
The Health Foundation measured a different quantity: for each trust and month, the median daily number of patients ready but not discharged, divided by occupied beds — the share of beds held by people who could already be elsewhere. Period 1 December 2021 to 31 May 2026; 25 adopting trusts and 91 non-adopting, with twelve months of data either side of adoption for 21 of the 25; names and adoption dates supplied by NHS England.
Result: no meaningful difference between adopters and non-adopters, in levels or in trend. In the 21 trusts with a complete series the share falls from around 14.5% one year before adoption to 12.5% four months after, rises to about 14% at nine months and then declines: the gradual improvement was already under way and continues at broadly the same rate. Of the case study cited by NHS England — a North East trust credited with a 37% fall in delay days for long stays — the analysis states it cannot replicate the figure from published data: it finds 3,575 delay days a month between February and November 2023 against 3,725 in January-February 2024: +4.2%.
What the analysis does not say
The document is descriptive and says so: a causal evaluation “would require additional data and is outside the scope” of the work. It attributes no causes and judges no software. The covering letter is explicit: “This is not a criticism of OPTICA — it may well be a helpful tool that is valued by staff who use it.” It closes on the most interesting hypothesis: the tool may act on one factor — communication between staff — while others, such as social care capacity, limit its effect on the final indicator.
The regulator did not judge the software
The OSR goes to the wording, not the technology. On 6 June 2026, it records, NHS England added to its methods page, in every section resting on before/after comparisons, the caveat “we cannot therefore draw conclusions about cause and effect as other variables have not been controlled for”: it calls the addition welcome and registers the commitments made — the same caveat on the main benefits page, explicit labelling of before/after comparisons in public communications, an independent academic evaluation commissioned from Imperial College with Health Foundation analysts invited to contribute, publication of trust-level disaggregated data, and a clarification that case studies are written by the organisations using the platform, not by NHS England. It expects figures cited by ministers or in official briefings to be communicated “in a clear, accurate manner”.
The question almost nobody asks
Between “the tool does not work” and “the organisation did not change what the data suggested” the difference is enormous, and public data cannot separate them. What follows describes no real organisation: these are recurring patterns, declared as such.
- A baseline not agreed beforehand. If the starting point is reconstructed afterwards, the comparison window becomes a choice, and every choice yields a different number. Record and date it before go-live, with the criterion by which the system counts as “adopted”.
- An outcome deduced from the contract, not the ward. The indicator must be chosen with the people who will use the tool. A tender asks for “fewer delays”; the ward knows whether the bottleneck is internal communication or a place missing outside the hospital.
- A correct prediction that moves nothing. A system can identify early, and accurately, who will get stuck. If no process changes downstream — nobody rebooking, reallocating, reserving — the indicator stays put: the model is right and the number does not move.
- A number without its context is not proof. A before/after comparison among adopters describes a trend; only comparison with non-adopters separates the effect from the background.
Where the parties stand
The Health Foundation published analysis and letter, copied to the chair of the relevant parliamentary committee and to the Chief Executive of NHS England. NHS England has added the caveat, keeps figures and methodology online, and has made the commitments listed by the OSR. As at the date of this article no specific public reply to the 26 July analysis appears to have come from NHS England, Palantir, NECS or the trust named in the case study.
The point.
Neither document shows the tool is useless, nor the opposite: they show the claimed benefit was not verified with a design capable of carrying it. A lesson about method, not about a product — and it applies to us in exactly the same way. A serious supplier accepts three conditions before signature: an outcome chosen with the people who will work with the system, a baseline recorded and dated before go-live, and the final measurement made by a third party. Not in writing, they remain unsellable as proof.
We stay with what we can guarantee: two modes, and only two. On-premise, with the AI installed in the customer’s own environment, or on a dedicated cloud reserved for the single customer, reached over a dedicated VPN, with the data centre in Italy and the premises staffed directly by us. That is the architectural part: how it is built, what it means in health care, and the chain that assigns an appointment, already taken apart.
About to start a system without having fixed the outcome to be measured? Let us talk it through in a thirty-minute session: we begin with the baseline.
Sources
- The Health Foundation — Analysis of the relationship between adoption of OPTICA and delayed discharge from hospital (26 July 2026), with the analysis and the 22 July letter to the OSR
- Office for Statistics Regulation — NHS England Federated Data Platform (FDP): presentation and communication of information and performance metrics (22 July 2026)
- NHS England — NHS Federated Data Platform uptake and benefits
- NHS England — Methods for the published uptake and benefits information