Does a human check always find the hidden data in a file?
7 min read
On 30 July 2026 the House of Commons Defence Committee published “Shifting heaven and earth? The Afghan data breach and resettlement schemes” — First Report of Session 2026–27, HC 69: the fullest account yet published of how a spreadsheet exposed the data of thousands of people whose lives, on the UK government’s own assessment, were at risk, and of how the Ministry of Defence (MOD) kept that breach hidden for nearly two years. The report contains one admission that matters more than any other line in it, for anyone handling sensitive data in a company or a public body: even a second human check, the Ministry itself says, would probably not have found the data that caused the disaster.
A hundred and fifty apparent names, 18,500 real ones
In February 2022 a member of MOD personnel sent a spreadsheet to a trusted third party outside government. It appeared to contain information on around 150 applicants to the Afghan Relocations and Assistance Policy (ARAP), the scheme for Afghans who had worked alongside UK forces. It actually included detailed personal information relating to more than 18,500 applications. “This information was hidden to the casual viewer,” the Committee writes, “and we understand that the person sending the file was both unaware that it was included and apparently unaware that spreadsheet files can contain hidden data of this kind, let alone how to check for it.”
From that moment the Ministry lost all control of it: “Once the file left government systems the MOD no longer controlled who could access or share it.” The breach was not discovered for eighteen months, and not through an internal check: in August 2023 part of the dataset appeared in a Facebook group.
The contradiction
Giving evidence to the Committee, Sir Ben Wallace — then Secretary of State for Defence — said the breach could only have occurred if “operating procedures involving additional checks on sending data outside the MOD had been ignored,” adding that “someone definitely did not do their job.” The Ministry’s written evidence says the opposite: the breach took place “while officials were following agreed processes.” Asked about the discrepancy, the MOD told the Committee it was “unlikely the hidden data would have been found even if a second check had been made.”
The Committee’s verdict is blunt: “This is a stark admission of continuing cultural failure.” And it adds a detail that removes any technical excuse: “it is relatively straightforward to ensure that there is no hidden data in an Excel file, and central government guidance on how to protect against this known concern was well publicised at the time of the data breach” — “the MOD itself had issued local guidance on the same point.” It was “well established that spreadsheets could contain hidden or non-obvious data, including hidden rows, columns, worksheets, metadata, formulas and cached source data”; the Information Commissioner’s Office had published detailed guidance “since at least 2015”; by June 2021, central government guidance “required checks for hidden tabs, columns and rows before sharing or publication.”
Why a human check is not enough
Both versions can be true at once, and that is exactly the point. A written rule that no machine enforces is not a security measure: it is a hope. The check the MOD was relying on depended on a person opening a file and looking — but hidden data is, by definition, what you cannot see by opening a file. A check that runs on the file itself before it leaves, rather than on the person sending it, works deterministically, every time, and leaves a trail. That is what was missing here.
There is a second, equally structural point: the perimeter of control ended where the file ended. Once it left, there was no revocation, no audit, no way to pull it back. Tellingly, the breach surfaced not through an internal check but through a Facebook group, eighteen months later. The same principle governs the log of a company AI assistant: the question is not whether the policy exists, but whether something — a system, not a person — enforces it before the data leaves, and leaves a record of having done so.
The third element is the tool itself. David Williams, the Permanent Secretary at the time, described the working environment as “essentially a combination of the ad hoc use of spreadsheets on SharePoint sites.” The dedicated system — the Defence Afghan Caseworking System — arrived only in May 2022, three months after the breach; ARAP had been in development since September 2020. The Committee ties the three together: “The breach was not just an isolated mistake by an individual; it was a foreseeable systemic failure,” arising “from the combination of inappropriate tools, weak operating procedures, insufficient training, poor organisational continuity, and an inadequate culture of data protection and accountability.” More bluntly: “The MOD handled sensitive immigration casework using tools and controls not appropriate for a life-endangering dataset at any scale.”
Two years of silence
The government treated the incident as potentially life-threatening, and from September 2023, for nearly two years, a superinjunction prevented public reporting of both the breach and the government’s response, while major policy and spending decisions were taken in secret. The Committee draws Recommendation 62 from this: “Government should mandate and enforce minimum standards for skills, process, tools, controls, independent assurance and testing for datasets where compromise could plausibly risk life.” The report itself flags what it is: a committee report, with recommendations to government, which has two months to respond.
The other side of the story
This needs saying with the same precision as everything else. The Committee accepts that the fall of Kabul in August 2021 was a genuine crisis, and that the Ministry could not have foreseen the speed of the Taliban advance. But it adds that those pressures “help to explain how these weaknesses developed; they do not excuse their continuation into 2022.” This is not an indictment of the United Kingdom: the same combination — sensitive data in a spreadsheet, a check entrusted to a person, no automated verification before the file leaves — exists in organisations of every country. This report is rare precisely because someone wrote it up in full, with names and citations, instead of leaving it as an unspoken suspicion.
How we address this
The ARAP case has nothing to do with artificial intelligence, but the lesson is exactly the kind of system we build. The first axis is compliance you can produce on demand: a check that runs on the client’s documents and systems before a file leaves — what it actually contains, how many people are inside it, what is hidden — leaving a trail with the date, the file, the outcome and the recipient. Not a form signed once, not a policy someone has to remember to apply, but a system that reads the content every time.
Here the case for on-premises deployment is particularly strong: a check that has to read the content of confidential documents — payslips, personnel files — should not run on an external service the client does not control. We always work in two modes: on-premises, on autonomous machines that do not require deep integration into the client’s network, or on a dedicated CSIDIA cloud, with a dedicated VPN and a data centre in Italy, on premises we directly guard. Always with shared management, because almost no organisation already has someone in-house to administer a system like this — the same principle that governs the credentials of whoever administers a company AI system: designation and the audit trail are written before the system is switched on, not reconstructed after an incident.
The second axis is what makes the first one worth having. The same system that checks a file before it leaves also brings the organisation’s scattered data — archives, management systems, documents, sensors — into a single operational model, where a dataset of that sensitivity does not live in a spreadsheet attached to an email but has one home, with governed access, on which AI agents execute decisions with an operator in command: for large enterprises, defence, government and healthcare. Compliance is the way in; the decision-making system is what we sell.
Could you say today how many files in your organisation hold data that no one has ever properly checked? Half an hour with one of our experts is enough for the first map.