Operational notes Scenarios

"The customer's edge should never become training data": true, but not because of the contract clause

7 min read

A dry-stone arch, built without mortar, spanning a gap between two rock walls, in black and white
The arch stands because of how it is built, not because anyone promises it will. That is the difference between an architecture and a clause.

On 3 August 2026 Palantir Technologies files with the SEC, as Exhibit 99.1 to Form 8-K, its press release of second-quarter 2026 results. Not a speech from a stage, not a post: an attachment to a mandatory filing with a regulator, next to figures the company must be able to defend in court. The headline leaves no doubt about the tone: 149% growth in US commercial revenue and 93% growth overall, “crushing consensus expectations”. But the line that matters most is not in the headline. It is in the CEO’s statement, where Alex Karp writes:

“Demand for AI sovereignty has now been unleashed. And Palantir is the only company that has demonstrated it can transform tokens into actual economic value. Our customers trust us to provide them with maximal control over their operations, data, and decisions. Their competitive advantage should never become the training data for future models. This quarter was otherworldly: our U.S. commercial revenue grew 149% year-over-year, our overall revenue grew 93% year-over-year, and our Rule of 40 score climbed to 155%. The sovereign AI revolution makes us very optimistic about the future.”

A customer’s competitive edge should never become the training data for future models. Not a line from a stage, in front of an audience that applauds regardless: a line filed with a regulator. And it is, by the direct statement of the person who runs this company, the foundation CSIDIA is built on. From the same release, we have already told a different angle — how much weight a single counterparty carries in the quarter’s revenue. Here the task is different: not to repeat Karp’s line, but to make it precise. In the version circulating outside an SEC filing, it gets used in a way that does not hold — and whoever uses it that way in front of an informed procurement office gets taken apart.

The numbers around the line

Karp cites them himself, in order: US commercial revenue at $764 million, up 149% year-on-year; overall revenue at $1.935 billion, up 93%; Rule of 40 — growth plus margin — climbing to 155%. The filing adds the rest: total US revenue at $1.573 billion, up 115%; full-year 2026 guidance raised to 82% growth, with US commercial guidance alone raised to at least 134%. Figures a listed company puts in writing knowing the market will check them every quarter. We cite them not to prove Palantir is right, but because Karp’s line sits in the middle of these numbers, not in an interview detached from consequences.

The version that falls apart in a meeting

The reading that circulates most often, simplified, runs like this: you pay a subscription to a closed model and, by using it, you train the competitor that will compete with you tomorrow. It is a striking argument, which is why it keeps being repeated. But it is contestable, and whoever brings it into a negotiation risks losing credibility on everything else. We checked, on 4 August 2026, by reading the published terms of the leading closed-model providers, that on API and enterprise plans training on customer data is excluded by contract: the major labs state they do not use inputs and outputs from these product lines to train future models, and offer zero-retention plans. We stay deliberately generic on attribution — “the leading providers state” — because that is not the point on which we want to be taken on trust: anyone who wants to know which provider says what can read it for themselves. In the subsequent analyst call the theme is understood to have been developed at greater length; we do not report it, since the only transcript in circulation is automated, with obvious errors, and unverifiable against an official source.

Why the argument holds anyway — three points

First: a clause is not an architecture. “We do not train on your data” is a contractual promise, not a physical constraint. It gets rewritten with an update to the terms of service, a change of ownership of the company that signed it, a change of jurisdiction. In the meantime the data has already left the customer’s perimeter: it sits on infrastructure the customer does not administer, under a promise the customer cannot verify from outside. An architecture in which the data never leaves does not need anyone to believe it: the guarantee is in the structure, not in the word given.

Second, and this is the point no clause covers: the real edge does not depend on that promise. Even at zero training, all the work that makes a model good at an organisation’s specific trade remains: targeted training, uploaded documents, task-by-task evaluations, corrections, everyday use that accumulates signal. That work improves a model the customer does not own, cannot inspect, cannot take elsewhere, and whose price and future availability it does not decide. The day it switches provider, that competence does not follow: it starts again from zero.

Third: someone else can decide for you, regardless of the supplier. From 10 August 2026, Commission Implementing Regulation (EU) 2026/1755 lets the European Commission demand that a provider of a general-purpose AI model give access to the model’s weights, to its hosting infrastructure, and to “all levels of access granted to employees of the provider” — and require it to disable logging of that access. It is not a hypothesis: it is a power written into a text already in force, exercisable on a model an organisation depends on and has no say over. No contractual clause overrides a Commission decision taken under the AI Act.

What we do, at our scale

We do the same thing, at a different scale and for different customers. Open-weight models, trained on the customer’s specific domain, run inside its own perimeter: on premises, on self-contained machines that need no deep integration into the existing network, or on a dedicated CSIDIA cloud, with a dedicated VPN and a data centre in Italy, in premises we guard directly. This is not a promise: we are already in the field, with systems live at a number of enterprise and other organisations, and the first installation goes live within weeks, not quarters.

The part that really closes the argument is different: we are multi-model, and we say so explicitly because it is proof that this is not a way of locking the customer to us. The model gets swapped whenever the customer wants — even just to benchmark it against another, or test it on real tasks. What makes an organisation good is not the model it uses: it is its ontology, its data, its processes. Those remain its own — including in relation to us: if the customer switches provider, it takes the installation with it, because it was never ours.

The other side, owed in fairness

Palantir says this because it sells the alternative to a closed, metered model. We say it for the same reason: we sell the same alternative. The reader has every right to discount both sources — ours included — and this article asks them to believe neither. It asks them to verify the argument, which stands on its own: the terms of closed-model providers can be read, the European regulation is published in the Official Journal, and the question of where competence actually accumulates is one anyone can put to themselves, with any provider — us included.

In fairness: Palantir is a huge, listed company, with $9.2 billion in cash; its numbers are not ours, and we do not borrow credibility from them. A company growing 93% in a quarter does not prove its thesis about data ownership is correct: it proves a growing part of the market is buying it. Two different facts, and only the second is certain.

How it applies, in practice

The first axis is the one you can show an inspector: compliance controls on a third party’s model — who has had access, what was sent to it, which clauses cover what — run on the customer’s own documents and systems, with the audit trail ready. The second is what makes the first worth having: the same system ties together the organisation’s scattered data — archives, business systems, documents, sensors — into a single operating model, trained on its specific domain, on which AI agents execute decisions with a human operator in command. For large enterprises, defence, government and healthcare, who accumulates the competence being built is the same question as who stays indispensable five years from now.

Want to know where the competence your organisation builds every time it uses an AI model is actually going — and whether it would stay with you if you switched provider? Thirty minutes, at no cost, to map it out.

Sources