Operational notes Observatory

OpenAI and the publishers: the real stakes are the evidence, not the copyright

4 min read

Facade of a courthouse with columns
The copyright case has become a case about evidence. That is where it will be decided.

The biggest AI copyright case has shifted ground: the argument is no longer only whether training a model on protected texts is lawful, but what happened to the evidence. On 9 July 2026 the group of publishers led by the New York Times filed a motion seeking sanctions against OpenAI in the copyright litigation. It is still an open matter, with serious allegations on one side and denials on the other — we report it for the verified facts and for the lesson it already offers, one that concerns anyone handling data within a business.

The facts, in order (and with both versions)

  • The sanctions motion (9 July 2026): the publishers argue that OpenAI was able to search its own data and logs for protected content, but concealed this capability from the court for roughly two years.
  • The tools cited: according to reported testimony from an OpenAI engineer, the company had built internal search tools — including a project called “Giraffe” — capable of flagging when ChatGPT’s responses closely tracked existing texts.
  • The data deletion: the publishers claim OpenAI deleted billions of conversations after a preservation order issued by the court.
  • OpenAI’s position: the company denies any wrongdoing, defends its conduct, and cites user privacy protection as the reason behind its data-handling choices.

The context raises the stakes further: a few days later (14 July) Google was hit with a new lawsuit brought by publishers over training Gemini on protected books, and looming in the background is the $1.5 billion settlement with which Anthropic closed its own case over content used in training — the largest payout in the history of US copyright law. We take no position on the merits: we await the rulings. What interests us is what the case already teaches.

The heart of the accusation is not training: it is spoliation of evidence — the claim that relevant data was deleted after a preservation order. For any business, the translation is immediate: once litigation exists (or is reasonably foreseeable), a duty to preserve relevant data kicks in, and “routine” deletion becomes a serious legal problem. Adopting AI multiplies the places where that data lives — logs, histories, prompts, outputs. Knowing what is retained, where, and for how long is not technical housekeeping: it is governed legal exposure. It is the same principle we set out for agents acting on your systems: without a complete record and a retention policy, the “afterwards” is ungovernable.

Lesson 2 — The provenance of training data is a risk the customer inherits too

If the model your service is built on was trained on content of disputed provenance, that risk does not stay entirely upstream: it resurfaces in contractual indemnities, in service continuity, in reputation. This is why the Data Act and the AI Act push for transparency on training — the summary of content used is not paperwork, it is your vendor’s risk profile. Asking for it, and putting it in writing in the contract, is the difference between a governed risk and one inherited in silence.

Lesson 3 — “User privacy” and “retention obligations” must be designed together

OpenAI’s defence rests (in part) on privacy: deleting data to protect users. It is a real tension — the GDPR and data minimisation push towards deletion, while evidentiary preservation obligations push towards retention. The lesson is not to pick a side: it is to design in advance a retention policy that reconciles the two, with explicit exceptions for legal holds. Whoever decides this only after the judge’s order has already lost control of the narrative.

What to do, inside your organisation

  1. Map where your AI data lives: logs, prompts, outputs, histories — in your own systems and in your vendors’.
  2. Define a retention policy with a legal hold procedure: who triggers it, what it freezes, how it is documented.
  3. Demand from vendors a summary of training data and contractual indemnities on provenance.
  4. Keep a human in control over critical actions and deletions: it is the rule we never waive, and here it protects you legally too.

Want to understand which AI data you should be retaining — and what to ask your vendors about provenance? Half an hour with one of our experts to map it out.

Sources