Operational notes Observatory

What must a model’s technical report tell you before you install it?

6 min read

Page of an open technical manual, with numbered drawings of the machining stages of a pivot and a caption for each step, in black and white
You fit the part with the drawing beside you. For a frontier model the drawing is the technical report: if it does not say enough, the part stays opaque.

On 27 July 2026, at 13:31 UTC, the first commit to the moonshotai/Kimi-K3 repository on Hugging Face made 96 weight files, some 1.56 terabytes, available for download. A little over ninety minutes later, on GitHub, a 47-page PDF appeared: the technical report. We have already covered the model. The point holds for any open model: of those two artefacts the first is opaque by construction, the second is the only one you can read.

A weight file declares nothing about itself. Everything you will tell a regulator or a workforce representative about how that system was built and measured has to come from the document that accompanies it. It is worth knowing what such a document contains, and what it almost never contains.

The release, as far as it can be verified

  • The date is in the commits, not on the card. The repository shows as created on 13 June 2026, but its history begins on 27 July: before that nothing was public. To date an adoption, the model card is not evidence.
  • A single specimen. No base checkpoint, no smaller sizes, no intermediate versions. One format only — MXFP4 weights, MXFP8 activations — with quantisation obtained during training rather than applied afterwards.
  • The licence is its own text, the “Kimi K3 License”: a broad permission with two thresholds — a separate agreement for anyone reselling inference above 20 million US dollars of revenue over twelve months, attribution in the interface above 100 million monthly users. We have written about how to read a weights licence.
  • The report sits in a repository, not on arXiv: no persistent identifier, no peer review, replaceable without leaving a trace. Archive it with its hash on the day you read it.

That the model at the centre of the debate over restrictions on open weights is now a downloadable file is news. Not the useful part.

What the report declares

We read it. On architecture it is generous: 93 layers, 896 experts of which 16 are active per token, 104 billion activated parameters out of 2.78 trillion. On evaluations it is, for the genre, unusually rigorous: it states the sampling parameters, which agent harness was used for each model, and that some third-party scores are taken from public listings “as of July 23, 2026”. It also declares its weaknesses: on research-level reasoning it writes that these remain “a key direction for improvement”, and names the suites where it lags.

Then comes a chapter you do not expect in an engineering document: a measurement of its own offensive cyber capability. On an internal set of 36 exploit-development tasks, measured by the laboratory itself, the model solves 14, and the report calls that “a lower bound” on real capability. It then cites a joint assessment by the UK AI Security Institute and the US CAISI: there the figure that matters to a company is not the score, it is the sentence stating that the model’s safeguards “did not prevent” attempts at exploit development.

What is missing

The section on training data runs to fifteen lines. It names four text domains — web, code, mathematics, knowledge — plus a vision corpus, and the methods: heuristics, classifier-based quality scoring, deduplication, rephrasing with fidelity verification. Nothing else: not how much text, not where it comes from, not the date it stops at, not on what basis it was used, not whether it contains personal data.

There is no limitations section: the limits declared are score gaps, not warnings about where the system should not be placed. Nothing on refusal behaviour, on harmful outputs, on discrimination or fairness. And not one line on what happens to any of it if the model is retrained — the first thing you will do.

This is not one laboratory’s failing: it is the genre. A technical report is an engineering document, written for people who build models, not the product documentation of a supplier.

The breaking point is item (d)

Here the argument stops being academic. Article 1-bis of Legislative Decree 152/1997 requires the employer to tell the worker and the union representatives the logic and operation of the system (item c), the categories of data and main parameters used to programme or train it (item d) and the level of accuracy, robustness and cybersecurity with the metrics used and their potentially discriminatory impacts (item f).

Put the report next to that list. Item (c) you can cover, better than with almost any commercial product. Item (f) halfway: the metrics are there, with conditions and dates, but their potentially discriminatory impacts are not addressed. Item (d) you cannot: four domain names are not categories of data in any useful sense, and the training parameters are not published.

This is the asymmetry open weights introduce. With a supplier those details go into the contract, and if they are missing the breach has an owner. With a file downloaded under a unilateral “AS IS” licence there is no counterparty to ask: you have gained the right to run the model and lost anyone to hold to an obligation. In front of the inspector the duty still rests with you.

What to look for, and what to ask for if it is not there

  • Origin and scope of the data. If the report does not give them, the written request goes to whoever integrates the model for you — sources, period covered, personal data — and the statement goes into the contract.
  • Conditions of the measurements. A score without who measured it, with which harness and on what date is not a measurement: it goes into no document of yours.
  • Admitted weaknesses. A report that declares none is less trustworthy, not more. If there are none, ask for the list of discouraged use cases.
  • Durability after adaptation. Ask whether the measured behaviours survive fine-tuning. Without an answer, every number describes a model you stop having the moment you adapt it.
  • Your own test. No published evaluation covers your documents, your language, your process: results reach your case only through a test set built by you.

How we solve it

Documentation is of little use if the model runs where you cannot look: a report says how a system was built, not who sees your data while you use it. That is why we work in two modes and no others: on-premise in the client’s own environment, or on our dedicated cloud — reserved for the single client, accessed over a dedicated VPN, with the data centre resident in Italy and premises we staff ourselves. In both, whatever the report does not declare you measure yourself, on your data, inside your perimeter. It is the method by which we keep enterprise AI under demonstrable control: architecture, not a commercial promise.

Do you have to get an open-weights model approved and nobody can say what data it was trained on? Half an hour with one of our specialists: we read the technical report together and list what must be asked for in writing.

Sources