Hardware and models

Open-weight AI on dedicated hardware, inside the perimeter you control.

Before you sign, you should know four things: where the compute runs, who wrote the software, who holds the model weights, and what leaves your perimeter. Here are the four answers — and the limits, stated on the same page.

Black-and-white close-up of the fins of an aluminium heat sink, side-lit against a dark background

With a subscription service, three things leave. Not one.

  • The documents.

    Prompts and attachments reach a system you do not administer. Italian Legislative Decree 63/2018 protects know-how as a trade secret only if the holder takes reasonably adequate steps to keep it secret: a manual carrying calibration parameters, thresholds and maintenance cycles, uploaded to a generic assistant, quietly stops meeting that condition. And the GDPR, articles 24 and 32, requires you to know and govern the processing: an undeclared use is not governed.

  • The metadata of the work.

    It is not only documents that leave. Whoever selects and annotates your examples sees which datasets matter, in what volumes and in what order of priority: that is metadata describing your roadmap. In June 2025 Google, OpenAI and xAI walked away from Scale AI within days of Meta taking a stake — not because anything had been proven, but because the supplier’s ownership structure had changed.

  • The ability to verify.

    A use limit you cannot check technically is a promise, not a control. If the usage rules can change with an update on the supplier’s website, your limits are a web page. Clauses become verifiable when the infrastructure is under the buyer’s control: the logs are yours, and you audit when you choose.

Three layers of ownership: the machine, the software, the model.

The machine.

On-premise, the machine is yours: the AI runs on infrastructure you already control, and administration stays with your IT. On the dedicated cloud the environment is reserved for a single client — never shared, never a generic public cloud — reached over a dedicated VPN, with the data centre resident in Italy and premises we staff directly. We tell buyers to ask anyone offering capacity “in Italy” for three names: who owns the infrastructure, who operates the service, who controls the operator. We give ours before you sign.

The software.

The method is already built — data, ontology, agents, human control — and what is made to measure is only your domain, the one part nobody can reuse. This is not a licence to configure: it is a system built around your processes, added on top of the business systems you already run rather than replacing them. Public buyers already hold the lever: Italian Legislative Decree 36/2023, article 30(2), requires the contracting authority to secure the source code, the documentation and everything else needed to understand how the system works. It is a clause you can only exercise when the system is under the buyer’s control.

The model.

Open-weight models: a copy of the weights downloaded and archived inside the perimeter where it runs, not a remote endpoint. Archived alongside them: a cryptographic hash of every file, the exact revision reference, the tokeniser, the configuration, the chat template, the model card and the licence text as it stood on the day of download. The other side belongs in the same paragraph: open weights arrive AS IS, with no warranty of title and no warranty of non-infringement, and clearing third-party rights falls on whoever installs them. The licence tells you what you may do with them. It does not tell you where to run them.

A weights file inside your perimeter has no remote switch.

The precedent is recent. On 31 October 2025 Google pulled Gemma from AI Studio, hours after a letter from Senator Marsha Blackburn: the public surface disappeared in an evening, the already-downloaded weights did not. A service reached over a remote interface, by contrast, can stop or change behaviour without your consent and without useful notice.

It is worth being precise about the instruments being named this year. An Entity List designation acts on exports, re-exports and transfers to the listed party; a procurement ban binds the buyer. These are levers on future flows, not orders to delete copies already distributed. The risk is called supply, not seizure, and it is managed in advance: a second candidate from another jurisdiction, already downloaded and already tested.

The reverse, in full. A frozen model does not improve: no new checkpoints, no fixes to unwanted behaviour, no adaptation to new formats and languages. And the security fixes that matter almost never sit in the weights: they sit in the engine that runs them. CVE-2025-32444 affects vLLM, CVSS score 9.8, remote code execution, fixed in version 0.8.5. Freeze the environment so nothing breaks, and you freeze its defects too.

Four steps. You climb only when the step below has been measured and found short.

Each step has three lines: what it involves, what it needs from you, and when it is not the answer.

  1. 00

    The open model, as it is.

    An open-weight model chosen for the task, quantised and sized to the machine you have. Memory is counted in bytes, not parameters: a model of around 30 billion parameters fits a 48 GB GPU at 8-bit and a 24 GB one at 4-bit; for a 128-billion dense model Mistral states four H100 cards, six to eight in production; a 2.8-trillion-parameter model needs over 1.4 TB of memory just to load the weights and more than sixty accelerators across nodes — within reach of a hyperscaler, not of a mid-sized company. These are vendor-stated figures, not ours.

    The real documents to judge it against: real prompts, real documents, expected answers. That is the first thing you build, before you even pick the model.

    This is where most cases end. For classifying, extracting, summarising and checking documents against requirements, a well-chosen open model is often more than enough.

  2. 01

    Context, not weights.

    Your data connected, the ontology that makes it readable by an agent, guardrails on what the AI may touch. It is the layer the platform is built on, and the work this site has always described.

    The sources you already own, connected read-only at the start. No training data.

    It does not change the weights, so it invalidates none of the vendor’s measurements and survives a change of model without friction. That is why it comes before tuning, not after.

  3. 02

    Tuning on your own examples.

    This is the step we take into production most often, and it is work we have already done. The model is adapted to the language and formats of your organisation. The cost is not in the compute: it is in preparing the examples. Examples selected, annotated and corrected by experts are the slow part of the chain, the part no download replicates — and whoever does that work sees your priorities. That is why the work stays inside the perimeter, and why the curated corpus stays yours even when the model changes.

    A curated corpus and an evaluation set of your own. After tuning, the vendor’s measurements no longer apply: only yours do, on your documents. How many examples it takes depends on the task, and you find out by measuring.

    Not every model adapts well: some are quantised during training and have no full-precision checkpoint to start from. The licence decides whether you may, and on what terms. And a small quantity does not mean a small risk: the UK AI Security Institute has found that defences introduced through fine-tuning can be undone with a few dozen training examples.

  4. 03

    Pre-training a model of your own.

    Building the model instead of adapting it. The only order of magnitude we can publish is a European public project: Soofi S, 31.6 billion parameters, roughly 26.68 trillion tokens, from 24 March to 13 May 2026, on up to 512 NVIDIA B200 GPUs for about 253,000 GPU-hours. Seven weeks of a national cluster, paid for once by a state budget.

    A corpus in the order of trillions of tokens, with documented provenance. And the understanding that whoever pre-trains becomes the provider of a model, with the documentation duties that follow — no longer only a deployer.

    It makes sense only when the domain is genuinely unrepeatable and the corpus already exists. In the large majority of cases it is not the answer, and we say so before assessing it, not after.

You move up a step only when the one below has been measured and found short. If the value is not measurable, it stops there: that holds for the operational trial, and it holds for this ladder.

What it costs, and the line that never reaches the invoice.

The comparison that counts is not against the cheapest API on the market: it is against the bill you are already paying. Send your documents to a closed frontier model and you pay by consumption, on a price list you do not negotiate, for a spend that rises with use and has no ceiling — the better the system works, the more it costs, and the project’s success becomes its budget problem. APIs for Chinese open-weight models cost up to ten times less than frontier ones; open weights on hardware you own take the cost per token to zero. What remains is power, maintenance and depreciation: your own lines, and forecastable.

The most expensive line, though, never reaches the invoice. Paying by the token hands the supplier the thing that makes your business different from the rest: which questions you ask, against which documents, how often, and in what order of priority. It is the competitive edge no price list shows, it has no cost line, and once it has left the perimeter it does not come back. That is why this page talks about ownership before it talks about price.

Finally, the nature of the spend changes: a capital line you put on the balance sheet, depreciate and forecast, instead of a variable cost set by a price list that can change with an update. We state the limit anyway, because it is real. Measured against the cheapest API in circulation — on DeepSeek’s list, $3,999 buys over 14 billion output tokens — a single desktop machine does not pay for itself. That arithmetic holds for one user and for the lowest price on the market, not for a company sending its documents to a frontier model. The threshold is worked out on your real volume: if it does not add up, we tell you first.

The model can be replaced. The dependencies that remain, we name.

None of these disappears by bringing the system in-house. What changes is that they are named, that they can be replaced one at a time, and that the question to put to any supplier — us included — is not which model you use, but what it costs to change it.

EU compliance does not stand in the way of controlling the stack: it requires it.

Several rules ask for evidence that can only be produced when the infrastructure is under the control of the party accountable for it. The AI Act, article 26(6), requires the deployer to keep the system logs “to the extent that such logs are under their control”, for at least six months: if the supplier does not hand them over, the duty cannot be met. Article 25(4) requires a written agreement setting out the information, capabilities and technical access needed for compliance. Italian Legislative Decree 36/2023, article 30(2), gives the contracting authority a right to the source code and to how the system works. Legislative Decree 63/2018 protects a trade secret only where reasonably adequate steps have been taken.

The limit of our own argument, stated here: keeping the processing inside the perimeter is one of those measures, not compliance itself. It makes others possible — your own logs, audits whenever you want, usage rules that do not change without going through you — but it has to be described and traced: who has access, over which channel, where the document is stored.

And this is not self-sufficiency. Giving up the best models in the world on principle is a tariff your competitors will not pay. The strategy we see working is layered: data and the operating model in-house, under European law, non-negotiable; open models in-house where the data cannot leave; frontier models over an API with enterprise guarantees where you need maximum capability — behind an architecture that keeps them replaceable. You decide what leaves and what does not.

Where the data demands it, the architecture is not chosen: it is inherited from the classification. For public bodies, residence in Italy does not qualify a service: qualification against the class of data has to be verified before the tender, not at signature.

What we do not promise.

What you are left holding.

  • The environment that runs the model, in the mode you chose, and the inventory of what it takes to stand it up somewhere else: weights with a cryptographic hash and the exact revision reference, tokeniser, configuration, templates, model card and the licence text as of the day of download.
  • The recipe to rebuild the environment: versions of engine, drivers and libraries, startup parameters, quantisation procedure.
  • Your evaluation set — real prompts, real documents, expected answers — which is the only proof that a replacement model does the same job.
  • The ontology and the curated data, in open formats and owned by you: the part no download replicates, and the part that stays yours when the model changes.
  • The logs of actions and approvals, kept where you can read them, retain them and produce them.
  • If you choose the dedicated cloud, the three names behind the infrastructure it runs on: owner, service operator, and who controls the operator. Before you sign.

Two delivery modes.

On-premise, in your own environment

The AI runs on infrastructure you already control. Documents never cross the boundary of your network, and administration stays with your IT department.

Dedicated cloud, in Italy

An environment reserved for a single client, a dedicated VPN, a data centre resident in Italy, in premises we staff ourselves. Nothing is shared with other clients.

In both cases the models are dedicated and closed, disconnected from the open web: nothing they read feeds third-party services. They are open-weight models, with the weights archived inside the perimeter where they run: the version you use changes when you decide it does.

Frequently asked questions.

  • Can we run an AI model inside our own company without sending documents to a cloud service?

    Yes, and it is one of the two modes we deliver in: on-premise, with the AI installed on infrastructure you already control. The other is the dedicated cloud, an environment reserved for a single client, reached over a dedicated VPN, with the data centre resident in Italy. In both cases the model is dedicated and closed, disconnected from the open web, and documents do not leave the perimeter. The choice between the two turns on who administers the infrastructure and on the class of the data — not on price.

  • What hardware do you need to run an open-weight model in a company?

    Memory is counted in bytes, not parameters, and these are vendor-stated figures. A model of around 30 billion parameters takes roughly 32 GB at 8-bit and roughly 16 GB at 4-bit: it fits a single 48 GB professional GPU, or a 24 GB one in the more compressed build. For a 128-billion dense model the vendor states four H100 cards, and six to eight in production. A 2.8-trillion-parameter model needs over 1.4 TB of memory just to load the weights and more than sixty accelerators across nodes: within reach of a hyperscaler, not of a mid-sized company. Quality after quantisation has to be checked on your own documents, because it degrades differently depending on the task.

  • Can a supplier revoke the weights of an open model?

    A weights file already copied inside your perimeter has no remote switch: it keeps working even if the supplier changes strategy or a repository takes it down. The precedent is 31 October 2025, when Gemma was pulled from AI Studio hours after a letter from a US senator: the public surface disappeared in an evening, the already-downloaded weights did not. Entity List designations and procurement bans act on future flows; they do not order the deletion of copies already distributed. The risk is called supply, not seizure. The reverse has to be said too: a frozen model does not improve, and the security fixes that matter sit in the engine that runs the weights, not in the weights.

  • Is on-premise AI cheaper than a cloud subscription?

    Against what you pay a closed frontier model today, yes, for three reasons that compound. Cost per token goes to zero: the weights run on hardware you own, and what remains is power, maintenance and depreciation — forecastable lines instead of consumption billing on a price list you do not negotiate and that rises with use. The know-how stays inside: which questions you ask, and against which documents, reaches nobody. And the spend changes nature, from variable cost to a capital line on your balance sheet. The limit is real and we state it: against the cheapest API in circulation — on DeepSeek’s list, $3,999 buys over 14 billion output tokens — a single desktop machine for a single user does not pay for itself. The threshold is worked out on your real volume, and if it does not add up we tell you first.

  • Is the dedicated cloud a public cloud? Where are the data centres?

    It is not a generic public cloud and it is not shared: it is an environment reserved for a single client, reachable solely over a dedicated VPN, with the data centre resident in Italy and premises we staff directly. It involves no access by non-EU entities. That said, a data centre in Italy does not by itself qualify a service: for public bodies what counts is the class of the data and the qualification of the service, to be verified before the tender. And the three names — who owns the infrastructure, who operates the service, who controls the operator — we give before you sign.

  • Does fine-tuning on our data make the model ours?

    It makes the model better suited to your task, not “yours” and not better in absolute terms. Three things to know first: the cost is not in the compute but in preparing the examples, which is the slow part no download replicates; after tuning the vendor’s stated measurements no longer apply and you need an evaluation set of your own; and a few dozen examples are enough to shift behaviour, including where you did not intend it. The weights licence decides whether tuning is permitted and on what terms. The curated corpus, by contrast, stays yours and is reused when the base model changes.

  • Is an open-weight model on dedicated hardware GDPR-compliant?

    Compliance is not a property of the model: it is a property of the processing. Keeping the processing inside a perimeter you control is one of the measures articles 24 and 32 ask for, and it makes others possible — your own logs, audits whenever you want, usage rules that do not change without going through you. On its own it is not enough: it has to be described and traced — who has access, over which channel, where the document is stored. The reverse also holds: several of the proofs the GDPR and the AI Act require can only be produced when the infrastructure is under the control of the party accountable for it.

The first step

Operational from week one.

A real use case, on your data, in production. Then it grows, week after week.

30 minutes video call €150 free July promotion
Start an operational trial

It starts with a session with our engagement expert. Your data stays yours, always.