Operational notes Observatory

DeepSeek V4 runs on a $4,000 PC: what changes for on-premise AI

7 min read

A processor with hundreds of gold pins, photographed close up against a soft backlight
A frontier model now fits inside a case under a desk: here, sovereignty is measured in memory, not server halls.

On 22 July 2026, Framework and AMD previewed a detail that matters more than any benchmark: a mini-ITX desktop with a single Ryzen AI Max+ PRO 495 processor and 192 GB of unified memory, running DeepSeek V4-Flash locally — a Chinese open-weight model, MIT-licensed, 284 billion parameters — without a rack, a cluster or a server hall. Until a few days ago, “frontier open-weight model” meant the 2.8 trillion parameters of Kimi K3 and the dozens of accelerators needed just to load it into memory. Now the same class of model fits inside a case that sits under a desk. For anyone weighing on-premise AI in an enterprise or a public body, the question changes shape: no longer “how many millions does this need”, but “what are the real numbers, and what risks does the hardware alone not solve”.

The facts, in order

  • Before 22 July 2026 — DeepSeek published DeepSeek-V4-Flash and DeepSeek-V4-Pro on Hugging Face: a Mixture-of-Experts architecture, MIT licence with no revenue threshold, unlike the modified licence under which Mistral limits free self-hosting to companies below a certain size. According to the model card, V4-Flash has 284 billion total parameters and 13 billion active per token — 256 routed experts plus one shared expert, six active at a time — with a context window of up to 1,048,576 tokens, roughly one million.
  • 22-23 July 2026 — Framework and AMD previewed a Framework Desktop configuration with a Ryzen AI Max+ PRO 495 APU and 192 GB of unified memory: in the demonstration, covered by specialist outlets including Wccftech, VideoCardz and Notebookcheck, the machine runs DeepSeek V4-Flash locally at 8-bit quantisation. Independent throughput figures (tokens per second) have not yet been published. This is a preview, not a shipping product: the price and availability date of this specific configuration have not yet been announced. The previous configuration in the same line, with 128 GB of memory, was listed from $3,999.
  • DeepSeek’s cloud price list, on the official API documentation, remains among the lowest on the market: for V4-Flash, $0.0028 per million tokens on a cache hit, $0.14 on a cache miss, $0.28 for output; for V4-Pro, $0.003625, $0.435 and $0.87 respectively.
  • The regulatory backdrop is moving in the opposite direction from technical availability: Italy’s data protection authority, the Garante, had already blocked the DeepSeek app in January 2025 and opened an inquiry into its data handling; Belgium banned the platform for public officials in late 2025; several US states had already excluded it from government devices during the same year. None of these restrictions touch the licence, though: the weights remain downloadable and can be run entirely offline, outside any service hosted in China.

The sum almost nobody does: tokens or capital?

Faced with a model that runs on a PC rather than a cluster, the temptation is to conclude that on-premise finally wins on price too. The arithmetic says otherwise. At DeepSeek’s official price list, $3,999 — the price of the previous configuration in the same Framework line — buys over 14 billion output tokens of V4-Flash on the cloud service. That is a volume the vast majority of organisations will not reach in a full year of use, even with intensive usage. If the criterion is cost per token, the cloud remains almost always cheaper: local hardware adds power, maintenance, space, and depreciates. The case for bringing a model like this in-house was never about the price of a single token. It is about not sending your prompts and documents to a service hosted outside your own perimeter, having control over which version of the model runs and when it changes, and being able to operate without a continuous internet connection — conditions that other vendors sell as “sovereignty” inside a commercial contract, and that here come free with the licence instead.

A word on how narrow that sum is, because it needs saying. The comparison above sets a desktop machine for a single user against the lowest price list on the market: DeepSeek’s, which this same note lists among the cheapest in circulation. Within those bounds the cloud wins, and that stays true. It is not, however, the sum a company does, because a company does not pay DeepSeek: it pays a closed frontier model, where the same operations cost up to ten times as much, by consumption, on a price list it does not negotiate and that rises with use without a ceiling. Change the comparator and the scale and the sign of the result flips: open weights on hardware you own take the cost per token to zero and leave power, maintenance and depreciation, which are forecastable lines. And the most expensive line sits outside the calculation entirely, because it never reaches the invoice: paying by the token hands the supplier which questions the business asks and against which documents. The argument in full is on the hardware and open-weight models page.

What the hardware does not solve

A PC costing a few thousand dollars lowers the barrier to entry; it does not remove the risks. The MIT licence carries no support, no independent security audit of the published weights, no guarantee of the model’s behaviour in production: whoever adopts it inherits the responsibility of verifying it, exactly as with the other Chinese open-weight models already under evaluation at European companies. And local execution alone does not resolve the most awkward procurement constraint: several administrations and regulated sectors exclude suppliers based on the country of origin of the model’s developer, not only on where inference actually runs. A specification that bans “Chinese cloud services” may not explicitly ban “Chinese weights run on your own servers” — but an attentive reviewer will notice, and the clause can still arrive, in a public tender or in an enterprise customer’s audit. Checking this before investing in hardware costs an email to your legal team; discovering it afterwards costs an entire project.

What to do, in practice

  • Do not justify on-premise deployment on token savings alone: calculate the real cost — hardware, power, maintenance — against your actual usage volume, not against an extreme case.
  • Wait for official pricing and availability of the 192 GB configuration before treating it as a purchase plan: today it is still a preview.
  • Before adopting a Chinese-origin open-weight model, check your organisation’s or sector’s procurement rules on “country of origin of the supplier”, not only on where the infrastructure sits.
  • Treat a permissive licence as a starting point, not a guarantee: plan your own security and quality testing, because no supplier will do it for you.
  • If the use case handles classified or particularly sensitive data, isolate the environment with a genuine air gap — not merely the absence of a default connection — and verify that no optional telemetry function calls back out.

A frontier model running inside a case under a desk is a real piece of news, and it changes the arithmetic for many organisations that had previously ruled out on-premise deployment on grounds of scale. It does not, on its own, change the discipline that every supplier — European, American or Chinese — should face before entering production: licence verification, testing on your own data, traceability of the dependency chain. It is the same logic behind an architecture that treats every model as a replaceable component, not a permanent commitment.

Weighing an open-weight model for a use case with real data-control requirements? Let’s talk in a thirty-minute session: we’ll map out cost, licence and risk together before you have to do it in production.

Sources