Operational notes Observatory

Kimi K3: Microsoft is testing it for Copilot. What changes for decision-makers

6 min read

A storage module pulled from a server rack in a data centre hall
The real cost of a 2.8-trillion-parameter “open” model is measured in racks, not in licence fees.

On 16 July 2026 Moonshot AI put Kimi K3 online: 2.8 trillion parameters, announced as the largest model ever released with open weights. Five days later, on 21 July, came the more interesting news of the launch itself: Microsoft is reportedly testing Kimi K3 inside Copilot, with the stated aim of cutting up to $600 million in inference costs by shifting workloads currently handled by OpenAI and Anthropic. If the world’s largest enterprise buyer of AI models is the one running the evaluation, the question for decision-makers at European companies is no longer whether Chinese open-weight models have matured. It is under what conditions.

The facts, in order

  • 16 July 2026 — Moonshot releases Kimi K3 via app and API: a Mixture-of-Experts architecture with Kimi Delta Attention (KDA) and Attention Residuals, 896 experts of which 16 are active per token, a 1-million-token context window, native visual understanding. The full weights are expected by 27 July 2026; at launch the model is available only through a hosted service, with API pricing of $0.30 per million tokens on a cache hit, $3 on a cache miss, and $15 on output.
  • 17 July 2026 — International press reports that Kimi K3 matches or exceeds leading American models on several tool-use, web search and agentic coding benchmarks, while lagging on others (general knowledge, some of the more demanding coding tests). No independent third-party verification is available yet: the figures circulating so far are those published by Moonshot itself.
  • 20–21 July 2026 — Moonshot temporarily suspends new sign-ups to the service: “demand has pushed our capacity close to the limit,” the company writes, adding that existing subscribers are unaffected and that it will reopen “in batches” as soon as additional capacity is added.
  • 21 July 2026 — It emerges that Microsoft has added Kimi K3 to Azure and is evaluating whether to use it for some Copilot functions currently based on OpenAI and Anthropic models, with an estimated saving of up to $600 million in inference costs. Microsoft clarifies that this is a test, not a decision to replace anything: the model still has to meet the product’s quality, reliability, security and latency requirements. Microsoft has previously evaluated DeepSeek and earlier versions of Kimi: the stated strategy is a mix of vendors, not a single supplier.

The signal isn’t the model — it’s who is evaluating it

Chinese companies such as DeepSeek, Qwen and GLM have already entered the evaluations of many European businesses for cost reasons. What’s different about Kimi K3 is this: here, the one running the numbers is a company that needs to save money on an API less than almost anyone else — and it is doing so anyway, because the price gap has become wide enough to justify the test. It’s a market signal, not a green light. Microsoft itself says as much: quality, reliability, security and latency are checked first — only then, perhaps, comes a decision. That is the discipline that often gets skipped when a company gets excited about a benchmark headline rather than a test on its own data.

The on-premise arithmetic, before you believe it

A 2.8-trillion-parameter model, even compressed to 4 bits, weighs around 1.4 terabytes just to load the weights into memory. No single GPU, however advanced, comes close to that figure: Moonshot itself recommends configurations with 64 or more accelerators spread across multiple nodes, and that estimate does not yet include the cache and activations for a 1-million-token context, which push the figure considerably higher. All of this, it should be said, is pre-release guidance: no validated configuration exists until the weights are public. The point for decision-makers remains the same: “weights you can download for free” does not equal “cheap to run.” The hardware bill to host the flagship model in-house is a capital investment on the scale a hyperscaler can absorb, not a mid-sized company. For most organisations the sensible choice remains what it has been for months: smaller or quantised variants for narrowly defined use cases, not the entire flagship model, inside an architecture that treats the model as a replaceable component.

A licence that doesn’t exist yet

At launch, Moonshot had not published the definitive licence text for Kimi K3. Its predecessor, Kimi K2, used a modified MIT-style licence with its own commercial conditions. Until the text is public, any production decision is provisional by definition. And even once the licence is published, an open checkpoint is not automatically an open operating environment: reproducing the behaviour Moonshot has claimed requires the same serving stack — specific attention kernels, the prompt format needed to exploit the cache, and management of the reasoning trace across turns. For anyone using the hosted version, this means negotiating in writing — not trusting the product page alone — data use for model improvement, retention periods, storage location and audit rights: the same discipline terms of use always require, whatever the supplier.

The political risk doesn’t wait for your decision

In Washington, export-style restrictions on Chinese open-weight models are already under discussion; Microsoft’s own test has been described as a possible target of political attention. For a European company the risk stacks on two fronts: not only “Beijing will restrict access to its best models,” but also “Washington will restrict the use of Chinese ones,” especially if they run through a US supplier’s infrastructure. This is the kind of dependency to write into a specification before signing, not to discover after the next political announcement.

What to do, in practice

  • If you are evaluating Kimi K3 for use cases involving confidential or personal data, wait for the definitive licence text and have it checked by counsel before any production use.
  • Don’t confuse “free weights” with “guaranteed savings”: calculate the real hardware spend — aggregate memory, nodes, bandwidth — before considering on-premise deployment cost-effective.
  • To start, use smaller or quantised variants on a narrowly defined, verifiable use case, not the entire flagship model.
  • If you use the hosted version, negotiate data use, retention, storage location and audit rights in writing.
  • Monitor US–China regulatory developments: an export restriction can change the availability of a model you have already adopted from one day to the next.

The Kimi K3 case does not settle the question of whether Chinese open-weight models are worthwhile for a European company. It reopens it more seriously, because now even those with the least urgency to save money are asking it too. The answer, for your organisation, depends on data no benchmark publishes: what your hardware really costs, what your licence says in full, who else in your supply chain depends on the same supplier.

Evaluating an open-weight model for a concrete use case and want the numbers worked out properly before you commit? Let’s talk in a thirty-minute session.

Sources