Ling-3.0-tiny: MIT licence declared in the tags, no licence file in the weights repository
7 min read
inclusionAI/Ling-3.0-tiny appeared on Hugging Face on 10 August 2026: the API’s createdAt reads 2026-08-10T02:44:06Z, last modified 2026-08-18T12:16:27Z, the day before our check. We queried it ourselves on 19 August: 9,990 downloads, 321 likes, 43 files in the repository, of which 32 .safetensors shards. Declared tags: safetensors, bailing_hybrid, text-generation, conversational, custom_code, license:mit, region:us. inclusionAI is the open-source brand of Ant Group — the Alipay company, part-owned by Alibaba but distinct from the team that publishes Qwen — for its own research effort: the family is called Ling, heir to the earlier internal name «Bailing», which survives in the declared architecture, BailingMoeV3ForCausalLM.
A model built to run outside the datacentre
The README declares 7.9 billion total parameters and 1.4 billion activated per token (1.14 billion excluding embeddings): an MoE with 128 routed experts, 8 activated per token plus 1 shared expert, across 24 layers. config.json confirms hidden_size 1536, 16 attention heads, max_position_embeddings 131072 — 128K native, which the card says YaRN extends to a “256K YaRN context”. The usedStorage the API declares is 15.8 GB, consistent with the parameter count in BF16.
The card doesn’t present it as a cluster model: it states it was «validated on NVIDIA DGX Spark, Apple Silicon MacBook, and Mac mini», with 100-105 tokens/s on DGX Spark and 86-90 tokens/s on an M4 Pro MacBook in FP8, roughly 8.34 GiB peak memory at 8K context — the lab’s own numbers, not measured by us. The point here isn’t the declared speed: it’s the hardware threshold, low enough that a technical team no longer needs a datacentre GPU budget to bring an agentic-class model inside its own perimeter.
The declared licence, the file that isn’t there
Here begins the part we checked line by line. The API’s cardData field reports "license": "mit"; the license:mit tag appears among the public ones. We looked for the file that normally accompanies that declaration — LICENSE, LICENSE.md, LICENCE, License.txt, in several capitalisation variants: HTTP 404 on every one. Among the 43 files the API lists, none is named «licence» in any form: only README.md, the configuration files, the 32 shards and the tokenizer files. The «mit» declaration exists; the text that substantiates it — in the weights repository — does not.
This isn’t the case of DeepSeek-V4-Pro-0813, where the LICENSE file was there in full. Nor is it the opposite case, of a download with no licence reference at all: here the licence is declared, under a precise name. It’s a third case: the label is there, the document that ought to sit behind the label is missing from the repository that declares it.
The full text exists, but elsewhere: in the code repository inclusionAI/Ling-V2 on GitHub — a different project, not this model’s weights — a LICENSE file carries the full MIT License text, headed «Copyright (c) 2025 inclusionAI»: «Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the “Software”), to deal in the Software without restriction, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software». It is not, however, the document attached to the weights repository: it’s the text of a different project. Whoever archives this model has, today, no licence file to keep alongside the weights: they have a tag, and a reconstruction.
We also looked for the second kind of surprise we often find: a usage policy, a NOTICE file, an attachment that narrows what the tag promises. NOTICE and USAGE_POLICY.md return 404 just as LICENSE does, and the README’s text contains none of the phrases that usually flag a constraint — «usage policy», «acceptable use», «shall not», revenue or user thresholds. There is no second document, here or anywhere else in the repository.
Access and channels: gated false, verified without credentials
The gated field is false, verified without authenticating: config.json returns a 307 redirect to the public cache, not a 401; the API responds 200. The model is also reachable via a free hosted endpoint on OpenRouter (inclusionai/ling-3.0-tiny:free) and through ModelScope: we have not checked whether the hosted version offers functions absent from the downloadable weights, nor the reverse.
An operational detail: Ollama support is not in the official release — it lives in a pull request not yet merged (ollama/ollama#17643), requires a build from source and is, verbatim, «limited to running via MLX on Apple Silicon», verified on an M4 Pro Mac with 48 GB of unified memory. Anyone reading «open weights, runs on a Mac mini» and expecting an ollama pull command is conflating two distinct stages: the weights are downloadable today, the tool many teams would use to run them doesn’t officially support them yet.
We also tried the links cited on the card — the SGLang cookbook, the GitHub pull request, the hosted images, the OpenRouter page, the ModelScope organisation: all return 200. No phantom document this time: the check is worth doing regardless, even when the result is negative.
The perimeter, not just the download
Downloadable weights without the compute that makes them runnable aren’t sovereignty: here the compute required is unusually low, and that’s exactly why the reasoning flips. If an agentic model runs on a Mac mini, the threshold that used to hold back an unauthorised download inside a department — GPUs, budget, approval — drops close to zero. The model-register problem doesn’t ease: it gets worse, because more people inside the organisation can download and run a model without anyone knowing. The same applies to the broader choice between Western and Chinese open-weight models: the question is never just «does it work», it’s «who downloaded it, when, under which archived licence».
That’s why the licence line on its own isn’t enough: you need the artefact’s version and digest, the channel and date of acquisition, the licence archived in the text of that day — not recalled from memory months later, when it may have changed or, as here, when the file behind the tag simply doesn’t exist — and the approved uses, with who approved them.
What we don’t know
We have not tried the model: no claim about quality, behaviour or real-world reliability. The card declares a score of 25 on the Artificial Analysis Intelligence Index and 16 on the Agentic Index: figures from the card, not verified by us and not reported as our own. We don’t know the minimum hardware for production use beyond what’s declared for short contexts; we don’t know what data it was trained on or who answers for that; we make no legal qualification, AI Act included. We don’t know whether the licence will stay this way, nor whether a LICENSE file will appear in future: what we report is the state verified on 19 August 2026.
See the service · Talk to an engineer
The two axes, applied to this case
Complying: the register of models in production becomes a check that runs on the client’s own systems — for each model, the artefact’s version and digest, channel and date of acquisition, licence archived in the text of the day it was downloaded (here: an «mit» tag, not a file), approved uses and who approved them. With the dated record ready to show an inspector, a client in a tender, or a board.
Deciding: the same system brings models, data, contracts, archives and documents together into a single operating model, on which AI agents execute decisions with a human operator in command, for large enterprises, defence, government and healthcare. The lower the hardware threshold for running a model drops, the more an organisation needs a system that always knows what’s running where, under which archived licence and who authorised it — not a ban nobody follows, a register everyone can consult. Always in two modes, on-premise on the client’s own self-contained machines, or dedicated cloud with a dedicated VPN and a data centre in Italy, always with shared management.
From the first session, at no cost, comes the dated register of the models you have in service: for each one, how it was obtained, whether the channel is still open, what you have kept — including the boxes that stay blank. It stays yours even if we don’t go on to work together.