Scale AI inside Meta: when everyone's supplier chooses a side
5 min read
Some acquisitions shift market share, and some acquisitions reveal how a market really works. Meta’s investment in Scale AI — $14.3 billion for a 49% stake, announced last June — belongs to the second category. Scale AI was the quiet supplier to half the industry: the company that recruits and coordinates hundreds of thousands of annotators to label, correct and refine the data used to train large models. It worked for Google, for OpenAI, for xAI, for the Pentagon. Then one of its clients became a 49% shareholder, and the founder moved across to lead its technical direction. Within days, the other clients walked away. Eight months on, the picture is clear enough to draw the lessons from it.
The facts, in order
- November 2024: Meta opens the Llama models to US national security agencies and their contractors, removing the ban on military use from its licence terms. The same day, Scale AI unveils Defense Llama, a version of Llama 3 fine-tuned for operational planning and intelligence, distributed through its government platform. The two companies’ paths begin to intertwine.
- 12–13 June 2025: Meta invests $14.3 billion for 49% of Scale AI (without voting rights), valuing it at roughly $29 billion. Founder Alexandr Wang steps down from running the company — staying on its board — and moves to Meta with a small group of executives; by the end of June, Zuckerberg puts him in charge of the new Superintelligence Labs as Chief AI Officer. Scale states that it will remain independent and that client data remains protected.
- 14–18 June 2025: Google — the largest client, with roughly $200 million in expected annual spend — suspends projects within hours and prepares its exit; OpenAI confirms it is winding down the relationship; xAI also freezes activity. The reason given is always the same: whoever prepares your training data sees your priorities, and that company is now half-owned by a competitor.
- 16 July 2025: Scale AI cuts 14% of its staff (200 people, plus 500 external contractors), mostly in labelling. Interim CEO Jason Droege cites over-rapid growth and announces a repositioning towards enterprise and public-sector applications.
It should be said: neither party breached any contract, and no improper data transfers have been documented. Meta describes it as an industrial investment in a strategic supplier; Scale points to robust internal barriers. But the market did not wait for proof: it priced in the structural risk, and walked away.
Lesson 1: training data is the contested asset, not the models
Models get copied, distilled, and are outdated within six months. Curated data — examples selected, annotated and corrected by experts — is the slow, costly part of the pipeline, the part that cannot be replicated with a download. Meta did not buy algorithms: it bought the factory that produces the raw material, and the person who had built it. For any business adopting AI, the translation is direct: defensible value does not lie in the model you use, but in your operational data and the preparation work you have invested in it. It is why our platform is built so that this asset remains the client’s property, in open formats: it is the one piece that no supplier should ever be able to walk off with.
Lesson 2: a supplier’s neutrality does not survive a change of ownership
Google and OpenAI did not leave because Scale had done anything wrong: they left because its ownership structure had changed. A supplier that serves everyone is neutral only for as long as it belongs to no one; the day a competitor takes a stake, every shared pipeline becomes a potential observation channel — which data you label, in what volumes, with what priorities. Contractual assurances existed; they were not enough. Anyone buying critical technology should treat a change of control at a supplier as a contractual event in its own right: a right of withdrawal, verifiable segregation obligations, an exit plan already drafted. Afterwards, you negotiate from a position of weakness.
Lesson 3: a supplier’s trajectory is part of the risk
The case also has a geopolitical dimension. With Defense Llama and the opening of its models to US security agencies, the Scale–Meta pipeline has become bound to a specific national mission. That is a legitimate choice — and for US government clients, an advantage. But for a European business or public administration, it means that your data or model supplier can shift its strategic centre of gravity without asking permission. Operational sovereignty is not a slogan: it means knowing where the data resides, under which jurisdiction, and what happens if the supplier’s ownership or mission changes. These are questions to ask before signing, not after the announcement.
What to do, if you depend on data or model suppliers
- Map what the supplier sees: datasets, volumes, annotation priorities. These are metadata that reveal your roadmap.
- Insert a change-of-control clause: withdrawal, data return, segregation audits — with defined timeframes and formats.
- Keep ontology and curated data in-house: the preparation work must remain your own asset, exportable and documented.
- Avoid single-supplier dependency where the risk is systemic: the queue of clients leaving Scale showed just how quickly an alternative can become necessary.
- Assess jurisdiction, not just price: ownership, data location, exposure to third-country security missions.
Want to measure how much of your AI supply chain depends on suppliers you do not control? Half an hour with one of our experts for an initial concentration-risk map.