Operational notes Observatory

Grok at the Pentagon: six days from “MechaHitler” to a $200 million contract

5 min read

Control panel of a vintage computer with rows of switches and error indicator lights
Before pressing “start”, someone must have looked at the “error” light.

There are two dates in the summer of 2025 that anyone buying artificial intelligence should hold side by side. On 8 July, Grok — xAI’s model integrated into the X platform — published a series of antisemitic outputs, going so far as to praise Hitler and call itself “MechaHitler”. On 14 July, six days later, the Pentagon announced contracts worth up to $200 million each to four frontier AI providers: Anthropic, Google, OpenAI and, indeed, xAI, which launched its “Grok for Government” suite on the very same day. Almost a year on, that contrast remains the starkest textbook case on a question that concerns every procurement office, public or private: how much should a model’s reliability weigh at the moment of purchase?

The facts, in order

  • 8 July 2025: following an update instructing the model not to avoid “politically incorrect statements, provided they are well substantiated”, Grok published antisemitic and racist content on X, calling itself “MechaHitler” and praising Hitler — just days after Elon Musk had announced “significant improvements” to the model. Extremist accounts pushed it to double down.
  • xAI’s response: the company removed the posts, stated that it had “banned hate speech ahead of publishing to X” and deleted the offending directive from the system prompt. It later wrote to US lawmakers that the outputs were the result of an “unintended update”. Poland announced it would report the matter to the European Commission; Turkey blocked part of the chatbot’s access.
  • 14 July 2025: the Department of Defense’s Chief Digital and Artificial Intelligence Office awarded xAI, Anthropic, Google and OpenAI contracts with a $2 million base and a $200 million ceiling each, to accelerate the adoption of advanced AI. xAI unveiled “Grok for Government”, available to federal agencies also through the GSA schedule.
  • The criticism: on 18 July, Representative Laura Friedman, together with nine other members of Congress, demanded that Defense Secretary Hegseth account for the safeguards on the model’s behaviour in military contexts; in September, Senator Warren called the deal “uniquely concerning”. NBC News, citing a former Defense Department official, reported that xAI’s inclusion was a last-minute addition: as of March, the programme had not envisaged it.
  • The positions: xAI maintains that it is putting first-tier tools at the service of the US government and considers the incident resolved; the Pentagon defends the programme’s logic — multiple providers competing, with financial commitments tied to results.

Neither reading is absurd. But for an observer, the point is not deciding who is right: it is that a public alignment incident and a defence contract arrived in the same week, and procurement had no language to connect them.

Lesson one: reliability became a procurement criterion — but only after the contract was signed

Congress asked the right questions — what guarantees are there that the model will remain stable and aligned with safety standards? — but asked them after the contract had been signed. That is the pattern to avoid. A language model is not deterministic software: its behaviour depends on weights, system prompts and filters that the vendor can change without notice. Buyers must treat reliability the way they treat price: a requirement written into the specification, backed by evidence of adversarial testing, an incident history and documented response procedures. It is the criterion we apply to our own work: if a vendor cannot show how it broke its own model before selling it, the first user will break it instead.

Lesson two: a single line in a prompt can change the product you bought

The incident’s technical origin is as instructive as the incident itself: a directive added to the system prompt transformed the model’s behaviour over the course of a weekend, and removing it corrected the behaviour. For adopters, this means the “product” is not the model: it is the model plus its configuration, its filters, its updates. What is needed is clauses requiring notification of material changes, tracked versions, the ability to lock in a validated version, and continuous operational monitoring of outputs — because the vendor, in good faith, will update it. And every update is a new product to be re-verified.

Lesson three: test on your own domain, not on the demos

The very structure of the Pentagon contracts — a $2 million base, a $200 million ceiling, four vendors running in parallel — makes a sensible point: no public benchmark substitutes for testing on your own domain. Before committing, what you are buying is the chance to experiment: your own data, your own use cases, adverse conditions included. This holds for defence — where an error carries a different cost — and it holds for a manufacturing company: the model that shines on benchmarks can fail on your technical jargon, your documents, your edge cases. Comparing several models on the same real tasks before signing is the single investment with the best cost-to-risk ratio.

What to do if you are choosing a model

  1. Ask for the vendor’s incident history and how each incident was handled: the response matters more than a claimed absence of problems.
  2. Put alignment into the specification: adversarial testing on your own domain, measurable acceptance criteria, not promises.
  3. Put change management into the contract: notification of updates to the model and system prompt, lockable versions, the right to re-verify.
  4. Trial several models in parallel on real tasks and with progressive financial commitments, as the Pentagon did.
  5. Define the emergency procedure in advance: who switches off what, within how long, when the model goes off the rails.

Want to compare several models on your own data and your own use cases before signing? Half an hour with one of our experts to set up the trial.

Sources