When Anthropic’s chief national security officer and former Biden administration export control architect Tarun Chhabra recently highlighted model distillation as an emerging national security concern, one detail stood out. Chhabra identified Chinese AI developer Zhipu among the companies that have allegedly distilled capabilities from frontier American models. The warning fits neatly into Washington’s narrative: Chinese firms are racing to close the AI gap by leveraging innovations pioneered by Silicon Valley.

But only hours after Chhabra’s comments, that narrative became considerably more complicated. Mira Murati’s Thinking Machines — which raised $2 billion on the strength of her reputation as OpenAI’s former chief technology officer — revealed that its first foundation model was built in part using Chinese models. Inkling’s architecture drew on DeepSeek-V3, and its post-training process incorporated synthetic data generated by Moonshot AI’s Kimi K2.5. This is how the open-weight/source model world works: One company builds on top of another’s innovation.

The juxtaposition is striking. One of America’s leading frontier companies is warning that Chinese laboratories are learning from U.S. models at precisely the moment one of Silicon Valley’s newest flagship ventures openly acknowledges learning from Chinese ones. That apparent contradiction says far more about the state of frontier AI than either announcement.

American firms … are more willing to incorporate advances from China’s rapidly improving open-weight ecosystem.”

The word “distillation” has rapidly become one of the most politically charged terms in AI and now in geopolitics, driven by accusations from leading U.S. closed-source model developers that Chinese companies are able to stay close to the frontier by leveraging this practice. Increasingly, it is presented almost interchangeably with intellectual property theft. But distillation is neither new nor particularly exotic. It has been part of machine learning for more than a decade. A larger “teacher” model generates outputs that are then used to train a smaller “student” model capable of reproducing much of the original model’s behavior at a fraction of the computational cost.

Today, however, the technique has expanded well beyond simply compressing large models. Synthetic data generated by one model routinely becomes training material for another. Frontier laboratories increasingly rely on teacher models to improve reasoning, coding, multilingual performance, and post-training alignment. In many cases, the most valuable training data is no longer collected from humans at all — it is generated by other AI systems. This practice is not confined to China.