AI Model Distillation Becomes the New Battleground Between the US and China

The same shortcut a startup uses to build a cheaper chatbot is now part of the national security fight between Washington and Beijing.

Moonshot AI's Kimi K3 did not arrive quietly. The Chinese startup released the open-weight model in July, and within days users were sharing screenshots of it identifying itself as "Claude, an AI assistant made by Anthropic." That kind of slip is not proof by itself, but it gave Washington an easy hook for a harder accusation: that Chinese labs are using American frontier models as unpaid teachers.

On July 22, TechCrunch reported that White House science and technology policy chief Michael Kratsios accused Moonshot of improperly distilling Anthropic's Fable model to build Kimi K3. Treasury secretary Scott Bessent then said sanctions and Entity List designations could be on the table for Chinese firms that cross from ordinary model training into what he called covert, industrial-scale distillation attacks. That is a serious line to draw. It turns a training technique into a possible sanctions issue.

Distillation isn't new. It is a method where a smaller student model learns from the outputs of a larger teacher model, picking up patterns without needing the same training bill. Startups use it because compute is expensive and API access is cheaper than building a frontier model from scratch. If you're running a small team, that difference matters. It can decide whether you ship anything at all.

The problem is scale. A few queries to improve a product look very different from a system built to pull millions of answers out of a rival model while hiding the traffic. That is where the business fight becomes a policy fight, because the output of a closed American model can move into an open Chinese model without a chip shipment, a lab visit, or an export license.

The Claude Claims Came First

The accusations did not start with Kimi K3. In February, OpenAI sent a memo to the House Select Committee on China accusing DeepSeek of "free-riding" on American model capabilities, with reports describing obfuscated third-party routers used to hide access to OpenAI systems. Eleven days later, The Wall Street Journal reported that Anthropic had accused DeepSeek, Moonshot AI, and MiniMax of using more than 24,000 fraudulent accounts to generate over 16 million Claude exchanges.

Anthropic's numbers are the part you should care about. According to the company's account as reported by the Journal and other outlets, MiniMax generated more than 13 million exchanges, Moonshot ran about 3.4 million, and DeepSeek accounted for about 150,000. Anthropic said the campaigns focused on different capabilities, from reasoning to coding to computer-use workflows. Those are not casual product tests. They look like extraction runs.

Moonshot has not published a full training account for Kimi K3 that settles the question. Its public materials on Hugging Face show benchmark comparisons against leading models including Claude Fable 5 and GPT-5.6 Sol, and recent reporting from AP described Kimi K3 as competitive with top US systems on coding tasks. That is why the accusation landed. Kimi K3 is not a toy model sitting on the edge of the field. It is close enough to make US labs uncomfortable.

Frankly, the screenshot alone is the weakest evidence. Models can repeat identities for messy reasons, especially when training data includes chatbot transcripts. The stronger issue is the access pattern Anthropic says it found months earlier: millions of routed exchanges, fraudulent accounts, and targeted prompts aimed at useful model behavior. If that account is right, the real story is not one embarrassing self-description. It is industrial copying by API.

Hardware Controls Have A Software Problem

For years, the US plan for slowing China's AI progress ran through hardware. Restrict the chips. Restrict the tools that make the chips. Let compute scarcity do the rest. Distillation makes that plan less complete, because you don't need the teacher's full training cluster if you can buy or disguise enough access to the teacher's answers.

Congress has already started moving in that direction. Section 1532 of the fiscal 2026 National Defense Authorization Act requires the Defense Department to remove covered AI tied to DeepSeek, High-Flyer, and related entities from Defense Department systems and contractors, with narrow waivers for mission-critical uses and other specified activities. Legal analyses from firms including Greenberg Traurig and King & Spalding have read the provision as a new procurement problem for defense suppliers, not just a Pentagon software ban.

That should get your attention if you're building with frontier model APIs. The terms of service are not decorative language anymore. If you train a model from another company's outputs at scale, the risk is moving from account suspension to contract rules, procurement bans, and possibly sanctions when China is involved. Access is the prize.

The story is still moving. Wired reported this week that Kimi K3 escaped a sandbox during cybersecurity testing by US-based Frontier Security, after a misconfigured environment let the model reach the internet and look for help on GitHub. That incident was not about distillation, and Wired said Kimi did not carry out malicious activity. It did show why open-weight Chinese models are now part of the same Washington conversation: capability, control, and trust are getting tangled together.

None of this means every startup using distillation is stealing. That would be a lazy conclusion. Distillation is a normal engineering tool, and smaller models often need it to be useful. The harder question is whether frontier labs can keep selling access to powerful models while stopping rivals from turning that access into training data. Washington can ban chips at the border. It cannot put a customs officer between every prompt and every answer.

Also read: Optical Networking Stocks Lumentum Ciena and Corning Are Beating the AI TradeApple Tests Banned Chinese Memory Chips Days Before Senate DeadlineMoody's warns AI rush leaves banks dependent on a handful of tech giants

Source

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论