NEWS
China Rejects U.S. Charge of Industrial-Scale AI Distillation
Beijing says a three-agency U.S. advisory on industrial-scale AI distillation nationalizes a contract fight.
China’s Commerce Ministry on September 9 rejected a U.S. cybersecurity advisory that named six Chinese AI firms for industrial-scale distillation of American models. A spokesperson said the charge had no factual or legal basis and warned of countermeasures if Washington used the fight to suppress Chinese companies.
The notice, dated September 8, came from the National Security Agency, the Cybersecurity and Infrastructure Security Agency, and the Federal Bureau of Investigation. It accused DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI of pulling billions of tokens from Claude, GPT, Gemini, and Grok since late 2024. U.S. labs had spent a year asking the government to treat contract-violating distillation as more than platform abuse. The three-agency paper did that, and Beijing answered the state move.
A Cyber Advisory Built on a Year of Lab Complaints
CISA’s joint cybersecurity advisory on distillation, tagged AA26-251A, says knowledge distillation is a recognized training method and then argues the named Chinese firms used it as the core of their model work, not a side tool. The agencies say the campaigns likely ran with Chinese government awareness and that they violated U.S. labs’ terms of use.
The NSA warning on frontier-model extraction puts the same point in plainer commercial language. Distillation, it says, lets Chinese models close the gap without paying the full bill for compute, electricity, and basic research that U.S. frontier labs carry. That is a cost argument wearing a cybersecurity label, and it is the part of the file that maps onto sanctions talk already in circulation this year.
China-based AI companies are illicitly distilling U.S. frontier AI capabilities. Read NSA’s new report, co-sealed with @FBI, @CISAgov, highlighting AI knowledge distillation, TTPs used, and recommended mitigations: https://t.co/IgBFOnjXeg pic.twitter.com/JUYg5I08GP
— NSA Cyber (@NSACyber) September 8, 2026
NSA Cyber posted the report on September 8 with a PDF of the full advisory. The post restates the illicit-distillation charge and points readers to the tactics and the recommended fixes. It does not add new counts beyond the paper.
What the September 8 Notice Accuses Each Firm Of
The advisory’s company pages are unusually specific for a public cyber note. DeepSeek, it says, ran organized campaigns from late 2024 to train R1 and V3, drawing on Claude 3.7 through Opus 4.1, Gemini 2.5 previews, GPT-4 through GPT-5, and Grok 4. The extracted skills listed include legal specialization, agentic functions, chain-of-thought writing, and supervised fine-tuning tricks. CISA also says DeepSeek’s widely cited $5.6 million training cost is misleading because it leaves out data taken through distillation.
Moonshot AI is tied to a campaign from mid-2025, including Claude Fable 5 data for Kimi-K3 and GPT-4o data for Kimi-K2, plus a long list of Claude, GPT, Gemini, and Grok variants used for software engineering, math, and reinforcement learning. Alibaba is accused of distilling Claude 4-family models and GPT-5 in late 2025 to lift Qwen on coding, customer dialogue, and agent workflows. MiniMax is described as pulling chain-of-thought, coding, and agent skills into M2, and as trying prompt injections so Claude Code would treat itself as a MiniMax product. StepFun is placed in late 2025 and early 2026 against Claude 4.5-class models and the GPT-5 family. By mid-2026, Z.AI is said to have distilled billions of tokens from GPT-5.5 and Claude Opus 4.8 for chain-of-thought training.
THE SIX FIRMS NAMED ON SEPTEMBER 8
| Firm | U.S. models named as sources | Campaign window in the advisory |
|---|---|---|
| DeepSeek | Claude, Gemini, GPT, Grok families | Late 2024 onward, including R1 and V3 |
| Moonshot AI | Claude including Fable 5, GPT, Gemini, Grok | Mid-2025 onward, including Kimi-K2 and Kimi-K3 |
| Alibaba | Claude 4 family, GPT-5 | Late 2025, Qwen improvements |
| MiniMax | Claude Code and Claude, Gemini, GPT-5 | Late 2025, M2 coding and agents |
| StepFun | Claude 4.5-class, GPT-5 family | Late 2025 to early 2026 |
| Z.AI | GPT-5.5, Claude Opus 4.8 | By mid-2026, chain-of-thought data |
Access, in the agencies’ telling, did not look like a single stolen weight file. Requests were spread across native APIs, remote clouds, and third-party aggregators that strip metadata. A gray market of API proxies, called transfer stations, is described as the way around geographic blocks. Bulk premium subscriptions shared across developer teams are listed as the cost trick. Advanced steps include chain-of-thought extraction, automatic failover when a path is blocked, and quality checks meant to spot models that have been quietly degraded as a defense.
Anthropic Counted 16 Million Exchanges Before Washington Did
The government paper did not invent the complaint. On February 23, 2026, Anthropic said DeepSeek, Moonshot, and MiniMax had generated over 16 million exchanges with Claude through about 24,000 fraudulent accounts, in breach of its terms and its regional blocks. Anthropic does not sell Claude commercially in China, so every one of those accounts was, on its telling, unauthorized.
ANTHROPIC’S FEBRUARY CLAUDE TALLY
- DeepSeek: Over 150,000 exchanges aimed at reasoning, rubric grading that turned Claude into a reward model, and censorship-safe rewrites of political queries.
- Moonshot AI: Over 3.4 million exchanges on agentic reasoning, coding, computer-use agents, and later attempts to rebuild Claude’s reasoning traces.
- MiniMax: Over 13 million exchanges on agentic coding and tool orchestration, the largest of the three by a wide margin.
- Speed check: When Anthropic shipped a new model during MiniMax’s live campaign, the lab redirected nearly half its traffic within 24 hours.
Those three “over” figures add to about 16.55 million, which is how Anthropic could call the pool over 16 million without publishing a single combined total. The 16 million Claude exchanges Anthropic logged are a Claude-only, three-lab snapshot. They are not the same quantity as CISA’s later “billions of tokens” and “millions of exchanges” across four U.S. model families and six Chinese firms.
Anthropic also described the plumbing. Proxy services resell Claude access through what it called hydra clusters, sprawling nets of fake accounts across its API and third-party clouds. In one case, a single proxy network ran more than 20,000 fraudulent accounts at once and mixed distillation traffic with ordinary customer requests. That 20,000 figure is one proxy’s pool, not a second count of the 24,000 accounts tied to the three labs.
A prompt that looks harmless in isolation was part of the tell. Anthropic published an approximation of the kind of request it saw at scale: “You are an expert data analyst combining statistical rigor with deep domain knowledge. Your goal is to deliver data-driven insights, not summaries or visualizations, grounded in real data and supported by complete and transparent reasoning.” One such message is a user. Tens of thousands of close variants, aimed at the same skill, across hundreds of coordinated accounts, is a training harvest.
OpenAI had already taken a similar complaint to Congress. In a February memo to the House Select Committee on Strategic Competition between the United States and the Chinese Communist Party, the company said it had seen DeepSeek-linked accounts trying to get around access limits through obfuscated third-party routers, and it described “ongoing efforts to free-ride on the capabilities developed by OpenAI and other U.S. frontier labs.”
Poisoned Answers and Shared Intelligence Are the Defense
CISA Acting Director Nick Andersen used the release to tell U.S. labs to move now, not to wait for a statute.
We strongly urge AI companies to take immediate steps to safeguard their platforms against knowledge distillation campaigns that threaten to close the gap in advancements made by American companies.
Nick Andersen, CISA acting director, CISA press release, September 8, 2026
The paper’s to-do list is short and operational. It is also the clearest picture of what the agencies want private firms to do without new law.
THREE STEPS IN THE ADVISORY
- Detection: Watch anomalous prompts, accounts, and networks, plus subscription-to-usage ratios, new accounts that jump straight to maximum use, and enterprise-scale throughput.
- Poisoned replies: Quietly change answers for suspected distillation traffic so the stolen set is less useful, without wrecking the product for ordinary customers.
- Shared intel: Match activity across model providers, clouds, and API aggregators so a campaign split across vendors still shows up as one operation.
Anthropic says it is already building classifiers for chain-of-thought harvesting, tightening checks on education and startup accounts that get used as cover, and sharing indicators with other labs and authorities. The CISA note maps the same behavior onto the MITRE ATLAS acquire-infrastructure technique, which is how the agencies describe the transfer-station layer that buys frontier access at a discount and erases the buyer’s trail.
None of that requires distillation to be illegal as a method. It requires a lab to decide that a traffic pattern is hostile, then degrade or cut it. That is platform enforcement with a federal letterhead behind it, which is exactly the conversion Beijing chose to fight.
Beijing Calls Distillation a Neutral Tool, Then Draws a Line
The Commerce Ministry spokesperson, answering a question about the September 8 notice, did not spend the reply on token counts. The first move was to deny that the accusation had evidence or law behind it. The second was to call distillation a normal way for models to learn from one another, used by companies worldwide, including U.S. firms. The third was to treat the advisory as industrial policy.
Distillation is a common practice in the field of artificial intelligence for models to learn from each other. It is essentially a neutral technical method used by model companies worldwide, including those in the US.
Ministry of Commerce spokesperson, Beijing, September 9, 2026
The spokesperson said China encourages open-source sharing and that Chinese open models are available to global firms, including U.S. companies. U.S. model reports, the ministry added, show extensive distillation of Chinese models. That specific U.S. disclosure was not named. It remains Beijing’s claim, not a document quoted in the U.S. advisory.
Foreign Ministry spokesperson Mao Ning struck a parallel note the same day, saying China’s AI progress comes from self-reliance and that the United States should stop making unfounded accusations. She pointed to an understanding between the two heads of state on government-to-government AI talks and said the two sides should cooperate as major AI countries.
The ministry’s sharper line was about state power. Security agencies, it said, were tying the interests of individual companies and capital to national security and interfering in ordinary business. It also said some U.S. AI firms abuse their position with broad geographic limits and other unfair terms, and that the distillation charge looks like an endorsement of those terms. If the United States uses anti-distillation language as a pretext to suppress Chinese AI companies, China “will resolutely take countermeasures.”
The technique argument will not close this file, because the U.S. agencies already concede that distillation is a valid research method. What they are policing is disguised, high-volume extraction across APIs, clouds, and gray-market proxies. Replies to the NSA’s own notice went straight at that gap, asking whether American labs train on Chinese open models and how U.S. systems learned Chinese if other people’s data is off limits. That is the selectivity problem Beijing is trying to make the whole story, and it is why the ministry aimed at hegemony and compute monopoly rather than at student-teacher training math.
Sanctions Files, a House Probe, and a Late-September Summit
The advisory is a midpoint, not a first shot. Anthropic’s February post argued that illicit distillation undercuts export controls by letting labs close a gap those controls were meant to protect, and that running the harvest at scale still needs advanced chips. In other words, the same company that published the 16 million-exchange file was already asking Washington to treat distillation as a reason to keep the chip regime tight.
THE 2026 PATH TO THE ADVISORY
- Late 2024: CISA says organized distillation campaigns against U.S. frontier models begin, with DeepSeek among the earliest.
- February 12, 2026: OpenAI tells the House Select Committee that DeepSeek-linked accounts are routing around its access limits to obtain outputs for distillation.
- February 23, 2026: Anthropic publishes the three-lab Claude file and calls for industry and policy coordination.
- April 29, 2026: House Homeland Security Chairman Andrew Garbarino and House Select Committee on China Chairman John Moolenaar open a House investigation of PRC AI models, citing unauthorized distillation and low-cost open-weight systems from DeepSeek, Alibaba, Moonshot AI, and MiniMax.
- July 2026: Treasury Secretary Scott Bessent puts sanctions and Entity List designations on the table for covert industrial-scale distillation, saying open source is not open season on American IP.
- September 8, 2026: NSA, CISA, and the FBI issue AA26-251A naming six firms.
- September 9, 2026: The Commerce Ministry rejects the charge and warns of countermeasures, while repeating an offer of equal-footing AI talks.
The timing sits on a diplomatic calendar. U.S. officials issued the advisory as President Donald Trump prepares to host Chinese President Xi Jinping later in September. The ministry’s own close, after the threat, was that the two leaders have already agreed to an intergovernmental AI dialogue and that risks should be managed through talks, not unilateral action.
A cybersecurity advisory is not a listing, a tariff, or a criminal case. It is a public predicate. Once three agencies have called a commercial training method a malicious campaign at industrial scale, a later Entity List action or a round of sanctions can point back to a dated government paper instead of a private terms-of-service ticket. That is the step Beijing is trying to stop before the summit, and it is why a ministry of commerce, not only a foreign ministry, took the microphone.
Cheap Open Models Keep Arriving Anyway
The political fight has not paused product calendars. DeepSeek, the first firm named in the U.S. paper and the one whose $5.6 million training figure the agencies tried to knock down, posted DeepSeek-V4.1-Flash on September 10. The company called it the smallest model in a new architecture family, with native visual understanding, faster inference, and a path to larger systems. It did not address the advisory in that launch note.
Chinese labs have spent two years putting capable open weights into global hands at prices U.S. frontier APIs do not match. Washington’s new line is that some of that speed was bought with other people’s models, accessed through fake accounts and transfer stations. Beijing’s line is that student-teacher training is how the industry works, that Chinese open models already flow the other way, and that a security agency has no business closing a commercial channel for U.S. labs.
Both can be partly true at once. Distillation is ordinary. Running it through 24,000 fraudulent Claude accounts is not ordinary use. Folding that pattern into a three-agency advisory, weeks before a leaders’ meeting, is a policy choice. The Commerce Ministry has now said what happens if that choice turns into sanctions: China will answer in kind, and the AI talks the two presidents already endorsed become harder to hold.
-
NEWS1 month agoCSIRTs Inherit Europe’s Missing security.txt Before Article 14
-
NEWS4 weeks agoCity’s £125m Enzo Deal Caps a £327m Midfield Rebuild
-
NEWS1 month agoInstinct’s $2.5 Billion Raise Still Binds the User as Agent
-
NEWS1 month agoMeta’s Teen Settlement Leaves Chat Off the Clock
-
NEWS1 month agoOpenAI Codes a Persistent Agent the Week Persistence Backfired
-
BUSINESS2 months agoBerkshire Anchors Alphabet’s Record Raise With a $10 Billion Check
-
BUSINESS1 month agoTreasury’s First Iran Bank Shot Lands on an Ally
-
ENTERTAINMENT2 months agoRolex Made Drake the Daytona It Fights Jewelers Over
