AI Model Report

Open Source · SEPTEMBER 17, 2026

US pushes to cut Chinese labs off from frontier APIs — and the cheap-inference tier hangs in the balance

A Sept. 8 NSA/CISA/FBI advisory names six Chinese labs distilling Claude, GPT, Gemini, and Grok at industrial scale. If Washington closes the pipes, the price competition that made AI viable for small operators loses its floor.

By Lars Iverson · Open source & model weights · September 17, 2026

On Sept. 8, the NSA, CISA, and FBI jointly published Advisory AA26-251A naming six Chinese labs, DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI, as running "industrial-scale" knowledge-distillation campaigns against Claude, GPT, Gemini, and Grok since late 2024. The advisory characterizes the extraction, "billions of tokens across millions of exchanges," as "the core — not merely a supplement — of their AI development strategy," and recommends U.S. labs "subtly alter responses for suspected malicious distillation." The Register read the document, correctly, as a predicate for further U.S. action.

The market read it the same way. Bloomberg reports Z.AI and MiniMax shares have each fallen more than 30% this month, erasing roughly $33 billion in combined value.

The numerical spine of the U.S. case was assembled in February, when Anthropic disclosed over 16 million exchanges across three campaigns using roughly 24,000 fraudulent accounts: 150,000+ tied to DeepSeek, 3.4 million to Moonshot AI, and 13 million to MiniMax, which "pivoted within 24 hours" to a newly released Claude model mid-campaign through a proxy network of "more than 20,000 fraudulent accounts simultaneously." The advisory says chain-of-thought traces, agentic tool use, coding behavior, and SFT/RL training data were all extracted, likely with Chinese government awareness.

Beijing's response has been notably flat. China's foreign ministry called the warnings "fear-mongering." Chinese labs contacted by CNBC declined to comment. Alex Lu of LSY Consulting dismissed Anthropic's framing as "marketing communication," which is both a fair read of the incentives and irrelevant to the policy vector already in motion.

Here's the piece worth sitting with. The sub-cent-per-million-token pricing tier that a five-to-thirty-person business now takes for granted, DeepSeek V4.1-Flash at $0.15/M off-peak, Z.AI's 320B-A18B GLM-5.3-Flash, and architectures like DeepSeek's causal encoder-decoder that keep dropping agentic inference costs, exists because Chinese labs have been racing U.S. frontier prices down. If the U.S. case establishes that this capability was bootstrapped on distilled outputs, and the response is API cutoffs, Entity List additions, or export restrictions, the competitive pressure that produced those prices is what erodes first.

That's not a prediction that inference bills will spike next month. It's an observation about which layer of the stack is actually load-bearing. The advisory doesn't need to succeed on its narrower legal claims to succeed on the pricing one. It needs only to make continued access to U.S. frontier outputs a policy question rather than a market one. On Sept. 8, it did.

Sources