Top US Science Advisor Accuses Chinese AI Company Moonshot of Copying Anthropic’s Model Using Banned Chips
By admin | Jul 23, 2026 | 4 min read
White House science advisor Michael Kratsios has accused the Chinese company Moonshot of developing its Kimi K3 model—currently the largest openly available open-weight large language model—by copying Anthropic’s Fable LLM, while also using chips that are prohibited from export to China. “Large-scale, covert industrial distillation aimed at stealing proprietary U.S. technology and undermining American research is unacceptable,” Kratsios wrote, as discussions about potentially banning Chinese open-weight models have stirred significant controversy in the AI industry. Moonshot did not respond to inquiries regarding its training methods, and Kratsios did not provide additional evidence to support his claims.
Kratsios’ statement echoed earlier comments from Treasury Secretary Scott Bessent, who remarked, “We are finding watermarks of our U.S. large language models on many of the Chinese models, and that’s unacceptable.” The exact nature of these watermarks remains unclear, and the Treasury Department did not respond to a request for clarification. However, experts have expressed doubt that distillation—the process of querying an LLM to understand its internal workings and replicate its capabilities—is responsible for the advanced performance of Kimi K3. As one researcher noted, “There’s just not even frankly time, right. Fable’s only been publicly available since July 1st. You can’t distill that much data, train a model, and release it in two weeks.”
Nathan Lambert, an AI researcher at the Allen Institute for AI, shared his perspective in a recent podcast: “I’ve been of the opinion that distillation has become less and less impactful over time as the Chinese models get closer to the frontier and the training regime shifts to reinforcement learning. If it were the case, everyone would be easily able to catch up to a GLM or to a K3 by using its data for distillation. But we have not, or we won’t see this, from supervised fine-tuning alone.”
Distillation involves systematically querying a target model to generate data for post-training. This can include asking the model to explain its chain-of-thought reasoning or using its prompts and responses to train a new model through supervised fine-tuning (SFT). It is during this fine-tuning process that a model might appear to be a clone of another, like Claude. According to Lambert, fine-tuning is where “the model picks up its manners.”
Yet Lambert argues that the advantages of SFT are diminishing as models grow more complex. Replicating capabilities similar to Fable would likely require reinforcement learning techniques, which often involve a larger model grading the responses of a smaller one and adjusting accordingly. These advanced methods demand substantial infrastructure; large reinforcement learning runs may need tens of millions of agents. Using a frontier lab’s API for such purposes “would be insanely expensive and potentially it would probably be a time bottleneck because these models are pretty slow and to be frank might not even give you a performance uplift.”
It is plausible that earlier frontier models contributed to Kimi’s development. Earlier this year, Anthropic publicly accused Moonshot, DeepSeek, and MiniMax of systematically distilling its models, claiming it discovered millions of exchanges between its models and users from those companies, identified through IP addresses and metadata. These queries were “distinct from normal usage patterns, reflecting deliberate capability extraction rather than legitimate use.” However, distillation is widely practiced across the AI industry, not just in China. Elon Musk testified earlier this year that his company SpaceXAI distilled OpenAI models to develop Grok, calling the practice common. The line between distillation and creating synthetic datasets can be quite blurry. As researcher Hancock observed, “In general, Americans are understating the technical expertise of these Chinese teams. One of the founders of Moonshot was a CMU PhD student. These are legitimate researchers and engineers doing solid work. If American models ground to a halt, I think China’s progress would slow, but would still continue. They’re not just riding coattails here.”
Separating distillation from the other part of Kratsios’ accusation—that Moonshot obtained advanced Nvidia Grace Blackwell 300 chips and accessed GB300-equipped servers in Thailand—is also challenging. These chips are banned from export to China, but a black market exists, according to Sam Bresnick, a research fellow at Georgetown’s Center for Security and Emerging Technology. In May, the founder of Supermicro, a U.S. server builder, was indicted for smuggling advanced chips into China. Bresnick stated, “I am a proponent of know your customer laws for data centers across the world. If you are letting a company conduct huge training runs on your state-of-the-art hardware, there needs to be a reporting mechanism for who that company is and what they’re doing.”
In 2024, President Joe Biden’s Department of Commerce proposed federal know-your-customer rules for data centers, but no further progress appears to have been made under Donald Trump. Exporters shipping advanced chips abroad are still expected to ensure they are only used for approved purposes.
Comments
Please log in to leave a comment.
No comments yet. Be the first to comment!