Experts say exploiting Anthropic’s Legend is not how Kimi K3 was good

White House science adviser Michael Kratsios said Moonshot, the Chinese company behind the Kimi K3, the world’s largest open-source LLM, built its model by copying Anthropic’s Fable LLM while using unwiped chips to ship to China.
“Large-scale, covert industrialization aimed at stealing American proprietary technology and undermining American research is unacceptable,” Kratsios wrote, amid reported discussions about curbing China’s open-minded models that have invaded the AI sector. Moonshot did not respond to questions about its training process, and Kratsios did not elaborate on the sources of his suspicions.
Kratsios’ tweet echoed Treasury Secretary Scott Bessent’s comments that “we are getting watermarks of our major US language brands on many Chinese models, and that is unacceptable.” It’s not clear what those watermarks include, and the Treasury Department did not respond to a query.
However, experts doubt that distillation-the process of questioning the LLM to determine its inner workings and copy its skills-is the cause of the advanced skills demonstrated by Kimi K3.
“I don’t think you get a model as strong as this and right on the heels of Fable doing the distillation so strongly,” Braden Hancock, a researcher at the Laude Institute and founder of Snorkel AI, told TechCrunch. “There’s not even a timeline, is there? Fable has been publicly available since July 1st. You can’t compile that much data, train a model, and release it in two weeks.”
“I was of the opinion that the distillation had little effect over time as the Chinese models approached the border and training regime. [reinforcement learning]”said Nathan Lambert, an AI researcher at the Allen Institute for AI, in a podcast released yesterday.”[I]if that were the case, everyone would easily be able to access GLM or K3 by using its filtering data. But we have never, nor will we ever see this, in a surveillance-only configuration.”
Distillation requires the lab to systematically query its target model to generate data that can be used for post-training. Sometimes this involves explicitly asking the model to reveal its own train of thought to understand how it solves problems. Sometimes, the information and feedback from the model is used to train a new model in a process called supervised fine-tuning, or SFT.
This fine-tuning process can lead to a model apparently created by a third-party company called Claude. Fine tuning is where, in Lambert’s opinion, “the model takes its habits.”
But Lambert says the benefits of SFT are becoming less and less important as the models become more complex. Practicing skills such as fiction may require reinforcement learning strategies. In most cases, that means having the agent range from the larger model to the smaller model’s responses, and adjust based on the distance.
More advanced techniques also require more significant infrastructure. A large run of reinforcement learning can require tens of millions of agents. Using the frontier lab API to do that “would be ridiculously expensive and potentially a time bottleneck because these models are slow and frankly may not even give you a lift in performance.”
It seems that previous models may be influencing Kimi; Anthropic publicly accused Moonshot, DeepSeek and MiniMax of disassembling its models earlier this year. Anthropic said it found millions of exchanges between its models and users that identified them to those companies through IP addresses and other meta data. Those questions “were different from the usual patterns of use, which reflect the intentional discharge of power rather than legitimate use.” Anthropic did not respond to TechCrunch’s questions about the Fable distillation.
However, distillation seems to be common among AI companies, not only in China. Elon Musk testified earlier this year that his company SpaceXAI had broken OpenAI models to develop Grok, and that this practice was common in the industry. The line between distillation and developing synthetic data sets, for example, can be blurred.
“[I]”In general, Americans underestimate the technology of these Chinese teams,” said Hancock. “One of the founders of Moonshot was a CMU PhD student. These are legitimate researchers and engineers doing solid work. …if the American models stop, I think China’s progress will be slow, but it will continue. They’re not just riding on the coattails here.”
It is also difficult to separate the distillation of the second part of Kratsios’s comment – that Moonshot received improved Nvidia Chips, Grace Blackwell 300s, and reached GB300 equipped servers in Thailand. Those chips are banned from being shipped to China, but there is a black market, according to Sam Bresnick, a researcher at Georgetown’s Center for Security and Emerging Technologies. In May, the founder of Supermicro, an American server maker, was charged with smuggling advanced chips to China.
“I’m a proponent of know-your-customer rules for data centers around the world,” Bresnick said. “If you’re letting a company run massive training on your modern hardware, there should be a way to report who that company is and what they’re doing.”
President Joe Biden’s Commerce Department has proposed federal know-your-customer rules for data centers by 2024, but no further progress appears to have been made under Donald Trump. Exporters sending advanced chips abroad, however, must ensure that they are used only for authorized purposes.
If you shop through links in our articles, we may earn a small commission. This does not affect our editorial independence.



