Experts Claim Kimi K3’s Success Isn’t Attributed to Anthropic’s Fable.
Michael Kratsios, the science advisor at the White House, stated that Moonshot, the Chinese firm behind the Kimi K3—currently the largest open-weight LLM—developed its model by mimicking Anthropic’s Fable LLM, using chips that are banned from being exported to China.
“Covert, large-scale industrial distillation aimed at appropriating proprietary U.S. technology and undermining American research is intolerable,” Kratsios remarked, as talks surrounding a potential ban on Chinese open-weight models continue to stir controversy in the AI landscape. Moonshot has not responded to questions about its training methods, and Kratsios did not offer further details to support his claims.
Kratsios’ comments echoed those of Treasury Secretary Scott Bessent, who noted, “We are detecting signatures of our U.S. large language models within various Chinese models, and that is unacceptable.” The precise nature of these signatures remains uncertain, and the Treasury Department has not replied to requests for clarification.
However, experts are skeptical that distillation—the process of querying an LLM to comprehend its functionalities and replicate its capabilities—can wholly account for the remarkable performance shown by Kimi K3.
“It’s hard to believe that such a powerful model could be released so quickly just through distilling Fable,” noted Braden Hancock, a researcher at the Laude Institute and co-founder of Snorkel AI, during a discussion with TechCrunch. “There simply wasn’t enough time. Fable was made public only on July 1st; it’s impossible to distill a large dataset, train a model, and launch it within two weeks.”
“I’m beginning to think that the effectiveness of distillation is diminishing as Chinese models progress closer to the cutting edge and training techniques evolve towards [reinforcement learning],” Nathan Lambert, an AI researcher at the Allen Institute for AI, said during a podcast released yesterday. “[If it were that simple, anyone could quickly catch up to a GLM or a K3 through distillation. Yet, we haven’t seen this; it’s not merely about supervised fine-tuning.]”
Successfully executing distillation necessitates the ability to systematically query the target model to extract useful post-training information. This can involve prompting the model to elucidate its reasoning processes for solving problems. Alternatively, interactions with the model can be leveraged to improve a new model through a technique known as supervised fine-tuning (SFT).
This fine-tuning method can create a model that seems to have been developed by a third party claiming to be Claude. According to Lambert, during fine-tuning, the “model picks up its nuances.”
However, Lambert argues that the benefits of SFT are waning as models become increasingly complex. Achieving capabilities similar to Fable would likely necessitate reinforcement learning techniques, which usually involve a larger model assessing the responses of a smaller model and adjusting its strategy based on that evaluation.
Advanced methods also require substantial infrastructure. Large-scale reinforcement learning operations may include millions of agents, and relying on a leading lab’s API for such tasks “could be prohibitively expensive and potentially time-consuming, as these models are relatively slow and may not provide a noticeable performance improvement.”
It seems likely that earlier leading models may have influenced the creation of Kimi. Earlier this year, Anthropic publicly accused Moonshot, DeepSeek, and MiniMax of systematically distilling its models. Anthropic claimed it found millions of interactions between its models and users affiliated with those companies through IP addresses and other metadata, describing these interactions as “diverging from normal usage patterns, indicating intentional capability extraction rather than legitimate use.” Anthropic did not respond to TechCrunch’s inquiries regarding the distillation of Fable.
Nonetheless, distillation is recognized as a common practice among AI companies worldwide, not just those in China. Earlier this year, Elon Musk testified that SpaceXAI had distilled OpenAI models to create Grok, asserting that such practices are common in the industry. The line between distillation and generating synthetic datasets is often quite blurred.
“Generally speaking, Americans are underestimating the technical capabilities of these Chinese teams,” Hancock noted. “One of Moonshot’s founders was a PhD student at CMU. These are credible researchers and engineers producing high-quality work. …If American models reach a deadlock, China’s progress may slow, but it will continue. They’re not merely following in others’ footsteps.”
Distinguishing distillation from Kratsios’ other claim—that Moonshot acquired advanced Nvidia chips, like the Grace Blackwell 300s, along with access to GB300-equipped servers in Thailand—is also complicated. These chips are prohibited for export to China, but a black market supposedly exists, according to Sam Bresnick, a research fellow at Georgetown’s Center for Security and Emerging Technology. In May, the founder of Supermicro, a U.S. server manufacturer, faced charges for smuggling advanced chips into China.
“I advocate for know-your-customer laws that apply to data centers globally,” Bresnick commented. “If a company is using your cutting-edge hardware for extensive training activities, there should be a system to report who that company is and what they are doing.”
In 2024, President Joe Biden’s Department of Commerce proposed federal know-your-customer regulations for data centers, but no significant advancements have been reported since Donald Trump’s administration. Nonetheless, exporters dealing in advanced chips are currently expected to ensure their products are used exclusively for authorized purposes.
When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.


