OTHER

Experts Claim Kimi K3’s Success Is Not Linked to Anthropic’s Fable.

Michael Kratsios, the scientific advisor at the White House, highlighted that Moonshot, the Chinese firm responsible for the Kimi K3—now the largest open-weight LLM—developed its model by emulating Anthropic’s Fable LLM and employing chips that are not permitted for export to China.

“The covert and extensive industrial distillation aimed at seizing proprietary U.S. technology and undermining American research is completely unacceptable,” Kratsios stated, as debates over a possible ban on Chinese open-weight models continue to ignite controversy in the AI field. Moonshot has not responded to inquiries about its training methods, and Kratsios did not offer further evidence to substantiate his claims.

Kratsios’ remarks echoed those of Treasury Secretary Scott Bessent, who stated, “We are detecting signatures of our U.S. large language models within several Chinese models, which is intolerable.” The specific nature of these signatures remains ambiguous, and the Treasury Department has yet to address requests for more information.

However, experts express skepticism that distillation—the technique of querying an LLM to grasp its functionalities and replicate its abilities—could fully account for the exceptional performance exhibited by Kimi K3.

“It’s hard to believe that such a potent model could be launched so quickly just by distilling Fable,” said Braden Hancock, a researcher at the Laude Institute and co-founder of Snorkel AI, in a conversation with TechCrunch. “There simply wasn’t enough time. Fable was released only on July 1st; it’s unrealistic to distill a large dataset, train a model, and introduce it within two weeks.”

“I’m beginning to suspect that the effectiveness of distillation is waning as Chinese models evolve and training methods shift toward [reinforcement learning],” Nathan Lambert, an AI researcher at the Allen Institute for AI, remarked during a podcast aired yesterday. “[If it were that straightforward, anyone could quickly bridge the gap with a GLM or a K3 through distillation. Yet, we haven’t seen this; it’s not just about supervised fine-tuning.]”

Successfully carrying out distillation requires the capability to systematically query the target model to extract valuable post-training insights. This might involve prompting the model to elaborate on its reasoning processes when solving problems. Additionally, interactions with the model can be utilized to refine a new model via a technique called supervised fine-tuning (SFT).

This fine-tuning strategy can produce a model that seems to have been crafted by a third party representing Claude. According to Lambert, during fine-tuning, the “model absorbs its nuances.”

Nonetheless, Lambert argues that the benefits of SFT are diminishing as models become increasingly sophisticated. Achieving capabilities comparable to Fable would likely necessitate reinforcement learning methods, which generally involve a larger model assessing the responses of a smaller model and adjusting its approach based on that evaluation.

Advanced techniques also require substantial infrastructure. Large-scale reinforcement learning initiatives may involve millions of agents, and relying on a leading lab’s API for such tasks “could be prohibitively expensive and potentially time-consuming, as these models are relatively slow and may not deliver a significant performance boost.”

It appears likely that earlier leading models may have played a role in the development of Kimi. Earlier this year, Anthropic publicly accused Moonshot, DeepSeek, and MiniMax of systematically distilling its models. Anthropic claimed to have identified millions of interactions between its models and users linked to these companies through IP addresses and other metadata, labeling these interactions as “deviating from normal usage patterns, indicating intentional capability extraction rather than legitimate use.” Anthropic did not respond to TechCrunch’s inquiries regarding the distillation of Fable.

However, distillation is recognized as a prevalent practice among AI companies worldwide, not solely those in China. Earlier this year, Elon Musk testified that SpaceXAI had distilled OpenAI models to create Grok, asserting that such practices are commonplace in the industry. The line between distillation and generating synthetic datasets is often quite blurred.

“In general, Americans are underestimating the technical abilities of these Chinese teams,” Hancock noted. “One of Moonshot’s founders was a PhD student at CMU. These are credible researchers and engineers producing high-quality output. …If American models hit a standstill, China’s progress may slow, but it won’t halt. They’re not simply lagging behind.”

Distinguishing distillation from Kratsios’ other claim—that Moonshot obtained advanced Nvidia chips, like the Grace Blackwell 300s, and access to GB300-equipped servers in Thailand—is also complex. These chips are banned for export to China, yet a black market reportedly exists, according to Sam Bresnick, a research fellow at Georgetown’s Center for Security and Emerging Technology. In May, the founder of Supermicro, a U.S. server manufacturer, faced charges for smuggling advanced chips into China.

“I advocate for know-your-customer laws that apply to data centers globally,” Bresnick stated. “If a company is using your cutting-edge hardware for extensive training activities, there should be a mechanism to report who that company is and what they are doing.”

In 2024, President Joe Biden’s Department of Commerce proposed federal know-your-customer regulations for data centers, but no significant progress has been reported since the Trump administration. Nevertheless, exporters dealing in advanced chips are currently expected to ensure their products are utilized solely for authorized purposes.

When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.