French Startup ZML Unveils Free Tool to Boost Inference Across Multiple AI Chips
Nvidia continues to have a strong foothold in the market, even as new competitors and alternatives are emerging from different sectors.
ZML, a forward-thinking AI startup based in France and backed by Turing Award winner Yann LeCun, has introduced inference-performance software that allows a range of open-source large language models to run on multiple chips, including Nvidia, AMD, Google’s TPU, Apple Metal, and Intel Arc.
With the launch of their LLM inference server, ZML/LLMD, the company aims to break down existing barriers and enhance the performance of various chips for AI applications, occasionally achieving speeds that surpass standard benchmarks, as noted by ZML founder Steeve Morin in a discussion with TechCrunch.
As AI becomes more integrated into our daily and professional lives, optimizing inference—the process of managing prompts—has gained critical importance, often overshadowing model training. However, challenges related to software and infrastructure can lead to vendor lock-in, according to Morin.
Maximizing performance across different chips is not just a technical challenge; it has the potential to disrupt the market, especially amidst growing concerns about AI costs.
ZML aims to provide companies and cloud providers with the flexibility to utilize a mix of chips, some of which may be more cost-effective or energy-efficient. “Our objective is to enable users to build their own systems and achieve real efficiency gains that encourage widespread AI adoption,” Morin explained.
This software could empower innovative AI chip manufacturers, many located in Europe. Morin mentioned companies like Axelera, Fractile, Kalray, OLIX, Q.ANT, SiPearl, SpiNNcloud, and VSORA as examples. Nevertheless, he emphasizes ZML’s focus on collaborating with these firms on new initiatives.
Morin is optimistic about Nvidia’s future, as he attributes its success to a well-established supply chain. He told TechCrunch that ZML has cultivated a strong partnership with the leading AI chip producer, which is gearing up for an influx in inference demand.
The increased investment in inference has led some analysts to refer to it as an “inference gold rush.” Consequently, ZML faces competition from companies like Baseten, which recently reached a $13 billion valuation; Inferact, founded by the creators of the open-source project vLLM; and RadixArk, which is commercializing SGLang.
While both vLLM and SGLang present competitive challenges for LLMD, Morin’s ambitions for ZML are broader. “We are at a point where we co-design silicon,” he said. He also highlighted the agility of ZML’s compact team of just 20, viewing this as a key factor in the startup’s ability to innovate quickly, with more releases on the way.
This efficient team is also notably well-funded for its size. Morin, who previously held the position of VP of engineering at Zenly—a company acquired by Snapchat for a significant amount in 2017—has raised $20 million from various venture capital firms, including Harry Stebbings’ 20VC, >commit, AALVC, Drysdale Ventures, Xavier Niel’s Kima Ventures, Kindred Capital, LocalGlobe, and Puzzle Ventures.
Unlike ZML’s initial public offering, an inference-focused ML framework launched in 2024 and updated in March, ZML/LLMD is not open-source. However, it is being provided as a free resource for collecting usage insights. “I prefer to evaluate the landscape and generate revenue where it’s most impactful, rather than stifle my growth by being excessively greedy too soon,” Morin remarked.
It remains uncertain when ZML/LLMD will shift to a paid model or how its adoption will progress. Nevertheless, the startup’s cap table indicates that key figures are paying attention, including Dagger and Docker founder Solomon Hykes, Clément Delangue and Julien Chaumond from Hugging Face, and LeCun from AMI Labs. This suggests that European AI startups can indeed succeed locally. “I couldn’t have founded ZML anywhere but in Paris,” Morin commented.
When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.


