French Startup ZML Unveils Free Tool for Optimizing Inference on Diverse AI Chips
Nvidia continues to hold a strong market position, even as new competitors and alternatives emerge across various sectors.
ZML, a groundbreaking AI startup based in France and backed by Turing Award winner Yann LeCun, has introduced software aimed at boosting inference performance. This enables a variety of open-source large language models to run on multiple hardware systems, including Nvidia, AMD, Google’s TPU, Apple Metal, and Intel Arc.
With the launch of their LLM inference server, ZML/LLMD, the company seeks to address existing challenges and enhance chip performance for AI tasks, often achieving speeds that surpass standard benchmarks, as noted by ZML founder Steeve Morin in an interview with TechCrunch.
As AI becomes more integral to both everyday and professional settings, optimizing inference—the processing of prompts—has become vital, increasingly overshadowing model training. Nevertheless, software and infrastructure challenges could result in vendor lock-in, as Morin pointed out.
Improving performance across various chips not only poses a technical challenge but also holds the potential to disrupt the market, especially amidst rising concerns regarding AI costs.
ZML aims to provide businesses and cloud service providers with the flexibility to leverage a diverse range of chips, some of which may offer better cost efficiency or energy savings. “Our objective is to empower users to build their own systems and achieve real efficiency gains that encourage widespread AI adoption,” Morin stated.
This software could pave the way for innovative AI chip manufacturers, many of which are located in Europe. Morin mentioned companies like Axelera, Fractile, Kalray, OLIX, Q.ANT, SiPearl, SpiNNcloud, and VSORA but reiterated ZML’s commitment to collaborating with these firms on new projects.
Morin remains optimistic about Nvidia’s future, crediting its success to a well-established supply chain. He revealed to TechCrunch that ZML has solidified a strong partnership with the leading AI chip producer, which is preparing for an increase in inference demand.
The heightened emphasis on inference has led some analysts to label it an “inference gold rush.” As a result, ZML faces competition from companies like Baseten, which recently secured a valuation of $13 billion; Inferact, initiated by the creators of the open-source project vLLM; and RadixArk, which is bringing SGLang to the marketplace.
While both vLLM and SGLang pose significant competition for LLMD, Morin has broader visions for ZML. “We are currently co-designing silicon,” he stated. He also stressed the adaptability of ZML’s compact team of just 20, viewing this nimbleness as essential for the startup’s rapid innovation, with more releases anticipated shortly.
This efficient team is also well-funded for its size. Morin, a former VP of engineering at Zenly (which was acquired by Snapchat for a substantial amount in 2017), has raised $20 million from various venture capitalists, including Harry Stebbings’ 20VC, >commit, AALVC, Drysdale Ventures, Xavier Niel’s Kima Ventures, Kindred Capital, LocalGlobe, and Puzzle Ventures.
Unlike ZML’s earlier public offering, a focused inference ML framework launched in 2024 and updated in March, ZML/LLMD is not open-source. However, it is available as a complimentary tool for collecting usage data. “I prefer to evaluate the landscape and generate revenue where it matters most rather than hinder my growth by being overly greedy too soon,” Morin expressed.
It remains uncertain when ZML/LLMD will convert to a paid model or how its integration will progress. However, ZML’s cap table shows that influential individuals are taking notice, including Dagger and Docker founder Solomon Hykes, Clément Delangue and Julien Chaumond from Hugging Face, and LeCun from AMI Labs. This suggests a positive outlook for European AI startups to succeed locally. “I couldn’t have launched ZML anywhere but in Paris,” Morin reflected.
When you make a purchase through links in our articles, we may receive a small commission. This does not influence our editorial independence.


