French AI startup ZML launches free software to speed inference across many chips
ZML, a French artificial‑intelligence startup backed by Turing Award laureate Yann LeCun, has unveiled its latest product, ZML/LLMD. The new software is positioned as a cost‑saving solution for deploying large language models, promising to reduce the computational resources required for inference and training.
ZML/LLMD introduces a suite of optimizations that streamline memory usage and accelerate inference on commodity hardware. By partitioning models across multiple GPUs and employing dynamic quantization, the platform can deliver near‑real‑time performance while cutting GPU hours by up to 40 % according to the company’s benchmarks. The release follows a growing industry push to make AI more accessible, as the high cost of running state‑of‑the‑art models remains a barrier for many enterprises.
If the performance claims hold up in broader deployments, ZML/LLMD could reshape how small and mid‑size firms adopt advanced language models. The startup plans to partner with cloud providers for early pilots and is preparing a public demo later this quarter to showcase its cost‑efficiency gains.