One inference server that runs the same open-weight LLM on Nvidia, AMD, Google TPU, Apple Metal or Intel, with no Python anywhere in the path.
One inference server that runs the same open-weight LLM on Nvidia, AMD, Google TPU, Apple Metal or Intel, with no Python anywhere in the path. The stack is written in Zig on MLIR/OpenXLA and compiles the model ahead of time into a single self-contained binary, so the thing you ship is the thing that runs. The free product is LLMD — a self-contained OpenAI-compatible LLM server shipped as a Docker image per platform (and a Homebrew formula on Apple Silicon).
| Industry | AI & Machine Learning |
| Funding | |
| Stage | Series A |
| Website | Visit ZML |
Where ZML sits against the other names we cover on this beat. Each line is that company’s verdict, not a summary of it.
The 2026-07-30 print settled four of the five questions the prior dossier said it could
Cash $78.2B