They make AI models run on the device itself — phones, wearables, Raspberry Pis, microcontrollers — and hand a query off to a cloud model only when the small…
They make AI models run on the device itself — phones, wearables, Raspberry Pis, microcontrollers — and hand a query off to a cloud model only when the small model knows it is out of its depth. Three products: Cactus Engine (a source-available C++ runtime and quantizer), Cactus Hybrid (a router post-trained *into* the model weights that scores its own confidence and escalates), and Needle 2 (their own 45M-parameter, 14MB tool-calling model that fits in 28MB of RAM).
| Industry | AI & Machine Learning |
| Website | Visit Cactus Compute |
Where Cactus Compute sits against the other names we cover on this beat. Each line is that company’s verdict, not a summary of it.