Inference crossover calculator
How many months until running AI on a Mac you own pays for itself against paying for API calls — at your hours, your query rate and your electricity.
Connor runs local models on the M4 Mini the defaults describe. The wattage is a wall reading, not a spec sheet.
The API wins. The machine would pay for itself in month 295, past its 36-month write-down — you would replace it before it broke even.
Cumulative cost, month by month
GBP · list prices · 36-month write-down- Where local starts winning
- Never — not at any number of hours in a day, at this query rate.
- What a query costs you
- £0.0023 on the API · 22s of machine time
Your numbers in the datacenter’s terms
SemiAnalysis, 14 September 2026 compared a Jetson Thor on a desk against a B300 in a rack. Once utilization is counted — a rack kept 90% busy, a device 40% — sending the work to the cloud costs 54 cents for every on-device dollar, and for a device used 1–2 hours a day it falls to 12%. One datacenter GPU does the work of about 7 devices. Those are their numbers. These are yours:
- Your utilization
- 17%
- API per local pound
- 17p
- Machines for the load
- 1
What this counts, and what it leaves out
- The API side is list price. No prompt caching, no batch discount, no committed-use deal. If you have one, your API month is smaller than the number here.
- The local side is the sticker price over three years, plus electricity under load. Not counted: the model being worse, your time setting it up, a RAM upgrade, or the hours the box idles at a few watts.
- Capacity is the number the datacenter maths hides. A rack never runs out of headroom; a Mini does. Each preset carries a tokens-a-second rate for a model that fits it, and if your demand needs more than one machine the bill says so.
- The frame is SemiAnalysis, 14 September 2026. Their numbers are a Jetson Thor against a B300 — datacenter silicon on both sides. This page translates the shape of that argument, utilization first, onto hardware a person buys.