Inputs

Monthly tokens (in + out) 100M

In millions. Slide to your typical monthly throughput across all calls.
Model class medium (30B–70B)
small (7B–13B) — chat, classification, summaries medium (30B–70B) — agent workflows, code, RAG large (70B+ MoE/dense) — frontier-ish reasoning
Determines tok/sec per mini and which frontier API tier is the fair comparison.
Latency P95 target (ms first token) 150 ms
Tighter latency = more headroom = more minis. Default 150ms covers most chat + RAG.
Peak/avg ratio 3.0×
If peak hour is 3× the average, you size for peak to keep latency.
Horizon (months) 36 mo
36mo matches typical Mac Mini service life. Sets the TCO window below.

Cluster Ops sizing

Mac Mini count (peak) ·

Capex (one-time) ·
Monthly opex (power + colo + ops) ·
Cost per million tokens (Cluster Ops) ·
TCO over horizon ·

vs frontier API — same horizon

GPT-4o

·
·
Claude Sonnet
·
·
GPT-4o-mini
·
·
Break-even vs cheapest comparable API ·
Show the math
computing…

Where the API wins. Low monthly volume (under ~30M tokens/month for medium-class), bursty traffic without sustained peak, frontier-quality requirements where on-prem is not yet a fair comparison (full-precision GPT-5/Opus class). The calculator says so when it’s honest.

Where on-prem wins. Sustained high volume (~100M+ tokens/mo medium-class), data-residency requirements (HIPAA / GDPR / public-sector), or workflows where the latency to a co-located mini beats round-trip to a hyperscaler region. Cluster Ops productizes this: hot-swap models, automated eviction, hardened deploys, monthly cost-per-token report.

Assumptions are visible (click "Show the math" above) and biased pessimistic for on-prem — we assume entry-level M4 Pro mini at $2,000, $0.15/kWh power, $5/mo colo amortized, and 1 ops-hour/mini/month at $200/hr. Plug in your own numbers; the page is pure client-side JS and the math is in plain view.

See the Cluster Ops subscription → Methodology vs frontier API / cloud GPU

Scope this with us →