Self-host or pay per token?
Put your monthly workload through a tracked API route and through a GPU at the hourly rate you actually pay. The tool shows both monthly totals, the break-even GPU rate, the memory the model needs, and whether your measured throughput can keep up.
- API prices
- 144 tracked routes
- GPU price
- Yours, never assumed
- Memory check
- Weights + margin
- Quality claim
- None
Describe the workload and the GPU
API prices are tracked automatically. Your GPU rate and measured throughput are yours to enter.
Monthly cost side by side
Same workload, two ways to run it.
Two honest totals, several things they leave out.
The API total is requests × tokens × the route's listed per-token prices, with no prompt caching. The GPU total is your hourly rate × GPUs × hours per month; the break-even rate is the API total divided by those GPU-hours. Weight memory is parameters × bits ÷ 8 plus a 20% margin; KV cache is excluded because the architecture is unknown here.
Not modelled: engineering and operations time, failover, egress, storage, quantisation quality loss, API rate limits and enterprise discounts. A cheaper monthly number on either side is a starting point for that conversation, not its end.