Home / Rank #9

Cheaper & open models / self-hosting

Typical savings50–95% per token on suitable tasks
EffortHigh for self-hosting, low for switching

Frontier-model prices keep falling, and small models (Haiku-class, GPT-mini-class, Nova, open Llama/Qwen/Mistral weights) now handle classification, extraction, and routine drafting at a tiny fraction of frontier price. For high-volume, well-scoped tasks, a fine-tuned small model regularly beats a prompted frontier model on cost and matches it on quality.

Self-hosting open weights (with quantization) makes sense past sustained volume thresholds — but be honest about GPU, ops, and eval costs; the API price war means the crossover point is higher than most teams assume.

How to do it

  1. Benchmark your top-volume tasks on one tier down (and two tiers down) from your current model.
  2. Fine-tune a small model on tasks with clear ground truth and high volume.
  3. For self-hosting, price the full picture: GPUs, autoscaling headroom, ops time, and eval maintenance.
  4. Re-benchmark quarterly — model prices and quality shift fast enough to change the answer.

Frequently asked questions

When does self-hosting pay off?

Rules of thumb vary, but sustained six-figure annual API spend on stable workloads is where serious evaluation starts. Below that, falling API prices usually beat owning GPUs.

Tools for this method

Multi-provider model marketplace

OpenRouter

One API and one bill across hundreds of models from every major lab — the fastest way to arbitrage the model price war and A/B che…

Next method: #10 LLM gateways & spend-tracking tools