Not a flowchart: a computation. State your constraints, and a live Pareto frontier of cost versus capability recomputes which models are even worth arguing about.
Why "which model?" is math, not a flowchart
Dominated
A model is dominated when another model is at least as capable AND at least as cheap, and strictly better on one of the two. A dominated model is never the right answer, whatever your taste: something else beats it on both axes.
On the frontier
The Pareto frontier is what remains after eliminating dominated models: every step up in capability costs real money. Your only genuine decision is where on that frontier your task sits. Constraints shrink the frontier first; preference picks last.
The capability index and prices in this guide are editable assumptions: defaults as of July 2026, list-price style estimates and generic capability tiers, not vendor-verified benchmark results. Context window, latency class, and hosting are fixed dataset attributes. Prices and scores go stale; the point is the method. Edit a capability score or price in the table below and every dot, line, and recommendation recomputes exactly.
The tradeoff explorer
Tighten a constraint and watch models drop out (they turn dim with an ×). The pink dots are the recomputed Pareto frontier of the survivors. Click a dot, or use the Select buttons in the table below, to inspect a model.
The minimum capability your task can tolerate. Raising it kills the cheap end.
Max blended price per million tokens (3:1 input:output mix). Lowering it kills the frontier end.
Latency ceiling
Interactive UIs usually need medium or faster; batch jobs tolerate slow.
How much you must fit in one request: long documents and big codebases push this up.
Privacy / hosting
Open weights means you can run it on your own hardware, so data never leaves your infrastructure.
models passing
11/11
on the frontier
8
cheapest frontier
$0.30
The dataset: editable assumptions
Capability index and prices are editable assumptions (defaults as of July 2026). Edit a value and the chart, frontier, and statuses recompute.
| model | hosting | latency | context | capability | $ in / Mtok | $ out / Mtok | blended | status | |
|---|---|---|---|---|---|---|---|---|---|
| Claude Opus (frontier) | Hosted API | Slow | 200K | $30.0 | ● frontier | ||||
| Claude Sonnet | Hosted API | Medium | 1M | $6.00 | ● frontier | ||||
| Claude Haiku | Hosted API | Fast | 200K | $2.00 | passes, dominated | ||||
| GPT frontier tier | Hosted API | Slow | 400K | $17.5 | ● frontier | ||||
| GPT mini tier | Hosted API | Fast | 128K | $0.70 | ● frontier | ||||
| Gemini Pro tier | Hosted API | Medium | 1M | $5.63 | ● frontier | ||||
| Gemini Flash tier | Hosted API | Fast | 1M | $0.85 | passes, dominated | ||||
| Llama 70B class (open) | Open weights | Medium | 128K | $0.90 | passes, dominated | ||||
| Mistral small (open) | Open weights | Fast | 128K | $0.30 | ● frontier | ||||
| Qwen 72B class (open) | Open weights | Medium | 256K | $0.75 | ● frontier | ||||
| DeepSeek reasoning (open) | Open weights | Slow | 128K | $0.96 | ● frontier |