TL;DR: Choosing a model is a trade among five axes: capability, latency, cost, privacy, and openness - plus two settings you tune after choosing, how much reasoning effort to buy per request and how much context you need. Rank the axes for your specific task, use benchmarks only to build a shortlist, and decide with your own evals on your own data. Then design for change - the ranking of models shifts every few months, and the right answer is one you can swap.
How it works
Start from the task, not the leaderboard. Capability: how hard is the hardest thing the model must do reliably? Complex multi-step reasoning and agentic work justify a frontier model; classification, extraction, and summarization usually do not. Latency: a chat interface needs tokens streaming in well under a second, while a nightly batch job does not care - and smaller models are consistently faster. Cost: pricing is per token but not at a single rate - cached input, batched requests and long context are billed differently, so the bill tracks how you call the model as much as how often; a model ten times cheaper that succeeds 95 percent as often wins for most high-volume steps. Privacy: if data cannot leave your infrastructure, hosted APIs are off the table, which turns the choice into a local LLM or self-hosted deployment question. Openness: open weights buy control, fine-tuning freedom, and immunity from provider deprecation, at the price of running or renting the serving yourself. Jurisdiction alone is less often disqualifying than it was: the major providers now let you pin where inference runs, per request or per workspace.
Few applications need one answer. The emerging default is a portfolio: a frontier model for the hardest steps, a mid-tier model for routine ones, and sometimes a tiny model for classification and routing. That routing pattern - send each request to the cheapest model that handles it - is a real lever, though no longer the first one to reach for: on agent workloads, caching the repeated prefix and trimming the prompt typically save more than any model swap, and turning a model's reasoning effort down often saves more than changing models at all. Routing still falls out naturally once you stop asking "which model is best" and start asking "which model does this step need" - and some of it is now a product rather than a project, with providers retrying a refused request on another model inside one API call and gateways exposing an automatic choice as a model name you can call.
The process, then: rank the five axes for your task; use benchmarks nearest your task's shape to cut the field to two or three candidates; then run your own evals on your own data, because leaderboard rank predicts your workload poorly. Finally, assume the answer expires - literally: providers publish retirement dates for named model versions and give notice before requests start failing, and a model released one autumn can be scheduled for retirement by the next. Models improve, prices drop, and providers deprecate, so keep the model behind an abstraction - a gateway, a router, or just a config value - and re-run the evals when the landscape moves. Choosing well once matters less than being cheap to re-choose.
Where it sits in the AI stack
Model choice is the decision that connects your requirements to everything downstream - the provider you integrate, the costs you carry, and the evals that keep the choice honest:
Key tools and implementations
-
Benchmark leaderboards
Public rankings for the shortlisting step - a coarse filter, never the final call.
-
Provider playgrounds
Browser consoles for trying candidate models on real examples before writing any code.
-
Eval frameworks
Tools that score candidate models against your own test set - the evidence the decision should rest on.
-
Gateways and routers
Abstraction layers that put many providers behind one API, making the chosen model cheap to swap - and increasingly doing the routing, caching and spend tracking for you.
Related entries
- Model benchmarks Standardized tests that score model capabilities, useful for rough comparison but easy to overfit and game.
- Inference provider A company that runs AI models on its own hardware and sells access through an API.
- Local LLM Running a language model entirely on hardware you control instead of calling a hosted API.
- Open-weights models Models whose trained weights are published for anyone to download, run locally, and fine-tune.
- Evals (AI evaluations) Structured tests that score an AI system's outputs so teams can measure quality and catch regressions.