Every week a new model launches, tops a benchmark, and sets off a round of 'is this the new best AI?' It's the wrong question, and chasing it will keep you permanently one launch behind. I spend my days looking at the hardware and economics underneath these systems, and from down there the truth is plain: there is no 'best model,' the same way there is no 'best vehicle.' A cargo ship and a motorcycle are both correct answers — to completely different questions. The skill that actually compounds is learning to name the job first, then reach for the class of tool built for it.
Name the job, not the model
Almost every task you'd hand an AI falls into one of a few shapes, and each shape has a matching tool. The mistake isn't picking the wrong model; it's picking a model at all before you've said out loud what shape the work is. Do that first and the choice mostly makes itself — and it keeps making itself correctly after the current leaderboard has been replaced twice. The five branches below are the ones I actually use, in the order they come up. They are mutually exclusive on purpose: take the first that fits and stop reading, because a task that seems to match three of them is usually a task you haven't finished defining.
Name the shape, get the tool
That router is deliberately boring, and boring is the point. Every one of those branches has been true for two years and will be true for two more, because they describe the shape of work rather than the state of the market. Notice how little the question 'which model won this week' has to do with any of them: the private-data branch doesn't care, the long-document branch cares about one spec, and the high-volume branch cares mostly about price. Only one branch out of five is even about capability, and it is the branch you take least often.
There's no best model, the same way there's no best vehicle. A cargo ship and a motorcycle are both right answers to different questions. Name the trip first.
The tiers exist for a reason — use them
Here's the part my desk cares about most, and where the most money gets wasted. The big labs now sell *families*, not models — a flagship, a mid-tier and a budget tier — and the gap between the ends of the ladder is much larger than most people using them assume. Our own [Scoreboard](#/scoreboard) carries vendor list prices next to an independent capability score for each model, which makes the trade legible in a way a leaderboard ranking never does. Read the two columns together rather than sorting by either one, and the shape of the decision changes.
Vendor list price, per million output tokens
Now put the scores beside those prices, because the prices alone are only half the trade. Claude Fable 5 lists at $50 per million output tokens and scores 60 on the independent index our board tracks. GPT-5.6 Sol lists at $25 and scores 59. GPT-5.6 Terra lists at $15 and scores 55. GPT-5.6 Luna lists at $6 and scores 51. Gemini 3.5 Flash lists at $2 and scores 50. Read that ladder from the bottom: the entire distance from the cheapest model on the board to the most expensive one is ten points of measured capability and twenty-five times the price. GLM-5.2 sits at $2 as well, on the same 51 as Luna. The spread within a single family is smaller but still enormous — Sol and Luna are the same product line, and one costs about four times the other.
What each tier is actually for
| Budget GPT-5.6 Luna, GLM-5.2, Gemini 3.5 Flash | Mid GPT-5.6 Terra, Claude Sonnet 5 | Flagship Claude Fable 5, GPT-5.6 Sol | |
|---|---|---|---|
| List price, per million output tokens | $2 | $15 | $50 |
| Independent index score | 51 | 55 | 60 |
| Share of your actual work | Most of it | The awkward middle | The hard tenth |
| Best at | Summarising, sorting, reformatting, first drafts | Real work you'll edit — code, analysis, long drafts | Multi-step reasoning where a wrong step ruins the answer |
| Where it fails you | Long chains of logic; it loses the thread and stays confident | Genuinely novel problems with no near-neighbour | Your budget, and your patience with the latency |
| Escalate when | Output is wrong in a way you can name | You've rewritten the same section twice | Nothing above it to escalate to — verify by hand instead |
The branch that overrides the others
One row of that table deserves separating out, because it isn't a capability judgment at all and it gets made by accident more than any other. If the data is genuinely private, regulated, or belongs to someone who hasn't agreed to this, the tier question is moot: you run it somewhere you control or you don't run it. That branch overrides the score column, the price column and your deadline, and the reason people skip it is structural rather than careless — the capability question is interesting and the data question is boring, so attention goes to the interesting one. Make it the first question rather than the last. It takes about five seconds to answer and it is the only decision on this page that you cannot reverse afterwards.
The habit: draft cheap, escalate hard
Here's the workflow that beats memorizing any leaderboard. Start a task on a fast, cheap model. If the output's good enough — and it will be, more often than you expect — you're done, for a fraction of the cost and the wait. If it visibly struggles, *then* escalate the same prompt to a frontier model. You'll quickly build an intuition for which of your own tasks need the heavy machinery and which never did. That intuition is durable in a way that 'which model is best this week' never will be: the launches will keep coming, the names will keep changing, and the job-to-tool mapping underneath will keep being the thing that actually matters.
Sort your own work into tiers
- Scroll back and write down the recurring jobs. Don't list what you meant to use it for.
- Volume, hard reasoning, long document, code, or private. Every task takes exactly one tag — the first that fits.
- Change the model you land on when you open a new chat. This single setting is the whole intervention.
- Don't escalate pre-emptively. You are gathering evidence about your own work, not testing the model.
- Turn it into a rule you could hand to someone else: a named failure, not a feeling.
- Move the identical prompt up a tier. Changing the prompt and the model at once tells you nothing about either.
The rule you write in step five is the real output of the whole exercise, and it's the part people skip. 'It feels off' is not a rule; it's a mood, and a mood escalates on the days you're anxious rather than on the days the work is hard. 'It invents library methods that don't exist' is a rule. 'It drops the third and fourth constraints when I give it more than two' is a rule. Both are things you can check in ten seconds and both are things a colleague could apply without asking you what you meant. Once you have two or three of those written down, tier selection stops being a judgment call you make forty times a day and becomes something closer to a reflex.
Five expensive habits
Two limits on everything above. The prices and scores are a snapshot — ours was updated July 29, 2026 — and the whole point of this beat is that they move; treat the ordering as current rather than permanent, and re-read the board rather than remembering it. And the independent score is a single aggregate over a set of benchmarks, which means it is a decent proxy for general capability and a poor proxy for your particular job. A model two points lower on the index can be plainly better at the specific thing you do all day. Several rows on our own board carry no score at all, because the only figures their vendors published were self-reported, and self-reported numbers are not established capability.
So use the ladder for what it's good for: sizing the trade, not settling it. Ten points of measured capability across twenty-five times the price tells you, unambiguously, that most work does not belong at the top of the board — and that is the decision worth getting right, because you make it dozens of times a day and it compounds quietly in both directions, in what you spend and in what you spend waiting. Which model sits in first place this month is the decision you make about twice a year, and it is the one everybody spends their attention on. Get the daily one right and the annual one stops mattering very much: you will already have a cheap default that handles the bulk of your work and a frontier model you know the shape of, and a new launch becomes a thing you evaluate on a quiet afternoon rather than a thing that resets your habits.
- There is no best model — name the job first, then pick the tool class.
- Cheap-fast for volume work, frontier for hard reasoning, local for private data.
- Our board: $2 to $50 per million output tokens, for a ten-point spread in score.
- The routine 90% runs fine on cheap tiers; escalate only the hard 10%.
- Prices and scores here are a snapshot, not a constant — they are not stable week to week.
