For most of the last two years, your AI bill has been a lie. Not yours specifically — everyone's. The price you paid to run a task through a frontier model was never what it cost to serve you. The gap was covered by venture capital and by the big cloud providers, content to lose money on every call while they fought for market share. That was the deal, whether you knew you were in it or not: cheap, capable AI, paid for by someone else.

That deal is ending, and you can see it in the invoices. Which means the way most small businesses buy AI, one model for everything and usually the dearest one, has just turned from a quirk into a budgeting problem, and the businesses that sort it out first will be the ones still smiling at month-end.

On 1 June, GitHub raised the multipliers on its Copilot plans. The frontier models like GPT-5.4 and Gemini 3.1 Pro went from a 1× premium request to 6×. Claude Opus jumped from 7.5× to 27×. Developers running agentic sessions started reporting bills ten to fifty times higher than the month before, for the same work. Around the same time Anthropic brought in weekly rate limits on its Pro and Max plans, with a separate, tighter cap on its most capable model — not because demand fell, but because the compute to serve everyone at the old rate simply wasn't there. How far the repricing runs from here, I don't think anyone can honestly say; the direction, though, has stopped being in doubt.

Bar chart showing GitHub Copilot premium-request multipliers rising on 1 June 2026: GPT-5.4 and Gemini 3.1 Pro from 1x to 6x, Claude Opus 4.7 from 7.5x to 27x
One repricing, two messages: routine-model work barely moved, while sending the same job to a frontier model from 1 June started costing several times more.

I've been making half of this argument for nine months. Back then the point was that most of what your team does with AI runs perfectly well on the cheapest model available, and that defaulting everything to the flagship was a management failure dressed up as a technical default. That was true when the flagship was cheap. What's changed is the price. The old argument was that you don't need to. The new one is that you can't afford to.

The maths isn't complicated, which is exactly why it's so easy to ignore until the bill arrives. Think about what you actually run through AI on a normal day: summarising email threads, pulling figures out of invoices, drafting first-pass replies, formatting reports, tidying meeting notes. None of that needs the cleverest model on the market; the cheapest one handles it without breaking a sweat. The work that genuinely earns a frontier model, the nuanced client deliverable or the call that touches money or legal exposure, is a thin slice of the week. Call it eighty-twenty, give or take; you won't know your own split until you've looked.

When the top model cost barely more than the cheap one, mixing the two up cost you a rounding error. Now that it costs six, ten, twenty times more per call, every task you send to the wrong tier is a small overpayment that compounds quietly, until month-end makes it loud.

Your 90-day routing playbook

So what do you actually do about it? You've got a quarter, and the work breaks into four moves.

  1. Audit what you're spending (week one). Pull last month's AI invoice and find the three workloads that ate the most tokens. Most owners have never looked, and the first look is usually a shock: the expensive model turns out to be doing work the cheap one could do in its sleep. This is the same discipline that let a Mark Cuban-backed cheese company, Rebel Cheese, point an AI agent at its shipping invoices and save four hundred thousand dollars in a year — a quarter of a million of it in overcharges the firm only caught after one frantic holiday season. You can't route what you haven't measured. If what you find makes you want to rebuild the budget from the ground up rather than just trim it, I've set out a framework for exactly that.
  2. Set a cheap-first default (weeks two to four). Every task starts on the cheapest model that might handle it, and only moves up a tier when it genuinely needs to. Frontier access becomes the exception you reach for on purpose, not the default you fall into by accident. If the selector in your AI tool means nothing to you, it's worth learning to tell the model apart from the product wrapped around it, because that little toggle is the part of the bill you haven't been choosing. I run my own work this way, cheap model first and escalate only when the task earns it, and the thing that still surprises me is how rarely it does. If you want the mechanics, how to wire it up and who owns it once it's running, I've written about that separately.
  3. Run bake-offs on real work (month two). Take your actual tasks, not benchmarks, and run them side by side, cheap model against expensive, then look at the output honestly. If the cheap model gets it eighty per cent right, you've found your routing rule. The other twenty per cent tells you precisely where the frontier model still earns its multiplier — the spend you should stop apologising for. It's how Cost Plus Drugs ended up having Claude produce a weekly price-comparison report across its twenty-five most expensive drugs: a recurring, well-scoped job that used to need software or a dedicated person, now done in minutes in a browser.
  4. Give someone the job of keeping it fresh (ongoing). Models change every few weeks, so today's bargain is next quarter's overpriced default. Appoint a model sommelier, someone whose job is to know which model is the right pour for which task and to revisit the routing as the menu changes. It needn't be a full role; it just needs to be someone's actual job, because routing without an owner drifts back to one-model-for-everything inside a week.

The wider shift is that the cost of AI has stopped being someone else's problem. For eighteen months the labs absorbed it; now they're handing it back, politely, through multipliers, rate limits and repricing nobody bothered to announce. The businesses that come out of the next eighteen months in good shape will be the ones who built routing discipline before the bill forced their hand, not the ones still explaining, come December, why a line item they never watched managed to triple.