Your AI bill has stopped being a licence and started being a meter
AI has stopped behaving like software and started behaving like a raw material. Budget it as a variable cost, cap it and meter only what pays its way.
Most of us are still buying AI the way we buy software: a seat, a monthly fee, a figure you set in January and stop looking at. That was perfectly sensible while AI was something a person sat in front of and typed at. It stops being sensible the moment an agent starts running on its own, by the hour, spending your money while you're asleep. The unit is wrong. What you're buying now behaves far more like a raw material, something you consume by the job, and raw materials have no business sitting on the same line of your P&L as the Office licences.
Two very large companies found this out in public this year. Somewhere around April, Uber's engineers finished off the company's entire 2026 AI coding budget — not their department's share of it, the lot. The tool was Claude Code, adoption inside the engineering org went from about a third to more than four-fifths in the space of a month, and the cost of running a single engineer landed somewhere between $500 and $2,000 a month. Microsoft looked at much the same arithmetic and went the other way: it cancelled most of the internal Claude Code licences in one of its divisions at the end of June, at roughly $2,000 per engineer per month, and pointed everyone back at Copilot, which, conveniently, it owns.
Both stories get filed under "AI has got expensive". The truer reading is duller and more useful. The tool worked, the people used it, and the meter ran — and nobody had a budget shaped like a meter. A Copilot Business seat costs $19 a month. A metered engineer at Uber cost several hundred times that. Those two numbers are not two prices for the same thing.

What changed underneath
The practical consequence for you is that a task no longer has a price; it has a running cost. Aaron Levie of Box put his finger on why earlier this year: we went from chat tools, which were cheap and had small context windows, to agents with enormous context windows that keep track of long-running work, and models that cost an order of magnitude more to run. Every time an agent takes another turn, it reads the whole conversation again — the files it opened, the tool calls it made, the plans it drew up and then abandoned. The tenth turn is carrying everything that came before it.
So the cost of a job is the price of asking, multiplied by however long the agent decides to keep going. And you don't set that number. Whoever typed the instruction sets it, more or less by accident.
A seat licence has one genuinely lovely property: it's bounded. However enthusiastic your team gets, a person can only use a tool so hard in a working day. There's a ceiling built into human stamina, and every software vendor since the 1980s has priced against it. An agent has no such ceiling. I once watched one spend the best part of an afternoon, and rather more of my money than I'd have authorised in advance, chasing a bug that turned out to be a typo I'd made myself — and the only reason it stopped was that I happened to be in the room.
Small is, for once, an advantage
Now for the encouraging part. Uber's chief operating officer said out loud that he can't yet draw a line between all this Claude Code spend and anything a passenger would notice. "That link is not there yet," he said. Read that again, because it's remarkable. A company that size has burned a year's budget and can't tell you what it bought.
You can. If you run a business of five or fifty people, there are perhaps a dozen jobs in the week that AI is doing at any scale, and you could name most of them off the top of your head. You know whose work it touched, roughly what it saved, and whether the client noticed. That link, spend to output, is the whole game — and it's the one thing a 5,000-person engineering organisation can't see and you can. It's an advantage that lasts exactly as long as you bother to use it.
What to do about it
The instinct when the bill jumps is to go hunting for a cheaper model, and there's real money in that; I've written before about rewiring the budget when the price of the frontier collapses and about routing work to the cheapest model that can handle it. Both still hold. But routing answers which model. This is a question of what kind of cost, and no amount of clever routing fixes a line item sitting in the wrong place on the page.
- Split the budget in two. One fixed line for the seats — the chat subscriptions, the copilots, the things a human drives at human speed. One variable line for metered agent work, which sits with your other variable costs, next to the revenue it's meant to be producing. It'll look uncomfortable there. Good; overheads get renewed, inputs get justified.
- Set the ceilings before you need them. Uber has since capped spend at $1,500 per tool per month, per engineer. Sensible, and about four months too late. You want three caps: a monthly ceiling for the business, a per-person ceiling so no single enthusiast can eat the quarter, and a per-workflow ceiling on anything that runs unattended. Set alerts at half and three-quarters, and be clear about who is allowed to raise a cap mid-month. Every serious provider will let you do all of this from one settings page. Almost nobody opens it.
- Meter only what you can point at. The test is simple: if you can't name the revenue, the margin, the turnaround time or the service outcome that a metered job improves, it doesn't get metered. Summarising email, tidying notes, formatting a report — that's routine work, it runs perfectly well on a cheap fixed tier, and it shouldn't be anywhere near an agent with a running meter.
- Learn what a job costs you. Not what the month costs; what one instance of the work costs, end to end. Cost per completed job is the number that tells you whether an agent is earning its keep, and hardly anyone tracks it, because the invoice arrives as a single frightening lump with no story attached.
The bit I can't tell you
I don't know where this settles. Per-token prices are still falling, quite fast, and yet real spend keeps climbing, because as soon as something gets cheaper we use dramatically more of it. That's Jevons' paradox doing its usual work, and it means the two curves can go in opposite directions for years. Whether your bill in 2028 is smaller than today's, I honestly couldn't say, and I'd be wary of anyone who tells you they can.
What isn't in doubt is the shape of the cost. It has come off the fixed line and onto the variable one, and it will stay there, because agents that run by the hour can't be priced by the seat. Most businesses will work this out the way Uber did, which is by running out of money in April and asking the questions in May.
Let's get into the habit now, while the numbers are still small enough to be embarrassing rather than fatal.
