The $3 frontier: what DeepSeek V4 does to your 2026 AI budget
DeepSeek V4-Pro just matched Opus 4.7 on coding at a seventh of the cost. The four-bucket framework for rewriting your 2026 AI budget — by Q3 ideally.
Open your procurement spreadsheet. A workload that runs at $25 per million output tokens on Claude Opus 4.7 now has a credible alternative at $3.48 with comparable or better coding performance. That’s not a forecast or a thought experiment. DeepSeek V4-Pro shipped at that price on 24 April, and on the headline coding benchmarks buyers actually quote — Codeforces, LiveCodeBench, SWE-bench Verified, the proxies for how a model performs on production-scale coding and agentic work — V4-Pro is ahead on the first two (3206 against GPT-5.4’s 3168, and 93.5 at the top of the LiveCodeBench table) and statistically tied with Opus 4.7 on the third (80.6 against 80.8). The premium-frontier line in your 2026 AI budget needs rewriting, not renegotiating, by Q3.

Amazon committed $25 billion to Anthropic on 21 April. Google committed up to $40 billion three days later — $10 billion immediate cash at a $350 billion valuation, $30 billion more conditional on performance targets, per Bloomberg’s reporting on the deal terms. Anthropic absorbed roughly $65 billion in fresh pledges in four days, on the way to fundraising talks reported at a $900 billion valuation. The American capital stack means the frontier providers can absorb a price war if they choose to start one. They won’t lose this fight on cash. But your renewal cycle won’t wait for them to spend it, and the gap you’re staring at right now in your spreadsheet — $25 versus $3.48 for a token of comparable output — is the gap you’ve got to make a decision about. Whether it holds at that size in twelve months is genuinely hard to call. It’s the gap that exists today.
So the practical question is which workloads actually move. Treating this as a single decision (keep paying frontier prices, or move everything to the cheaper tier) and then picking one is a mistake. The right move is to segment your workloads and price them honestly. There are four categories worth thinking about.
Four workloads, four answers
1. Workloads where US frontier pricing still earns its keep. These are the ones where the marginal model quality genuinely matters: legal review where the wrong answer costs a settlement, security analysis where missing a vulnerability matters more than the API bill, customer-facing copy where the brand can't afford a hallucinated claim, anything externally exposed in a regulated workflow where the model's output reaches a customer or a regulator without a human in the loop. The tokens are usually low-volume and the downside risk is high. If anything, you should pay extra here. The price of the API premium is small set against the cost of the mistake the cheaper model makes once a quarter.2. Workloads where the DeepSeek-class tier is now obviously correct. Think high-volume, structured, low-stakes work where the answer is wrong-or-right and the model either delivers it or doesn’t. Bulk classification. Document tagging. First-pass code review on internal commits. Generating synthetic test data. Routine summarisation. Batch back-office work where the output goes into a queue rather than out to a customer. Anything where you’re spending five figures a month on tokens that could be spent at a seventh of the rate without changing the business outcome. Move them. The procurement question here is purely about timing — how fast you can do the data residency and weights-provenance work to make it allowed.
3. Workloads in the messy middle. This is most of them, honestly, and this is where the planning work happens. Internal-facing agents that draft emails, populate CRM fields, summarise meetings, or run multi-step research with a human reviewing the result before it’s used: none of these is high-stakes per call, but the aggregate quality matters because it shapes how staff perceive the AI tool over time. Rarely is the right answer to switch outright. The work is to test whether the cheaper model degrades the experience at the level your users actually notice. For a lot of workloads it won’t, and the savings are real. For some it will, and the user trust costs more than the saved tokens. You’ll only know which is which by running both models side-by-side on a representative slice of your traffic for a couple of weeks.
4. Workloads governance won’t allow. This is the hidden tax, and it’s the part of the conversation procurement readers should think hardest about. Running a Chinese-trained model means absorbing a buying conversation about data residency (where do the inference requests get processed and where are the logs stored), weights provenance (who trained the base model and on what), sanctions risk, model-card and incident-disclosure obligations, and what your customers’ contracts say about cross-border processing. For some industries — UK financial services, NHS-adjacent health, defence-supply-chain — the answer will be no, regardless of price. For others it’s genuinely unclear how this will land; there’s no precedent for a Chinese model getting widespread enterprise approval in the UK or EU yet, and procurement is going to make that call industry by industry. The number worth putting in the budget here is the cost of running the parallel governance assessment to find out which workloads are actually allowed to move. That work needs to start in May, not Q3.
This is what ‘rewrite, not renegotiate’ actually means in practice. Last year’s AI budget assumed frontier US pricing across the board, with the only debate being how much of that line to spend. This year’s budget needs four lines instead of one: frontier-justified, cheap-tier-obvious, cheap-tier-pending-test, and governance-blocked, with separate volume assumptions and a separate scenario column for each. If you negotiate your 2026 enterprise contract on the old shape of that line item, you’ll be locked into a price you no longer need to be paying for half of what’s flowing through it.
Cheaper tokens widen the 93/7 gap
There's a wider connection too. The case that AI isn't a bubble — that agents are running far more inference per task than chatbots ever did, so demand keeps catching up with supply — was already strong, and I argued the case at length in this earlier piece. The price of those tokens has just dropped by something close to a factor of seven for a chunk of what agents actually do. The buyer benefits twice over: more calls per task, cheaper calls per token.The risk is that the saving disappears straight back into infrastructure rather than into the work that actually shifts AI maturity. I’ve called this the 93/7 problem — most enterprises currently spend 93% of their AI budget on infrastructure and only 7% on the people who’d make that infrastructure useful. Saved API budget that doesn’t move into training your team and rebuilding your processes is just margin you’ve handed back to the model providers. The whole point of the cost relief is to fund the unphotogenic work that shifts the maturity needle.
Three actions before your next AI renewal
The most urgent piece of work is the workload inventory. Ask your AI lead to produce one by the end of May, sorted into the four categories above, with current monthly token spend on each. Without that you're guessing about the shape of the saving and the size of the move. After that, get your procurement and legal teams scoping the data residency and weights-provenance work for the cheap-tier-obvious bucket: even if the answer comes back as no, you need to know it before Q3. And in parallel, pause any 2026 enterprise renewal conversation that locks you into a single-tier frontier price for more than nine months. The shape of the budget you're committing to is no longer the right shape.What’s really happened over the last fortnight is that the frontier has split into tiers. There is still a top-of-line tier where US pricing buys real edge, and that tier still earns its keep on the workloads that justify it. There is now a credible second tier at a seventh of the price. And there is a third tier of governance friction sitting on top of both. The single ‘AI’ line item that worked in 2025 has stopped being the right shape, and the procurement teams who notice that first will spend the next nine months in a much more comfortable position than the ones who don’t.
