Choosing an AI assistant for your business this year runs into something odd almost immediately: every demo is the same demo. The agent reads your files, drives your browser, remembers what you did yesterday and hands back finished work instead of suggestions. Claude does it. So do Microsoft's Copilot, OpenAI's Codex and Perplexity's Computer. The screen changes colour and the logo changes; the pitch doesn't.

There's a reason for the sameness, and there's no clear winner coming that you'd be wise to wait for. The New Stack counted six labs shipping essentially the same knowledge-worker agent inside four months: Microsoft, Anthropic, Google, OpenAI, Amazon and Perplexity. Anthropic's Claude Cowork opened the run in January; Perplexity, Microsoft and OpenAI had all shipped their own versions by April. Same idea, near enough the same season.

So the question most of us are still asking — which one is best? — has gone stale. When the tools have converged like this, “best” is a benchmark word, and benchmarks change hands every few weeks. What you actually want to know is duller: which of these fits the way my business already works? That's the whole article. The feature list has stopped being a useful way to choose, and most of what should replace it has nothing to do with the model at all.

How we got here matters, because it tells you the sameness is real rather than marketing. Claude Code, Anthropic's coding tool, showed what a capable model could do once you wrapped it in software that runs jobs and edits files rather than only answering questions: real work, not just talk about it. That wrapper is what the trade calls an agentic harness. Developers fell for it hard, and every lab watching asked the same question: why should this stay a developer tool?

So every lab made the same bet: take the tool developers loved and point it at everyone else's work too. Anthropic's commercial lead, Kate Jensen, put it plainly to CNBC: “We expect that every knowledge worker will feel that way about Cowork,” the way engineers already feel about Claude Code. By the account of Boris Cherny, who runs Claude Code, the agents aren't staying in their lane. One bet, six labs, one product.

The four that actually matter

So if the feature list is a dead end, what do you choose on? Four things, none of them glamorous.

  1. Workflow fit. Where does the work actually happen in your business? If you run on Google Workspace, the assistant that reads a Google Doc without a fight beats the one with the higher score on some leaderboard; if your people live in Microsoft, the reverse holds. Fit is mostly about how little has to change for the thing to be useful on Monday morning, and the tool that asks everyone to move house first tends to get abandoned by Wednesday.
  2. Switching cost. How hard is it to leave once you've started? For most small teams you won't switch often anyway (running thousands of agents in parallel is a big-company concern, not yours), but the prompts your people refine, the connections you wire up and the habits they build are real investment, and far lighter to carry elsewhere if the tool doesn't trap your data and history inside itself. I went through the layers of this in the part you're not choosing: a tool that lets you swap the model underneath and connect outward to your other software leaves you far less exposed to one vendor's roadmap and one vendor's pricing.
  3. Data gravity. Your data has weight. The more of your real work — documents, customers, accounts, the last three years of email — that already sits in one company's orbit, the more an assistant from that same company can reach without you building bridges to it. It's why the cleverest standalone tool can lose to the merely-good one parked next to your stuff: in the demo everything connects, but in your business only what's already within reach does.
  4. Which assistant your team will actually open. The tool that gets used beats the tool that benchmarks well and sits idle, every time. A team that genuinely lives in one assistant, day in, day out, gets far more from it than a team hopping between three because each topped a different leaderboard that month. I've argued before that the gains come from depth of use rather than breadth of tooling: the firm that masters one capable tool out-produces the one with a drawer full of half-learned ones.

Put those four together and the decision gets refreshingly small. Start with whichever assistant already sits closest to where your work lives (the Microsoft stack if you're a Microsoft shop, Google's if you're a Google one), and give it the real jobs for a fortnight. If it reaches your files and your people actually open it, you're done. Only go looking further if there's a specific job it genuinely can't do, and then weigh that one gap against everything you'd lose by moving. That's the whole filter, and it has nothing to do with which model is winning this week.

Don't wait for the winner

There's a strong pull to hold off — to wait until the field settles and an obvious champion emerges. It won't. The labs are leapfrogging each other on a monthly cycle and they'll keep doing it; whoever's top of the table today is back in second place by the next release. Waiting for the winner becomes a way of never choosing at all. And the cost of not choosing is real: your team carries on doing by hand the work an assistant could have absorbed months ago, while you wait for a verdict that never arrives.

Choosing early isn't risk-free, mind. You might back a tool that slips behind the pack, and it's hard to say which ones will still be ahead in a year. But fit and data gravity tend to hold their value even as the models trade places, so a sensible early choice usually ages better than it looks on the day.

The freedom buried in all this sameness is that you've been let off a decision you were never well placed to make. You don't have to identify the cleverest model — you couldn't reliably pick it out anyway, and it would change next month regardless. You only have to find the one that fits how your business already works and that your team will genuinely use. Settle on that, and you can let the benchmark wars carry on without you.