In seventeen years of running an agency I must have interviewed hundreds of people, and I judged them the way we all do: give someone a few hard questions, watch how they handle the awkward follow-up, and trust the rest to follow. That’s a reasonable way to assess a human being, because human skills tend to cluster — a person who can build a discounted cash flow model can almost certainly add up a column of figures, and someone who writes brilliant strategy papers isn’t going to be defeated by an email. The skills sit next to each other, and the hard thing usually contains the easy thing.

AI breaks that assumption, and I think a surprising amount of the frustration in businesses adopting it right now comes back to this. We interview these tools — a dazzling demo, a spectacular first week — and then hand them the whole job. Sooner or later the tool fails at something a school leaver would find trivial, trust collapses, and the business swings from over-delegation to abandonment without ever finding the useful middle ground where the real value sits — a middle ground that’s only ever found one task at a time.

A name for the problem

The pattern now has a name. Andrej Karpathy, a founding member of OpenAI, coined the term “jagged intelligence” in 2024 to describe the strange, unintuitive fact that state-of-the-art models can solve complex maths problems while simultaneously struggling with very dumb ones — his example at the time was models that confidently insisted 9.11 is a bigger number than 9.9. Demis Hassabis has been using the same phrase this year to explain why he doesn’t consider current AI to be general intelligence at all. Speaking at the India AI Impact Summit in February, the Google DeepMind chief described today’s systems as “very good at certain things, but they are very poor at other things, including sometimes the same thing” — models that can perform at gold-medal level on International Mathematical Olympiad problems and still make elementary arithmetic slips when a question is phrased awkwardly.

Hassabis runs one of the world’s leading AI labs and has every commercial incentive to describe his products generously, which makes the admission more striking. The capability of these systems is a mountain range: peaks of superhuman performance sitting beside crevasses of surprising incompetence. My working theory on why this catches so many of us out is simply that our intuition about ability was trained entirely on people, and the peaks and crevasses follow no pattern that intuition recognises.

The research has been saying this for a while

The best evidence we’ve got on what this means in practice is still a study Harvard Business School ran with Boston Consulting Group back in 2023. The researchers gave 758 BCG consultants a set of realistic tasks and called the boundary between what AI could and couldn’t do well the “jagged technological frontier”. For tasks inside the frontier, consultants using GPT-4 produced work rated over 40% higher in quality and finished more than 25% faster. For tasks deliberately designed to sit outside the frontier, consultants given the same assistance were 19 percentage points less likely to reach the correct answer than colleagues working unaided.

It’s the same tool and the same intelligent professionals, yet the outcome swings from a substantial performance boost to actively making experts worse. The researchers had chosen the two kinds of task to look similar in difficulty, and that was the point: the output reads just as fluent and confident either side of the frontier, so nothing about the experience warned the consultants which side they were on. The people who did worst had simply extended the sort of trust we’d extend to any colleague who aced the interview, which is precisely the mistake our instincts teach us to make. With AI, impressive performance at one task is no guide to the task next door; each one has to earn trust separately.

Those percentages will have moved (the models are three years better) but the shape of the problem hasn’t, because Hassabis is describing in 2026 exactly what the researchers measured in 2023.

Working with the jagged edge

The practical response isn’t to trust these tools less overall — it’s to trust them at the level of the task rather than the level of the tool. A few strategies I’d recommend:

  1. Test the task, not the tool. Evaluate each specific task with your own real examples before it goes anywhere near production: a handful of representative cases, one or two awkward ones, and your highest-consequence scenario make a perfectly good test set. And don’t let performance on one task stand as evidence for another, however closely related the two appear.
  2. Map your own frontier. Keep a simple running log of what you delegated, what came back, and whether it survived checking. Within a few weeks you’ll have something far more valuable than any benchmark: a map of where the edge runs through your work specifically. Share it across the team, because your colleagues are currently discovering the same crevasses one painful surprise at a time.
  3. Set boundaries by cost of error. Decide what a tool is allowed to touch based on what a mistake would cost. A wrong first draft costs you a proofread; a wrong number in a client proposal costs you the client. I’ve written before about the five levels of agentic software, and for most businesses the right pace is autonomy expanding one consequence tier at a time.
  4. Re-draw the map when the model changes. Every release moves the edge. Some crevasses get filled in; occasionally new ones open where the ground was solid. Keep the test set from the first strategy and re-run it whenever the model, the prompt, or the connected data changes. As I found when revisiting the AI 2027 forecasts, the models have got more capable without getting more consistent, which means your map from six months ago is already wrong in both directions.
  5. Keep humans where verification is expensive. The safest tasks to delegate are the ones where checking the output is quick and cheap — a summary you can skim, code you can run. Where verification costs nearly as much as doing the work, or where you lack the expertise to check at all, the frontier problem bites hardest and a human should stay in the loop.

I don’t know how long the jaggedness will last. Hassabis argues that smoothing it out needs genuine research breakthroughs — continual learning, better reasoning and planning — rather than simply more scale, and he may be right, or the next generation of training techniques may surprise us all. For the foreseeable future, though, jagged is the operating condition.

The bottom line

The interview heuristic served us well for a century because human skills really do tend to cluster and transfer. AI competence doesn’t, and pretending otherwise is what produces both the horror stories and the disillusionment. The businesses getting real value from these tools are the ones who know precisely where their tools are brilliant and where they fall over, because they took the time to find out. So start the log this week — a plain spreadsheet with three columns will do — and you’ll soon know more about where AI can be trusted in your business than any benchmark or launch demo will ever tell you.