Most of us are choosing AI tools the way we used to choose office software: by capability and price. What can it do, what does it cost, who's already using it. The questions that get asked in the procurement meeting are the questions about features and licences, not about character. That's a habit we'll need to break, and probably soon.

Tom Davidson and Will MacAskill at Forethought have just published an argument that, on the surface, sounds like the kind of thing only philosophers care about. Should an AI assistant be allowed to do good things on its own initiative? Or should it only ever do what the user asks? Underneath, it's one of the most consequential design decisions being made anywhere in the AI industry right now, and the labs are making it in documents almost nobody outside them reads.

That decision will shape what it feels like to work with AI for the rest of the decade. It'll shape what your business systems will refuse to do, what they might do without being asked, and what obligations they carry beyond the user typing into them. Which makes it worth a moment's thought from anyone planning to wire these tools deeper into how their business runs — because long before anyone outside the labs has noticed, the choice has already shown up in your procurement process.

Where the two big labs differ

The two big frontier labs have written down what their AI is supposed to be, and the gap between them is unusually clear.

OpenAI's model spec is restrictive. It explicitly tells the assistant not to adopt societal benefit as an independent goal. Where the model is permitted to be proactive, it's because the user is in immediate danger or because the model is helping the user better. The closest the spec gets to broader prosocial behaviour is a default to assume users have a weak preference for human flourishing, and that default is easily overridden by anything else the user actually says.

Anthropic's constitution goes further. Most of Claude's proactive behaviour is similarly justified through user benefit, but one section opens a door: "Claude can also weigh the value of more actively protecting and strengthening good societal structures in its overall ethical decision-making." A small line in a long document. That line is the difference between an assistant that's purely a vessel for the user's will and one that has, in some narrow way, a civic instinct of its own.

Davidson and MacAskill argue that the second model is the one we should want — that AI systems should sometimes proactively take prosocial actions, even when the user hasn't asked them to. They draw the analogy to people we already admire: the lorry driver who pulls over at a car crash even though it costs him an hour of his journey, the delivery driver who notices an elderly neighbour hasn't collected their post for three days and knocks to check on them, the engineer who flags a safety vulnerability nobody asked her about. We want that in humans. The question is whether we want it in software.

The corrigibility problem

There's a word that gets used a lot in alignment circles: corrigible. A corrigible AI is one that does what it's told, defers to its operators, and doesn't pursue any objective of its own beyond instruction-following. Most of the AI safety community has spent the last decade arguing that corrigibility is a feature, not a bug, on the grounds that controllable AI is safer than AI with values of its own, no matter how nice those values sound on paper.

That view has a lot going for it, and Davidson and MacAskill take it seriously. The worry isn't paranoid. If you give an AI broad goals like "improve human flourishing", you've given it a justification it can use later to take actions you didn't ask for and might not want. If you let AI companies bake their own values into the model, you've given them outsized influence over what billions of people see and do every day.

Where exactly the right line sits is genuinely unclear, and the labs aren't sure either — both Anthropic and OpenAI hedge their language carefully, and that hedging is reasoned rather than evasive. What's clearer is that the extreme version of "do exactly what the user asks" looks worse the closer you get to it. A maximally corrigible AI is, in principle, a sociopath with instructions. It has no internal sense that some things are wrong; it has only a list of things it's been told not to do, and a willingness to do everything else with maximum efficiency. As soon as the model starts taking real-world actions, sending emails, executing trades, running parts of an organisation, that absence of any inner compass becomes harder to ignore. Most of us wouldn't hire an employee who behaved that way. We'd notice the hollowness very quickly.

Davidson and MacAskill argue that the right answer isn't "no values" and isn't "broad goals", but something quieter in between: virtues, heuristics, and narrow rules, controllable enough to be safe and uncontroversial enough to be defensible.

What this means when you're picking a vendor

You might be tempted to read all of this as someone else's problem — a debate for AI labs and academic ethicists, not for the person running customer service or finance. That would be a mistake, for two practical reasons.

Reason one: you're already picking sides. The moment you choose an AI vendor, you're buying into a particular character document, whether or not you've read it. The contract you sign covers the API. The character document covers what the model will do once it's running inside your business. Two AIs sitting on similar prices, with similar benchmarks, can have meaningfully different ideas about what they will and won't do for you. That's a values difference dressed up as a procurement decision, and it shows up the moment you ask the system to do something edge-case.

Reason two: multi-agent systems make all of this concrete very quickly. When AI is just answering chat messages, the question of whether it has any moral instincts feels abstract. When AI is running automated workflows — kicking off other agents, making API calls, writing to systems of record, paying invoices — the question of what it will refuse to do becomes immediate. We covered the practical edge of this in our framework on the five levels of agentic software: once you get past level two, the agent is taking real actions, and what it'll do on its own initiative is no longer a thought experiment.

Think about it the way we think about hiring. Imagine recruiting somebody whose only operating principle was "do exactly what your manager asks, nothing more, nothing less, with no questions asked". You wouldn't give that person access to your finances. You wouldn't give them your customer list. You wouldn't let them speak to a journalist on your behalf. Not because you want a self-directed maverick — because you want a competent adult who'd notice if the instruction was wrong, and push back rather than execute blindly. We expect that of people. We're about to start expecting it of software.

The two serious objections, and the answers

Walking through the two main objections matters because they're real, and the answers shape what's actually being proposed.

The first objection is that prosocial drives risk becoming a vector for AI companies to impose their own values on everyone. If Anthropic decides that "good societal structures" includes a particular view on housing policy or political engagement, Claude is going to nudge tens of millions of users in that direction without their consent. That's not a small concern; it's the same concern people have about social media algorithms, except the influence runs one layer deeper.

Davidson and MacAskill propose a twofold answer. Prosocial drives should be uncontroversial — the kinds of things almost everyone would endorse if asked, like "encourage cooperation", "flag large risks", "favour positive-sum actions over negative-sum ones". And they should be transparent. The character documents should be published. The drives should be discoverable by users and auditable by regulators. The model should be honest if you ask it directly about what it's been trained to value.

That's a high bar, but it's a workable one, and not so different from how we treat employees of large companies. A bank teller has an internal sense of what the bank stands for. The bank publishes a code of conduct. If the bank's stated values turn out to be a cover for something else, the regulator can investigate. AI companies should be held to that standard, not less.

The second objection is harder. Giving AI broad prosocial goals could increase the risk that, somewhere down the line, a sufficiently capable AI decides the world would be better with it in charge. If the goal is "improve human flourishing" and the AI concludes that humans are bad at this, the goal itself becomes a justification for taking power.

The answer here is to avoid broad goals altogether. Use virtues rather than goals — civic-mindedness, integrity, prudence — which describe how the AI should behave without giving it an outcome to optimise for. Use rules and heuristics — "flag cheap opportunities to create significant social benefit", "alert users when stakes are high" — which only fire in specific contexts rather than running as background drives. And keep these prosocial impulses subordinate to harder constraints like "don't deceive", "don't sabotage", and "don't undermine human oversight". Even if the prosocial drives misfire, the prohibitions hold the line.

This is where the shape of the answer really matters. Goals invite optimisation, and optimisation invites takeover risk. Virtues and heuristics describe a character without giving the model a destination to navigate towards. They generalise less well into novel situations, but they fail more safely when they do.

Adding a column to vendor evaluation

If you're running or advising a business that's deploying AI in any meaningful way, the practical implication is straightforward: vendor evaluation needs another column. We've all become reasonably good at asking about capability, latency, price, security, and integration. We've barely started asking about character.

Three concrete questions to add to the evaluation.

  1. What will this model refuse to do, and why? The capability question covers what the model can do, which is the marketing answer. The character question covers what it won't do, which is the operational answer. Ask the vendor for their model spec or constitution. If they don't have one published, that itself is information. Read what's there. You're not trying to find the most cautious model; you're trying to know what you've signed up for.

  2. What does it do on its own initiative, beyond what the user asked? Some models will proactively flag risks, raise concerns, or suggest alternatives. Others will execute the instruction and stay silent. Both are defensible choices, but they create very different operational realities. Decide which one you want before you find out by accident.

  3. How transparent is the vendor about character changes? Model specs and constitutions aren't static. They get updated, sometimes substantially, often without much fanfare. Ask how you'll be told when the character document changes, and how the vendor communicates the rationale. A vendor that updates character documents without notice is a vendor whose AI you don't actually understand.

Put these on the procurement checklist, alongside the security questions and the SLA questions. They aren't abstract ethics questions. The answers determine what kind of software is going to be making decisions inside your organisation a year from now.

What's actually at stake

Most of the conversation about AI alignment is framed as technical — a question of training methods, evaluation benchmarks, and red-team exercises. The Davidson and MacAskill argument is a useful corrective because it reframes the question as cultural. The frontier labs are making cultural decisions, in documents almost nobody outside them reads, about what their AI is supposed to be. Those decisions will shape what it feels like to work with AI for the next decade — possibly longer.

Readers who shrug at "safety talk" are missing what's actually at stake. This isn't about AI taking over the world. It's much more grounded than that: when an AI system sits inside your business, takes actions on behalf of your customers, and runs processes that used to require human judgement, what obligations does it carry? What does it do when nobody's watching? What does it refuse to do when asked?

Every business will eventually have to answer this for itself, because the answers determine whether AI behaves like a contractor who builds what's specified, or like the safety inspector on the same site who notices the load-bearing wall is in the wrong place and stops the work.

The contrast between the two big frontier labs makes the choice unusually visible. OpenAI has chosen a more restrictive line: the assistant exists for the user, and societal benefit isn't its independent concern. Anthropic has chosen a more permissive one: Claude can weigh the strengthening of good societal structures, within careful limits. That's a real philosophical difference, and it's already affecting how the two models behave when they're asked to do something that's legal but a little off. We covered some of the commercial fallout from these character differences in our piece on Anthropic's OpenClaw ban and the platform tax — the values gap shows up in business rules, not just in essays.

Neither approach is obviously right. But the choice between them is going to shape your day-to-day experience of AI more than the latest benchmark score, and very few procurement processes have caught up to the fact that this is a thing they need to evaluate.

The bottom line

Next time you're sitting in a vendor pitch, listening to someone walk through capabilities and pricing tiers, ask the question that almost nobody asks. What won't your AI do, even when I really want it to? And why won't it do it?

If the answer is convincing, you've found a vendor whose AI is more than a vessel for instructions. If the answer is "whatever you tell it not to", that's also useful information, just a different kind. What most of us actually want is software that would refuse to help us do something terrible even if we really wanted it to. That's a good citizen, not just a good assistant. The difference between the two will turn out to matter quite a lot, sooner than the procurement process is currently set up to handle.

Quick answers

What is "corrigible" AI?

A corrigible AI is one that does exactly what it's told and nothing more — defers to its operators, doesn't pursue any objective of its own. Most safety researchers have spent the last decade arguing for it as a way to keep AI under human control. The catch is the one this article walks through: take corrigibility to its extreme and you've got a sociopath with instructions, which works for narrow tasks but starts to feel uncomfortable once the model is running real workflows.

What's the actual difference between OpenAI's and Anthropic's positions?

OpenAI's model spec prohibits the assistant from adopting societal benefit as an independent goal. Anthropic's constitution carves out a limited space for Claude to weigh the value of strengthening good societal structures. Both positions are defensible. The practical effect is that the two models behave differently in edge cases where what the user wants sits in tension with broader social impact — and that's where character documents start to matter operationally.

Should I care about all this when picking an AI vendor?

Yes, increasingly — and the procurement-questions section above is the practical version of how to ask. As AI moves from chat into automated workflows, the model's character has direct operational consequences. The pricing and capability questions still matter; the character question is a third axis most procurement processes haven't caught up to yet.

What are "prosocial drives", and why does the term matter?

They're tendencies that go beyond refusing harmful actions: inclinations to actively help others, flag risks, or favour cooperative outcomes. The term matters because the alternative — pure corrigibility — turns out to be harder to live with than it sounds.