A writer at The New Stack did something more useful than another feature review. She built a fake set of accounts — seven months of figures for an imaginary software consultancy — and buried twenty deliberate problems in them. Some were obvious: the company had lost money every single month, a cumulative net loss of nearly $135,000. Others were subtle enough that you'd need a trained eye to catch them. Then she pointed Claude for Small Business at the file and asked it to find everything that was wrong.

It found seventeen of the twenty, in under six minutes. Every easy problem, every medium one, and five of the eight hardest. If you were marking it like a school test, that's a solid 85%. Which sounds like a pass, and that's exactly where I'd stop you.

But 85% isn't a mark you can live with when the missing 15% is the part that matters most. The three problems Claude didn't catch weren't random. They were the most forensic items on the list — the ones a seasoned accountant spots by asking why something looks too clean. Interest income that comes in at exactly the same figure for seven months running doesn't add up wrong; it adds up too neatly, which is what tells an expert that someone keyed it in by hand rather than a real bank paying it out.

And the tool won't tell you what it missed. The gaps don't arrive flagged as gaps. They simply aren't in the summary, aren't in the deck, aren't in the email that goes off to your accountant. A silent omission is a very different animal from a visible mistake, and it's the reason I'd test this thing hard before trusting it with anything that counts.

What it actually does

Claude for Small Business launched in May. It's essentially a set of connectors that let Claude work inside the tools you already run: QuickBooks, HubSpot, Canva, PayPal, Google Workspace and a handful of others, all from within Claude Cowork. You toggle it on, connect your accounts, and pick a job. It'll reconcile your books, chase overdue invoices, draft a campaign or close the month, and it asks you to approve the plan before anything sends, posts or pays.

The speed isn't marketing. In that same test, Claude read nine tabs of financial data, wrote a plain-English summary, built an eighteen-slide deck in Canva and drafted a covering email, all in around twenty minutes. That's work that would take most of us the better part of a week. It even noticed things nobody had planted: a commission plan paying out on bookings rather than gross profit, and a typo in the file name. For an owner doing the books at nine at night, that's real value, and I don't want to undersell it.

Where it falls short

Claude is strong on the questions that have a definite answer: is this number wrong, does this total reconcile, has that cost line jumped. It's weaker on the questions that start with a hunch. Spotting that seven identical monthly figures are suspicious, rather than simply correct, is a different kind of thinking, and it's one the tool doesn't reliably do yet. It's one test, mind, and I wouldn't hang everything on a single fake P&L. But one test is enough to show you the kind of mistake you need to plan for — the quiet kind, the sort that never puts its hand up.

Stacked bar chart: Claude caught all 12 easy and medium planted accounting problems and 5 of 8 hard ones, missing 3
The easy and medium problems all came back; three of the eight hardest didn't — and those were the forensic ones.

For most of what an owner does day to day, catching the obvious and the moderately tricky is plenty. But the moment the stakes rise — a year-end, a funding round, a sale of the business — that forensic 15% is exactly the part you were paying an accountant to find. And you won't learn from Claude's output that it's gone missing, because a tool that drops something still hands you a report that looks complete.

Stress-test it before you trust it

The good news is that the test itself is the template. You don't have to take a vendor's demo or a reviewer's verdict on faith. You can run the same experiment on your own numbers, and it takes an afternoon.

  1. Feed it books where you already know the answers. Take a closed prior year, or a month your accountant has already signed off. Give it the same brief the reviewer did: read every tab before drawing conclusions, flag every anomaly and inconsistency no matter how small, and say what it would ask you if it were your finance director. Then set its list beside the one you already hold in your head.
  2. Salt the file deliberately. Copy a real set of accounts and plant, say, eight to ten known problems across a range of difficulty: a duplicated invoice, a margin that collapses, a suspiciously round figure that never moves, a cost line that stops for no reason. Write each one on a separate sheet before you start, then count how many come back. The easy ones will; what you care about is its hit rate on the hard ones.
  3. Write down what it didn't say. The valuable part isn't the list Claude produces; it's the gap between that list and reality. Keep a running note of the categories it misses — the forensic, too-clean, why-is-this-so-flat kind — because those are exactly what it'll miss on live data, and those are the ones you check by hand at every close.

That gap between the machine's answer and the real one is where your own judgement earns its keep. It's the judgement gap I've written about before, and the reason these tools reward the owners who already know their numbers and quietly punish the ones who don't. Run the experiment once and you stop guessing about whether the tool is any good. You know — for your books, in your hands.

Scope what it can reach

The second discipline is about access rather than accuracy. These connectors reach into live systems: your accounts, your CRM, your outgoing email. So accuracy is only half the question. The other half is reach — what Claude can touch when it gets something wrong.

Anthropic has set sensible defaults here. Your existing permissions carry over, so if a member of staff can't see something in QuickBooks today, they can't see it through Claude either, and nothing sends or pays without your say-so. Lean on that, and widen its remit in stages rather than all at once. Let it read and analyse first, and nothing more. Once that's earned your trust, let it draft (invoices, summaries, campaign copy) but stop it short of sending. Only then let it act on your explicit approval, one task at a time, and keep it clear of anything that actually moves money for a good while yet. The approval step is where all of this lives or dies, so treat it as a real review and not a rubber stamp; a human in the loop is worth nothing the moment the human stops actually looking.

It's the same instinct behind grounding an AI in your own data instead of its training, which is what I wrote about with NotebookLM as a strategy copilot. A tool working from your real numbers, under your real permissions, with you checking its output, is one you can build on. Hand it the keys on trust alone and you've got a liability wearing a productivity badge.

The bottom line

None of this is a reason to pass on Claude for Small Business. It's a reason to bring it in the way you'd bring in a fast, capable, slightly overconfident new hire — give it real work, check what it gives you back against things you already know, and widen its remit as it earns your confidence. Handle it that way and you get most of the upside with very little of the risk, which is about the best any of us can ask of a tool this new.