Shadow agents: the enterprise governance crisis nobody planned for
Shadow AI agents running on staff laptops with live credentials to company systems. KiloClaw named the crisis — and most governance is lagging behind.
On 1 April, Kilo quietly launched KiloClaw for Organizations. The pitch was simple, and slightly awkward if you run enterprise IT: more than 25,000 developers were already using the consumer version of KiloClaw at work, often on personal machines, connected to company Slack, GitHub, calendars and internal APIs through credentials they had added themselves.
In other words, Kilo wasn't selling an enterprise agent product. It was selling a clean-up service for an agent fleet most companies didn't know they already had.
The framing in the press coverage was neat: the era of shadow AI is ending. I don't think that's true. Shadow AI isn't ending. It's being discovered, counted, packaged, and rebranded as a procurement opportunity.
A few weeks ago I would have framed this as an infrastructure story. I'm not so sure any more. The more interesting story is what was already running, on personal kit, before the enterprise product arrived to triage it.
The Claude Code leak changed the risk
For most of 2025, shadow AI looked like ChatGPT tabs. Employees quietly pasting documents, customer records, and draft emails into a free LLM to speed something up. The risk was real but fairly easy to understand: a human copied something in, the model produced something back. You could trace where the data went because a human was sitting in the loop.
Shadow agents are different. They don't just answer. They act. A shadow agent is a long-running process with credentials, a loop, and a goal. It has access to the user's GitHub, calendar, documents, possibly production services. It's making API calls autonomously, on infrastructure the company neither owns nor audits.
The Claude Code leak at the end of March is what made this legible. On 31 March, Anthropic pushed version 2.1.88 of their npm package with a 59.8 MB source map accidentally bundled in. The map pointed at an unprotected archive in their own Cloudflare storage. Within hours, mirror repos were everywhere. The full unobfuscated TypeScript source of Claude Code — thousands of files, hundreds of thousands of lines — was suddenly available to anyone who wanted to inspect it.
GitGuardian's State of Secrets Sprawl 2026, out two weeks before the leak, reported that Claude Code-assisted commits leaked secrets at 3.2% versus a 1.5% baseline. More than double. At the time that felt like a chart you'd file under "things to monitor." After the leak, the meaning shifted. Researchers and bad actors now had the harness. They knew exactly how the tool parsed repos, which files it trusted, and which commands it would run without confirmation. The attack surface of every shadow agent running that code became much more specific overnight.
Shadow IT, with credentials and autonomy
The pattern underneath is familiar from every previous wave of shadow IT. It just got faster.
Employees find the official tools slow, locked down, or missing entirely. They find something better that runs on their laptop. They wire it up to their work accounts because that's where the data is. They get their job done. Nobody complains because the output is good. Everyone else on the team starts asking how they did it.
None of this is malicious — it's just what humans do when the official option is friction and the unofficial one is superpowers. The Gartner number floating around says 40% of enterprises will report a security or compliance incident traceable to unauthorised AI by 2030. I'm mildly suspicious of that kind of precision in a forecast running four years out, but the direction is hard to argue with. IBM's 2026 breach cost data puts shadow-AI-involved breaches about $670,000 above the baseline. That at least suggests somebody is starting to count.
What's different with agents is the persistence. A ChatGPT tab closes when you shut the lid. An agent loop doesn't. Some of them — KiloClaw, Claude Code, the various Claw-harnesses — are explicitly designed to run overnight, pick up scheduled work, and message the user when something interesting happens. Useful, yes. But it also turns a personal productivity hack into a standing exposure.
I underestimated this too
I didn't properly see this coming either.
When I wrote about Claude Managed Agents as an infrastructure moment earlier this month, the frame was that the friction around running agents had collapsed and that was mostly good news for the developer. It still is. What I didn't push hard enough on was that once the friction collapses, the agent shows up everywhere — including in places nobody at the top of the organisation has clocked. The managed version is the well-lit case. The shadow version is what was running first.
That's the pattern with every productivity technology. It lands in the laps of the people doing the work before the people running the org have noticed. IT used to get the luxury of a procurement cycle to catch up. With agents the gap between "my developer tried a tool" and "our data is running through a third-party harness on their home Wi-Fi" is about fifteen minutes.
Why procurement won't solve this on its own
The tempting response, if you're running governance, is to lock everything down. Approved tools list. Blocklist the unapproved ones. Training module for the staff. Done by Christmas.
That might have worked for SaaS shadow IT a decade ago. It won't work here.
The tools the developers are using are, most of the time, strictly better than the approved alternatives. A blunt blocklist tells your best engineers that management has decided they should be less effective at their jobs. They will find a way around it, and the way around it will be even less visible than what they were doing before.
And the agents aren't going to stay contained to engineers. Finance teams are already using AI to reconcile ledgers. Sales ops people are wiring up agents to triage leads. Legal juniors run contract review through Claude at two in the morning because the deadline is at nine. The BYOAI footprint is already cross-functional and growing faster than any central function's ability to enumerate it.
KiloClaw's commercial pitch — SSO, centralised billing, policy controls, BYOK — is a tacit admission that the game has changed. You don't prevent shadow AI at this point. You absorb it. You offer people the thing they're already using, with authentication and logging bolted on, and you hope the delta in friction is small enough that they come willingly.
That's how this pattern always resolves, eventually. It resolved this way for cloud storage (Dropbox, then sanctioned cloud drives), for messaging (WhatsApp groups, then Teams and Slack), for personal devices (the corporate laptop people ignored, then BYOD with MDM). Each time the unofficial tool won enough ground that the only rational play was to give people a sanctioned version of it.
Agents will go the same way. The CIOs treating shadow agents as a first-class category of IT risk — alongside cloud misconfigurations and unpatched endpoints — will adapt. The ones still treating "AI policy" as a question of restricting ChatGPT on the corporate network are going to have a rough 2026.
What the CIO reading this should actually do
Some of this is boring blocking and tackling that most security teams already know how to do. Audit outbound API calls. Look for token usage patterns that don't match interactive work — long-running, late-night, weekend spikes are the tell. Enumerate which SaaS apps have "connected apps" tokens in user accounts and at least know what's been authorised.
The harder part is who you're actually looking at. The real risk isn't the junior developer running Claude Code on their MacBook. It's that the person running the most-used shadow agent in the company is probably one of your best people, solving a real problem the organisation hasn't resourced. Treating them like a compliance threat will cost you the person. They're probably telling you something useful: where the sanctioned stack is too slow, too limited, or simply absent.
This connects to a point I made in the 93/7 problem a few days back. Nine out of every ten pounds of AI spend is going into infrastructure. The tiny remainder is meant to cover the people side — training, workflow redesign, role reshaping. Shadow agents are what you get when organisations push infrastructure onto people without rebuilding the path through. The tool lands, the approved workflow for it doesn't, so people build their own. The agent is the clue. It shows where the operating model hasn't caught up.
The honest version of agent governance isn't "we'll catch them and punish them." It's "we'll figure out what they're trying to do, and we'll either give them a sanctioned way to do it or admit we don't have one yet."
The fleet is already running
I could be wrong about this — and I'm genuinely not sure how fast the window is closing. But I think most large organisations are running out of time to know what's happening inside their own agent fleet, faster than the procurement cycles that would let them react. KiloClaw for Organizations and the products that will follow it are triage, not prevention. The fleet is already running. The data has already moved. The question is whether your product teams bring you inside the tent or whether you find out in a post-mortem.
I'm not arguing against agents. Of course companies should use them. The productivity gains are real. The business cases are real. Waiting while competitors move faster isn't a serious strategy.
But pretending the unmanaged version isn't already running is worse.
KiloClaw didn't create that crisis. It just built the first product to admit it out loud.
