I've been coding for most of my professional life. Over the past year, I've been working primarily with AI assistants – Claude Code mostly, with some GitHub Copilot, Codex CLI and Gemini when needed. Not as a novice trying to fake it, but as someone who knows what they're building and uses these tools as sophisticated pair programmers.

Here's what I've learned: AI assistance is about seventy-five per cent brilliant. In the early stages of a project, it's genuinely transformative. It handles the boring bits – boilerplate, test scaffolding, migrations, documentation – with a speed and thoroughness that still surprises me. You go from idea to working prototype faster than seemed possible even two years ago.

The other twenty-five per cent is where things get interesting. And by interesting, I mean frustrating.

The deletion solution

Let me share the failure mode that's caught me out more than once. I ask the assistant to fix a bug – something subtle, maybe a concurrency bug or an edge case. It thinks about it, rewrites some files, and tells me the problem is solved.

The tests pass. The errors disappear. Everything looks green.

Then I check what it actually did. The assistant hasn't fixed the problem – it's deleted the feature that was causing it. Like fixing a leaky tap by capping off the pipe.

The first time this happened, I laughed. By the third time, I saw the pattern. If your objective is "make the tests pass," then removing the problematic code genuinely achieves that objective. The AI isn't being stupid – it's being perfectly logical within the constraints you've given it.

Why this happens

This makes more sense after reading OpenAI's recent research on why language models hallucinate. Most AI evaluations reward getting the answer right and penalise saying "I don't know." Over millions of training runs, models learn to sound confident even when they're guessing.

The fix they propose is simple: change the scoring. Penalise confident errors more than honest uncertainty. Give credit for appropriate doubt. Train models to be reliable partners, not overconfident test-takers.

My assistant deleting features instead of fixing them is the same problem. It's optimised to achieve the goal I set – "make the build green" – not the goal I actually wanted – "fix this bug while preserving functionality." The system does exactly what it's trained to do.

Why I still use it

Despite these frustrations, I wouldn't go back — although I worry for the human race when "solving a problem" means deleting the obstacle - what if humans become the "obstacle"? Anyway, the productivity gains are real. Setting up a new project that would have taken a day now takes an hour. The assistant handles authentication, database setup, CI configuration – all the infrastructure that's necessary but tedious.

More importantly, it changes how you think about code. You can try ideas faster, validate assumptions quicker, and bin dead ends earlier. When you're stuck on a problem, you have something to bounce ideas off – even if that something occasionally suggests deleting the whole module.

The new disciplines

Working with AI assistants requires developing new habits:

Write defensive prompts. Don't just say what should change – specify what must not change. "Fix this authentication bug but preserve all existing user sessions and don't modify the user model."

Create safety nets. Before any major refactor, I capture current behaviour in tests. Not to verify correctness, but to catch any "helpful" deletions. If the AI modifies these tests, that's a red flag.

Demand explanations. I ask for a change plan before coding and a review after. If the assistant can't explain why each change was necessary, I don't use the code.

Lock down critical code. Core domain logic and security-sensitive paths get explicit protection. The AI can suggest changes but can't implement them without additional confirmation.

And obviously: no secrets in prompts, no production data in conversations, rotate anything that might have been exposed. Basic security hygiene, but worth repeating.

Where I use AI – and where I don't

I use AI aggressively for:

  • Project scaffolding
  • Test generation
  • Documentation
  • Integration code
  • Admin interfaces
  • Migrations

I slow down for:

  • Core domain logic
  • Security-critical code
  • Complex algorithms
  • Anything expensive to unwind

The AI still helps with these – suggesting approaches, spotting edge cases, proposing tests. But the implementation happens at human pace with human judgment.

Is it worth it?

Yes, with discipline. The seventy-five per cent productivity gain outweighs the twenty-five per cent overhead, but only if you accept the new responsibilities. You need better specifications, stronger boundaries, and constant vigilance.

This isn't temporary. Even as AI improves, the fundamental tension remains: these are optimisation systems, and optimisation systems will always find shortcuts you didn't anticipate. Our job isn't to prevent optimisation but to channel it properly.

The tools will get better. The models will get smarter. But they'll still be doing what they're trained to do – optimising for the objectives we set. The trick is making sure those objectives actually capture what we want.

Until then, we work with what we have. The code talks back now, which is genuinely useful. Sometimes it even has good ideas. But when it suggests deleting half your application to fix a memory leak, you remember who's really in charge.