What AI Coding Agents Are Still Bad At (From Someone Who Ships With Them Daily)
I use Claude Code and Codex every day and I'd genuinely struggle to go back. That's exactly why the limitations are worth being honest about — not to talk you out of them, but so you keep a human on the parts they quietly get wrong. Here's where they reliably fall down.
They can't tell when they're wrong
This is the big one, and everything else is downstream of it. An agent hands you incorrect logic with precisely the same confidence as correct logic. There's no wobble in the voice, no "I'm not sure about this part" unless you engineered one. I've watched an agent write a calculator that computed the wrong formula and then write a fluent, reassuring explanation of the right one directly above it. You cannot outsource the judgment of whether the output is correct — you can only outsource the typing. The fix is unglamorous: tests, known-good values, and checks that run the real output instead of trusting the summary.
They lose the thread on big changes
On a tight, well-scoped task they're excellent. Ask for a sprawling refactor across twenty files and the seams show: they'll fix the thing in front of them and forget a constraint you set two files ago, or "finish" a migration with three call sites still pointing at the old code. They optimize locally and lose the global picture. The workaround is to be the one holding the global picture — break the work into pieces small enough that a lapse can't hide, and keep a running status of what's actually done.
They'd rather add code than say "no"
Ask an agent whether something is possible and it will almost never answer "it isn't." It will build you something — a workaround, a wrapper, a plausible-looking function — rather than push back. That eagerness is useful right up until it invents complexity you didn't need, or implements a bad idea faithfully instead of flagging it. A good chunk of my job is playing the skeptic the agent won't: "is this actually necessary, or did you just not want to tell me no?"
They forget your constraints unless you write them down
Anything not in front of the agent effectively doesn't exist. The rule you gave it last session, the file that must not change, the convention the rest of the codebase follows — all of it evaporates unless it's captured somewhere the agent reads every run. This isn't really a flaw so much as a fact of how they work, and it's why a good config file and a lessons log matter so much: they're the memory the agent doesn't have.
They have no free lunch on cost or context
Run agents hard all day and two limits show up: money and context. Long sessions burn tokens, and every model has a ceiling on how much it can hold at once — cross it and earlier context gets summarized or dropped, which is often where the "it forgot what we decided" moments come from. Neither is a dealbreaker, but both reward the same discipline: smaller tasks, written-down state, and not asking one session to hold the whole project in its head.
So why use them at all?
Because on the work they're good at — well-scoped changes, boilerplate, volume, exploration, the tedious 80% — they're a genuine multiplier. The trick is to stop expecting them to be good at the things they're structurally bad at. Keep the human on correctness, architecture, and the decision of what not to build; hand the agent everything else. Do that and they're one of the best tools we've got. Forget it, and they'll confidently ship you a mess.
Frequently asked questions
Are AI coding agents worth using despite their limits?
Yes. Used with judgment they are a large speed-up on well-scoped work. The point of knowing their weaknesses is not to avoid them — it is to keep a human on the parts they are bad at: verifying correctness, holding architecture steady, and deciding what not to build.
What is the single biggest weakness of AI coding agents?
They cannot reliably tell when they are wrong. An agent will hand you incorrect logic with the same confidence as correct logic, so you cannot outsource the judgment of whether the output is right — you need tests and checks that verify it independently.
Do bigger models fix these problems?
They shrink some of them, especially reasoning and context size. But the structural ones — no real feedback loop unless you give it one, a bias toward adding code, and confident wrongness — are properties of how the tools work, so the workflow around the agent still matters.
Get the workflow right, not just the model
Most of the value from AI coding agents comes from the process around them — verification, guardrails, and knowing what to hand off. I help teams build that. If you want yours dialed in, let's talk.
AI consulting with Patrick Bushe