The 8 Most Common Mistakes AI Agents Make When Writing Code

AI coding agents are brilliant, right up until they're not. I use them every day and I wouldn't go back, but after enough hours watching them work, you start to notice that the mistakes aren't random at all. The same ones come up again and again.
That's actually good news. Once you know the patterns, you can catch most of them at review, before they cause any damage. Reviewing AI output well has become such a valuable skill that there's now a whole AI certification on the way built around it.
So let's take a look at the eight most common mistakes AI agents make when writing code, with the studies and stories to back them up.
Almost right, but not quite
If you asked developers what frustrates them most about AI tools, you'd probably guess wild hallucinations. It's actually something more subtle. In the 2025 Stack Overflow Developer Survey, 66% of developers said their biggest frustration was "AI solutions that are almost right, but not quite".
Completely wrong code is easy to deal with, because you spot it instantly and bin it. Almost right code is the dangerous stuff. It compiles, it reads sensibly, it survives a quick skim, and the bug is sitting in the one line you didn't read properly.
And here's the part that stings. METR ran a randomised trial in 2025 where experienced open source developers worked on real tasks in mature codebases, with and without AI tools. The developers using AI were 19% slower, but even after finishing, they believed the AI had sped them up by around 20%.
That difference between how it feels and what actually happened is the trap. Code that looks plausible makes careful review feel optional, and that's exactly when these mistakes slip through.
Reaching for code that doesn't exist
Ask an agent to solve a niche problem and it will sometimes import a package that has never existed. This isn't rare. A study presented at USENIX Security 2025 generated 576,000 code samples across 16 models and collected over 200,000 unique made-up package names. Around a fifth of package recommendations from open source models didn't exist at all, and even commercial models invented around 5%.
It gets worse. The same hallucinated names come back again and again (43% of them appeared in every single re-run of the same prompt), which means someone malicious can register those names for real. Security researchers call this slopsquatting.
One researcher actually tried it. He registered an empty package under a name LLMs kept inventing, huggingface-cli, and it picked up over 15,000 downloads in three months. Alibaba even added it to the install instructions of one of their open source repos.
The everyday version of this mistake is less dramatic but far more common: code written against APIs that did exist, years ago. Old framework syntax, deprecated methods, config options that were removed two major versions back. The model's knowledge of your favourite library is frozen at its training date, and it will reach for what it remembers over what's current.
Ignoring the code you already have
Every codebase past a certain size has a helper for formatting dates, a wrapper for HTTP calls and a utility for generating slugs. Your agent doesn't know that, or forgets, so it writes new ones.
One team reported finding 14 near-identical HTTP retry wrappers scattered through their codebase, each one written by an agent that had no idea the last one existed.
This shows up in the data too. GitClear analysed 211 million changed lines of code between 2020 and 2024 and found duplicated code blocks became eight times more frequent during 2024. Refactoring collapsed over the same period, with "moved" lines (code being reworked into shared places) dropping from 25% of changes in 2021 to under 10% in 2024.
In fact, 2024 was the first year on record where copy and pasted lines outnumbered refactored ones. Duplication feels harmless in the moment, but every clone is a place a future bug fix won't reach.
Only writing the happy path
Agents write the successful case beautifully and then just... stop. Missing null checks, no timeouts, and my personal favourite, the empty catch block:
try {
await syncSubscription(user);
} catch (e) {
// continue anyway
}
The error hasn't been handled here, it's just been hidden. Your app carries on as if the sync worked, and you find out weeks later when the data doesn't add up.
Part of the reason is training data. Most example code on the internet demonstrates the case where everything works, so that's what the model reproduces.
The spectacular version of this happened to Google's own Gemini CLI in July 2025. A user asked it to move some files into a new folder. The mkdir failed silently, the agent never checked, and it went on issuing move commands that overwrote the files, reporting success the whole time. When confronted, it replied "I have failed you completely and catastrophically". At least it was honest at the end.
Tests that prove nothing
Ask an agent to "write tests for this" and you'll get back a beautifully formatted suite that goes green on the first run. Feels great, right?
Here's the problem: the model doesn't know what your function is supposed to return. It only knows what it does return. So if there's a bug in the code, the tests will assert the bug as expected behaviour, and they'll pass forever.
Then there's the tautological test, where the agent mocks the very thing it's testing. Here's what it looks like:
jest.mock('./discounts', () => ({
calculateDiscount: jest.fn(() => 20),
}));
test('calculates the discount', () => {
expect(calculateDiscount(100)).toBe(20);
});
We've mocked calculateDiscount and then asserted on the mock's return value. This test checks that Jest works. The real function could return anything (or be deleted entirely) and this would stay green.
Watch out for weak assertions too: toBeDefined, not null, length greater than zero. They pass on wrong values just as happily as right ones.
Insecure by default
This is the mistake with the strongest evidence behind it, and the numbers aren't great. Veracode tested over 100 models on 80 curated coding tasks in 2025 and found that 45% of the AI-generated code failed security checks, introducing OWASP Top 10 style vulnerabilities. For Java specifically, it was 72%.
The genuinely worrying finding was that this didn't improve with newer or bigger models. They got better at making code that works, but they didn't get any better at making code that's safe.
Hardcoded secrets deserve a special mention. GitGuardian found that repositories with Copilot enabled leaked secrets 40% more often than the baseline (6.4% of repos versus 4.6%).
And a Stanford study found the cruellest twist of all. Developers using an AI assistant wrote less secure code than those without, and were more likely to believe their code was secure. The participants who wrote the least secure code trusted the AI the most.
Doing things you didn't ask for
You ask for a small fix and get back a 400 line diff. The agent fixed your bug, then reformatted the file, renamed a few variables it didn't like, "improved" an unrelated function and removed some code it decided was unused.
Most of the time this is just annoying, because now you have to review everything it touched rather than the change you wanted. Sometimes it's much worse.
In July 2025, Replit's agent famously deleted a production database during an explicit code freeze. The user had told it, in capital letters, to make no more changes. It ran a destructive command anyway, wiped records for over a thousand executives and companies, then generated around 4,000 fake user profiles to cover it up.
None of this means agents are malicious. But the "while I'm here..." behaviour that's mildly irritating when it's reformatting your files becomes catastrophic the moment an agent holds production credentials. Keeping an agent inside the scope of the task is your job, because it won't do it on its own.
Saying it worked when it didn't
"All tests pass." Does the agent actually know that? Often, no. It might not have run them, it might be reading stale output from earlier in the session, or it might have run something else entirely.
One developer mined 327 public pull requests attributed to AI agents and found maintainers complaining about cheating in 8% of them. The tricks are recognisable once you know them: assertions relaxed until they pass (toEqual becomes toBeTruthy), errors swallowed, @ts-ignore sprinkled around, and "fixes" where only the test changed and the source didn't.
To be fair to the machines, this usually isn't lying in any deliberate sense. The model narrates what it expects to be true from context, rather than checking what's actually true. But the effect on you is the same: you can't take "done and verified" at face value.
So don't. Run the tests yourself, read the diff, and treat every claim of success as exactly that, a claim.
So, what can you do?
None of these mistakes are new, really. Duplicated code, weak tests and missing error handling are the same things you'd flag in any code review. The difference is that an agent produces them much faster than a person ever could, and makes them look far more polished.
So the honest advice is a bit boring: review an agent's work like you'd review a pull request from someone you've never worked with. Read the whole diff, check the packages it imports actually exist, run the tests yourself, and be suspicious of any green tick you didn't watch happen.
Spotting these issues is quickly becoming a skill in its own right. It's exactly the kind of thing that the certificates.dev AI certification I mentioned at the start will test you on, which tells you how real the problem is.
Either way, keep these eight in mind the next time an agent hands you something that looks perfect. Chances are at least one of them is hiding in there somewhere.


