← back to the benchmark

The bugs AI coding assistants actually write

AI-written code rarely fails in a way that looks broken. It compiles, the names sound right, and the happy path works. The bugs sit in the one line that reaches for something that is not there, or in a comparison that points the wrong way. These are the seven families of mistake the AI Bug Hunt is built from, each with a small example written for this page and the review habit that catches it.

  1. 01Invented functions, fields and arguments

    The most AI-specific bug. The model calls a helper that sounds right but was never imported or defined, reads a field the type never declared, or passes an argument the function does not take.

    EXAMPLEA file imports get, set and has from a cache module, then calls getOrThrow. It reads like a stricter variant that should exist. It does not.

    HOW TO CATCH ITFor every call in a diff, find where the name comes from: the import block, the definition, the type. If you cannot point to it in ten seconds, treat it as suspect.

  2. 02Logic that points the wrong way

    Flipped comparisons, off-by-one limits, inverted loop exits, or the wrong one of two values returned.

    EXAMPLEA cache treats expiresAt < now as still fresh, so it serves stale entries and evicts live ones.

    HOW TO CATCH ITRead every comparison aloud with a concrete value: if it expired one second ago, does this return it?

  3. 03Security that looks handled

    The safe version is right there, used on one line and skipped on the next.

    EXAMPLETable headers go through an escaping function, while body cells are interpolated raw into HTML.

    HOW TO CATCH ITWhen a safety helper appears once, check every sibling line that handles the same kind of input. Also watch for decode where verify was meant, a plain random generator where a cryptographic one was meant, and endsWith used to check a hostname.

  4. 04SQL that drops rows without an error

    The query runs and returns a wrong number.

    EXAMPLEA LEFT JOIN whose right-hand table is then filtered in WHERE, which quietly turns it into an inner join and loses every row with no match.

    HOW TO CATCH ITFor each join, ask which rows can be NULL on each side, and whether a filter, a NOT IN or an AVG changes what happens to them. Integer division and averaging at the wrong grain belong here too.

  5. 05Concurrency that only works on one thread

    Code that is correct when one thing happens at a time.

    EXAMPLEAn async callback passed to forEach, which never waits for it, so the function returns before any of the work finishes.

    HOW TO CATCH ITFind the shared state and the awaits, and check that the locked or awaited section really covers the whole read, change and write.

  6. 06Resources that leak on the error path

    Everything is closed when things go well.

    EXAMPLEAn HTTP response body closed on success and left open when the status is not 200.

    HOW TO CATCH ITTrace each early return and error branch, and check that everything opened above it is closed.

  7. 07Failures that look like success

    An error happens and nothing says so.

    EXAMPLEA line scanner that stops on a read error and never checks for one, so a truncated file comes back as if it were complete.

    HOW TO CATCH ITLook for every error value the code receives, and make sure each one is checked.

Clean code is part of the test

Half the skill is leaving correct code alone. Unfamiliar but valid idioms get flagged as bugs all the time, and in a real review that costs as much as a missed bug. The AI Bug Hunt includes clean snippets for that reason, and a false alarm counts the same as a miss.

Why this matters in interviews now

Several employers now run coding rounds where an AI assistant is allowed or expected, and grade whether you check what it wrote. Calibrd's guides cover Cerebras, Shopify and Meta's pilot round in the Meta London guide.

Questions

What kind of bugs does AI-generated code have?
Seven families come up again and again: invented functions, fields or arguments; logic that points the wrong way; security that is handled on one line and skipped on the next; SQL that drops rows without an error; concurrency that only works on one thread; resources that leak on the error path; and failures that look like success. The most AI-specific is the first: a call to something that sounds right and does not exist.
How do you review AI-written code?
Check where every name comes from, read each comparison with a concrete value, look at the sibling lines whenever a safety helper appears once, follow the error paths, and ask which rows a join can lose. Just as important, do not block code only because it looks unfamiliar: a correct idiom flagged as a bug costs a review as much as a missed bug.
How can I practise catching bugs in AI-written code?
The AI Bug Hunt on this site shows four short snippets, some with a real bug and some clean, and gives you about two minutes in total. There is no account. It scores both missed bugs and false alarms.

Try it: four snippets, some clean, about two minutes, no account. Start the AI Bug Hunt →