Code Review Checklist for AI-Generated Code: 12 Things to Verify
96% of developers don't fully trust AI-generated code, yet only 48% always check it before committing. Here are the 12 things to verify in every AI-written PR, and which ones a tool can take off your plate.
Tired of slow code reviews? AI catches issues in seconds. You decide what gets published.
Code Review Checklist for AI-Generated Code: 12 Things to Verify
TL;DR: AI does not write bad code — it writes plausible code. The lines compile, the linter passes, the variable names look reasonable, and on a fast PR review it slides through. That is the exact reason bugs ship. This checklist covers 12 specific things that AI gets wrong differently than humans do, with a code example for each, a comparison table showing which items Git AutoReview catches automatically, and a copy-paste version for your team's PR template.
The real problem with AI-generated code is not quality. It is that the failure mode looks like the success mode. A human author who hands you a confused PR usually leaves obvious signs: inconsistent naming, abandoned comments, a TODO that says "fix this." AI strips those signals out by default, which means your standard review pass approves things it would normally catch. What reaches your codebase is a slow leak of subtle bugs.
The verification gap
Sonar's 2026 State of Code survey put a number on something most teams already feel. Across more than 1,100 developers, 96% said they do not fully trust that AI-generated code is functionally correct, while only 48% said they always check it before committing. Those two figures sit badly next to each other. Roughly half the industry is shipping code it openly does not trust.
The volume behind that gap keeps climbing. The same survey found AI already writes 42% of committed code, with developers expecting 65% by 2027, and 72% of people who have tried these tools now reach for them every day. Stack Overflow's 2025 survey, published in February 2026, shows the two curves moving in opposite directions: usage rose to 84% while trust dropped from 40% to 29% in a single year. Developers are not growing more confident with practice. They are growing less confident and using the tools more anyway.
Review capacity is what absorbs the difference, and it is buckling. LinearB's 2026 benchmark, built on 8.1 million pull requests from 4,800 teams across 42 countries, found AI-generated PRs wait more than 16 hours on average before a reviewer even picks them up, against roughly 200 minutes for human-authored ones. Of the AI PRs that do get reviewed, 32.7% are merged. For human PRs the figure is about 84.5%. They arrive bigger, too, at over 400 lines against 157 for unassisted work. A queue where the biggest changes wait longest and then get rejected two times out of three is a review process running past its limit. Our PR review time benchmark breaks down the rest of that cycle-time picture.
So the useful question here is not whether AI writes bad code, but which specific things go wrong often enough to earn a permanent place in your review pass.
Why AI-generated code fails differently than human code
GitClear's Maintainability Gap report, published in January 2026 off 623 million analyzed changes from 2023 onward, traces where the quality actually goes. Copy-pasted lines rose from 9.4% of changed code in 2022 to 15.7% in the first half of 2026. Over the same stretch, moved code — the signature of refactoring, of someone tidying up rather than adding — collapsed from 21% to 3.8%. Duplicated blocks went from 40.3 to 73.0, an 81% jump. Put plainly, the industry stopped reorganizing code and started copying it, and reviewers are the only checkpoint left that notices.
CodeRabbit's State of AI vs Human Code Generation report, out in December 2025, is the closest thing to a controlled comparison anyone has published. Across 470 open-source pull requests (320 AI-co-authored, 150 human-only), AI changes averaged 10.83 issues per PR against 6.45 for human ones, roughly 1.7x. The category breakdown is the part worth memorizing, because it tells you where to point your attention: error-handling gaps and naming inconsistencies each came in near 2x the human rate, logic and correctness issues 75% higher, readability more than 3x worse, security findings up to 2.74x, and excessive I/O operations about 8x more common. Their own conclusion is that AI makes the familiar categories of mistake far more often, rather than inventing new ones.
All of which puts the reviewer in an odd position. Only 48% of developers always check AI output before committing, so on a good share of pull requests the reviewer, rather than the author, ends up being the first person to read that code with full attention. The checklist below is written for that position, not the one we would prefer to be in.
The core failure is alignment. AI optimizes for syntactic correctness and surface plausibility — code that parses, passes type checks, and matches the pattern of training-data examples. It does not optimize for semantic correctness, architectural fit, or whether the result actually solves the problem the ticket described. Those last three are exactly what code review exists to verify, and they are the parts a human reviewer has to do manually because the AI cannot self-check them.
The 12-item checklist
1. Requirement alignment — does the code actually do what the ticket asks?
AI reads the ticket literally and fills in the gaps with its own assumptions. "Add export button" turns into CSV export of the current page only — not Excel, not all users, not the format your finance team actually needs. The gap between ticket intent and code behavior is where most "this works but it's wrong" PRs come from.
// Ticket: "Export user data"
// AI generated: exports only current page
function exportUsers() {
return currentPageUsers.map(u => u.toCSV());
// Missing: pagination, all users, format preference
}
What to check: read the ticket once, read the code, ask whether a non-technical stakeholder would call this "done."
2. Hallucinated package names (slopsquatting)
Slopsquatting exists because the list of names to squat on is enormous. USENIX Security's 2025 paper, We Have a Package for You!, measured exactly how enormous: 576,000 generated code samples run through 16 models, with commercial LLMs inventing a package name at least 5.2% of the time and open-source models doing it at 21.7%. Across those samples, the researchers cataloged 205,474 distinct invented names, and not one of them exists on npm or PyPI — which makes every entry a free package name waiting for an attacker to grab it. It works because developers paste AI-suggested code into package.json without checking whether the package it names is real.
{
"dependencies": {
"react-data-fetcher": "^2.1.0",
"mongoose-validator-utils": "^1.0.0"
}
}
What to check: for every new dependency, run npm info <package-name> (or the equivalent for your registry), look at the download count, the GitHub repo, and the publish date. A package that appeared three weeks ago with no downloads is a red flag regardless of how reasonable the name sounds.
3. Cross-file side effects (what changed outside the diff?)
AI sees the diff. It does not see the rest of your codebase. A rename that looks clean in the PR can break 14 import paths in files that are not part of the change, and TypeScript or your linter might not catch it if the imports are dynamic or wrapped in conditional logic. For TypeScript-specific patterns AI catches and misses, see our TypeScript code review breakdown.
// AI renamed: formatDate → formatDateTime
// Clean in this file, but 8 other files still import formatDate
export function formatDateTime(date: Date): string { ... }
What to check: grep the old name across the codebase before approving. For function renames, search for both the import and the call site. For config changes, look at every build script that reads the file.
4. Hardcoded credentials and secrets
Claude Code-assisted commits leak secrets at 3.2% against a 1.5% baseline, according to GitGuardian's State of Secrets Sprawl 2026 report — published in March and built on the 28.65 million new hardcoded secrets the company found during 2025. Read that comparison precisely, because the framing matters: the baseline covers every public GitHub commit, not a clean human-only control, and GitGuardian is explicit that developers still decide what gets accepted and pushed. Even with that caveat, the gap is roughly double. The mechanism is dull and consistent. AI fills in placeholder values it saw during training, and some of those placeholders were real keys somebody pushed to a public repo years ago.
OPENAI_API_KEY = "sk-proj-abc123..."
DATABASE_URL = "postgresql://user:password123@prod-db:5432/myapp"
What to check: any string literal that looks like a key, token, or credential. Tools that scan for entropy patterns catch most of them, but the review pass should still flag anything that is not loaded from environment variables.
5. Error handling completeness
CodeRabbit's PR study put error handling gaps at 2x the human baseline. AI tends to write the happy path cleanly and skip the failure modes — network errors, null returns, partial responses, timeouts. The function "works" in development where nothing fails, then breaks in production where everything does.
async function fetchUser(id) {
const response = await fetch(`/api/users/${id}`);
const data = await response.json();
return data.user;
}
What to check: every await, every external call, every place data crosses a trust boundary. If there is no try/catch and no response.ok check, the code is incomplete.
6. Logic correctness (not just syntax)
The most dangerous AI bugs are the ones where every line is grammatically correct and the overall logic is wrong. Off-by-one errors in date ranges, inverted conditionals, wrong operator precedence — the linter cannot help here, the type checker cannot help here, and a fast review will miss it because the code looks right.
def is_eligible(user):
return not user.is_premium and user.subscription_active
# Should be: user.is_premium and user.subscription_active
What to check: read the code out loud. If the function name says "is eligible" and the code returns true for non-premium users, something is off. Trace one real input through the function by hand.
7. Naming and consistency
CodeRabbit measured naming inconsistencies at 2x the human rate. AI generates code in its own naming style and ignores the conventions of the surrounding file. You end up with snake_case, camelCase, and PascalCase instances of the same kind of thing in the same module.
const user_service = new UserService();
const UserRepo = new UserRepository();
const getuser = async () => {};
What to check: scan for naming style mismatches before merging. If the project uses camelCase for variables, every new variable should be camelCase. Configure a linter rule if you can — most teams cannot maintain this manually.
8. Dead and unreachable code
AI often generates belt-and-suspenders code that the type system or runtime would never execute. Redundant null checks after an early return, fallback branches that cannot be reached, unused variables that compile but add cognitive load to every future reader.
function processPayment(amount: number): Result {
if (amount <= 0) throw new Error('Invalid amount');
if (amount <= 0) return { error: 'Invalid' }; // Dead
// ...
}
What to check: every conditional branch — could the runtime ever reach it? Every variable declaration — does it get used? A reviewer with five minutes can catch most of this by reading top to bottom and asking "why is this here?"
9. Test coverage for new paths
AI rarely writes tests for the edge cases that matter. The happy path gets a test. Error paths often do not. The new code might pass existing tests by changing what those tests actually verify — a subtle form of test rot that takes weeks to surface.
What to check: every new function should have at least one test. Error paths need tests too, not just success paths. If existing tests pass after a refactor, look at whether they pass for the right reasons or whether the refactor accidentally weakened the assertions.
10. Console.log and debug artifacts
AI leaves debugging artifacts everywhere. console.log statements with PII in the output, commented-out blocks with TODO markers that were never meant to ship, debugger keywords that crash production builds. These are individually small and collectively a noise problem in production logs.
async function processOrder(order) {
console.log('Processing order:', order);
console.log('User:', JSON.stringify(order.user)); // PII in logs
// TODO: add validation here
return await db.orders.create(order);
}
What to check: grep for console.log, print(, debugger, and TODO in the diff. None of them belong in production code unless your team has explicit policy for it.
11. Security vulnerabilities (OWASP Top 10)
Veracode ran more than 150 models through their 2026 GenAI Code Security Report, and only 55% of AI-generated code cleared basic security tests. The number that should change how you review is their observation that this rate has barely moved since 2023, through several model generations, while syntax correctness climbed above 95%. Newer models write code that compiles more reliably and secures itself no better.
The failure is also lopsided by category rather than evenly spread, which is what makes it reviewable. SQL injection passes 82% of the time and insecure cryptography 86%. Cross-site scripting passes 15%. Log injection passes 13%. If you have limited attention for a given PR, spend it on output encoding and on anything written into logs.
def get_user(username):
query = f"SELECT * FROM users WHERE name = '{username}'"
return db.execute(query)
What to check: parameterized queries, sanitized HTML output, authorization checks on every new endpoint, validation on every input that crosses a trust boundary. Tools help here, but the reviewer still has to verify that the tools are configured to scan the new code.
12. Architectural fit
AI writes code that works in isolation. It does not know that your team made a deliberate choice to route all DB access through a service layer, that this React component is in the presentation tier and should not call Supabase directly, or that the middleware pattern you established three months ago exists for a reason.
// In a React component — AI added a direct Supabase call
const { data } = await supabase.from('users').select('*');
// Should go through: userService.getUsers()
What to check: does the new code follow the project's existing patterns for its layer? If the codebase has a service layer, repository layer, or controller pattern, the new code should respect those boundaries. This is the item where Deep Review (full codebase exploration) catches things diff-only review misses.
Which items can be automated?
Some of these 12 items need a human reading the ticket and the code together. Others are pattern matches that an automated tool catches faster and more consistently than a tired reviewer at 4 PM on a Friday. Here is the honest split:
| Item | Manual | Git AutoReview | How |
|---|---|---|---|
| 1. Requirement alignment | Manual | Jira Integration | Reads ticket, flags code-ticket gaps |
| 2. Hallucinated packages | Manual | Manual | Check npm/PyPI directly |
| 3. Cross-file side effects | Partial | Deep Review | Explores full codebase, follows imports |
| 4. Hardcoded secrets | Manual | 20+ security rules | Catches API keys, passwords, tokens |
| 5. Error handling | Manual | Deep Review | Traces async paths across files |
| 6. Logic errors | Manual | Quick Review | Flags inverted conditions, off-by-one |
| 7. Naming consistency | Manual | Quick Review | Compares to codebase conventions |
| 8. Dead code | Manual | 20+ rules | Flags unreachable code, unused vars |
| 9. Test coverage | Manual | Deep Review | Checks coverage for new code paths |
| 10. Debug artifacts | Manual | 20+ rules | console.log, debugger, TODO detection |
| 11. Security (OWASP) | Manual | 20+ rules | SQL injection, XSS, validation |
| 12. Architectural fit | Manual | Deep Review | Checks layer boundaries, patterns |
Git AutoReview's 20+ built-in rules catch items 4, 7, 8, 10, and 11 on every PR. Deep Review adds items 3, 5, 9, and 12 through full codebase exploration. Jira Integration covers item 1. The rest stay with the human reviewer where they belong.
Install Free Extension →
The copyable checklist
Drop this into your PR template, your Notion doc, or wherever your team keeps review process documentation. The wording is intentionally short so reviewers can scan it without losing focus on the code.
## AI-Generated Code Review Checklist
- [ ] 1. Requirement alignment — code does what the ticket actually asked
- [ ] 2. No hallucinated packages — every new import verified on registry
- [ ] 3. Cross-file side effects — renames/refactors don't break files outside the diff
- [ ] 4. No hardcoded secrets — keys/tokens/passwords loaded from env
- [ ] 5. Error handling complete — try/catch on async, response.ok checks, null guards
- [ ] 6. Logic correct — traced one real input through every new function
- [ ] 7. Naming consistent — matches surrounding file's conventions
- [ ] 8. No dead code — every branch reachable, every variable used
- [ ] 9. Tests for new paths — error cases tested, not just happy path
- [ ] 10. No debug artifacts — console.log, debugger, TODO stripped
- [ ] 11. Security clean — parameterized queries, sanitized HTML, validated inputs
- [ ] 12. Architectural fit — respects existing layer boundaries and patterns
On GitHub, save this as .github/PULL_REQUEST_TEMPLATE.md and it auto-fills every new PR. For more PR template patterns, including GitLab and Bitbucket setups, see our full guide.
Where this connects to the rest of your review process
The 12 items above sit on top of standard code review practice, not in place of it. The GitHub code review best practices guide covers the broader workflow — PR size targets, review SLAs, the metrics that matter. The VS Code PR review guide walks through three ways to run reviews inside the editor, including the AI-assisted approach where most of this checklist gets automated. For teams running GitHub Copilot Code Review, the June 2026 pricing change adds a billing wrinkle worth reading about before your invoice arrives. If you are running this checklist as a solo developer with no peer reviewer in the loop, our solo developer code review playbook covers the cognitive reason self-review fails and the four methods that actually work — this 12-item list pairs with method 2 in that piece.
If your team is already drowning in PRs because AI cranked up the write side without changing the review side — that is the exact problem Git AutoReview was built to solve. Free plan covers 10 reviews per day with no credit card. Team plan handles unlimited reviews for $14.99 flat across the entire team.*
* Git AutoReview subscription price only. AI compute costs of approximately $2–5/month per developer are billed directly by your AI provider (Anthropic, Google, or OpenAI). CodeRabbit and Qodo bundle AI compute into their per-user price.
Tired of slow code reviews? AI catches issues in seconds. You decide what gets published.
Frequently Asked Questions
How is reviewing AI-generated code different from reviewing human code?
What is slopsquatting in AI code review?
How often does AI-generated code have security vulnerabilities?
What percentage of developers use AI coding tools?
Can AI tools review AI-generated code?
How do I add this checklist to our PR template?
Try it on your next PR
AI reviews your code for bugs, security issues, and logic errors. You approve what gets published.
Free: 10 AI reviews/day, 1 repo. No credit card.
Related Articles
OpenCode vs Claude Code: An Honest Comparison (2026)
OpenCode vs Claude Code compared on licence, models, cost and PR review. Verified GitHub and pricing data, plus the archived repo everyone still quotes.
Open VSX: AI Code Review in Cursor, VSCodium & Windsurf (2026)
VS Code forks cannot reach Microsoft's Marketplace. Here is the licence reason, and how to install AI code review in Cursor, VSCodium and Windsurf.
Jira to Pull Request: Closing the Loop Between Tickets and Code Review 2026
Most teams mark Jira tickets Done before the PR gets a real review. Here's how to wire Jira to GitHub, GitLab, and Bitbucket so ticket context drives code review — and nothing ships unverified.
Get the AI Code Review Checklist
25 PR bugs AI catches that humans miss — with real code examples. Free PDF, sent instantly.
One-click unsubscribe. We never share your email.