How to verify AI-generated code
Updated
A confident explanation is a reason to inspect a change, not evidence that it works. Verification starts by stating the behavior you expect in a form that could fail. The useful question is which observation would show the proposed fix is wrong.
This guide describes a review method you can use with an AI agent or by yourself. It does not require a particular model. Keep the task's existing tests and conventions as your starting point, then add the smallest check needed to cover the missing behavior.
Write the expectation before the fix
Read the requirement and the affected call path. Specify inputs, outputs, and any behavior that must stay intact. Ask whether the proposed check would fail on the original implementation. A test written to match the new code can pass while preserving the original bug.
For example, suppose a helper should remove duplicate integers while preserving their first appearance. Sorting a set would remove duplicates but change their order. Checking only that the result contains distinct numbers would miss that mistake. The order is part of the contract, so it needs an assertion.
Example: check order, empty input, and zero
This is a small synthetic example, not a solution to a benchmark task. Save it as test_unique.py and run python -m pytest test_unique.py. The first assertion checks order, the second checks empty input, and the third protects zero from accidental filtering. Pytest reports a failed assertion with the observed values.
Replace the helper's return value with sorted(set(values)) and the order check fails. That failure demonstrates that the check can distinguish one plausible wrong implementation. It doesn't establish correctness for a larger program or for types outside the stated integer contract.
# Synthetic example: integers, first occurrence wins.
def unique_in_order(values):
return list(dict.fromkeys(values))
def test_unique_in_order():
assert unique_in_order([2, 1, 2]) == [2, 1]
assert unique_in_order([]) == []
assert unique_in_order([0, 0, 1]) == [0, 1]Expand from the failing case
For a repository fix, run the targeted check first, then tests for callers or neighboring behavior affected by the edit. Use the project's runner and the relevant test file. The example's pytest command belongs to that example. Other projects may need their own runner, configuration, or environment.
Read the output closely enough to establish that tests actually ran. Collection errors, skipped cases, and missing dependencies are different from passing assertions. Keep the exact command and result. If you can't run a relevant check, say which behavior remains unverified rather than replacing the result with the agent's prediction.
Review the patch independently
Inspect working-tree changes with git diff and staged changes with git diff --cached. Check untracked files separately. Confirm that the change fixes the source behavior and hasn't weakened tests, changed unrelated configuration, or added a dependency without a task-specific need. A clean diff is easier to reason about than a broad rewrite.
In a live practice round, the graded-test action checks the task's recorded required and existing tests. Submission produces a separate grading result. Read both outcomes and any environment notes. Passing the evaluated tests supports those observed behaviors. It doesn't prove all inputs are correct, that the code is secure, or that every test in the repository passed.