Your tests are green. Here are three ways the code is still wrong
Agents write my code and an AI reviewer reads every push. In 13 days it blocked 32 of 315 changes, most of them bugs green tests could not see.
· Jeremy
A test checks the path its author imagined. A reviewer that reads every change goes looking for the others.
On the afternoon of 6 August a change to one part of my product went through every test. All 1,041 of them, across three suites, came back green, so I pushed it. The next push (the next attempt to send code online) was blocked for a bug none of them could see.
That push touched how my product chooses which AI model answers a request. A catch-all setting (one meant for wherever nothing is specified) would have switched the deep-think model on for every free user. An AI reviewer that reads every change before it leaves my machine caught it. Free users would have got the expensive deep-think model for nothing, and the bill would have landed on me.
Why did no test see it? The deep-think model is deliberately left out of the free plan. But the code checked the catch-all setting before it reached that rule. Nobody had set it anywhere the code runs, so every test ran without it. The note saved with the fix calls it the kind of setting somebody sets on the evening of an outage.
The free-plan bug was the first of three ways I have watched green tests lie: a setting nobody has set, a hand-off between two parts of the code that changes shape, and a database clause that can never fire. In each case the bug lived on a path nobody tested.
Hacker News is the forum where programmers trade war stories. There, after an AI reviewer found half a dozen bugs, a commenter called hinkley put it more bluntly: "in part because the tests were written by an optimist, and I mean that as a dig."
My reviewer, which my AI coding agents and I built, runs on my machine on every git push. Alongside the tests and a type check that scans for mismatched data, a model reads the changed lines and ranks each finding high, medium or low. Only high stops the push.
I no longer write my product's code myself, since agents do, so the review is the first thing that reads it. It exists, in my words, "to avoid sending shit." I call it the gate.
Is a review on every push worth it? The gate blocked 32 of its 315 reviews, each on at least one high finding. I read the 35 high findings behind those blocks, and where a follow-up change existed, I read that too. By my reading, about 25 of them were logic bugs my tests would not have caught. So mostly yes, with the limits below.

Each row shows what the author pictured on the left and what the code did on the right.
What I would do, and what each item rests on. Items 1 and 2 come from the gate's written rules and its log:
- Block only on a bug the change itself introduces, stated with a one-sentence failure scenario.
- Judge a block by its reason, not its verdict: one unchanged push got "block", then "pass" minutes later.
My reasoning, which I did not measure:
- Run a reviewer on every push where nobody else reads the changes: a check that runs on
git push, sends the changed lines to a model, and stops the push on a high finding. If CI (an automatic test run on a shared server) also checks the code, treat the gate as a second check, not the only one. - Keep a list of what it missed. I have no such list.
- Read the medium and low findings even though they never stop a push. I have not sorted mine.
When a hand-off changes shape
The second way came on the evening of 6 August, after the free-plan block. A customer of my product can pin a model: lock in the one they choose. Each model comes from a provider (the company that supplies it), and the customer pays that provider through their own account.
One function in the code that chooses the model returned a model's name with its provider attached, and the next function read that to find out which provider to call. A refactor, a rewrite that should change nothing visible, made the first function return the name alone. After it, a pinned model went to the default provider instead. The gate blocked the push that carried it, and nothing would have shown an error anywhere.
The customer's usage would have landed on the wrong account, and the first sign would have been the invoice.
When a database clause can never fire
The third way needs no customer, only a batch of data and a clause that looks like protection. On 27 September, a batch write stored links (connections between two of a user's records). It gave each link a fresh random id and relied on a database clause, "on conflict do nothing," which tells the database to skip a row whose id already exists. A fresh random id never matches an existing one, so the clause could never fire, and a link written twice would be stored twice. The tests were green.
The gate rated it high, and its finding says it reproduced the duplicates against a real database. The fix drew two more highs, for duplicates inside a single batch, and the third version passed. Without the gate, every repeated link would have sat twice in a user's records, and every test would still be green.
Isn't it just a tool that always finds something?
To a sceptic, a reviewer that needs three versions before it passes looks like a tool told to find problems, and such a tool always finds some. It does write plenty of notes, but most are ranked medium or low and never stop a push. It is told to block on one level only: a bug the change introduces, with a way it fails. Most pushes went through without a block, as the chart below shows. Most blocks were followed within 20 minutes by a change to the flagged file and a review that passed.

The top bar splits the 315 reviews by verdict, and the two bars below count the same 32 blocks by two different questions.
Not every block was a logic bug. Some highs repeated what a red test or the type check already reported, and one came from a test I ran on the gate on purpose. A few blocks look like noise to me. Some fell on other people's code pulled into a branch (a separate line of work). On those pushes I had only edited documentation. One more block flipped to pass with no change.
Shouldn't a person be doing this review?
The objection to the whole idea comes from another Hacker News commenter, trjordan, who wrote that "if you've gotten to the point where you're relying on AI code review to catch bugs, you've lost the plot." In the commenter's view, a pull request (the proposal teammates read before a change is merged) exists "to share knowledge and to catch structural gaps." For a team, that is right. Here no person wrote the change, so no author already understands it. On projects without CI, the gate is the only automatic review. I still read the pull request before I merge.
The bill and the blind spot
The reviews have a price: a few minutes of waiting on each push, and a little money. The reviews I could trace, about half of them, cost 55 cents in total, mostly because the gate's model is usually a free one. That leaves out the fallback (the model used when the free one fails), a Claude model from Anthropic, paid by subscription. So the true bill is higher, by an amount I have not measured.

The top bar shows how little of the gate's output I have sorted, and the bottom bar shows my reading of the part I did.
I have no list of bugs the gate passed that later shipped, so I cannot say what it misses. The counts above are mine, from one project, over 13 days, made by the person who built the gate. The findings ranked medium or low were never sorted. The August scenes predate the gate's log, so they rest on the notes saved with each change. You will have the same blind spot until you keep a list of what it missed.
What a green test proves
Green tests prove the path you pictured, not the others. Let a reviewer that did not write the tests look for the rest. The fix for the misrouted pinned model was saved ten minutes after the block. The note saved with it, translated from French, says: "Found by the gate's reviewer, which refused the previous push, rightly."