Stickybit.← NotebookPortuguêsEssay · tests and AI agents · Sep 30, 2026 ·
Essay

Green is silence, not a certificate.

When the test pipeline turns green, it is saying one thing only: no test failed. That is worth a lot if the tests could have failed, and almost nothing if they could not. On the most used leaderboard for AI coding agents, hundreds of accepted fixes were wrong, and the green did not see it.

Specimen · a shipping function, its tests and eight planted defects

The rule: orders of R$ 199 or more ship free. Below that, R$ 15 plus R$ 2 per kilo, with the weight rounded up.


        
        
Planted defects: does the suite notice?
–the pipeline
–planted defects caught
–the real defect

The code really runs in your browser: each planted defect is a copy of the function with a single change, and each test is executed against it. A defect only counts as caught if the suite passes on the original code and fails on the copy.

Before the argument

Four words for this page.

Green pipeline

The dashboard that runs the tests on every change and turns green when none failed. It is the signal that lets the change ship.

A test that can fail

A test that would turn red if the code were wrong in the way that matters. Only that kind turns green into information.

Planted defect

A copy of the code with a deliberate mistake, like a flipped sign. If the tests do not notice it, they would not notice the real one either.

Yardstick

Whatever decides whether the work passed. When whoever is graded can touch the yardstick, the result measures the yardstick, not the work.

What green says

A detector with no battery is silent too.

A smoke detector that stays quiet can mean two things: there is no fire, or the detector does not work. From outside, the silence is the same. The only way to know which is to test the detector: put smoke near it and see whether it beeps.

Software tests are the same. A failing test says a lot: something is wrong, and it points to where. A passing test says almost nothing, unless it would have failed if the code were wrong. Green is a sum of silences, and each silence is worth what its test was worth.

The common mistake is not having too few tests. It is reading silence as a certificate: the report says "passed all tests", and the reader hears "it is correct". Those are different sentences, and the gap between them is where defects live.

red there is a defect,almost certainly,and the test points to it green there is no defect,orthe test could notsee this defect THE QUESTION GREEN DOES NOT ANSWER could this test have failed?
Red and green are not symmetric. Red already carries the information; green only does if someone has checked that the test can fail.
The agents' leaderboard

Passed by the tests, wrong in fact.

SWE-bench is the most cited leaderboard for comparing AI coding agents. Each task is a real issue filed in a public project. The agent writes the fix, and it counts as resolved if it passes the tests chosen for that task. The whole leaderboard is made of greens.

In 2025, two groups went to check those greens. One ran the developers' own full test suite on the accepted fixes: 7.8% passed the task's tests and failed the full suite. And 29.6% of the accepted fixes behaved differently from the fix written by the project's people. Altogether, the resolution rate was inflated by 6.2 percentage points.

The other group added tests where they were missing and found 176 wrong fixes accepted in one version of the leaderboard and 169 in another. With the yardstick fixed, 40.9% of the entries changed position in one of the tables. None of those fixes had lied: the tests simply could not fail for them.

SWE-BENCH, RECHECKED IN 2025 7.8%of accepted fixes fail thedevelopers' full suite 29.6%behave differently fromthe project's own fix 176 + 169wrong fixes accepted,in two leaderboard versions 40.9%of positions changed whenthe tests were strengthened
Numbers from two independent studies (Wang, Pradel and Liu; and UTBoost, published at ACL 2025). None of them is our measurement.

There is an earlier problem, filed in SWE-bench's own repository: in some setups, the agent can edit the test files that will later judge it. When whoever is graded can touch the yardstick, green measures the yardstick.

Why tests go slack

A strict test charges now. A slack one charges someone else, later.

This part is our reading, not a measurement. A test that is too strict fails on code that worked: it blocks the release, someone loses an afternoon finding out why, and the test gets the blame. A slack test lets a defect through: the release ships on time, and the bill arrives weeks later, in production, for another team or for the customer.

The cost of strictness is immediate and has an owner; the cost of slackness is delayed and diffuse. Without anyone deciding anything, day-to-day pressure pushes suites toward slack: the annoying test gets loosened or deleted, the one that never fails stays. Over time, green gets more frequent and worth less.

The fix is to make slackness expensive for whoever chooses it: measure, every so often, whether the tests can still fail. That is what the specimen does.

WHEN THE BILL ARRIVES, AND FOR WHOM Overly strict test fails today blocks the release; the team pays,right away, with a name on it Slack test passes today the release ships on time defect in production another team paysor the customer TODAYWEEKS LATER day-to-day pressure pushes the suite toward the bottom row
Sketch, our reading: neither cost was measured here. What changes between the rows is not the size of the bill, but when it arrives and who gets it.
Planting the defect

Putting smoke near the detector.

The way to test the tests has existed for decades: plant defects on purpose. Make a copy of the code with one small mistake, like flipping a sign or changing a number, and run the suite. If it turns red, the defect was caught. If it stays green, the suite cannot see that kind of mistake. The share of defects caught is the suite's grade, not the code's.

In the specimen, the three tests checked by default are what a common suite tends to have: runs without error, one typical case below the threshold, one above. It stays green and catches five of the eight planted defects. The real defect gets through: the code charges shipping on an order of exactly R$ 199, and none of those tests looks at the threshold.

We tried to take this to real changes and ran into reach. In a partial sample, from a single project, off-the-shelf defect-planting tools almost never managed to plant anything inside the lines the change touched. The idea is good; the shelf tools do not reach where the typical change operates.

EIGHT COPIES, ONE MISTAKE EACH the suite turned red:mistake caught stayed green:mistake invisible 5 of 8 is the suite's grade, not the code's.Its green holds for those five kindsof mistake, and only for them.
The specimen's result with the default suite. Here the color is the suite running on the defective copy: red is good.
The measured case

A test written from the code protects the defect.

It is now common to ask an AI model to write the tests for a change. We measured that in August, on two public Go projects. We picked 19 changes we knew contained a defect, because it was fixed later, and 10 clean changes of the same size. For each one, an AI agent, without knowing which group it was in and without seeing the fix, wrote tests from the code and the change description.

The generated tests caught none of the 19 defects. The reason is the same as the snapshot of the current code in the specimen: whoever writes the test by looking at the code records what the code does, mistake included. The test is born agreeing with the defect.

Two details matter. In the only two cases where the change description said one thing and the code visibly did another, the test failed and pointed to a real defect. And on the 10 clean changes the method barely invented alarms: it fired once, and rightly. It is good for acquitting, not for accusing. It is the same split we found when one AI checks another, in an earlier entry of this notebook.

GroupChangesGenerated test turned redReading
With a known defect190Caught none. The test inherited what the code does.
Clean, same size101Barely invents alarms, and the one it raised was real.
Among the 29, description and code visibly disagreed22The only place the generated test bites.

Our measurement, August 2026, with the losing criterion recorded beforehand: the bet that generated tests would recover defects died. The sample is small, and some tests did not compile against the fixed version, which is a limit of how we checked, not of the method in real use.

What to do

Four rules to make green mean something again.

  1. Ask whether the test can fail

    Every so often, plant defects and see which get through. The suite's grade is worth more than the test count or the percentage of lines covered.

  2. Write the test from the rule, not the code

    A test born from the specification disagrees with the code when the code is wrong. One born from the code agrees with it, mistakes included.

  3. Whoever is graded does not touch the yardstick

    Neither a person nor an AI agent changes the tests that will judge their own change without someone seeing it. A test file changed together with the code is a warning in itself.

  4. Call green by its name

    In the report, "no alarm in tests X, Y and Z", not "approved". The difference in words is the difference between what was measured and what one would like to be true.

What we do not know yet

Where this could be wrong.

The leaderboard numbers are someone else's

The SWE-bench percentages come from two 2025 studies, each with its own method. We checked the published summaries; we did not redo the math.

Our sample is small

29 changes, two projects, one language. And what we call a "known defect" depends on linking each fix to the change that caused it, which sometimes goes wrong.

Planting defects has a cost and a blind spot

Running the suite for each planted defect is expensive, and some planted defects do not change behavior, so no test could catch them. The suite's grade is an estimate, not a verdict.

The specimen is a toy

A three-line function makes the mechanism visible. In real code, the defects that escape tend to live where there is not even a simple way to test.

Talk to us

Does this apply to your case?

Tell us in two lines what you need to decide or measure. The first conversation is to see whether measurement solves your case, and if it does not, we say so.

Talk on WhatsApp algorithms@stickybit.com.br Stickybit · Porto Alegre, Brazil, since 2004

← Research notebook · The cheap one acquits · stickybit.com.br

Sources