Four words for this page.
The dashboard that runs the tests on every change and turns green when none failed. It is the signal that lets the change ship.
A test that would turn red if the code were wrong in the way that matters. Only that kind turns green into information.
A copy of the code with a deliberate mistake, like a flipped sign. If the tests do not notice it, they would not notice the real one either.
Whatever decides whether the work passed. When whoever is graded can touch the yardstick, the result measures the yardstick, not the work.
A detector with no battery is silent too.
A smoke detector that stays quiet can mean two things: there is no fire, or the detector does not work. From outside, the silence is the same. The only way to know which is to test the detector: put smoke near it and see whether it beeps.
Software tests are the same. A failing test says a lot: something is wrong, and it points to where. A passing test says almost nothing, unless it would have failed if the code were wrong. Green is a sum of silences, and each silence is worth what its test was worth.
The common mistake is not having too few tests. It is reading silence as a certificate: the report says "passed all tests", and the reader hears "it is correct". Those are different sentences, and the gap between them is where defects live.
Passed by the tests, wrong in fact.
SWE-bench is the most cited leaderboard for comparing AI coding agents. Each task is a real issue filed in a public project. The agent writes the fix, and it counts as resolved if it passes the tests chosen for that task. The whole leaderboard is made of greens.
In 2025, two groups went to check those greens. One ran the developers' own full test suite on the accepted fixes: 7.8% passed the task's tests and failed the full suite. And 29.6% of the accepted fixes behaved differently from the fix written by the project's people. Altogether, the resolution rate was inflated by 6.2 percentage points.
The other group added tests where they were missing and found 176 wrong fixes accepted in one version of the leaderboard and 169 in another. With the yardstick fixed, 40.9% of the entries changed position in one of the tables. None of those fixes had lied: the tests simply could not fail for them.
There is an earlier problem, filed in SWE-bench's own repository: in some setups, the agent can edit the test files that will later judge it. When whoever is graded can touch the yardstick, green measures the yardstick.
A strict test charges now. A slack one charges someone else, later.
This part is our reading, not a measurement. A test that is too strict fails on code that worked: it blocks the release, someone loses an afternoon finding out why, and the test gets the blame. A slack test lets a defect through: the release ships on time, and the bill arrives weeks later, in production, for another team or for the customer.
The cost of strictness is immediate and has an owner; the cost of slackness is delayed and diffuse. Without anyone deciding anything, day-to-day pressure pushes suites toward slack: the annoying test gets loosened or deleted, the one that never fails stays. Over time, green gets more frequent and worth less.
The fix is to make slackness expensive for whoever chooses it: measure, every so often, whether the tests can still fail. That is what the specimen does.
Putting smoke near the detector.
The way to test the tests has existed for decades: plant defects on purpose. Make a copy of the code with one small mistake, like flipping a sign or changing a number, and run the suite. If it turns red, the defect was caught. If it stays green, the suite cannot see that kind of mistake. The share of defects caught is the suite's grade, not the code's.
In the specimen, the three tests checked by default are what a common suite tends to have: runs without error, one typical case below the threshold, one above. It stays green and catches five of the eight planted defects. The real defect gets through: the code charges shipping on an order of exactly R$ 199, and none of those tests looks at the threshold.
We tried to take this to real changes and ran into reach. In a partial sample, from a single project, off-the-shelf defect-planting tools almost never managed to plant anything inside the lines the change touched. The idea is good; the shelf tools do not reach where the typical change operates.
A test written from the code protects the defect.
It is now common to ask an AI model to write the tests for a change. We measured that in August, on two public Go projects. We picked 19 changes we knew contained a defect, because it was fixed later, and 10 clean changes of the same size. For each one, an AI agent, without knowing which group it was in and without seeing the fix, wrote tests from the code and the change description.
The generated tests caught none of the 19 defects. The reason is the same as the snapshot of the current code in the specimen: whoever writes the test by looking at the code records what the code does, mistake included. The test is born agreeing with the defect.
Two details matter. In the only two cases where the change description said one thing and the code visibly did another, the test failed and pointed to a real defect. And on the 10 clean changes the method barely invented alarms: it fired once, and rightly. It is good for acquitting, not for accusing. It is the same split we found when one AI checks another, in an earlier entry of this notebook.
| Group | Changes | Generated test turned red | Reading |
|---|---|---|---|
| With a known defect | 19 | 0 | Caught none. The test inherited what the code does. |
| Clean, same size | 10 | 1 | Barely invents alarms, and the one it raised was real. |
| Among the 29, description and code visibly disagreed | 2 | 2 | The only place the generated test bites. |
Our measurement, August 2026, with the losing criterion recorded beforehand: the bet that generated tests would recover defects died. The sample is small, and some tests did not compile against the fixed version, which is a limit of how we checked, not of the method in real use.
Four rules to make green mean something again.
Ask whether the test can fail
Every so often, plant defects and see which get through. The suite's grade is worth more than the test count or the percentage of lines covered.
Write the test from the rule, not the code
A test born from the specification disagrees with the code when the code is wrong. One born from the code agrees with it, mistakes included.
Whoever is graded does not touch the yardstick
Neither a person nor an AI agent changes the tests that will judge their own change without someone seeing it. A test file changed together with the code is a warning in itself.
Call green by its name
In the report, "no alarm in tests X, Y and Z", not "approved". The difference in words is the difference between what was measured and what one would like to be true.
Where this could be wrong.
The leaderboard numbers are someone else's
The SWE-bench percentages come from two 2025 studies, each with its own method. We checked the published summaries; we did not redo the math.
Our sample is small
29 changes, two projects, one language. And what we call a "known defect" depends on linking each fix to the change that caused it, which sometimes goes wrong.
Planting defects has a cost and a blind spot
Running the suite for each planted defect is expensive, and some planted defects do not change behavior, so no test could catch them. The suite's grade is an estimate, not a verdict.
The specimen is a toy
A three-line function makes the mechanism visible. In real code, the defects that escape tend to live where there is not even a simple way to test.
Does this apply to your case?
Tell us in two lines what you need to decide or measure. The first conversation is to see whether measurement solves your case, and if it does not, we say so.
← Research notebook · The cheap one acquits · stickybit.com.br
- Wang, Pradel and Liu, "Are 'Solved Issues' in SWE-bench Really Solved Correctly? An Empirical Study" (arXiv, 2025): 7.8%, 29.6% and the 6.2-point inflation.
- Yu et al., "UTBoost: Rigorous Evaluation of Coding Agents on SWE-Bench" (ACL 2025): 176 and 169 wrong fixes accepted; 40.9% of positions changed on SWE-bench Lite.
- The problem of the agent being able to edit the tests that judge it is filed in the SWE-bench repository's issues.
- Our measurements, August 2026: AI-agent-generated tests for 29 changes in two public Go projects, and the attempt to plant defects in real changes (partial sample, stopped before the end).
- The specimen runs in your browser: one function, six tests and eight copies with one planted defect each.