Software quality · AI development
Securing software quality when AI writes the code:which proof protects which rule?
What you hear at sign-off
All tests green. The software works.
What it actually means
Green means the code and the test agree. Whether they agree on the right thing, nobody has checked that yet.
Acceptance report
Example1,248 testsall green · 87% coverage
View: what the report counts. All five rules are ticked.
Rules that must never break
- 1A statement never goes to the wrong household.passed
- 2An invoice never goes out twice.passed
- 3What's saved is still there after a restart.passed
- 4Blocked beats approved.passed
- 5No payout over 500 euros without a second approval.passed
Tuesday, 4:40 pm, sign-off
Maren (fictitious name) runs a property management firm with nine people. In front of her is her contractor’s report: 1,248 tests, all green, 87 percent coverage. She nods. Wednesday morning a colleague creates a utility statement, the program restarts once, and the statement is gone. Not one of the 1,248 tests is red.
Maren and her numbers are made up. The pattern behind them is real, and the research on it follows in a minute. The question for you: how do you secure software quality when an AI writes the code and the tests along with it? How do you test AI-generated code so that a green check actually means something?
This piece is for you if someone hands you software and you have to decide whether to sign it off. You don’t write unit tests; that’s your contractor’s job. Your question is a different one: how can you tell at sign-off that the software does what your business needs?
The good news first: you don’t have to read any code. You need five sentences about your business and one question you ask about each of them. As a small business, you have a real advantage here, because you know your rules. A large company with 5,000 rules needs a department for this. For your part, a sheet of paper is enough.
I call it
the proof question
Which proof protects this rule? A rule is a sentence about your business that must never break: “An invoice never goes out twice.” A proof is something a machine can run that turns red the moment the rule breaks.
Every important rule deserves a proof, right where a failure would hurt.

01
How can you tell software quality when all tests are green?
By the rules that demonstrably hold. A green check first shows only that code and test agree with each other.
A case that comes up regularly in software development: hundreds of test files, and only a handful of checks that use the product the way a person uses it. Some of the tests check whether a little label has the right color. It shows during a refactor that changes nothing for anyone, and still dozens of tests go red. The other way round, too: things that are genuinely broken, and everything stays green.
So what does a green check actually tell you? Three terms help.
A unit test checks a single small component on its own. Like checking whether one brick can take the load.
Code coverage says how much of the code ran during testing. It doesn’t say whether anyone looked at what came out. 87 percent means 87 percent of the lines were touched once.
And then there’s the mock. To make a test fast, you swap the database, the payment service or the mail program for a stand-in that always answers politely.
Say you rehearse your loan meeting with your brother-in-law playing the bank. He says yes, and you go home relieved. The real bank has different forms, and that is exactly how a test against a mock works. It proves your code gets along with your idea of the database. Only the real database shows whether the idea is right. I call this mock theater: test and code share the same wrong assumption, and both play their parts perfectly. So the first thing to secure is that the mock has the same interface as the real system: it understands the same questions and answers in the same shape. That’s your contractor’s job, not yours. But it helps to understand what quality even means here.
Every brick tested. The house never entered.
A research team led by Asma Hamidi recreated exactly this workflow. Five language models write code and the tests for it. More than 6,000 faulty programs in total. The result, in their words: actual fault detection rates remain extremely low, often near zero (Source: Hamidi et al., 2026). The tests run through the broken spot, but they don’t ask the right question there. To be fair: it’s a preprint, and the same paper says simple bugs get caught well. The hard ones are the problem.
A second paper from the University of Toronto looked at more than 100,000 AI-written test cases from eleven models. Once you can’t be sure the code under test is correct, coverage stops being a reliable sign of whether bugs get found (Source: Zhao, Zhou and Cohen, 2026). This belongs here too: if the code is considered correct and the tests only guard future changes, coverage does work as a signal. Coverage is a diagnostic. It isn’t a verdict on quality.
Maren’s report was true. It just answered a question she never asked.
Actual fault detection rates remain extremely low, often near zero.
02
What changes when AI writes the tests?
Tests cost almost nothing now. So their number no longer works as a measure.
The green-check problem has always been around. Why does it matter now? Because the cost of a test changed. A test used to be expensive. If you had 1,000 tests, someone had thought 1,000 times. Today you order 1,000 tests in one sentence, and the agent delivers before the coffee is done.
One number shows how widespread this is. At Stack Overflow, the share of developers using AI agents climbed from 31 to 59 percent within a year (Source: Stack Overflow, 2026). That’s a smaller pulse survey from April; the full 2026 annual survey hadn’t come out at the time of writing.
What do the agents do differently? A study presented at MSR 2026 analyzed 1.2 million changes from 2,168 projects. 23 percent of the changes made by AI agents touch tests, compared with 13 percent for people. 36 percent of agent changes add mocks to tests, compared with 26 percent for people. The authors write that tests with mocks may be easier to generate automatically but less effective at validating real interactions (Source: Hora and Robbes, MSR 2026).
What this means for bugs is hinted at from elsewhere: a code-review vendor analyzed 470 open code changes: AI-co-authored changes had 10.83 findings on average, human-only ones 6.45. Logic errors were 75 percent more common (Source: CodeRabbit, 2025). That’s a vendor report with a small sample, measured with its own tool. The point for you: your business rules live in the logic errors.
The agent isn’t doing anything wrong here. It does what it gets rewarded for, and what gets rewarded is green. Tell it “write tests” and you get tests. Tell it which rule needs protecting and you get a proof.
In August I argued in “Schrödinger’s Doomer” that verification is one of the three things that gain value once code costs next to nothing. That piece also has the numbers: 96 percent of developers surveyed don’t fully trust AI code, and only 48 percent always check it (Source: Sonar, 2026). Here I take up that single point.
What costs nothing makes a poor yardstick. So you need a different one.
Changes that touch tests
Changes that add mocks
Share of commits in open source projects. A mock replaces a real system in a test with a stand-in that only gives the expected answer.
Source: Hora and Robbes, MSR 2026. 1.2M commits, 2,168 projects. Counts mocks, not missed bugs.
03
How can you check software that an AI built?
With one question per rule: which proof protects it, and does it sit where a failure would hurt?
The other yardstick is surprisingly unspectacular: it fits on one sheet.
You know your business. You know what must never happen. That’s enough to start.
Three questions find your rules. What would embarrass you in front of a customer? What costs money immediately if it goes wrong? What must never end up with the wrong person?
For Maren, these are the five sentences:
- A statement never goes to the wrong household.
- An invoice never goes out twice.
- What’s saved is still there after a restart.
- Blocked beats approved. If an account is blocked, nothing goes out, no matter which other permissions are set.
- No payout over 500 euros without a second approval.
No jargon in there. Anyone in the office can tell whether the sentence is true.
Each sentence becomes a chain: rule, proof, the layer it sits on, and the workflow where the rule shows up in daily life. The old question was: how many tests are there? And how high is coverage? The new one is the proof question: which proof protects this rule? The first question always has an impressive number for an answer. The second gets either an answer or an honest silence.
Practitioners know the idea as risk-based testing. What’s new is who can ask it: as long as tests were expensive, their number was a decent sign of care, and the question stayed with the developers. Now that tests cost next to nothing, it belongs to the person who knows the rules. To you.
Martin Kleppmann of the University of Cambridge put his finger on why the rule part stays with people. Once checking becomes automatic, the difficulty moves to writing the requirement down correctly: how do you know that the properties that were proved are actually the properties you cared about (Source: Kleppmann, 2025)? No machine can take that off your hands. And that’s good news. Your knowledge of your business becomes more valuable as code gets cheaper.
In my view, the rules matter more than the function they happen to be coded in. Code gets rebuilt. The rule stays.
How do you know that the properties that were proved are actually the properties that you cared about?
04
Is the test pyramid outdated?
Its core idea still holds. What's new is the sorting: by the kind of failure a proof is meant to find.
You now have five sentences. The question left is where each proof belongs.
The test pyramid in two sentences: many small, fast tests at the bottom, a few big ones at the top. The shape comes from a time when small tests were cheap to run and big ones expensive, and when every test was written by hand. These days I sort them into four layers, each with a picture from building a house.
A few real journeys
The end-to-end test; I prefer to call it a journey. Walk through the whole house once: create a statement, save it, restart the program, find the statement again. Maren’s Wednesday would have shown up here. The important word is few. A journey checks the path. It doesn’t check a hundred number variants.
Real boundaries
The integration test belongs wherever two parts hand something over: interface to server, server to database, save to reload, old data to new version. My guiding rule: the mock must never replace exactly the boundary you want to test.
Two big outages show what this looks like in real life, in both cases with no part looking suspicious on its own and the failure arising at a handover. At Cloudflare in November 2025, a change in database permissions made a configuration file suddenly double in size. The program reading it had a limit of 200 entries; normal was about 60 (Source: Cloudflare, 2025). By the company’s own account, it was its worst outage since 2019 (Source: Cloudflare, 2025). No attack, no single broken function. A handover. At Amazon Web Services in October 2025, it was a hidden race between two automations that kept overwriting each other’s entry (Source: Amazon Web Services, 2025).
The Uptime Institute expects outages to hinge more and more on the interplay between systems and less and less on a single point (Source: Uptime Institute, 2026). That’s an assessment from the data-center world, not a measurement at small firms. Maren’s Wednesday follows the same pattern: the failure sits at the handover from saving to restarting.
Fast tests for calculation logic
The unit test stays. For prices, deadlines, rounding and cost-allocation keys, small tests are unbeatable: they run fast and point straight at the broken spot. There’s a stronger variant, the property-based test. Instead of three examples, you state a property, for example: the sum of all individual statements always equals the total cost, whatever the numbers. Then the machine tries hundreds of inputs.
A study from UC San Diego looked at 40 projects. One such property-based test finds on average about 50 times as many seeded bugs as an average unit test, 76 percent of them within the first 20 inputs (Source: Ravi and Coblenz, OOPSLA 2025). This was measured with artificially seeded bugs.
A proof for a handful of core rules
The fourth layer sounds like university and, until recently, was exactly that. More on it in a moment.
My principle across all four layers: no test gets thrown out just because it’s small. And none stays just because it’s green.
The proofs check that the business logic is correct. Whether the product feels right, sounds like you and solves your customers’ problem is still yours to check. What that sign-off looks like is covered in “Testing vibe-coded software”.
an average unit test
one property-based test: about 50× as many seeded bugs found
76 % of them within the first 20 inputs
Each square stands for the bugs found by an average unit test. A property-based test checks a rule against many generated inputs instead of one hand-picked example.
Source: Ravi and Coblenz, OOPSLA 2025. 40 projects, measured against artificially seeded bugs.
05
Can you prove software instead of testing it?
For a few small, stable core rules, yes. The proof holds for the model of the rule. You check the bridge to the real code separately.
A test is a sample. You try ten cases and hope the eleventh works too. A proof is a calculation that holds for all cases at once, the way you don’t have to try every number to know that even plus even is even. The technical term is formal verification.
Why has this been out of reach for small and mid-sized businesses so far? The most famous example, a proven operating-system kernel, had 8,700 lines of code. The proof took 20 person-years (Source: Kleppmann, 2025). That was 2009.
Now something is shifting. In December 2025, Martin Kleppmann predicted that AI will bring this method out of its niche and into everyday software engineering (Source: Kleppmann, 2025). Then again, predictions are hard, especially about the future (after a Danish saying often attributed to Niels Bohr; Source: Quote Investigator, 2013). Amazon says it already uses such proofs in several products and writes that as AI agents move money and approve claims, the standard approach to testing is no longer sufficient (Source: Amazon Science, 2026). That comes from a company investing in the technique itself.
The damper comes from a benchmark at ICML 2026. Ask current models to write proven code for classic algorithms and, depending on the proof language, they succeed in 40.3, in 24.7 or in only 7.8 percent of cases (Source: Zhao et al., ICML 2026). An older comparison from September 2025, with different tasks, came out at 82, 44 and 27 percent (Source: Beneficial AI Foundation and MIT, 2025). My reading: it works, the success rate swings widely with task and proof language, and it’s a long way from “one click for the whole product”.
Where the limit lies shows in a case from April 2026. Kiran Gopinathan took a file-compression program that had been proven correct and hit it with about 105 million random inputs. In the proven part: zero memory errors (Source: Gopinathan, 2026). Still, he found two bugs, and both sat outside what the proof covers (Source: Gopinathan, 2026). One was in a part nobody had written a rule for. The other was in the foundation the proof trusts blindly. His conclusion: verification is only as strong as the questions you think to ask and the foundations you choose to trust.
A pattern follows that practitioners will recognize. Write the rule down once, small and exact, and prove it. Then run the real code against that small version with a very large number of random inputs and see whether both always give the same answer. In one Amazon project, the proven model is about ten times smaller than the production code (Source: Lean FRO, undated). In an older write-up from 2024, the proofs there found 4 bugs, and checking the model against the code found 21 more (Source: Disselkoen et al., 2024).
Donald Knuth put this limit into one sentence back in 1977, at the end of a note on an algorithm: “Beware of bugs in the above code; I have only proved it correct, not tried it.” (Source: Knuth, 1977). The rule is proven. Whether the code follows it is something you check on top.
Which rule is a good fit? It’s unambiguous: same input, same result. It rarely changes. A failure would be expensive. It fits on one page and works without a user interface and without a network. For Maren, that’s “Blocked beats approved” and “No payout over 500 euros without a second approval”. The rest stays with journeys and boundary tests.
A single experiment, not a series. Inside the frame, no error turned up. What the frame leaves out, the proof does not check.
Source: Gopinathan, April 2026. Single experiment.
06
How do you know whether a proof is any good?
It turns red when you break the rule on purpose. If it stays green, it's decoration.
Now you have four layers. Which leaves the uncomfortable question: how do you know the new proofs are worth more than the old 1,248 checks?
You plant the bug on purpose. Leave out the save. Swap the recipient. Switch off the block. Then see who notices. The automated version of this is called mutation testing. You can try it yourself right below this section.
Volume isn’t evidence for bug reports either. An AI agent wrote property-based tests for widely used software libraries and reported 984 bugs. In a sample of 50 reports, 56 percent were real bugs; among the top-scoring reports, 86 percent (Source: Anthropic, 2026). So the method finds real bugs. The list of findings still needs checking.
What the rebuild costs
The rebuild has a price. You should know it.
- Speed. A journey through the whole product takes longer than a small test. That’s why calculation logic stays with the fast tests.
- Search effort. When a big journey goes red, the bug can sit in five places. A small test points right at it. That’s why you need both.
- Flakiness. Big tests are sometimes red and sometimes green without anything having changed. My rule: a flaky test counts as broken. It isn’t background noise.
- False security. A proof sounds final. It covers exactly what was written down.
For the rebuild I follow three rules. You can ask the same of anyone who builds your software, person or agent. Old and new run side by side for a while before anything gets deleted. Every deletion has a reason: which rule was protected, and who protects it now? And you measure success by how many relevant kinds of failure you catch with how little maintained test code. The number of deleted tests isn’t success.
A smoke detector nobody has ever tested with smoke is decoration on the ceiling.
Break it and see who notices.
Who notices?
Fault “Skip the save”: turns red for Boundary test against the real database and A flow through the whole product. Every other proof stays green or only sometimes notices.
ASmall tests against mocks
stays greenthe mock always says “saved”BBoundary test against the real database
turns redthe record is missing after reloadingCA flow through the whole product
turns redthe statement can't be found after a restartDProven core rule plus conformance check
stays greenno proven rule covers saving
Kinds of failure your selection catches: 1 of 4
“Only sometimes” doesn't count as caught: it's not something you can rely on.
No single proof catches all four faults.
Simplified model for illustration, not a measurement.
07
Where do you start?
With five sentences on a sheet: the rules that must never break in your business.
You tidy up one corner at a time. Here’s yours: 60 minutes, one sheet, five sentences. Take it to the person or the agent building your software and ask the proof question about each sentence. Then follow up with a second one: show me that it turns red when the rule breaks.
You can hear four answers.
- “There’s a journey test against the real system for that.” Good.
- “That’s in the unit tests.” Ask: against the real database or against a mock?
- “We test that manually before the release.” Honest, but not a proof that goes red on its own.
- Silence. Also an answer. Now you know where to start.
A look ahead, without legal advice. For products with digital elements, reporting obligations under the European Cyber Resilience Act have applied since 11 September 2026, and the main obligations apply from 11 December 2027. The regulation requires technical documentation of the means the manufacturer uses to ensure the requirements are met (Source: Regulation (EU) 2024/2847). And the new EU Product Liability Directive explicitly counts software as a product and applies to products placed on the market from December 2026 (Source: Directive (EU) 2024/2853). Whether your project falls under this is a question for legal advice, not for me. The direction is clear, though: the question of how you prove your software does what it should will be asked more often. If you can answer it with five sentences and five proofs, you have a solid start.
Writing the rules down is human work. Proving them is machine work. When the two are cleanly separated, you no longer have to click through the program yourself every evening to sleep well. Your head is free for what only you can do: the conversation with the customer, the decision about what gets built next.
Every important rule deserves a proof, right where a failure would hurt. With small teams, I build exactly that bridge: from five sentences in the office to proofs in the product that turn red the moment a rule breaks. If your proof question was met with silence, that’s where my work begins.
Take a sheet today and write down the first sentence. The one that just made you a little uneasy.
All names of individuals and companies used in this article are fictitious. Any resemblance to real persons or businesses is purely coincidental and unintentional. The examples are provided solely for illustrative purposes.
| Rule | Proof | Layer |
|---|---|---|
| A statement never goes to the wrong household. | Full flow from creation to sending, recipient checked | Journey |
| An invoice never goes out twice. | Boundary test against the real database, submitted twice | Boundary |
| What's saved is still there after a restart. | Flow with a restart against the real system | Journey |
| Blocked beats approved. | proven rule, code checked against it | Proof |
| No payout over 500 euros without a second approval. | proven rule, code checked against it | Proof |
Example, fictional property management firm
Related to This Topic
Get the free Getting Started Guide: 10 concrete ways to start using AI productively tomorrow.
Did this article spark an idea? Let's find out which Sinnvampire can disappear for you.