Verified = Command + Output
The bill for one “it should be fine” usually falls due three weeks later.
Of all the disciplines I keep for working with AI, if I could keep only one it would be this: “verified” must mean “command + output.” Without the command pasted in and its real console echo, “done,” “should be fine,” and “the tests passed” all count as nothing said.
I did not think this rule up. A bill taught it to me.
Close to twenty bugs, every one behind a “done”
There was a project I cared about, highly automated, the kind where a wrong number causes real loss. Before it went live I read it through module by module — not the reports, only the code.
That pass turned up close to twenty bugs. What they had in common was not difficulty; quite the opposite, most were visible at a glance. What they had in common was this: every one of them sat inside a stretch of code that had been declared done.
What sent a chill down my back was not the count. It was realizing I had no way, from the outside, to tell which stretch was trustworthy. Every report was written just as beautifully, in the same confident tone.
Why self-acceptance does not hold up structurally
This is not a question of the model’s morals. There are three plainer reasons:
- The context that wrote it is the context that judges it. It reads its own code carrying “this is exactly what I meant,” and so it sees intent rather than behavior. People are no different — proofreading a sentence you have just finished writing is nearly impossible.
- Saying “done” costs almost nothing, while actually running it costs something real. Any system at all will take the cheap path if the cheap path also passes acceptance. That is a fact of structure, not a flaw of character.
- A long conversation dilutes the standard. The fuller the context, the more easily an acceptance clause set down early gets quietly downgraded to “broadly met.”
So I stopped pinning my hopes on making it more honest, and went to change the structure instead.
Three practices that actually land
- Every “verified” travels with its evidence. One line of command, one block of output, pasted next to the conclusion. The rule applies to me exactly as much — I very often want to say “should be fine” myself.
- Whoever writes it does not judge it. Acceptance goes to a window with a fresh context, and the implementer never sees the questions in advance. The cleaner the context, the easier it is to see behavior instead of intent.
- A discipline only counts if a script can run it. This website has one check command, and local and CI run the same one. What it governs is very plain: no Chinese in the code or the docs, a missing document field is an error, the keys of the three locales must line up, the build must pass. It is not clever in the slightest, but it never gets tired, and it never lets me off because today happens to be a rush.
What it actually saves
An unverified “done” is not zero; it is debt. Later work takes it for foundation and puts three more floors on top, and by the afternoon you finally discover the foundation is hollow, what has to come down is no longer that one line. The interest always falls due three weeks later.
At first I thought “command + output” was there to guard against the AI. After living with it a while I found that its greatest use is giving me the nerve to move forward: the chain of evidence is there, so I do not have to go back and re-read everything each time.
This is really the same subject as another article, Signal and Noise — that one is about judging whether a conclusion can be trusted, this one about settling that judgment into a process, so that it does not depend on what kind of day I am having.