Method

How I ship a 600,000-line product without reading the code

I don't review AI output line by line. Nobody can at this scale. Instead I spend my time on two things: the decisions only a person can make, and the checks that prove the AI got it right.

1. Direct I set intent, design, rules 2. Build AI agents write the code 3. Verify tests, rules, fuzzers, me 4. Ship only what passes goes out Every miss becomes a new check so the same bug can't come back fails passes

1. Direct: decide what only a person can

AI is very good at writing code to a clear spec. It's bad at knowing what the spec should be. So I own:

I also make the agents ask me questions before they start, so ambiguity gets resolved by a person instead of guessed by a model.

2. Give the agents a map and rules

Agents forget everything between sessions. So the repository itself carries the knowledge:

3. Verify: checks that don't depend on me reading code

The obvious question: if AI wrote the tests, why trust them?

Because a test only counts once it has been seen to fail. The rule is: break the thing on purpose, watch the test go red, then restore it. A test that can't fail is worse than no test. Every one of the 41 fixes in my last big bug sweep shipped with a test that was proven to catch the original bug.

4. Close the loop

When a bug gets past the checks, fixing the code is only half the job. The other half is adding the check that would have caught it. Over time the loop gets harder to fool, and my time goes into new decisions instead of old bugs.

Maximum leverage per decision: I spend my time where only a person can, and build systems so everything else gets done and verified without me.