How I ship a 600,000-line product without reading the code
I don't review AI output line by line. Nobody can at this scale. Instead I spend my time on two things: the decisions only a person can make, and the checks that prove the AI got it right.
1. Direct: decide what only a person can
AI is very good at writing code to a clear spec. It's bad at knowing what the spec should be. So I own:
- What to build — from running a service business, not from guessing.
- The shape of the system — what the major parts are and how they connect. For example, splitting the business rules into a shared package so the offline phone and the server can't disagree.
- The data model — the part that's most expensive to get wrong, so it's the part I review closely.
- The interface — AI still builds weak UI. I direct it screen by screen and test it by hand on the phone and the web.
I also make the agents ask me questions before they start, so ambiguity gets resolved by a person instead of guessed by a model.
2. Give the agents a map and rules
Agents forget everything between sessions. So the repository itself carries the knowledge:
- An agent handbook: where things live, how a feature is put together, which rules are non-negotiable.
- Design docs for every subsystem. When an agent learns something hard-won, it writes it into the doc that owns the topic — not into a memory that can go stale.
- Every claim is tagged by how it's known: measured, inferred, or reported. "I think it works" and "I ran it" are never written the same way.
3. Verify: checks that don't depend on me reading code
- Architecture rules (140 of them) fail the build when code crosses a boundary it shouldn't.
- Behavior tests — a third of the codebase.
- Fuzzers that throw random, hostile scenarios at sync and security and check that the important promises still hold.
- One gate command that runs all of it. Work isn't done until it passes.
- Me, using the real app on real devices, the way a customer would.
The obvious question: if AI wrote the tests, why trust them?
Because a test only counts once it has been seen to fail. The rule is: break the thing on purpose, watch the test go red, then restore it. A test that can't fail is worse than no test. Every one of the 41 fixes in my last big bug sweep shipped with a test that was proven to catch the original bug.
4. Close the loop
When a bug gets past the checks, fixing the code is only half the job. The other half is adding the check that would have caught it. Over time the loop gets harder to fool, and my time goes into new decisions instead of old bugs.
Maximum leverage per decision: I spend my time where only a person can, and build systems so everything else gets done and verified without me.