Arvid Kahl runs adversarial agent reviews with a third as arbiter
His answer to DHH's claim that nobody will hand review every line of agent code.
DHH said the quiet part out loud first. "It takes a while to build the confidence, but there's no future where you're manually reviewing every line of agent code," he posted, adding that what replaces it is adversarial agent reviews, automated testing and maybe a spot check. From prompt to production.
Arvid Kahl answered with the actual rules he now works by. His list is the most concrete version of this anyone has put up.
humans need to think of themselves as validators, not proofreaders
The rules
Adversarial reviews at every step, ideally with a third agent acting as arbiter between the two. A max coverage test suite from unit to end to end, no exceptions, with TDD baked into the agentic flow rather than bolted on afterwards. Human review, he argues, matters more for the non-code work, particularly architecture and decision-making.
Then the part he calls contextmaxxing, which runs in two directions. In the repo, docs, ARDs, customer ICP descriptions, runbooks and schema commentary all have to live inside the codebase, not merely attached to it. In the process, he tracks every agentic conversation and the changes it made, and attaches that to the ticket or PR.
He also flags dev and testing database content as a bigger deal than it used to be. Dry runs, workflow simulations, edge case hunting and performance measurement should happen before anything hits production, and seeds are not dummy data. They need to be representative of production, he says, if not in volume then in shape.
The role change
The sharpest line is about what humans are for now. Kahl says code that works best for agents does not look best to humans, so the job shifts from reading output to checking that it holds. "Comprehension is now a function of our tools, not a state of our minds," he writes, calling it weird, novel and kind of scary.
The business read is that this is a process claim, not a results claim. Neither post comes with a defect rate or a shipped-bugs number, so what you have is two experienced builders describing how they work, not evidence that it is safer. The test suite is doing the load bearing. If yours is thin, none of the rest applies.
You need adversarial agent reviews, you need automated testing, and maybe you spot check. That's it. From prompt to production!

