One person runs a software factory: agents plan, implement, review and open the PR — the human sits on approvals and steering, not in the loop. The same discipline as everywhere on this shelf: nothing merges on an agent's word alone, and every number below is counted, not claimed.
Real stage names from the running system — no invented screenshots. Press play to walk one spec through it.
Why a rival vendor in seat 04: measured on this factory's own production data, one model family reviewing its own code raised a blocking concern in 0% of rounds; a rival vendor on the same code — 43%. Review that cannot disagree is not review.
Counted on 2026-08-23. Each card names its own recount, so these rot loudly, not quietly.
~140,000 lines of TypeScript, one person. Planned, implemented, reviewed and merged through the loop above.
471 spec documents. A spec is its own pull request; the implementation is a separate one — acceptance criteria are never an afterthought.
Same model family reviewing its own code: 0% blocking verdicts. A rival vendor on the same code: 43%. So the reviewer seat is always a rival.
The point of the factory is not removing the human — it is moving them to the two seats where a human is irreplaceable.
Nothing merges without a click, and the click is informed: the PR arrives with the spec it implements, the reviewer's verdicts, and CI state attached. The same shape as this lab's publishing gate — verified first, approved second, shipped third.
Review findings are banked, not re-litigated: a concern raised once binds the follow-up work until it is fixed or explicitly waived. The review loop debounces to the branch head so agents converge instead of circling. The human writes specs and priorities; the loop turns them into merged code.
The factory itself is not exposed publicly — it holds write access to real repositories, and autonomy is earned, even here. What you can touch is this shelf: the same verify-then-approve discipline, live, in the cockpit. The factory's fuller story lives at codecot.com.