One engineer, a room full of agents, and a single discipline running through every project: let the machine do the work, but make it prove what it produced before anyone acts on it. An engineer's working notebook: each method below is explained, animated, and running live — press the buttons.
counted by the machine itself · nothing on this strip is typed by hand
each with its own checkable receipt · press the buttons
A model that answers data questions directly can hallucinate your numbers.
The question compiles to a contract from a whitelist; the engine executes; the model never touches a figure.
Prose that sounds right ships wrong numbers.
Recompute-or-reject: every figure is recomputed from source before a human ever sees it, and a human clicks before it ships.
A method you can only show once is a trick.
The same compute→narrate→verify→approve discipline in five different use cases, each with its money moment.
Models fill gaps by guessing, convincingly.
EXPLAIN before SELECT: the model states what it will claim and where from — before it is allowed to claim it.
Agents that merge on their own word are a liability.
A spec goes in; a reviewed pull request comes out. Gates, multi-model review, a human on approvals — not in the loop.
A model does not judge its own family.
The reviewer seat always belongs to a rival vendor — measured on production data, not assumed.
Fifteen agent sessions and a factory, one person.
One coordinating session dispatches to named peers, opens specs and launches rounds, and verifies what comes back against disk and git rather than against the report.
Capability without restraint is a demo, not a system.
Autonomy is earned in tiers: reading is free; spending waits behind a gate; the seat that pages a human never opens at all.
Every quarter the models get better and the prompt gets rented for less. What does not commoditise is the layer that makes a generated answer safe to act on — the substrate a system accumulates, and the verification that lets someone use its output without checking it by hand.
I don't build assistants that sound confident. I build systems whose output you can ship without re-reading it — because a deterministic gate already did.
Fifteen years inside regulated financial engineering taught me what a wrong number costs. Two years building agent infrastructure taught me how to stop one from reaching a human. Every system here sits on that seam: the place where autonomy meets accountability, and has to be earned.
the summary register · every line checkable · each lives on its method card above
A source figure was withheld; the model derived it from the others — arithmetically correct. The gate rejected it anyway: correct is not verified.
A frontier model in every seat of a multi-agent committee made each agent flawless — and changed the group’s behaviour by nothing (+1.7 pts, p=0.62).
One model family reviewing its own code: 0% blocking verdicts. A rival vendor on the same code: 43%.
10 runs published · 100% of shipped figures recomputed from source · 0 unverifiable numbers reached a reader.
A spec goes in; a reviewed pull request comes out. ~140,000 lines of TypeScript, 471 specs, one person.
11 pipeline runs: 10 published with every figure verified; 1 timed out waiting for a human and shipped nothing. The system fails closed.
the methods above, packaged · prices on the page
Ten productions we built and run ourselves.
Each with the process it handles, where the gate sits, what a human still decides, plus numbers with the command that recounts them.
Three ways in, with the price written down.
A process review, a solution with a gate your people run themselves, or a monthly line that keeps it running.
where the methods run · context, not the showcase