codecot · applied AI
How to build with AI, now

The model is not the system. The system is what earns your trust.

One engineer, a room full of agents, and a single discipline running through every project: let the machine do the work, but make it prove what it produced before anyone acts on it. An engineer's working notebook: each method below is explained, animated, and running live — press the buttons.

01
The engine computes
Deterministic, testable, no model in the numbers.
02
The model narrates
An LLM turns computed truth into language.
03
A gate verifies
Every figure recomputed from source, or rejected.
04
A human approves
Autonomy is earned, not granted.

New here? Take the four-stop route →

What the machine did this fortnight

counted by the machine itself · nothing on this strip is typed by hand

and what is not shown here, on purpose

  • Revenue, waitlist. No such data exists.
  • Visitors. Not measured; request logging is off.
  • Uptime, latency. Not measured; there is no canary.
  • Gate refusals. Currently zero, and a counter reading zero would argue the gate is decorative.
  • Pipeline runs. Eleven: a number that undercuts rather than supports.

The methods

each with its own checkable receipt · press the buttons

How I see it

Building has become cheap. Being able to act on the output has not.

Every quarter the models get better and the prompt gets rented for less. What does not commoditise is the layer that makes a generated answer safe to act on — the substrate a system accumulates, and the verification that lets someone use its output without checking it by hand.

I don't build assistants that sound confident. I build systems whose output you can ship without re-reading it — because a deterministic gate already did.

Fifteen years inside regulated financial engineering taught me what a wrong number costs. Two years building agent infrastructure taught me how to stop one from reaching a human. Every system here sits on that seam: the place where autonomy meets accountability, and has to be earned.

Receipts

the summary register · every line checkable · each lives on its method card above

A source figure was withheld; the model derived it from the others — arithmetically correct. The gate rejected it anyway: correct is not verified.

the withhold demo · reproducible live→ method

A frontier model in every seat of a multi-agent committee made each agent flawless — and changed the group’s behaviour by nothing (+1.7 pts, p=0.62).

the multi-agent study · 57 paired games · $8

One model family reviewing its own code: 0% blocking verdicts. A rival vendor on the same code: 43%.

462 review rounds · production data, 2026-08-19→ method

10 runs published · 100% of shipped figures recomputed from source · 0 unverifiable numbers reached a reader.

the lab ledger · ~$0.003/run→ method

A spec goes in; a reviewed pull request comes out. ~140,000 lines of TypeScript, 471 specs, one person.

the factory · counted 2026-08-23→ method

11 pipeline runs: 10 published with every figure verified; 1 timed out waiting for a human and shipped nothing. The system fails closed.

Step Functions execution history→ method

What this is for sale as

the methods above, packaged · prices on the page

Built with these methods

where the methods run · context, not the showcase

FactoryThe agent platform behind the SDLC method: plans, implements, reviews and opens the PR.the method →
Applied-AI labThis site’s backend: the publishing pipeline, verified commentary, queries and the agent — on AWS.the cockpit →
CountersignThe thesis under everything: verified AI output for regulated finance, in development.
Hiring DeskScreening and interview support with an anti-souffleur integrity layer.
Astro Enginepyswisseph computations under ASR-verified voice; daily generated shorts.
VoiceBridgeSelf-hosted STT/TTS bridge: faster-whisper, multi-client, self-updating.+ Tangent, in development →
Movie FactoryMedia pipeline with a capability router across local and cloud generators.
Research LabMeasured multi-agent studies — predictions registered before the data, nulls beside every number.
Chat LibraryThree sources, 74k messages, full-text search over the owner’s working conversations.
WhiteRabbitContext studio — the design of record for the cockpit line.