Axiom three in our thesis says trust scales with verification, not with intelligence. This paper is the engineering consequence: one loop, four steps, implemented in every velofy product.
propose → the agent renders exactly what it intends to do,
in a form a human or machine can inspect
confirm → an accountable party approves THIS proposal,
not a category of proposals
read back → the system re-fetches what actually landed
and diffs it against the promise
audit → all four moments land in a tamper-evident record
Propose: the payload is the interface
In Numera, our accounting console, nothing reaches a ledger without first rendering as a readable proposal: the full entry, debits and credits by account, the period it touches, and the effect on balances. Journal entries are balance-checked before they can even be proposed. The same discipline appears in Kestrel as a dry-run diff preview with atomic per-file rollback, and in heyIAS as a scored rubric with per-dimension marks exposed.
The design rule: the proposal must be checkable by something weaker than the agent that wrote it.
Confirm: accountability is a person or a gate, chosen deliberately
Confirmation takes a different shape per product. Accounting keeps a human: every confirmed write records who confirmed it, and filings carry a chartered accountant's signature. Coding automates the gate: Kestrel's exit code is zero only when the test suite ran green. Evaluation splits the difference: dual-pass scoring at temperature zero, low-confidence answers routed to a human queue, and a sampled slice re-checked by experts weekly.
Auto-approve toggles deserve suspicion everywhere. Numera's defaults off, and even then per-session only.
Read back: catch the wrong-account post
The step most teams skip. After a confirmed write lands, Numera re-fetches the posted record from the platform and diffs it against the proposal line by line, per account, per currency, per amount. Totals tying is not success: a post to the wrong account with correct totals passes every naive check. Our read-back catches it, and when a platform's API returns no account codes, the verdict says "checked totals + line count".
Writes are idempotent: a deterministic control id means a retried confirmation returns already-applied instead of double-posting. Posting into a closed or locked period is refused before the proposal ever renders. Unwinding is a documented operation, reversing entries that swap every debit and credit and cite the original.
Audit: a ledger for the ledger
Every executed action lands in an append-only record chained with HMACs, each row signing the previous, so edits, insertions, renumberings and deletions are detectable, and a one-click packet exports a period's writes with their confirmations, read-back verdicts and row hashes for an auditor. We disclose the residual openly in-product: an attacker who deletes tail rows and their seals in lockstep defeats an on-box chain.
The same loop, three products
| Propose | Confirm | Read back | Audit | |
|---|---|---|---|---|
| Numera | Balanced-entry draft, balance-checked pre-render | Named human, maker ≠ checker | Line-level re-fetch diff per account/currency | HMAC-chained ledger + exportable packet |
| heyIAS | Rubric scores with per-dimension marks | Confidence threshold → human queue; weekly expert sample | Agreement rate vs expert re-checks, published | Score provenance kept per answer |
| Kestrel | Dry-run diff, atomic rollback | Test/build gate decides | Verification command output is the read-back | Run logs + failure reports, reproducible |
Rules in code, not prompts
Underneath the loop sits one more rule: anything that must hold is versioned code, never a prompt instruction. Debits equal credits is a validator. GST box arithmetic is a validator. Summit's expression interpreter enforcing a CSP-safe allowlist is a validator.
See also: three markets, one thesis for why this loop exists, and taste as a constraint for the same discipline applied to interfaces.