Source: https://velofy.co/research/verified-agents/

Published: 2026-08-23T00:00:00.000Z
Updated: 2026-08-23T00:00:00.000Z

[RESEARCH](https://velofy.co/research/) / ARCHIVE

# The verified agent loop

Propose, confirm, read back, audit: the one engineering pattern every velofy product implements, and what it looks like in code.

Velofy · 23 Aug 2026

From the archive. This page preserves the original publication; product plans and capabilities may have changed. See [current open source work](https://velofy.co/open-source/).

Axiom three in [our thesis](https://velofy.co/research/first-principles/) says trust scales with verification, not with intelligence. This paper is the engineering consequence: one loop, four steps, implemented in every velofy product.

```
propose   →  the agent renders exactly what it intends to do,
             in a form a human or machine can inspect
confirm   →  an accountable party approves THIS proposal,
             not a category of proposals
read back →  the system re-fetches what actually landed
             and diffs it against the promise
audit     →  all four moments land in a tamper-evident record
```

## Propose: the payload is the interface

In Numera, our accounting console, nothing reaches a ledger without first rendering as a readable proposal: the full entry, debits and credits by account, the period it touches, and the effect on balances. Journal entries are balance-checked before they can even be proposed. The same discipline appears in Kestrel as a dry-run diff preview with atomic per-file rollback, and in heyIAS as a scored rubric with per-dimension marks exposed.

The design rule: **the proposal must be checkable by something weaker than the agent that wrote it.**

## Confirm: accountability is a person or a gate, chosen deliberately

Confirmation takes a different shape per product. Accounting keeps a human: every confirmed write records _who_ confirmed it, and filings carry a chartered accountant's signature. Coding automates the gate: Kestrel's exit code is zero only when the test suite ran green. Evaluation splits the difference: dual-pass scoring at temperature zero, low-confidence answers routed to a human queue, and a sampled slice re-checked by experts weekly.

Auto-approve toggles deserve suspicion everywhere. Numera's defaults off, and even then per-session only.

## Read back: catch the wrong-account post

The step most teams skip. After a confirmed write lands, Numera re-fetches the posted record from the platform and diffs it against the proposal line by line, per account, per currency, per amount. Totals tying is not success: a post to the wrong account with correct totals passes every naive check. Our read-back catches it, and when a platform's API returns no account codes, the verdict says _"checked totals + line count"_.

Writes are idempotent: a deterministic control id means a retried confirmation returns _already-applied_ instead of double-posting. Posting into a closed or locked period is refused before the proposal ever renders. Unwinding is a documented operation, reversing entries that swap every debit and credit and cite the original.

## Audit: a ledger for the ledger

Every executed action lands in an append-only record chained with HMACs, each row signing the previous, so edits, insertions, renumberings and deletions are detectable, and a one-click packet exports a period's writes with their confirmations, read-back verdicts and row hashes for an auditor. We disclose the residual openly in-product: an attacker who deletes tail rows _and_ their seals in lockstep defeats an on-box chain.

## The same loop, three products

|  | Propose | Confirm | Read back | Audit |
| --- | --- | --- | --- | --- |
| **Numera** | Balanced-entry draft, balance-checked pre-render | Named human, maker ≠ checker | Line-level re-fetch diff per account/currency | HMAC-chained ledger + exportable packet |
| **heyIAS** | Rubric scores with per-dimension marks | Confidence threshold → human queue; weekly expert sample | Agreement rate vs expert re-checks, published | Score provenance kept per answer |
| **Kestrel** | Dry-run diff, atomic rollback | Test/build gate decides | Verification command output is the read-back | Run logs + failure reports, reproducible |

## Rules in code, not prompts

Underneath the loop sits one more rule: anything that _must_ hold is versioned code, never a prompt instruction. Debits equal credits is a validator. GST box arithmetic is a validator. Summit's expression interpreter enforcing a CSP-safe allowlist is a validator.

### Verdict

The loop is not free: read-back plumbing per platform, idempotency keys, chained audit storage, sampled human review. Anyone can rent a model; almost nobody will do the work of proving the model right.

See also: [three markets, one thesis](https://velofy.co/research/first-principles/) for why this loop exists, and [taste as a constraint](https://velofy.co/research/taste/) for the same discipline applied to interfaces.

[← Taste is a constraint, not a vibe](https://velofy.co/research/taste/)
