Source: https://velofy.co/research/first-principles/

Published: 2026-08-23T00:00:00.000Z
Updated: 2026-08-23T00:00:00.000Z

[RESEARCH](https://velofy.co/research/) / ARCHIVE

# Three markets, one thesis

How velofy picks what to build: decompose a market to irreducible truths, derive the product from them, publish the derivation. Accounting, exam prep, coding.

Velofy · 23 Aug 2026

From the archive. This page preserves the original publication; product plans and capabilities may have changed. See [current open source work](https://velofy.co/open-source/).

Most product decisions are defended after the fact. We run the argument the other way: start from truths that would still hold if every current vendor vanished tonight, and permit only the products those truths force into existence.

This document is our working example: three markets, one shared derivation. Where the evidence is someone else's, we say whose.

## Four axioms

Everything below derives from four claims we hold across every market we enter.

1.  **Customers don't buy software; they buy the disappearance of a recurring chore.** Nobody wants a ledger app, a mock-test portal or a code assistant. They want books closed, ranks improved, and pull requests merged.
2.  **Wherever a human works as the connector between systems, there is a wedge for an agent.** Provided the output stays verifiable. The accountant re-keying Tally entries and the coach hand-scoring answers are both connectors.
3.  **Trust scales with verification, not with intelligence.** Every product needs a sign-off layer a human or a machine can check. A smarter agent with no receipt is worth less than a dumber one with an audit trail.
4.  **At zero stage, distribution beats features.** Each product ships one free artifact that is genuinely useful before it asks for money.

## The pattern

Stated once, the thesis reads: **remove the human who connects systems, replace them with a verified agent loop, keep an accountable check at the point of trust.** Applied three ways:

| Market | Old connector | Verified replacement | Sold outcome |
| --- | --- | --- | --- |
| **Accounting** | SMB owner ↔ accountant ↔ Tally | Ledger agent + signing chartered accountant | "Filed, daily" |
| **Exam prep** | Coach ↔ aspirant ↔ evaluator | Rubric AI + sampled expert review | "Rank trajectory" |
| **Coding** | Developer ↔ editor ↔ CI | Harness agent + test gate | "Green in one shot" |

The conclusion is the table above.

## Case one: accounting

Three truths survive decomposition. **N1:** a company's books are a deterministic function of its money events: bank transactions, invoices, payroll. **N2:** the historical cost of accounting is not computation. It is data capture and error recovery: chasing receipts, reconciling statements, re-keying between systems. **N3:** compliance deadlines are rigid and liability requires a licensed professional, so the correct architecture is _agent does the work, accountant signs off_.

The strongest external validation comes from Last Accounting Company (YC S26, Finland): one general ledger as the operating layer, sources in once, an agent doing the manual work, accountants signing off, $80M+ in transaction value processed. Their gap is geography: Finland/EU-only, built around Procountor/Fennoa/Netvisor replacements. Nobody dominant applies the model to Indian SMBs: GST filings, Tally and Zoho Books ecosystems, e-invoicing mandates.

Derived decisions: sell the outcome ("books closed and filed, daily"), never seats. Connect banks, Razorpay or Stripe, POS and the email inbox once; reconcile continuously. Build one India-native feature no global player has: **GSTR-2B input-tax-credit mismatch detection**: credit lost because a supplier did not file, found and chased automatically. The wedge artifact (axiom four): a free GST-readiness audit of a prospect's last quarter.

One deliberate negative decision: we will not build another ledger app for people who like their ledger app. Read-only integration with Tally and Zoho for switchers; replacement only for clients who want zero involvement.

## Case two: exam preparation

India's government-exam prep market is not a content market. **H1:** the outcome is a function of syllabus coverage, practice- with-feedback quality, and long-horizon retention. **H2:** the scarce resources are personalized feedback, especially Mains answer evaluation, and honest progress measurement. **H3:** prep spans one to three years; attrition is driven by unmeasured progress and isolation, not lack of videos. **H4:** aspirants distrust marketing but trust analytics computed from their own data.

The derived product leads with the feedback loop, not content: Mains answers evaluated against published rubrics within minutes, percentile against a real cohort, a revision queue grown from the user's own errors, and a rank trajectory band updated after every mock. The free artifact is a diagnostic mock with a weak-topic map.

## Case three: coding agents

**K1:** a coding agent's value is P(first-shot correctness) times developer time saved. **K2:** base models commoditize through the API; durable value sits in the harness: context assembly, safe file operations, and a verification loop. **K3:** "one-shot" must mean something operational: the process exits zero with tests green, or prints an honest failure report. **K4:** setup costs more than one command and one environment variable lose the audience.

K3 is the whole product. Hence Kestrel's rule, no success claim without a green verification step, and its honesty clause: if a task implies no verification command, refuse politely and suggest one.

## What the method kills

Orrery is not an RSS reader with AI summaries: models in the ranking path would violate verifiability; ranking is corroboration, independence and novelty, computed deterministically, and it dies if independent-source detection can't beat a precision bar. Latch holds gig notifications mid-trip and ranks by rupees per effective kilometre, on-device, with no auto-accept, because the rider, not the app, must remain the accountable check. Both scope documents name the milestone that kills them.

## How we work

*   Dossier before code. Every product gets a first-principles teardown with external evidence attached.
*   Rules that must hold go in code, not prompts. A GST validation rule or a debit-equals-credit check is versioned code; the model may propose, it may not assert.
*   Every product shares one harness discipline: the verified loop described in [its own paper](https://velofy.co/research/verified-agents/).
*   Publish the misses. The blocked Reddit APIs and the dead ends belong in the same document as the wins.

### The one-sentence test

A market qualifies if a paid human currently works as a connector between systems, the connector's output can be verified mechanically at the moment of trust, and we can name the free artifact that earns distribution.

Companion pieces: [the verified agent loop](https://velofy.co/research/verified-agents/) and [taste as a constraint](https://velofy.co/research/taste/).

[← The feedback loop exam prep never had](https://velofy.co/research/feedback-loop/)[An accounting operating layer for India →](https://velofy.co/research/india-ledger/)
