Most product decisions are defended after the fact. We run the argument the other way: start from truths that would still hold if every current vendor vanished tonight, and permit only the products those truths force into existence.
This document is our working example: three markets, one shared derivation. Where the evidence is someone else's, we say whose.
Four axioms
Everything below derives from four claims we hold across every market we enter.
- Customers don't buy software; they buy the disappearance of a recurring chore. Nobody wants a ledger app, a mock-test portal or a code assistant. They want books closed, ranks improved, and pull requests merged.
- Wherever a human works as the connector between systems, there is a wedge for an agent. Provided the output stays verifiable. The accountant re-keying Tally entries and the coach hand-scoring answers are both connectors.
- Trust scales with verification, not with intelligence. Every product needs a sign-off layer a human or a machine can check. A smarter agent with no receipt is worth less than a dumber one with an audit trail.
- At zero stage, distribution beats features. Each product ships one free artifact that is genuinely useful before it asks for money.
The pattern
Stated once, the thesis reads: remove the human who connects systems, replace them with a verified agent loop, keep an accountable check at the point of trust. Applied three ways:
| Market | Old connector | Verified replacement | Sold outcome |
|---|---|---|---|
| Accounting | SMB owner ↔ accountant ↔ Tally | Ledger agent + signing chartered accountant | "Filed, daily" |
| Exam prep | Coach ↔ aspirant ↔ evaluator | Rubric AI + sampled expert review | "Rank trajectory" |
| Coding | Developer ↔ editor ↔ CI | Harness agent + test gate | "Green in one shot" |
The conclusion is the table above.
Case one: accounting
Three truths survive decomposition. N1: a company's books are a deterministic function of its money events: bank transactions, invoices, payroll. N2: the historical cost of accounting is not computation. It is data capture and error recovery: chasing receipts, reconciling statements, re-keying between systems. N3: compliance deadlines are rigid and liability requires a licensed professional, so the correct architecture is agent does the work, accountant signs off.
The strongest external validation comes from Last Accounting Company (YC S26, Finland): one general ledger as the operating layer, sources in once, an agent doing the manual work, accountants signing off, $80M+ in transaction value processed. Their gap is geography: Finland/EU-only, built around Procountor/Fennoa/Netvisor replacements. Nobody dominant applies the model to Indian SMBs: GST filings, Tally and Zoho Books ecosystems, e-invoicing mandates.
Derived decisions: sell the outcome ("books closed and filed, daily"), never seats. Connect banks, Razorpay or Stripe, POS and the email inbox once; reconcile continuously. Build one India-native feature no global player has: GSTR-2B input-tax-credit mismatch detection: credit lost because a supplier did not file, found and chased automatically. The wedge artifact (axiom four): a free GST-readiness audit of a prospect's last quarter.
One deliberate negative decision: we will not build another ledger app for people who like their ledger app. Read-only integration with Tally and Zoho for switchers; replacement only for clients who want zero involvement.
Case two: exam preparation
India's government-exam prep market is not a content market. H1: the outcome is a function of syllabus coverage, practice- with-feedback quality, and long-horizon retention. H2: the scarce resources are personalized feedback, especially Mains answer evaluation, and honest progress measurement. H3: prep spans one to three years; attrition is driven by unmeasured progress and isolation, not lack of videos. H4: aspirants distrust marketing but trust analytics computed from their own data.
The derived product leads with the feedback loop, not content: Mains answers evaluated against published rubrics within minutes, percentile against a real cohort, a revision queue grown from the user's own errors, and a rank trajectory band updated after every mock. The free artifact is a diagnostic mock with a weak-topic map.
Case three: coding agents
K1: a coding agent's value is P(first-shot correctness) times developer time saved. K2: base models commoditize through the API; durable value sits in the harness: context assembly, safe file operations, and a verification loop. K3: "one-shot" must mean something operational: the process exits zero with tests green, or prints an honest failure report. K4: setup costs more than one command and one environment variable lose the audience.
K3 is the whole product. Hence Kestrel's rule, no success claim without a green verification step, and its honesty clause: if a task implies no verification command, refuse politely and suggest one.
What the method kills
Orrery is not an RSS reader with AI summaries: models in the ranking path would violate verifiability; ranking is corroboration, independence and novelty, computed deterministically, and it dies if independent-source detection can't beat a precision bar. Latch holds gig notifications mid-trip and ranks by rupees per effective kilometre, on-device, with no auto-accept, because the rider, not the app, must remain the accountable check. Both scope documents name the milestone that kills them.
How we work
- Dossier before code. Every product gets a first-principles teardown with external evidence attached.
- Rules that must hold go in code, not prompts. A GST validation rule or a debit-equals-credit check is versioned code; the model may propose, it may not assert.
- Every product shares one harness discipline: the verified loop described in its own paper.
- Publish the misses. The blocked Reddit APIs and the dead ends belong in the same document as the wins.
Companion pieces: the verified agent loop and taste as a constraint.