Velofy LabDelhi, India
AI research and open source tools for real work.
An applied AI lab. AI research, applied AI, and the harnesses and tool calls that make agentic coding better. We build Troy, the browser for agents, and open source tools we measure.
- Papers
- 46
- Measured runs
- 04
- Open questions
- 70
- Open source tools
- 06
01Products
Troy, the browser for agents.
Get TroyA Chromium browser for macOS and Windows that you use like any other, with one difference: open the debugging port and an AI agent attaches over CDP to the tabs you are already signed into, instead of a fresh browser with no logins. It ships with a Claude Code plugin and skills that teach an agent to work in your session. Free and MIT licensed. Read the documentation.
02Signal
A trained baseline sets the bar at 77 labels.
On a 77-way Banking77 slice, a TF-IDF model trained on 10,003 examples answers in 0.19 ms and beats both zero-shot decision models. At 5 options the picture changes, so option count is a first-order variable.
Read the reportRandom guessing: 0.014. Source: Decision Bench v0.
03Papers
Research, with the evidence attached.
All 46 papers- GLM-5.3 and open-weight cyber capability: what the evidence supports
- Decision Bench v0, first run: three systems on Banking77
- Decision Bench experiments: temperature refit and the option-count sweep
- Decision Bench experiments: near-miss distractors, per-bucket temperature, and shortlists
- Laya: An Open Typed-Decision Engine, Examined
- Jev and the Typed-Decision Landscape
- Typed Decisions Without Generation: A Survey
- Abstention and Honest Uncertainty for Small Decision Models
- Which Encoder Backbone Should a Small Decision Model Start From?
- Maestro Flash: A Non-Autoregressive Multilingual Model for Typed Decisions and Structured Outputs
04Tools
Open source you can run today.
All projects05Index
Areas
Overview- Agent systemsContext, tools, and feedback for software that takes action.
- Applied AIUseful AI at the boundary of unstructured information and working systems.
- AI infrastructureThe retrieval and developer tools underneath agent workflows.
- Human interfacesHow people and agents interact with software, and what makes that interaction clear.
Journal
All notes- The open loop: fast decision models in cloud operationsSmall decision models are ten times cheaper and ten times faster. Whether that makes monitoring better or worse depends on a question nobody has measured yet.
- The same weights are not the same modelAn endpoint is weights plus numerics, decoding, parsers, limits and operations. Each layer can change the answer. How to reason about it, measure it and buy it.
- Building Velofy in the openVelofy brings together applied AI, agent systems, and the tools around them.