Daniel Kahneman split thinking into two systems. System 1 is fast, automatic and runs on recognition. System 2 is slow, deliberate and runs on reasoning. Most of what a frontier model does when you ask it something hard is System 2, and it bills accordingly.
A lot of infrastructure work does not look like System 2. Is this alert real. Which runbook step applies. Roll back or hold. Is this line dead or just slow. These have a closed set of answers and a deadline attached. You do not reason your way to them so much as recognise them.
That observation is the whole pitch for small decision models in operations, and the economics behind it are real. What I am less sure about is whether putting one into a monitoring system makes the system better or worse. This post is about why that is a genuinely open question, and what you would have to measure to close it.
The economics are not in doubt
A typed-decision model takes some state and a declared question, and returns a label, a score or a boolean with a confidence value. It does not generate prose. The independent benchmark we survey in what a million decisions costs measured cost and latency on a 77-way banking intent task:
| Contender | Cost per million decisions | p50 latency |
|---|---|---|
| Jev | $70 | 264 to 276 ms |
| gpt-5.4-nano | $190 | 684 to 782 ms |
| glm-5.3-flash | $210 | 2.0 to 3.5 s |
| claude-sonnet-4.6 | $2,480 | 1.8 to 2.6 s |
Ten times cheaper than the next thing, and ten to sixteen times faster than models in thinking mode. The latency is flat from 2 options to 255, which matters if your runbook has forty branches. Self-hosted encoders go further still: the Laya project reports 33 ms for a single question and 7.2 ms per question batched on one T4 card.
So the case for speed is settled. A decision that used to cost a tenth of a cent and most of a minute now costs a thousandth of a cent and a quarter of a second. The interesting question is what you do with that, and whether the answer is always “more”.
Where fast decisions are obviously good
It helps to split operational decisions into two kinds.
In the first kind, the decision does not change the thing being measured. Deciding which team owns an alert does not alter the alert. Classifying whether a log line is worth indexing does not alter the log. Deciding whether this page is the same incident as that one, or which service a trace belongs to, or whether a stack trace matches a known issue: none of these feed back into the signal.
Call these open loop. Here the case for a fast model is strong and mostly uncomplicated. A wrong answer costs some human attention and can be escalated. Latency is pure loss, so removing it is pure gain. If you are routing a hundred thousand alerts a month, the difference between $70 and $2,480 per million decisions is a real line item, and the difference between 250 ms and three seconds is the difference between enrichment that happens before a human opens the ticket and enrichment that happens after.
If I were putting a small decision model into production tomorrow, it would be here, and I would expect it to work.
Where it gets interesting
In the second kind, the action changes the signal. An autoscaler scales up, load per instance falls, the metric that triggered the scaling moves. A remediation agent restarts a service, the health check changes. Traffic gets shifted, latency on the remaining path rises. The decision is inside a feedback loop with the system it is deciding about.
Call these closed loop. And here is the thing that gives me pause: in a closed loop, the latency of the controller is not simply a cost. It is part of the dynamics.
Operations has known this for a long time, and the evidence is sitting in the defaults of tools everyone runs.
Nagios stores the results of the last 21 checks of a host or service, works out what percentage of them were state transitions, and if that percentage crosses a threshold it declares the service to be flapping and suppresses further notifications. The twenty-one-check window is not a performance compromise. It exists so that the system refuses to react to a signal until it has watched it for a while.
Kubernetes does the same thing one layer up. The horizontal pod autoscaler applies a stabilisation window to scale-down decisions, five minutes by default, configurable through --horizontal-pod-autoscaler-downscale-stabilization. During that window it does not use the current recommendation. It uses the highest recommendation it has seen across the window. A brief dip in traffic will not scale you down, because the controller has been deliberately built to distrust anything it has only seen once.
Two different layers, two different decades, the same conclusion: in a control loop, slow down the decision on purpose.
That is the sentence I keep returning to. The pitch for System 1 models in operations is “make the decision ten to sixteen times faster”. A quarter of a century of operations practice says that the delay in these specific loops was load-bearing. Those two statements are in direct tension and I have not seen anyone test which one wins.
Why a fast model might make it worse
The mechanism is not exotic. If a controller acts faster than the system it controls can settle, it acts on signal that has not finished responding to its last action. It overcorrects, then corrects the overcorrection, and the loop oscillates. The fix taught in every control course is to add damping, and one way to damp a loop is to make the controller slower.
A decision model at 250 ms, dropped into a remediation loop, removes most of the damping that was there. The replacement damping has to come from somewhere else: a stabilisation window, a rate limit, a cooldown, a hysteresis band. All of which are ways of reintroducing the delay you just paid to remove.
The calibration evidence makes this worse rather than better. The abstention survey collects three findings that are uncomfortable in this context. An independent benchmark found Jev admitting ignorance on 49.7% of forced-uncertainty items where frontier models admit on 97.3 to 100%. Its expected calibration error measured 0.246 against a frontier range of 0.039 to 0.122. And AbstentionBench, across 20 models and 20 datasets, reports that abstention is unsolved, that scale does not fix it, and that reasoning fine-tuning actually degrades it by 24% on average.
An overconfident controller in a feedback loop is the bad case. It will not hesitate at the moment hesitation is the correct move, because hesitation is exactly the behaviour that is missing.
So there is a plausible story where small decision models improve alert triage enormously and make auto-remediation measurably less stable, in the same deployment, for the same reason.
What nobody is measuring
Here is the gap. Every benchmark I have read for these models, our own survey included, measures single-shot accuracy, latency and cost against a static dataset. Banking intents, 77 ways, graded once. That is a reasonable way to measure a classifier.
It is not a way to measure a controller. None of them put the model inside a loop with the system it is classifying and watch what the loop does over an hour.
If you wanted to settle this, the experiment is not complicated:
- Stability, not accuracy. Put the model in a simulated autoscaling or remediation loop against a workload with realistic noise. Count state changes per hour against a slow baseline. If the fast controller oscillates where the slow one converges, that is the finding.
- Time to stable state. Measure how long the system takes to settle after an injected fault, not whether the individual classifications were right. A sequence of correct decisions can still produce an unstable loop.
- Does abstention damp it. Give the model a confidence threshold and an escalation path, then check whether abstaining under uncertainty restores the stability that speed removed. My guess is that it does, and that this turns out to be the main reason abstention matters operationally rather than any argument about honesty.
- Cost per resolved incident. Not cost per decision. A controller that is ten times cheaper per decision and makes forty times as many decisions is not cheaper. We sketched the denominator problem in the cost-per-accepted-decision design and this is the operational version of it.
Point four is the one I would run first, because it is the one most likely to produce an unflattering result, and unflattering results are the useful kind.
Where this leaves me
Open loop decisions: the case is strong, the economics are measured, and I would ship it.
Closed loop decisions: unknown. The prior from operations practice is that the delay was doing work, and the calibration numbers suggest the fast models are currently least trustworthy in exactly the situation where trust matters most. That is not an argument against trying. It is an argument for measuring the loop rather than the classifier, which is not what anyone is doing today.
I am going to build the simulation and find out. The result will go up here whichever way it falls, and if the answer is that a 250 ms controller thrashes a loop that a slow one holds steady, that is worth knowing before someone puts one in front of a production autoscaler.
Decision models deserve study in their own right, and operations is a better laboratory for them than a static intent dataset, precisely because operations is where a wrong answer argues back.
Sources for the numbers above: what a million decisions costs and abstention and honest uncertainty, both of which cite the underlying independent benchmarks. The Nagios flap detection behaviour and the Kubernetes stabilisation window are documented in their respective project documentation, read on 4 October 2026.