SearchBook a call

JOURNAL / Engineering

The same weights are not the same model

An endpoint is weights plus numerics, decoding, parsers, limits and operations. Each layer can change the answer. How to reason about it, measure it and buy it.

When you call a model by name, you are not calling a set of weights. You are calling an endpoint: someone’s servers, software and settings wrapped around those weights. Two endpoints with the same model name can give different answers to the same prompt, at different speeds and at different prices.

This holds for open-weight models served by many hosts. It also holds for closed models sold through more than one cloud. Model names change every few months. This structure does not, so it is worth understanding once.

What an endpoint is made of

Start from what has to happen between your request and the reply. There are six layers, and each one can differ between hosts.

Weights. The file itself. Hosts can serve different revisions, and a model name rarely tells you which.

Numerics. How the arithmetic is done. Weights can be stored at 16, 8 or 4 bits per number. Lower precision is called quantisation. It saves memory and money and can cost accuracy. The kernels that do the multiplication and the precision of the cache also vary. Floating-point addition gives slightly different results in a different order, so even the batch your request lands in can change the last digit.

Decoding. How the next token is chosen. Hosts set their own defaults for temperature, for how much reasoning the model does when you do not say, and for the maximum reply length. Some use a small draft model to speed up generation.

Text plumbing. Models read and write text, and software around them turns it into structure. A chat template turns your messages into the exact string the model was trained on. A reasoning parser separates thinking from the answer. A tool-call parser turns the model’s output into a function name and arguments. A small mistake in any of these changes what the model sees or what you get back, while the weights are untouched.

Limits. The context window, the output cap, and the wall-clock timeout. These look like service terms. They are also accuracy settings, for a reason covered below.

Operations. Which machine serves you, whether the deployment changed last night, what happens when a replica is busy, and how prompt caching is applied.

Layer What differs between hosts How it shows up
Weights Revision served Behaviour shifts with no announcement
Numerics Precision, kernels Small, broad loss of accuracy
Decoding Sampling and reasoning defaults Shorter thinking, more variance
Text plumbing Template and parsers Broken tool calls, lost reasoning
Limits Output cap, timeout Hard questions cut off mid-answer
Operations Routing, updates Results that drift week to week

The useful habit is to treat quality as a property of the whole stack. The weights set a ceiling. Every layer above them can only match it or lower it.

One snapshot

The numbers in this section will age. The pattern in them is the point.

Artificial Analysis re-runs three test sets against each host of a model and reports the result as a percentage of a reference deployment it runs itself at native precision. The sets cover tool calling (500 tasks), hard reasoning (250) and long-context reasoning (25). On 3 October 2026 its page for GLM-5.3 listed 26 endpoints, of which 19 had a score.

Unless the note says otherwise, each endpoint charged the maker’s list price of $1.40 per million input tokens and $4.40 per million output tokens.

Endpoint Accuracy against reference What the benchmarker noted
Modal 100.0% Matches
Nebius, 4-bit 99.7% Matches
Z.ai, the model’s maker 98.7% Matches
DeepInfra 98.6% Matches. The cheapest scored endpoint, at $0.56 and $2.50
Crusoe, 4-bit 97.2% Matches
Together AI 96.4% About 11% fewer reasoning tokens
Databricks 96.3% Output capped at 65,536 tokens. Requests end at about 10 minutes
Novita 95.6% Output capped at 131,072 tokens
Baseten, fast tier 94.2% A 20-minute limit cuts off the longest answers. Priced at $2.10 and $6.60

Four things follow.

Most of the spread is noise. The confidence interval on each score is three to five points either way. Only 4 of the 19 scored endpoints were flagged as below the reference. The maker’s own endpoint scored 98.7%, not 100%. A table like this ranks endpoints less precisely than it appears to.

The losses that had an explanation came from limits and plumbing. Both 4-bit endpoints that were scored matched the reference. The notes on the lower scores describe output caps, time limits, fewer reasoning tokens, and tool calls with mangled or extra parameters. On one endpoint a parameter named size came back with stray characters attached. The usual suspect, quantisation, was not what the evidence pointed to.

Price did not track accuracy. Most hosts charged the same list price. The cheapest scored endpoint came in at 98.6%. A fast tier at one and a half times the list price scored lowest.

Speed varied far more than accuracy. Output speed differed by about ten times between the slowest and fastest host. Accuracy differed by six points.

Why a timeout is an accuracy setting

A reasoning model thinks in output tokens. On the hard reasoning set, the reference used about 45,000 reasoning tokens per task before it answered.

Now do the arithmetic. At 75 tokens per second, 45,000 tokens take ten minutes. A host with a ten-minute timeout and that speed will cut off a typical hard question, and a faster host will still cut off the long tail. The reply comes back truncated or empty, and it is scored as wrong. Nothing about the model changed.

The same applies to an output cap. If thinking counts against the cap, a cap is a limit on how hard a question the endpoint can answer.

This gets more important over time. Each generation of models thinks for longer. Limits that were generous for last year’s model become the main source of error for this year’s.

How to measure an endpoint

Public tables tell you where to look. They cannot tell you about your workload. The method is short.

  1. Pick a reference. Use the maker’s own endpoint, or your own deployment at native precision.
  2. Use your own tasks. A hundred or more, the same set sent to each endpoint.
  3. Find the noise floor first. Run the reference twice. The gap between those two runs is the smallest difference you can believe anywhere else.
  4. Report intervals. A three-point gap on a hundred tasks is not a finding.
  5. Read the counters before the scores. Finish reason, output tokens, reasoning tokens, and the share of tool calls that parse. These show the cause. A drop in reasoning tokens or a rise in length-limited finishes explains a score before you have to guess.
  6. Split by task type. Tool calling fails in the parsers. Long reasoning fails at the limits. Long context fails at the context window. One average hides all three.
  7. Repeat on a schedule. Endpoints change without a version number.

A harness that lets you change the base URL without changing anything else makes this cheap. That is one reason Kestrel treats the endpoint as configuration and keeps the rest of the run fixed.

How to buy

Price per token is the wrong unit. The unit that matters is cost per accepted result.

If a failed task is caught and retried, cost per accepted result is the price of an attempt divided by the success rate. An endpoint that is 30% cheaper and four points less accurate costs 0.70 divided by 0.92, which is 0.76, against 1.00 divided by 0.96, which is 1.04. The cheaper endpoint wins easily.

If a failure is not caught, the sum changes. A wrong answer that reaches a customer or a ledger can cost far more than the call. Then the four points matter and the 30% does not.

So the question to settle first is what a failure costs and whether you would notice it. We work through that calculation in more detail in cost per accepted decision.

Speed is worth paying for in two cases: a person is waiting, or a timeout is cutting answers off. Otherwise it is a number on a chart.

Questions to ask a host

  • At what precision are the weights served, and with which quantisation method?
  • What are the output cap and the request timeout, and does reasoning count against them?
  • What is the default reasoning effort when I do not set one?
  • Which chat template and tool-call parser versions are in use?
  • Which revision of the weights is this, and will I be told when it changes?
  • Can I pin a deployment?

A host that answers these plainly is giving you most of what a benchmark would.

What stays true

The model, the hosts and the prices in the snapshot will all be replaced. What will not change is that an endpoint is a stack, that every layer of the stack can alter the answer, that limits become errors as models think longer, and that the only measurement that settles the question is your own tasks against a reference with the noise floor known.

Source for the snapshot: Artificial Analysis, GLM-5.3 providers, read on 3 October 2026.

← All journal entries