Four words that decide the bill.
A piece of a word the model reads or writes. "Unbelievable" becomes about three. It is the unit everything is billed in.
Everything you send the model on each call: instructions, the conversation history, documents, what the tools returned.
Everything the model writes, including the reasoning it does before answering. It usually costs five times more per token than input.
A discount on repeated stretches the provider has already seen in earlier calls. Here, 90% cheaper. Cheap is not free: it is billed again on every call.
"40% fewer" is the good average.
In September, Fireworks launched Ember-1: an open model (Kimi K3) retrained to reason more briefly. The promise is "K3 quality with 40% fewer tokens".
Fireworks' own table, across five public tests, shows reductions from 5.9% to 51.9%. The middle one is 23.7%. On the test closest to real customer support, an airline agent, the reduction was 5.9%. The 40% figure comes from an internal index and from a test with an unnamed customer, where the answer dropped from 49.3k to 29.9k tokens with a practically equal score (0.753 vs 0.751).
None of this is a lie. It is just that "40% fewer tokens" talks about one kind of token, the kind the model writes. The bill you pay has two.
| Test | Tasks | Score before | Ember-1 score | Fewer tokens |
|---|---|---|---|---|
| Terminal tasks (Terminal Bench 2.1) | 89 | 80.9% | 82.0% | 51.9% |
| Coding with conversation (SWE-Interact) | 75 | 21.3% | 20.0% | 32.5% |
| Long coding tasks (DeepSWE 1.1) | 113 | 66.4% | 75.2% | 23.7% |
| Fixing real bugs (SWE-bench Verified) | 500 | 93.2% | 92.2% | 15.5% |
| Airline customer support (τ²-Bench) | 50 | 64% | 66% | 5.9% |
Quality held: four of the five score differences amount to one task more or less, which is noise. The only big difference (DeepSWE, +8.8 points) does not pass a simple statistical test. "The same quality" holds up; "it got better" does not.
Every round resends everything.
An agent is a program that talks to the model several times to finish a task: it asks, calls a tool, reads the result, asks again. The model keeps no memory between calls. So on every round the agent resends everything: the instructions, the documents and the whole history so far.
The history grows every round, and what was sent in the first is paid for again in the second, the third, the fifteenth. What the model writes in each round stays roughly the same size. With few rounds, output dominates; with many, input swallows the bill.
The cache helps: the repeated stretch comes out 90% cheaper. But it makes repetition cheaper; it does not forgive a bloated context. And a model that "thinks less" touches none of this: it only trims the thin band on top.
Short question, hard answer.
The two examples in the specimen up top, with Fireworks' prices, show the two extremes.
One hard question (classifying a tricky case, solving a calculation, planning): one round, little input, lots of reasoning. Almost all the cost is output, and cutting 40% of it saves 39% of the cost. The promise is kept almost in full.
A 15-round support conversation, with long instructions, a growing history and cache: the same model switch saves 17%. And that assumes the 40% cut, when Fireworks' own support test measured 5.9%.
There is a gain the money math does not show: time. The model writes token by token, so writing less means answering sooner. When someone is waiting at the screen, that can be worth more than the money.
The cheaper token lost on the task.
IBM measured 417 agent tasks with two models. The one with the cheaper token cost almost twice as much per task: US$ 155 against US$ 79. The reason is the same as on this page: in agent work, input repeats, and whoever makes better use of the cache wins the bill, not whoever has the smaller price tag.
The same holds the other way around. A model can have cheap tokens and spend many more tokens to reach the answer, or need more rounds. The price that matters is the cost per finished task, measured on your own workload, with the same effort level on both sides.
And the worst spending shows up in no table: the agent that gets stuck in a loop. A Brazilian case published in August: an API returned broken data, the agent retried without stopping, and each attempt carried the previous error in its history. The first cost a thousand tokens; the fifteenth, twenty thousand. A single user's session cost US$ 800.
Four rules for an agent's bill.
Measure cost per task, not per token
Add up input, cache, output and rounds for one finished task. It is the only number that compares models, versions and vendors honestly.
Look at the input first
How many rounds, how much is resent in each, how much hits the cache. In agents, that is where most of the money is, and that is where good design saves.
Cut reasoning where output dominates
One round, short question, hard answer; or when someone is waiting for the answer on screen. Outside that, a model that "thinks less" barely moves the bill.
Set a ceiling before you switch it on
A limit on rounds and tokens per task, enforced before the call, not discovered later on the invoice. A retry loop grows faster than your ability to notice it.
Where this could be wrong.
The calculator is a simple model
We assume the history grows a thousand tokens per round and that reasoning is not resent. Real agents vary a lot; the right math is your own workload, measured call by call.
The Ember-1 numbers are theirs
One run per test, 50 to 500 tasks, and the customer test does not say who the customer is or what the task is. The weights have not been released for us to measure.
Prices change every week
The same week as the launch, GPT-6 Sol and Claude Opus 5.5 came out, with price cuts. The ratio between input and output can change, and the whole math with it.
Each vendor counts tokens its own way
Some bill reasoning together with output, others separately; the cache has its own rate and different rules. Comparing without normalizing this measures the price list, not the cost.
← Research notebook · stickybit.com.br
- Fireworks: Ember-1 (test table, prices used, customer test) · Hacker News discussion
- IBM Research, "Model routing is simple until it isn't" (Hugging Face, July 2026): 417 AppWorld tasks, cost per task versus price per token.
- BIX Tecnologia, "we put AI agents in production for 6 months" (TabNews, in Portuguese, August 2026): the US$ 800 retry loop.
- The calculator and figure math is ours, with the prices above and the assumptions stated in the specimen.