Stickybit.← NotebookPortuguêsThe math · AI cost · Sep 28, 2026
The math

Fewer tokens is not lower cost.

A new model promises "the same quality with 40% fewer tokens". That sounds 40% cheaper. On a single hard question, it almost is. On an agent that talks to the model fifteen times to finish a task, the savings shrink to a fraction of that, because what weighs on the bill is not what the model writes: it is what you send it on every round.

Specimen · what one agent task costs
input: what you sendoutput: what it writeswhat the 40% cuts
–cost per task today
–with the model that thinks less
–actual money saved
–less waiting while it writes

Prices used by Fireworks itself: US$ 3 per million input tokens, US$ 0.30 with cache, US$ 15 per million output tokens. Our assumptions: each round adds a thousand tokens to the history (the visible answer plus what the tools returned), reasoning is not resent, and the model writes about 80 tokens per second.

Before the math

Four words that decide the bill.

Token

A piece of a word the model reads or writes. "Unbelievable" becomes about three. It is the unit everything is billed in.

Input

Everything you send the model on each call: instructions, the conversation history, documents, what the tools returned.

Output

Everything the model writes, including the reasoning it does before answering. It usually costs five times more per token than input.

Cache

A discount on repeated stretches the provider has already seen in earlier calls. Here, 90% cheaper. Cheap is not free: it is billed again on every call.

What was promised

"40% fewer" is the good average.

In September, Fireworks launched Ember-1: an open model (Kimi K3) retrained to reason more briefly. The promise is "K3 quality with 40% fewer tokens".

Fireworks' own table, across five public tests, shows reductions from 5.9% to 51.9%. The middle one is 23.7%. On the test closest to real customer support, an airline agent, the reduction was 5.9%. The 40% figure comes from an internal index and from a test with an unnamed customer, where the answer dropped from 49.3k to 29.9k tokens with a practically equal score (0.753 vs 0.751).

None of this is a lie. It is just that "40% fewer tokens" talks about one kind of token, the kind the model writes. The bill you pay has two.

promise: 40%middle test: 23.7%
Terminal tasksTerminal Bench 2.1
51.9%
Coding with conversationSWE-Interact
32.5%
Long coding tasksDeepSWE 1.1
23.7%
Fixing real bugsSWE-bench Verified
15.5%
Airline supportτ²-Bench
5.9%
Reduction in tokens written by the model on each test, against the original model. Fireworks' numbers, one run per test; the tests have 50 to 500 tasks.
TestTasksScore beforeEmber-1 scoreFewer tokens
Terminal tasks (Terminal Bench 2.1)8980.9%82.0%51.9%
Coding with conversation (SWE-Interact)7521.3%20.0%32.5%
Long coding tasks (DeepSWE 1.1)11366.4%75.2%23.7%
Fixing real bugs (SWE-bench Verified)50093.2%92.2%15.5%
Airline customer support (τ²-Bench)5064%66%5.9%

Quality held: four of the five score differences amount to one task more or less, which is noise. The only big difference (DeepSWE, +8.8 points) does not pass a simple statistical test. "The same quality" holds up; "it got better" does not.

Why input rules

Every round resends everything.

An agent is a program that talks to the model several times to finish a task: it asks, calls a tool, reads the result, asks again. The model keeps no memory between calls. So on every round the agent resends everything: the instructions, the documents and the whole history so far.

The history grows every round, and what was sent in the first is paid for again in the second, the third, the fifteenth. What the model writes in each round stays roughly the same size. With few rounds, output dominates; with many, input swallows the bill.

The cache helps: the repeated stretch comes out 90% cheaper. But it makes repetition cheaper; it does not forgive a bloated context. And a model that "thinks less" touches none of this: it only trims the thin band on top.

ROUND16k in28.5k in311k in413.5k in516k in618.5k inAdding up all six: 73.5k input tokens, 12k output.instructions + docshistoryoutputthe 40% cut
Six rounds of an agent. The gray and light-blue parts are input, resent every time; the dark part is what the model writes. The 40% cut only reaches the dark part.
Where it really pays

Short question, hard answer.

The two examples in the specimen up top, with Fireworks' prices, show the two extremes.

One hard question (classifying a tricky case, solving a calculation, planning): one round, little input, lots of reasoning. Almost all the cost is output, and cutting 40% of it saves 39% of the cost. The promise is kept almost in full.

A 15-round support conversation, with long instructions, a growing history and cache: the same model switch saves 17%. And that assumes the 40% cut, when Fireworks' own support test measured 5.9%.

There is a gain the money math does not show: time. The model writes token by token, so writing less means answering sooner. When someone is waiting at the screen, that can be worth more than the money.

One hard question1 round · 2k input · 20k writteninput US$ 0.01 · output US$ 0.30savings: 39%Customer support15 rounds · 12k resent · 80% cached · 800 writteninput US$ 0.24 · output US$ 0.18savings: 17%
Where each task's cost comes from, and how much of it the model that thinks 40% less manages to cut. Our math, with Fireworks' prices.
The list price is not the cost

The cheaper token lost on the task.

IBM measured 417 agent tasks with two models. The one with the cheaper token cost almost twice as much per task: US$ 155 against US$ 79. The reason is the same as on this page: in agent work, input repeats, and whoever makes better use of the cache wins the bill, not whoever has the smaller price tag.

The same holds the other way around. A model can have cheap tokens and spend many more tokens to reach the answer, or need more rounds. The price that matters is the cost per finished task, measured on your own workload, with the same effort level on both sides.

And the worst spending shows up in no table: the agent that gets stuck in a loop. A Brazilian case published in August: an API returned broken data, the agent retried without stopping, and each attempt carried the previous error in its history. The first cost a thousand tokens; the fifteenth, twenty thousand. A single user's session cost US$ 800.

GPT-4.1· cheaper tokenUS$ 155Claude Sonnet· pricier tokenUS$ 79total cost of the 417 tasks
Total cost of the same 417 agent tasks (IBM Research, July 2026). Ranking by token price and ranking by task cost come out reversed.
What to do

Four rules for an agent's bill.

  1. Measure cost per task, not per token

    Add up input, cache, output and rounds for one finished task. It is the only number that compares models, versions and vendors honestly.

  2. Look at the input first

    How many rounds, how much is resent in each, how much hits the cache. In agents, that is where most of the money is, and that is where good design saves.

  3. Cut reasoning where output dominates

    One round, short question, hard answer; or when someone is waiting for the answer on screen. Outside that, a model that "thinks less" barely moves the bill.

  4. Set a ceiling before you switch it on

    A limit on rounds and tokens per task, enforced before the call, not discovered later on the invoice. A retry loop grows faster than your ability to notice it.

What we don't know yet

Where this could be wrong.

The calculator is a simple model

We assume the history grows a thousand tokens per round and that reasoning is not resent. Real agents vary a lot; the right math is your own workload, measured call by call.

The Ember-1 numbers are theirs

One run per test, 50 to 500 tasks, and the customer test does not say who the customer is or what the task is. The weights have not been released for us to measure.

Prices change every week

The same week as the launch, GPT-6 Sol and Claude Opus 5.5 came out, with price cuts. The ratio between input and output can change, and the whole math with it.

Each vendor counts tokens its own way

Some bill reasoning together with output, others separately; the cache has its own rate and different rules. Comparing without normalizing this measures the price list, not the cost.

← Research notebook · stickybit.com.br

Sources