Back to System Design Index

AI InfrastructureSeptember 202614 min read

LLM Inference Cost Per Task: Prefill, Decode plus Cache Math

Teams compare vendors by sticker price while the bill hides in the task shape. Output tokens cost several times input tokens. Retries rebill the whole turn. This postmortem prices one support task end to end plus scales it to a worked million-task estimate.

TL;DR: Price per task, not per token. Decode dominates the base. Prefix cache halves the input leg. Retries plus eval sampling multiply everything. The worked model below shows each step with explicit assumptions you can replace.

Modeled Tasks
1M
Input Per Task
8,000 Tokens
Output Per Task
600 Tokens
Prefix Hit Rate
55 Percent

By Mukul Kumar Mishra · Research-led architecture teardown · Published September 29, 2026

LLM inference cost waterfall from task tokens through prefix cache plus batching plus retries to total
Figure 1. The task enters with 8,600 tokens. Cache discounts the prefix. Retries rebill the remainder. The bill hides in the boundary.

1. Every Task Has a Meter Running

A task is not one API call. A support task in this model carries 8,000 input tokens of history plus retrieved context plus 600 output tokens of answer. The meter runs on four legs: full-price input for the uncached prefix, discounted input for cache hits, full-price output for every generated token plus a multiplier for retries with eval sampling on top. Miss one leg plus the forecast misses by multiples, not percents.

Vendors publish per-million-token rates that look comparable until task shape enters. A short chat task with 400 input tokens behaves nothing like an agent task with 8,000. The bill hides in the boundary between what you assume the task looks like plus what production actually sends. Measure mean tokens per task first. Everything below depends on that number.

The mental model: tokens are the raw material. Cache is the discount. Retries are the loan. Price all three before comparing any two vendors.

2. Prefill Versus Decode: Output Dominates

Prefill processes input tokens in parallel. Decode generates output tokens one by one, holding the full context in memory for each step. That serial work is why decode prices run several times higher than input prices at every vendor. In the worked model the 600 output tokens cost more than the 8,000 input tokens before cache enters. Teams that trim prompts while ignoring answer length optimize the cheap leg.

Cap answer length first. Strip reasoning traces from user-facing turns. Route short answers to smaller models where quality holds. Each saved output token is worth several saved input tokens at every price tier. The inference course dedicates three lessons to this split because it decides every downstream choice.

3. Prefix Cache: The Discount You Already Own

Support tasks repeat the same system prompt plus retrieved context across turns. Prefix caching reprices those repeated tokens at roughly one tenth of full input price. At a 55 percent hit rate on the 8,000-token prefix, 4,400 tokens bill at the cached rate while 3,600 bill at full price. The input leg falls by nearly half with zero quality change. Stable prefixes plus deterministic retrieval order raise the hit rate further.

Design for hits. Freeze the system prompt. Order context blocks identically. Keep tool definitions stable across turns. Then measure the hit rate in production instead of assuming it. A ten-point swing in hit rate moves the million-task total more than switching vendors at adjacent tiers.

LeverWhat changesEffect on bill
Stable prefixIdentical prompt orderHit rate climbs
Shorter answersFewer decode tokensBase cost falls fastest
Batched servingShared prefill plus packingPer-task cost compresses
Budgeted retriesCaps plus backoffMultiplier stays near 1x

4. Batching plus the RPS Envelope

One request at a time wastes the accelerator. Batched serving packs concurrent tasks through shared prefill plus denser decode, trading added latency for lower cost per task. The worked model assumes a modest batching benefit of 20 percent off the base, which matches a mid-throughput service with latency SLOs intact. Push batching harder only while tail latency stays inside budget. Past that point the saved dollars buy user churn.

Latency SLOs cap the batch window. Voice tasks with 75ms first-audio budgets batch less than overnight batch jobs. Know the envelope before promising the discount. The course lesson on batching shows where the curve bends for interactive traffic.

5. The Retry Amplifier plus the Eval Tax

Failed attempts rebill the full turn. At a 12 percent task failure rate with one retry each, 120,000 extra turns bill on top of the base million. With no backoff during an incident the policy prices near 6x calm load, which this house prices before approving any retry config. Run your own numbers in the retry-storm calculator before the incident instead of after.

Eval sampling adds a second multiplier. Scoring 5 percent of tasks with a judge running triple depth adds roughly 15 percent of base cost as quality overhead. That spend is correct when it gates bad deploys. It is waste when scores go unread. Budget evals like capacity plus review their catches on the same schedule.

6. Worked Model: One Million Tasks Priced

Assumptions, all labeled estimates: 8,000 input tokens plus 600 output tokens per task. Prefix hit rate 55 percent. Illustration rates of $3 per million full-price input tokens plus $0.30 cached plus $15 output. Batching benefit 20 percent. Retry overhead 12 percent. Eval sampling overhead 15 percent. These rates illustrate the arithmetic. They are not quotes. Replace them with live numbers from the pricing pages below.

Base per task: full input 3,600 tokens at $3 per million equals $0.0108. Cached input 4,400 tokens at $0.30 per million equals $0.00132. Output 600 tokens at $15 per million equals $0.009. Base total equals $0.02112 per task. After the 20 percent batching benefit the base lands near $0.0169. Retry overhead adds 12 percent for $0.0189. Eval sampling adds 15 percent for a final estimate near $0.0218 per task. At one million tasks the modeled month lands near $21,800. Change any assumption plus the total follows linearly except cache, which compounds.

The ordering matters more than the total. Output length first. Cache hit rate second. Retries third. Vendor sticker price last. A team that halves answer length plus lifts cache hits by ten points beats any vendor switch between adjacent tiers without migrating a single call.

7. The Verdict: Meter the Task, Not the Token

Benchmark tables rank vendors. Bills rank task shapes. The team that measures mean tokens per task, pins the prefix for cache hits, caps answer length plus budgets retries will outspend a team chasing the cheapest sticker price every quarter. The bill hides in the boundary between assumed tasks plus production tasks. Close that gap first.

Start today with three numbers from your own logs: mean input tokens, mean output tokens plus retry rate. Paste them into the token-bill calculator with live rates. Then take the full workshop in LLM Inference Economics to route, batch plus gate what you measured.

Frequently Asked Questions

What dominates LLM inference cost per task?

Output tokens dominate because decode prices run several times higher than input prices while retries plus eval sampling multiply the base. Price output first, then cache the repeated prefix.

How does prefix cache change the bill?

A cache hit reprices repeated prefix tokens at the discounted cached rate instead of full input price. At a 55 percent hit rate on an 8,000-token prefix, the input leg falls by nearly half before any other optimization.

Why do agent retries cost more than model choice?

Each failed attempt rebills the full turn at the worst hour with no discount. Price every retry policy at 6x calm load before approving it, because the incident collects that loan.

How do I price my own workload?

Measure mean input plus output tokens per task, your prefix hit rate, batch size plus retry rate, then paste live vendor rates into the token-bill calculator. The worked model above shows every step so you can swap in your numbers.

Sources and Method

Pricing mechanics follow vendor documentation linked below. All counts, rates plus totals are labeled estimates from the assumption table in section 6, not invoices from any provider. Token-meter behavior follows the production teardown in Cursor and Claude Code: Token Limits Postmortem.