Skip to content

Tokens, Models and Prices: What an LLM Call Actually Costs

Site Console Site Console
8 min read Updated Oct 10, 2026 AI & Tools 0 comments

The Unit You Are Billed In Is Not the Unit You Think In

You think in requests. A user asks a question, your service answers, that is one unit of work. It is how you reason about every other API you have integrated.

Language models bill by the token, in both directions, at rates that differ by roughly six to one, with a discount for repeated prefixes that you only receive if you have arranged your prompt to earn it. The same feature, built two different ways, can differ in monthly cost by two orders of magnitude while behaving identically.

This post is the arithmetic. It is the least glamorous post in the series and the one most likely to change a decision you were about to make.


What a Token Is

A token is a chunk of text produced by a byte-pair encoding tokenizer. Common English words are usually one token. Longer or rarer words split into several. Whitespace and punctuation carry their own weight, and the leading space before a word is typically part of that word's token.

The rough rule for English prose is about four characters per token, or roughly three-quarters of a word. That rule is useful for a back-of-envelope estimate and misleading everywhere else. Code tokenizes worse than prose because identifiers and punctuation fragment. Non-English text tokenizes considerably worse, and Vietnamese with full diacritics can consume several times as many tokens as the same meaning in English — which matters a great deal if your users write in Vietnamese and you are paying per token.

Two practical consequences. Your context window is measured in tokens, not characters, so a document that "should fit" might not. And your bill is measured in tokens, so a prompt that reads as short may not be.


Input Costs, Output Costs More

Here is the current OpenAI text pricing, per million tokens, verified on 8 October 2026:

Model

Input

Cached input

Output

GPT-5.6 Sol (flagship)

$5.00

$0.50

$30.00

GPT-5.6 Terra (mid)

$2.00

$0.20

$12.00

GPT-5.6 Luna (small)

$0.20

$0.02

$1.20

Three things in that table are worth more than the numbers themselves.

Output costs six times input. Generation is the expensive half. This means a prompt engineering decision that saves you input tokens is worth far less than one that shortens the answer. "Reply in one sentence" is a cost control, not just a style note.

Cached input is ten times cheaper than fresh input. If the beginning of your prompt is byte-identical across calls — a system prompt, a label taxonomy, few-shot examples, a style guide — the provider can serve it from cache. The requirement is that it is a prefix: the constant material has to come first, before anything that varies. Putting the user's question above your instructions forfeits the discount entirely.

The spread between tiers is 25 to 1. Sol to Luna is twenty-five times on input and twenty-five times on output. No prompt optimization you will ever write comes close to the leverage of using a smaller model that still passes your evaluation.

Two modifiers sit on top. Batch processing takes 50% off input and output for work that does not need an immediate answer. Data residency adds 10%. And the prices above apply to contexts under 270K tokens; beyond that, longer-context rates apply.

Embeddings are priced separately and are cheap by comparison — on the order of a few cents per million tokens for the small model and well under a dollar for the large one. Those figures move, so read them from the pricing page when you size a retrieval corpus rather than trusting a number in a blog post, including this one.


The Trap: Reasoning Tokens Bill as Output

Reasoning models produce internal tokens before they answer. You never see them. You pay for them, at the output rate, which is the expensive rate.

This is the single most common source of a bill that makes no sense. A model answers in thirty tokens and you are billed for three hundred, because it thought for two hundred and seventy. LangChain surfaces this, and it is worth looking at before you commit to a reasoning model for a high-volume path:

response = model.invoke("Hello!")
print(response.usage_metadata)
# {
#   'input_tokens': 8,
#   'output_tokens': 304,
#   'total_tokens': 312,
#   'input_token_details': {'audio': 0, 'cache_read': 0},
#   'output_token_details': {'audio': 0, 'reasoning': 256}
# }

Eight tokens in, and 304 billed out — of which 256 were reasoning. The visible answer was 48 tokens.

Note input_token_details.cache_read in the same structure. That field is how you confirm prompt caching is actually working rather than assuming it. Zero means you are paying full price for your prefix, and the usual cause is that something variable crept in above the constant part.


Estimate Before You Send, Measure After

Two different numbers, and you need both.

For an estimate before the call, tiktoken runs the tokenizer locally:

# pip install tiktoken
import tiktoken

encoding = tiktoken.get_encoding("o200k_base")
tokens = encoding.encode("Explain retrieval in one sentence.")
print(len(tokens))

tiktoken.encoding_for_model("gpt-4o") is the more convenient form, and it raises a KeyError for models it does not yet know — which newly released models often are. Naming the encoding directly avoids that, but you should confirm which encoding your model uses rather than assuming; it is the kind of detail that changes between model families.

For the number you are billed on, read usage_metadata from the response, as above. That is the provider's own count and it is authoritative. A local estimate is for budgeting and for refusing to send a request that will not fit; it is not the invoice.

The habit worth forming is logging usage_metadata from day one, aggregated by feature. Without it, "our AI costs went up" is a mystery. With it, it is a line item.


What a Real Feature Costs

Tutorials stop at the API call. Budgets do not. So here is a full worked example.

A support-ticket classifier. Each call sends a system prompt of about 1,200 tokens — instructions, the label taxonomy, a handful of few-shot examples, all identical every time — plus roughly 400 tokens of ticket text, and gets back about 30 tokens of JSON with a label and an urgency score. The service handles 2,000 tickets a day, so 60,000 calls a month: 96 million input tokens and 1.8 million output tokens.

Configuration

Monthly cost

Sol, no caching

$534.00

Terra, no caching

$213.60

Luna, no caching

$21.36

Luna + prompt caching

$8.40

Luna + caching + Batch

$4.20

The same feature, with the same behaviour, costs $534 or $4.20 depending on four decisions — a factor of 127.

Read the gaps rather than the totals. Choosing Luna over Sol removes 96% of the cost, and it is one string in your code. Caching the system prompt removes 61% of what is left, and it requires only that the constant part sits at the front. Batch halves the remainder, and it costs you latency you may not need.

The order matters: model choice dominates everything else combined. Which is an argument for a workflow rather than a guess.


Choosing a Model Without Guessing

Start at the bottom of the tier list, not the top. Build a small evaluation set — fifty to a hundred real inputs with the answers you want — and run the cheapest model against it. If it passes, you are done, and you have just saved 96% over the instinct to reach for the flagship.

When it fails, look at how it fails before escalating. Failures of format and consistency are usually fixed by a better prompt, a schema, or a few examples, all of which are cheaper than a bigger model. Failures of reasoning — multi-step inference, genuine ambiguity, long-range synthesis — are what you actually buy a larger model for.

And the choice is not global. Routing and classification can run on the small tier while the one step that needs real reasoning calls up a tier. Because the model is a configuration value in LangChain rather than a dependency woven through your code, mixing tiers across a pipeline costs you nothing structurally — which is one of the better arguments for the abstraction in the first place.

This post was written against pricing published on 8 October 2026, tiktoken 0.14.0 and langchain-core 1.6.7. Prices move more often than APIs do; read them from the provider before you commit to a number.


🧭 What's Next

  • Post 4: The OpenAI API Before LangChain — now that you know what you are paying for, it is worth seeing what the framework is wrapping. The next post calls the API directly: the system, user and assistant roles, a chatbot with a deliberately sarcastic personality, and what temperature, max tokens and streaming actually control.

Related

Leave a comment

Sign in to leave a comment.

Comments