veridive is now an applied AI company. Looking for the answer engine?Looking for the answer engine? What happened

veridive TR Start a project Menu

Field notesEconomics

What is a token? The unit behind AI costs and limits, explained.

Tokens are how language models read, write and bill. They explain context limits, why long documents cost more, why output is priced differently from input, and why the same text can cost differently across languages and models.

veridive5 min read

Every model price list, context limit and usage report counts the same unit, and it isn’t the word. It is the token. You don’t need to know how models work inside to buy them well, but you do need to know what a token is, because it explains many of the surprises in an AI bill.

Each term below gets a plain definition and the question it should make you ask. The short version: tokens are how models read, write and bill, so budget per finished task, not per token.

What is a token?

A token is the unit of text a language model reads and writes: often a whole short word, sometimes a piece of a longer word, a punctuation mark or a space. The tokenizer, the part of the system that splits text into tokens, usually differs from one model family to the next.

Here is an illustrative split of one sentence. Every model splits text differently, so treat it as a picture of the idea, not a measurement:

Please · check · the · attached · invoice · .

Six tokens for five words and a period. Common English words are often one token each; longer or rarer words break into pieces, and numbers, codes and unusual names can break into many.

Ask: how many tokens does a typical case in our workflow take, measured with the model we plan to use?

Why do models bill by the token?

Input tokens are everything you send: instructions, the question, retrieved documents and earlier conversation. Output tokens are everything the model writes back, including, for some models, intermediate reasoning that is billed even when you don’t see it.

Models bill by the token because the work scales with them. Every input token has to be processed, and every output token is generated one step at a time, which is why output is often priced higher than input. Some providers also charge less for input they have already processed, such as a long set of instructions repeated on every call.

That makes context the hidden driver. In a system that answers from your documents, each question sends the retrieved passages as input. Send ten long passages instead of three short ones and the input per task multiplies, even though the user typed one line. Long conversations do the same, because the history is sent again with every turn.

Ask: how many input and output tokens does one finished task use, including retrieved context, history and retries?

What is a context window, and why does it matter?

The context window is the most tokens a model can consider at once: instructions, input, retrieved text and its own output together. Anything that doesn’t fit is cut off or has to be left out.

A long contract may not fit whole. A long chat pushes its beginning out. Even inside the window, a very long input can make it harder for the model to use the one paragraph that matters, so bigger is not automatically better, and a full window costs more per call. The usual design answer is retrieval: send the relevant passages, not the whole library.

Ask: what happens when our longest document doesn’t fit, and would we notice?

Why can the same text cost differently?

The same paragraph can produce a different bill on two models for several reasons at once:

  • Different tokenizers split it into different numbers of tokens.
  • Different prices per input and output token.
  • Different habits: one model answers in three lines, another in ten.
  • Hidden overhead: instructions, tool descriptions and formatting rules resent with every call.
  • Retries and reasoning: failed attempts and intermediate reasoning are billed too.

So a price per million tokens compares badly across models. Cost per finished task, measured on the same examples, compares well.

Ask: what is the cost per task on our own evaluation set, not the price per token?

What does this mean for Turkish and other languages?

Turkish builds meaning by adding suffixes to a root, so one word form can carry what English says in several words. Tokenizers built mostly on English text often split those forms into more pieces than English words.

An illustrative split: faturalarımızdan, “from our invoices”, might become fatura · lar · ımız · dan, one word and four tokens. The effect differs by model, and it also touches Arabic and other languages written in non-Latin scripts. The practical consequence: the same message can cost more in Turkish, and a Turkish document fills the context window sooner.

Don’t assume a ratio. Take a sample of real Turkish and English cases, run them through each candidate model and read the token counts from the usage the provider reports. Our note on Turkish and AI covers the other design questions.

Ask: what are the measured tokens per case for our Turkish text, on each model we are considering?

How should a buyer use this?

Budget per finished task, not per token. A task is the unit the business cares about: one invoice drafted, one ticket answered, one contract checked. Its cost is the tokens of every call it needs, retries included, plus review time and infrastructure. Measure it in a pilot on real cases, in each language, and compare models on cost per task at the quality you need; our note on forecasting LLM running costs shows the sheet.

Nobody buys tokens. You buy finished tasks, and tokens are one line of their cost.

Two of our services pick this up: data and AI foundations compares models on your own examples, and AI reliability keeps cost per task under a ceiling after launch.

Ask an assistant about this note

EconomicsCostsModels

veridive

Field notes are written and reviewed by veridive. How we write them

Questions

Questions about this note

What is a token in AI?

A token is the unit of text a language model reads and writes: often a short word, sometimes part of a longer word, a punctuation mark or a space. Models process, limit and bill text in tokens, so the number of tokens in a request and its answer determines both what it costs and whether it fits in the model’s context window.

Does Turkish text use more tokens than English?

Often, yes, but it depends on the model. Turkish builds meaning with suffixes, so one word form can carry what English says in several words, and many tokenizers split those forms into several pieces. The same message can then cost more and fill the context window sooner. Measure it on your own Turkish text with each model you are considering.