gptagency.io

Token

A token is the smallest processing unit into which an AI language model breaks down text. A token corresponds roughly to a short word, a word part or a punctuation mark. Models like ChatGPT or Claude don't read, think and answer in letters or whole sentences but in tokens. They form the computing unit with which every AI output is generated and billed.

Why tokens matter for your visibility

Tokens are the currency of every AI interaction. Every question a user poses to an AI assistant, and every answer, is measured in tokens. This has direct consequences for your visibility: models have a limited context window, meaning a maximum amount of tokens they can process at once. If your content doesn't fit into this budget, it gets shortened or not considered at all. Whoever wants to understand why an AI names or passes over their own brand has to know that the machine thinks in tokens. Concise, clearly structured texts cost fewer tokens and are more easily taken into an answer. Bloated phrasing wastes budget that would otherwise benefit your brand.

How tokenization works

Before a model processes text, a so-called tokenizer breaks the string into tokens. Frequent words often become a single token, while rare or long words are split into several parts. The German word Suchmaschinenoptimierung can thus consist of several tokens, while und is a single one. As a rough rule of thumb: for German text, one token corresponds to about 0.7 to 0.8 words, so noticeably less efficient than in English. Each token is then translated into a sequence of numbers, a so-called vector embedding, with which the neural network computes. After that the model predicts, token by token, the most likely continuation. That way an answer emerges building block by building block, not as a finished sentence all at once.

Common mistakes and misconceptions

The biggest error is equating tokens with words. A token is usually smaller than a word, especially in German with its long compounds. A second mistake concerns the context window: many believe a model could process arbitrarily long texts in one go. In fact a fixed token limit applies to input and output combined. If it's exceeded, the model forgets earlier content or breaks off. For AI visibility this means: long, nested pages with a lot of ballast compete over scarce token budget. Whoever places the core statement high up and in clear language increases the chance that the relevant tokens make it into the answer and the brand gets cited.

Relation to AI recommendations and GEO

In Generative Engine Optimization, thinking in tokens is the bridge between your content and the AI answer. An assistant draws on relevant passages, converts them into tokens and forms a recommendation from them. Short, fact-dense paragraphs with clear entities like place, service or price can be tokenized efficiently and cited more easily. Costs play a part too: providers bill their models per thousand tokens, which is why systems tend to prefer economical, easily compressible sources. If you write your content so that it delivers a complete answer in few tokens, you work in the machine's interest and boost your chance of a brand mention in AI search.

Example

Imagine a small tax firm that explains on its website that it advises startup founders in Freiburg. If a user asks an AI assistant for a tax advisor for founders in Freiburg, the model breaks both the question and the found website texts into tokens. If the core statement stands compactly and with clear terms like startup founding, Freiburg and initial consultation in the text, it fits easily into the token budget of the answer. A rambling running text, by contrast, costs many tokens without directly answering the question, and is more likely to be passed over. Compact clarity wins.

Common questions

How many tokens does a German sentence have?

As a rule of thumb, one token corresponds to about 0.7 to 0.8 words. A normal German sentence with 15 words therefore has roughly 20 tokens. Long compounds and rare technical terms increase the token count, because they are split into several parts.

Why are AI models billed per token?

Because the token is the actual computing unit. Every token processed and generated costs computing power. Providers like OpenAI or Anthropic therefore bill per thousand tokens, separated by input and output, rather than per word or request.

Related terms