Tokens, context and hallucinations: why the model gets it wrong

Basics6 min read

When an AI model behaves strangely, one of three phenomena is usually behind it: a token limit, limited context or a hallucination.

Token: the unit the model counts in

The model does not read letters or whole words, but tokens, which are pieces of text. A token can be a whole word, part of one or a single character. In English it is roughly three quarters of a word, while in many other languages there are more tokens per word. Length limits and service costs are usually counted in tokens.

Context window: working memory

The model sees only what fits in its context window: your instructions, earlier answers and pasted files. When a conversation is very long, the oldest parts may be cut off or lose weight. The result: the model "forgets" what was agreed at the start, repeats mistakes or changes the style of the code.

Hallucination: confident untruth

The model generates an answer that sounds good even when there is nothing behind it. In programming, typical examples are a function that does not exist in a given library, a nonexistent parameter, an invented package or syntax from a different version of a tool. This happens because the model is meant to produce plausible text, not to check facts.

How to protect yourself:

Why a good prompt helps

Clear context leaves less room for guessing, and less guessing means fewer hallucinations. We cover how to do this in the guide on writing prompts for code.

Rule of thumb: the more important the information, the less you should rely on the model's memory and the more on checking it at the source.