Tokens, context and hallucinations: why the model gets it wrong

When an AI model behaves strangely, one of three phenomena is usually behind it: a token limit, limited context or a hallucination.
Token: the unit the model counts in
The model does not read letters or whole words, but tokens, which are pieces of text. A token can be a whole word, part of one or a single character. In English it is roughly three quarters of a word, while in many other languages there are more tokens per word. Length limits and service costs are usually counted in tokens.
Context window: working memory
The model sees only what fits in its context window: your instructions, earlier answers and pasted files. When a conversation is very long, the oldest parts may be cut off or lose weight. The result: the model "forgets" what was agreed at the start, repeats mistakes or changes the style of the code.
- Start a new conversation for a new task.
- Remind it of the most important constraints.
- Paste only the pieces of code that are needed.
- During long work, ask for a short summary of what has been agreed and start a new conversation from it.
Hallucination: confident untruth
The model generates an answer that sounds good even when there is nothing behind it. In programming, typical examples are a function that does not exist in a given library, a nonexistent parameter, an invented package or syntax from a different version of a tool. This happens because the model is meant to produce plausible text, not to check facts.
How to protect yourself:
- run the code instead of taking its word for it,
- check function and library names in the official documentation,
- ask the model to state the assumptions its answer rests on,
- for important conclusions, ask again in a new conversation or in a different tool.
Why a good prompt helps
Clear context leaves less room for guessing, and less guessing means fewer hallucinations. We cover how to do this in the guide on writing prompts for code.


