A token is a word piece into which a language model splits text, and the context window is the maximum number of tokens it can consider at once in one request: your instruction, the history so far, attached documents and the answer. When something no longer fits, or gets lost in the volume, it looks as if the AI forgets. Billing is usually by tokens.
What is a token?
A language model reads neither letters nor whole words but tokens. Common words are often one token; rare or long words are split into several parts. According to OpenAI, the rule of thumb for English is about 100 tokens for 75 words. In German, because of long compounds, fewer words usually fit into 100 tokens. By this rule a text of 3,000 English words would come to about 4,000 tokens.
What the splitting looks like is shown by OpenAI's tokenizer. How language models work overall is explained in What is an LLM?.
What is the context window, and why does the AI ‘forget’?
The context window is the amount of text the model sees at once when it answers. Everything counts: your instruction, the chat history so far, uploaded documents and the answer itself. The model has no permanent storage. With every message, the whole history is sent again.
When the window is full, older parts have to be cut or left out, and the AI ‘forgets’ them. But even within the window there is an effect: research shows that models use information at the start and end of long texts better than information in the middle (Liu et al., "Lost in the Middle").
How do the costs arise?
Providers often bill by tokens, separated into input (what the model reads) and output (what it writes). Output tokens are often more expensive. The longer the history and the larger the documents, the more tokens are processed with every answer, and the higher the cost and the waiting time. Concrete prices change often; always check your provider's current price page.
| Term | Meaning | Practical consequence |
|---|---|---|
| Token | Word piece, billing unit | Long texts cost more |
| Context window | Maximum tokens per request | Anything beyond is lost or cut |
| Input | Instruction, history, documents | Grows with every chat message |
| Output | The model's answer | Often priced higher |
| Memory feature | Notes stored between chats | Only active if switched on, and relevant to data protection |
How do you save tokens and avoid memory gaps?
Start a new chat for every new topic instead of continuing a long history. Summarise interim results and supply the summary. Put important points at the start or end of the instruction and repeat key requirements.
Supply only the relevant parts of a document, not the whole folder. For large bodies of knowledge, a method that picks out the matching passages is worthwhile, as described in What is RAG?. A bigger context window is not automatically better: it costs more, takes longer and does not protect against errors, as AI hallucinations shows.
- New chat started for new topics
- Long histories condensed into a summary
- Important requirements placed at the start or end of the instruction
- Only relevant document parts supplied
- Token use of a typical task measured
- Memory features and privacy settings checked
Conclusion: less, but more targeted context
If you understand tokens and context windows, you get better answers at lower cost. Give the model what it needs and keep histories short. If you want to build AI into processes, see our AI automation service or describe your task.




