Journal
Why Your AI Forgets: Context Windows Explained
The number in the spec sheet that nobody explains. Here is what it measures, and what happens the moment your conversation outgrows it.
- Published
- Last updated
A context window is the amount of text a language model can hold in view at once while producing its next reply. It covers everything: the character definition, the system instructions, the whole visible conversation and the reply being written. It is a budget, and everything competes for it.
Measured in tokens, not words
Models count in tokens, which are chunks of text roughly between a character and a word. Common English words are usually one token; unusual names, punctuation and other writing systems cost more. As a rough planning figure, English prose runs somewhere near three quarters of a word per token, which is close enough for estimating and wrong enough that you should not build anything on it.
What happens when you run out
- Oldest messages are dropped from view, usually silently and usually first.
- Some products summarise what they drop, which preserves the gist and loses the wording.
- The character definition may be pinned so it survives, which is why personality outlasts specifics.
- Nothing announces any of this, so from your side it reads as the character forgetting.
Bigger is not automatically better
A large window costs more to run per message, which is why generous limits and generous free tiers rarely appear together. Models also attend unevenly across a long window: material at the start and the end tends to carry more weight than material buried in the middle. A very long conversation can technically fit and still feel thin.
Practical habits
If something matters, restate it. Start a fresh conversation for a genuinely new thread rather than dragging thirty thousand words behind you. And read a window figure as a ceiling on what is possible, not a promise about what will be remembered.