An assistant answers a second question about your document quickly. Tomorrow, it remembers that you prefer short answers. These can look like one kind of memory. They solve different problems.
What fits this request
A context window is the token budget for a model request. Tokens are pieces of the model’s input or output; they are not necessarily whole words. Input, output and, on some models, reasoning share this budget. A conversation stored by an app is not automatically all inside the current window.
What the cache retains
KV cache is saved numerical work inside attention-based inference: keys and values calculated from earlier tokens. Reusing them avoids calculating those same keys and values again as the answer grows. New tokens still require computation. This is a mechanism for efficient generation, not a profile of your preferences.
Here is an illustrative sequence, not a measured experiment or a promise about a particular chat app. Assume an ordinary causal-transformer service, unchanged model settings, and a document that fits its context budget.
First request
First request: you provide a long equipment manual and ask, “How do I restart it?” With no reusable cache, the service processes the input before generating an answer. It can keep the resulting KV states for reuse.
Second request
Second request: “Which warnings should I check before restarting?” The app must make the relevant earlier material available, for example through conversation state or resent history. If eligible cached state matches the unchanged opening of the request, prefix caching can skip processing that part again. vLLM documents this benefit for repeated document questions and continuing conversations; it mainly saves input processing, not the work of generating a new answer.
An edit near the beginning
Third request: you correct an early sentence in the manual and ask again. With exact-prefix caching, only the unchanged opening before the edit is a candidate for reuse. Actual reuse extends only to an eligible cache boundary within that opening. Keeping most words identical elsewhere does not restore the same prefix. OpenAI’s API also applies cache eligibility and boundary rules: identical visible text alone does not guarantee a hit.
“Prompt cache” therefore names a service-level reuse feature; “KV cache” names the underlying computed state in these implementations. The concepts connect, but neither term specifies every provider’s retention, routing or conversation handling.
Memory tomorrow
Fourth request, tomorrow: you start a new chat. A product memory feature may retrieve your preference for short answers even if yesterday’s computation cache is unavailable. ChatGPT documents personalization from available past context when Memory is enabled, with availability varying by account and settings. It does not promise to retain every detail. Other products implement continuity differently.
Select information, or reuse computation?
Which calculations can be reused?
Eligible cached states from an unchanged prefix
Reuse computationWhich past context should return?
Available preferences and context, subject to settings
Select contextNext: where to keep it
The practical question is two questions: what information is available for this answer, and which computations can be reused? That distinction is the doorway to our SK hynix article: once cached numerical work exists, where should a server put it?