An AI server specification is incomplete if the memory discussion ends at HBM capacity. Where a long conversation’s context is stored, and how quickly it can be retrieved, also belong in the design. Recent material from SK hynix makes that question tangible.
First, what does the cache do?
Imagine asking a follow-up question about a long manual. A KV cache keeps numerical states calculated from earlier tokens so that work can be reused. The relevant context must still be available, and reuse conditions must be met. This has a different role from product memory that carries your preferences into another conversation. With that distinction in place, the storage question in this article becomes clearer: which memory tier should hold reusable computation?
From components to tiers
The company’s September 28 report on TSMC OIP described exhibits spanning HBM, server DRAM and enterprise SSDs. The event itself took place on September 23. The presentation covered several tiers, from memory close to the accelerator to storage. The event date and publication date should remain distinct.
Where should context live?
Its September 17 AI Infra Summit report provided a more specific example. SK hynix described HBF as a proposed tier between HBM and SSD and exhibited a structural model. It described SALT-KV as dividing cached computation results by context and placing them across HBM, DRAM and SSD according to reuse value and storage cost. A displayed model and commercial supply are different stages.
Connections have a cost
An October 2 guest article by ETH Zurich professor Onur Mutlu, published in the SK hynix newsroom, also argues for reducing data movement. It notes that separating and connecting resources can raise latency and energy costs if communication becomes excessive. A disclaimer identifies the views as the author’s, not necessarily the company’s position.
V’s analysis · A broader purchasing question
V’s analysis: the purchasing proposition from this Korean memory maker is expanding from chip capacity to system-level data placement. That matters because software choices need to be tested alongside component choices. If adding storage makes responses slower, capacity alone is an incomplete measure of success.
Test short requests and long conversations separately
A useful trial would separate two kinds of requests: many users asking short questions at once, and long conversations that later revisit earlier material. With the model and required answer quality held constant, record time to first response, subsequent output speed, total power consumption and the equipment configuration. This is V’s proposed evaluation design.
In particular, separate a first request with an empty cache from a repeat request whose context remains available. A single average hides their differences. Recording the wait introduced by moving cache data between tiers, alongside any gain in request capacity, makes the trade-off more concrete. Fix the conversation lengths and concurrent load before deciding how much HBM to buy.
This analysis reviews official exhibition reports and a guest essay. Commercial availability, pricing and performance in a customer environment were not independently verified.