If an AI protects more of the original question, do its older working entries become safer too? With a limited allocation, those goals can collide. Reserving space for one kind of information changes the competition for everything left.

Random Attention protects the input and newest entries, randomly retaining the rest. Appendix D supplies the formula.

Give the question more of 1,024 slots

In the September 4 code, CLI 1024 means total entries after compaction: 960 plus 64 newest. September 28 v2 defines K differently.

V turned that allocation into a hypothetical ledger. Reserve 512 slots for input and 64 for the newest entries. That leaves 1,024 − 512 − 64 = 448 randomly retained slots. A slot here holds the key–value entry for one token position in a particular layer and KV head. It is not a character or a complete fact.

Consider the repeating phase after the ledger has filled. Another 64 arrivals bring it to 1,088 entries just before compaction. Exclude the newest 64 and the 512 protected input entries: 512 candidates remain. Keeping 448 gives any particular eligible entry a 448/512 = 7/8 chance of surviving this draw. “1,024” is not the maximum length at every instant.

For hypothetical protected allocations of 128, 512 and 768 slots, one old entry survives 16 eligible draws with probabilities 30.55%, 11.81% and 1.00%.
Blue P: protected input. Ochre R: randomly retained generated entries. Gray: newest 64. V’s hypothetical allocation calculation. Bars show 1,024 slots after compaction, not probabilities or bytes. Percentages track one entry in one layer and KV head, eligible in all 16 rounds. They are not accuracy rates.

A small risk gets repeated

Conditional on still being present, that entry faces the same 7/8 chance at the next draw. With fresh uniform draws, survival through 16 rounds in which it is eligible is (7/8) to the 16th power, about 11.81%. A discarded entry does not return in this calculation.

Keep the ledger size unchanged but reserve only 128 input slots: 832 random slots remain, giving about 30.55% survival over 16 eligible rounds. Reserve 768 instead: 192 random slots remain, and survival falls to about 1.00%. These three input allocations are chosen illustrations, not observed benchmark cases. Older generated entries bear the cost of the larger protected allocation.

Sixteen eligible rounds is not an entry’s age or another name for 1,024 tokens since its birth. Time spent in the newest buffer, or waiting for the ledger to fill initially, does not count as exposure here. We observe whether it remains immediately after each eviction draw.

This calculation answers a question about one entry. Keeping all the entries needed for a sentence requires a different calculation. Slots compete within each draw, so multiplying their survival probabilities as if they were independent is also wrong. Presence in a cache does not establish that a model can read the content and answer correctly.

V’s view: decide what must survive first

V would start by choosing what must be protected and which losses are acceptable. A task whose final answer depends on a value stated once, far earlier, deserves more than an average performance score.

In Table 3, once-stated passcode retrieval after 57 rounds was 0% for Random Attention versus 83.6% for R-KV.

That countercase does not make another method best for every task. V would test once-stated critical values, repeated values and long inputs separately in the intended workload. Under matched total allocation and protection, compare final correctness, critical-value retrieval and latency together. Repeated savings without unacceptable critical-value losses would support random selection. Otherwise, change the protected allocation or selection rule.

Change the numbers yourself

The only execution for this article was hypothetical allocation arithmetic using standard Python. No model, GPU or upstream experiment code was run, and no v2 result was replicated.

Download the calculator and CSV results. Run python random-attention-retention.py to print CSV. B is the total immediately after compaction, r the newest buffer, and P the protected input allocation; the examples require 0 ≤ P < B − r.

For the distinction between computation storage and conversation memory, continue with How is AI caching different from memory?.