ELAI S.r.l.

Why does text already read keep using memory in an AI assistant?

A guided derivation of KV-cache memory: context, concurrent requests, shared heads and capacity limits, with reproducible counts and explicit assumptions.

Why does text already read keep using memory in an AI assistant?

The model fits in memory, but four conversations do not

A business assistant answers a short question correctly. Then four people send long documents, and available memory runs out before the answers finish. The model file has not changed. What is growing? One important component is context memory: during generation, the model retains representations of processed text for reuse. We want to understand their cost and why the number of conversations matters alongside their length.

Abstract. We derive key-value cache cost for a causal decoder with dense attention from data shapes. A synthetic example with 32 layers and eight KV heads needs 4 GiB for four 8,192-token sequences. We compare head sharing, common prefixes, quantization and a capacity budget including weights and workspace. All numbers are counts, not allocated-memory or latency measurements. They explain the constraint but do not certify that any particular GPU or runtime can execute the workload.

What is retained, and why

An autoregressive model generates one token at a time using previous context. In attention, the current token produces a query, a representation used to compare against the past. Processed tokens provide keys for that comparison and values to combine. These are not sentences copied into an address book, but numerical vectors inside the model. Different attention heads perform parallel comparisons with different representations.

In the causal decoder considered here, earlier tokens cannot see future tokens. With the model and position conventions fixed, their keys and values can be reused in later steps. The KV cache retains them for each layer. It avoids recomputing them but does not eliminate comparisons between the new query and the past. Previously read text therefore still has a cost: some saved recomputation is replaced by memory storage.

Count numbers before counting gigabytes

Let L be the number of cached layers, B concurrent sequences, T retained tokens per sequence, H_KV key-value heads, d each head’s dimension and s bytes per number. Assume equal length T and identical configuration across layers, without prefix sharing or compression. One layer’s K tensor contains B×T×H_KV×d elements, and V contains the same number.

M_KV = 2 × L × B × T × H_KV × d × s [byte]

The factor 2 counts K and V; L repeats the count across layers. Every other factor describes a tensor dimension. This is cache-data memory, not total process memory. It excludes learned weights, temporary activations, vocabulary logits, memory-manager metadata and runtime reservations. A useful formula must also say what it leaves out.

The numerical case: four 8,192-token documents

Choose L=32, B=4, T=8,192, H_KV=8, d=128 and s=2 bytes per element. These are illustrative parameters, not specifications of a named model. Substitution gives 4,294,967,296 bytes, exactly 4 GiB. One GiB is 2³⁰ bytes, unlike a decimal GB of one billion bytes. The same calculation gives 1 GiB for one sequence.

Each extra token costs 2×32×8×128×2=131,072 bytes per sequence, or 128 KiB. If all four conversations grow by one token, the total rises by 512 KiB. T covers what the runtime retains: the prompt plus generated tokens still kept, not just the initial input. A budget based only on the question can be exceeded while a long answer is completed.

Sequences BTokens TKV headsCache GiB
1819281
4819284
481923216
41638488

Why KV heads matter, not only query heads

Not every query head necessarily owns separate keys and values. In conventional multi-head attention, MHA, the counts match. In grouped-query attention, GQA, several query heads share a KV head. Our comparison keeps 32 query heads while reducing KV heads from 32 to 8: cache falls from 16 to 4 GiB. This follows directly from the number of distinct vectors retained. It does not mean three quarters of an arbitrary model’s tensors can be deleted without changing behaviour.

Ainslie and colleagues published GQA at EMNLP 2023. We read methods, experiments and limitations in arXiv v3, dated 23 December 2023: encoder-decoder T5 experiments, TPUv4 timings and adaptation after conversion. We do not transfer those results to a GPU or decoder-only model, or describe GQA as new in 2026; our calculation concerns memory.

Theoretical count with 32 layers, four sequences, head dimension 128 and two bytes per element. Lines compare only 8 and 32 KV heads. Weights and workspace are excluded; these are not measured memory curves.
Theoretical count with 32 layers, four sequences, head dimension 128 and two bytes per element. Lines compare only 8 and 32 KV heads. Weights and workspace are excluded; these are not measured memory curves.

The slope answers how much extending context costs. Doubling T doubles this memory component, as does doubling B. We are not plotting a T×T attention-score matrix. Avoiding materialization of that matrix and retaining a KV cache are separate issues: efficient attention does not automatically make stored history free.

A complete budget changes the admissible request count

Add a hypothetical six-billion-parameter model, all stored at two bytes. Weights alone occupy 12 billion bytes, about 11.1759 GiB. Also assume a 2 GiB reserve for the rest of execution and 16 GiB available capacity. These are chosen numbers, not measurements. Four sequences total 11.1759+2+4≈17.1759 GiB, exceeding the budget. Three total about 16.1759, still too much; two total 15.1759.

Saying two sequences “fit the arithmetic” does not promise they run. Actual reserve depends on runtime, kernels, prompt processing, intermediate shapes, fragmentation and other allocations. The count is a planning condition to compare with measured peak usage. We ran no model, selected no GPU and measured no peak; JSON explicitly calls the result arithmeticFits.

Are four copies of the same prefix always necessary?

Four conversations may start with identical instructions and the same document. If a runtime supports safe sharing of an immutable prefix, count that prefix once and suffixes separately. Let P=4,096 common tokens out of T=8,192 total per sequence. Without sharing we store BT=32,768 positions; ideal sharing gives:

T_stored = P + B(T − P) T_stored = 4096 + 4×4096 = 20480 M_shared = 2 L H_KV d s T_stored = 2.5 GiB

The theoretical saving is 1.5 GiB from the original 4. Identical text on screen is insufficient: tokens, weights and adapters, positions and attention conditions determining the cache must match. Changing earlier instructions may also change later representations. Suffixes must separate as responses diverge. We implemented no such runtime; the count gives the ideal benefit for this case before metadata and allocation blocks.

Quantizing the cache does not halve all memory

Initially each KV number uses two bytes. Assuming an eight-bit representation reduces data payload alone from 4 to 2 GiB. Quantization may also need scales and metadata. Our purely accounting scheme adds one two-byte scale per 64 values and no zero point: overhead is 2/64 bytes per value. Total becomes 2×(1+2/64)=2.0625 GiB, not exactly 2.

This is a storage formula, not an evaluated quantization algorithm. We quantized no real tensors, measured no answer errors and demonstrated no speedups. Temporary dequantized copies may affect peak memory. Halving KV also leaves weights and reserve unchanged in our comparison, so total percentage savings are smaller. A concrete method needs numerical and task-level validation.

What changes in the runtime, and what the count does not imply

Hugging Face documents dynamic, static and sliding-window caches, mask handling and a generation loop, which we read. Dynamic caches grow; static caches may reserve capacity early; windows bound retained history. We did not execute those components. T must reflect actual allocated length and the layers involved.

A shorter window is not free compression of identical information: it removes direct access to earlier tokens according to architecture and policy. Moving cache outside accelerator memory changes location, not data volume, and may add transfers. Likewise, four times less KV data does not guarantee four times lower latency: projections, attention computation, bandwidth, synchronization and implementation all contribute to elapsed time.

The answer and the protocol needed for real measurements

Previously read text uses memory because the model keeps representations useful to later tokens, per layer and sequence. Our example reaches 4 GiB before weights; adding the rest makes four requests exceed the hypothetical 16 GiB budget. Operational planning must consider length, concurrency and cache representation together, not just downloaded model size. This is a design consideration for business assistants, not a measured EL-AI infrastructure report.

A real test would fix checkpoint, tokenizer, weight and KV precision, GPU, runtime version, attention backend, request count, input/output lengths and allocation strategy. It would separately measure allocated and reserved memory, prompt and generation peaks, latency and throughput. That protocol is proposed, not executed. Our package instead contains deterministic integer counts, no seed, GiB assertions and figure data: each configuration needs O(1) arithmetic operations without simulating tensors.

Primary sources and code

Hugging Face Transformers — How caching works.

Ainslie, Lee-Thorp, de Jong, Zemlyanskiy, Lebrón, Sanghai — GQA, arXiv v3, 23 December 2023.

GQA — EMNLP 2023, ACL Anthology.

GiB = 2**30
layers, kv_heads, head_dim, bytes_per_value = 32, 8, 128, 2
def cache_bytes(batch, tokens):
    return 2 * layers * batch * tokens * kv_heads * head_dim * bytes_per_value
for batch in (1, 2, 3, 4):
    kv = cache_bytes(batch, 8192)
    total = 6_000_000_000 * 2 + 2 * GiB + kv
    print(batch, kv/GiB, total/GiB, total <= 16*GiB)

Code, data, and instructions · JSON. Educational calculations executed with Python 3.14.0; figures with Matplotlib 3.11.2. AI-assisted analysis, without claiming peer review or human review. Original illustrative ImageGen cover: it does not document EL-AI people, premises, or installations. Sources accessed 3 October 2026.