← All models

MODEL EXPLORER / MODERN LLMS

Eight queries, fewer key/value heads

Which attention heads can share a key/value cache?

Curated by TensorVizMHA / MQA / GQA · 4-step tour
Static architecture
Which attention heads can share a key/value cache?
Opening the interactive graph…
Architecture snapshot · No Python requiredDownload graph ↓

RECORDED NUMERICAL EXAMPLE

Keep queries, share keys and values.

Head connections from the recorded configuration; cache bytes calculated as 2 × batch × tokens × KV heads × head width × bytes per value. This CPU reference does not measure decoding speed.

All eight query heads remain distinct. Lines show which K/V head each query consults. K/V weights are group means of the same MHA weights, so outputs need not match. Cache size is an analytical estimate for one layer, batch 1, head width 2 and float32, excluding temporary expansion and framework overhead.

Query headsShared key/value headsQ 0Q 1Q 2Q 3Q 4Q 5Q 6Q 7KV 0KV 1KV 2KV 3KV 4KV 5KV 6KV 7
Lines show nonzero connections. Exact values are available below.
Connection weights
Query headsShared key/value headsWeight
Q 0KV 01
Q 1KV 11
Q 2KV 21
Q 3KV 31
Q 4KV 41
Q 5KV 51
Q 6KV 61
Q 7KV 71
Last token output · four-token forward
-0.2111
0.1168
0.5017
-0.1425
0.186
-0.2192
-0.374
-0.00523
-0.004842
0.06731
-0.06666
-0.01772
0.1159
-0.06247
0.1221
-0.1313

Bars share one scale within this example. Values are rounded to four significant digits.

Analytical K + V cache · bytes
16384
KV heads
8
Query heads per KV head
1
Verified by the local recipe
  • MHA with duplicated K/V weights matches grouped attention for 8, 2 and 1 KV heads
  • Future token changes cannot affect earlier causal outputs
  • Compact per-group K/V calculation reproduces the final query output
  • Analytical KV bytes scale with the number of KV heads

Download the source and run.py to reproduce these checks. Choosing an example here replays recorded values.

ABOUT THIS EXAMPLE

Compare MHA, GQA and MQA while keeping eight query heads. Follow the head mapping and calculate compact KV-cache storage.

Component · 2023

MHA / MQA / GQA

Eight query heads of width two with 8, 2 or 1 KV heads. Transparent repeat_interleave CPU implementation; no optimized attention kernel.

Source, capture & limitations +

Untrained width-16 comparison. Averaging K/V weights is an illustrative conversion without the paper's uptraining. Cache estimates exclude temporary repeated tensors, attention activations and allocator overhead; no quality or latency claim.

Original TensorViz teaching example. PyTorch provides the underlying operators.

No separate redistribution license has been declared for these project examples.

Content revision
803295b72f9f4cbd
Source SHA-256
2d7dbb8bead241d951d07146b0a15f05cd19c7ae0d0bb5a90bf1ad1dd974aa16
Captured
2026-09-17 · Python 3.13.13 / Torch 2.7.1

Reduced causal MHA/GQA/MQA CPU forwards, independent per-head cache calculation, and duplicated-weight equivalence. Analytical cache estimates are not runtime measurements.

Read the model manifest ↗