Agents don’t just talk. They read.
Imagine a coding agent that emits a 12-token search command and gets back 20,000 tokens of logs. Before it writes its next short answer, the model must prefill those logs: process the incoming tokens and prepare memory for future attention. Repeat over many turns and the input dwarfs the generated output.
5.02×
Fewer estimated prefill FLOPs than Hybrid SWA at 1M tokens; 2.92× fewer than HySparse. Not an end-to-end latency claim.
2.69 GB
FP8 KV cache at 1M tokens versus 12.09 GB for Hybrid SWA and 6.72 GB for HySparse.
+19.81 points
Mean post-training RULER-v2 score over HySparse across evaluated lengths; not a 19.81% relative gain.
Figures 3–4 and §4.2 in the original PDF. Comparisons use 80B-A3B MoE configurations trained with the same data and schedules, but the attention-head and head-dimension configurations differ (Table 1).
What does the cache actually remember?
In a transformer attention layer, each token’s hidden vector is projected to a query Q, key K and value V. A query scores old keys, normalizes the scores, then mixes their values. The KV cache stores past K and V so they do not need recomputation when another token arrives.
Full attention: every query can read every previous token. For N input tokens, the attention interactions grow roughly as N² per full layer. Sliding-window attention (SWA): a query can read only the newest W tokens; it is cheaper but may miss an old fact. Sparse attention: read k selected older tokens and a recent window; its selected attention work is approximately N(k+W), though selection, full layers, projections and feed-forward/MoE computation remain.
KV cache ≠ model weights
Weights are the trained rules. The KV cache is per-request memory of previous tokens. A long conversation increases cache storage even when model weights do not change.
Prefill ≠ decode
Prefill ingests an existing prompt and prepares cache entries. Decode generates new tokens one at a time. Here, the cache-building pass can skip the cross-decoder, but generating outputs still needs the cross-decoder.
Two decoders, two levels of sharing.
The evaluated model has 49 transformer layers. Figure 2 depicts a front self-decoder (first 25 layers) mixing SWA and occasional full attention, and a back cross-decoder mixing full attention and sparse attention. The two kinds of sharing solve different problems.
Where do back-half keys come from?
From an earlier self-decoder full-attention layer’s input hidden states, through the target cross-decoder full layer’s own K/V projections. Each target layer has distinct projected K/V; it does not simply alias the same K/V arrays. Its queries still come from its current cross-decoder hidden states (§3.2).
Where do sparse-layer keys come from?
From the full layer within their hybrid block. Sparse followers reuse that layer’s cached K/V and the positions selected by its exact full-attention scores, instead of maintaining independent long-lived caches (§3.3–3.4).
Adapted from Eq. 1, §3.2: H denotes the input hidden states of a paired full-attention layer. One source full layer may provide hidden states for several target full layers. Diagram is conceptual; see Figure 2 of the PDF for the exact architecture.
Precise retrieval without a separate indexer.
In HySparse, the full layer scores attention and chooses blocks of 64 tokens; selecting one important token also buys 63 neighbors. HySparse2 chooses individual tokens instead. Its full-attention layers serve as trained “oracle” indexers for the following sparse layers: no separately distilled indexer module is required (§§2.3, 3.4).
The 128 newest tokens are forced into selection, even when their scores are low. This replaces HySparse’s separate gated SWA branch. Why is that important? A separate local branch would need keys projected from cross-decoder states, which depend on prior cross-decoder states; even a short local window produces a cascading dependence on an increasingly long suffix. Reusing only bridged/full-layer KV removes that dependency, so new input can finish its cache-building pass after the self-decoder (§3.3).
In the reported configuration, HySparse2 uses MQA (64 Q heads / 1 KV head) rather than the other two baselines’ GQA (64 Q heads / 4 KV heads); this contributes to its smaller cache and faster token-sparse kernels. Full and sparse layers use NoPE, while self-decoder SWA layers use partial RoPE. These are meaningful architectural differences, not controlled single-variable changes (Table 1, §4.1).
What actually improved?
Experiments compare 80B-A3B MoE designs, pretrained on roughly 500B tokens at 32k and lightly post-trained on roughly 100B additional tokens with agent data and context extended to 256k. The following are reported by the authors, not reproduced here.
| Measurement | Hybrid SWA | HySparse | HySparse2 |
|---|---|---|---|
| KV cache at 1M tokens, FP8 | 12.09 GB | 6.72 GB | 2.69 GB |
| Relative prefill FLOPs at 1M (normalize HySparse2 = 1) | 5.02× | 2.92× | 1× |
| RULER-v2 at 256k, post-training | 35.74 | 32.61 | 58.45 |
| RULER, pretraining (Table 2) | 88.71 | 84.89 | 90.77 |
| NoLiMa, pretraining (Table 2) | 30.13 | 40.27 | 49.76 |
Measured/estimated under the paper’s configuration, §4.2, Figure 4. Bars visualize reported numbers, not a browser simulation.
Ablations: what does each design decision buy?
Token vs. block
At equal 1,024-token global budget and ≤32k context, token selection improves RULER-v2 49.56 → 56.13 and two-needle MRCR-v2 12.94 → 21.08 (Table 3). NoLiMa moves 40.27 → 38.43, so it is not a universal win.
Forced vs. gated local window
Forced local selection enables the clean early exit, but compared with a separate gated SWA branch, pretraining MRCR-v2 drops 27.66 → 22.67 and GSM8K 64.52 → 59.44 in this ablation; RULER-v2 improves 53.66 → 55.98 (Table 4).
Bridging tested on separate 290B-A8B models gives RULER 96.32 without vs 96.01 with bridging; DROP falls 71.37 → 68.17, while LongPPL improves 3.6053 → 3.4202 (Table 5). This supports a quality/efficiency trade-off rather than proving zero cost to quality.
Spend a tiny attention budget yourself.
This illustrative toy models a 16-token history in four 4-token blocks. Scores are fabricated, not generated by HySparse2. Choose token-level or block-level selection and adjust the total global token budget. The newest four tokens form an additional forced window; green = globally selected, gold = forced-local, pink outline = missed relevant evidence.
Scenario: a useful fact buried in a tool log
For the block mode, each selected 4-token block costs four units of global budget. The demo excludes causal masks, softmax, per-query variation, learned score generation and exact model computation; forced-local tokens do not consume the global budget in this display.
Why this helps
If important evidence lies in multiple distant blocks, block selection spends capacity on irrelevant neighbors; token selection can buy individual facts instead. But if all useful tokens are adjacent, choosing the block might be just as good. A true system uses scores produced by full attention, and those scores may themselves be wrong.
Twenty lines of algorithmic intuition.
This pseudocode captures data flow, not an executable reproduction: real training must make bridged projections, exact-score selection, causal masking, attention kernels and model quality work jointly.
# Toy abstraction; do not treat this as production inference code.
def prefill(new_tokens):
self_states = run_self_decoder(new_tokens) # SWA + rare full attention
for full_layer_j, source_i in bridge_pairs:
H = self_states.input_to_full_layer[source_i]
cache[full_layer_j].append(
K=full_layer_j.Wk(H), V=full_layer_j.Wv(H)
)
# No cross-decoder forward pass needed to build its KV caches.
def decode_one_token(token):
h = run_self_decoder(token)
for block in cross_decoder_blocks:
q = block.full_layer.Wq(h)
scores = causal_scores(q, cache[block.full_layer].K)
recent = newest_visible_positions(window=128)
old = visible_positions_excluding(recent)
selected = recent | top_k(scores[old], k=1024)
h = block.full_layer(h, cache[block.full_layer])
for sparse_layer in block.followers:
h = sparse_layer.attend(h, cache[block.full_layer], selected)
return output_head(h)The powerful idea — and its boundaries.
What this establishes
On the reported matched-data training setup, sharing KV across decoder halves and hybrid blocks, plus finer token retrieval, jointly gives lower modeled prefill FLOPs and smaller FP8 KV cache while improving many long-context results.
What it does not establish
The paper reports FLOPs and cache estimates, not an across-hardware wall-clock speedup or overall serving cost. The 1M comparison is a cost analysis; the shown post-training retrieval evaluation reaches 256k, not 1M. AgentPPL is internal, limiting independent verification.
Also keep attribution honest: HySparse2’s MQA, head dimensions, local-window design and token selection change together relative to baselines. Ablations isolate some components but reveal trade-offs. A full-attention layer remains computationally expensive, even if it is rare; extreme-length accuracy and engineering depend on selection and kernels (§§3.4, 4).
Primary source: Jianyu Wei, Yizhao Gao et al., HySparse2: Hybrid Sparse Attention with Two-Level KV Sharing, arXiv:2609.26368v1, 22 September 2026. All empirical claims above are attributed to that PDF; the interactive lab and pseudocode are teaching illustrations.