Paper field guide · 28 September 2026 · Efficiency

HySparse2: read a million tokens without paying every layer’s bill.

A visual, first-principles guide to hybrid sparse attention, two-level KV sharing, and why an agent can prefill through half the network yet decode through the whole network.

Read original PDF ↗DAIR.AI paper page ↗Wei, Gao et al. · Xiaomi LLM-Core · arXiv:2609.26368v1
01 / The big picture

Agents don’t just talk. They read.

Imagine a coding agent that emits a 12-token search command and gets back 20,000 tokens of logs. Before it writes its next short answer, the model must prefill those logs: process the incoming tokens and prepare memory for future attention. Repeat over many turns and the input dwarfs the generated output.

The paper’s idea: keep a few layers that see everything, let most layers read only high-value tokens plus recent tokens, share the resulting memory across layers, and build the cache for the back half from states produced by the front half. Source: §§1–3.
Less work

5.02×

Fewer estimated prefill FLOPs than Hybrid SWA at 1M tokens; 2.92× fewer than HySparse. Not an end-to-end latency claim.

Less memory

2.69 GB

FP8 KV cache at 1M tokens versus 12.09 GB for Hybrid SWA and 6.72 GB for HySparse.

Better retrieval

+19.81 points

Mean post-training RULER-v2 score over HySparse across evaluated lengths; not a 19.81% relative gain.

Figures 3–4 and §4.2 in the original PDF. Comparisons use 80B-A3B MoE configurations trained with the same data and schedules, but the attention-head and head-dimension configurations differ (Table 1).

02 / First principles

What does the cache actually remember?

In a transformer attention layer, each token’s hidden vector is projected to a query Q, key K and value V. A query scores old keys, normalizes the scores, then mixes their values. The KV cache stores past K and V so they do not need recomputation when another token arrives.

attention(q, K, V) = softmax(q Kᵀ / √d) V

Full attention: every query can read every previous token. For N input tokens, the attention interactions grow roughly as N² per full layer. Sliding-window attention (SWA): a query can read only the newest W tokens; it is cheaper but may miss an old fact. Sparse attention: read k selected older tokens and a recent window; its selected attention work is approximately N(k+W), though selection, full layers, projections and feed-forward/MoE computation remain.

KV cache ≠ model weights

Weights are the trained rules. The KV cache is per-request memory of previous tokens. A long conversation increases cache storage even when model weights do not change.

Prefill ≠ decode

Prefill ingests an existing prompt and prepares cache entries. Decode generates new tokens one at a time. Here, the cache-building pass can skip the cross-decoder, but generating outputs still needs the cross-decoder.

Important distinction: exiting prefill after the self-decoder does not mean the system can produce good final predictions without the cross-decoder. It builds the cross-decoder’s reusable keys and values; decoding subsequently runs the rest of the model (§3.5).
03 / Architecture

Two decoders, two levels of sharing.

The evaluated model has 49 transformer layers. Figure 2 depicts a front self-decoder (first 25 layers) mixing SWA and occasional full attention, and a back cross-decoder mixing full attention and sparse attention. The two kinds of sharing solve different problems.

Outer / KV Bridging

Where do back-half keys come from?

From an earlier self-decoder full-attention layer’s input hidden states, through the target cross-decoder full layer’s own K/V projections. Each target layer has distinct projected K/V; it does not simply alias the same K/V arrays. Its queries still come from its current cross-decoder hidden states (§3.2).

Inner / KV Reuse

Where do sparse-layer keys come from?

From the full layer within their hybrid block. Sparse followers reuse that layer’s cached K/V and the positions selected by its exact full-attention scores, instead of maintaining independent long-lived caches (§3.3–3.4).

Kⱼᶜʳᵒˢˢ = Wⱼᴷ Hᵢˢᵉˡᶠ ; Vⱼᶜʳᵒˢˢ = Wⱼⱽ Hᵢˢᵉˡᶠ ; Qⱼᶜʳᵒˢˢ = Wⱼᑫ Hⱼᶜʳᵒˢˢ

Adapted from Eq. 1, §3.2: H denotes the input hidden states of a paired full-attention layer. One source full layer may provide hidden states for several target full layers. Diagram is conceptual; see Figure 2 of the PDF for the exact architecture.

04 / How the trick works

Precise retrieval without a separate indexer.

In HySparse, the full layer scores attention and chooses blocks of 64 tokens; selecting one important token also buys 63 neighbors. HySparse2 chooses individual tokens instead. Its full-attention layers serve as trained “oracle” indexers for the following sparse layers: no separately distilled indexer module is required (§§2.3, 3.4).

The 128 newest tokens are forced into selection, even when their scores are low. This replaces HySparse’s separate gated SWA branch. Why is that important? A separate local branch would need keys projected from cross-decoder states, which depend on prior cross-decoder states; even a short local window produces a cascading dependence on an increasingly long suffix. Reusing only bridged/full-layer KV removes that dependency, so new input can finish its cache-building pass after the self-decoder (§3.3).

Concrete mental model: a folder has 1,000 pages. A full-attention “librarian” can inspect the whole folder and tag the important sentences. Five sparse readers per block borrow the librarian’s tags and the same stored page representations, while always retaining the newest notes. Each reader still asks its own question using its own query.

In the reported configuration, HySparse2 uses MQA (64 Q heads / 1 KV head) rather than the other two baselines’ GQA (64 Q heads / 4 KV heads); this contributes to its smaller cache and faster token-sparse kernels. Full and sparse layers use NoPE, while self-decoder SWA layers use partial RoPE. These are meaningful architectural differences, not controlled single-variable changes (Table 1, §4.1).

05 / Read the results carefully

What actually improved?

Experiments compare 80B-A3B MoE designs, pretrained on roughly 500B tokens at 32k and lightly post-trained on roughly 100B additional tokens with agent data and context extended to 256k. The following are reported by the authors, not reproduced here.

MeasurementHybrid SWAHySparseHySparse2
KV cache at 1M tokens, FP812.09 GB6.72 GB2.69 GB
Relative prefill FLOPs at 1M (normalize HySparse2 = 1)5.02×2.92×1×
RULER-v2 at 256k, post-training35.7432.6158.45
RULER, pretraining (Table 2)88.7184.8990.77
NoLiMa, pretraining (Table 2)30.1340.2749.76
At 1 million tokens · cache footprint

Measured/estimated under the paper’s configuration, §4.2, Figure 4. Bars visualize reported numbers, not a browser simulation.

Ablations: what does each design decision buy?

Token vs. block

At equal 1,024-token global budget and ≤32k context, token selection improves RULER-v2 49.56 → 56.13 and two-needle MRCR-v2 12.94 → 21.08 (Table 3). NoLiMa moves 40.27 → 38.43, so it is not a universal win.

Forced vs. gated local window

Forced local selection enables the clean early exit, but compared with a separate gated SWA branch, pretraining MRCR-v2 drops 27.66 → 22.67 and GSM8K 64.52 → 59.44 in this ablation; RULER-v2 improves 53.66 → 55.98 (Table 4).

Bridging tested on separate 290B-A8B models gives RULER 96.32 without vs 96.01 with bridging; DROP falls 71.37 → 68.17, while LongPPL improves 3.6053 → 3.4202 (Table 5). This supports a quality/efficiency trade-off rather than proving zero cost to quality.

06 / Interactive lab

Spend a tiny attention budget yourself.

This illustrative toy models a 16-token history in four 4-token blocks. Scores are fabricated, not generated by HySparse2. Choose token-level or block-level selection and adjust the total global token budget. The newest four tokens form an additional forced window; green = globally selected, gold = forced-local, pink outline = missed relevant evidence.

Scenario: a useful fact buried in a tool log

For the block mode, each selected 4-token block costs four units of global budget. The demo excludes causal masks, softmax, per-query variation, learned score generation and exact model computation; forced-local tokens do not consume the global budget in this display.

Why this helps

If important evidence lies in multiple distant blocks, block selection spends capacity on irrelevant neighbors; token selection can buy individual facts instead. But if all useful tokens are adjacent, choosing the block might be just as good. A true system uses scores produced by full attention, and those scores may themselves be wrong.

07 / Minimal concept prototype

Twenty lines of algorithmic intuition.

This pseudocode captures data flow, not an executable reproduction: real training must make bridged projections, exact-score selection, causal masking, attention kernels and model quality work jointly.

# Toy abstraction; do not treat this as production inference code.
def prefill(new_tokens):
    self_states = run_self_decoder(new_tokens)  # SWA + rare full attention
    for full_layer_j, source_i in bridge_pairs:
        H = self_states.input_to_full_layer[source_i]
        cache[full_layer_j].append(
            K=full_layer_j.Wk(H), V=full_layer_j.Wv(H)
        )
    # No cross-decoder forward pass needed to build its KV caches.

def decode_one_token(token):
    h = run_self_decoder(token)
    for block in cross_decoder_blocks:
        q = block.full_layer.Wq(h)
        scores = causal_scores(q, cache[block.full_layer].K)
        recent = newest_visible_positions(window=128)
        old = visible_positions_excluding(recent)
        selected = recent | top_k(scores[old], k=1024)
        h = block.full_layer(h, cache[block.full_layer])
        for sparse_layer in block.followers:
            h = sparse_layer.attend(h, cache[block.full_layer], selected)
    return output_head(h)
Timing nuance: this pseudocode shows the cache-sharing relationship, not the paper’s exact execution schedule. Each actual layer has its own queries and attention behavior; Eq. 1 gives the bridging projections and Figure 2 sketches index reuse.
08 / What to take away

The powerful idea — and its boundaries.

What this establishes

On the reported matched-data training setup, sharing KV across decoder halves and hybrid blocks, plus finer token retrieval, jointly gives lower modeled prefill FLOPs and smaller FP8 KV cache while improving many long-context results.

What it does not establish

The paper reports FLOPs and cache estimates, not an across-hardware wall-clock speedup or overall serving cost. The 1M comparison is a cost analysis; the shown post-training retrieval evaluation reaches 256k, not 1M. AgentPPL is internal, limiting independent verification.

Also keep attribution honest: HySparse2’s MQA, head dimensions, local-window design and token selection change together relative to baselines. Ablations isolate some components but reveal trade-offs. A full-attention layer remains computationally expensive, even if it is rare; extreme-length accuracy and engineering depend on selection and kernels (§§3.4, 4).

One-sentence takeaway: HySparse2 lets the front half construct reusable KV memory for the back half, while a few full-attention layers guide many sparse layers to exactly chosen old tokens plus a guaranteed recent window.

Primary source: Jianyu Wei, Yizhao Gao et al., HySparse2: Hybrid Sparse Attention with Two-Level KV Sharing, arXiv:2609.26368v1, 22 September 2026. All empirical claims above are attributed to that PDF; the interactive lab and pseudocode are teaching illustrations.