Skip to content

Sparse Attention

Topic archive1 matches

Back to homeGEO summary endpoint

2026-09-21

Technology

  • RBS-Attention training-free method optimizes long-context LLM prefill: To address the prefill bottleneck in long-context large language model inference, researchers proposed RBS-Attention, a training-free sparse-prefill method. The approach uses a centroid base branch to capture average relevance and a rescue branch to identify blocks at risk of underestimation due to mean dilution.

    AI ResearcharXiv

    Permalink