FlashPrefill V2 paper reports large long-context serving speedups with block-sparse attention
An arXiv paper describes FlashPrefill V2, a block-sparse attention system for long-context LLM serving, and reports speedups of up to 47.26x over FlashAttention-2 at 128K context on NVIDIA H20 GPUs.