TL;DR
A memory management algorithm that reduces GPU memory waste from the Key-Value (KV) cache during large language model inference by allocating cache space in non-contiguous pages.
PagedAttention draws inspiration from virtual memory paging in operating systems to solve the memory fragmentation bottleneck in model serving. Instead of pre-allocating contiguous blocks of GPU memory for the maximum possible generation length of each request, it partitions the Key-Value cache into fixed-size physical pages. This allows the system to allocate memory dynamically as tokens are generated, sharing cache pages across requests and slashing memory waste by up to ninety percent.
Why this matters for your business
By dramatically optimizing GPU memory allocation, PagedAttention enables businesses to run inference at significantly higher throughput and batch sizes, reducing operational serving costs.