PagedAttention

Paged attention

Infrastructure

Deployment

Soft glowing orange and yellow light with a gradient blending into black background.
TL;DR
A memory management algorithm that reduces GPU memory waste from the Key-Value (KV) cache during large language model inference by allocating cache space in non-contiguous pages.

In depth

PagedAttention draws inspiration from virtual memory paging in operating systems to solve the memory fragmentation bottleneck in model serving. Instead of pre-allocating contiguous blocks of GPU memory for the maximum possible generation length of each request, it partitions the Key-Value cache into fixed-size physical pages. This allows the system to allocate memory dynamically as tokens are generated, sharing cache pages across requests and slashing memory waste by up to ninety percent.

Why this matters for your business

By dramatically optimizing GPU memory allocation, PagedAttention enables businesses to run inference at significantly higher throughput and batch sizes, reducing operational serving costs.

Ready to scale AI across your organization?

Talk to an AI expert