TL;DR
The algorithmic process of summarizing, filtering, or pruning conversation history and session data to fit a long-running AI agent's state within a fixed model token window.
Unlike human brains, large language models do not natively store ongoing memories and must ingest entire past conversations with each prompt, leading to token overflow during long-running sessions. Context compaction resolves this constraint by dynamically summarizing old dialogues, discarding redundant information, or utilizing loss-aware token-level pruning. This ensures that the agent retains critical state, system prompts, and tool permissions, while keeping the operating context lightweight and highly efficient.
Why this matters for your business
This technique dramatically lowers inference latency and API costs while preventing models from losing core instructions or experiencing performance degradation over extended operations.