- Keeping a long agent loop inside the context window by summarizing the turns that no longer change the outcome.
- Cutting the input tokens of every following turn, which otherwise resend the full history.
- Compacting on demand instead of waiting for a threshold to be crossed.
Overview
Context compaction summarizes the older items of a conversation through a model and replaces them with one compaction item, keeping the most recent items verbatim. It is a history rewrite, not prompt compression or truncation. The summary is written to preserve what a continuing agent needs:- Decisions made and their rationale.
- Files read or modified, and their current state.
- Tool outcomes that still matter.
- Open tasks and unresolved questions.
Quick start
POST /v3/router/responses/compact summarizes a conversation and returns the item to continue from. The same operation is available on the OpenAI-compatible base as POST /v1/responses/compact.
The SDK examples call
orq.responses.compact, which needs an SDK built from the 4.16 API line. The current stable SDKs do not ship it.Request fields
The endpoint has no threshold, and a short conversation is compacted whole. A longer one keeps the most recent ~20,000 tokens verbatim out of the summary, and the response contains only the compaction item, so keep the reserved recent turns before continuing from the item.
Automatic compaction
Setcontext_management on a continued Responses request and the gateway compacts as the conversation grows:
- It loads the stored history from
previous_response_id. - Everything older than the reserve is summarized into one item.
- The item is prepended to the response
output.
The pass reads history the gateway loaded, so the response being continued must be stored.
store defaults to true; a first turn sent with store: false has no history to compact. See Statefulness.Entry fields
context_management takes a list of entries. One compaction entry carries:
The pass runs only on history the gateway loaded, whether from
previous_response_id or a conversation. Items sent in the body of the current request are not part of it, so a single oversized request is not reduced.
Threshold and reserve
The pass runs only when the loaded history exceedscompact_threshold. Everything older than the reserve window is sent to the summarizer; the reserve itself stays verbatim.
Continuing from a compaction item
Send the compaction item back in theinput of a later Responses request to continue from the summary. The gateway reads an item by its shape, not by where it came from:
- A
compactionitem with nocontent, acreated_by, and anencrypted_contentexpands into a developer message, so the model reads its text as context. - Any other item keeps its own envelope and is passed through untouched.
Provider-native compaction
An OpenAI model can compact through the provider instead of the gateway. Everycompaction entry is forwarded to OpenAI with the request, whatever provider_native says, so set the flag to true to turn off the gateway’s own pass: the provider then applies its compaction alone and returns the item in its own format.
The setting applies to OpenAI models. On every other model the gateway performs the pass.
Prompt caching
Compaction is a model call of its own, and it carries the request’sprompt_cache_key when one is set, so the summarization request takes part in prompt caching like any other call. See Prompt caching.