Skip to main content
Use Cases
  • Keeping a long agent loop inside the context window by summarizing the turns that no longer change the outcome.
  • Cutting the input tokens of every following turn, which otherwise resend the full history.
  • Compacting on demand instead of waiting for a threshold to be crossed.

Overview

Context compaction summarizes the older items of a conversation through a model and replaces them with one compaction item, keeping the most recent items verbatim. It is a history rewrite, not prompt compression or truncation. The summary is written to preserve what a continuing agent needs:
  • Decisions made and their rationale.
  • Files read or modified, and their current state.
  • Tool outcomes that still matter.
  • Open tasks and unresolved questions.
Summarization is lossy. A turn the summary omits cannot be recovered from the compaction item, so keep the original transcript wherever the detail matters. Compaction applies to the Responses API, which keeps conversation state server-side. Reach for it when a conversation’s history outgrows the model’s context window, or when its input tokens dominate the bill.

Quick start

POST /v3/router/responses/compact summarizes a conversation and returns the item to continue from. The same operation is available on the OpenAI-compatible base as POST /v1/responses/compact.
The SDKs take a message’s content as a string, while the REST body takes content parts.
The SDK examples call orq.responses.compact, which needs an SDK built from the 4.16 API line. The current stable SDKs do not ship it.

Request fields

The endpoint has no threshold, and a short conversation is compacted whole. A longer one keeps the most recent ~20,000 tokens verbatim out of the summary, and the response contains only the compaction item, so keep the reserved recent turns before continuing from the item.

Automatic compaction

Set context_management on a continued Responses request and the gateway compacts as the conversation grows:
  • It loads the stored history from previous_response_id.
  • Everything older than the reserve is summarized into one item.
  • The item is prepended to the response output.
The pass reads history the gateway loaded, so the response being continued must be stored. store defaults to true; a first turn sent with store: false has no history to compact. See Statefulness.

Entry fields

context_management takes a list of entries. One compaction entry carries: The pass runs only on history the gateway loaded, whether from previous_response_id or a conversation. Items sent in the body of the current request are not part of it, so a single oversized request is not reduced.

Threshold and reserve

The pass runs only when the loaded history exceeds compact_threshold. Everything older than the reserve window is sent to the summarizer; the reserve itself stays verbatim.

Continuing from a compaction item

Send the compaction item back in the input of a later Responses request to continue from the summary. The gateway reads an item by its shape, not by where it came from:
  • A compaction item with no content, a created_by, and an encrypted_content expands into a developer message, so the model reads its text as context.
  • Any other item keeps its own envelope and is passed through untouched.

Provider-native compaction

An OpenAI model can compact through the provider instead of the gateway. Every compaction entry is forwarded to OpenAI with the request, whatever provider_native says, so set the flag to true to turn off the gateway’s own pass: the provider then applies its compaction alone and returns the item in its own format.
The setting applies to OpenAI models. On every other model the gateway performs the pass.

Prompt caching

Compaction is a model call of its own, and it carries the request’s prompt_cache_key when one is set, so the summarization request takes part in prompt caching like any other call. See Prompt caching.

Limitations