> ## Documentation Index
> Fetch the complete documentation index at: https://docs.orq.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Context compaction for long conversations

> Summarize the older turns of a Responses API conversation into a single compaction item so a long history stays inside the model's context window.

**Use Cases**

* Keeping a long agent loop inside the context window by summarizing the turns that no longer change the outcome.
* Cutting the input tokens of every following turn, which otherwise resend the full history.
* Compacting on demand instead of waiting for a threshold to be crossed.

***

## Overview

Context compaction summarizes the older items of a conversation through a model and replaces them with one compaction item, keeping the most recent items verbatim. It is a history rewrite, not prompt compression or truncation.

The summary is written to preserve what a continuing agent needs:

* Decisions made and their rationale.
* Files read or modified, and their current state.
* Tool outcomes that still matter.
* Open tasks and unresolved questions.

Summarization is lossy. A turn the summary omits cannot be recovered from the compaction item, so keep the original transcript wherever the detail matters.

Compaction applies to the [Responses API](/ai-gateway/features/responses-api), which keeps conversation state server-side. Reach for it when a conversation's history outgrows the model's context window, or when its input tokens dominate the bill.

## Quick start

`POST /v3/router/responses/compact` summarizes a conversation and returns the item to continue from. The same operation is available on the OpenAI-compatible base as `POST /v1/responses/compact`.

<CodeGroup>
  ```bash cURL theme={"theme":{"light":"github-light","dark":"github-dark"}}
  curl -X POST https://my.orq.ai/v3/router/responses/compact \
    -H "Authorization: Bearer $ORQ_API_KEY" \
    -H "Content-Type: application/json" \
    -d '{
      "model": "openai/gpt-4o",
      "input": [
        { "type": "message", "role": "user", "content": [{ "type": "input_text", "text": "We chose Postgres for the managed failover." }] },
        { "type": "message", "role": "assistant", "content": [{ "type": "output_text", "text": "Recorded." }] },
        { "type": "message", "role": "user", "content": [{ "type": "input_text", "text": "Add the migration." }] }
      ]
    }'
  ```

  ```typescript TypeScript theme={"theme":{"light":"github-light","dark":"github-dark"}}
  import { Orq } from "@orq-ai/node";

  const orq = new Orq({ apiKey: process.env.ORQ_API_KEY });

  const compaction = await orq.responses.compact({
    model: "openai/gpt-4o",
    input: [
      { type: "message", role: "user", content: "We chose Postgres for the managed failover." },
      { type: "message", role: "assistant", content: "Recorded." },
      { type: "message", role: "user", content: "Add the migration." },
    ],
  });
  ```

  ```python Python theme={"theme":{"light":"github-light","dark":"github-dark"}}
  import os

  from orq_ai_sdk import Orq

  orq = Orq(api_key=os.environ["ORQ_API_KEY"])

  compaction = orq.responses.compact(
      model="openai/gpt-4o",
      input=[
          {"type": "message", "role": "user", "content": "We chose Postgres for the managed failover."},
          {"type": "message", "role": "assistant", "content": "Recorded."},
          {"type": "message", "role": "user", "content": "Add the migration."},
      ],
  )
  ```
</CodeGroup>

The SDKs take a message's content as a string, while the REST body takes content parts.

<Note>
  The SDK examples call `orq.responses.compact`, which needs an SDK built from the 4.16 API line. The current stable SDKs do not ship it.
</Note>

```json theme={"theme":{"light":"github-light","dark":"github-dark"}}
{
  "id": "resp_01m47nq46dmgc1rqa7wmgje0hh",
  "object": "response.compaction",
  "output": [
    {
      "type": "compaction",
      "id": "cmpct_01m47nq46d2v7e3cq1xyh073xw",
      "encrypted_content": "### Decisions Made\n- Database: Postgres, chosen for the managed failover.\n\n### Outstanding Tasks\n- Add the migration.",
      "created_by": "openai/gpt-4o"
    }
  ],
  "created_at": 1791259021,
  "usage": { "input_tokens": 156, "output_tokens": 49, "total_tokens": 205 }
}
```

### Request fields

| Parameter | Type | Required | Description |
| - | - | - | - |
| `model` | string | Yes | Model that writes the summary. |
| `input` | string or array | No | Conversation to compact, as a string or a list of items. |
| `prompt_cache_key` | string | No | Cache key for the summarization call. |

The endpoint has no threshold, and a short conversation is compacted whole. A longer one keeps the most recent \~20,000 tokens verbatim out of the summary, and the response contains only the compaction item, so keep the reserved recent turns before continuing from the item.

## Automatic compaction

Set `context_management` on a continued Responses request and the gateway compacts as the conversation grows:

* It loads the stored history from `previous_response_id`.
* Everything older than the reserve is summarized into one item.
* The item is prepended to the response `output`.

<Note>
  The pass reads history the gateway loaded, so the response being continued must be stored. `store` defaults to `true`; a first turn sent with `store: false` has no history to compact. See [Statefulness](/ai-gateway/features/responses-api#statefulness).
</Note>

<CodeGroup>
  ```bash cURL theme={"theme":{"light":"github-light","dark":"github-dark"}}
  # Second turn of a stored conversation. Once the history loaded from
  # previous_response_id passes compact_threshold tokens, the older items
  # are replaced by one compaction item.
  curl -X POST https://my.orq.ai/v3/router/responses \
    -H "Authorization: Bearer $ORQ_API_KEY" \
    -H "Content-Type: application/json" \
    -d '{
      "model": "openai/gpt-4o",
      "previous_response_id": "resp_01m47nq46dmgc1rqa7wmgje0hh",
      "context_management": [
        { "type": "compaction", "compact_threshold": 1000 }
      ],
      "input": "Confirm the decisions so far."
    }'
  ```
</CodeGroup>

### Entry fields

`context_management` takes a list of entries. One `compaction` entry carries:

| Parameter | Type | Required | Description |
| - | - | - | - |
| `type` | string | Yes | `compaction` is the supported entry. |
| `compact_threshold` | integer | No | Token count the loaded history must exceed before the pass runs. Without it, the entry does nothing. OpenAI models require at least `1000`. |
| `provider_native` | boolean | No | Let an OpenAI model compact through the provider instead of the gateway. See [Provider-native compaction](#provider-native-compaction). |

The pass runs only on history the gateway loaded, whether from `previous_response_id` or a conversation. Items sent in the body of the current request are not part of it, so a single oversized request is not reduced.

### Threshold and reserve

The pass runs only when the loaded history exceeds `compact_threshold`. Everything older than the reserve window is sent to the summarizer; the reserve itself stays verbatim.

| `compact_threshold` | Recent tokens kept verbatim |
| - | - |
| 40000 or more | 20000 |
| Below 40000 | Half the threshold |

## Continuing from a compaction item

Send the compaction item back in the `input` of a later Responses request to continue from the summary. The gateway reads an item by its shape, not by where it came from:

* A `compaction` item with no `content`, a `created_by`, and an `encrypted_content` expands into a developer message, so the model reads its text as context.
* Any other item keeps its own envelope and is passed through untouched.

## Provider-native compaction

An OpenAI model can compact through the provider instead of the gateway. Every `compaction` entry is forwarded to OpenAI with the request, whatever `provider_native` says, so set the flag to `true` to turn off the gateway's own pass: the provider then applies its compaction alone and returns the item in its own format.

```json theme={"theme":{"light":"github-light","dark":"github-dark"}}
{
  "context_management": [
    { "type": "compaction", "compact_threshold": 1000, "provider_native": true }
  ]
}
```

<Note>
  The setting applies to OpenAI models. On every other model the gateway performs the pass.
</Note>

## Prompt caching

Compaction is a model call of its own, and it carries the request's `prompt_cache_key` when one is set, so the summarization request takes part in prompt caching like any other call. See [Prompt caching](/ai-gateway/features/prompt-caching).

## Limitations

| Limitation | Impact | Workaround |
| - | - | - |
| The pass covers loaded history only | Items sent in the current request are never compacted, so a single oversized request is not reduced | Continue the conversation with `previous_response_id` |
| `compact_threshold` is inert alone | Without the field, the entry does nothing | Set the threshold, or call the compact endpoint |
| A threshold below 1000 fails on OpenAI models | The provider rejects the request with `integer_below_min_value` | Use `1000` or more |
| Summarization is lossy | Detail the summary omits is not recoverable from the compaction item | Keep the original transcript, or raise the threshold so recent turns stay verbatim |
| Item provenance is not verified | A `compaction` item shaped like the gateway's own expands into a developer message, and `encrypted_content` holds plaintext while `created_by` is caller-supplied, so text of any origin can reach the model in that role | Echo back only items the gateway returned, and treat the summary as caller-controlled text |


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.