> ## Documentation Index
> Fetch the complete documentation index at: https://docs.orq.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Tool search server tool

> Keep large tool libraries out of the prompt and let the model search for the tools it needs with the orq:tool_search server tool.

The `orq:tool_search` tool defers function tools from the prompt until the model asks for them. Mark function tools with `defer_loading: true`, and the model receives only the search tool plus the non-deferred tools. When the model needs a capability, it searches the deferred tools with a regular expression, and the matching tools become callable on the next turn.

Deferring tools reduces prompt tokens and keeps the tool list stable for prompt caching. It pays off from roughly ten tools or 10k tokens of tool definitions. The search runs on the **AI Gateway**, so it works with any model that supports tool calling.

## Quick start

The examples use the client configuration from the [Server tools overview](/ai-gateway/features/server-tools).

<CodeGroup>
  ```bash cURL theme={"theme":{"light":"github-light","dark":"github-dark"}}
  curl -X POST https://my.orq.ai/v3/router/responses \
    -H "Authorization: Bearer $ORQ_API_KEY" \
    -H "Content-Type: application/json" \
    -d '{
      "model": "openai/gpt-5.4-mini",
      "input": "Book me a table for four tomorrow night and add it to my calendar.",
      "tools": [
        { "type": "orq:tool_search" },
        {
          "type": "function",
          "name": "book_restaurant",
          "description": "Reserve a restaurant table for a party at a given date and time.",
          "parameters": {
            "type": "object",
            "properties": {
              "party_size": { "type": "integer", "description": "Number of guests." },
              "datetime": { "type": "string", "description": "Reservation time as an ISO 8601 timestamp." }
            },
            "required": ["party_size", "datetime"]
          },
          "defer_loading": true
        },
        {
          "type": "function",
          "name": "create_calendar_event",
          "description": "Add an event to the user'"'"'s calendar.",
          "parameters": {
            "type": "object",
            "properties": {
              "title": { "type": "string", "description": "Event title." },
              "start": { "type": "string", "description": "Start time as an ISO 8601 timestamp." }
            },
            "required": ["title", "start"]
          },
          "defer_loading": true
        }
      ]
    }'
  ```

  ```typescript TypeScript theme={"theme":{"light":"github-light","dark":"github-dark"}}
  const response = await client.responses.create({
    model: 'openai/gpt-5.4-mini',
    input: 'Book me a table for four tomorrow night and add it to my calendar.',
    tools: [
      { type: 'orq:tool_search' },
      {
        type: 'function',
        name: 'book_restaurant',
        description: 'Reserve a restaurant table for a party at a given date and time.',
        parameters: {
          type: 'object',
          properties: {
            party_size: { type: 'integer', description: 'Number of guests.' },
            datetime: { type: 'string', description: 'Reservation time as an ISO 8601 timestamp.' },
          },
          required: ['party_size', 'datetime'],
        },
        defer_loading: true,
      },
    ] as any,
  });

  console.log(response.output);
  ```

  ```python Python theme={"theme":{"light":"github-light","dark":"github-dark"}}
  response = client.responses.create(
      model="openai/gpt-5.4-mini",
      input="Book me a table for four tomorrow night and add it to my calendar.",
      tools=[
          {"type": "orq:tool_search"},
          {
              "type": "function",
              "name": "book_restaurant",
              "description": "Reserve a restaurant table for a party at a given date and time.",
              "parameters": {
                  "type": "object",
                  "properties": {
                      "party_size": {"type": "integer", "description": "Number of guests."},
                      "datetime": {"type": "string", "description": "Reservation time as an ISO 8601 timestamp."},
                  },
                  "required": ["party_size", "datetime"],
              },
              "defer_loading": True,
          },
      ],
  )

  print(response.output)
  ```
</CodeGroup>

On Chat Completions, set `defer_loading` next to `function` on the tool object.

## How it works

1. The model sees `orq_tool_search` and the tools that are not deferred.
2. When it needs a capability, it calls `orq_tool_search` with a regular expression, for example `calendar|event`.
3. The **AI Gateway** matches the pattern case-insensitively against the names, descriptions, and parameter names and descriptions of the deferred tools. A pattern of `weather` finds a tool whose only mention of weather is in a parameter description.
4. The matching tools are added to the model's tool list. The model calls them like any other function tool, and the application handles the call as usual.

Revealing a tool does not change the tools already in the conversation, so the cached prompt prefix stays intact. On a continuation with `previous_response_id` or `conversation`, tools revealed in earlier turns are offered from the start.

## Configuration

| Parameter | Type | Required | Default | Description |
| - | - | - | - | - |
| `type` | string | Yes | | Must be `orq:tool_search`. |
| `max_results` | integer | No | `5` | Maximum tools revealed per search. Accepted range: 1 to 50. |

Set `defer_loading: true` on any `function` tool to hide it. Only function tools can be deferred; `orq:tool_search` itself is never deferred so the model always has a tool to call.

## Rules

* A request with a deferred tool must include `orq:tool_search`, otherwise it fails with `tool_search_required`.
* `tool_choice` must be omitted, `auto`, or `none`. Forcing a tool call conflicts with tool search and fails with `tool_choice_conflict`.
* A pattern is at most 200 characters. A malformed pattern is returned to the model as a tool error, not as a request failure.
* A call to a deferred tool that no search revealed is rejected as a tool that was not offered.

## Agents

Agents can use tool search as well. Add `orq:tool_search` to the agent's tools and set `configuration.defer_loading` on the custom tools to hide:

```json theme={"theme":{"light":"github-light","dark":"github-dark"}}
{
  "tools": [
    { "type": "orq:tool_search", "configuration": { "max_results": 5 } },
    { "type": "http", "key": "book_restaurant", "configuration": { "defer_loading": true } },
    { "type": "http", "key": "create_calendar_event", "configuration": { "defer_loading": true } }
  ]
}
```

Without `orq:tool_search` on the agent, `defer_loading` has no effect and every tool stays visible.

## Cost and usage

Tool searches have no additional charge. The number of searches appears at `usage.server_tool_use.tool_search_requests`. Each search is an `orq:tool_search` output item whose result lists the revealed tools.
