service_tier. The AI Gateway passes the value through and reports back the tier that served the request.
Use Cases
- Run latency-sensitive work ahead of standard traffic with
fast. - Move background work onto discounted spare capacity with
flex. - Record the tier that served each request, since billing follows it.
Quick Start
Setservice_tier in the request body. The AI Gateway accepts six values and sends the request to the provider on that tier, where the model sells it.
service_tier works on the Responses API and on Chat Completions.
Accepted values
Any other value is rejected before the provider call, on both endpoints:
Tier availability
A model sells a tier only where its pricing declares one, so the set differs per model and changes as providers revise their line-ups. A tier appears as a pricing variant on the model, in the model catalog, and on List Models. Readservice_tier back rather than assuming the tier was honoured; a model that does not sell the requested tier does not fail the request, it runs on the standard tier and reports default.
Reading back the served tier
Every response carriesservice_tier, naming the tier that served the request:
fast is reported as priority, and auto, scale and default are reported as default. A request that sets no tier is reported as default.