Skip to main content
Use Cases
  • Distributing traffic across multiple provider accounts to stay within per-key rate limits.
  • A/B testing providers by routing a configurable percentage of traffic to each.
  • Spreading traffic across multiple providers to reduce blast radius from a single provider outage.
  • Maximizing throughput when one provider’s capacity is a bottleneck.
  • Minimizing response time by routing to the fastest available model.

Quick Start

Distribute requests across multiple providers using weighted routing.

Configuration

Weight Calculation:
  • Weights are normalized: [0.4, 0.8][33%, 67%].
  • Higher weight = more traffic.
  • Minimum weight: 0.001.
  • Weight is ignored for round robin selection; every model in the list still receives an equal share of traffic.
A matching Routing Rule with its own target models configured overwrites this load_balancer parameter entirely, regardless of what the request sends. The Routing Rule’s own strategy and models are used instead; the two configurations are never merged.
load_balancer selects one model per request. It does not retry a failed call against another model in the pool. The top-level model is used as-is only when load_balancer cannot produce a selection, for example an empty models list. To fail over when a call errors, pair load_balancer with Fallbacks.

Latency-based routing

Route each request to the model with the lowest recently observed latency, instead of a fixed traffic split.
How selection works:
  • Latency is tracked per model as a running average that weights recent calls more heavily, so the selection adapts quickly when a provider speeds up or slows down.
  • With fresh data (samples from the last 5 minutes) for every configured model, the lowest-latency model is selected. A model with no data, or stale data, is probed again instead of being written off.
  • Ten percent of selections explore the full pool by configured weight to keep latency measurements fresh. Failed calls are penalized so a fast failure does not look like a fast success.
  • Configured weights decide exploration and near-ties: with the two fastest models within 0.5 ms of each other, the higher weight wins. With clear data, the lowest latency always wins.

See also: Organization-level load balancing

To apply load balancing across your organization without changing request code, use Routing Rules to configure Fallback, Latency, Weighted, and Round Robin strategies at the workspace level.

Weight-based routing

Split traffic across models by percentage weights instead of latency or a fixed rotation.

Patterns

Use cases

Round Robin routing

Rotate through the configured models evenly, one request at a time, instead of splitting traffic by weight or latency.
How selection works:
  • Each request selects the next model in the list in turn, wrapping back to the first model after the last one. With two models, requests alternate; with three, the rotation cycles through all three.
  • Weights must be included in the request but do not affect selection: round robin is purely rotational, so every model receives an equal share of traffic over time regardless of the weights.
  • Round robin only selects a model per request; it does not retry or fail over a failed call.
Round robin selects one model per request but does not fail over when that model errors. To add failover on top of round robin, pair load_balancer with Fallbacks.

Code examples

Monitoring

Track these metrics for optimal load balancing:
Key Metrics:
  • Traffic distribution: Actual vs expected percentages.
  • Cost per model: Monitor spending across providers.
  • Response times: Compare latency by model.
  • Error rates: Track failures by provider.
With latency_based, response time is already the selection signal. Manual weight tuning for performance is not needed; adjust weights only to influence exploration, cold-start probing, and near-tie behavior.

Troubleshooting

Uneven distribution
  • Check if weights are normalized correctly.
  • Verify sufficient request volume (min 100 requests for accuracy).
  • Monitor over longer time periods.
Unexpected costs
  • Track actual vs expected cost distribution.
  • Monitor for expensive model overuse.
  • Set up cost alerts per provider.
Performance issues
  • Check latency differences between models.
  • Monitor for provider-specific slowdowns.
  • Adjust weights based on performance data.
All traffic going to one model with latency_based
  • Expected once one model is consistently fastest. Ten percent of requests still explore the rest of the pool to keep their latency data fresh.
  • Confirm load_balancer.type is set to the intended strategy if an even split was expected instead.
Selection is slow to adapt after a deploy or restart
  • Expected. Latency history is in-memory: after a restart, every model is treated as unknown until fresh samples are collected.

Limitations

  • Probabilistic routing: Short-term traffic may not match exact weights.
  • Minimum volume needed: Requires sufficient requests for statistical accuracy.
  • Response variations: Different models may return varying output quality.
  • Cost complexity: Managing billing across multiple providers.
  • Provider dependencies: Requires API access to all models.
  • In-memory latency state: With latency_based, latency history is in-memory and does not persist across restarts.

Advanced weight-based usage

Environment-specific weights:
Dynamic weight adjustment:
With other features: