Skip to main content

Routing and Reliability

Multiple upstreams for the same model allow traffic distribution and use of alternative nodes during failures. This page explains routing configuration and application-side handling of limits and interruptions.

Configure upstreams​

Edit a service under Admin Console → AI Gateway → Commercial API. Add each endpoint, its credentials, and weight, then inspect health status. Select upstreams for the same target model capability and check that their request parameters and response formats are compatible.

StrategyConfiguration and use
Round robinDistribute requests across upstreams for independent calls, using weights to set allocation
Session hashingPrefer the same node for a repeated identifier, useful when the upstream can reuse context caches

Session hashing accepts a configured HTTP header, such as X-Session-Id. Without a session header, distribution uses caller identity. Hash replicas represent virtual nodes in the hash ring; the administration guide lists a default of 64. See Commercial API Administration.

Session hashing provides routing affinity, not conversation storage. Applications still send the required context. Failed nodes or configuration changes can cause requests to move to another node.

Health checks and failover​

The gateway provides upstream health checks, circuit breaking, isolation, and request-level failover. Effective failover requires a correctly configured and authorized alternative upstream.

In a test environment, verify:

  1. Normal availability and expected traffic distribution.
  2. Health and distribution changes when an upstream fails.
  3. Routing behavior after the upstream recovers.

Check failover for different errors and request stages in a test environment. After a stream is interrupted, preserve received content and use the application’s business state to decide whether to start a new request.

Rate limits, timeouts, and retries​

TPM controls token usage per unit time, separately from cumulative spending limits. See Limits, Metering, and Costs. Thresholds and scope depend on deployment configuration.

Applications should set an acceptable request timeout and distinguish authorization, allowance, rate-limit, and upstream failures:

  • Fix permission or spending-limit errors before retrying.
  • Use bounded retries with backoff for recoverable failures, including a total time budget.
  • Inspect partial output and business state before restarting an interrupted stream.
  • For tools that perform writes, implement idempotency in the application and tool service to prevent duplicate operations during retries.

After cancellation, interruption, or timeout, check the upstream request’s final state and reconcile charges against consumption records.

Monitoring, Logs, and Troubleshooting · Performance and Cost Optimization