Skip to main content

Performance and Cost Optimization

Evaluate AI services using task quality, response speed, and consumption together. After changing routing, inference settings, or models, compare the same workload rather than reducing price at the expense of business results.

Establish a baseline​

Use representative requests and record model/version, input and output size, success rate, TTFT, total response time, and charges. Keep request samples and traffic conditions consistent when comparing results before and after a change.

GoalInformation to inspect
Reduce initial waitTTFT, Prompt Cache, cold starts, and instance load
Improve success rateHealth, distribution, error types, and rate limits
Reduce chargesActual usage, applicable price, duplicate requests, and task quality

Routing and cache affinity​

For repeated context or conversations, evaluate whether session hashing helps the upstream reuse Prompt/KV Cache. Compare cache metrics and TTFT while checking for uneven load.

The upstream inference service provides Prompt/KV Cache to reuse computation for input context. Check the upstream’s supported cache behavior and pricing before using it.

Self-hosted inference configuration​

Public inference with zero minimum replicas can scale down while idle, so a later request requires a cold start. For latency-sensitive workloads, compare the benefit and resource cost of keeping replicas running. See Public Inference Administration.

Dedicated endpoints allow resource, runtime, and replica changes; some require stopping and restarting. See Endpoint Settings and Inference Runtimes. After adjusting inference parameters, use gateway metrics to observe cache hit rates and latency.

Validate changes​

Check business examples and use Model Evaluation for additional offline validation. Record the change and compare gateway metrics with charges.

Check model quality with offline evaluation and test requests first, then observe response behavior and costs in the application. Available routing and inference options are those exposed by the deployment.