Performance and Cost Optimization
Evaluate AI services using task quality, response speed, and consumption together. After changing routing, inference settings, or models, compare the same workload rather than reducing price at the expense of business results.
Establish a baseline
Use representative requests and record model/version, input and output size, success rate, TTFT, total response time, and charges. Keep request samples and traffic conditions consistent when comparing results before and after a change.
| Goal | Information to inspect |
|---|---|
| Reduce initial wait | TTFT, Prompt Cache, cold starts, and instance load |
| Improve success rate | Health, distribution, error types, and rate limits |
| Reduce charges | Actual usage, applicable price, duplicate requests, and task quality |
Routing and cache affinity
For repeated context or conversations, evaluate whether session hashing helps the upstream reuse Prompt/KV Cache. Compare cache metrics and TTFT while checking for uneven load.
The upstream inference service provides Prompt/KV Cache to reuse computation for input context. Check the upstream’s supported cache behavior and pricing before using it.
Self-hosted inference configuration
Public inference with zero minimum replicas can scale down while idle, so a later request requires a cold start. For latency-sensitive workloads, compare the benefit and resource cost of keeping replicas running. See Public Inference Administration.
Dedicated endpoints allow resource, runtime, and replica changes; some require stopping and restarting. See Endpoint Settings and Inference Runtimes. After adjusting inference parameters, use gateway metrics to observe cache hit rates and latency.
Validate changes
Check business examples and use Model Evaluation for additional offline validation. Record the change and compare gateway metrics with charges.
Check model quality with offline evaluation and test requests first, then observe response behavior and costs in the application. Available routing and inference options are those exposed by the deployment.