AI Gateway Overview
CSGHub AI Gateway provides a common entry point for AI services, connecting self-hosted inference and external model providers. Developers use gateway API keys to call enabled services. Administrators manage upstream connections, routing, content checks, and usage, and use metrics and request logs to investigate issues.
Why AI Gateway
When several teams use models from different sources, AI Gateway helps them:
- Connect services: Access public inference, dedicated endpoints, and commercial model APIs through a gateway.
- Manage access and spending: Use personal or organization credentials, set spending limits, and manage access to model services.
- Improve reliability: Configure multiple upstreams for a model and manage traffic with routing, health checks, and failover.
- Understand usage: Review request volume, charges, time to first token, cache hit rate, and upstream health.
See Core Capabilities for the capability map and task guides.
Service sources and request flow
| Source | Who runs the model | Starting point |
|---|---|---|
| Commercial API | External model provider | An administrator configures the endpoint, provider credentials, and model name |
| Public inference (Serverless API) | Platform administrator | Call an enabled model, or have an administrator create public inference |
| Dedicated inference endpoint | Endpoint owner on the platform | Create an endpoint and follow its API instructions |
The application sends a request to the gateway. The gateway routes it according to identity, service configuration, and routing rules. The upstream performs inference. Model repositories manage files and versions; inference endpoints manage runtime resources; the gateway manages access to services.
Applications use gateway API keys to invoke services. Administrators use provider credentials to connect upstreams. Obtain the API address and model name from the service page before making a request.
Start with a task
| Task | Guide |
|---|---|
| Make your first request | Quickstart |
| Connect a self-hosted or external model | Connect and Call Models |
| Set up personal or team credentials | Identity and Access |
| Manage content checks and request data | Content Safety and Data Handling |
| Configure upstream routing and failover | Routing and Reliability |
| Set limits or understand charges | Limits, Metering, and Costs |
| Review service health and investigate errors | Monitoring, Logs, and Troubleshooting |
Edition Notes
- Community Edition (CE) provides OpenAI-compatible model access. Available interfaces depend on the deployment version, inference runtime, and model capabilities.
- Enterprise Edition (EE) adds organizational management and extensions for Agent runtimes, sandboxes, the MCP Gateway control plane, and optional LLM-based content guards.
- Commercial operating deployments also involve pricing, consumption records, top-ups, and invoices. See Operating Paid Model Services. These workflows are separate from internal usage management.
Management screens and optional components depend on the installed edition and licensed scope. Before calling a model, review its supported parameters and confirm that the caller has access to the service.