Skip to main content

AI Gateway Overview

CSGHub AI Gateway provides a common entry point for AI services, connecting self-hosted inference and external model providers. Developers use gateway API keys to call enabled services. Administrators manage upstream connections, routing, content checks, and usage, and use metrics and request logs to investigate issues.

Why AI Gateway​

When several teams use models from different sources, AI Gateway helps them:

  • Connect services: Access public inference, dedicated endpoints, and commercial model APIs through a gateway.
  • Manage access and spending: Use personal or organization credentials, set spending limits, and manage access to model services.
  • Improve reliability: Configure multiple upstreams for a model and manage traffic with routing, health checks, and failover.
  • Understand usage: Review request volume, charges, time to first token, cache hit rate, and upstream health.

See Core Capabilities for the capability map and task guides.

Service sources and request flow​

SourceWho runs the modelStarting point
Commercial APIExternal model providerAn administrator configures the endpoint, provider credentials, and model name
Public inference (Serverless API)Platform administratorCall an enabled model, or have an administrator create public inference
Dedicated inference endpointEndpoint owner on the platformCreate an endpoint and follow its API instructions

The application sends a request to the gateway. The gateway routes it according to identity, service configuration, and routing rules. The upstream performs inference. Model repositories manage files and versions; inference endpoints manage runtime resources; the gateway manages access to services.

Applications use gateway API keys to invoke services. Administrators use provider credentials to connect upstreams. Obtain the API address and model name from the service page before making a request.

Start with a task​

TaskGuide
Make your first requestQuickstart
Connect a self-hosted or external modelConnect and Call Models
Set up personal or team credentialsIdentity and Access
Manage content checks and request dataContent Safety and Data Handling
Configure upstream routing and failoverRouting and Reliability
Set limits or understand chargesLimits, Metering, and Costs
Review service health and investigate errorsMonitoring, Logs, and Troubleshooting

Edition Notes​

  • Community Edition (CE) provides OpenAI-compatible model access. Available interfaces depend on the deployment version, inference runtime, and model capabilities.
  • Enterprise Edition (EE) adds organizational management and extensions for Agent runtimes, sandboxes, the MCP Gateway control plane, and optional LLM-based content guards.
  • Commercial operating deployments also involve pricing, consumption records, top-ups, and invoices. See Operating Paid Model Services. These workflows are separate from internal usage management.

Management screens and optional components depend on the installed edition and licensed scope. Before calling a model, review its supported parameters and confirm that the caller has access to the service.