Skip to main content

CSGHub Release Notes v2026.09

1 AI Gateway High Availability & Intelligent Traffic Governance​

Comprehensively enhanced AI Gateway traffic governance and high availability in high-concurrency environments, introducing bounded request backpressure queues and intelligent semantic routing.

  • Traffic Governance & Backpressure Queuing: Introduced bounded request queues, Admission Control, and Concurrency Control to provide effective backpressure protection during sudden traffic surges, preventing upstream overloads and crashes.
  • Multi-Upstream Balancing & Configuration Visualization: Optimized queue-mode load-balancing strategies across multiple upstreams and surfaced key real-time queue operating parameters on the API UI.
  • Intelligent Semantic Router: Added semantic routing capabilities to automatically route requests to optimal model endpoints based on user intent and text embeddings, while fine-tuning underlying default inference parameters.

2 Token Factory & Unified Multi-Model Operations (EE)​

Centered on Token asset governance and multi-model management, unifying internal deployed models and commercial APIs with enhanced quotas and reasoning model thinking-flow metering.

  • Upstream Capacity & Token Budget Control: Supported fine-grained capacity management for Model and Upstream instances, establishing Token Quota and Token Budget mechanisms across User, Org, and API Key dimensions.
  • Unified Model Management & Seamless Navigation: Overhauled the gateway Handle layer orchestration to uniformly govern self-hosted models (Serverless/Dedicated Instances) and external commercial APIs, enabling fast two-way navigation between Serverless and Commercial API interfaces.
  • Reasoning Model Thinking Stream & Token Metering: Fully supported thinking processes and tool calls for new-generation reasoning models (with reasoning_content pass-through) and added reasoning_tokens output to the platform metering and billing system.
  • Token Factory Operations Dashboard & Reports: Added Token Factory operations statistical reports and a frontend monitoring dashboard for administrators to monitor Token usage trends and upstream distribution.

3 Multi-Level Organizations & Fine-Grained Authorization (EE)​

Supported hierarchical multi-level organizational trees and modern relation-based authorization to meet enterprise-level permission isolation and multi-tier team collaboration needs.

  • Hierarchical Multi-Level Organization Tree: Supported multi-tier parent-child organization relationships with permission inheritance and intuitive web-based creation and maintenance in the Admin console.
  • OpenFGA Authorization Engine Integration: Abstracted a unified permission verification layer, establishing fine-grained direct authorization from models, datasets, and code repositories to organizations and individuals, with smooth migration and query index optimization.
  • Asset Adaptation & Search Refactoring for Multi-Level Orgs: Refactored global search for both single and multi-level organizations, fully adapting Spaces creation, public/private visibility, and member read/write permissions.
  • One-Click Personal Asset Migration to Organizations: Enabled regular users to migrate personal datasets to specified organizations in one click for unified team governance.

4 Model Ecosystem & Infrastructure Optimization​

Broadened hardware and open-source model framework support, refining multi-version model evaluation and high-capacity storage.

  • Multi-Version Model Evaluation Comparison: Introduced a multi-version evaluation comparison UI, expanded evaluation dimensions, and improved log and error viewer displays.
  • Latest Inference Framework & Multi-Modal Extensions: Upgraded underlying vLLM runtime to support the latest models such as GLM-5.3-Flash and Qwen series; integrated PaddleX framework with one-click deployment for PaddleOCR-VL multi-modal models.
  • External LFS & Large File Storage: Supported storing Git LFS and large files directly into external S3-compatible object storage to prevent Pod storage exhaustion, accompanied by underlying PVC storage quota limits.
  • Public Agent Infrastructure & Image Registry: Designed an enterprise-grade high-concurrency and multi-user agent framework, supported dynamic injection of PIP_INDEX_URL for rapid sandbox startup, and added a unified image management center.

5 Admin Console Interaction Refinements (EE)​

Elevated administration efficiency and everyday interactions across multiple management scenarios.

  • Cluster Name Customization in Admin Console: Supported dynamic configuration and modification of cluster names directly from the frontend UI after deployment.
  • Instant Search in Admin Sidebar: Added a quick search input at the top of the admin sidebar to instantly locate target entries across extensive menu hierarchies.
  • Localized Export Reports: Standardized report headers and status fields to localized Chinese when Chinese language is selected.
  • Responsive Layout & Filter Streamlining: Optimized model list filters for narrow screens and mobile devices; removed repetitive industry selection prompts on login and eliminated duplicate sorting options.

6 Key Security Hardening & System Stability​

Resolved high-risk security vulnerabilities and enhanced platform availability across cloud environments.

  • High-Risk Vulnerability Fixes: Fixed the CNVD-reported Cookie forgery administrator authentication bypass vulnerability; implemented frontend debouncing and rate limiting on Space start buttons to prevent abuse.
  • Multi-Cloud Database & Storage Compatibility: Removed strict runtime dependency on PostgreSQL timescaledb extension for smooth deployment on AWS RDS; fixed evaluation report download failures when S3 buckets block public access.
  • Lifecycle & Gateway Stability Fixes: Fixed inference instance deletion leaks where ksvc/Pods remained uncleaned; resolved AI Gateway MCP proxy nil-pointer panics due to uninitialized deployer, and fixed Prompt Caching errors.
  • Frontend Memory & Resource Optimization: Fixed memory leaks during long-running media playback to prevent site access disruptions.

7 Billing & Invoicing Workflow Enhancements​

Streamlined invoicing and recharge closure.

  • Direct Invoicing After Recharge: Upgraded invoice management workflows, allowing users to apply for invoices immediately after account recharge.
  • Standardized Billing Timestamps: Standardized recharge weekly report and billing timestamps to Beijing Time (UTC+8) to prevent cross-timezone reconciliation discrepancies.

Historical Versions​

Click Release List to view all historical versions.