LiteLLM
The shared LLM proxy and gateway used by AI services.
Component Category
LLM proxy and gateway
Component Description
LiteLLM provides a unified gateway for accessing language models and model providers through a consistent interface.
Why It Is Used
In BullSequana AI, LiteLLM simplifies how AI services call models, apply shared access patterns, and manage routing and credentials across different providers and endpoints.
Learn More
Position In AI
LiteLLM is an important internal AI dependency, but it should not usually be treated as the primary long-term integration surface for product developers. The preferred stable product-facing contract is the BSQAI API.
Deployment notes
- Requests through the gateway route and the proxy share a 600-second budget, so long generations and streaming responses are not cut off by shorter infrastructure timeouts.
- A dedicated Redis StatefulSet (
litellm-redis) backs a response cache and keeps virtual-key authentication state consistent across workers and replicas. - Worker count is pinned per pod, replicas spread across nodes, and failed upstream calls retry automatically.
- The image is mirrored in the platform's private registry.
Interacts With
AI Web Portal, for user-facing model access.BSQAI API, for application-level LLM calls.Model Installer, for shared model access patterns.Redis, for response caching and shared authentication state throughlitellm-redis.Prometheus, which receives LiteLLM request and service metrics through callbacks.OpenTelemetryandMLflow, which receive backend-owned GenAI traces for workflows that call LiteLLM.KServeandvLLM, which provide private inference services through platform-managed runtime profiles.External model providers, where applicable.