Inference
KServe model serving, private gateways, runtime profiles, and GPU-aware scheduling.
BullSequana AI 1.3.0 separates model serving into an inference Argo CD parent application. Higher-level AI services choose and expose models; this layer operates their Kubernetes workloads.
Main inference components
| Component | Main role |
|---|---|
| KServe | Reconciles LLMInferenceService and predictive InferenceService workloads |
| vLLM | Executes high-throughput language-model inference |
| KAI Scheduler | Places and quotas GPU workloads |
| Envoy AI Gateway | Provides the private inference data path |
| Gateway API Inference Extension | Adds model-aware routing to the gateway path |
Runtime profiles
The platform supplies CPU and NVIDIA GPU vLLM profiles for 1, 2, 4, and 8 GPUs. A profile defines the image, resources, model-loading behavior, and runtime arguments for a KServe deployment. Operators can tune or add profiles for site hardware.
Model artifacts can come from Hugging Face, a PVC, the platform object store, or an optional dedicated [model_storage] S3 endpoint. Direct-S3 profiles stream Safetensors from approved S3-compatible storage.
Request path
Model workloads are private. Applications authenticate to the BSQAI API, which routes governed model calls through LiteLLM and the internal inference gateway to KServe. Workloads are not intended to have their own public route.
Operational boundary
Foundation owns the serving control plane, scheduling, networking, storage integration, and telemetry. AI owns model registration, tenant authorization, model discovery, and the stable application APIs.