Inference

KServe model serving, private gateways, runtime profiles, and GPU-aware scheduling.

Agentic Friendly

BullSequana AI 1.3.0 separates model serving into an inference Argo CD parent application. Higher-level AI services choose and expose models; this layer operates their Kubernetes workloads.

Main inference components

ComponentMain role
KServeReconciles LLMInferenceService and predictive InferenceService workloads
vLLMExecutes high-throughput language-model inference
KAI SchedulerPlaces and quotas GPU workloads
Envoy AI GatewayProvides the private inference data path
Gateway API Inference ExtensionAdds model-aware routing to the gateway path

Runtime profiles

The platform supplies CPU and NVIDIA GPU vLLM profiles for 1, 2, 4, and 8 GPUs. A profile defines the image, resources, model-loading behavior, and runtime arguments for a KServe deployment. Operators can tune or add profiles for site hardware.

Model artifacts can come from Hugging Face, a PVC, the platform object store, or an optional dedicated [model_storage] S3 endpoint. Direct-S3 profiles stream Safetensors from approved S3-compatible storage.

Request path

Model workloads are private. Applications authenticate to the BSQAI API, which routes governed model calls through LiteLLM and the internal inference gateway to KServe. Workloads are not intended to have their own public route.

Operational boundary

Foundation owns the serving control plane, scheduling, networking, storage integration, and telemetry. AI owns model registration, tenant authorization, model discovery, and the stable application APIs.

On this page