KServe
Model-serving control plane for KServe LLM and predictive inference workloads.
Component Category
AI serving / model inference
Component Description
KServe is the Kubernetes-native model-serving framework used by BullSequana AI 1.3.0. The platform installs KServe 0.20.0 and its LLM inference APIs, then uses LLMInferenceService resources for backend-managed model deployments. Predictive serving runtimes remain available through KServe InferenceService resources.
Why It Is Used
In BullSequana AI, KServe provides declarative model lifecycle, autoscaling, runtime configuration, workload status, and an internal inference data path.
Learn More
Deployment notes
KServe resources run in the kserve namespace and are owned by the separate inference Argo CD parent application. The layer includes KServe CRDs and resources, LLM inference CRDs, KAI Scheduler, Envoy AI Gateway, the Gateway API Inference Extension, runtime profiles, and a private internal gateway.
The shipped vLLM profiles include CPU and NVIDIA GPU configurations for 1, 2, 4, and 8 GPUs. Optional profiles stream Safetensors from the platform object store or from a dedicated [model_storage] S3 endpoint.
Administrators curate the deployment choices exposed by the guided model setup through Portal Model Presets.
Model workloads have no public route. Requests pass through the BSQAI API and LiteLLM before reaching the internal KServe gateway.
Interacts With
BSQAI API, which creates deployments and exposes workload health, pod, and log operations.LiteLLM, which routes authenticated model requests to the private inference path.vLLM, which serves LLM workloads through platform runtime profiles.KAI Scheduler, which places and quotas GPU workloads.Envoy AI Gateway, which provides the internal model-serving data path.Rook-Cephor external S3-compatible storage, which supplies model artifacts.Grafana AlloyandGrafana, which collect and display inference telemetry.