vLLM
High-throughput LLM inference engine in the Foundation layer.
Component Category
Inference / LLM serving engine
Component Description
vLLM is a high-throughput and memory-efficient inference engine designed for serving large language models.
Why It Is Used
In BullSequana AI Foundation, vLLM powers efficient LLM inference with strong performance characteristics for production workloads, especially where throughput and GPU utilization matter.
Learn More
Deployment notes
vLLM is not deployed as a standalone platform service. KServe creates vLLM workloads from LLMInferenceService resources and operator-managed runtime profiles. BullSequana AI 1.3.0 ships vLLM 0.28 GPU and 0.28.0-x86_64 CPU images, with profiles for CPU and 1, 2, 4, or 8 NVIDIA GPUs.
Some profiles can stream Safetensors directly from S3-compatible model storage. Inference remains private behind LiteLLM and the internal inference gateway.
Interacts With
KServe, which creates and manages vLLM model-serving workloads.Model Installerand other serving workflows, which deploy or operate models on top of vLLM-backed runtimes.