vLLM

High-throughput LLM inference engine in the Foundation layer.

Agentic Friendly

Component Category

Inference / LLM serving engine

Component Description

vLLM is a high-throughput and memory-efficient inference engine designed for serving large language models.

Why It Is Used

In BullSequana AI Foundation, vLLM powers efficient LLM inference with strong performance characteristics for production workloads, especially where throughput and GPU utilization matter.

Learn More

Deployment notes

vLLM is not deployed as a standalone platform service. KServe creates vLLM workloads from LLMInferenceService resources and operator-managed runtime profiles. BullSequana AI 1.3.0 ships vLLM 0.28 GPU and 0.28.0-x86_64 CPU images, with profiles for CPU and 1, 2, 4, or 8 NVIDIA GPUs.

Some profiles can stream Safetensors directly from S3-compatible model storage. Inference remains private behind LiteLLM and the internal inference gateway.

Interacts With

  • KServe, which creates and manages vLLM model-serving workloads.
  • Model Installer and other serving workflows, which deploy or operate models on top of vLLM-backed runtimes.

On this page