# vLLM (/docs/foundation/components/vllm)



## Component Category [#component-category]

Inference / LLM serving engine

## Component Description [#component-description]

vLLM is a high-throughput and memory-efficient inference engine designed for serving large language models.

## Why It Is Used [#why-it-is-used]

In BullSequana AI Foundation, vLLM powers efficient LLM inference with strong performance characteristics for production workloads, especially where throughput and GPU utilization matter.

## Learn More [#learn-more]

* [vLLM documentation](https://docs.vllm.ai/en/stable/)
* [vllm-project/vllm on GitHub](https://github.com/vllm-project/vllm)

## Deployment notes [#deployment-notes]

vLLM is not deployed as a standalone platform service. KServe creates vLLM workloads from `LLMInferenceService` resources and operator-managed runtime profiles. BullSequana AI 1.3.0 ships vLLM 0.28 GPU and 0.28.0-x86\_64 CPU images, with profiles for CPU and 1, 2, 4, or 8 NVIDIA GPUs.

Some profiles can stream Safetensors directly from S3-compatible model storage. Inference remains private behind LiteLLM and the internal inference gateway.

## Interacts With [#interacts-with]

* `KServe`, which creates and manages vLLM model-serving workloads.
* `Model Installer` and other serving workflows, which deploy or operate models on top of vLLM-backed runtimes.
