Milvus

The vector database used for retrieval and semantic search in AI.

Agentic Friendly

Component Category

Vector database

Component Description

Milvus is the vector database used by AI for storing embeddings and supporting semantic retrieval workflows. It is built for large-scale vector search and similarity matching across embedding-driven applications. BullSequana AI uses one shared Milvus service with platform-managed collections scoped to individual tenants.

Why It Is Used

In BullSequana AI, Milvus provides the vector storage and similarity search capabilities needed for retrieval-augmented generation and other embedding-based experiences. Tenant-scoped collection names let the BSQAI API enforce the active tenant boundary before it reads, searches, or changes vector data.

Learn More

Tenant isolation

The BSQAI API derives each collection name from the canonical tenant UUID, embedding model, and embedding dimension. Ingestion ensures that the derived collection exists, while a reindex creates a new target collection and promotes it only after the reindex completes successfully.

Collection operations expose only collections in the active tenant's namespace. Search uses the active collection recorded for that tenant, so a foreign or inactive collection identifier is returned as not found.

Deployment notes

BullSequana AI 1.3.0 deploys Milvus 2.5.16 in distributed mode. A three-member etcd cluster stores coordination metadata, Kafka provides the message queue, and the configured S3-compatible object store persists Milvus data. The in-cluster storage path uses the milvus bucket in Rook-Ceph RGW; deployments using external object storage use the configured provider instead.

Interacts With

  • BSQAI API, for vector queries and retrieval logic.
  • Kafka (Strimzi), which provides the message queue backend for Milvus (3 replicas).
  • etcd, which stores Milvus coordination metadata in a three-member cluster.
  • Rook-Ceph or external S3-compatible storage, which persists Milvus data.
  • OAuth2 Proxy, which protects the Attu administration UI with Keycloak SSO.
  • Docling, whose document processing outputs feed vector embedding workflows.
  • LiteLLM, for embedding-producing model flows upstream of storage.
  • AI Web Portal, where administrators inspect platform-managed collections for the selected tenant.

On this page