BSQAI API
The tenant-aware backend API for AI applications, models, tools, speech, and knowledge workflows.
Component Category
API backend for GenAI, agents, and tenant administration
Component Description
BSQAI API is the main backend service for AI applications. It integrates identity, tenant selection, model access, retrieval, files, storage, tools, speech, workflows, and observability behind product-facing HTTP and WebSocket APIs. Model installation is part of the same backend and deploys models through KServe.
Why It Is Used
In BullSequana AI, the BSQAI API gives applications one authenticated and authorized integration boundary. Clients do not need direct contracts with LiteLLM, KServe, Docling, Milvus, or other internal services.
Learn More
API surfaces
The default public endpoint is https://api.<platform-domain>. The service exposes:
- OpenAI-compatible Responses and Models routes under
/v1. - Anthropic-compatible Messages routes under
/anthropic. - tenant administration under
/v1/admin/tenants. - model installation and download tracking under
/v1/model-installer. - KServe workload health and logs under
/v1alpha/kserve-runtime. - MCP Server, credential, OAuth, and gateway routes.
- web-search provider configuration and search execution.
- batch and real-time speech-to-text and real-time text-to-speech.
JWT users with multiple Keycloak Organization memberships select a tenant with X-Tenant-Id. API keys are bound to one tenant when created.
AI-output provenance
Generative HTTP responses include X-AI-Generated and X-AI-Provenance. The structured provenance value records the resolved model, generation time, and backend producer. Browsers can read both headers through CORS.
Model installation
Model registration creates or updates a KServe LLMInferenceService and registers the corresponding model route in LiteLLM. runtime_config_name selects a platform-managed CPU, GPU, or direct-S3-streaming runtime profile. Supported source schemes are hf://, pvc://, and s3://.
See Model Installer API.
Interacts With
Keycloak, for authentication, Organization membership, and service identities.OpenFGA, for tenant-scoped authorization and sharing decisions.LiteLLM, for governed model routing.KServeandvLLM, for model deployment and execution.MCP GatewayandMCP Lifecycle Operator, for governed tool calls and managed MCP Server workloads.Milvus, for vector storage and semantic retrieval.Docling, for document conversion and extraction workflows.MLflow, for model metadata and GenAI traces.Rook-Cephor external S3-compatible storage, for files, models, and artifacts.PostgreSQL, for tenant-scoped application state.Temporal, for tenant lifecycle, file processing, and other durable workflows.OpenTelemetry, for application traces and metrics.