# BSQAI API (/docs/ai/components/bsqai-api)



## Component Category [#component-category]

API backend for GenAI, agents, and tenant administration

## Component Description [#component-description]

BSQAI API is the main backend service for AI applications. It integrates identity, tenant selection, model access, retrieval, files, storage, tools, speech, workflows, and observability behind product-facing HTTP and WebSocket APIs. Model installation is part of the same backend and deploys models through KServe.

## Why It Is Used [#why-it-is-used]

In BullSequana AI, the BSQAI API gives applications one authenticated and authorized integration boundary. Clients do not need direct contracts with LiteLLM, KServe, Docling, Milvus, or other internal services.

## Learn More [#learn-more]

* [API reference](/docs/reference)
* [AI reference architecture](/docs/ai/reference-architecture)

## API surfaces [#api-surfaces]

The default public endpoint is `https://api.<platform-domain>`. The service exposes:

* OpenAI-compatible Responses and Models routes under `/v1`.
* Anthropic-compatible Messages routes under `/anthropic`.
* tenant administration under `/v1/admin/tenants`.
* model installation and download tracking under `/v1/model-installer`.
* KServe workload health and logs under `/v1alpha/kserve-runtime`.
* MCP Server, credential, OAuth, and gateway routes.
* web-search provider configuration and search execution.
* batch and real-time speech-to-text and real-time text-to-speech.

JWT users with multiple Keycloak Organization memberships select a tenant with `X-Tenant-Id`. API keys are bound to one tenant when created.

## AI-output provenance [#ai-output-provenance]

Generative HTTP responses include `X-AI-Generated` and `X-AI-Provenance`. The structured provenance value records the resolved model, generation time, and backend producer. Browsers can read both headers through CORS.

## Model installation [#model-installation]

Model registration creates or updates a KServe `LLMInferenceService` and registers the corresponding model route in LiteLLM. `runtime_config_name` selects a platform-managed CPU, GPU, or direct-S3-streaming runtime profile. Supported source schemes are `hf://`, `pvc://`, and `s3://`.

See [Model Installer API](/docs/ai/model-as-a-service/model-installer-api).

## Interacts With [#interacts-with]

* `Keycloak`, for authentication, Organization membership, and service identities.
* `OpenFGA`, for tenant-scoped authorization and sharing decisions.
* `LiteLLM`, for governed model routing.
* `KServe` and `vLLM`, for model deployment and execution.
* `MCP Gateway` and `MCP Lifecycle Operator`, for governed tool calls and managed MCP Server workloads.
* `Milvus`, for vector storage and semantic retrieval.
* `Docling`, for document conversion and extraction workflows.
* `MLflow`, for model metadata and GenAI traces.
* `Rook-Ceph` or external S3-compatible storage, for files, models, and artifacts.
* `PostgreSQL`, for tenant-scoped application state.
* `Temporal`, for tenant lifecycle, file processing, and other durable workflows.
* `OpenTelemetry`, for application traces and metrics.
