BSQAI API

The tenant-aware backend API for AI applications, models, tools, speech, and knowledge workflows.

Agentic Friendly

Component Category

API backend for GenAI, agents, and tenant administration

Component Description

BSQAI API is the main backend service for AI applications. It integrates identity, tenant selection, model access, retrieval, files, storage, tools, speech, workflows, and observability behind product-facing HTTP and WebSocket APIs. Model installation is part of the same backend and deploys models through KServe.

Why It Is Used

In BullSequana AI, the BSQAI API gives applications one authenticated and authorized integration boundary. Clients do not need direct contracts with LiteLLM, KServe, Docling, Milvus, or other internal services.

Learn More

API surfaces

The default public endpoint is https://api.<platform-domain>. The service exposes:

  • OpenAI-compatible Responses and Models routes under /v1.
  • Anthropic-compatible Messages routes under /anthropic.
  • tenant administration under /v1/admin/tenants.
  • model installation and download tracking under /v1/model-installer.
  • KServe workload health and logs under /v1alpha/kserve-runtime.
  • MCP Server, credential, OAuth, and gateway routes.
  • web-search provider configuration and search execution.
  • batch and real-time speech-to-text and real-time text-to-speech.

JWT users with multiple Keycloak Organization memberships select a tenant with X-Tenant-Id. API keys are bound to one tenant when created.

AI-output provenance

Generative HTTP responses include X-AI-Generated and X-AI-Provenance. The structured provenance value records the resolved model, generation time, and backend producer. Browsers can read both headers through CORS.

Model installation

Model registration creates or updates a KServe LLMInferenceService and registers the corresponding model route in LiteLLM. runtime_config_name selects a platform-managed CPU, GPU, or direct-S3-streaming runtime profile. Supported source schemes are hf://, pvc://, and s3://.

See Model Installer API.

Interacts With

  • Keycloak, for authentication, Organization membership, and service identities.
  • OpenFGA, for tenant-scoped authorization and sharing decisions.
  • LiteLLM, for governed model routing.
  • KServe and vLLM, for model deployment and execution.
  • MCP Gateway and MCP Lifecycle Operator, for governed tool calls and managed MCP Server workloads.
  • Milvus, for vector storage and semantic retrieval.
  • Docling, for document conversion and extraction workflows.
  • MLflow, for model metadata and GenAI traces.
  • Rook-Ceph or external S3-compatible storage, for files, models, and artifacts.
  • PostgreSQL, for tenant-scoped application state.
  • Temporal, for tenant lifecycle, file processing, and other durable workflows.
  • OpenTelemetry, for application traces and metrics.

On this page