Use Local Models via API
How developers should connect to BullSequana AI models from local tools and applications
When developing against BullSequana AI, the preferred integration point is the BSQAI API.
Do not integrate directly with LiteLLM for application development. LiteLLM is an internal platform component and may change over time. The BSQAI API is the stable interface that BullSequana AI exposes to developers.
Recommended Background
This page is primarily for application developers and AI engineers integrating external tools or custom applications with the platform.
Helpful experience includes:
- API-based application development
- bearer-token or API-key authentication
- local development with Docker when reproducing platform-adjacent flows
- basic understanding of OpenAI-compatible clients and SDKs
What To Use
Use the API endpoint exposed by your BullSequana AI deployment:
export BSQAI_BASE_URL="https://api.<platform-domain>/v1"BullSequana AI 1.3.0 uses api as the default public hostname prefix. A deployment can expose a site-specific hostname, so use the endpoint supplied by the platform operator when it differs from this example.
This API already handles the platform integration behind the scenes:
- model routing
- authentication
- platform policy enforcement
- future component evolution behind the API boundary
Authentication Options
The backend supports both of these:
JWT bearer tokensBullSequana API keysin thesk-bsq-...format
For interactive user access, JWT is a good fit.
For local developer tools, scripts, IDE plugins, and long-lived integrations, API keys are usually the cleaner option when your deployment exposes API key management.
How to pass credentials
API keys can be sent in either of these ways:
Authorization: Bearer sk-bsq-...— the same header used for JWTsX-Api-Key: sk-bsq-...— a dedicated API-key header
Both are accepted. JWT tokens always use Authorization: Bearer <jwt-token>.
When using OpenAI-compatible SDKs or tools, the Authorization: Bearer form is the most convenient because most clients already send the configured key in that header. The X-Api-Key header is available as an alternative when the Authorization header is reserved for another purpose or when you want an explicit separation between JWT and API-key authentication.
Requires platform 1.2.1 or later for the Bearer form
On platform versions before 1.2.1, API keys sent as Authorization: Bearer are rejected at the gateway with Jwt is not in the form of Header.Payload.Signature — the request never reaches the BSQAI API. On those versions, use the X-Api-Key header instead.
Option 1: Authenticate with JWT
If your platform uses Keycloak-backed interactive access, obtain a JWT and pass it as the bearer token.
Example:
export BSQAI_TOKEN="<jwt-token>"
curl "$BSQAI_BASE_URL/models" \
-H "Authorization: Bearer $BSQAI_TOKEN"Select a tenant for JWT requests
The Keycloak organization claim contains the tenants available to the user. When the user belongs to more than one organization, select one tenant by sending its Organization UUID:
export BSQAI_TENANT_ID="<keycloak-organization-uuid>"
curl "$BSQAI_BASE_URL/models" \
-H "Authorization: Bearer $BSQAI_TOKEN" \
-H "X-Tenant-Id: $BSQAI_TENANT_ID"The header is optional when the JWT has exactly one tenant membership. Browser WebSocket clients use the tenant_id query parameter because they cannot set custom headers during the handshake.
In local backend development, the sibling coreai-llm-backend repo already includes a helper script:
./scripts/get_jwt_token.sh --print-tokenOption 2: Create and use an API key
API keys are created through the BSQAI API itself and are returned only once when created.
First call the API with a valid JWT:
curl -X POST "https://api.<platform-domain>/v1/api-keys" \
-H "Authorization: Bearer <jwt-token>" \
-H "Content-Type: application/json" \
-d '{
"name": "Local Development"
}'Typical response shape:
{
"id": "a37e6b2e-728e-4ffc-aea4-e5415616a96a",
"name": "Local Development",
"key": "sk-bsq-v1-...",
"masked": "sk-bsq-v1-********-e5b4a3",
"created_at": "2024-01-15T10:30:00Z"
}Then use that key for requests. Both header styles are accepted:
export BSQAI_API_KEY="sk-bsq-v1-..."
# Option A: Bearer header (works with most OpenAI-compatible SDKs)
curl "$BSQAI_BASE_URL/models" \
-H "Authorization: Bearer $BSQAI_API_KEY"
# Option B: Dedicated API-key header
curl "$BSQAI_BASE_URL/models" \
-H "X-Api-Key: $BSQAI_API_KEY"An API key is bound to the active tenant at creation time. It cannot be reused to access a different tenant.
Discover Available Models
Before wiring an application, check which models your deployment exposes:
Available models are also visible in the Portal.
If you want the API-level source of truth for the current environment, use:
curl "$BSQAI_BASE_URL/models" \
-H "Authorization: Bearer $BSQAI_API_KEY"This is the correct source of truth for model names in your environment.
Use the OpenAI-Compatible API
The BSQAI API implements the OpenAI Responses API:
/v1/models/v1/responses
That means SDKs and tools that speak the Responses API can be pointed to BullSequana AI with only a base URL and token change. The official OpenAI SDKs use the Responses API by default.
The legacy Chat Completions endpoint (/v1/chat/completions) is not available. Clients that only implement Chat Completions receive 404 Not Found; they must be configured for (or updated to support) the Responses API. Note that the Responses API takes an input field, not a messages array — sending messages returns a user_message_missing error.
Responses API example (curl)
curl -X POST "$BSQAI_BASE_URL/responses" \
-H "Authorization: Bearer $BSQAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "<model-name-from-/v1/models>",
"input": [
{ "role": "user", "content": "Hello from BullSequana AI" }
]
}'Responses API example (Python)
from openai import OpenAI
client = OpenAI(
api_key="sk-bsq-v1-...",
base_url="https://api.<platform-domain>/v1",
)
response = client.responses.create(
model="<model-name-from-/v1/models>",
input="Explain the deployment model of BullSequana AI."
)
print(response)Inspect AI-output provenance
Generative HTTP responses include headers that identify AI-generated output:
X-AI-Generated: true
X-AI-Provenance: model="<resolved-model>", generated-at="<timestamp>", producer="coreai-llm-backend"The headers are present on Responses API output, Anthropic-compatible Messages, portal chat responses, and chat-history reads that contain generated output. Streaming responses send them before the first event. Error responses and non-generative endpoints do not include them.
IDE And Tooling Integrations
For tools such as OpenCode or custom internal applications, point the tool to the BSQAI API, not to LiteLLM directly.
The correct pattern is:
- provider type: OpenAI (Responses API)
- base URL: your BSQAI API base URL
- bearer token: JWT or
sk-bsq-...API key
The tool must use the OpenAI Responses API for the platform's model names. Continue, for example, only uses the Responses API for OpenAI's own model naming (o-series, GPT-5+) and routes all other model names to Chat Completions, so it cannot currently connect to the BSQAI API (see Continue).
Recommended Integration Rule
Use this rule for all developer-facing integrations:
Application or tool -> BSQAI API -> platform components
Not:
Application or tool -> LiteLLM
That keeps the developer contract stable even if the platform team changes the internal inference or proxy layer later.