Use Local Models via API

How developers should connect to BullSequana AI models from local tools and applications

Agentic Friendly

When developing against BullSequana AI, the preferred integration point is the BSQAI API.

Do not integrate directly with LiteLLM for application development. LiteLLM is an internal platform component and may change over time. The BSQAI API is the stable interface that BullSequana AI exposes to developers.

This page is primarily for application developers and AI engineers integrating external tools or custom applications with the platform.

Helpful experience includes:

  • API-based application development
  • bearer-token or API-key authentication
  • local development with Docker when reproducing platform-adjacent flows
  • basic understanding of OpenAI-compatible clients and SDKs

What To Use

Use the API endpoint exposed by your BullSequana AI deployment:

export BSQAI_BASE_URL="https://api.<platform-domain>/v1"

BullSequana AI 1.3.0 uses api as the default public hostname prefix. A deployment can expose a site-specific hostname, so use the endpoint supplied by the platform operator when it differs from this example.

This API already handles the platform integration behind the scenes:

  • model routing
  • authentication
  • platform policy enforcement
  • future component evolution behind the API boundary

Authentication Options

The backend supports both of these:

  • JWT bearer tokens
  • BullSequana API keys in the sk-bsq-... format

For interactive user access, JWT is a good fit.

For local developer tools, scripts, IDE plugins, and long-lived integrations, API keys are usually the cleaner option when your deployment exposes API key management.

How to pass credentials

API keys can be sent in either of these ways:

  • Authorization: Bearer sk-bsq-... — the same header used for JWTs
  • X-Api-Key: sk-bsq-... — a dedicated API-key header

Both are accepted. JWT tokens always use Authorization: Bearer <jwt-token>.

When using OpenAI-compatible SDKs or tools, the Authorization: Bearer form is the most convenient because most clients already send the configured key in that header. The X-Api-Key header is available as an alternative when the Authorization header is reserved for another purpose or when you want an explicit separation between JWT and API-key authentication.

Requires platform 1.2.1 or later for the Bearer form

On platform versions before 1.2.1, API keys sent as Authorization: Bearer are rejected at the gateway with Jwt is not in the form of Header.Payload.Signature — the request never reaches the BSQAI API. On those versions, use the X-Api-Key header instead.

Option 1: Authenticate with JWT

If your platform uses Keycloak-backed interactive access, obtain a JWT and pass it as the bearer token.

Example:

export BSQAI_TOKEN="<jwt-token>"

curl "$BSQAI_BASE_URL/models" \
  -H "Authorization: Bearer $BSQAI_TOKEN"

Select a tenant for JWT requests

The Keycloak organization claim contains the tenants available to the user. When the user belongs to more than one organization, select one tenant by sending its Organization UUID:

export BSQAI_TENANT_ID="<keycloak-organization-uuid>"

curl "$BSQAI_BASE_URL/models" \
  -H "Authorization: Bearer $BSQAI_TOKEN" \
  -H "X-Tenant-Id: $BSQAI_TENANT_ID"

The header is optional when the JWT has exactly one tenant membership. Browser WebSocket clients use the tenant_id query parameter because they cannot set custom headers during the handshake.

In local backend development, the sibling coreai-llm-backend repo already includes a helper script:

./scripts/get_jwt_token.sh --print-token

Option 2: Create and use an API key

API keys are created through the BSQAI API itself and are returned only once when created.

First call the API with a valid JWT:

curl -X POST "https://api.<platform-domain>/v1/api-keys" \
  -H "Authorization: Bearer <jwt-token>" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "Local Development"
  }'

Typical response shape:

{
  "id": "a37e6b2e-728e-4ffc-aea4-e5415616a96a",
  "name": "Local Development",
  "key": "sk-bsq-v1-...",
  "masked": "sk-bsq-v1-********-e5b4a3",
  "created_at": "2024-01-15T10:30:00Z"
}

Then use that key for requests. Both header styles are accepted:

export BSQAI_API_KEY="sk-bsq-v1-..."

# Option A: Bearer header (works with most OpenAI-compatible SDKs)
curl "$BSQAI_BASE_URL/models" \
  -H "Authorization: Bearer $BSQAI_API_KEY"

# Option B: Dedicated API-key header
curl "$BSQAI_BASE_URL/models" \
  -H "X-Api-Key: $BSQAI_API_KEY"

An API key is bound to the active tenant at creation time. It cannot be reused to access a different tenant.

Discover Available Models

Before wiring an application, check which models your deployment exposes:

Available models are also visible in the Portal.

If you want the API-level source of truth for the current environment, use:

curl "$BSQAI_BASE_URL/models" \
  -H "Authorization: Bearer $BSQAI_API_KEY"

This is the correct source of truth for model names in your environment.

Use the OpenAI-Compatible API

The BSQAI API implements the OpenAI Responses API:

  • /v1/models
  • /v1/responses

That means SDKs and tools that speak the Responses API can be pointed to BullSequana AI with only a base URL and token change. The official OpenAI SDKs use the Responses API by default.

The legacy Chat Completions endpoint (/v1/chat/completions) is not available. Clients that only implement Chat Completions receive 404 Not Found; they must be configured for (or updated to support) the Responses API. Note that the Responses API takes an input field, not a messages array — sending messages returns a user_message_missing error.

Responses API example (curl)

curl -X POST "$BSQAI_BASE_URL/responses" \
  -H "Authorization: Bearer $BSQAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "<model-name-from-/v1/models>",
    "input": [
      { "role": "user", "content": "Hello from BullSequana AI" }
    ]
  }'

Responses API example (Python)

from openai import OpenAI

client = OpenAI(
    api_key="sk-bsq-v1-...",
    base_url="https://api.<platform-domain>/v1",
)

response = client.responses.create(
    model="<model-name-from-/v1/models>",
    input="Explain the deployment model of BullSequana AI."
)

print(response)

Inspect AI-output provenance

Generative HTTP responses include headers that identify AI-generated output:

X-AI-Generated: true
X-AI-Provenance: model="<resolved-model>", generated-at="<timestamp>", producer="coreai-llm-backend"

The headers are present on Responses API output, Anthropic-compatible Messages, portal chat responses, and chat-history reads that contain generated output. Streaming responses send them before the first event. Error responses and non-generative endpoints do not include them.

IDE And Tooling Integrations

For tools such as OpenCode or custom internal applications, point the tool to the BSQAI API, not to LiteLLM directly.

The correct pattern is:

  • provider type: OpenAI (Responses API)
  • base URL: your BSQAI API base URL
  • bearer token: JWT or sk-bsq-... API key

The tool must use the OpenAI Responses API for the platform's model names. Continue, for example, only uses the Responses API for OpenAI's own model naming (o-series, GPT-5+) and routes all other model names to Chat Completions, so it cannot currently connect to the BSQAI API (see Continue).

Use this rule for all developer-facing integrations:

Application or tool -> BSQAI API -> platform components

Not:

Application or tool -> LiteLLM

That keeps the developer contract stable even if the platform team changes the internal inference or proxy layer later.

On this page