Model Installer API
Register model sources and deploy them as KServe LLM inference services.
The Model Installer API turns an approved model location into a platform-managed inference deployment. It is part of the BSQAI API and is available under /v1/model-installer.
What registration does
POST /v1/model-installer/register_model:
- validates the source and runtime settings;
- creates or updates a KServe
LLMInferenceServicein thekservenamespace; - waits for the model-serving workload to become available; and
- registers the model route in LiteLLM for use through the BSQAI API.
Applications should use the BSQAI API rather than call the private KServe endpoint directly.
Supported model locations
The url field accepts these source schemes:
| Scheme | Use |
|---|---|
hf://<organization>/<model> | Download model artifacts from Hugging Face |
pvc://<claim>/<optional-path> | Use artifacts already stored on a persistent volume |
s3://<bucket>/<path> | Read artifacts from configured S3-compatible model storage |
Other URL schemes are rejected by the current registration API.
Runtime selection
Use runtime_config_name to select a platform-managed LLMInferenceServiceConfig. BullSequana AI ships CPU and NVIDIA GPU profiles for 1, 2, 4, and 8 GPUs. Sites can also expose profiles that stream Safetensors directly from the platform object store or from the optional [model_storage] endpoint.
The legacy resourceProfile field remains available as a compatibility mapping for existing clients. New integrations should use runtime_config_name.
targetRequests controls the target request concurrency used by KServe autoscaling. The platform operator determines which profiles and limits are available at a site.
Register a model
The exact optional fields depend on the selected engine and runtime profile. A typical vLLM request is:
{
"name": "mistral-small",
"namespace": "kserve",
"engine": "VLLM",
"features": ["TextGeneration"],
"url": "hf://mistralai/Mistral-Small-3.2-24B-Instruct-2506",
"runtime_config_name": "vllm-gpu-1",
"minReplicas": 0,
"maxReplicas": 2,
"targetRequests": 1,
"model_mode": "chat",
"max_input_tokens": 8192,
"max_output_tokens": 2048
}Use the runtime profile names supplied by your platform operator; they can differ from the example.
Repository imports
The same API includes asynchronous import and repository operations:
| Endpoint | Purpose |
|---|---|
POST /v1/model-installer/download_hf_model | Import a Hugging Face model into the platform repository |
POST /v1/model-installer/download_s3_model | Copy a model from S3-compatible storage into the repository |
POST /v1/model-installer/register_s3_model | Register an existing S3 artifact in MLflow without copying it |
GET /v1/model-installer/downloads | List background downloads |
GET /v1/model-installer/downloads/{download_id}/status | Read download state |
GET /v1/model-installer/downloads/{download_id}/logs | Read download logs |
DELETE /v1/model-installer/delete_model | Delete a repository model |
DELETE /v1/model-installer/unregister_model/{name}/{namespace} | Remove an inference deployment and route |
Import first when the model needs a platform-managed MLflow artifact and version history. Register directly when the source URL is already the desired serving source.
Portal flow
The Portal Models page uses the same endpoints. Easy Setup applies an operator-managed preset; Advanced Setup exposes the source, KServe runtime profile, scaling, timeouts, and model metadata. Both flows deploy to kserve and show the generated LLMInferenceService configuration before submission.
Long-running imports appear in the model-download status UI, where users can inspect progress, logs, and retry failed work.
Authentication and tenant scope
Model-management endpoints require an authenticated user with the relevant permission. When a JWT belongs to more than one Keycloak Organization, include X-Tenant-Id. API keys already carry their tenant binding.
Verify a deployment
Use the BSQAI API KServe runtime routes to inspect the resulting workload:
GET /v1alpha/kserve-runtime/health/{namespace}/{name}
GET /v1alpha/kserve-runtime/pods/{namespace}/{name}
GET /v1alpha/kserve-runtime/pods/{namespace}/{name}/{pod}/logs
GET /v1alpha/kserve-runtime/pods/{namespace}/{name}/{pod}/describeCluster operators can also inspect the Kubernetes resource directly:
kubectl get llminferenceservice -n kserve
kubectl describe llminferenceservice <model-name> -n kserve