# External S3 object storage (/docs/deployment/playbooks/external-s3-object-storage)



BullSequana AI 1.3.0 can use a customer-managed S3-compatible service instead of the Rook-Ceph object gateway. The platform renders the endpoint, CA trust, and per-component Kubernetes secrets, while the storage operator owns the accounts, buckets, access policies, availability, and recovery process.

This playbook covers the primary `[object_storage]` provider. It does not require the optional `[model_storage]` endpoint and does not migrate existing objects automatically.

The storage operator provides the external-storage configuration through one approved handoff. A BullSequana AI deployment operator maps that information into the protected site configuration and renders the sealed Kubernetes secrets. Portal users and tenant administrators do not configure these credentials.

## Understand the storage boundaries [#understand-the-storage-boundaries]

| Configuration               | Purpose                                                                                             | Required |
| --------------------------- | --------------------------------------------------------------------------------------------------- | -------- |
| `[object_storage]`          | Primary store for files, model artifacts, logs, traces, backups, pipelines, and other platform data | Yes      |
| `[model_storage]`           | Separate, read-only-capable repository used by the direct-S3 KServe runtime profile                 | No       |
| RWO and RWX storage classes | Kubernetes persistent volumes for databases and stateful workloads                                  | Yes      |

Selecting `external-s3` replaces the platform's object-store endpoint. It does not provide block or shared-file storage. Before disabling the Ceph components, configure working RWO and RWX storage classes from another supported storage provider.

## 1. Confirm provider compatibility [#1-confirm-provider-compatibility]

Collect and test the following information with the storage operator:

* a private endpoint reachable from cluster workloads
* a browser-reachable DNS name for presigned downloads
* the endpoint scheme, explicit port, region, and addressing style
* the CA certificate when HTTPS uses a private trust chain
* support for the required bucket, object, multipart-upload, and lifecycle operations
* the provider's server-side encryption and versioning behavior
* how one service account can access a bucket owned by another account
* the provider's presigned-request handling, especially `Content-Type`

The public host is supplied without a scheme. BullSequana AI constructs browser-facing URLs with HTTPS, so its certificate and DNS name must be valid for users as well as reachable from any workload that uses that URL.

## 2. Prepare accounts and buckets [#2-prepare-accounts-and-buckets]

The named application buckets and their credentials must exist before rendering the platform. The storage operator includes their account-to-bucket mapping in the site handoff. The platform does not create them through Rook `ObjectBucketClaim` resources.

The default deployment uses these buckets when all corresponding components are enabled:

| Consumer            | Default bucket or prefix                  | Purpose                                     |
| ------------------- | ----------------------------------------- | ------------------------------------------- |
| Grafana Loki        | `loki-chunks`, `loki-ruler`, `loki-admin` | logs, rules, and Loki administration data   |
| Grafana Tempo       | `tempo-traces`                            | distributed traces                          |
| CloudNativePG       | `cnpg-backups`                            | base backups and WAL archives under `cnpg/` |
| Milvus              | `milvus`                                  | vector-store data                           |
| MLflow              | `mlflow`                                  | model and run artifacts                     |
| BSQAI API           | `default`, `default-artifacts`            | application files and artifacts             |
| KServe              | `kserve`                                  | model-serving artifacts                     |
| Airbyte integration | `airbyte`                                 | configured S3 integration data path         |
| Argo Events         | `argo-events`                             | event payloads and artifacts                |
| Kubeflow Pipelines  | `kubeflow-pipelines`                      | central pipeline artifacts                  |
| Lakekeeper          | `lakekeeper`                              | governed table and warehouse objects        |

Kubeflow can also create tenant-derived `user-*` and `space-*` buckets. Its storage account needs the provider permissions required to create and manage those prefixed buckets when per-profile storage is enabled.

Use a dedicated account for each enabled consumer where the provider supports it. This limits the impact of a compromised credential and allows each account to receive only the bucket and object operations that its component needs.

A provider can instead supply one shared account for several or all consumers. That account must have access to every required bucket. The deployment operator maps the same pair into the applicable component credential slots. This is supported, but it creates a broader access boundary.

### Shared MLflow access [#shared-mlflow-access]

The `mlflow` bucket crosses component boundaries: MLflow owns the artifact store, the BSQAI API and Model Installer write model artifacts, and KServe reads them during model loading. Grant that access on the storage provider before deployment.

Set `[object_storage.capabilities] cross_account` to describe how the provider enforces this boundary:

* `policy` — bucket policies grant access to the other accounts
* `posix` — the service also enforces filesystem ownership or group access
* `none` — cross-account sharing is not available

With `posix`, a bucket policy alone might not be sufficient. Configure the provider's filesystem ownership and group permissions as well. With `none`, use a reviewed provider-side alternative, such as one credential shared only by the required consumers.

## 3. Configure the endpoint and capabilities [#3-configure-the-endpoint-and-capabilities]

Put site-specific values in the platform environment file rather than changing the shipped defaults. This example uses HTTPS and a provider that supports bucket policies, SSE-S3, and versioning:

```bash
OBJECT_STORAGE_PROVIDER=external-s3
OBJECT_STORAGE_SCHEME=https
OBJECT_STORAGE_HOST=s3.internal.example.com
OBJECT_STORAGE_PORT=443
OBJECT_STORAGE_REGION=eu-west-1
OBJECT_STORAGE_PATH_STYLE=true
OBJECT_STORAGE_PUBLIC_HOST=s3.example.com

# Leave empty when the endpoint chains to a public CA already trusted by the platform.
OBJECT_STORAGE_CA_PEM="<PEM-encoded private CA certificate>"

OBJECT_STORAGE_CAPABILITIES_SSE=true
OBJECT_STORAGE_CAPABILITIES_PRESIGN_SIGNS_CONTENT_TYPE=false
OBJECT_STORAGE_CAPABILITIES_CROSS_ACCOUNT=policy
OBJECT_STORAGE_CAPABILITIES_STS_FEDERATION=false
OBJECT_STORAGE_CAPABILITIES_STS=false
OBJECT_STORAGE_CAPABILITIES_VERSIONING=true
OBJECT_STORAGE_CAPABILITIES_MANAGE_BUCKET_CORS=false
```

Set the capability values to match the service that was tested; they are not preference flags.

| Capability                   | Meaning for an external provider                                                                                 |
| ---------------------------- | ---------------------------------------------------------------------------------------------------------------- |
| `sse`                        | Enables SSE-S3 requests only when the provider supports them                                                     |
| `presign_signs_content_type` | Records whether the provider validates `Content-Type` on presigned uploads                                       |
| `cross_account`              | Selects `policy`, `posix`, or `none` for shared-bucket access                                                    |
| `sts_federation`             | Must be `false`; Keycloak-to-RGW federation is specific to Rook-Ceph                                             |
| `sts`                        | Keep `false` unless the provider is verified to intersect an assumed-role session policy with the account policy |
| `versioning`                 | Enables behavior that depends on bucket versioning only when the provider supports it                            |
| `manage_bucket_cors`         | Allows BullSequana AI to replace the bucket's complete CORS document                                             |

Use `OBJECT_STORAGE_PATH_STYLE=true` when the provider expects `https://host/bucket/key`. Use `false` only when wildcard DNS and certificates support virtual-hosted requests such as `https://bucket.host/key`.

## 4. Hand off storage credentials [#4-hand-off-storage-credentials]

The storage operator provides one site bundle through the approved secret-handoff process. Include:

* each provider-issued access-key and secret-key pair
* the account-to-bucket mapping and allowed operations
* the account that can write model artifacts to the shared `mlflow` bucket
* any additional policy or filesystem-group assignments used for cross-account access

The BullSequana AI deployment operator stores this bundle through the site's secret-management process and maps it into the protected platform environment. The platform currently uses separate configuration slots because it renders a Kubernetes Secret for each component. It does not require every slot to contain a different provider account.

Components that are disabled need no credential mapping. For the exact deployment variables, see [Credential mapping reference](#credential-mapping-reference).

## 5. Configure browser CORS [#5-configure-browser-cors]

Presigned file downloads originate from the portal in the user's browser. Each affected external bucket must allow `GET` and `HEAD` from the portal origin.

A provider-neutral policy has this shape:

```json
{
  "CORSRules": [
    {
      "AllowedOrigins": ["https://portal.example.com"],
      "AllowedMethods": ["GET", "HEAD"],
      "AllowedHeaders": ["*"],
      "ExposeHeaders": ["ETag"],
      "MaxAgeSeconds": 3600
    }
  ]
}
```

Prefer `OBJECT_STORAGE_CAPABILITIES_MANAGE_BUCKET_CORS=false` and let the storage operator own this policy. Set it to `true` only when the storage owner explicitly delegates the complete CORS document to BullSequana AI. `PutBucketCors` replaces the existing document rather than merging a rule into it.

## 6. Disable unused Ceph components [#6-disable-unused-ceph-components]

When the site does not use Rook-Ceph for object, block, or shared-file storage, disable its components:

```bash
ROOK_CEPH_COMPONENT_ENABLED=false
ROOK_CEPH_CLUSTER_COMPONENT_ENABLED=false
CEPH_CSI_DRIVERS_COMPONENT_ENABLED=false
```

Do not disable the CSI drivers until replacement RWO and RWX storage classes are configured and tested. Selecting `external-s3` alone does not create those classes.

## 7. Validate and render [#7-validate-and-render]

Before rendering, test each writable account against every bucket it is expected to use. When the provider supplies one shared account, repeat the bucket checks with that account:

```bash
export AWS_ACCESS_KEY_ID=<component-access-key>
export AWS_SECRET_ACCESS_KEY=<component-secret-key>
export AWS_DEFAULT_REGION=eu-west-1
export S3_ENDPOINT=https://s3.internal.example.com:443

aws --endpoint-url "$S3_ENDPOINT" s3api head-bucket --bucket <component-bucket>
aws --endpoint-url "$S3_ENDPOINT" s3 cp ./s3-smoke-test.txt "s3://<component-bucket>/s3-smoke-test.txt"
aws --endpoint-url "$S3_ENDPOINT" s3 cp "s3://<component-bucket>/s3-smoke-test.txt" ./s3-smoke-test.downloaded.txt
aws --endpoint-url "$S3_ENDPOINT" s3 rm "s3://<component-bucket>/s3-smoke-test.txt"
```

For a read-only account, replace the write and delete operations with `head-object` and a download of a known test object. Also test the MLflow bucket with the BSQAI API/Model Installer writer and the KServe reader credentials.

Render the complete desired state:

```bash
bsqai apply all -r
```

Rendering fails closed when the provider, host, scheme, cross-account mode, STS federation setting, or required component credentials are inconsistent. Treat those errors as configuration failures; do not replace missing external credentials with generated placeholder values.

Review the generated manifests before pushing them. Confirm that:

* endpoints include the intended explicit port
* the public hostname is correct for presigned URLs
* private CA material is mounted where required
* external credentials are rendered as Sealed Secrets
* Rook bucket and user resources are absent for external-S3 consumers
* disabled Ceph applications are absent from the generated app-of-apps output

## 8. Verify after deployment [#8-verify-after-deployment]

Run an end-to-end check for every enabled storage path:

1. Upload and download a file through the portal, then verify the browser reports no CORS error.
2. Process a file for RAG and confirm that Milvus search returns it.
3. Create and verify a CloudNativePG backup in `cnpg-backups`.
4. Confirm Loki logs and Tempo traces continue to appear in Grafana.
5. Import or deploy a model and verify the MLflow writer and KServe reader paths.
6. Run a Kubeflow pipeline and confirm its artifacts, if Kubeflow is enabled.
7. Run an Airbyte S3 integration and an Argo Events artifact flow when those integrations are enabled.

Inspect provider audit logs for denied requests, unexpected bucket access, TLS failures, or requests sent to the public endpoint when the private endpoint was expected.

## Migration and rollback [#migration-and-rollback]

Changing `OBJECT_STORAGE_PROVIDER` changes where the platform looks for objects; it does not copy data from Rook-Ceph or rewrite existing object references.

Before switching an existing deployment:

1. stop or quiesce writers
2. copy every required bucket, object version, and metadata record
3. validate object counts, sizes, checksums, retention, and restore procedures
4. update the endpoint, capabilities, credentials, and Ceph component settings together
5. render and review the complete desired state
6. keep the previous store read-only until application-level validation succeeds

Rollback requires the previous endpoints, credentials, manifests, and object contents to remain available. Do not delete or reuse the old buckets as part of the initial cutover.

## Credential mapping reference [#credential-mapping-reference]

This mapping is performed by the BullSequana AI deployment operator. Set the pair for every enabled S3 consumer in the protected platform environment. A dedicated provider account is recommended for each consumer. When the provider supplies one shared account, map the same pair into each required slot.

```bash
GRAFANA_LOKI_S3_ACCESS_KEY=<access-key>
GRAFANA_LOKI_S3_SECRET_KEY=<secret-key>
TEMPO_S3_ACCESS_KEY=<access-key>
TEMPO_S3_SECRET_KEY=<secret-key>
MILVUS_S3_ACCESS_KEY=<access-key>
MILVUS_S3_SECRET_KEY=<secret-key>
MLFLOW_S3_ACCESS_KEY=<access-key>
MLFLOW_S3_SECRET_KEY=<secret-key>
KSERVE_RESOURCES_S3_ACCESS_KEY=<access-key>
KSERVE_RESOURCES_S3_SECRET_KEY=<secret-key>
AIRBYTE_S3_ACCESS_KEY=<access-key>
AIRBYTE_S3_SECRET_KEY=<secret-key>
ARGO_EVENTS_S3_ACCESS_KEY=<access-key>
ARGO_EVENTS_S3_SECRET_KEY=<secret-key>
KUBEFLOW_S3_ACCESS_KEY=<access-key>
KUBEFLOW_S3_SECRET_KEY=<secret-key>
LAKEKEEPER_S3_ACCESS_KEY=<access-key>
LAKEKEEPER_S3_SECRET_KEY=<secret-key>
LLM_BACKEND_S3_ACCESS_KEY=<access-key>
LLM_BACKEND_S3_SECRET_KEY=<secret-key>
CNPG_BACKUPS_ACCESS_KEY_ID=<access-key>
CNPG_BACKUPS_ACCESS_SECRET_KEY=<secret-key>
```

The Model Installer has two additional configuration fields for writing to the `mlflow` bucket. These can contain the same provider-issued writer credential used elsewhere when its policy grants the required access.

```bash
LLM_BACKEND_MLFLOW_S3_ACCESS_KEY_ID=<access-key-with-mlflow-bucket-access>
LLM_BACKEND_MLFLOW_S3_SECRET_ACCESS_KEY=<secret-key>
```

Do not commit these values. The platform seals the rendered Kubernetes secrets, but the source credentials remain in the site's secret-management process.

## Troubleshooting [#troubleshooting]

* `AccessDenied` during rendering or startup — verify the mapped credential and its bucket policy for the affected component.
* TLS verification failure — supply the private CA through `OBJECT_STORAGE_CA_PEM` or fix the endpoint certificate chain.
* Requests use the wrong port — set `OBJECT_STORAGE_PORT` explicitly and verify the rendered endpoint.
* Browser download fails while backend access works — verify `OBJECT_STORAGE_PUBLIC_HOST`, its certificate, and the bucket CORS policy.
* Presigned upload fails only on one provider — verify `OBJECT_STORAGE_CAPABILITIES_PRESIGN_SIGNS_CONTENT_TYPE` against that provider's signing behavior.
* MLflow import succeeds but model loading fails — test the `mlflow` bucket with the KServe credential and review cross-account enforcement.
* PostgreSQL backups fail — verify the `cnpg-backups` bucket, the two `CNPG_BACKUPS_*` credentials, retention support, and CA trust.

## Related pages [#related-pages]

* [Configuration model](/docs/deployment/configuration-model)
* [Deployment sequence](/docs/deployment/playbooks/deployment-sequence)
* [Storage and databases](/docs/foundation/storage-and-databases)
* [PostgreSQL backup and disaster recovery](/docs/deployment/playbooks/backup-and-disaster-recovery)
* [Rook-Ceph RGW S3 access](/docs/deployment/playbooks/rook-ceph-rgw-s3-access)
