# Working with data (/docs/use/data)



BullSequana AI provides several data paths rather than one universal pipeline. Start from the outcome, then choose the managed service and storage boundary that fit the source, latency, governance, and consumer.

## Choose the data path [#choose-the-data-path]

| Goal                                                                | Start here                                                                       | Main platform services                                                                                                            |
| ------------------------------------------------------------------- | -------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------- |
| Load data from an application, database, file, or API on a schedule | Define an Airbyte source, destination, and connection                            | [Airbyte](/docs/data/components/airbyte), object storage, PostgreSQL                                                              |
| Publish and consume event streams                                   | Define topics, producers, consumers, retention, and ownership                    | [Kafka](/docs/data/components/kafka)                                                                                              |
| Explore data or build repeatable processing                         | Open a managed development environment, then package the work as a pipeline      | [Developer Workspace](/docs/data/developer-workspace), Kubeflow Pipelines, [Spark Operator](/docs/data/components/spark-operator) |
| Build governed analytical tables                                    | Create the approved project and warehouse boundary before writing Iceberg tables | [Lakekeeper](/docs/data/components/lakekeeper), Spark, S3-compatible object storage                                               |
| Build dashboards and explore governed datasets                      | Connect an approved data source and publish through the correct tenant or team   | [Superset](/docs/data/components/superset)                                                                                        |
| Turn documents into knowledge for AI                                | Upload through the portal library and wait for indexing to complete              | [Files and RAG](/docs/ai/rag), Docling, Milvus, object storage                                                                    |
| Track experiments and register models                               | Log the run, parameters, metrics, and artifacts before model hand-off            | [MLflow](/docs/data/developer-workspace/mlflow)                                                                                   |

For a guided data-engineering journey, use [Build a data and ML workflow](/docs/get-started/build-data-workflow).

## Understand where data lives [#understand-where-data-lives]

| Data                                                       | System of record in the platform                                             |
| ---------------------------------------------------------- | ---------------------------------------------------------------------------- |
| Files, pipeline artifacts, model artifacts, and table data | Configured S3-compatible object storage: in-cluster Rook-Ceph or external S3 |
| Application metadata and transactional state               | PostgreSQL through CloudNativePG-managed services                            |
| Iceberg catalog, projects, and warehouse metadata          | Lakekeeper, with table objects in S3-compatible storage                      |
| Embeddings and retrieval payloads                          | Milvus                                                                       |
| Event streams                                              | Kafka                                                                        |
| Experiment runs, metrics, and registered-model metadata    | MLflow, with artifacts in configured object storage                          |

Do not build an endpoint by concatenating a host and path. Use the endpoint, addressing style, TLS, public-presign host, and capability settings defined by the platform's object-storage configuration.

## Typical flow [#typical-flow]

<Mermaid
  chart="flowchart TD
    SRC[&#x22;Approved sources&#x22;] --> INGEST[&#x22;Airbyte or application ingestion&#x22;]
    SRC --> EVENTS[&#x22;Kafka event streams&#x22;]
    SRC --> WORKSPACE[&#x22;Developer Workspace&#x22;]
    INGEST --> STORE[&#x22;S3-compatible storage or PostgreSQL&#x22;]
    EVENTS --> PIPE[&#x22;Kubeflow or Spark processing&#x22;]
    WORKSPACE --> PIPE
    STORE --> PIPE
    PIPE --> TABLES[&#x22;Governed Iceberg tables through Lakekeeper&#x22;]
    PIPE --> MLFLOW[&#x22;MLflow experiments and models&#x22;]
    TABLES --> BI[&#x22;Superset analytics&#x22;]
    STORE --> RAG[&#x22;Files and RAG&#x22;]
    RAG --> AI[&#x22;Grounded AI experiences&#x22;]"
/>

The exact path can be batch, streaming, or interactive. Production work should make inputs, outputs, schedules, retry behavior, data quality, and ownership explicit rather than relying on a notebook or one-off transfer.

## Tenant and team boundaries [#tenant-and-team-boundaries]

Keycloak Organizations establish tenant membership. Inside a tenant, team and personal workspaces define who can operate notebooks, pipelines, artifacts, and governed data assets. In the 1.3.0 Data stack:

* a Keycloak Organization maps to a Lakekeeper project with a default warehouse
* a personal user bucket and each shared-space bucket map to additional warehouses in that tenant project
* the Kubeflow profile reconciler creates the user's personal bucket, ensures buckets for the user's space memberships, and prepares their Lakekeeper warehouses
* deleting a user Profile archives only its personal bucket; shared-space buckets persist with the space
* Tuples records ownership and parent-space relationships for spaces, buckets, datasets, and tables
* service credentials are injected into managed workloads rather than copied into notebooks or source code

Read [Multi-tenancy and tiered RBAC](/docs/data/multi-tenancy-and-tiered-rbac) before sharing data or promoting a workflow.

## Production checklist [#production-checklist]

Before a workflow becomes operational, confirm:

* the source, accountable owner, classification, and permitted purpose are recorded
* the tenant and team access boundary is tested with an allowed and a denied user
* schemas, freshness, retention, and data-quality expectations are measurable
* credentials come from the platform's approved secret path
* another authorized user can reproduce the processing from versioned code or a pipeline definition
* failures, retries, artifacts, logs, and downstream impact are visible
* backup, restore, and deletion behavior match the system of record

Use [Data](/docs/data) for the capability map and [Data component catalog](/docs/data/components) when you need implementation detail.
