Working with data

Choose the right BullSequana AI data path for ingestion, processing, governance, analytics, and AI knowledge.

Agentic Friendly

BullSequana AI provides several data paths rather than one universal pipeline. Start from the outcome, then choose the managed service and storage boundary that fit the source, latency, governance, and consumer.

Choose the data path

GoalStart hereMain platform services
Load data from an application, database, file, or API on a scheduleDefine an Airbyte source, destination, and connectionAirbyte, object storage, PostgreSQL
Publish and consume event streamsDefine topics, producers, consumers, retention, and ownershipKafka
Explore data or build repeatable processingOpen a managed development environment, then package the work as a pipelineDeveloper Workspace, Kubeflow Pipelines, Spark Operator
Build governed analytical tablesCreate the approved project and warehouse boundary before writing Iceberg tablesLakekeeper, Spark, S3-compatible object storage
Build dashboards and explore governed datasetsConnect an approved data source and publish through the correct tenant or teamSuperset
Turn documents into knowledge for AIUpload through the portal library and wait for indexing to completeFiles and RAG, Docling, Milvus, object storage
Track experiments and register modelsLog the run, parameters, metrics, and artifacts before model hand-offMLflow

For a guided data-engineering journey, use Build a data and ML workflow.

Understand where data lives

DataSystem of record in the platform
Files, pipeline artifacts, model artifacts, and table dataConfigured S3-compatible object storage: in-cluster Rook-Ceph or external S3
Application metadata and transactional statePostgreSQL through CloudNativePG-managed services
Iceberg catalog, projects, and warehouse metadataLakekeeper, with table objects in S3-compatible storage
Embeddings and retrieval payloadsMilvus
Event streamsKafka
Experiment runs, metrics, and registered-model metadataMLflow, with artifacts in configured object storage

Do not build an endpoint by concatenating a host and path. Use the endpoint, addressing style, TLS, public-presign host, and capability settings defined by the platform's object-storage configuration.

Typical flow

The exact path can be batch, streaming, or interactive. Production work should make inputs, outputs, schedules, retry behavior, data quality, and ownership explicit rather than relying on a notebook or one-off transfer.

Tenant and team boundaries

Keycloak Organizations establish tenant membership. Inside a tenant, team and personal workspaces define who can operate notebooks, pipelines, artifacts, and governed data assets. In the 1.3.0 Data stack:

  • a Keycloak Organization maps to a Lakekeeper project with a default warehouse
  • a personal user bucket and each shared-space bucket map to additional warehouses in that tenant project
  • the Kubeflow profile reconciler creates the user's personal bucket, ensures buckets for the user's space memberships, and prepares their Lakekeeper warehouses
  • deleting a user Profile archives only its personal bucket; shared-space buckets persist with the space
  • Tuples records ownership and parent-space relationships for spaces, buckets, datasets, and tables
  • service credentials are injected into managed workloads rather than copied into notebooks or source code

Read Multi-tenancy and tiered RBAC before sharing data or promoting a workflow.

Production checklist

Before a workflow becomes operational, confirm:

  • the source, accountable owner, classification, and permitted purpose are recorded
  • the tenant and team access boundary is tested with an allowed and a denied user
  • schemas, freshness, retention, and data-quality expectations are measurable
  • credentials come from the platform's approved secret path
  • another authorized user can reproduce the processing from versioned code or a pipeline definition
  • failures, retries, artifacts, logs, and downstream impact are visible
  • backup, restore, and deletion behavior match the system of record

Use Data for the capability map and Data component catalog when you need implementation detail.

On this page