Data
Connect, govern, process, analyze, and operationalize data on BullSequana AI.
Data is the product domain for moving information from source systems into governed pipelines, workspaces, models, analytics, and AI experiences. It connects data engineering and data science without requiring users to operate the underlying platform components.
Choose an outcome
Build a data workflow
Follow the path from a source to an operated data or ML result.
Open a Developer Workspace
Use managed development environments, pipelines, and MLflow.
Build and run pipelines
Upload, execute, schedule, and inspect pipeline runs.
Work with data tools
Use the platform's data-facing interfaces and services.
Capability map
Integrate and move data
Data integration connects applications, databases, object stores, and external services to governed platform workflows. Connectors define the source and destination; orchestration controls when work runs; operational signals show whether transfers are healthy.
Use the Airbyte component reference when diagnosing an existing connector implementation.
Engineer and orchestrate data
Lakehouse and pipeline capabilities organize catalog-backed tables, processing jobs, artifacts, schedules, and dependencies. Developer Workspace pipelines provide the practical path for building and operating reusable workflows.
Govern and discover data
Catalog and authorization capabilities describe schemas, ownership, permissions, lineage, and the tenant or team allowed to use an asset. Multi-tenancy and tiered RBAC explains how those boundaries apply across Data.
Develop models and experiments
Developer Workspace provides notebook-first and Python-first environments. MLflow records experiments, artifacts, and model registrations so that development work can move into governed AI serving without losing its history.
Start with track experiments and register models.
Stream and analyze
Streaming moves events through Kafka-backed paths. Analytics and BI turn governed data into dashboards and operational views. These paths share catalog, access-control, observability, and storage responsibilities with batch workflows.
Natural-language data access
Natural-language experiences combine a governed semantic model, controlled query execution, citations, and permission-aware answers. They sit across Data and AI: Data protects the meaning and access boundary; AI provides the conversational experience.
Typical data-to-production flow
- Identify the source, owner, classification, and expected refresh behavior.
- Choose a connector or approved ingestion path.
- Develop transformations or models in a managed workspace.
- Package the work as a repeatable pipeline with artifacts and parameters.
- Track experiments and register models when machine learning is involved.
- Apply tenant, team, catalog, and data-access policies before production use.
- Monitor freshness, failures, consumption, and downstream impact.