# Build a data and ML workflow (/docs/get-started/build-data-workflow)



This path is for data engineers, analytics engineers, data scientists, and machine-learning engineers using managed workspaces, pipelines, and MLflow.

## Outcome [#outcome]

You can identify an approved source, develop against it in the correct tenant and team, package repeatable work, track artifacts or models, and establish an operational owner.

## 1. Define the data contract [#1-define-the-data-contract]

Before choosing a tool, record:

* source and accountable owner
* classification and permitted purposes
* tenant and team allowed to use it
* schema and expected volume
* refresh or event behavior
* freshness, quality, and retention expectations
* downstream consumers and failure impact

This contract determines whether the workflow needs batch integration, streaming, interactive development, analytics, or model training.

## 2. Enter through an approved source path [#2-enter-through-an-approved-source-path]

Use a managed connector or approved object, database, or event interface. Avoid copying credentials or datasets into personal environments. Use [Working with Data](/docs/use/data) for user-facing entry points and the [Data component catalog](/docs/data/components) only when implementation detail is necessary.

## 3. Develop in a managed workspace [#3-develop-in-a-managed-workspace]

[Developer Workspace](/docs/data/developer-workspace) provides the development environment, pipeline tooling, and MLflow entry points. Choose an environment that matches the workload and use injected platform credentials instead of embedding secrets in notebooks or source code.

## 4. Turn exploration into a pipeline [#4-turn-exploration-into-a-pipeline]

Package repeatable processing as a pipeline with explicit inputs, outputs, parameters, artifacts, and failure behavior. [Pipeline concepts](/docs/data/developer-workspace/pipelines/concepts) explains the execution model; [upload a pipeline](/docs/data/developer-workspace/pipelines/upload-a-pipeline) and [runs](/docs/data/developer-workspace/pipelines/runs) cover operation.

Use recurring runs only after one execution is reproducible and observable.

## 5. Track experiments and models [#5-track-experiments-and-models]

When the workflow trains or evaluates a model, use [MLflow experiment tracking](/docs/data/developer-workspace/mlflow/track-experiments) for parameters, metrics, and artifacts. Use [register models](/docs/data/developer-workspace/mlflow/register-models) to create a governed hand-off into model lifecycle and serving.

## 6. Apply governance and operations [#6-apply-governance-and-operations]

BullSequana AI carries tenant and team authorization into notebooks, pipelines, artifacts, models, logs, and analytics. Verify the workspace memberships and resource assignments for this workflow, then assign an owner for freshness, failures, capacity, and downstream changes.

Use [Multi-tenancy and tiered RBAC](/docs/data/multi-tenancy-and-tiered-rbac), [Security and compliance](/docs/security), and [Administration and operations](/docs/administration).

## You are done when [#you-are-done-when]

* the source and permitted use are documented
* another authorized user can reproduce the workflow
* inputs, outputs, artifacts, and failures are visible
* model experiments and registrations retain traceable context
* production schedules and capacity have an owner
* tenant and team boundaries are tested end to end
