Build a data and ML workflow
Move from a governed source to a repeatable pipeline, experiment, or production result.
This path is for data engineers, analytics engineers, data scientists, and machine-learning engineers using managed workspaces, pipelines, and MLflow.
Outcome
You can identify an approved source, develop against it in the correct tenant and team, package repeatable work, track artifacts or models, and establish an operational owner.
1. Define the data contract
Before choosing a tool, record:
- source and accountable owner
- classification and permitted purposes
- tenant and team allowed to use it
- schema and expected volume
- refresh or event behavior
- freshness, quality, and retention expectations
- downstream consumers and failure impact
This contract determines whether the workflow needs batch integration, streaming, interactive development, analytics, or model training.
2. Enter through an approved source path
Use a managed connector or approved object, database, or event interface. Avoid copying credentials or datasets into personal environments. Use Working with Data for user-facing entry points and the Data component catalog only when implementation detail is necessary.
3. Develop in a managed workspace
Developer Workspace provides the development environment, pipeline tooling, and MLflow entry points. Choose an environment that matches the workload and use injected platform credentials instead of embedding secrets in notebooks or source code.
4. Turn exploration into a pipeline
Package repeatable processing as a pipeline with explicit inputs, outputs, parameters, artifacts, and failure behavior. Pipeline concepts explains the execution model; upload a pipeline and runs cover operation.
Use recurring runs only after one execution is reproducible and observable.
5. Track experiments and models
When the workflow trains or evaluates a model, use MLflow experiment tracking for parameters, metrics, and artifacts. Use register models to create a governed hand-off into model lifecycle and serving.
6. Apply governance and operations
BullSequana AI carries tenant and team authorization into notebooks, pipelines, artifacts, models, logs, and analytics. Verify the workspace memberships and resource assignments for this workflow, then assign an owner for freshness, failures, capacity, and downstream changes.
Use Multi-tenancy and tiered RBAC, Security and compliance, and Administration and operations.
You are done when
- the source and permitted use are documented
- another authorized user can reproduce the workflow
- inputs, outputs, artifacts, and failures are visible
- model experiments and registrations retain traceable context
- production schedules and capacity have an owner
- tenant and team boundaries are tested end to end