Kubeflow
Pipelines, notebooks, and experiment management for ML, DL, and GenAI workflows.
Component Category
Machine learning / pipelines and notebooks
Component Description
Kubeflow is a Kubernetes-native platform for building and managing ML, deep learning, and GenAI workflows. In BullSequana AI it provides Kubeflow Pipelines for orchestrating experiments, a notebook server manager (Developer Workspace) for interactive development environments (JupyterLab, VSCode), and a profile system for multi-tenant namespace isolation. It is deployed as raw Kustomize manifests (not a Helm chart), based on upstream tag 26.03 with local overlays.
Why It Is Used
In BullSequana AI, Kubeflow provides the experimentation and pipeline execution layer for the Data tier. It gives data scientists and AI engineers a structured environment to build, run, and track ML, DL, and GenAI workflows with pipeline versioning, artifact tracking through Rook Ceph S3 storage, and per-profile namespace isolation. Kubeflow complements MLflow (experiment tracking) and Argo Workflows (general-purpose automation) by focusing specifically on pipeline orchestration and notebook-based development.
Learn More
Deployment notes
Kubeflow deploys in the proai tier at sync wave 6. It depends on two components that must be running first:
- Kyverno (sync wave 5) — enforces ClusterPolicy and GeneratingPolicy resources for profile-level RBAC
- Metacontroller (sync wave 5) — provides the DecoratorController CRD used by Kubeflow Pipelines
Kubeflow uses Kustomize format instead of Helm. The chart_name field is intentionally empty, so ArgoCD applies the raw manifests from Git instead of pulling a chart from the OCI registry.
Storage uses a shared central pipeline-artifact bucket. When a user Profile is reconciled, the profile bucket provisioner creates a personal user-* bucket and ensures a space-* bucket for each of that user's space memberships. It records ownership and parent-space relationships through Tuples and prepares a Lakekeeper warehouse for each bucket in the tenant project. Deleting the user Profile archives only the personal bucket; space buckets persist because they belong to the space. Object storage can be Rook-Ceph RGW or the configured external S3-compatible provider.
Authentication is handled through an internal OAuth2 Proxy instance configured against Keycloak, with access controlled through the COREAI-KUBEFLOW-ADMIN-GROUP and COREAI-KUBEFLOW-USERS-GROUP groups.
Kubeflow also deploys an Istio 1.28 Ambient service-mesh layer. Istio CNI and ztunnel provide the node-level data plane, while istiod manages policy and a dedicated ingress gateway receives the public Kubeflow route. The Kubeflow application namespaces use Ambient mode. Istio RequestAuthentication and AuthorizationPolicy resources validate Keycloak JWTs and restrict access between the ingress gateway, OAuth2 Proxy, Kubeflow services, and profile namespaces.
Custom container images are built for the central dashboard, Jupyter web app, and KFP frontend, stored in the platform registry.
Interacts With
Kyverno, which enforces the ClusterPolicy resources shipped by Kubeflow for profile RBAC and pipeline access controls.Metacontroller, which provides the DecoratorController CRD required by Kubeflow Pipelines.Keycloak, which provides SSO authentication through an internal OAuth2 Proxy instance.Istio, which provides the Ambient data plane, ingress boundary, JWT validation, and service authorization policies.Rook-Cephor external S3-compatible storage, which stores pipeline artifacts and user data.TuplesandOpenFGA, which record and authorize personal and shared-space object-storage relationships.LakekeeperandSpark Operator, which support catalog-backed distributed data processing from Kubeflow workflows.MLflow, which handles experiment tracking alongside Kubeflow's pipeline execution.trust-manager, which distributes CA trust bundles into Kubeflow namespaces for TLS verification.