Spark Operator

Kubernetes operator for tenant-scoped Apache Spark applications.

Agentic Friendly

Component Category

Distributed data processing

Component Description

Spark Operator reconciles SparkApplication resources into Apache Spark driver and executor workloads on Kubernetes. It watches tenant namespaces so teams can run data-processing jobs without operating a separate Spark cluster.

Why It Is Used

In BullSequana AI, Spark supplies distributed processing for data preparation, lakehouse tables, and machine-learning pipelines. The operator makes those jobs declarative, observable, and compatible with GitOps and Kubernetes resource controls.

Learn More

Deployment notes

BullSequana AI 1.3.0 deploys Spark Operator 2.5.2, watches all namespaces, uses images mirrored in the private registry, and exposes operator metrics for platform monitoring.

Interacts With

  • Lakekeeper, which supplies the Iceberg REST Catalog.
  • Rook-Ceph or external S3-compatible storage, which stores lakehouse data and job artifacts.
  • Kubeflow, which can orchestrate Spark-backed ML and data pipelines.
  • Grafana and Prometheus, which expose operator and workload telemetry.

On this page