# Spark Operator (/docs/data/components/spark-operator)



## Component Category [#component-category]

Distributed data processing

## Component Description [#component-description]

Spark Operator reconciles `SparkApplication` resources into Apache Spark driver and executor workloads on Kubernetes. It watches tenant namespaces so teams can run data-processing jobs without operating a separate Spark cluster.

## Why It Is Used [#why-it-is-used]

In BullSequana AI, Spark supplies distributed processing for data preparation, lakehouse tables, and machine-learning pipelines. The operator makes those jobs declarative, observable, and compatible with GitOps and Kubernetes resource controls.

## Learn More [#learn-more]

* [Spark Operator user guide](https://spark.kubeflow.org/en/latest/user-guide/index.html)
* [kubeflow/spark-operator on GitHub](https://github.com/kubeflow/spark-operator)

## Deployment notes [#deployment-notes]

BullSequana AI 1.3.0 deploys Spark Operator 2.5.2, watches all namespaces, uses images mirrored in the private registry, and exposes operator metrics for platform monitoring.

## Interacts With [#interacts-with]

* `Lakekeeper`, which supplies the Iceberg REST Catalog.
* `Rook-Ceph` or external S3-compatible storage, which stores lakehouse data and job artifacts.
* `Kubeflow`, which can orchestrate Spark-backed ML and data pipelines.
* `Grafana` and `Prometheus`, which expose operator and workload telemetry.
