# Speaches (/docs/ai/components/speaches)



## Component Category [#component-category]

AI inference / speech processing

## Component Description [#component-description]

Speaches is an OpenAI-compatible speech server for streaming transcription, translation, and speech generation. BullSequana AI deploys it as a backend-owned service that provides the internal speech endpoint used by the BSQAI API.

## Why It Is Used [#why-it-is-used]

In BullSequana AI, Speaches provides one runtime for batch and real-time speech-to-text and text-to-speech workflows. FasterWhisper handles transcription, while pre-cached Piper and Kokoro voices support speech generation without requiring runtime model downloads.

## Learn More [#learn-more]

* [Speaches documentation](https://speaches.ai/)
* [speaches-ai/speaches on GitHub](https://github.com/speaches-ai/speaches)

## Deployment notes [#deployment-notes]

The Speaches service deploys inside the BSQAI API release as a separate single-replica Kubernetes `Deployment` and internal `ClusterIP` service. The default CPU configuration uses int8 FasterWhisper inference. An environment override selects the CUDA image, float16 inference, and one NVIDIA GPU.

The deployment uses a `Recreate` rollout strategy because its 10Gi model-cache volume is `ReadWriteOnce`. A pre-deployment Job fills that cache from a digest-pinned model image by default, which keeps speech startup reproducible in restricted-network environments. The cache contains the selected Whisper model and the configured TTS voices.

## Interacts With [#interacts-with]

* `BSQAI API`, which exposes authenticated batch and real-time STT and TTS endpoints.
* `FasterWhisper`, which performs speech transcription inside Speaches.
* `AI Web Portal`, which uses speech services for microphone input and generated audio.
* `NVIDIA GPU Operator`, which provides GPU scheduling support when CUDA mode is enabled.
