Speaches

Speech inference service used by the BSQAI API for transcription and speech generation.

Agentic Friendly

Component Category

AI inference / speech processing

Component Description

Speaches is an OpenAI-compatible speech server for streaming transcription, translation, and speech generation. BullSequana AI deploys it as a backend-owned service that provides the internal speech endpoint used by the BSQAI API.

Why It Is Used

In BullSequana AI, Speaches provides one runtime for batch and real-time speech-to-text and text-to-speech workflows. FasterWhisper handles transcription, while pre-cached Piper and Kokoro voices support speech generation without requiring runtime model downloads.

Learn More

Deployment notes

The Speaches service deploys inside the BSQAI API release as a separate single-replica Kubernetes Deployment and internal ClusterIP service. The default CPU configuration uses int8 FasterWhisper inference. An environment override selects the CUDA image, float16 inference, and one NVIDIA GPU.

The deployment uses a Recreate rollout strategy because its 10Gi model-cache volume is ReadWriteOnce. A pre-deployment Job fills that cache from a digest-pinned model image by default, which keeps speech startup reproducible in restricted-network environments. The cache contains the selected Whisper model and the configured TTS voices.

Interacts With

  • BSQAI API, which exposes authenticated batch and real-time STT and TTS endpoints.
  • FasterWhisper, which performs speech transcription inside Speaches.
  • AI Web Portal, which uses speech services for microphone input and generated audio.
  • NVIDIA GPU Operator, which provides GPU scheduling support when CUDA mode is enabled.

On this page