Skip to content

feat(whisper): OpenAI-compatible transcription server, probes that survive load - #10

Merged
eksrha merged 1 commit into
mainfrom
feat/whisper-openai-server
Oct 2, 2026
Merged

eksrha merged 1 commit into
mainfrom
feat/whisper-openai-server

Conversation

@eksrha

@eksrha eksrha commented Oct 2, 2026

Copy link
Copy Markdown
Contributor

Summary

The whisper image wrapped an upstream ASR service that only serves /asr, so /v1/audio/transcriptions through the router returned 404, and the /docs probe blocked behind inference.

Changes

  • deploy/whisper/server.py: faster-whisper behind FastAPI — /v1/audio/transcriptions, /v1/audio/translations, /v1/models, /health. Inference runs in a worker thread, /health never queues. Formats json/text/verbose_json/srt/vtt, upload limit, default language.
  • Containerfile.whisper: python base, av pinned (newer PyAV breaks faster-whisper), model baked and loaded from a local path (fully offline), non-root.
  • Chart: probes on /health, new values cpuThreads, language, maxUploadMB, maxConcurrent; seccomp RuntimeDefault (ctranslate2 4.x has no executable stack); router.extraEnv to tune timeouts.
  • Renovate: drop the rule for the removed upstream image. README/AGENTS updated.

Testing

  • Built the image; large-v3-turbo, German mp3 and wav transcribed correctly, directly and through the router
  • /health answers in 1-2 ms during a 3.7 min transcription
  • Container runs with --user 1000 --cap-drop ALL --security-opt no-new-privileges and default seccomp
  • helm lint and helm template with whisper.enabled=true
  • Measured: ~2.0 GiB RSS peak, ~1x real-time on 2 threads

Related Issues

Closes #9

…rvive load

The previous image wrapped an ASR web service that only exposes /asr, so
the router's /v1/audio/transcriptions route returned 404 and the liveness
probe (/docs) shared the event loop with inference.

- deploy/whisper/server.py: faster-whisper behind FastAPI with
  /v1/audio/transcriptions, /v1/audio/translations, /v1/models, /health;
  inference runs in a worker thread, /health never queues
- Containerfile.whisper: python base, av pinned, model baked and loaded
  from a local path (offline); default seccomp profile works
- chart: probes on /health, cpuThreads/language/maxUploadMB/maxConcurrent
  values, RuntimeDefault seccomp, router.extraEnv for timeout tuning
- drop the Renovate rule for the removed upstream image

Verified locally with large-v3-turbo: German mp3/wav via server and via
router, /health 1-2 ms during a 3.7 min transcription, container runs
non-root with cap-drop ALL and no-new-privileges.
@eksrha
eksrha merged commit 239e221 into main Oct 2, 2026
2 checks passed
@eksrha
eksrha deleted the feat/whisper-openai-server branch October 2, 2026 19:27
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Whisper backend does not serve /v1/audio/transcriptions and its probes fail under load

1 participant