feat(whisper): OpenAI-compatible transcription server, probes that survive load - #10
Merged
Merged
Conversation
…rvive load The previous image wrapped an ASR web service that only exposes /asr, so the router's /v1/audio/transcriptions route returned 404 and the liveness probe (/docs) shared the event loop with inference. - deploy/whisper/server.py: faster-whisper behind FastAPI with /v1/audio/transcriptions, /v1/audio/translations, /v1/models, /health; inference runs in a worker thread, /health never queues - Containerfile.whisper: python base, av pinned, model baked and loaded from a local path (offline); default seccomp profile works - chart: probes on /health, cpuThreads/language/maxUploadMB/maxConcurrent values, RuntimeDefault seccomp, router.extraEnv for timeout tuning - drop the Renovate rule for the removed upstream image Verified locally with large-v3-turbo: German mp3/wav via server and via router, /health 1-2 ms during a 3.7 min transcription, container runs non-root with cap-drop ALL and no-new-privileges.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
The whisper image wrapped an upstream ASR service that only serves
/asr, so/v1/audio/transcriptionsthrough the router returned 404, and the/docsprobe blocked behind inference.Changes
deploy/whisper/server.py: faster-whisper behind FastAPI —/v1/audio/transcriptions,/v1/audio/translations,/v1/models,/health. Inference runs in a worker thread,/healthnever queues. Formats json/text/verbose_json/srt/vtt, upload limit, default language.Containerfile.whisper: python base,avpinned (newer PyAV breaks faster-whisper), model baked and loaded from a local path (fully offline), non-root./health, new valuescpuThreads,language,maxUploadMB,maxConcurrent; seccompRuntimeDefault(ctranslate2 4.x has no executable stack);router.extraEnvto tune timeouts.Testing
/healthanswers in 1-2 ms during a 3.7 min transcription--user 1000 --cap-drop ALL --security-opt no-new-privilegesand default seccomphelm lintandhelm templatewithwhisper.enabled=trueRelated Issues
Closes #9