fix(whisper): avoid repetition loops on long audio, add startup warmup - #20
Merged
Merged
Conversation
Pass faster-whisper's temperature fallback schedule unless the request sets temperature explicitly, keep condition_on_previous_text off by default and transcribe audio longer than WHISPER_BATCH_THRESHOLD_S (default 35 s) with BatchedInferencePipeline. Run a short synthetic clip through VAD and decoder at startup and report /health ready only afterwards. The request log gains path, batch_size, temperature_mode and condition_on_previous_text. Closes #17
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #17
What
temperatureform field the server now passes faster-whisper's fallback schedule (0.0, 0.2 ... 1.0, default compression-ratio/log-prob/no-speech thresholds) instead of the single float0.0. An explicit value (including 0) is used as given.condition_on_previous_textis off by default (whisper.conditionOnPreviousText/WHISPER_CONDITION_ON_PREVIOUS_TEXT).WHISPER_BATCH_THRESHOLD_S(default 35 s, chartwhisper.batchThresholdSeconds) goes throughBatchedInferencePipeline(WHISPER_BATCH_SIZE, default 8,whisper.batchSize). Audio is decoded once, the duration comes from the decoded samples. Shorter audio uses the regulartranscribe().WHISPER_WARMUP, default true);/healthreturns 503 until model and warmup are done. A failed warmup is only logged.whisper_requestlog gainspath(standard|batched),batch_size,temperature_mode(fallback|fixed),condition_on_previous_text. Prompt/hotwords logic and response formats are unchanged.Notes
temperature=0get no fallback on the standard path (explicit value wins).Measured locally (large-v3-turbo, int8, 4 threads, beam 1, VAD, FLEURS de clips)
Tests
pytest (stubbed model): path selection, threshold, temperature fallback vs explicit, condition flag, log fields, response formats on both paths, health/warmup lifecycle.