Skip to content

fix(whisper): avoid repetition loops on long audio, add startup warmup - #20

Merged
eksrha merged 1 commit into
mainfrom
feat/whisper-robust-long-audio
Oct 4, 2026
Merged

eksrha merged 1 commit into
mainfrom
feat/whisper-robust-long-audio

Conversation

@eksrha

@eksrha eksrha commented Oct 4, 2026

Copy link
Copy Markdown
Contributor

Closes #17

What

  • Temperature: without an explicit temperature form field the server now passes faster-whisper's fallback schedule (0.0, 0.2 ... 1.0, default compression-ratio/log-prob/no-speech thresholds) instead of the single float 0.0. An explicit value (including 0) is used as given.
  • condition_on_previous_text is off by default (whisper.conditionOnPreviousText / WHISPER_CONDITION_ON_PREVIOUS_TEXT).
  • Audio longer than WHISPER_BATCH_THRESHOLD_S (default 35 s, chart whisper.batchThresholdSeconds) goes through BatchedInferencePipeline (WHISPER_BATCH_SIZE, default 8, whisper.batchSize). Audio is decoded once, the duration comes from the decoded samples. Shorter audio uses the regular transcribe().
  • Warmup: after loading the model a 2 s synthetic clip runs through VAD and decoder (WHISPER_WARMUP, default true); /health returns 503 until model and warmup are done. A failed warmup is only logged.
  • whisper_request log gains path (standard|batched), batch_size, temperature_mode (fallback|fixed), condition_on_previous_text. Prompt/hotwords logic and response formats are unchanged.

Notes

  • The batched pipeline in faster-whisper 1.2.1 always needs VAD and uses only the first temperature; this is documented in the README.
  • Clients that always send temperature=0 get no fallback on the standard path (explicit value wins).

Measured locally (large-v3-turbo, int8, 4 threads, beam 1, VAD, FLEURS de clips)

  • Three 55-88 s clips: no repetition loop (88 s clip: 132 words out for 132 reference words, previously 256 vs 120), WER about 3 % on long clips instead of 42 %.
  • 88 s clip in 23.5 s, 56 s clip in about 16 s (previously 27 s); short clips unchanged (about 8 s for 8.7 s of audio).
  • First request after start: 8.4 s without warmup, 7.7 s with it. Warmup itself takes about 8 s at startup.

Tests

pytest (stubbed model): path selection, threshold, temperature fallback vs explicit, condition flag, log fields, response formats on both paths, health/warmup lifecycle.

Pass faster-whisper's temperature fallback schedule unless the request sets
temperature explicitly, keep condition_on_previous_text off by default and
transcribe audio longer than WHISPER_BATCH_THRESHOLD_S (default 35 s) with
BatchedInferencePipeline. Run a short synthetic clip through VAD and decoder
at startup and report /health ready only afterwards. The request log gains
path, batch_size, temperature_mode and condition_on_previous_text.

Closes #17
@eksrha
eksrha merged commit 2bdce6c into main Oct 4, 2026
3 checks passed
@eksrha
eksrha deleted the feat/whisper-robust-long-audio branch October 4, 2026 20:08
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Whisper-Server: Wiederholungsschleifen bei langem Audio vermeiden

1 participant