Bug
In ed3468f3881abb9e7b6c7d404f75049aecccb04f, the tested
moonshine-streaming-tiny-Q8_0.gguf model's adapter positional table implies an
offline input capacity of approximately 81.92 seconds at 16 kHz, but
transcribe-cli reports:
On the CPU backend, my tested 81-second input succeeds, while the tested
82-second input reproducibly aborts inside GGML instead of returning a clean
input-length error.
The crashing measurements below were collected using a binary built from the
unmodified tested revision, before applying any local diagnostic changes.
Tested revision
ed3468f3881abb9e7b6c7d404f75049aecccb04f
Model
moonshine-streaming-tiny-Q8_0.gguf
Downloaded from the model location referenced by this repository:
handy-computer/moonshine-streaming-tiny-gguf
SHA-256 of the tested model:
930e4622ad3a24158b91406c30c977fa6a26b34cb32d6ac3e57cfb23383a869e
Backend/configuration:
Reproduction
Start with a speech recording at least 82 seconds long.
Create deterministic PCM16 mono 16 kHz clips:
ffmpeg -y -i input.wav -t 81 -ar 16000 -ac 1 -c:a pcm_s16le 81.wav
ffmpeg -y -i input.wav -t 82 -ar 16000 -ac 1 -c:a pcm_s16le 82.wav
At 16 kHz these should contain:
81.wav: 1,296,000 samples
82.wav: 1,312,000 samples
Run the 81-second clip:
build/bin/transcribe-cli \
-m moonshine-streaming-tiny-Q8_0.gguf \
--threads 2 \
--backend cpu \
--timestamps none \
81.wav
On my system, this completes successfully.
Then run the 82-second clip:
build/bin/transcribe-cli \
-m moonshine-streaming-tiny-Q8_0.gguf \
--threads 2 \
--backend cpu \
--timestamps none \
82.wav
On my system, this reproducibly aborts with:
ggml-cpu/ops.cpp:4886:
GGML_ASSERT(i01 >= 0 && i01 < ne01) failed
I also reproduced the same abort with 100 s, 120 s, and 140 s inputs.
Why 81.92 seconds looks significant
For the tested model, the adapter uses a learned positional embedding with
4096 positions.
The encoder reduces the input through two stride-2 stages. For the tested
model geometry, one adapter position corresponds to 320 input samples:
4096 × 320 / 16000 = 81.92 seconds
Equivalently, the resulting encoder rate is 50 positions per second:
4096 / 50 = 81.92 seconds
This predicts:
81 s -> 4050 encoder positions
82 s -> 4100 encoder positions
A 4096-row positional table has valid row IDs 0..4095.
In the tested revision, the adapter path constructs absolute position IDs:
pos_ids[i] = abs_frame_offset + i;
and passes them into ggml_get_rows().
I could not find a check in the original revision that rejects an adapter
position at or beyond the positional-table size before this lookup.
The observed GGML failure is consistent with an out-of-range adapter
positional lookup. The CPU get_rows assertion checks that the requested row
index is below the source tensor's row count.
At the same time, moonshine_streaming_max_audio_ms() derives the reported
~1024-second value from dec_max_position_embeddings and an estimated output
token rate:
4096 tokens / ~4 tokens per second = ~1024 seconds
That appears to describe the decoder output budget rather than the earlier
input-side adapter positional constraint.
Expected behavior
For one-shot/offline input, if this adapter table is the binding hard input
constraint, input beyond that limit should be rejected before the out-of-range
embedding lookup rather than aborting the process.
For example, a result such as:
TRANSCRIBE_ERR_INPUT_TOO_LONG
would be consistent with the repository's existing hard-input-limit contract.
max_audio_ms should also represent a duration consistent with the actual
binding input-side constraint.
I am not asserting what the correct streaming API behavior should be. The
existing streaming contract has separate committed-text/truncation semantics,
so that may require different handling from the one-shot path.
Additional diagnostic
As a local diagnostic only, I added a guard around the positional/input limit.
With that guard:
81 s -> succeeds
82 s -> clean INPUT_TOO_LONG
The guard also prevents the corresponding abort when a streaming session
eventually reaches an invalid adapter position.
This is supporting evidence for the diagnosis, not a proposal that
INPUT_TOO_LONG is necessarily the correct upstream streaming behavior.
I have not included the local patch here because I wanted to report the
reproducible failure and let the maintainers determine the intended API and
implementation behavior.
Bug
In
ed3468f3881abb9e7b6c7d404f75049aecccb04f, the testedmoonshine-streaming-tiny-Q8_0.ggufmodel's adapter positional table implies anoffline input capacity of approximately 81.92 seconds at 16 kHz, but
transcribe-clireports:On the CPU backend, my tested 81-second input succeeds, while the tested
82-second input reproducibly aborts inside GGML instead of returning a clean
input-length error.
The crashing measurements below were collected using a binary built from the
unmodified tested revision, before applying any local diagnostic changes.
Tested revision
Model
Downloaded from the model location referenced by this repository:
SHA-256 of the tested model:
Backend/configuration:
Reproduction
Start with a speech recording at least 82 seconds long.
Create deterministic PCM16 mono 16 kHz clips:
At 16 kHz these should contain:
Run the 81-second clip:
On my system, this completes successfully.
Then run the 82-second clip:
On my system, this reproducibly aborts with:
I also reproduced the same abort with 100 s, 120 s, and 140 s inputs.
Why 81.92 seconds looks significant
For the tested model, the adapter uses a learned positional embedding with
4096 positions.
The encoder reduces the input through two stride-2 stages. For the tested
model geometry, one adapter position corresponds to 320 input samples:
Equivalently, the resulting encoder rate is 50 positions per second:
This predicts:
A 4096-row positional table has valid row IDs
0..4095.In the tested revision, the adapter path constructs absolute position IDs:
and passes them into
ggml_get_rows().I could not find a check in the original revision that rejects an adapter
position at or beyond the positional-table size before this lookup.
The observed GGML failure is consistent with an out-of-range adapter
positional lookup. The CPU
get_rowsassertion checks that the requested rowindex is below the source tensor's row count.
At the same time,
moonshine_streaming_max_audio_ms()derives the reported~1024-second value from
dec_max_position_embeddingsand an estimated outputtoken rate:
That appears to describe the decoder output budget rather than the earlier
input-side adapter positional constraint.
Expected behavior
For one-shot/offline input, if this adapter table is the binding hard input
constraint, input beyond that limit should be rejected before the out-of-range
embedding lookup rather than aborting the process.
For example, a result such as:
would be consistent with the repository's existing hard-input-limit contract.
max_audio_msshould also represent a duration consistent with the actualbinding input-side constraint.
I am not asserting what the correct streaming API behavior should be. The
existing streaming contract has separate committed-text/truncation semantics,
so that may require different handling from the one-shot path.
Additional diagnostic
As a local diagnostic only, I added a guard around the positional/input limit.
With that guard:
The guard also prevents the corresponding abort when a streaming session
eventually reaches an invalid adapter position.
This is supporting evidence for the diagnosis, not a proposal that
INPUT_TOO_LONGis necessarily the correct upstream streaming behavior.I have not included the local patch here because I wanted to report the
reproducible failure and let the maintainers determine the intended API and
implementation behavior.