Skip to content

Moonshine Streaming Tiny aborts above ~81.92s despite reporting 1024s max audio #180

Description

@koppimashin

Bug

In ed3468f3881abb9e7b6c7d404f75049aecccb04f, the tested
moonshine-streaming-tiny-Q8_0.gguf model's adapter positional table implies an
offline input capacity of approximately 81.92 seconds at 16 kHz, but
transcribe-cli reports:

max audio: 1024.0 s

On the CPU backend, my tested 81-second input succeeds, while the tested
82-second input reproducibly aborts inside GGML instead of returning a clean
input-length error.

The crashing measurements below were collected using a binary built from the
unmodified tested revision, before applying any local diagnostic changes.

Tested revision

ed3468f3881abb9e7b6c7d404f75049aecccb04f

Model

moonshine-streaming-tiny-Q8_0.gguf

Downloaded from the model location referenced by this repository:

handy-computer/moonshine-streaming-tiny-gguf

SHA-256 of the tested model:

930e4622ad3a24158b91406c30c977fa6a26b34cb32d6ac3e57cfb23383a869e

Backend/configuration:

CPU
2 threads

Reproduction

Start with a speech recording at least 82 seconds long.

Create deterministic PCM16 mono 16 kHz clips:

ffmpeg -y -i input.wav -t 81 -ar 16000 -ac 1 -c:a pcm_s16le 81.wav
ffmpeg -y -i input.wav -t 82 -ar 16000 -ac 1 -c:a pcm_s16le 82.wav

At 16 kHz these should contain:

81.wav: 1,296,000 samples
82.wav: 1,312,000 samples

Run the 81-second clip:

build/bin/transcribe-cli \
  -m moonshine-streaming-tiny-Q8_0.gguf \
  --threads 2 \
  --backend cpu \
  --timestamps none \
  81.wav

On my system, this completes successfully.

Then run the 82-second clip:

build/bin/transcribe-cli \
  -m moonshine-streaming-tiny-Q8_0.gguf \
  --threads 2 \
  --backend cpu \
  --timestamps none \
  82.wav

On my system, this reproducibly aborts with:

ggml-cpu/ops.cpp:4886:
GGML_ASSERT(i01 >= 0 && i01 < ne01) failed

I also reproduced the same abort with 100 s, 120 s, and 140 s inputs.

Why 81.92 seconds looks significant

For the tested model, the adapter uses a learned positional embedding with
4096 positions.

The encoder reduces the input through two stride-2 stages. For the tested
model geometry, one adapter position corresponds to 320 input samples:

4096 × 320 / 16000 = 81.92 seconds

Equivalently, the resulting encoder rate is 50 positions per second:

4096 / 50 = 81.92 seconds

This predicts:

81 s -> 4050 encoder positions
82 s -> 4100 encoder positions

A 4096-row positional table has valid row IDs 0..4095.

In the tested revision, the adapter path constructs absolute position IDs:

pos_ids[i] = abs_frame_offset + i;

and passes them into ggml_get_rows().

I could not find a check in the original revision that rejects an adapter
position at or beyond the positional-table size before this lookup.

The observed GGML failure is consistent with an out-of-range adapter
positional lookup. The CPU get_rows assertion checks that the requested row
index is below the source tensor's row count.

At the same time, moonshine_streaming_max_audio_ms() derives the reported
~1024-second value from dec_max_position_embeddings and an estimated output
token rate:

4096 tokens / ~4 tokens per second = ~1024 seconds

That appears to describe the decoder output budget rather than the earlier
input-side adapter positional constraint.

Expected behavior

For one-shot/offline input, if this adapter table is the binding hard input
constraint, input beyond that limit should be rejected before the out-of-range
embedding lookup rather than aborting the process.

For example, a result such as:

TRANSCRIBE_ERR_INPUT_TOO_LONG

would be consistent with the repository's existing hard-input-limit contract.

max_audio_ms should also represent a duration consistent with the actual
binding input-side constraint.

I am not asserting what the correct streaming API behavior should be. The
existing streaming contract has separate committed-text/truncation semantics,
so that may require different handling from the one-shot path.

Additional diagnostic

As a local diagnostic only, I added a guard around the positional/input limit.

With that guard:

81 s -> succeeds
82 s -> clean INPUT_TOO_LONG

The guard also prevents the corresponding abort when a streaming session
eventually reaches an invalid adapter position.

This is supporting evidence for the diagnosis, not a proposal that
INPUT_TOO_LONG is necessarily the correct upstream streaming behavior.

I have not included the local patch here because I wanted to report the
reproducible failure and let the maintainers determine the intended API and
implementation behavior.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions