Moonshine Streaming Tiny โ€” Mandarin Chinese

Mandarin Chinese streaming speech recognition, 27.0M parameters. Same architecture as moonshine-ai/moonshine-streaming-tiny, trained for Mandarin Chinese with a 12,288-entry Mandarin Chinese tokenizer.

Moonshine Streaming pairs a 50 Hz time-domain audio frontend with a sliding-window Transformer encoder, so it transcribes incrementally rather than waiting for an utterance to finish. It is intended for on-device use on edge-class hardware.

Checkpoint identity

This repository is a conversion of one specific training checkpoint, recorded here because the weights behind a language move as later stages win:

Checkpoint zh12k_tiny_stageC_e3_20260824T003347Z.safetensors
Stage C (read-speech mix), epoch 3
Architecture slinkier_prime_adapted
Tokenizer tokenizer_zh12k.json, vocab 12,288
Snapshot taken 2026-08-24
Parameters 27.0M

If you need reproducibility, pin the revision of this repository rather than tracking main.

Usage

pip install --upgrade transformers datasets[audio]
from transformers import MoonshineStreamingForConditionalGeneration, AutoProcessor
import torch

model = MoonshineStreamingForConditionalGeneration.from_pretrained(
    "moonshine-ai/moonshine-streaming-tiny-zh"
).eval()
processor = AutoProcessor.from_pretrained("moonshine-ai/moonshine-streaming-tiny-zh")

inputs = processor(audio, return_tensors="pt", sampling_rate=16000)

# Cap the output length. Like other seq2seq ASR models this one can fall into a
# repetition loop, and short or noisy clips are where it happens.
seq_lens = inputs.attention_mask.sum(dim=-1)
max_new_tokens = int((seq_lens * 6.5 / 16000).max().item()) + 2

generated = model.generate(**inputs, max_new_tokens=max_new_tokens)
print(processor.batch_decode(generated, skip_special_tokens=True)[0])

Pass the attention_mask. The encoder applies its per-layer sliding windows only when it is given one; called without a mask it attends over the whole utterance instead, which is a different model from the one that was trained. The processor returns the mask, so the snippet above is the safe form. The processor also pads audio to a whole number of 80-sample frames, which the frontend requires.

Architecture

Encoder 6 layers, width 320, 8 heads, sliding windows (16, 4) on the first two and last two layers and (16, 0) between
Decoder 6 layers, width 320, 8 heads, RoPE over 32 of each head's 40 dimensions
Frontend 50 Hz features, CMVN, asinh compression, two causal stride-2 convolutions
Adapter learned absolute positional embeddings before the decoder

The lookahead layers give roughly 80 ms of lookahead; the intermediate layers have none.

Training data

Trained on a large-scale automatically labeled Mandarin corpus:

  • Podcast crawl, roughly 91,700 hours, pseudo-labeled.
  • YouTube crawl, roughly 13,200 hours, pseudo-labeled.

The crawled transcripts are pseudo-labels: they were produced by running a Whisper-family teacher model over crawled audio, not by human transcription. The model therefore inherits the teacher's error modes, including its handling of proper nouns, numerals and code-switching. No human-verified transcript was used for the bulk of training.

Evaluation

Mandarin Chinese is scored on character error rate with spaces removed (cer_nospace), never WER. It is written without spaces, so tokenization differences alone can read as several hundred percent WER while the characters are correct. Every other language in this family except Japanese is scored on WER; Mandarin and Japanese are the two exceptions.

suite_zh is FLEURS Mandarin (read news) and WenetSpeech TEST_NET (internet- sourced spontaneous speech). The spontaneous panel is the one that resembles the training corpus, and it is much the harder of the two.

Seeded 400-utterance sample, batch 1

Batch 1 is the honest number for deployment. Batched evaluation zero-pads short clips up to the longest in the batch, and that trailing silence flatters the model.

Panel CER
fleurs_zh 12.78
wenetspeech_net 19.77
macro 16.276

This repository against the training checkpoint

These weights were converted from the neo training checkpoint, and the conversion was checked by measurement rather than inspection: same seeded sample, same batch size, same normalizer. A conversion that loads and emits plausible text can still have a permuted weight mapping, which only a score catches.

fleurs_zh wenetspeech_net macro
Training checkpoint 12.78 19.77 16.276
This repository 12.50 19.71 16.110

398/400 and 395/400 transcripts are byte-identical.

The quantized build we ship

The .ort package served to the Moonshine deployment library is quantized to int8 from these same weights, and scores 16.063 against 16.276 for the float checkpoint on the same sample under the same stopping rule -- a difference of -0.213, which is inside the noise of a 400-clip sample and should not be read as the quantized build being better or worse. That build is a different artifact from this repository, which is float32.

Limitations

  • Machine-labeled training data. See above; the model reproduces its teacher's mistakes as well as its strengths.
  • Repetition loops on short clips. Like other seq2seq ASR models this one can fall into a repetition loop, and short or noisy clips are where it happens. Cap the output length, as the usage snippet does.
  • Evaluated on 2 panels only. No evaluation of telephony, children's speech, heavy dialect, or noisy far-field conditions.

Out-of-scope use

Not intended for non-consensual surveillance, speaker identification, or high-stakes decisions.

License

MIT.

Downloads last month
4
Safetensors
Model size
27M params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support