MOSS-Transcribe-Diarize-Whisper-Encoder

The audio encoder of OpenMOSS-Team/MOSS-Transcribe-Diarize, repackaged as a standard transformers Whisper encoder (WhisperModel config: d_model 1024, 24 layers, 80 mel bins, 307 M parameters) so that it can feed a frozen LLM through a DuplexJev connector. The weights are those of the original model's whisper_encoder, unchanged (model.encoder.*); the tokenizer files are copied from openai/whisper-small only so that the Whisper processor loads. Reloading gives outputs identical to the original encoder (max difference 0).

In a linear probe on mean-pooled last-layer states (3,000 train / 800 test clips) it separates speaker gender at 99.0% and four-way emotion at 88.2%, close to Qwen3-ASR-0.6B (99.2 / 90.9) and well above Whisper-large-v3-turbo (88.4 / 82.8).

Used by the DuplexJev-B[-Para]-MOSS-Transcribe-* connectors; Decider.from_pretrained("adventists-ai/DuplexJev-...") fetches it automatically. Source: github.com/adventists-ai/duplexjev.

License and attribution

Apache-2.0, as the original model. The weights are MOSS-Transcribe-Diarize by MOSI.AI / OpenMOSS. Please cite:

@misc{moss_transcribe_diarize_2026,
  title={MOSS Transcribe Diarize Technical Report},
  author={{MOSI.AI}},
  year={2026},
  eprint={2601.01554},
  archivePrefix={arXiv},
  primaryClass={cs.SD},
  url={https://arxiv.org/abs/2601.01554}
}
Downloads last month
15
Safetensors
Model size
0.3B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for adventists-ai/MOSS-Transcribe-Diarize-Whisper-Encoder

Finetuned
(24)
this model

Paper for adventists-ai/MOSS-Transcribe-Diarize-Whisper-Encoder