emotion2vec+ base for Core ML

emotion2vec+ base by Ziyang Ma et al. (emotion2vec), converted to Core ML for on-device speech emotion recognition on iOS 18+ and macOS 15+.

It hears a voice and gives a probability for each of nine classes: angry, disgusted, fearful, happy, neutral, other, sad, surprised, unknown. A 3 s clip takes about 25 ms on an Apple-silicon GPU.

Files

Folder Weights Size Compared with the PyTorch original
Emotion2Vec8bit.mlmodelc 8-bit, per channel 89 MB same top label on every test clip; probabilities within ~0.05 on quiet audio
Emotion2Vec.mlmodelc 16-bit, 32-bit feature extractor 186 MB probabilities within 0.013

Both are compiled Core ML ML Programs, ready for MLModel(contentsOf:).

Input and output

  • audio (input): Float32 [1, N], 16 kHz mono samples in [-1, 1], with N from 8,000 to 160,000 (0.5–10 s). Pass audio at its true length and don't pad it: the model has no padding mask, so added silence changes the result. Split longer audio into chunks of 10 s or less.
  • probs (output): Float32 [1, 9], softmax probabilities in the order angry, disgusted, fearful, happy, neutral, other, sad, surprised, unknown.
  • logits (output): the same scores before the softmax.

The waveform normalisation, frame averaging and classification head are part of the model, so plain 16 kHz samples are all it needs.

Run it on CPU and GPU (.cpuAndGPU). The Neural Engine is no faster for this model, and loading it there costs a one-time compile of about 20 s.

Silence isn't neutral to this model: pure silence comes back as "surprised". Skip quiet windows (below about −42 dBFS) instead of classifying them.

Use from Swift

Emotion2VecKit, our Swift package for these files (to be published), downloads them, checks each against its SHA-256, and wraps the model:

let model = try await Emotion2Vec.load()
let result = try await model.classify(fileAt: url)
print(result.top, result.confidence)   // happy 0.82

It also resamples any audio format, handles long recordings, and has a live stream for microphone input.

Conversion

Converted from the original checkpoint with coremltools 9.0 and torch 2.7.0, then checked against FunASR's own output on speech, quiet speech, silence, noise and very short clips.

Licence and credit

These weights are a conversion of emotion2vec+ base and keep its licence, the FunASR Model Open Source License Agreement. Please credit the original work:

@article{ma2023emotion2vec,
  title={emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation},
  author={Ma, Ziyang and Zheng, Zhisheng and Ye, Jiaxin and Li, Jinchao and Gao, Zhifu and Zhang, Shiliang and Chen, Xie},
  journal={arXiv preprint arXiv:2312.15185},
  year={2023}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for happyface-studio/emotion2vec-plus-base-coreml

Finetuned
(1)
this model

Paper for happyface-studio/emotion2vec-plus-base-coreml