emotion2vec+ base for Core ML
emotion2vec+ base by Ziyang Ma et al. (emotion2vec), converted to Core ML for on-device speech emotion recognition on iOS 18+ and macOS 15+.
It hears a voice and gives a probability for each of nine classes: angry, disgusted, fearful, happy, neutral, other, sad, surprised, unknown. A 3 s clip takes about 25 ms on an Apple-silicon GPU.
Files
| Folder | Weights | Size | Compared with the PyTorch original |
|---|---|---|---|
Emotion2Vec8bit.mlmodelc |
8-bit, per channel | 89 MB | same top label on every test clip; probabilities within ~0.05 on quiet audio |
Emotion2Vec.mlmodelc |
16-bit, 32-bit feature extractor | 186 MB | probabilities within 0.013 |
Both are compiled Core ML ML Programs, ready for MLModel(contentsOf:).
Input and output
audio(input): Float32[1, N], 16 kHz mono samples in [-1, 1], with N from 8,000 to 160,000 (0.5–10 s). Pass audio at its true length and don't pad it: the model has no padding mask, so added silence changes the result. Split longer audio into chunks of 10 s or less.probs(output): Float32[1, 9], softmax probabilities in the order angry, disgusted, fearful, happy, neutral, other, sad, surprised, unknown.logits(output): the same scores before the softmax.
The waveform normalisation, frame averaging and classification head are part of the model, so plain 16 kHz samples are all it needs.
Run it on CPU and GPU (.cpuAndGPU). The Neural Engine is no faster for this
model, and loading it there costs a one-time compile of about 20 s.
Silence isn't neutral to this model: pure silence comes back as "surprised". Skip quiet windows (below about −42 dBFS) instead of classifying them.
Use from Swift
Emotion2VecKit, our Swift package for these files (to be published), downloads them, checks each against its SHA-256, and wraps the model:
let model = try await Emotion2Vec.load()
let result = try await model.classify(fileAt: url)
print(result.top, result.confidence) // happy 0.82
It also resamples any audio format, handles long recordings, and has a live stream for microphone input.
Conversion
Converted from the original checkpoint with coremltools 9.0 and torch 2.7.0, then checked against FunASR's own output on speech, quiet speech, silence, noise and very short clips.
Licence and credit
These weights are a conversion of emotion2vec+ base and keep its licence, the FunASR Model Open Source License Agreement. Please credit the original work:
@article{ma2023emotion2vec,
title={emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation},
author={Ma, Ziyang and Zheng, Zhisheng and Ye, Jiaxin and Li, Jinchao and Gao, Zhifu and Zhang, Shiliang and Chen, Xie},
journal={arXiv preprint arXiv:2312.15185},
year={2023}
}
Model tree for happyface-studio/emotion2vec-plus-base-coreml
Base model
emotion2vec/emotion2vec_plus_base