OpenVoice V2 — Core ML
Voice Cloning
Zero-shot voice conversion. Clone a speaker from ~10s reference audio.

Core ML conversion of myshell-ai/OpenVoice for on-device inference on iPhone, iPad and Mac. Converted with coremltools; the packages are stateless, so all sequencing and buffering lives in your Swift code.
| Task | audio to audio |
| Upstream | myshell-ai/OpenVoice |
| Packages | 2 |
| Download size | 58 MB |
| Minimum iOS | 17.0 |
| Peak RAM | ~500 MB |
Files
| File | Size | Compute units | SHA-256 |
|---|---|---|---|
OpenVoice_SpeakerEncoder.mlpackage.zip |
1 MB | cpuAndGPU |
c3f2a96aaf5ecb5c… |
OpenVoice_VoiceConverter.mlpackage.zip |
57 MB | cpuAndGPU |
ef3ce8a2d1564aef… |
| Total | 58 MB |
compute_units is not a suggestion -- it is the configuration the conversion was verified against. Moving a package to a different compute unit can silently change the numerics (FP16 attention overflow) or crash on the GPU.
Download
hf download mlboydaisuke/coreml-zoo --include "openvoice/*" --local-dir ./openvoice
unzip './openvoice/openvoice/*.zip' -d ./openvoice
Use in Swift
import CoreML
let config = MLModelConfiguration()
config.computeUnits = .cpuAndGPU // as converted — see the table above
// Unzip the .mlpackage, drop it into your Xcode target and Xcode compiles it
// at build time:
let model = try OpenVoice_SpeakerEncoder(configuration: config)
// ...or compile a downloaded .mlpackage at runtime:
let compiled = try await MLModel.compileModel(at: mlpackageURL)
let model = try MLModel(contentsOf: compiled, configuration: config)
This model is split into 2 Core ML packages that are driven in sequence from Swift. Load them one at a time, copy the outputs out of the
MLMultiArraybuffers and release each model before loading the next — two large Core ML models resident at once will OOM on an iPhone.
Demo
- Sample app —
sample_apps/OpenVoiceDemo, a standalone SwiftUI project. - Models Zoo — this model is downloadable and runnable inside the Models Zoo app on the App Store, no build required.
Conversion
- Script:
convert_openvoice.py - Pitfalls hit during conversion (FP16 overflow, ANE buffer limits, stride handling):
docs/coreml_conversion_notes.md - Model index: CoreML-Models
License
The conversion inherits the upstream license: MIT.
Credits
- Upstream authors: myshell-ai/OpenVoice, 2023
- Core ML conversion: john-rocky (Daisuke Majima)
- Downloads last month
- 9