GigaAM v3 (e2e CTC) β Core ML
Core ML conversion of Sber's GigaAM v3 v3_e2e_ctc for Apple silicon. Runs on the Neural
Engine, produces Russian text with punctuation and capitalization out of the box.
Attribution
| Original model | GigaAM v3 by Sber (salute-developers), MIT β weights on ai-sage/GigaAM-v3 |
| Core ML conversion | Ivan Ushakov (Inhum) |
Converted from the original PyTorch checkpoint (gigaam.load_model("v3_e2e_ctc")), not from
an intermediate ONNX export. Weights are redistributed under the original MIT license.
Files
| File | Size | SHA-256 |
|---|---|---|
gigaam_v3_e2e.mlpackage/Data/com.apple.CoreML/model.mlmodel |
469 KB | 8e0c43329503b3b2fc41f6e594f44d7b35acdd0bd61a91bf5741557d716be957 |
gigaam_v3_e2e.mlpackage/Data/com.apple.CoreML/weights/weight.bin |
421 MiB | 706153a1e14e2b26f328f394f3eded2689865eaf2d653c021006a4f3ab3a0ec6 |
gigaam-melfb.f32 |
40 KB | mel filterbank, 64 Γ 161 float32, little-endian |
gigaam-window.f32 |
1.3 KB | analysis window, 320 float32, little-endian |
gigaam-vocab.txt |
1 KB | 256 sentencepiece pieces, one per line, line number = token id |
The three small files are not optional. The model takes mel features, not audio, and emits token ids β without the filterbank, the window and the vocabulary the package cannot be used.
Input / output
| Name | Shape | Type | Notes | |
|---|---|---|---|---|
| in | features |
[1, 64, 2499] |
float32 | log-mel, fixed 25 s window |
| in | feature_lengths |
[1] |
int32 | real frame count before padding |
| out | log_softmax |
[1, 625, 257] |
float32 | log probabilities, 256 pieces + blank at id 256 |
| out | _to_copy_6 |
[1] |
int32 | encoded length, valid frames in the output |
Output names are generated by the converter and carry no meaning. Matching them by rank β the 3-D tensor is the logits β is more robust than matching by name.
Preprocessing
The feature extractor must reproduce the checkpoint's own configuration exactly:
sample rate 16000 Hz, mono
n_fft 320
win_length 320 (Hann, periodic)
hop_length 160
center false
n_mels 64 (HTK scale, 0β8000 Hz, unnormalized)
power 2
output log(clamp(mel, 1e-9, 1e9))
frames floor((samples - 320) / 160) + 1
Do not compute the mel filterbank from a formula. GigaAM ships its own filterbank inside the
checkpoint, it overwrites the standard torchaudio one on load, and it matches no standard
recipe. gigaam-melfb.f32 and gigaam-window.f32 are those exact tables.
Audio is padded with zeros to the 25 s window; the unpadded frame count goes into
feature_lengths. Longer audio has to be chunked by the caller.
Decoding
Greedy CTC: argmax per frame β collapse repeats β drop blank (id 256) β concatenate pieces,
where a leading β marks a word boundary and becomes a space. Token id 0 is <unk> and is
skipped. The sentencepiece library is not needed at run time β the flat vocabulary is enough.
How it was converted
Environment: Python 3.13, torch==2.7.*, torchaudio==2.7.*, coremltools>=9, gigaam from
its git repository. Target ct.target.macOS14, weights in fp16, format .mlpackage.
The conversion is not a one-liner, and most of the work is in getting a clean graph out of PyTorch:
- Export
forward_for_export(features, lengths)(encoder + head), notforward()β the latter expects a raw waveform and calls the preprocessor internally. - Do not call
transcribe()before tracing. Its@inference_modecaches rotary cos/sin as inference tensors and the trace fails afterwards. Warm the model up throughforward_for_exportunderno_gradinstead. - Trace inside
model.encoder.onnx_export_mode(). torch2.13 is incompatible with coremltools 9 (aten::Intbug) β use 2.7.x.torch.jit.tracefails even on 2.7 (the same integer bug in positional encoding). Usetorch.export.exportfollowed byep.run_decompositions({}).- The example input must be
.contiguous()(EXIR rejects a non-contiguous dim order), but do not put.contiguous()insideforwardβ that introduces an alias node. - The decomposed graph still contains 32
aten.aliasnodes, which coremltools cannot lower. They are no-ops: walk the FX graph,replace_all_uses_with(node.args[0]),erase_node, thenlint()andrecompile().
Dynamic input length did not survive conversion, hence the fixed 25 s window.
import gigaam, torch
import coremltools as ct
model = gigaam.load_model("v3_e2e_ctc", device="cpu"); model.eval()
wav, length = model.prepare_wav("sample_25s.wav")
with torch.no_grad():
features, feat_len = model.preprocessor(wav, length)
features = features.contiguous()
class W(torch.nn.Module):
def __init__(s, m): super().__init__(); s.m = m
def forward(s, features, feature_lengths):
return s.m.forward_for_export(features, feature_lengths.to(torch.long))
w = W(model).eval()
with model.encoder.onnx_export_mode(), torch.no_grad():
ep = torch.export.export(w, (features, feat_len.to(torch.int32)))
ep = ep.run_decompositions({})
gm = ep.graph_module
for node in list(gm.graph.nodes):
if node.op == "call_function" and "alias" in str(node.target):
node.replace_all_uses_with(node.args[0])
gm.graph.erase_node(node)
gm.graph.lint(); gm.recompile()
mlm = ct.convert(ep, minimum_deployment_target=ct.target.macOS14)
mlm.save("gigaam_v3_e2e.mlpackage")
Usage
import CoreML
let cfg = MLModelConfiguration()
cfg.computeUnits = .all // let the system pick ANE/GPU/CPU
let model = try MLModel(contentsOf: compiledURL, configuration: cfg)
// features: [1, 64, 2499] float32 log-mel, zero-padded to the 25 s window
// realFrames: frame count before padding
let lens = try MLMultiArray(shape: [1], dataType: .int32)
lens[0] = NSNumber(value: Int32(realFrames))
let out = try model.prediction(from: MLDictionaryFeatureProvider(
dictionary: ["features": features, "feature_lengths": lens]))
// 3-D output is the logits; the other one is the encoded length
A .mlpackage has to be compiled before use β either MLModel.compileModel(at:) once at
install time, or xcrun coremlcompiler compile gigaam_v3_e2e.mlpackage . ahead of time.
Measurements
MacBook Air, Apple silicon, 8 GB, 25 s of audio:
| PyTorch (CPU) | Core ML (ANE/GPU) | |
|---|---|---|
| Inference | ~1200 ms | 102 ms (β245Γ real time) |
| Artifact | 433 MB checkpoint | 422 MB .mlpackage |
| Peak RAM | ~1.1 GB | lower (not measured precisely) |
| argmax token match vs PyTorch | β | 98.4% (fp16 rounding) |
| Decoded text | reference | identical, punctuation included |
Limitations
- Fixed 25 s input window; dynamic shapes did not survive the conversion. Longer audio must be chunked, and chunk boundaries need de-duplication if the chunks overlap.
- Input is mel features, not audio. The feature extractor is the caller's responsibility and must match the tables shipped here.
- fp16 weights: 1.6% of argmax tokens differ from the PyTorch reference. Decoded text was identical on the samples tested, but this is not a guarantee.
- Russian only. English words are usually transliterated into Cyrillic.
- Greedy CTC only. No beam search, no language model, no timestamps, no diarization.
- Deployment target macOS 14. Not tested on iOS.
Used by
Voica β a macOS menu-bar dictation app; this model is its offline engine.
Support
A by-product of a personal project, published because it may be useful. Issues are read, but answers may be slow or absent, and feature requests are not accepted. There is no commitment to convert future GigaAM releases.
- Downloads last month
- 2
Model tree for inhum/gigaam-v3-coreml
Base model
ai-sage/GigaAM-v3