Instructions to use mlx-community/Step-Audio-EditX-8bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use mlx-community/Step-Audio-EditX-8bit with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] hf download mlx-community/Step-Audio-EditX-8bit --local-dir Step-Audio-EditX-8bit
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
Step-Audio-EditX β int8 MLX bundle
Step-Audio-EditX (StepFun, Apache-2.0): a 3 B audio LLM that
re-delivers an existing take β emotion, speaking style, inserted paralinguistics (laughter, sighs, breaths), denoise,
silence trim β in the same voice, plus zero-shot cloning. This is the
bf16 bundle with the step1 LM quantised to 8 bits
(affine, group 64, every Linear and the embeddings β model.safetensors 3.75 GB instead of 7.06 GB); the tokenizers,
flow, HiFT and CAM++ are the bf16 files unchanged. config.json carries
"quantization": {"bits": 8, "group_size": 64, "mode": "affine"}.
The quantised weights were written by the Swift consumer itself from the bf16 bundle (editx-gates --write-int8-bundle, quantised on MLX's CPU stream), so loading this bundle reproduces quantise-at-load exactly.
| file | component | source |
|---|---|---|
model.safetensors (3.75 GB, int8) + config.json |
step1 LM β 32 Γ 3072, 48 heads / 4 KV groups, sqrt-ALiBi | stepfun-ai/Step-Audio-EditX |
vq02.safetensors, vq06.safetensors, step-audio-tokenizer-assets.safetensors + configs |
the dual tokenizer (bf16) | FunASR Paraformer (FunASR model licence) / stepfun-ai/Step-Audio-Tokenizer |
flow-model.safetensors, flow-conditioner.safetensors, hift.safetensors, campplus.safetensors + configs |
flow, vocoder, speaker (bf16) | StepFun-trained / Alibaba CAM++ (Apache-2.0) |
tokenizer.json, tokenizer_config.json, frontend-config.json |
text tokenizer, mel front end | derived from tokenizer.model |
Consumers
- Swift: xocialize/mlx-step-audio-editx-swift β₯ 0.1.1 β
EditXPipeline.load(bundle:)reads the quantisation fromconfig.json. - Python:
mlx-speechloads int8 LM bundles of this layout (its--bits 8converter writes the same keys).
Measured (M5 Max, the Swift consumer, the E19 evaluation cues)
- Resident 6.2 GB phys_footprint after load (MLX active 5.7 GB); MLX peak 7.3β7.7 GB per edit; real-time factor 0.51β0.53 (bf16: 0.72β0.75); loads in 0.5 s.
- Content survives every edit (ASR error 0.000 on the four test edits), speaker similarity 0.83β0.97 to the source, DNSMOS 3.45β3.49 β the same bars as bf16 and the torch reference. Sampled tokens differ from bf16's at near-ties, as any quantised tier's do.
Licence
Apache-2.0 (StepFun) for the LM, flow, HiFT and the S3 tokenizer; the Paraformer encoder under the FunASR model licence; CAM++ Apache-2.0 (Alibaba / 3D-Speaker). Quantised and re-hosted by xocialize for mlx-community.
- Downloads last month
- -
Quantized
Model tree for mlx-community/Step-Audio-EditX-8bit
Base model
stepfun-ai/Step-Audio-EditX