Step-Audio-EditX β€” int8 MLX bundle

Step-Audio-EditX (StepFun, Apache-2.0): a 3 B audio LLM that re-delivers an existing take β€” emotion, speaking style, inserted paralinguistics (laughter, sighs, breaths), denoise, silence trim β€” in the same voice, plus zero-shot cloning. This is the bf16 bundle with the step1 LM quantised to 8 bits (affine, group 64, every Linear and the embeddings β€” model.safetensors 3.75 GB instead of 7.06 GB); the tokenizers, flow, HiFT and CAM++ are the bf16 files unchanged. config.json carries "quantization": {"bits": 8, "group_size": 64, "mode": "affine"}.

The quantised weights were written by the Swift consumer itself from the bf16 bundle (editx-gates --write-int8-bundle, quantised on MLX's CPU stream), so loading this bundle reproduces quantise-at-load exactly.

file component source
model.safetensors (3.75 GB, int8) + config.json step1 LM β€” 32 Γ— 3072, 48 heads / 4 KV groups, sqrt-ALiBi stepfun-ai/Step-Audio-EditX
vq02.safetensors, vq06.safetensors, step-audio-tokenizer-assets.safetensors + configs the dual tokenizer (bf16) FunASR Paraformer (FunASR model licence) / stepfun-ai/Step-Audio-Tokenizer
flow-model.safetensors, flow-conditioner.safetensors, hift.safetensors, campplus.safetensors + configs flow, vocoder, speaker (bf16) StepFun-trained / Alibaba CAM++ (Apache-2.0)
tokenizer.json, tokenizer_config.json, frontend-config.json text tokenizer, mel front end derived from tokenizer.model

Consumers

  • Swift: xocialize/mlx-step-audio-editx-swift β‰₯ 0.1.1 β€” EditXPipeline.load(bundle:) reads the quantisation from config.json.
  • Python: mlx-speech loads int8 LM bundles of this layout (its --bits 8 converter writes the same keys).

Measured (M5 Max, the Swift consumer, the E19 evaluation cues)

  • Resident 6.2 GB phys_footprint after load (MLX active 5.7 GB); MLX peak 7.3–7.7 GB per edit; real-time factor 0.51–0.53 (bf16: 0.72–0.75); loads in 0.5 s.
  • Content survives every edit (ASR error 0.000 on the four test edits), speaker similarity 0.83–0.97 to the source, DNSMOS 3.45–3.49 β€” the same bars as bf16 and the torch reference. Sampled tokens differ from bf16's at near-ties, as any quantised tier's do.

Licence

Apache-2.0 (StepFun) for the LM, flow, HiFT and the S3 tokenizer; the Paraformer encoder under the FunASR model licence; CAM++ Apache-2.0 (Alibaba / 3D-Speaker). Quantised and re-hosted by xocialize for mlx-community.

Downloads last month
-
Safetensors
Model size
4B params
Tensor type
U32
Β·
BF16
Β·
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for mlx-community/Step-Audio-EditX-8bit

Finetuned
(5)
this model