Fish Audio S2 Pro — int8 / int4 weights, blind-tested in Brazilian Portuguese

Built with Fish Audio. Quantized weights of fishaudio/s2-pro (revision 1de9996), with a blind listening test of four quantization arms. One arm is usable as-is, one is promising, one is here as a measured negative result.

⚠️ Licence — read this first

These are modified files of the Fish Audio Materials, redistributed under the Fish Audio Research License (full text in LICENSE.md; required notice and statement of modifications in NOTICE.txt). Research and non-commercial use only — any commercial use needs a separate licence from Fish Audio. All rights in the original model remain with 39 AI, INC. / Fish Audio. This repository is not affiliated with, sponsored by or endorsed by Fish Audio.

The files

file Slow AR (36 layers) Fast AR (4 layers) GiB verdict
s2-pro-w8.safetensors int8, per output channel int8, per output channel 4.74 use this one — tested as W8A8: no audible difference from bf16
s2-pro-w4g128.safetensors int4, asymmetric, group 128, RTN int8, per output channel 3.20 promising with bf16 activations (W4A16); broken with int8 activations (W4A8)

Original: 8.50 GiB bf16. Quantized: the 200 linear layers (wqkv, wo, w1, w2, w3). Kept in bf16: token embeddings (also the tied output head), all norms, the Fast AR output head. codec.pth, config and tokenizer are not included — they come from the official repo. Each file has a .quant.json listing every tensor and its format.

The int4 file is plain round-to-nearest without calibration (no AWQ/GPTQ) — the worst case for 4-bit.

Four arms, two files. The "A8" part is not stored in any file: it is a runtime choice (int8 activations, see activation_int8.py). So the arms map to files like this — there is no separate W4A8 file because it would be byte-identical to the int4 one:

arm tested weights file activations blind-test result
bf16 original fishaudio/s2-pro bf16 reference
W8A8 s2-pro-w8.safetensors int8 (activation_int8.py) ok — no audible difference
W4A16 s2-pro-w4g128.safetensors bf16 (default) ok — no audible difference
W4A8 s2-pro-w4g128.safetensors int8 (activation_int8.py) broken — 3/10 texts with a wrong phoneme

W8A16 (the int8 file with bf16 activations) was not tested.

How to use

The files are a compact storage format. reconstruct_bf16.py rebuilds a standard bf16 s2-pro checkpoint directory that stock fish-speech loads unchanged:

pip install torch safetensors huggingface_hub
python reconstruct_bf16.py --quant s2-pro-w8.safetensors --out ./s2-pro-w8-bf16

It downloads config.json, tokenizer and codec.pth from fishaudio/s2-pro at the pinned revision. This saves download and disk, not VRAM — the rebuilt model is bf16. The rebuilt weights are bit-identical to the checkpoints used in the test below (all 358 tensors, both files, verified).

The "A8" arms add int8 activations at runtime: activation_int8.py rounds the input of every quantized linear to int8 per row — the numerics of a per-token int8 GEMM, simulated (it does not speed anything up). A real int8 inference kernel was not part of this test.

The test

  • No LoRA, no fine-tuning — base s2-pro weights only, in all four arms.
  • Brazilian Portuguese only. English was not tested.
  • One voice reference (a community voice), one fixed seed, same text, same pipeline; only the weights / activation mode change between arms.
  • The model was fed the spoken form of each text (numbers and symbols already written out in words). So this measures what quantization does to speech, not to reading notation.

1. ASR (72 texts, 24 text families, Whisper large-v3-turbo)

arm exact + near-exact transcript wrong value (number/entity) total audio
bf16 79.2% 4/72 368.3 s
W8A8 79.2% 4/72 367.7 s
W4A8 82.0% 3/72 378.2 s
W4A16 77.8% 5/72 376.5 s

The ASR does not separate the arms (1–2 clips out of 72), and it ranked W4A8 best. The only hint: both W4 arms came out ~2.5% longer.

2. Blind listening (10 texts × 4 arms, order shuffled per text, key opened after)

# bf16 W8A8 W4A16 W4A8
1 ok ok ok ok
2 ok ok ok ok
3 ok¹ ok¹ ok¹ ok¹
4 ok ok ok ok
5 ok ok ok "setentes" instead of "setenta"
6 flicker at 6–7 s² ok ok ok
7 ok ok ok "batiu" instead of "bateu"
8 ok ok ok ok
9 ok ok ok ok
10 ok ok ok "E igual a zero" instead of "M igual a zero"

¹ All four said "registado" (the European-Portuguese form) instead of "registrado" — a trait of the base model, not of quantization. ² The only defect of the bf16 arm; it comes from sampling, not quantization.

Every audible error attributable to quantization landed on W4A8 (3/10 texts). W8A8 and W4A16: none. Voice quality and timbre were unchanged in all arms. 4-bit weights alone passed, int8 activations alone passed, the two together broke — each error is a single wrong phoneme, which is exactly what the ASR normalizes away. With 10 texts this is a strong signal, not proof (3/10 vs 0/10).

The 10 texts

The model read the spoken form; the written form is shown for reference.

# family written form spoken form given to the model
1 composite passo = 54 é o número que fechou o mês, e ninguém contestou. Vale a pena conferir 53ns antes de seguir. 14-46 ms é o número que fechou o mês, e ninguém contestou. O par certo/errado aparece o tempo todo nesse código. passo igual a cinquenta e quatro é o número que fechou o mês, e ninguém contestou. Vale a pena conferir cinquenta e três nanossegundos antes de seguir. catorze a quarenta e seis milissegundos é o número que fechou o mês, e ninguém contestou. O par certo barra errado aparece o tempo todo nesse código.
2 composite Passei 94MB para o time e ninguém questionou. O script aceita --host como argumento opcional. Ninguém reparou em 46-69 GB até o dia seguinte. A máquina antiga ficava em 192.168.15.90:80 antes da troca. Passei noventa e quatro megabytes para o time e ninguém questionou. O script aceita traço traço host como argumento opcional. Ninguém reparou em quarenta e seis a sessenta e nove gigabytes até o dia seguinte. A máquina antiga ficava em cento e noventa e dois ponto cento e sessenta e oito ponto quinze ponto noventa dois pontos oitenta antes da troca.
3 composite Deixei R$ 0,47 registrado para não esquecer depois. Mais de 1/10 dos testes passaram de primeira. A conta fechou em RTX 3060. Reserve 26-55 MB para não passar aperto. Deixei quarenta e sete centavos registrado para não esquecer depois. Mais de um décimo dos testes passaram de primeira. A conta fechou em erre tê xis trinta e sessenta. Reserve vinte e seis a cinquenta e cinco megabytes para não passar aperto.
4 phone number Salva aí o (51) 96327 0526 antes que eu esqueça. Salva aí o cinquenta e um, nove seis três dois sete, zero cinco dois seis antes que eu esqueça.
5 money E o número final ficou em R$ 3.104.670,00. E o número final ficou em três milhões cento e quatro mil seiscentos e setenta reais.
6 IP:port Ninguém reparou em 172.16.5.206:443 até o dia seguinte. Ninguém reparou em cento e setenta e dois ponto dezesseis ponto cinco ponto duzentos e seis dois pontos quatrocentos e quarenta e três até o dia seguinte.
7 lowercase format bfloat16 bateu com o que a gente tinha calculado antes. bê float dezesseis bateu com o que a gente tinha calculado antes.
8 chat abbreviation Ela mandou agr às duas da manhã. Ela mandou agora às duas da manhã.
9 version A conta fechou em CUDA 13.2.4. A conta fechou em CUDA treze ponto dois ponto quatro.
10 operator A documentação nem menciona M == zero, o que é um problema. A documentação nem menciona M igual a zero, o que é um problema.

Limitations

  • Portuguese only; English and other languages were not tested.
  • One voice, one seed; 72 texts for the ASR, 10 for the blind test, one listener.
  • Int8 activations were simulated; speed and VRAM of a real int8 kernel were not measured here.
  • The int4 variant is uncalibrated round-to-nearest; a calibrated 4-bit may do better.

Licence and attribution

Fish Audio Research License — see LICENSE.md and NOTICE.txt. Built with Fish Audio. Base model: fishaudio/s2-pro.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for JoaoZaokk/fish-s2-pro-quantized

Base model

fishaudio/s2-pro
Quantized
(12)
this model