V1 · Fish Audio S2-Pro · NVFP4 Balanced

A complete Blackwell-ready S2-Pro download: mixed native NVFP4/MXFP8 transformer weights, BF16 codec, tokenizer, pinned Fish Speech source, API server, web UI, and reproducibility checks.

Balanced V1 prioritizes useful English quantization fidelity and VRAM reduction. This is not yet an XPO3 release.

An XPO3 version is coming soon. Follow ajh-code on Hugging Face and Arands.com for release updates.

Original S2-Pro · Fish Speech · Arands.com · updates


Download

Component Purpose Size
model-*.safetensors Complete mixed NVFP4/MXFP8 transformer checkpoint 4.90 GB
codec.pth Complete BF16 S2-Pro DAC codec 1.87 GB
Tokenizer, runtime, and pinned source No separate base-model or codec download ~24 MB
Complete repository Weights, codec, runtime, source, and metadata 6.80 GB

All model and codec weights required by the server are in this repository. The root config.json preserves the S2-Pro architecture metadata and adds the mixed-precision policy, while Hugging Face metadata records this repository as a quantization of fishaudio/s2-pro.

Quick start

Tested on Linux x86-64, Python 3.12, CUDA 13.0, PyTorch 2.11.0+cu130, comfy-kitchen==0.2.22, and NVIDIA Blackwell SM120. The current native path is for GeForce RTX 50-series/SM120 GPUs; it is not a generic CUDA fallback.

hf download ajh-code/Fish-Audio-S2-Pro-NVFP4-Balanced \
  --local-dir fish-audio-s2-pro-nvfp4-balanced
cd fish-audio-s2-pro-nvfp4-balanced
./install.sh
./launch.sh

Open http://127.0.0.1:8080/ui for the bundled zero-shot web interface. The API listens on all interfaces by default; set TTS_HOST=127.0.0.1 if it should not be reachable from the local network. Protect or firewall the service before exposing it beyond a trusted network.

Docker Compose is the recommended clean deployment when Docker, the NVIDIA Container Toolkit, and a compatible driver are already configured:

docker compose up --build

This path passed a clean outer-Docker build, SM120 runtime launch, route/UI checks, and a zero-shot API smoke test on an RTX 5080 with CUDA 13.0.

The image excludes the 6.8 GB model payload and mounts the downloaded repository read-only, so rebuilding the runtime does not duplicate the weights inside the image.

Zero-shot voice cloning

Use a clean, consented 10–30 second reference with one speaker and supply its exact transcript:

python client.py \
  --url http://127.0.0.1:8080/v1/tts \
  --reference-audio reference.wav \
  --reference-text "The exact words spoken in reference.wav." \
  --text "A few notes as this story begins." \
  --seed 42 \
  --output result.wav

Equivalent JSON API call in Python:

import base64
from pathlib import Path

import requests

payload = {
    "text": "A few notes as this story begins.",
    "references": [{
        "audio": base64.b64encode(Path("reference.wav").read_bytes()).decode(),
        "text": "The exact words spoken in reference.wav.",
    }],
    "reference_id": None,
    "format": "wav",
    "streaming": False,
    "normalize": True,
    "max_new_tokens": 1024,
    "chunk_length": 200,
    "top_p": 0.9,
    "temperature": 0.9,
    "repetition_penalty": 1.1,
    "seed": 42,
    "use_memory_cache": "off",
}
response = requests.post("http://127.0.0.1:8080/v1/tts", json=payload, timeout=600)
response.raise_for_status()
Path("result.wav").write_bytes(response.content)

Useful endpoints:

Endpoint Purpose
GET /ui Bundled zero-shot web interface
GET /v1/health Service health
GET /v1/model Active release, quantization, and sampling metadata
POST /v1/tts Fish Speech-compatible TTS request; returns audio

Set TTS_API_KEY before launch to require bearer authentication. For an API key named secret, send Authorization: Bearer secret.

Quantization policy

S2-Pro has 180 projections in its 36-layer slow transformer. V1 uses:

Scope Stored/executed precision Count
Gate/up in layers 3–32 packed NVFP4 E2M1; W4A16 at M=1, W4A4 above M=1 60
Other slow-transformer projections native dynamic MXFP8 W8A8 120
Embeddings, tied text output, fast transformer/output, norms, RoPE, KV cache, sampling BF16/original precision
DAC codec arithmetic BF16

The English calibration scale is folded into the selected norm and packed gate/up tensors. It adds no runtime tensor or operation. The original BF16 slow-projection weights are not retained as a second copy. V1's transformer checkpoint is 46.25% smaller than the original transformer shards.

This is intentionally described as a mixed NVFP4/MXFP8 checkpoint. It is not a claim that every operation, activation, or weight in the end-to-end TTS stack runs at FP4.

Measured performance

Measurements below are local RTX 5080 results with the bundled compact BF16 codec path and a 3072-token cache. They are not universal performance claims.

Measurement V1 result
Loaded PyTorch allocation 5.350 GiB
15-sample short zero-shot peak 5.607–5.769 GiB
Held-out ~39-second generation peak 7.561 GiB
Median short/control real-time factor about 1.05–1.06
Median short/control time to first playable audio about 5–6.5 s
Semantic generation throughput about 20.4–20.6 frames/s

The matching compact-runtime BF16 control loaded at 9.273 GiB, so V1 reduced loaded PyTorch allocation by 42.31%. V1 is near real time on the RTX 5080, but the current ordinary single-speaker API waits for a complete semantic segment before playable audio. This release therefore does not claim agent-grade low-latency streaming. An RTX 5060 Ti focused zero-shot run measured about 2.06 RTF and is not a real-time path.

Limited blind voice-cloning test

The first blind comparison is encouraging, but deliberately small. It used one listener, one consented English reference speaker, and four matched BF16/V1 pairs: conversational, reflective, question-shaped, and long narrative prompts at seeds 7, 17, 123, and 42. Both models used the same reference, runtime path, temperature=0.9, top_p=0.9, and top_k=30.

Blind result BF16 Balanced V1
Speaker-likeness scores all four 5/5 all four 5/5
Mean reference-style likeness 4.50 / 5 4.50 / 5
Mean naturalness 3.75 / 5 3.75 / 5
Pair preference 1 2

The fourth pair was tied. The only severe artifact reported in the set was a deterministic BF16 pitch squeak in the long seed-42 sample; its V1 counterpart did not contain that excursion.

This test suggests that the quant did not cause a detectable speaker-identity loss for that reference. It is not a general MOS study or broad cloning qualification: more listeners, speakers, accents, recording conditions, and languages are still needed. The release therefore reports the result without claiming parity in every voice-cloning setting.

Validated scope

Gate Result
Native execution 60 NVFP4 and 120 MXFP8 projections execute through native SM120 paths
Standalone packaging Fresh load from these shards, without BF16 source projections, matched a frozen 64-frame code canary bit exactly
English automated gates Passed fixed-input signal/spectral, ASR, speaker-embedding, short/control, and held-out long-termination screens
Blind English clone identity Limited four-pair test above: every BF16 and V1 sample scored 5/5 speaker likeness; preferences were V1 2, BF16 1, tie 1
Multilingual Not qualified; use an MXFP8 or BF16 model when language coverage matters
Hardware NVIDIA Blackwell SM120 only in V1

The blind result supports quantization fidelity for that English reference; it does not establish universal cloning quality across voices, recording conditions, accents, or languages.

Known limitations

  • Fish S2-Pro itself sounded substantially flatter and less expressive than VoxCPM2 in our reference comparison. BF16 shared this behavior, so V1 does not treat it as NVFP4-specific damage and does not claim to fix it.
  • Inline emotion/style instructions change output trajectories but did not reliably repair the perceived flatness in the tested voice.
  • The objective speaker embedding saturated near 0.99 and failed to predict human preference; human listening remains required for new voices.
  • One matched long BF16 sample produced a deterministic pitch squeak while its V1 counterpart did not. This is evidence from one seed, not a claim that V1 is generally more artifact-free than BF16.
  • Long-form peak memory is materially higher than loaded memory. Do not market V1 as a sub-6-GiB operational model for arbitrary request lengths.

These bounded claims are why this package is Balanced V1, not an XPO3 speed/quality/size release. Follow ajh-code for the upcoming XPO3 version.

Validate the download

python validate_release.py

MANIFEST.json records the byte size and SHA-256 of every distributed file except itself. Validation also checks the safetensors index/header mapping, the 60/120 NVFP4/MXFP8 tensor counts, source pins, license/notice files, and runtime payload. Hashing the 6.8 GB package takes a little while.

For an additional hash check every time the service loads:

TTS_VERIFY_CHECKSUMS=1 ./launch.sh

License and attribution

Built with Fish Audio. This derivative is governed by the Fish Audio Research License. Research and non-commercial use are permitted subject to its terms. Commercial use requires a separate written license from Fish Audio; no commercial rights are granted by this repository. See Notice for the required attribution and exact change statement, and THIRD_PARTY_NOTICES.md for runtime dependencies.

Use only voices and recordings you have the right and consent to use.

Downloads last month
-
Safetensors
Model size
4B params
Tensor type
I32
·
BF16
·
F8_E4M3
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ajh-code/Fish-Audio-S2-Pro-NVFP4-Balanced

Base model

fishaudio/s2-pro
Quantized
(11)
this model