V1 · Fish Audio S2-Pro · NVFP4 Balanced
A complete Blackwell-ready S2-Pro download: mixed native NVFP4/MXFP8 transformer weights, BF16 codec, tokenizer, pinned Fish Speech source, API server, web UI, and reproducibility checks.
Balanced V1 prioritizes useful English quantization fidelity and VRAM reduction. This is not yet an XPO3 release.
An XPO3 version is coming soon. Follow ajh-code on Hugging Face and Arands.com for release updates.
Original S2-Pro · Fish Speech · Arands.com · updates
Download
| Component | Purpose | Size |
|---|---|---|
model-*.safetensors |
Complete mixed NVFP4/MXFP8 transformer checkpoint | 4.90 GB |
codec.pth |
Complete BF16 S2-Pro DAC codec | 1.87 GB |
| Tokenizer, runtime, and pinned source | No separate base-model or codec download | ~24 MB |
| Complete repository | Weights, codec, runtime, source, and metadata | 6.80 GB |
All model and codec weights required by the server are in this repository.
The root config.json preserves the S2-Pro architecture metadata and adds the
mixed-precision policy, while Hugging Face metadata records this repository as
a quantization of fishaudio/s2-pro.
Quick start
Tested on Linux x86-64, Python 3.12, CUDA 13.0, PyTorch 2.11.0+cu130,
comfy-kitchen==0.2.22, and NVIDIA Blackwell SM120. The current native path is
for GeForce RTX 50-series/SM120 GPUs; it is not a generic CUDA fallback.
hf download ajh-code/Fish-Audio-S2-Pro-NVFP4-Balanced \
--local-dir fish-audio-s2-pro-nvfp4-balanced
cd fish-audio-s2-pro-nvfp4-balanced
./install.sh
./launch.sh
Open http://127.0.0.1:8080/ui for the bundled zero-shot web interface. The
API listens on all interfaces by default; set TTS_HOST=127.0.0.1 if it should
not be reachable from the local network. Protect or firewall the service before
exposing it beyond a trusted network.
Docker Compose is the recommended clean deployment when Docker, the NVIDIA Container Toolkit, and a compatible driver are already configured:
docker compose up --build
This path passed a clean outer-Docker build, SM120 runtime launch, route/UI checks, and a zero-shot API smoke test on an RTX 5080 with CUDA 13.0.
The image excludes the 6.8 GB model payload and mounts the downloaded repository read-only, so rebuilding the runtime does not duplicate the weights inside the image.
Zero-shot voice cloning
Use a clean, consented 10–30 second reference with one speaker and supply its exact transcript:
python client.py \
--url http://127.0.0.1:8080/v1/tts \
--reference-audio reference.wav \
--reference-text "The exact words spoken in reference.wav." \
--text "A few notes as this story begins." \
--seed 42 \
--output result.wav
Equivalent JSON API call in Python:
import base64
from pathlib import Path
import requests
payload = {
"text": "A few notes as this story begins.",
"references": [{
"audio": base64.b64encode(Path("reference.wav").read_bytes()).decode(),
"text": "The exact words spoken in reference.wav.",
}],
"reference_id": None,
"format": "wav",
"streaming": False,
"normalize": True,
"max_new_tokens": 1024,
"chunk_length": 200,
"top_p": 0.9,
"temperature": 0.9,
"repetition_penalty": 1.1,
"seed": 42,
"use_memory_cache": "off",
}
response = requests.post("http://127.0.0.1:8080/v1/tts", json=payload, timeout=600)
response.raise_for_status()
Path("result.wav").write_bytes(response.content)
Useful endpoints:
| Endpoint | Purpose |
|---|---|
GET /ui |
Bundled zero-shot web interface |
GET /v1/health |
Service health |
GET /v1/model |
Active release, quantization, and sampling metadata |
POST /v1/tts |
Fish Speech-compatible TTS request; returns audio |
Set TTS_API_KEY before launch to require bearer authentication. For an API
key named secret, send Authorization: Bearer secret.
Quantization policy
S2-Pro has 180 projections in its 36-layer slow transformer. V1 uses:
| Scope | Stored/executed precision | Count |
|---|---|---|
| Gate/up in layers 3–32 | packed NVFP4 E2M1; W4A16 at M=1, W4A4 above M=1 |
60 |
| Other slow-transformer projections | native dynamic MXFP8 W8A8 | 120 |
| Embeddings, tied text output, fast transformer/output, norms, RoPE, KV cache, sampling | BF16/original precision | — |
| DAC codec arithmetic | BF16 | — |
The English calibration scale is folded into the selected norm and packed gate/up tensors. It adds no runtime tensor or operation. The original BF16 slow-projection weights are not retained as a second copy. V1's transformer checkpoint is 46.25% smaller than the original transformer shards.
This is intentionally described as a mixed NVFP4/MXFP8 checkpoint. It is not a claim that every operation, activation, or weight in the end-to-end TTS stack runs at FP4.
Measured performance
Measurements below are local RTX 5080 results with the bundled compact BF16 codec path and a 3072-token cache. They are not universal performance claims.
| Measurement | V1 result |
|---|---|
| Loaded PyTorch allocation | 5.350 GiB |
| 15-sample short zero-shot peak | 5.607–5.769 GiB |
| Held-out ~39-second generation peak | 7.561 GiB |
| Median short/control real-time factor | about 1.05–1.06 |
| Median short/control time to first playable audio | about 5–6.5 s |
| Semantic generation throughput | about 20.4–20.6 frames/s |
The matching compact-runtime BF16 control loaded at 9.273 GiB, so V1 reduced loaded PyTorch allocation by 42.31%. V1 is near real time on the RTX 5080, but the current ordinary single-speaker API waits for a complete semantic segment before playable audio. This release therefore does not claim agent-grade low-latency streaming. An RTX 5060 Ti focused zero-shot run measured about 2.06 RTF and is not a real-time path.
Limited blind voice-cloning test
The first blind comparison is encouraging, but deliberately small. It used one
listener, one consented English reference speaker, and four matched BF16/V1
pairs: conversational, reflective, question-shaped, and long narrative prompts
at seeds 7, 17, 123, and 42. Both models used the same reference, runtime path,
temperature=0.9, top_p=0.9, and top_k=30.
| Blind result | BF16 | Balanced V1 |
|---|---|---|
| Speaker-likeness scores | all four 5/5 | all four 5/5 |
| Mean reference-style likeness | 4.50 / 5 | 4.50 / 5 |
| Mean naturalness | 3.75 / 5 | 3.75 / 5 |
| Pair preference | 1 | 2 |
The fourth pair was tied. The only severe artifact reported in the set was a deterministic BF16 pitch squeak in the long seed-42 sample; its V1 counterpart did not contain that excursion.
This test suggests that the quant did not cause a detectable speaker-identity loss for that reference. It is not a general MOS study or broad cloning qualification: more listeners, speakers, accents, recording conditions, and languages are still needed. The release therefore reports the result without claiming parity in every voice-cloning setting.
Validated scope
| Gate | Result |
|---|---|
| Native execution | 60 NVFP4 and 120 MXFP8 projections execute through native SM120 paths |
| Standalone packaging | Fresh load from these shards, without BF16 source projections, matched a frozen 64-frame code canary bit exactly |
| English automated gates | Passed fixed-input signal/spectral, ASR, speaker-embedding, short/control, and held-out long-termination screens |
| Blind English clone identity | Limited four-pair test above: every BF16 and V1 sample scored 5/5 speaker likeness; preferences were V1 2, BF16 1, tie 1 |
| Multilingual | Not qualified; use an MXFP8 or BF16 model when language coverage matters |
| Hardware | NVIDIA Blackwell SM120 only in V1 |
The blind result supports quantization fidelity for that English reference; it does not establish universal cloning quality across voices, recording conditions, accents, or languages.
Known limitations
- Fish S2-Pro itself sounded substantially flatter and less expressive than VoxCPM2 in our reference comparison. BF16 shared this behavior, so V1 does not treat it as NVFP4-specific damage and does not claim to fix it.
- Inline emotion/style instructions change output trajectories but did not reliably repair the perceived flatness in the tested voice.
- The objective speaker embedding saturated near 0.99 and failed to predict human preference; human listening remains required for new voices.
- One matched long BF16 sample produced a deterministic pitch squeak while its V1 counterpart did not. This is evidence from one seed, not a claim that V1 is generally more artifact-free than BF16.
- Long-form peak memory is materially higher than loaded memory. Do not market V1 as a sub-6-GiB operational model for arbitrary request lengths.
These bounded claims are why this package is Balanced V1, not an XPO3
speed/quality/size release. Follow
ajh-code for the upcoming XPO3 version.
Validate the download
python validate_release.py
MANIFEST.json records the byte size and SHA-256 of every distributed file
except itself. Validation also checks the safetensors index/header mapping,
the 60/120 NVFP4/MXFP8 tensor counts, source pins, license/notice files, and
runtime payload. Hashing the 6.8 GB package takes a little while.
For an additional hash check every time the service loads:
TTS_VERIFY_CHECKSUMS=1 ./launch.sh
License and attribution
Built with Fish Audio. This derivative is governed by the
Fish Audio Research License. Research and non-commercial use
are permitted subject to its terms. Commercial use requires a separate
written license from Fish Audio; no commercial rights are granted by this
repository. See Notice for the required attribution and exact
change statement, and THIRD_PARTY_NOTICES.md for
runtime dependencies.
Use only voices and recordings you have the right and consent to use.
- Downloads last month
- -
Model tree for ajh-code/Fish-Audio-S2-Pro-NVFP4-Balanced
Base model
fishaudio/s2-pro