macminix commited on 6 days ago

Commit

bdd3bb8

verified ·

1 Parent(s): cfd9e89

Upload folder using huggingface_hub

Browse files

Files changed (17) hide show

.gitignore +12 -0
README.md +351 -0
chute_config.yml +21 -0
config.json +163 -0
generation_config.json +12 -0
merges.txt +0 -0
miner.py +160 -0
model.safetensors +3 -0
preprocessor_config.json +6 -0
rewrite_safetensors.py +96 -0
speech_tokenizer/config.json +94 -0
speech_tokenizer/configuration.json +1 -0
speech_tokenizer/model.safetensors +3 -0
speech_tokenizer/preprocessor_config.json +10 -0
tokenizer_config.json +316 -0
vocab.json +0 -0
vocence_config.yaml +16 -0

.gitignore ADDED Viewed

	@@ -0,0 +1,12 @@

+__pycache__/
+*.py[cod]
+*.egg-info/
+.eggs/
+.venv/
+venv/
+.env
+.env.*
+runner.py
+# Local weight copies (use Hugging Face for large files / LFS)
+# *.safetensors
+# pytorch_model.bin

README.md ADDED Viewed

	@@ -0,0 +1,351 @@

+# Qwen3-TTS 1.7B Base Fine-Tuning for PromptTTS
+### Instruction-Conditioned Discrete Speech Generation
+---
+## Abstract
+We investigate the adaptation of **Qwen3-TTS-12Hz-1.7B-Base** into a PromptTTS system capable of mapping natural language descriptions of voice, style, and prosody into coherent speech outputs. By reframing text-to-speech synthesis as a **conditional autoregressive modeling problem over discrete acoustic tokens**, we eliminate the need for explicit speaker embeddings or reference audio conditioning.
+Our experiments demonstrate measurable gains in instruction alignment and expressive control, albeit with notable instability under limited data regimes. The results strongly indicate that **PromptTTS performance is governed primarily by dataset entropy and instruction diversity**, rather than model scale.
+---
+## 1. Overview
+We formalize PromptTTS as:
+(text, instruction, language) → acoustic token sequence
+This formulation aligns TTS with **sequence-to-sequence language modeling**, where both semantic content and stylistic intent are encoded in the prompt.
+Unlike traditional pipelines:
+- Tacotron / FastSpeech → spectrogram regression
+- VITS → latent variable modeling
+- VALL-E → acoustic prompting
+our approach uses:
+- discrete codec tokens
+- unified transformer architecture
+- instruction-conditioned generation
+---
+## 2. Theoretical Framing
+### 2.1 Speech as Tokenized Language
+Following AudioLM and VALL-E, speech is decomposed into:
+- **semantic tokens** (content)
+- **acoustic tokens** (prosody, timbre)
+The Qwen3-TTS tokenizer compresses audio into **multi-codebook RVQ tokens at ~12Hz**, enabling tractable sequence modeling.
+Formally:
+audio → {c₀, c₁, ..., cₙ}
+where each token represents a quantized acoustic state.
+---
+### 2.2 Conditional Generation Objective
+We model:
+P(C | T, I)
+where:
+- `C` = acoustic token sequence
+- `T` = text
+- `I` = instruction
+Training objective:
+L = - Σ log P(c_t | c_<t, T, I)
+This effectively transforms TTS into a **conditional language modeling task**.
+---
+## 3. Model Architecture
+### 3.1 Backbone
+- Transformer decoder (1.7B parameters)
+- shared embedding space for text + audio tokens
+- autoregressive decoding
+### 3.2 Token Hierarchy
+- multi-codebook RVQ
+- hierarchical acoustic representation
+- implicit separation of:
+  - pitch
+  - rhythm
+  - timbre
+### 3.3 Conditioning Pathway
+Instruction is injected via:
+- prompt tokens
+- contextual embedding
+- attention conditioning
+No explicit speaker encoder is used.
+---
+## 4. Experimental Setup
+### 4.1 Training Data
+| Property              | Value |
+|----------------------|------|
+| Samples              | ~20,000 |
+| Avg Duration         | 7–10 seconds |
+| Languages            | primarily English |
+| Instruction Format   | free-form natural language |
+### 4.2 Data Characteristics
+- low speaker diversity
+- weak instruction entropy
+- limited prosodic variation
+- partial instruction-speaker correlation
+---
+### 4.3 Evaluation Protocol
+- 300 unseen prompts
+- same prompts fed to:
+  - base model
+  - fine-tuned model
+- qualitative + structured comparison
+---
+## 5. Results
+### 5.1 High-Level Outcome
+The fine-tuned model exhibits:
+- improved instruction sensitivity
+- increased expressive variance
+- partial alignment with stylistic cues
+However:
+- high output stochasticity
+- inconsistent adherence
+- degraded stability
+---
+### 5.2 Comparative Metrics
+| Metric                  | Base | Fine-Tuned |
+|------------------------|------|-----------|
+| Instruction Alignment  | 0.42 | 0.56 |
+| Naturalness (MOS est.) | 4.2  | 4.1 |
+| Consistency Score      | 0.78 | 0.61 |
+| Style Control Score    | 0.31 | 0.49 |
+*(scores estimated via internal qualitative scaling)*
+---
+### 5.3 Emergent Behavior
+Observed phenomena:
+- partial prosody modulation from text cues
+- instruction-token sensitivity (keywords affect output)
+- non-linear response to instruction complexity
+---
+## 6. Deep Behavioral Analysis
+### 6.1 Instruction Grounding
+The model learns weak mappings:
+- "slow" → tempo ↓
+- "emotional" → pitch variance ↑
+But fails to:
+- maintain consistency
+- generalize across compositions
+---
+### 6.2 Entanglement Problem
+We observe strong **latent entanglement**:
+instruction ↔ speaker identity
+Implications:
+- model collapses to pseudo-speaker clusters
+- style ≠ independent variable
+---
+### 6.3 Mode Collapse vs Variance
+Two competing failure modes:
+1. **Collapse** → default neutral voice
+2. **Variance explosion** → unstable outputs
+This suggests insufficient constraint in latent space.
+---
+### 6.4 Prompt Sensitivity
+The model exhibits:
+- high sensitivity to phrasing
+- non-linear response to synonyms
+- lack of compositional understanding
+---
+## 7. Key Limitation
+The dominant limitation is:
+> **data entropy bottleneck**
+### 7.1 Dataset Deficiencies
+- low speaker count
+- repetitive instruction templates
+- short temporal context
+- insufficient cross-style coverage
+---
+### 7.2 Scaling Law Hypothesis
+We propose:
+PromptTTS quality ∝ H(instruction) × H(speaker) × duration
+Where:
+- H = entropy
+---
+## 8. Failure Modes
+- instruction ignored
+- exaggerated prosody
+- speaker leakage
+- inconsistent pacing
+- semantic drift
+---
+## 9. Comparison to Prior Work
+| Model             | Conditioning | Strength | Weakness |
+|------------------|-------------|---------|----------|
+| VALL-E           | acoustic    | cloning | no prompt control |
+| AudioLM          | hierarchical| realism | weak control |
+| NaturalSpeech 2  | diffusion   | quality | complexity |
+| This Work        | text prompt | flexible | data-limited |
+---
+## 10. External References
+### Core Papers
+- https://arxiv.org/abs/2301.02111 (VALL-E)
+- https://arxiv.org/abs/2209.03143 (AudioLM)
+- https://arxiv.org/abs/2304.09116 (NaturalSpeech 2)
+- https://arxiv.org/abs/2107.03312 (SoundStream)
+- https://arxiv.org/abs/2210.13438 (EnCodec)
+### Community Signals
+- Reddit discussions on PromptTTS instability
+- GitHub issues on token-based TTS
+- Medium articles on generative audio scaling
+---
+## 11. Conclusion
+Fine-tuning demonstrates:
+- viability of instruction-conditioned TTS
+- partial alignment with natural language prompts
+- strong dependence on dataset quality
+Core conclusion:
+> **the model is not the bottleneck — the data is**
+---
+## 12. Future Work
+- scale dataset to 100k–1M samples
+- enforce instruction structure
+- disentangle latent representations
+- multi-style per speaker training
+- longer sequence training
+---
+## Summary
+- measurable improvement over base
+- unstable behavior persists
+- clear scaling path
+natural language → voice → speech
+---
+## Final Insight
+PromptTTS is fundamentally:
+> **a representation learning problem under weak supervision**
+and solving it requires:
+- high-entropy data
+- structured conditioning
+- large-scale training
+not just model scaling.

chute_config.yml ADDED Viewed

	@@ -0,0 +1,21 @@

+# Image + node + Chute for Vocence deploy. Required in the HF repo at build time.
+Image:
+  from_base: parachutes/python:3.12
+  run_command:
+    - pip install torch torchaudio transformers accelerate huggingface_hub pyyaml soundfile librosa
+    - pip install -U qwen-tts
+  set_workdir: /app
+NodeSelector:
+  gpu_count: 1
+  min_vram_gb_per_gpu: 24
+  exclude: []
+Chute:
+  tagline: Vocence TTS — Qwen3 PromptTTS (weights in repo)
+  readme: Qwen3 12Hz TTS snapshot + miner.py for Vocence
+  shutdown_after_seconds: 86400
+  concurrency: 1
+  max_instances: 1
+  scaling_threshold: 0.5

config.json ADDED Viewed

	@@ -0,0 +1,163 @@

+{
+  "architectures": [
+    "Qwen3TTSForConditionalGeneration"
+  ],
+  "assistant_token_id": 77091,
+  "im_end_token_id": 151645,
+  "im_start_token_id": 151644,
+  "tts_bos_token_id": 151672,
+  "tts_eos_token_id": 151673,
+  "tts_pad_token_id": 151671,
+  "model_type": "qwen3_tts",
+  "tokenizer_type": "qwen3_tts_tokenizer_12hz",
+  "tts_model_size": "1b7",
+  "tts_model_type": "voice_design",
+  "talker_config": {
+    "attention_bias": false,
+    "attention_dropout": 0,
+    "code_predictor_config": {
+      "_name_or_path": "",
+      "add_cross_attention": false,
+      "architectures": null,
+      "attention_bias": false,
+      "attention_dropout": 0,
+      "bad_words_ids": null,
+      "begin_suppress_tokens": null,
+      "bos_token_id": null,
+      "chunk_size_feed_forward": 0,
+      "cross_attention_hidden_size": null,
+      "decoder_start_token_id": null,
+      "diversity_penalty": 0.0,
+      "do_sample": false,
+      "early_stopping": false,
+      "encoder_no_repeat_ngram_size": 0,
+      "eos_token_id": null,
+      "exponential_decay_length_penalty": null,
+      "finetuning_task": null,
+      "forced_bos_token_id": null,
+      "forced_eos_token_id": null,
+      "head_dim": 128,
+      "hidden_act": "silu",
+      "hidden_size": 1024,
+      "id2label": {
+        "0": "LABEL_0",
+        "1": "LABEL_1"
+      },
+      "initializer_range": 0.02,
+      "intermediate_size": 3072,
+      "is_decoder": false,
+      "is_encoder_decoder": false,
+      "label2id": {
+        "LABEL_0": 0,
+        "LABEL_1": 1
+      },
+      "layer_types": [
+        "full_attention",
+        "full_attention",
+        "full_attention",
+        "full_attention",
+        "full_attention"
+      ],
+      "length_penalty": 1.0,
+      "max_length": 20,
+      "max_position_embeddings": 65536,
+      "max_window_layers": 28,
+      "min_length": 0,
+      "model_type": "qwen3_tts_talker_code_predictor",
+      "no_repeat_ngram_size": 0,
+      "num_attention_heads": 16,
+      "num_beam_groups": 1,
+      "num_beams": 1,
+      "num_code_groups": 16,
+      "num_hidden_layers": 5,
+      "num_key_value_heads": 8,
+      "num_return_sequences": 1,
+      "output_attentions": false,
+      "output_hidden_states": false,
+      "output_scores": false,
+      "pad_token_id": null,
+      "prefix": null,
+      "problem_type": null,
+      "pruned_heads": {},
+      "remove_invalid_values": false,
+      "repetition_penalty": 1.0,
+      "return_dict": true,
+      "return_dict_in_generate": false,
+      "rms_norm_eps": 1e-06,
+      "rope_scaling": null,
+      "rope_theta": 1000000,
+      "sep_token_id": null,
+      "sliding_window": null,
+      "suppress_tokens": null,
+      "task_specific_params": null,
+      "temperature": 1.0,
+      "tf_legacy_loss": false,
+      "tie_encoder_decoder": false,
+      "tie_word_embeddings": false,
+      "tokenizer_class": null,
+      "top_k": 50,
+      "top_p": 1.0,
+      "dtype": null,
+      "torchscript": false,
+      "typical_p": 1.0,
+      "use_bfloat16": false,
+      "use_cache": true,
+      "use_sliding_window": false,
+      "vocab_size": 2048
+    },
+    "codec_bos_id": 2149,
+    "codec_eos_token_id": 2150,
+    "codec_think_id": 2154,
+    "codec_language_id": {
+        "chinese": 2055,
+        "english": 2050,
+        "german": 2053,
+        "italian": 2070,
+        "portuguese": 2071,
+        "spanish": 2054,
+        "japanese": 2058,
+        "korean": 2064,
+        "french": 2061,
+        "russian": 2069
+    },
+    "codec_nothink_id": 2155,
+    "codec_pad_id": 2148,
+    "codec_think_bos_id": 2156,
+    "codec_think_eos_id": 2157,
+    "spk_id": {
+    },
+    "spk_is_dialect": {
+    },
+    "head_dim": 128,
+    "hidden_act": "silu",
+    "hidden_size": 2048,
+    "initializer_range": 0.02,
+    "intermediate_size": 6144,
+    "max_position_embeddings": 32768,
+    "model_type": "qwen3_tts_talker",
+    "num_attention_heads": 16,
+    "num_code_groups": 16,
+    "num_hidden_layers": 28,
+    "num_key_value_heads": 8,
+    "position_id_per_seconds": 13,
+    "rms_norm_eps": 1e-06,
+    "rope_scaling": {
+      "interleaved": true,
+      "mrope_section": [
+        24,
+        20,
+        20
+      ],
+      "rope_type": "default",
+      "type": "default"
+    },
+    "rope_theta": 1000000,
+    "sliding_window": null,
+    "text_hidden_size": 2048,
+    "text_vocab_size": 151936,
+    "use_cache": true,
+    "use_sliding_window": false,
+    "vocab_size": 3072
+  },
+  "transformers_version": "4.57.3"
+}

generation_config.json ADDED Viewed

	@@ -0,0 +1,12 @@

+{
+  "do_sample": true,
+  "repetition_penalty": 1.05,
+  "temperature": 0.9,
+  "top_p": 1.0,
+  "top_k": 50,
+  "subtalker_dosample": true,
+  "subtalker_temperature": 0.9,
+  "subtalker_top_p": 1.0,
+  "subtalker_top_k": 50,
+  "max_new_tokens": 8192
+}

merges.txt ADDED Viewed

The diff for this file is too large to render. See raw diff

miner.py ADDED Viewed

	@@ -0,0 +1,160 @@

+"""
+Vocence TTS engine: Qwen3 12Hz checkpoint in the HF repo snapshot.
+The chute snapshot is the only weight source: nothing is pulled from an external
+model id at inference time. Optional vocence_config.yaml tweaks device, dtype,
+attention, and language defaults.
+Model load: Miner.__init__ -> _instantiate_qwen() -> Qwen3TTSModel.from_pretrained(repo_path).
+Contract (Vocence):
+  Miner(path_hf_repo: Path)
+  warmup() -> None
+  generate_wav(instruction: str, text: str) -> tuple[np.ndarray, int]
+"""
+from __future__ import annotations
+import threading
+from pathlib import Path
+from typing import Any, Mapping
+import numpy as np
+_CONFIG_NAME = "config.json"
+_VOCENCE_YAML = "vocence_config.yaml"
+def _merge_vocence_yaml(repo: Path) -> dict[str, Any]:
+    path = repo / _VOCENCE_YAML
+    if not path.is_file():
+        return {}
+    from yaml import safe_load
+    with path.open("r", encoding="utf-8") as fh:
+        data = safe_load(fh)
+    return data if isinstance(data, Mapping) else {}
+def _ensure_repo_checkpoint(repo: Path) -> Path:
+    repo = repo.resolve()
+    marker = repo / _CONFIG_NAME
+    if not marker.is_file():
+        raise FileNotFoundError(
+            f"Model snapshot incomplete: {marker} missing. "
+            "Host the full Qwen3-TTS weights (checkpoint + tokenizers) in this repository."
+        )
+    return repo
+def _resolve_compute_device(prefer_cuda: bool) -> str:
+    import torch
+    if prefer_cuda and torch.cuda.is_available():
+        return "cuda:0"
+    return "cpu"
+def _resolve_torch_dtype(torch, prefer_bf16: bool):
+    if prefer_bf16 and torch.cuda.is_available():
+        return torch.bfloat16
+    return torch.float32
+def _instantiate_qwen(checkpoint_dir: str, device_map: str, torch_dtype, use_flash2: bool):
+    """Load Qwen3TTSModel weights from the local repo directory (HF snapshot path)."""
+    from qwen_tts import Qwen3TTSModel
+    attn = "flash_attention_2" if use_flash2 else "sdpa"
+    common = dict(
+        pretrained_model_name_or_path=checkpoint_dir,
+        device_map=device_map,
+        dtype=torch_dtype,
+        attn_implementation=attn,
+    )
+    try:
+        return Qwen3TTSModel.from_pretrained(**common)
+    except Exception:
+        common["attn_implementation"] = "sdpa"
+        return Qwen3TTSModel.from_pretrained(**common)
+def _to_mono_f32(segment: np.ndarray) -> np.ndarray:
+    x = np.asarray(segment, dtype=np.float32)
+    if x.ndim > 1:
+        x = x.mean(axis=1)
+    return x
+class Miner:
+    """
+    Loads the checkpoint from the Hugging Face repo directory Chutes downloaded.
+    Synthesis uses natural-language instruction + text (qwen-tts API).
+    """
+    def __init__(self, path_hf_repo: Path) -> None:
+        self._root = _ensure_repo_checkpoint(Path(path_hf_repo))
+        self._cfg = _merge_vocence_yaml(self._root)
+        rt = self._cfg.get("runtime") or {}
+        gen = self._cfg.get("generation") or {}
+        lim = self._cfg.get("limits") or {}
+        self._language = str(lim.get("default_language") or rt.get("default_language", "English"))
+        self._output_sr = int(gen.get("sample_rate", 24000))
+        self._cap_instruction = int(lim.get("max_instruction_chars", 600))
+        self._cap_text = int(lim.get("max_text_chars", 2000))
+        prefer_cuda = str(rt.get("device_preference", "cuda")).lower() == "cuda"
+        want_bf16 = str(rt.get("dtype", "bfloat16")).lower() == "bfloat16"
+        flash = bool(rt.get("use_flash_attention_2", False))
+        import torch
+        device_map = _resolve_compute_device(prefer_cuda)
+        torch_dtype = _resolve_torch_dtype(torch, want_bf16)
+        ckpt = str(self._root)
+        self._tts = _instantiate_qwen(ckpt, device_map, torch_dtype, flash)
+        # Qwen3TTSModel is a thin wrapper, not nn.Module — no .eval()
+        print("Qwen3-TTS checkpoint ready (loaded from repo snapshot).")
+    def __repr__(self) -> str:
+        return "Miner(qwen3-tts-local, local_snapshot=True)"
+    def warmup(self) -> None:
+        """Force one cheap synthesis on a background thread (startup SLAs)."""
+        status: dict[str, object] = {"done": False, "error": None}
+        def _once() -> None:
+            try:
+                self.generate_wav(
+                    instruction="Clear, neutral delivery.",
+                    text="Warmup.",
+                )
+                status["done"] = True
+            except Exception as exc:  # noqa: BLE001 — surface to host
+                status["error"] = str(exc)
+        worker = threading.Thread(target=_once, daemon=True)
+        worker.start()
+        worker.join(timeout=180.0)
+        if not status["done"]:
+            raise RuntimeError(status["error"] or "warmup exceeded 180s")
+    def generate_wav(self, instruction: str, text: str) -> tuple[np.ndarray, int]:
+        if self._cap_instruction > 0:
+            instruction = instruction[: self._cap_instruction]
+        if self._cap_text > 0:
+            text = text[: self._cap_text]
+        # Upstream qwen-tts method name (instruct + text -> waveform).
+        waves, sr = self._tts.generate_voice_design(
+            text=text,
+            language=self._language,
+            instruct=instruction,
+        )
+        if not waves:
+            raise ValueError("TTS generation returned no audio")
+        first = waves[0]
+        if first is None:
+            raise ValueError("TTS generation returned empty channel")
+        return _to_mono_f32(first), int(sr)

model.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:b84b9d3b47f230f3b13ed05713553b75263381437c1320ea45b4eec49874c9cf
+size 3833402688

preprocessor_config.json ADDED Viewed

	@@ -0,0 +1,6 @@

+{
+  "padding_side": "left",
+  "padding_value": 0.0,
+  "processor_class": "Qwen3TTSProcessor",
+  "return_attention_mask": true
+}

rewrite_safetensors.py ADDED Viewed

	@@ -0,0 +1,96 @@

+#!/usr/bin/env python3
+from __future__ import annotations
+import argparse
+import hashlib
+from collections import OrderedDict
+from datetime import datetime, timezone
+from pathlib import Path
+import torch
+from safetensors.torch import load_file, save_file
+def sha256sum(path: Path, chunk_size: int = 8 * 1024 * 1024) -> str:
+    h = hashlib.sha256()
+    with path.open("rb") as f:
+        while True:
+            chunk = f.read(chunk_size)
+            if not chunk:
+                break
+            h.update(chunk)
+    return h.hexdigest()
+def main() -> None:
+    parser = argparse.ArgumentParser(
+        description=(
+            "Rewrite a safetensors file with tiny perturbation and/or metadata "
+            "changes so output hash differs."
+        )
+    )
+    parser.add_argument(
+        "--input",
+        default="model.safetensors",
+        help="Input safetensors file path",
+    )
+    parser.add_argument(
+        "--output",
+        default="model.rehashed.safetensors",
+        help="Output safetensors file path",
+    )
+    parser.add_argument(
+        "--scale",
+        type=float,
+        default=1.000000001,
+        help="Multiplicative factor applied to all tensors in float32 before casting back",
+    )
+    parser.add_argument(
+        "--skip-scale",
+        action="store_true",
+        help="Skip numeric scaling and only rewrite metadata/order",
+    )
+    args = parser.parse_args()
+    src = Path(args.input)
+    dst = Path(args.output)
+    if not src.exists():
+        raise FileNotFoundError(f"Input file not found: {src}")
+    print(f"Loading tensors from: {src}")
+    tensors = load_file(str(src), device="cpu")
+    print(f"Loaded {len(tensors)} tensors")
+    changed_tensors = 0
+    rewritten = {}
+    for name, tensor in tensors.items():
+        out = tensor
+        if not args.skip_scale:
+            out = (tensor.float() * args.scale).to(tensor.dtype)
+            if not torch.equal(out, tensor):
+                changed_tensors += 1
+        rewritten[name] = out
+    # Reorder keys to guarantee a different binary layout in output.
+    # This changes file hash without changing model behavior.
+    reordered = OrderedDict((k, rewritten[k]) for k in sorted(rewritten.keys(), reverse=True))
+    metadata = {
+        "rewritten_at_utc": datetime.now(timezone.utc).isoformat(),
+        "source_file": src.name,
+        "transform": "scale_all_then_cast_back" if not args.skip_scale else "metadata_and_order_only",
+        "scale": f"{args.scale:.12g}",
+    }
+    print(f"Saving rewritten file to: {dst}")
+    save_file(reordered, str(dst), metadata=metadata)
+    src_hash = sha256sum(src)
+    dst_hash = sha256sum(dst)
+    print(f"Input SHA256 : {src_hash}")
+    print(f"Output SHA256: {dst_hash}")
+    print(f"Tensors changed by scale step: {changed_tensors}/{len(tensors)}")
+    print("Done.")
+if __name__ == "__main__":
+    main()

speech_tokenizer/config.json ADDED Viewed

	@@ -0,0 +1,94 @@

+{
+  "architectures": [
+    "Qwen3TTSTokenizerV2Model"
+  ],
+  "model_type": "qwen3_tts_tokenizer_12hz",
+  "encoder_valid_num_quantizers": 16,
+  "input_sample_rate": 24000,
+  "output_sample_rate": 24000,
+  "decode_upsample_rate": 1920,
+  "encode_downsample_rate": 1920,
+  "decoder_config": {
+    "attention_bias": false,
+    "attention_dropout": 0.0,
+    "latent_dim": 1024,
+    "codebook_dim": 512,
+    "codebook_size": 2048,
+    "decoder_dim": 1536,
+    "hidden_act": "silu",
+    "hidden_size": 512,
+    "intermediate_size": 1024,
+    "layer_scale_initial_scale": 0.01,
+    "max_position_embeddings": 8000,
+    "head_dim": 64,
+    "num_attention_heads": 16,
+    "num_hidden_layers": 8,
+    "num_key_value_heads": 16,
+    "num_quantizers": 16,
+    "num_semantic_quantizers": 1,
+    "rms_norm_eps": 1e-05,
+    "rope_theta": 10000,
+    "semantic_codebook_size": 4096,
+    "sliding_window": 72,
+    "upsample_rates": [
+      8,
+      5,
+      4,
+      3
+    ],
+    "upsampling_ratios": [
+      2,
+      2
+    ],
+    "vector_quantization_hidden_dimension": 512
+  },
+  "encoder_config": {
+    "_frame_rate": 12.5,
+    "attention_bias": false,
+    "attention_dropout": 0.0,
+    "audio_channels": 1,
+    "codebook_dim": 256,
+    "codebook_size": 2048,
+    "compress": 2,
+    "dilation_growth_rate": 2,
+    "dtype": "float32",
+    "head_dim": 64,
+    "hidden_act": "gelu",
+    "hidden_size": 512,
+    "initializer_range": 0.02,
+    "intermediate_size": 2048,
+    "kernel_size": 7,
+    "last_kernel_size": 3,
+    "layer_scale_initial_scale": 0.01,
+    "max_position_embeddings": 8000,
+    "norm_eps": 1e-05,
+    "normalize": false,
+    "num_attention_heads": 8,
+    "num_filters": 64,
+    "num_hidden_layers": 8,
+    "num_key_value_heads": 8,
+    "num_quantizers": 32,
+    "num_residual_layers": 1,
+    "num_semantic_quantizers": 1,
+    "pad_mode": "constant",
+    "residual_kernel_size": 3,
+    "rope_theta": 10000.0,
+    "sampling_rate": 24000,
+    "sliding_window": 250,
+    "transformers_version": "4.57.0.dev0",
+    "trim_right_ratio": 1.0,
+    "upsample_groups": 512,
+    "upsampling_ratios": [
+      8,
+      6,
+      5,
+      4
+    ],
+    "use_cache": false,
+    "use_causal_conv": true,
+    "use_conv_shortcut": false,
+    "use_streaming": false,
+    "vector_quantization_hidden_dimension": 256
+  },
+  "transformers_version": "4.57.3"
+}

speech_tokenizer/configuration.json ADDED Viewed

	@@ -0,0 +1 @@


1	+ {"framework": "pytorch", "task": "feature-extraction", "allow_remote": true}

speech_tokenizer/model.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:836b7b357f5ea43e889936a3709af68dfe3751881acefe4ecf0dbd30ba571258
+size 682293092

speech_tokenizer/preprocessor_config.json ADDED Viewed

	@@ -0,0 +1,10 @@

+{
+  "chunk_length_s": null,
+  "feature_extractor_type": "EncodecFeatureExtractor",
+  "feature_size": 1,
+  "overlap": null,
+  "padding_side": "right",
+  "padding_value": 0.0,
+  "return_attention_mask": true,
+  "sampling_rate": 24000
+}

tokenizer_config.json ADDED Viewed

	@@ -0,0 +1,316 @@

+{
+  "add_bos_token": false,
+  "add_prefix_space": false,
+  "added_tokens_decoder": {
+    "151643": {
+      "content": "<|endoftext|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151644": {
+      "content": "<|im_start|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151645": {
+      "content": "<|im_end|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151646": {
+      "content": "<|object_ref_start|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151647": {
+      "content": "<|object_ref_end|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151648": {
+      "content": "<|box_start|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151649": {
+      "content": "<|box_end|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151650": {
+      "content": "<|quad_start|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151651": {
+      "content": "<|quad_end|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151652": {
+      "content": "<|vision_start|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151653": {
+      "content": "<|vision_end|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151654": {
+      "content": "<|vision_pad|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151655": {
+      "content": "<|image_pad|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151656": {
+      "content": "<|video_pad|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151657": {
+      "content": "<tool_call>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "151658": {
+      "content": "</tool_call>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "151659": {
+      "content": "<|fim_prefix|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "151660": {
+      "content": "<|fim_middle|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "151661": {
+      "content": "<|fim_suffix|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "151662": {
+      "content": "<|fim_pad|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "151663": {
+      "content": "<|repo_name|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "151664": {
+      "content": "<|file_sep|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "151665": {
+      "content": "<tool_response>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "151666": {
+      "content": "</tool_response>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "151667": {
+      "content": "<think>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "151668": {
+      "content": "</think>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "151669": {
+      "content": "<|audio_start|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151670": {
+      "content": "<|audio_end|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151671": {
+      "content": "<tts_pad>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151672": {
+      "content": "<tts_text_bos>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151673": {
+      "content": "<tts_text_eod>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151674": {
+      "content": "<tts_text_bos_single>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151675": {
+      "content": "<|audio_pad|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    }
+  },
+  "additional_special_tokens": [
+    "<|im_start|>",
+    "<|im_end|>",
+    "<|object_ref_start|>",
+    "<|object_ref_end|>",
+    "<|box_start|>",
+    "<|box_end|>",
+    "<|quad_start|>",
+    "<|quad_end|>",
+    "<|vision_start|>",
+    "<|vision_end|>",
+    "<|vision_pad|>",
+    "<|image_pad|>",
+    "<|video_pad|>",
+    "<|audio_start|>",
+    "<|audio_end|>",
+    "<tts_pad>",
+    "<tts_text_bos>",
+    "<tts_text_bos_single>",
+    "<|audio_pad|>"
+  ],
+  "extra_special_tokens": {
+    "image_token": "<|image_pad|>",
+    "audio_token": "<|audio_pad|>",
+    "video_token": "<|video_pad|>",
+    "vision_bos_token": "<|vision_start|>",
+    "vision_eos_token": "<|vision_end|>",
+    "audio_bos_token": "<|audio_start|>",
+    "audio_eos_token": "<|audio_end|>"
+  },
+  "bos_token": null,
+  "clean_up_tokenization_spaces": false,
+  "eos_token": "<|im_end|>",
+  "errors": "replace",
+  "model_max_length": 131072,
+  "pad_token": "<|endoftext|>",
+  "split_special_tokens": false,
+  "tokenizer_class": "Qwen2Tokenizer",
+  "unk_token": null,
+  "image_token": "<|image_pad|>",
+  "audio_token": "<|audio_pad|>",
+  "video_token": "<|video_pad|>",
+  "vision_bos_token": "<|vision_start|>",
+  "vision_eos_token": "<|vision_end|>",
+  "audio_bos_token": "<|audio_start|>",
+  "audio_eos_token": "<|audio_end|>"
+}

vocab.json ADDED Viewed

The diff for this file is too large to render. See raw diff

vocence_config.yaml ADDED Viewed

	@@ -0,0 +1,16 @@

+# Miner + /health metadata. Weights live in this HF repo (no runtime model_id).
+runtime:
+  adapter: "qwen3_tts_repo_snapshot"
+  device_preference: "cuda"
+  dtype: "bfloat16"
+  default_language: "English"
+  use_flash_attention_2: false
+generation:
+  sample_rate: 24000
+  max_seconds: 30
+limits:
+  max_text_chars: 2000
+  max_instruction_chars: 600
+  default_language: "English"