The strongest cliché reduction in the family — as a 2.8 GB download.
Gemma-4-31B-StyleTune-Voice
60% fewer clichés. 21.7% shared vocabulary. One tensor. ~2.8 GB.
The 31B is the original StyleTune — the one where Gryphe first proved a single tensor can carry a whole writing style. This repo is that voice, extracted and portable. Pair it with the Voice tool and any Gemma 4 31B GGUF you already have becomes StyleTune-voiced.
The numbers that matter
From Gryphe's card — 200 diverse roleplay prompts, greedy 0.0, versus the base instruct:
| Voice | Clichés / 100 words | Reduction | Shared trigram vocab |
|---|---|---|---|
| 31B (this one) | 1.23 → 0.52 | −60% | 21.7% |
| 12B | 1.050 → 0.463 | −56% | 16.8% |
| 26B A4B V2 | 1.141 → 0.551 | −52% | 19.9% |
The 31B posts the largest cliché reduction of the family. If you want the strongest de-slop effect and have the VRAM for a 31B, this is the one.
Two steps
# 1. Get the voice tool (one-time): https://huggingface.co/Wiself/voice
python3 voice.py path
# 2. Cast it onto any Gemma 4 31B GGUF you already have
voice cast ./gemma-4-31b-it-Q4_K_M.gguf voice.safetensors --out ./voiced/gemma-4-31b-styletune.gguf
Run it:
llama serve -m ./voiced/gemma-4-31b-styletune.gguf --jinja
One file out, nothing extra at runtime.
The technique, in one paragraph
Gryphe froze all 30 transformer layers — every attention head, every MLP — and trained only the lm_head output projection: the last stop before text appears on your screen. One overnight run on consumer hardware. All reasoning, world knowledge, and instruction following stay in the untouched tensors; the voice lives entirely in the trained one. That separation is what makes it liftable — we take the trained tensor and nothing else.
What's inside
voice.safetensors— thelm_head.weighttensor, BF16, shape[262144, 5376], ~2.8 GBvoice.json— metadata: source, dtype, shape, architecture
Bit-for-bit identical to the tensor in Gryphe/Gemma-4-31B-StyleTune.
Compatibility
| Target | Works? |
|---|---|
| Any Gemma 4 31B GGUF (any quant) | ✅ |
| Abliterated / uncensored 31B GGUFs | ✅ (delta path if it loops — see 12B card) |
| Gemma 4 other sizes (9B, 12B, 26B) | ❌ shape mismatch — use the matching voice |
| Non-Gemma architectures | ❌ |
Notes
- Precision: stored at original BF16. Casting quantizes only the head to Q8_0 (near-lossless); all other tensors byte-copied.
- Sampler tips from Gryphe: temp 1.0, MinP 0.10, DRY sampler on. Gemma 4's native chat template applies automatically.
- Verify:
voice info voice.safetensors→lm_head.weight · [262144, 5376] · BF16.
Family
All StyleTune voices, same tool, same two-step flow:
- 12B voice · 26B A4B V2 voice · 26B A4B QAT voice · 31B voice (this repo)
Source models: Gryphe/Gemma-4-12B-StyleTune · Gryphe/Gemma-4-26B-A4B-StyleTune-V2 · Gryphe/Gemma-4-31B-StyleTune
Credits
Gryphe — the StyleTune technique, the MythoMax lineage discovery that lm_head carries style, and all benchmarks cited above. Anthracite, Latitude. Tool: Voice. Gemma 4 is Google's model under Gemma terms — check the source card before sharing voiced models.
- Downloads last month
- -
Model tree for Wiself/gemma-4-31B-Styletune-Voice
Base model
google/gemma-4-31B