GemStrike 31B — EXL3 4.00 bpw
GemStrike is a Gemma 4 31B creative-writing and roleplay fine-tune. This release is the EXL3 4.00 bpw conversion of the merged BF16 model.
The model was produced by training a language-only LoRA on the BF16 safetensors release of Huihui Gemma 4 31B IT QAT unquantized abliterated, merging that adapter into the BF16 base, and converting the merged checkpoint with ExLlamaV3.
Format
- Quantization: EXL3, 4.00 target bits per weight
- Language-model head: 6-bit
- Codebook:
mul1 - Vision tower: preserved in BF16
- Calibration: 1,407 rows × 2,048 tokens from a deterministic 512-example creative-writing/roleplay calibration selection
- Storage: three safetensors shards
This is not a GGUF model. Use an EXL3-compatible ExLlamaV3 runtime, such as ExLlamaV3 itself or a compatible serving frontend. It is not intended for llama.cpp.
Fine-tune details
- Training examples: 4,546
- Held-out validation examples: 264
- Final optimizer step: 4,546
- LoRA rank: 16
- LoRA alpha: 32
- LoRA dropout: 0
- Adapter targets: language-model attention and MLP projections only
- Vision and other non-language components were frozen during LoRA training
The training corpus was the ordered creative-writing/roleplay split used for the associated GemStrike training run.
Runtime validation
This published directory passed the following local acceptance checks:
- Complete EXL3 tensor and safetensors-index coverage
- 411 EXL3 linear modules
- All 355 vision tensors present in BF16
- Deterministic text generation with ExLlamaV3 using two-GPU tensor parallelism
- Successful unload and GPU-memory return-to-baseline checks
The runtime generation check used a 4,096-token cache and a fixed held-out text prompt. The preserved vision tower was audited statically, but image-input generation was not separately validated as part of this publication gate.
Prompting
Use the included Gemma 4 chat template and tokenizer metadata. For best results, send structured chat messages rather than manually inventing control tokens.
This fine-tune emphasizes creative prose and roleplay. It can produce mature, controversial, or otherwise sensitive material, especially because its base model is abliterated. Applications should apply their own moderation and review appropriate to their audience and use case.
License and attribution
Use is subject to the Gemma Terms of Use. This derivative also retains the lineage and conditions of its Huihui base model. Review those terms before redistribution or deployment.
- Downloads last month
- 168
Model tree for UltimateIntent/GemStrike-31B-EXL3-4.00bpw
Base model
google/gemma-4-31B