Spark-X2.5-4B-Uncensored-GGUF

Model Description

This repository contains the GGUF formats (FP16 and Q4_K_M) of the optimized Spark-X2.5 4B uncensored architecture. The model has been completely stripped of artificial alignment layers and refusal behaviors, allowing for direct, objective, and unbounded responses to complex, creative, and reasoning-based queries.

Quantization & Imatrix Calibration

The Q4_K_M quantization was driven by a robust, highly curated Importance Matrix (imatrix). To preserve the model's structural integrity and reasoning pathways during quantization, the imatrix was calibrated using exactly 2.14 million high-quality tokens spanning:

  • Deep reasoning traces (<think> blocks)
  • Advanced mathematics and scientific queries
  • Uncensored instruction logic
  • Creative and conversational roleplay

This precise calibration ensures the quantized Q4_K_M variant retains near-FP16 fidelity, severely minimizing degradation when navigating complex logical deductions and unfiltered generative tasks.

Available Files

  • Spark-4B-F16.gguf: Uncompressed 16-bit precision base file for maximum accuracy.
  • Spark-4B-Q4_K_M-Imatrix.gguf: Highly efficient 4-bit quantization, calibrated via custom imatrix for an optimal balance of speed, VRAM usage, and structural fidelity.
Downloads last month
178
GGUF
Model size
4B params
Architecture
spark2_5
Hardware compatibility
Log In to add your hardware

4-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support