kenosistron-bf16

The full-precision (BF16) merged checkpoint of disinfozone/kenosistron. Read that card for what this model is and how to run it. This repo is for those with 230 GB of patience: re-quantizers, researchers, and anyone who wants the weights before the 5-bit haircut.

What it is, mechanically:

  • nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 (hybrid Mamba2 + Attention + MoE, 120B total / ~12B active, 256k context)
  • + the kenosistron v3 LoRA (rank 8, alpha 32, all-linear, Tinker-trained on the disinfo.zone corpus), fully merged
  • + the retrained MTP speculative head spliced over the stock one (pure-KL self-alignment; greedy acceptance rose from 61.9% to 72.1%; see "The Speculative Head" on the main card)

All 41,233 tensors verified mapped at merge time. The adapter, head, quantization imatrix, and every script needed to reproduce this checkpoint from the public base are in disinfozone/kenosistron-lora; the ready-to-run 81 GB MLX quant is disinfozone/kenosistron.

NVIDIA's original base-model documentation (bias.md, explainability.md, custom modeling code) ships alongside the weights, as it did in the source repo. Sampler guidance from the main card applies unchanged: run it hot.

Downloads last month
31
Safetensors
Model size
124B params
Tensor type
F32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for disinfozone/kenosistron-bf16

Finetuned
(18)
this model
Quantizations
1 model