Instructions to use disinfozone/kenosistron-bf16 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use disinfozone/kenosistron-bf16 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="disinfozone/kenosistron-bf16", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("disinfozone/kenosistron-bf16", trust_remote_code=True) model = AutoModelForCausalLM.from_pretrained("disinfozone/kenosistron-bf16", trust_remote_code=True, device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use disinfozone/kenosistron-bf16 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "disinfozone/kenosistron-bf16" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "disinfozone/kenosistron-bf16", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/disinfozone/kenosistron-bf16
- SGLang
How to use disinfozone/kenosistron-bf16 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "disinfozone/kenosistron-bf16" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "disinfozone/kenosistron-bf16", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "disinfozone/kenosistron-bf16" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "disinfozone/kenosistron-bf16", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use disinfozone/kenosistron-bf16 with Docker Model Runner:
docker model run hf.co/disinfozone/kenosistron-bf16
kenosistron-bf16
The full-precision (BF16) merged checkpoint of disinfozone/kenosistron. Read that card for what this model is and how to run it. This repo is for those with 230 GB of patience: re-quantizers, researchers, and anyone who wants the weights before the 5-bit haircut.
What it is, mechanically:
nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16(hybrid Mamba2 + Attention + MoE, 120B total / ~12B active, 256k context)- + the kenosistron v3 LoRA (rank 8, alpha 32, all-linear, Tinker-trained on the disinfo.zone corpus), fully merged
- + the retrained MTP speculative head spliced over the stock one (pure-KL self-alignment; greedy acceptance rose from 61.9% to 72.1%; see "The Speculative Head" on the main card)
All 41,233 tensors verified mapped at merge time. The adapter, head, quantization imatrix, and every script needed to reproduce this checkpoint from the public base are in disinfozone/kenosistron-lora; the ready-to-run 81 GB MLX quant is disinfozone/kenosistron.
NVIDIA's original base-model documentation (bias.md, explainability.md, custom modeling code) ships alongside the weights, as it did in the source repo. Sampler guidance from the main card applies unchanged: run it hot.
- Downloads last month
- 31