Instructions to use Ninnix96/Qwengram-4B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Ninnix96/Qwengram-4B with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Ninnix96/Qwengram-4B:Q4_K_M # Run inference directly in the terminal: llama cli -hf Ninnix96/Qwengram-4B:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Ninnix96/Qwengram-4B:Q4_K_M # Run inference directly in the terminal: llama cli -hf Ninnix96/Qwengram-4B:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Ninnix96/Qwengram-4B:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf Ninnix96/Qwengram-4B:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Ninnix96/Qwengram-4B:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf Ninnix96/Qwengram-4B:Q4_K_M
Use Docker
docker model run hf.co/Ninnix96/Qwengram-4B:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use Ninnix96/Qwengram-4B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Ninnix96/Qwengram-4B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Ninnix96/Qwengram-4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Ninnix96/Qwengram-4B:Q4_K_M
- Ollama
How to use Ninnix96/Qwengram-4B with Ollama:
ollama run hf.co/Ninnix96/Qwengram-4B:Q4_K_M
- Unsloth Desktop
- Pi
How to use Ninnix96/Qwengram-4B with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Ninnix96/Qwengram-4B:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Ninnix96/Qwengram-4B:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use Ninnix96/Qwengram-4B with Docker Model Runner:
docker model run hf.co/Ninnix96/Qwengram-4B:Q4_K_M
- Lemonade
How to use Ninnix96/Qwengram-4B with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Ninnix96/Qwengram-4B:Q4_K_M
Run and chat with the model
lemonade run user.Qwengram-4B-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use Ninnix96/Qwengram-4B with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Ninnix96/Qwengram-4B:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Ninnix96/Qwengram-4B:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Ninnix96/Qwengram-4B with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Ninnix96/Qwengram-4B:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Ninnix96/Qwengram-4B:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Qwengram-4B
Frozen Qwen3.5-4B plus an R=1 reader at decoder IDX3/IDX11 (human layers 4/12, 12.5%/37.5% of 32 layers), with a learned global early multiplier and linear token-dependent later arbitration. The external Qwen3.8-Flash-Next PLE supplies the memory. The backbone and PLE remain frozen.
The early multiplier is a learned scalar bounded to [0.75, 1.75]. The later
multiplier is 0.5 * sigmoid(w @ RMSNorm(h_late) + b), computed per token.
Canonical endpoint: REAL-15M + linear750 (15,000,064 reader tokens; 749,568 calibration tokens). The matched 10M/15M study independently calibrated both readers. Calibrated 15M improves full-validation, all five domains and their mean over calibrated 10M with paired 95% NLL intervals below zero, without a statistically clear benchmark regression. Unresolved accuracy differences do not establish equivalence.
Frozen full-validation perplexity falls 2.784%, from 11.298811 to 10.984210. Raw 15M achieves a larger 2.970% reduction; calibration trades some aggregate LM gain for lower LAMBADA NLL. Matched 500K controls resolve REAL > PERMUTED > RANDOM > DISABLED. See the decision report for all four endpoints, paired intervals, controls and cross-scale provenance.
GGUF files
| File | Backbone precision |
|---|---|
| QwenGram-4B-BF16.gguf | BF16 |
| QwenGram-4B-Q8_0.gguf | Q8_0 |
| QwenGram-4B-Q6_K.gguf | Q6_K |
| QwenGram-4B-Q4_K_M.gguf | Q4_K_M |
All files contain the backbone plus 11 FP32 reader/arbiter tensors. These
extension tensors stay FP32 in every precision. reader.safetensors and
arbiter.pt contain the same canonical checkpoint separately. The PLE is a
required external file. SHA256.json records hashes and sizes.
Required PLE sidecar
Download Ivan Fioravanti's Q4_1 PLE GGUF.
Credit for this PLE conversion belongs to Ivan. Its SHA256 is
66db3ab390f4dd5063ecc89cc180f4713898577682347001bf64ab8e328527a1.
The approximately 32 GB file is mapped on the host; selected rows are
dequantized per token. It does not require a 32 GB GPU allocation.
Build and run
Use the Qwengram llama.cpp fork at commit 3616a858f2326e87ad8b48e1341a4e341b3dad73:
git clone https://github.com/Ninnix/llama.cpp-qwengram.git
cd llama.cpp-qwengram
git checkout 3616a858f2326e87ad8b48e1341a4e341b3dad73
cmake -S . -B build-qwengram-cpu -DCMAKE_BUILD_TYPE=Release -DLLAMA_BUILD_EXAMPLES=ON
cmake --build build-qwengram-cpu -j --target llama-completion
export QWENGRAM_PLE=/path/to/Qwen3.8-Flash-Next-PLE-Q4_1.gguf
build-qwengram-cpu/bin/llama-completion -m /path/to/QwenGram-4B-Q8_0.gguf -p 'The capital of France is' -n 16 -no-cnv -ngl 0
For Vulkan, build with -DGGML_VULKAN=ON. On the tested AMD BC-250, BF16, Q8_0, Q4_K_M with full Vulkan offload (-ngl 99) matched the corresponding CPU eight-token greedy continuation for The capital of France is. These short checks do not establish broad GPU parity. Q6_K differed at full offload but matched with -ngl 33, which keeps the first decoder layer on CPU. Use -ngl 33 for Q6_K on this tested device, or CPU (-ngl 0). The fallback is also a short generation check.
The fork supports the original 0.8B/2B readers at IDX2/IDX8 and the 2560-wide 4B reader at IDX3/IDX11. Hashing, row ordering and reader math are preserved. Stock upstream llama.cpp does not execute this custom reader. MTP and embedding-only inputs are unsupported. Vision has not been validated.
Frozen evaluation
Canonical REAL-15M + linear750, using the original FP8 PLE and frozen Kaggle suite.
| Metric | Frozen stock | Canonical Qwengram-4B |
|---|---|---|
| Full-validation NLL | 2.424697 | 2.396459 |
| Full-validation perplexity | 11.298811 | 10.984210 |
| General NLL | 2.530256 | 2.498815 |
| Code NLL | 1.221217 | 1.204523 |
| Math NLL | 1.212598 | 1.196319 |
| Scientific NLL | 1.894821 | 1.891064 |
| Multilingual NLL | 2.975046 | 2.950582 |
| Five-domain mean NLL | 1.966787 | 1.948261 |
| LAMBADA-1000 NLL | 1.350000 | 1.336395 |
| LAMBADA-1000 accuracy | 65.9% | 66.6% |
| HellaSwag-1000 accuracy | 54.8% | 55.1% |
All five domains improve over stock with paired NLL intervals below zero. LAMBADA NLL improves significantly; HellaSwag and LAMBADA accuracy gains are point estimates with intervals overlapping zero. Training used a single T4 with decoder activation checkpointing, with a 10.029 GiB measured peak. The study and evaluation artifacts preserve the protocol, per-item scores, paired bootstrap results and checkpoint identities. The historical cross-scale table retains the 2B canonical choice frozen when that study began.
GGUF runtime retention
The matched CPU test scores 8,128 tokens from the first 64 consecutive 256-token WikiText-2 raw test chunks, scoring the last 127 tokens per chunk. All runs use eight threads and context/batch/microbatch 256, with no warmup. Reader gain is NLL(stock) - NLL(Qwengram); retention divides each quantized gain by the BF16 gain. Paired 95% intervals use 10,000 resamples of 16 consecutive four-chunk blocks, seed 1234. The external PLE is Ivan Fioravanti's Q4_1 sidecar.
| Precision | Stock NLL | Qwengram NLL | Reader gain [95% CI] | Gain retention [95% CI] | Perplexity reduction vs stock |
|---|---|---|---|---|---|
| BF16 | 2.225750 | 2.198449 | 0.027301 [0.022650, 0.031917] | 100% | 2.69% |
| Q8_0 | 2.226990 | 2.199469 | 0.027521 [0.022892, 0.032133] | 100.8% [97.1%, 104.7%] | 2.71% |
| Q6_K | 2.230865 | 2.200943 | 0.029922 [0.023478, 0.036439] | 109.6% [99.4%, 118.6%] | 2.95% |
| Q4_K_M | 2.262792 | 2.231828 | 0.030964 [0.025884, 0.036537] | 113.4% [100.6%, 127.6%] | 3.05% |
BF16 has the lowest absolute Qwengram NLL in this test. Gain retention measures the added PLE benefit within each precision.
All 441 backbone tensors and nine tokenizer fields match the corresponding stock controls; all 11 extension tensors remain bit-exact FP32. See tensor verification, sidecar validation, generation checks and the matched report. The Q4_1 runtime test and the FP8 frozen evaluation above are separate benchmarks.
The target model revision is 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a.
Full provenance is in qwengram-4b.json, evaluation/, and runtime/.
Study-relative paths in the checkpoint definition resolve in the linked study repository.
This is an experimental text-generation release.
- Downloads last month
- 220
4-bit
6-bit
8-bit
16-bit