Instructions to use hiteshluke/arbiter-4b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use hiteshluke/arbiter-4b with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("unsloth/gemma-3-4b-it-unsloth-bnb-4bit") model = PeftModel.from_pretrained(base_model, "hiteshluke/arbiter-4b") - Notebooks
- Google Colab
- Kaggle
Arbiter v3.3 · 4B
A production System One decision model with a trained 24-slot pointer head
on top of a LoRA-adapted unsloth/gemma-3-4b-it. One forward pass per
decision, deterministic typed output across all three decision primitives:
noul— Boolean / True-Falsechoice— multiple-choice up to 16 options (A–P)score— ordinal 0–5 rating
Part of the Zyot Lab open decision-model lineup by Codekins Pvt Ltd.
Benchmarks
| Benchmark | Primitive | Accuracy |
|---|---|---|
| BoolQ (validation, n = 1,000) | noul (T / F) | 0.849 |
| ARC-Challenge (test, n = 500) | 4-choice | 0.738 |
| CommonsenseQA (validation, n = 500) | 5-choice | 0.706 |
Independent runs on an NVIDIA T4 (4-bit quantized inference via bitsandbytes).
Architecture
A single trained pointer head of 24 output slots sits on the last hidden state of the base model:
| Slots | Primitive |
|---|---|
| 0 – 1 | T / F (noul) |
| 2 – 17 | A – P (choice, up to 16 options) |
| 18 – 23 | 0 – 5 (score) |
The head is initialized from the base LM-head verbalizer rows (T, F,
A–P, 0–5) so step 0 reproduces a classical verbalizer readout. The
LoRA adapter + head are then trained jointly with per-slot masking, so each
decision type routes gradient only through the slots that are valid for
that prompt.
Frozen T / F slots. During training, gradients on slots 0 and 1 are clamped to zero via a backward hook — the model therefore cannot drift away from the base model's native boolean behavior. This preserves BoolQ baseline accuracy by construction.
Usage
from transformers import AutoTokenizer, AutoModelForCausalLM
from peft import PeftModel
from huggingface_hub import hf_hub_download
import torch, torch.nn as nn, json
BASE = "unsloth/gemma-3-4b-it"
REPO = "hiteshluke/arbiter-4b"
tok = AutoTokenizer.from_pretrained(BASE)
model = AutoModelForCausalLM.from_pretrained(BASE, dtype=torch.bfloat16, device_map="auto")
model = PeftModel.from_pretrained(model, REPO).eval()
# Load the trained pointer head
head_path = hf_hub_download(REPO, "head.pt")
head_meta = json.loads(open(hf_hub_download(REPO, "head_meta.json")).read())
hidden = model.config.text_config.hidden_size
head = nn.Linear(hidden, head_meta["num_slots"], bias=False).to(
device=model.device, dtype=torch.bfloat16
)
head.load_state_dict({"weight": torch.load(head_path, map_location=model.device)["proj.weight"]})
head.eval()
@torch.no_grad()
def decide(prompt: str, valid_slots: list[int]) -> int:
ids = tok(prompt, return_tensors="pt", truncation=True, max_length=1024).input_ids.to(model.device)
out = model(input_ids=ids, output_hidden_states=True, use_cache=False)
pooled = out.hidden_states[-1][0, -1]
logits = head(pooled.to(head.weight.dtype)).float().cpu().numpy()
return valid_slots[int(max(range(len(valid_slots)), key=lambda i: logits[valid_slots[i]]))]
# Example: BoolQ-style question
prompt = (
"State: The Eiffel Tower is in Paris, France.\n\n"
"Question: Is the Eiffel Tower in France?\n\n"
"Options:\nT. Yes / True\nF. No / False\n\nAnswer:"
)
slot = decide(prompt, valid_slots=[0, 1]) # 0 = T, 1 = F
print("Answer:", "T" if slot == 0 else "F") # -> T
Files
adapter_config.json+adapter_model.safetensors— LoRA adapter (r = 16, α = 32) on Gemma 3 4Bhead.pt— trained 24-slot pointer head (nn.Linear(hidden_size, 24), bf16)head_meta.json— slot layout + verbalizer token ids
Training
- Base:
unsloth/gemma-3-4b-it - LoRA: r = 16, α = 32, dropout = 0.05, targets q/k/v/o + gate/up/down
- Head:
nn.Linear(hidden_size, 24)in bf16, initialized from LM-head verbalizer rows - Data: curated subset of
SargeDev/jev-distill-corpus-v3at teacher-confidence ≥ 0.75 (noul + choice + score) - Objective: focal cross-entropy on the pointer head's logits over valid slots per row, sample-weighted by teacher confidence
- Frozen slots: 0 (
T) and 1 (F) — gradient zeroed via backward hook - Optimizer: AdamW, LR 5 × 10⁻⁵ cosine, warmup 200
- Batch: effective 16 (batch 2 × grad accum 8), max seq 768, bf16/fp16 on Kaggle T4
- Steps: 3,000 with EMA-averaged head over the last 5 eval checkpoints
Positioning
- System One pointer head — real trained classifier, not verbalizer readout
- All three Jev decision primitives in a single head (noul + choice + score)
- One forward pass per decision — no autoregressive generation, no CoT
- Deploys via standard
transformers+peft+ one smallhead.ptload
License
Apache-2.0 for the LoRA adapter, pointer head, and code in this repository.
Base unsloth/gemma-3-4b-it and google/gemma-3-4b-it are governed by the
Gemma license.
Citation
@misc{arbiter-v3-3-4b-2026,
title = {Arbiter v3.3 · 4B — a System One decision model on Gemma 3 4B},
author = {Codekins Pvt Ltd · Zyot Lab},
year = {2026},
url = {https://huggingface.co/hiteshluke/arbiter-4b}
}
- Downloads last month
- 34
Model tree for hiteshluke/arbiter-4b
Dataset used to train hiteshluke/arbiter-4b
Evaluation results
- accuracy on BoolQ (validation, n=1000)self-reported0.849
- accuracy on ARC-Challenge (test, n=500)self-reported0.738
- accuracy on CommonsenseQA (validation, n=500)self-reported0.706