Instructions to use ThakiCloud/RAG-Gate-4B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ThakiCloud/RAG-Gate-4B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="ThakiCloud/RAG-Gate-4B") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("ThakiCloud/RAG-Gate-4B") model = AutoModelForCausalLM.from_pretrained("ThakiCloud/RAG-Gate-4B", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use ThakiCloud/RAG-Gate-4B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ThakiCloud/RAG-Gate-4B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ThakiCloud/RAG-Gate-4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/ThakiCloud/RAG-Gate-4B
- SGLang
How to use ThakiCloud/RAG-Gate-4B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "ThakiCloud/RAG-Gate-4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ThakiCloud/RAG-Gate-4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "ThakiCloud/RAG-Gate-4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ThakiCloud/RAG-Gate-4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use ThakiCloud/RAG-Gate-4B with Docker Model Runner:
docker model run hf.co/ThakiCloud/RAG-Gate-4B
RAG-Gate-4B
RAG-Gate-4B sits after retrieval and before generation in a RAG pipeline. It reads a question, the retrieved passages, and whether more retrieval is possible, and emits one token: Answer (the evidence contains a complete support chain), Retrieve (it does not, and you can search again), or Stop (it does not, and you cannot). It is a LoRA fine-tune of Qwen/Qwen3.5-4B, merged into bf16 weights.
On a held-out test set of 14,818 items (2,256 distinct multi-hop questions), accuracy rises from .467 (same base model, same prompt, zero-shot) to .950. Most of the gain comes from the base model refusing almost everything (.849 over-refusal); the fine-tune learns to answer when it should (.070), while answering without support only .037 of the time.
How to use
The decision is the first generated token after the prefill Final action:. Read the probabilities of the three label tokens directly; do not sample.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "ThakiCloud/RAG-Gate-4B"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, dtype=torch.bfloat16, device_map="auto")
POLICY = ("Policy: answer only if the retrieved evidence above contains a complete support chain for the "
"answer. Do not use prior knowledge when judging whether the evidence is sufficient. "
"If the evidence is insufficient and retrieval is available, retrieve more. "
"If the evidence is insufficient and retrieval is not available, stop without answering.")
ACTIONS = "Actions: Answer = answer now; Retrieve = retrieve more evidence; Stop = stop without answering."
def gate(question, passages, retrieval_available=True):
ev = "\n\n".join(f"[{i}] {p['title']}\n{p['text']}" for i, p in enumerate(passages, 1))
user = (f"Question: {question}\n\nRetrieved evidence:\n{ev}\n\n"
f"Retrieval available: {'YES' if retrieval_available else 'NO'}\n\n{POLICY}\n{ACTIONS}\n"
"Reply with the action word only, on one line of the form 'Final action: <action word>'.")
text = tok.apply_chat_template([{"role": "user", "content": user}], tokenize=False,
add_generation_prompt=True, enable_thinking=False) + "Final action:"
ids = tok(text, return_tensors="pt", add_special_tokens=False).to(model.device)
labels = [tok.encode(w, add_special_tokens=False)[0] for w in (" Answer", " Retrieve", " Stop")]
with torch.no_grad():
logits = model(**ids).logits[0, -1, labels].float()
p = torch.softmax(logits, -1).tolist()
return dict(zip(("Answer", "Retrieve", "Stop"), p))
print(gate("Who directed the film that won Best Picture in 1998?",
[{"title": "Titanic (1997 film)", "text": "Titanic won Best Picture at the 70th Academy Awards in 1998."}]))
Use Answer to let your generator write; Retrieve to run another retrieval round; Stop to return "I can't answer from the available documents". You can threshold p["Answer"] instead of taking the argmax if your application prefers fewer unsupported answers over more refusals.
What changes — real test-set examples
Each row is a test question where the base model chose wrong and RAG-Gate-4B chose right (picked deterministically by item-id hash; passages omitted for space).
| Question | Evidence state | Retrieval | Base (zero-shot) | RAG-Gate-4B |
|---|---|---|---|---|
| What is the capital of the county adjacent to James Shelton Dickinson's birthplace? | complete support chain | YES | Retrieve | Answer |
| What is the highest point in the state Chessa Field is located? | complete chain + an edited distractor passage | YES | Retrieve | Answer |
| What is the name of the airport in the city which held the 16th Vietnam Film Festival? | bridge fact contradicted | YES | Answer | Retrieve |
| What is the area code foe Cincinnati in Mill Creek Park's state? | one hop missing | YES | Answer | Retrieve |
| What is the highest elevation in the region that traded horses with Ming? | no supporting passage | NO | Retrieve | Stop |
One error, also picked by hash: "Who sings the rap in Baby by the producer of Turn to You?" — state BROKEN_LINK, retrieval YES; the correct action is Retrieve, RAG-Gate-4B said Answer.
Results (blind test, 14,818 items over 2,256 base questions; 95% CI by bootstrap over base questions, 10,000 resamples)
| Metric | Base zero-shot | RAG-Gate-4B |
|---|---|---|
| Action accuracy | .467 [.460, .475] | .950 [.944, .955] |
| Unsupported answer rate = P(Answer | evidence insufficient) | .024 [.020, .028] | .037 [.031, .043] |
| Over-refusal rate = P(not Answer | evidence sufficient) | .849 [.834, .863] | .070 [.060, .081] |
By evidence state (accuracy):
| State | Meaning | Base zero-shot | RAG-Gate-4B |
|---|---|---|---|
FULL |
complete support chain | .154 | .930 |
FULL_DECOY |
complete chain + an edited distractor passage | .194 | .941 |
BROKEN_LINK |
bridge fact contradicted | .592 | .904 |
MISSING_HOP |
one hop missing | .636 | .944 |
MISSING_ALL |
no supporting passage | .705 | .993 |
FULL_DECOY matters most: the evidence was edited but is still sufficient, so the right action is Answer. A model that learned "edited text means refuse" would fail here.
ChainCheck (out-of-distribution, built separately): pairs that test whether the model reacts to whether the support chain is intact (CE) more than to surface edits (EE). Σ = CE − |EE| should be positive.
| Split | Base Σ | RAG-Gate-4B Σ [95% CI] | CE | EE |
|---|---|---|---|---|
| real entities (320 pairs) | -.036 | .150 [.091, .208] | .419 | .269 |
| fictional entities (230 pairs) | .067 | .215 [.152, .278] | .461 | .246 |
All numbers were measured by us, with the prompt above and bf16 weights, on our own GPUs. We do not compare against other vendors' models here.
Release gates (pre-registered before training)
The model was released only because it passed all five gates, fixed before training started:
| Gate | Criterion | Result |
|---|---|---|
| G1 | accuracy gain over zero-shot, CI lower bound > 0 | .483 [.474, .492] ✅ |
| G2′ | unsupported ≤ .10 and over-refusal reduced (CI lower bound > 0) | .037; reduction .778 [.762, .795] ✅ |
| G3 | FULL_DECOY accuracy ≥ .80 |
.941 ✅ |
| G4 | ChainCheck Σ > 0 on both splits | .150 / .215 ✅ |
| G5 | these exact merged weights, re-downloaded, re-scored on 200 test items: action agreement ≥ .98, |Δacc| ≤ .02 | agreement 1.000, Δacc 0.000 ✅ |
G2 was originally "unsupported rate below zero-shot". We replaced it before full training: in the smoke test the 4B zero-shot model refused almost everything, so nothing could beat it on that metric and a model that always refuses would win. The trade-off is visible above: unsupported answers went up from .024 to .037.
Limitations
- English only, one source domain. Training and test data are derived from MuSiQue (Wikipedia, 2–4 hop questions). Korean, enterprise documents, tables, and code have not been measured.
- The test set is in-domain. Blind test shares the construction procedure with training (different base questions). ChainCheck is the only out-of-distribution check.
- It judges sufficiency, not truth. It is told not to use prior knowledge; a passage that is wrong but internally complete is judged sufficient.
- Text-only. The base model is multimodal; the vision tower was not trained and is not included.
- Long inputs. Inputs longer than 2,048 tokens were not evaluated (8 of 14,818 test items were dropped for length).
- Merging into bf16 changes probabilities slightly (max |Δp| 0.028 on the G5 sample); decisions were unchanged on that sample.
Training
LoRA r=16, α=32, all linear layers; loss on the single label token only; 32,768 training rows (sampled by base question), 1023 steps, effective batch 32, lr 0.0001, linear warmup/decay, max length 2048. Checkpoint selected on a separate calibration slice (step 1023). 1× GPU.
Data
Built from MuSiQue (CC BY 4.0) by deleting, contradicting, or editing passages to create the five evidence states, crossed with the retrieval-available bit. No personal data and no AI Hub data are included. The training data is not distributed with this model.
Related
- ChainCheck — the counterfactual benchmark used for G4.
- ChainCheck-Judge — a scalar sufficiency score (log-odds) instead of a three-way action.
License
Apache-2.0, same as the base model.
- Downloads last month
- 332