RAG-Gate-4B

RAG-Gate-4B sits after retrieval and before generation in a RAG pipeline. It reads a question, the retrieved passages, and whether more retrieval is possible, and emits one token: Answer (the evidence contains a complete support chain), Retrieve (it does not, and you can search again), or Stop (it does not, and you cannot). It is a LoRA fine-tune of Qwen/Qwen3.5-4B, merged into bf16 weights.

On a held-out test set of 14,818 items (2,256 distinct multi-hop questions), accuracy rises from .467 (same base model, same prompt, zero-shot) to .950. Most of the gain comes from the base model refusing almost everything (.849 over-refusal); the fine-tune learns to answer when it should (.070), while answering without support only .037 of the time.

How to use

The decision is the first generated token after the prefill Final action:. Read the probabilities of the three label tokens directly; do not sample.

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

repo = "ThakiCloud/RAG-Gate-4B"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, dtype=torch.bfloat16, device_map="auto")

POLICY = ("Policy: answer only if the retrieved evidence above contains a complete support chain for the "
          "answer. Do not use prior knowledge when judging whether the evidence is sufficient. "
          "If the evidence is insufficient and retrieval is available, retrieve more. "
          "If the evidence is insufficient and retrieval is not available, stop without answering.")
ACTIONS = "Actions: Answer = answer now; Retrieve = retrieve more evidence; Stop = stop without answering."

def gate(question, passages, retrieval_available=True):
    ev = "\n\n".join(f"[{i}] {p['title']}\n{p['text']}" for i, p in enumerate(passages, 1))
    user = (f"Question: {question}\n\nRetrieved evidence:\n{ev}\n\n"
            f"Retrieval available: {'YES' if retrieval_available else 'NO'}\n\n{POLICY}\n{ACTIONS}\n"
            "Reply with the action word only, on one line of the form 'Final action: <action word>'.")
    text = tok.apply_chat_template([{"role": "user", "content": user}], tokenize=False,
                                   add_generation_prompt=True, enable_thinking=False) + "Final action:"
    ids = tok(text, return_tensors="pt", add_special_tokens=False).to(model.device)
    labels = [tok.encode(w, add_special_tokens=False)[0] for w in (" Answer", " Retrieve", " Stop")]
    with torch.no_grad():
        logits = model(**ids).logits[0, -1, labels].float()
    p = torch.softmax(logits, -1).tolist()
    return dict(zip(("Answer", "Retrieve", "Stop"), p))

print(gate("Who directed the film that won Best Picture in 1998?",
           [{"title": "Titanic (1997 film)", "text": "Titanic won Best Picture at the 70th Academy Awards in 1998."}]))

Use Answer to let your generator write; Retrieve to run another retrieval round; Stop to return "I can't answer from the available documents". You can threshold p["Answer"] instead of taking the argmax if your application prefers fewer unsupported answers over more refusals.

What changes — real test-set examples

Each row is a test question where the base model chose wrong and RAG-Gate-4B chose right (picked deterministically by item-id hash; passages omitted for space).

Question Evidence state Retrieval Base (zero-shot) RAG-Gate-4B
What is the capital of the county adjacent to James Shelton Dickinson's birthplace? complete support chain YES Retrieve Answer
What is the highest point in the state Chessa Field is located? complete chain + an edited distractor passage YES Retrieve Answer
What is the name of the airport in the city which held the 16th Vietnam Film Festival? bridge fact contradicted YES Answer Retrieve
What is the area code foe Cincinnati in Mill Creek Park's state? one hop missing YES Answer Retrieve
What is the highest elevation in the region that traded horses with Ming? no supporting passage NO Retrieve Stop

One error, also picked by hash: "Who sings the rap in Baby by the producer of Turn to You?" — state BROKEN_LINK, retrieval YES; the correct action is Retrieve, RAG-Gate-4B said Answer.

Results (blind test, 14,818 items over 2,256 base questions; 95% CI by bootstrap over base questions, 10,000 resamples)

Metric Base zero-shot RAG-Gate-4B
Action accuracy .467 [.460, .475] .950 [.944, .955]
Unsupported answer rate = P(Answer | evidence insufficient) .024 [.020, .028] .037 [.031, .043]
Over-refusal rate = P(not Answer | evidence sufficient) .849 [.834, .863] .070 [.060, .081]

By evidence state (accuracy):

State Meaning Base zero-shot RAG-Gate-4B
FULL complete support chain .154 .930
FULL_DECOY complete chain + an edited distractor passage .194 .941
BROKEN_LINK bridge fact contradicted .592 .904
MISSING_HOP one hop missing .636 .944
MISSING_ALL no supporting passage .705 .993

FULL_DECOY matters most: the evidence was edited but is still sufficient, so the right action is Answer. A model that learned "edited text means refuse" would fail here.

ChainCheck (out-of-distribution, built separately): pairs that test whether the model reacts to whether the support chain is intact (CE) more than to surface edits (EE). Σ = CE − |EE| should be positive.

Split Base Σ RAG-Gate-4B Σ [95% CI] CE EE
real entities (320 pairs) -.036 .150 [.091, .208] .419 .269
fictional entities (230 pairs) .067 .215 [.152, .278] .461 .246

All numbers were measured by us, with the prompt above and bf16 weights, on our own GPUs. We do not compare against other vendors' models here.

Release gates (pre-registered before training)

The model was released only because it passed all five gates, fixed before training started:

Gate Criterion Result
G1 accuracy gain over zero-shot, CI lower bound > 0 .483 [.474, .492] ✅
G2′ unsupported ≤ .10 and over-refusal reduced (CI lower bound > 0) .037; reduction .778 [.762, .795] ✅
G3 FULL_DECOY accuracy ≥ .80 .941 ✅
G4 ChainCheck Σ > 0 on both splits .150 / .215 ✅
G5 these exact merged weights, re-downloaded, re-scored on 200 test items: action agreement ≥ .98, |Δacc| ≤ .02 agreement 1.000, Δacc 0.000 ✅

G2 was originally "unsupported rate below zero-shot". We replaced it before full training: in the smoke test the 4B zero-shot model refused almost everything, so nothing could beat it on that metric and a model that always refuses would win. The trade-off is visible above: unsupported answers went up from .024 to .037.

Limitations

  • English only, one source domain. Training and test data are derived from MuSiQue (Wikipedia, 2–4 hop questions). Korean, enterprise documents, tables, and code have not been measured.
  • The test set is in-domain. Blind test shares the construction procedure with training (different base questions). ChainCheck is the only out-of-distribution check.
  • It judges sufficiency, not truth. It is told not to use prior knowledge; a passage that is wrong but internally complete is judged sufficient.
  • Text-only. The base model is multimodal; the vision tower was not trained and is not included.
  • Long inputs. Inputs longer than 2,048 tokens were not evaluated (8 of 14,818 test items were dropped for length).
  • Merging into bf16 changes probabilities slightly (max |Δp| 0.028 on the G5 sample); decisions were unchanged on that sample.

Training

LoRA r=16, α=32, all linear layers; loss on the single label token only; 32,768 training rows (sampled by base question), 1023 steps, effective batch 32, lr 0.0001, linear warmup/decay, max length 2048. Checkpoint selected on a separate calibration slice (step 1023). 1× GPU.

Data

Built from MuSiQue (CC BY 4.0) by deleting, contradicting, or editing passages to create the five evidence states, crossed with the retrieval-available bit. No personal data and no AI Hub data are included. The training data is not distributed with this model.

Related

  • ChainCheck — the counterfactual benchmark used for G4.
  • ChainCheck-Judge — a scalar sufficiency score (log-odds) instead of a three-way action.

License

Apache-2.0, same as the base model.

Downloads last month
332
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ThakiCloud/RAG-Gate-4B

Finetuned
Qwen/Qwen3.5-4B
Finetuned
(883)
this model

Datasets used to train ThakiCloud/RAG-Gate-4B

Collection including ThakiCloud/RAG-Gate-4B