gemma-2-2b-qa-DPO

DPO applied to thesreedath/gemma-2-2b-qa-sft, a QA-SFT model, to sharpen answer quality on closed-book (question only) prompts.

Results

metric before after
preference accuracy (held-out) 0.0 0.9
chosen-vs-rejected logprob margin 21.2157 25.5598

Trained 2 epochs, beta 0.1, 64 steps on 260 pairs (40 held out), in 125s.

How it was trained

DPO (Direct Preference Optimization) optimises the preference objective in closed form against a frozen copy of the starting model. No reward model and no sampling are involved at training time, which is why it is fast and stable.

Preference data

600 triplets (prompt / chosen / rejected), 300 grounded and 300 closed-book. This model trained on the closed_book half, because that is the distribution it was supervised-fine-tuned on.

The rejected answer in each pair carries exactly one deliberately induced flaw, drawn evenly from six modes that these models actually exhibit: a wrong figure, an unsupported claim, vagueness, rambling repetition, a partial answer, and a false refusal. An LLM judge then confirmed the ordering with the two answers shown in randomised positions, so it could not score well by always choosing the first. Pairs the judge disagreed with, or was unconfident about, were dropped.

Prompt format

This model expects gemma_chat:

<bos><start_of_turn>user\n{question}<end_of_turn>\n<start_of_turn>model\n

The format was established empirically, by scoring known-good answers under competing templates and comparing generation behaviour -- not assumed from the tokenizer.

Limits

600 preference pairs is a small budget for preference optimization. Expect sharper formatting, less padding and more consistent refusals -- not new capability. Preference labels come from an LLM judge, so the model inherits that judge's blind spots. Not legal or financial advice.

Downloads last month
125
Safetensors
Model size
3B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for nagbhaskar55/gemma-2-2b-qa-DPO

Finetuned
(2)
this model