Instructions to use nagbhaskar55/gemma-2-2b-qa-DPO with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use nagbhaskar55/gemma-2-2b-qa-DPO with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="nagbhaskar55/gemma-2-2b-qa-DPO") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("nagbhaskar55/gemma-2-2b-qa-DPO") model = AutoModelForCausalLM.from_pretrained("nagbhaskar55/gemma-2-2b-qa-DPO", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use nagbhaskar55/gemma-2-2b-qa-DPO with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "nagbhaskar55/gemma-2-2b-qa-DPO" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "nagbhaskar55/gemma-2-2b-qa-DPO", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/nagbhaskar55/gemma-2-2b-qa-DPO
- SGLang
How to use nagbhaskar55/gemma-2-2b-qa-DPO with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "nagbhaskar55/gemma-2-2b-qa-DPO" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "nagbhaskar55/gemma-2-2b-qa-DPO", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "nagbhaskar55/gemma-2-2b-qa-DPO" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "nagbhaskar55/gemma-2-2b-qa-DPO", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use nagbhaskar55/gemma-2-2b-qa-DPO with Docker Model Runner:
docker model run hf.co/nagbhaskar55/gemma-2-2b-qa-DPO
gemma-2-2b-qa-DPO
DPO applied to
thesreedath/gemma-2-2b-qa-sft, a QA-SFT model, to sharpen answer
quality on closed-book (question only) prompts.
Results
| metric | before | after |
|---|---|---|
| preference accuracy (held-out) | 0.0 | 0.9 |
| chosen-vs-rejected logprob margin | 21.2157 | 25.5598 |
Trained 2 epochs, beta 0.1, 64 steps on 260 pairs (40 held out), in 125s.
How it was trained
DPO (Direct Preference Optimization) optimises the preference objective in closed form against a frozen copy of the starting model. No reward model and no sampling are involved at training time, which is why it is fast and stable.
Preference data
600 triplets (prompt / chosen / rejected), 300 grounded and 300 closed-book. This model trained on the closed_book half, because that is the distribution it was supervised-fine-tuned on.
The rejected answer in each pair carries exactly one deliberately induced flaw, drawn evenly from six modes that these models actually exhibit: a wrong figure, an unsupported claim, vagueness, rambling repetition, a partial answer, and a false refusal. An LLM judge then confirmed the ordering with the two answers shown in randomised positions, so it could not score well by always choosing the first. Pairs the judge disagreed with, or was unconfident about, were dropped.
Prompt format
This model expects gemma_chat:
<bos><start_of_turn>user\n{question}<end_of_turn>\n<start_of_turn>model\n
The format was established empirically, by scoring known-good answers under competing templates and comparing generation behaviour -- not assumed from the tokenizer.
Limits
600 preference pairs is a small budget for preference optimization. Expect sharper formatting, less padding and more consistent refusals -- not new capability. Preference labels come from an LLM judge, so the model inherits that judge's blind spots. Not legal or financial advice.
- Downloads last month
- 125
Model tree for nagbhaskar55/gemma-2-2b-qa-DPO
Base model
thesreedath/gemma-2-2b-qa-sft