Model Card for reward

This model is a fine-tuned version of HuggingFaceTB/SmolLM-135M-Instruct on the HumanLLMs/Human-Like-DPO-Dataset dataset. It has been trained using TRL.

Quick start

from transformers import pipeline

sample = [
  {'content': 'Do you have a favorite type of music or artist?',
     'role': 'user'},
  {'content': "You know, I'm a big fan of indie-rock music. There's something about the raw, emotional vibe that really speaks to me. Arctic Monkeys are one of my all-time favorite bands - their lyrics are so clever and witty! But I'm also really into Tame Impala's psychedelic sound, it's like a trippy dream come true 😊. How about you, do you have a go-to genre or artist that gets you pumped up or relaxed? 🎵",
     'role': 'assistant'}
]
reward_model = AutoModelForSequenceClassification.from_pretrained('dzhuj/reward')
tokenizer = AutoTokenizer.from_pretrained('dzhuj/reward')
inputs_sample = tokenizer.apply_chat_template(sample, tokenize=False)
inputs_sample = tokenizer(inputs_sample, return_tensors="pt")
score = reward_model(**inputs_sample).logits[0].cpu().detach().item()

Training procedure

Visualize in Weights & Biases

This model was trained with Reward.

Framework versions

  • TRL: 0.15.2
  • Transformers: 4.49.0
  • Pytorch: 2.6.0
  • Datasets: 3.3.2
  • Tokenizers: 0.21.0
Downloads last month
11
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for dzhuj/reward

Finetuned
(193)
this model

Dataset used to train dzhuj/reward

Collection including dzhuj/reward