GPT-2 DPO LoRA

Fine-tuning model meta-llama/Llama-3.2-1B-Instruct menggunakan Direct Preference Optimization (DPO) dengan LoRA.

Base Model

meta-llama/Llama-3.2-1B-Instruct

Training

  • Method: DPO
  • LoRA Rank: 16
  • LoRA Alpha: 32
  • Epoch: 2
  • Learning Rate: 0.0001
  • Beta: 0.1
  • Max Sequence Length: 1024

Dataset

IndonesiaAI/dpo-dataset

sampling 20.000 data saja yang diambil,

Dataset digunakan dalam format:

  • Prompt
  • Chosen response
  • Rejected response

Training Objective

Model dioptimalkan agar memberikan preferensi lebih tinggi terhadap response chosen dibandingkan response rejected.

Downloads last month
217
Safetensors
Model size
1B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for RantiRepo/Llama3.2-1B-DPO-LoRA

Adapter
(656)
this model