HuggingFaceH4/ultrachat_200k
Viewer • Updated • 515k • 108k • 893
Fine-tuned LLaMA-3.1-8B using SFT instruction tuning with prompt masking (loss computed only on response tokens).
| Benchmark | Baseline | This Model |
|---|---|---|
| GSM8K | 16.4% | 32.7% |
| MMLU | 58.1% | 58.2% |
| SST Safety | 62.0% | 77.0% |
| AlpacaEval | 1.57% | 4.5% |
eval_baseline/: Baseline evaluation results (pre-finetuning Llama-3.1-8B)Part of CS336 Assignment 5 (SFT Instruction Tuning). See building-from-scratch/sft for details.
Base model
meta-llama/Llama-3.1-8B