LLaMA-2 7B Self-Aligned Final Model

Instruction-following model finetuned from meta-llama/Llama-2-7b-hf using QLoRA, implementing the self-alignment pipeline from ["Self-Alignment with Instruction Backtranslation"].

Description

This model is the final output of a 4-step self-alignment pipeline:

  1. Finetune a backward model to generate instructions from responses
  2. Use the backward model to generate instructions for LIMA responses (self-augmentation)
  3. Score and filter high-quality (instruction, response) pairs (self-curation)
  4. Finetune this model on the curated dataset

Training Data

  • Dataset: RubyXZZZ/lima-curated-backtranslation (self-curated high quality pairs, score >= 4)
  • Source: LIMA dataset responses with backward-model-generated instructions

Training Details

  • Base model: meta-llama/Llama-2-7b-hf
  • Method: QLoRA (4-bit quantization + LoRA)
  • LoRA: r=64, alpha=128, dropout=0.1
  • Learning rate: 1e-5 (linear decay)
  • Weight decay: 0.1

Citation

@article{li2023self,
  title={Self-alignment with instruction backtranslation},
  journal={arXiv preprint arXiv:2308.06259},
  year={2023}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for RubyXZZZ/llama2-7b-self-aligned-final