GPT-2 Fine-Tuned on OpenAssistant (OASST1 - 50% Subset)

This repository contains a GPT-2 model fine-tuned on a 50% random sample of the OpenAssistant Conversations Dataset (OASST1). The model was trained to generate human-like conversational responses based on instruction-response pairs extracted from the dataset's conversation trees.

📊 Model Details

  • Base Model: openai-community/gpt2 (124M parameters)
  • Architecture: GPT2LMHeadModel (Causal Language Modeling)
  • Language: English
  • License: Apache 2.0

🛠️ Training Details

Dataset Preparation

The model was trained on a reconstructed set of linear conversation threads from the OASST1 dataset. Deleted messages were filtered out to ensure higher quality interactions. To optimize training time and resource usage, exactly 50% of the total conversation threads were randomly sampled (seed 42).

Prompt Format: The model expects standard instruction tuning formatting:

### Human:
{user_message}

### Assistant:
{assistant_response}

Hyperparameters

Parameter Value
Learning Rate 5e-5
Batch Size 4 (with Gradient Accumulation: 4 -> Effective Batch Size: 16)
Epochs 3
Optimizer AdamW
Weight Decay 0.01
Warmup Steps 100
Precision fp16 (Mixed Precision)
Max Length 1024 tokens

Hardware

  • Trained using Google Colab (GPU)
  • Framework: PyTorch, Hugging Face transformers & datasets

🚀 Usage

You can use this model for text generation using the transformers library.

Python Code

from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model_name = "iko-01/gpt2_oasst1_50pct"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(model_name)

# Note: The model was trained with the ### Human: / ### Assistant: format
prompt = "### Human:\nWhat is the capital of France?\n\n### Assistant:\n"
inputs = tokenizer(prompt, return_tensors="pt")

outputs = model.generate(
    inputs.input_ids, 
    max_length=150, 
    do_sample=True, 
    top_p=0.95, 
    temperature=0.8,
    pad_token_id=tokenizer.eos_token_id
)

print(tokenizer.decode(outputs[0], skip_special_tokens=True))

⚠️ Limitations and Bias

  • Base Model Constraints: As it is based on GPT-2 (124M), the model may struggle with complex reasoning, factual accuracy, and long-context coherence compared to larger modern LLMs (like Llama 3 or Qwen).
  • Dataset Bias: The OASST1 dataset contains crowdsourced conversations, which may include hallucinations, biases, or toxic language present in the original data.
  • Prompt Format: The model strictly expects the ### Human: and ### Assistant: formatting to perform optimally.

📚 Citation

If you use this model or the underlying dataset, please cite the original OpenAssistant paper:

@article{kopf2023openassistant,
  title={OpenAssistant Conversations--Democratizing Large Language Model Alignment},
  author={Köpf, Andreas and Kilcher, Yann and von Rütte, Dimitri and Anagnostidis, Sotiris and Tam, Zhi Rui and Stevens, Kaleb and Abdullah, Abdullah and Berre, Armin and Bui, Duc Minh and Hofmann, Fabian and others},
  journal={arXiv preprint arXiv:2304.07327},
  year={2023}
}
Downloads last month
373
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for iko-01/gpt2_oasst1_50pct

Finetuned
(2290)
this model

Dataset used to train iko-01/gpt2_oasst1_50pct

Paper for iko-01/gpt2_oasst1_50pct