GPT-2 Small — multi-turn chat (toy)

Multi-turn chat fine-tune of submarat/gpt2-small-fineweb-edu-10b (124M GPT-2 reproduced from scratch on 10B FineWeb-Edu tokens). SFT on smol-smoltalk (conversational, built for small models) with a ChatML-style chat template and assistant_only_loss (loss on assistant turns across the whole conversation).

Unlike the single-turn SFT model (Alpaca), this one carries a chat_template and holds a running conversation.

It's a toy: 124M with a 1024-token context, so multi-turn coherence degrades quickly and it hallucinates. It holds the chat format and can reference recent turns — don't expect real conversational memory.

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

m = AutoModelForCausalLM.from_pretrained("submarat/gpt2-small-fineweb-edu-10b-chat")
tok = AutoTokenizer.from_pretrained("submarat/gpt2-small-fineweb-edu-10b-chat")

messages = [{"role": "user", "content": "Give me three tips for studying."}]
prompt = tok.apply_chat_template(messages, add_generation_prompt=True, tokenize=False)
ids = tok(prompt, return_tensors="pt").input_ids
out = m.generate(ids, max_new_tokens=100, do_sample=True, top_k=40, temperature=0.7,
                 repetition_penalty=1.3, pad_token_id=tok.eos_token_id)
print(tok.decode(out[0][ids.shape[1]:], skip_special_tokens=True))

Training

  • Base: submarat/gpt2-small-fineweb-edu-10b (124M)
  • Data: smol-smoltalk (40k conversations)
  • 2 epochs, batch 24, LR 2e-5 cosine, bf16, ctx 1024, assistant_only_loss
Downloads last month
51
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for submarat/gpt2-small-fineweb-edu-10b-chat

Quantized
(2)
this model

Dataset used to train submarat/gpt2-small-fineweb-edu-10b-chat

Space using submarat/gpt2-small-fineweb-edu-10b-chat 1