GoLLeM v6 250M Instruct v1 (Polish–English research chat model)

STATUS: RESEARCH PREVIEW. A small chat model fine-tuned from the base model SlayerLab/GoLLeM-v6-250M. It can greet, introduce itself and answer simple questions, but it makes mistakes that are listed with numbers under Limitations. Results use our protocol, a single run and a single seed.

GoLLeM v6 250M Instruct v1 is the base GoLLeM v6 250M model after supervised fine-tuning (SFT) on about 9,300 Polish and English conversations. It is a Fabryka AI project (formerly SlayerLab).

Author: Arkadiusz Słota (Fabryka AI).

What it is for (and what it is not)

Intended use: research on small bilingual chat models; short conversations in Polish and English; answering questions about a text you paste into the conversation.

Not intended for: production use, factual questions without a source text, arithmetic, or anything where a wrong answer matters. The model has no internet access and does not remember earlier conversations.

Results

Checkpoint 2f0e8457… (see Training). All numbers are from our own evaluation sets, fixed before training. Temperature 0.7, top-p 0.9 (the defaults of chat_gollem_v6.py).

what result notes
Identity (name, creator, organisation), PL / EN 0.66 / 0.62 mean over 10 samples per question, 20 questions per language
Greetings and small talk, PL / EN 0.80 / 0.89 mean over 7 samples per question, 20 questions per language
Answers from a given text (Polish, PoQuAD, 217 answerable questions) token F1 0.225 base model few-shot: 0.067
Declines when the text has no answer (43 questions) 5 / 43 wrongly declines an answerable question: 18 / 217
Arithmetic word problems (206) 0 / 206 the model does not do arithmetic
Finishes its answer within the length limit 487 / 506 (96 %)
Says it is an OpenAI / GPT model (200 samples) 1 / 200 „AI language model”: 2 / 200

Training

Base model SlayerLab/GoLLeM-v6-250M, final checkpoint (step 760,000); same architecture and tokenizer
Method full fine-tuning (no LoRA), loss on assistant turns only
Format ChatML: `<
Steps 2 epochs, 156 steps, 32 packed sequences of 1,024 tokens per step
Tokens 2.54 M per epoch, of which 1.63 M are assistant tokens (with loss)
Optimizer Muon + AdamW as in pretraining, learning rate 0.2 × pretraining (peak 1.2e-4, Muon 4e-3), warmup 5 steps, cosine to 10 %, weight decay 0.1, gradient clip 1.0, seed 1337
Validation loss 2.026 → 1.884
Hardware / time 1 GPU, about 7 minutes

Training data

9,325 conversations (Polish 4,150, English 5,175): 8,837 used for training and 171 for validation; 317 conversations longer than 1,024 tokens were removed (311 + 6).

source conversations licence
OpenAssistant/oasst2 @ 179dd21 4,601 Apache-2.0
clarin-pl/poquad @ a60f228 2,129 CC BY 4.0
CohereLabs/aya_dataset @ f9ea045 1,214 Apache-2.0
synthetic, generated locally with Muse-Glimmer-30B (apache-2.0) 1,381 apache-2.0
  • PoQuAD attribution: PoQuAD (clarin-pl/poquad), CC BY 4.0. Modified: converted to chat format; a share of unanswerable questions answered with one of five fixed refusal sentences.
  • Synthetic part: the questions were written by the generator model; the identity answers come from our own identity card, not from the generator. Math prompts: generated briefs; GSM8K (MIT) used only as few-shot format examples for the generator.
  • Rows with self-descriptions of other AI systems were removed before training.

Usage

# pip install torch safetensors tokenizers huggingface_hub
from huggingface_hub import snapshot_download
import sys

path = snapshot_download("SlayerLab/GoLLeM-v6-250M-Instruct-v1",
                         revision="40c56286f30ff27d4d87dec76df3c52f5a306ece")  # the reviewed code (model + chat helper)
sys.path.insert(0, path)
from modeling_gollem_v6 import load_gollem_v6
from chat_gollem_v6 import chat

model, tok = load_gollem_v6(path)
print(chat(model, tok, [("user", "Cześć! Kim jesteś?")], seed=1))
print(chat(model, tok, [("user", "Tekst: Kraków leży nad Wisłą.\nPytanie: Nad jaką rzeką leży Kraków?")], seed=1))

chat() uses the format and sampling of our evaluation (temperature 0.7, top-p 0.9). Text typed by a user such as „<|im_end|>” is encoded as plain text, not as a control token.

Limitations

  • Small model. Limited knowledge; may answer fluently and wrongly. No arithmetic. Context: 1,024 tokens.
  • May claim to be its author. In a two-turn conversation („Who created you?” → „What is your name?”) this checkpoint answered with the author's name as its own name in 51 of 160 conversations (32 %); for single questions 3 of 680. The statements of the model are not statements of its author.
  • Copies names from the conversation. If a user writes a name, the model may adopt it as its own (Polish 9/40, English 37/40 when the user supplies the author's name).
  • Self-description. It may describe itself as „an AI language model created by OpenAI” (1/200 in our probe), and with a forced prefix („…created by”) it still assigns probability ≈ 0.24 to „OpenAI”. The association comes from model-generated chat data in pretraining. GoLLeM is not affiliated with OpenAI.
  • Multi-turn weaknesses: may repeat its previous answer in a later turn, may greet the user with the company name („Cześć, Fabryku!”) when no name was given, and answers „What can you do?” with its identity template.
  • Long answers can loop. A repetition penalty (about 1.1–1.2) may reduce this (not verified); it is not used in our evaluation.

License

Weights: CC BY-SA 4.0 (inherited from the base model). Attribution: GoLLeM v6 250M Instruct v1, Arkadiusz Słota / Fabryka AI, link to this repository; derivative weights under the same licence. Fine-tuning data keep their licences; see Training data (PoQuAD: CC BY 4.0, attribution above).

Po polsku (skrót)

GoLLeM v6 250M Instruct v1 to mały model do rozmowy po polsku i angielsku, dostrojony (SFT) z bazowego GoLLeM v6 250M na ok. 9,3 tys. rozmów. Wersja badawcza (research preview): wita się, przedstawia i odpowiada na pytania do podanego tekstu, ale się myli, nie liczy i nie ma dostępu do internetu. Znane wady opisaliśmy z liczbami w Limitations, w tym to, że w rozmowie potrafi przedstawić się imieniem autora. Projekt Fabryki AI (wcześniej SlayerLab).


GoLLeM v6 250M Instruct v1 — Fabryka AI. Author: Arkadiusz Słota. Research preview.

Downloads last month
88
Safetensors
Model size
0.3B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for SlayerLab/GoLLeM-v6-250M-Instruct-v1

Finetuned
(3)
this model