Text Generation
Transformers
Safetensors
PyTorch
English
fineweb_decoder
causal-lm
custom-code
educational
custom_code

FineWeb 12M SFT

The 12.2M base model supervised fine-tuned for three epochs on 10,789 instruction and question-answer examples.

This is a small educational and experimental model. It is not a production assistant and should not be relied on for factual, medical, legal, financial, safety-critical, or other consequential advice.

Model details

Property Value
Parameters 12,194,048
Layers 8
Hidden width 256
Attention heads 4
SwiGLU hidden width 960
Vocabulary 16,384 BPE tokens
Context window 1,024 tokens
Architecture Decoder-only Transformer
Position encoding RoPE
Normalization RMSNorm
Embeddings Tied input/output embeddings

The architecture was implemented from scratch in PyTorch. This repository includes custom Transformers-compatible configuration and modeling files.

Training

The base model was trained from scratch on 2,883,059,712 packed tokens selected from the FineWeb-Edu sample-10BT corpus. Documents were selected deterministically by hashing document IDs rather than taking the first source shards. Text was tokenized with a locally trained byte-level BPE tokenizer, an <eos> token was appended after every document, and tokens were packed into fixed 1,024-token blocks.

Base training used AdamW, cosine learning-rate decay, gradient clipping, bfloat16 autocast, and a global batch size of 65,536 tokens on an NVIDIA RTX 3090.

Supervised fine-tuning

This model was initialized from fineweb-12m-base and trained for three epochs with AdamW at a learning rate of 2e-5. The SFT set contained 10,789 training examples and 800 validation examples. It combined short Dolly 15K instruction examples with extractive question answering examples from SQuAD. Loss was applied only to assistant response tokens.

Evaluation

Best SFT validation loss: 2.7581, improved from 3.2562 before SFT.

These losses are internal held-out validation measurements. Base and SFT losses use different objectives, and losses from the 8K and 16K tokenizers are not directly comparable. No broad academic benchmark suite or human preference evaluation was run.

Usage

Install the dependencies:

pip install -r requirements.txt

Run the included example after cloning this repository:

python generate.py --model . --prompt "User: Explain why the sky appears blue.

Assistant:"

Or load it through Transformers:

from transformers import AutoModelForCausalLM

model = AutoModelForCausalLM.from_pretrained(
    ".",
    trust_remote_code=True,
)

trust_remote_code=True is required because this is a custom architecture. Review configuration_fineweb.py and modeling_fineweb.py before loading code from the Hub.

Limitations

  • The model is very small by modern language-model standards and often produces incorrect, repetitive, incoherent, or fabricated text.
  • Its knowledge is limited to patterns learned from the training data; it has no live information retrieval, tools, memory, or reliable sense of current time.
  • The 1,024-token context window includes both the prompt and generated output.
  • FineWeb-Edu is English-focused web data and can contain errors, biases, sensitive material, and uneven topic coverage.
  • The SFT variants received a small amount of instruction tuning and are not robust chat assistants.
  • The model has not been evaluated for safety, fairness, memorization, or production deployment.

Training data and attribution

The source datasets are not redistributed in this model repository and remain subject to their respective licences and attribution requirements.

Licence

The model weights and original code in this repository are released under the Apache License 2.0. This does not replace or override the licences of the training datasets.

Downloads last month
17
Safetensors
Model size
12.2M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train PeterRabbit/fineweb-12m-sft

Collection including PeterRabbit/fineweb-12m-sft