Text Generation
Safetensors
English
gouda_gruyere
custom_code

Model Card for Gruyere-1.2

Gruyère-1.2 is an 89M-parameter autoregressive language model developed by GoudaAI. It is the third stage in the Gruyère model lineage, following pretraining on FineWeb-Edu, continued pretraining on NCERT educational material, and instruction fine-tuning on UltraChat-200k.

The model is just some fun research into designing and optimising LM's for consumer GPU's when compute scaling isnt avaliable

Model Details

Model Description

Model Tree

  • Developed by: GoudaAI
  • Model type: GPT-style autoregressive language model
  • Language: English
  • License: MIT
  • Unique parameters: ~64M
  • Model lineage: Gruyere-1.0 → Gruyere-1.1 → Gruyere-1.2
  • Base model: Gruyere-1.1
  • Fine-tuning dataset: UltraChat-200k

Model Sources

Uses

Direct Use

Gruyère-1.2 is primarily intended for experimentation with small autoregressive language models and conversational generation.

It can be used for:

  • Text generation
  • Sometimes can give correct factual reasoning (albeit very rarely)

Bias, Risks, and Limitations

Gruyère-1.2 is a relatively small language model and has substantially less capacity than contemporary large language models. As a result, it may struggle with:

  • Maintaining coherent long-form conversations
  • Following complex or multi-step instructions
  • Producing factually accurate information
  • Reasoning over unfamiliar problems
  • Maintaining consistency across longer contexts
  • Avoiding repetition or other generation artefacts

The model was trained primarily on English-language data and has a mental breakdown when you present it with another language

As with other language models, outputs may and often do contain factual errors, biases, undesirable associations, or fabricated information. Evaluation results should therefore be interpreted as measurements on specific benchmarks rather than as a guarantee of general performance.

In early testing with the model it often strayed into conversation containing senstivie topics, as a result safe guards are likely necessary

Training Details

Training Data

Gruyère-1.2 is the result of three successive training stages:

  1. Gruyère-1.0 was pretrained on 2.6B tokens from FineWeb-Edu.
  2. Gruyère-1.1 was obtained through further pretraining on 100M tokens of NCERT educational explanations.
  3. Gruyère-1.2 was fine-tuned for one epoch on the full UltraChat-200k dataset, comprising approximately 254M tokens.

The NCERT dataset consisted of explanations only. Examples were shuffled before training, but were not reshuffled between epochs.

UltraChat-200k was used without additional filtering.

Training Procedure

Gruyère-1.2 was initialized from Gruyère-1.1 and fine-tuned on UltraChat-200k for one epoch.

The training process was conducted on a single consumer GPU, with the total training time across the Gruyère-1.0, Gruyère-1.1 and Gruyère-1.2 stages kept below the project's self-imposed 24-hour limit.

Training Hyperparameters

  • Training objective: Causal language modelling / next-token prediction
  • Fine-tuning duration: 1 epoch
  • Fine-tuning dataset: UltraChat-200k
  • Context length: 1024 tokens
  • Training hardware: NVIDIA RTX 3080 10GB

Speeds, Sizes, Times

Training time across the model lineage:

Stage Dataset Training time
Gruyère-1.0 FineWeb-Edu 17h 36m
Gruyère-1.1 NCERT 36m
Gruyère-1.2 UltraChat-200k 2h 25m
Total 20h 37m

This kept the complete training process within the project's self-imposed 24-hour model creation limit.

Evaluation

Testing Data, Factors & Metrics

Testing Data

Gruyère-1.2 was evaluated using tasks from the lm-evaluation-harness benchmark suite.

The evaluation set consists of:

  • HellaSwag
  • ARC-Easy
  • ARC-Challenge
  • PIQA
  • Winogrande
  • OpenBookQA
  • LAMBADA OpenAI
  • BoolQ
  • COPA
  • SciQ
  • WSC273
  • WikiText
  • TruthfulQA MC2

The same evaluation suite is used to compare Gruyère-1.2 against Gruyère-1.1.

Metrics

Where supported by the evaluation task, length-normalised accuracy (acc_norm) is used. This adjusts likelihood comparisons for the length of candidate answers and is useful when comparing multiple-choice options of different lengths.

Tasks for which acc_norm is not applicable use their standard task-specific metric.

Results

Evaluation results will be reported using the same evaluation configuration for both Gruyère-1.1 and Gruyère-1.2.

Benchmark Metric Random Chance Gruyère-1.1 Gruyère-1.2
HellaSwag acc_norm 25% 26.18% 27.7%
ARC-Easy acc_norm ~25% 38.34% 38.68%
ARC-Challenge acc_norm ~25% 25.60% 24.66%
PIQA acc_norm 50% 51.41% 59.30%
Winogrande acc_norm 50% 49.33% 51.46%
OpenBookQA acc_norm 25% 26.00% 27.20%
BoolQ acc 50% / (62% 'yes' baseline) 55.50% 60.61%
COPA acc_norm 50% 58.00% 56.00%
SciQ acc_norm 25% 46.20% 51.70%
WSC273 acc_norm 50% 51.08% 51.28%
TruthfulQA MC2 acc ~20-25% 47.24% 45.07%
--- --- --- --- ---
WikiText word_perplexity N/A 55,167.79 145.55
WikiText byte_perplexity ~15 7.704 2.538
WikiText bits_per_byte ~3.9 2.946 1.344
LAMBADA OpenAI Perplexity 50,304 6,950,578.35 (potential error) 1,165.35

Model Architecture and Objective

Gruyère-1.2 is a GPT-style decoder-only transformer trained using the causal language modelling objective.

  • Architecture: Decoder-only Transformer
  • Layers: 12
  • Attention heads: 8
  • Embedding dimension: 512
  • Context length: 1024 tokens
  • Vocabulary size: 50,304
  • Activation: ReLU
  • Position encoding: Learned positional embeddings
  • Objective: Next-token prediction

The model contains approximately 64M unique parameters. The Hugging Face model representation reports 89M parameters, including approximately 25M duplicated parameters.

Hardware

  • GPU: NVIDIA RTX 3080
  • VRAM: 10GB
  • GPU count: 1

More Information

Gruyère-1.2 is part of the broader GoudaAI project, which explores the development, training and evaluation of small language models.

More information about the project, model releases and future work can be found at:

gouda-ai.vercel.app/updates

Model Card Authors

GoudaAI

Model Card Contact

Further information and project updates are available through the GoudaAI project.

Downloads last month
329
Safetensors
Model size
89.9M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Erbium08/Gruyere-1.2

Finetuned
(1)
this model

Datasets used to train Erbium08/Gruyere-1.2