cosmos

A small GPT-2 style language model trained from scratch on astronomy and cosmology, from public-domain books about the heavens.

Intended use

Research and education. It is a small model trained on a limited corpus, so expect repetitive or factually wrong output. It is not suitable for factual question answering.

Training

  • Architecture: GPT-2 style decoder, 4 layers, 4 heads, 256 hidden size, context 256
  • Tokenizer: byte-level BPE, vocabulary 8000
  • Objective: next-token prediction
  • Final validation loss: 6.171
  • Configuration: train_config.json in this repository

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

tok = AutoTokenizer.from_pretrained("clivern/cosmos")
model = AutoModelForCausalLM.from_pretrained("clivern/cosmos")
inputs = tok("astronomy and cosmology", return_tensors="pt")
out = model.generate(**inputs, max_new_tokens=60, do_sample=True, top_p=0.95)
print(tok.decode(out[0], skip_special_tokens=True))
Downloads last month
27
Safetensors
Model size
5.27M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support