license: apache-2.0

Dostoevsky GPT

A small GPT-style language model trained from scratch and fine-tuned on a corpus of works by Fyodor Dostoevsky.

Model Details

  • Architecture: GPT-style Transformer
  • Tokenizer: GPT-2 tokenizer
  • Task: Next-token prediction / text generation
  • Training: Pretrained on TinyStories, then fine-tuned on Dostoevsky
  • Checkpoint: dostoevsky_gpt.pt

Evaluation

Metric Result
Validation Loss 4.109
Validation Top-1 Next-Token Accuracy 28.72%
Validation Perplexity ~60.9

Usage

Install dependencies

pip install torch tiktoken
import torch
import tiktoken

# Load tokenizer
tokenizer = tiktoken.get_encoding("gpt2")

# Load checkpoint
checkpoint = torch.load(
    "dostoevsky_gpt.pt",
    map_location="cpu"
)

# Recreate the model architecture here
# using the configuration from the training code.

model = GPT(...)
model.load_state_dict(checkpoint)
model.eval()

# Generate text
prompt = "The sky was"

tokens = tokenizer.encode(prompt)
idx = torch.tensor([tokens])

with torch.no_grad():
    output = model.generate(
        idx,
        max_new_tokens=100,
        temperature=0.8,
        top_k=40
    )

print(tokenizer.decode(output[0].tolist()))
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support