license: apache-2.0
Dostoevsky GPT
A small GPT-style language model trained from scratch and fine-tuned on a corpus of works by Fyodor Dostoevsky.
Model Details
- Architecture: GPT-style Transformer
- Tokenizer: GPT-2 tokenizer
- Task: Next-token prediction / text generation
- Training: Pretrained on TinyStories, then fine-tuned on Dostoevsky
- Checkpoint:
dostoevsky_gpt.pt
Evaluation
| Metric | Result |
|---|---|
| Validation Loss | 4.109 |
| Validation Top-1 Next-Token Accuracy | 28.72% |
| Validation Perplexity | ~60.9 |
Usage
Install dependencies
pip install torch tiktoken
import torch
import tiktoken
# Load tokenizer
tokenizer = tiktoken.get_encoding("gpt2")
# Load checkpoint
checkpoint = torch.load(
"dostoevsky_gpt.pt",
map_location="cpu"
)
# Recreate the model architecture here
# using the configuration from the training code.
model = GPT(...)
model.load_state_dict(checkpoint)
model.eval()
# Generate text
prompt = "The sky was"
tokens = tokenizer.encode(prompt)
idx = torch.tensor([tokens])
with torch.no_grad():
output = model.generate(
idx,
max_new_tokens=100,
temperature=0.8,
top_k=40
)
print(tokenizer.decode(output[0].tolist()))
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support