jgalego/eca-queiros
Viewer • Updated • 62.1k • 20
Small GPT pre-trained from scratch on the collected works of the Portuguese novelist Eça de Queirós, from the jgalego/eca-queiros dataset. Byte-level BPE tokenizer trained on the same text. eca.py in this repo holds the model and doubles as the loader.
"As enxergas rijas fazem as almas fortes." (A Cidade e as Serras)
uv run https://huggingface.co/jgalego/eca/resolve/main/eca.py write "A estação de Ovar"
From Python, with eca.py on the path:
from eca import load, write
model, tokenizer = load("jgalego/eca")
print(write(model, tokenizer, "A estação de Ovar", count=100))
| Model | Decoder-only transformer, 5.3M parameters, 6 layers, width 256, 4 heads, context 256, vocabulary 2048 |
| Data | First 98% of the text: 3,505,536 tokens. Random windows of 256 tokens. |
| Steps | 3,000, batch size 64, AdamW, lr 0.001 with warmup and cosine decay, dropout 0.1 |
| Selection | Weights from the step with the best validation score (step 3,000) |
| Hardware | NVIDIA A10G, 0.03 h |
Validation is the last 2% of the text, 205,864 characters (70,135 tokens), scored in non-overlapping windows of 256 tokens. Bits per character is comparable across tokenizers.
| Loss per token | Bits per character |
|---|---|
| 3.3384 | 1.6408 |