Eça

Eça logo: a cream letter C with a red cedilla on a blueprint tile

Small GPT pre-trained from scratch on the collected works of the Portuguese novelist Eça de Queirós, from the jgalego/eca-queiros dataset. Byte-level BPE tokenizer trained on the same text. eca.py in this repo holds the model and doubles as the loader.

"As enxergas rijas fazem as almas fortes." (A Cidade e as Serras)

🚀 Usage

uv run https://huggingface.co/jgalego/eca/resolve/main/eca.py write "A estação de Ovar"

From Python, with eca.py on the path:

from eca import load, write

model, tokenizer = load("jgalego/eca")
print(write(model, tokenizer, "A estação de Ovar", count=100))

🏋️ Training

Model Decoder-only transformer, 5.3M parameters, 6 layers, width 256, 4 heads, context 256, vocabulary 2048
Data First 98% of the text: 3,505,536 tokens. Random windows of 256 tokens.
Steps 3,000, batch size 64, AdamW, lr 0.001 with warmup and cosine decay, dropout 0.1
Selection Weights from the step with the best validation score (step 3,000)
Hardware NVIDIA A10G, 0.03 h

📊 Results

Validation is the last 2% of the text, 205,864 characters (70,135 tokens), scored in non-overlapping windows of 256 tokens. Bits per character is comparable across tokenizers.

Loss per token Bits per character
3.3384 1.6408

⚠️ Limitations

  • Trained on about 10M characters, so it imitates the style of the novels and has no facts or reasoning to speak of.
  • Spelling follows 19th-century Portuguese, not the current orthography.
  • The validation text is the end of one book, so the score reflects that book more than the whole corpus.
Downloads last month
106
Safetensors
Model size
5.33M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train jgalego/eca

Collection including jgalego/eca