RUMI-8M
RUMI-8M is a small decoder-only Transformer language model trained from scratch as part of the RUMI project.
Model Information
- Model name: RUMI-8M
- Parameters: 8,459,904
- Architecture: Decoder-only Transformer
- Vocabulary size: 16,000
- Context length: 256
- Embedding dimension: 192
- Transformer layers: 12
- Attention heads: 6
- Feed-forward dimension: 768
- Dropout: 0.1
- Weight tying: enabled
- Tokenizer: Byte-Level BPE
Important
RUMI-8M is a base language model.
It is not instruction-tuned and is not intended to behave like ChatGPT. Its primary capability is language-model text continuation.
Load RUMI-8M
Install the package dependencies:
pip install torch tokenizers huggingface_hub