Instructions to use Al-Sila/Rawi with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Al-Sila/Rawi with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("fill-mask", model="Al-Sila/Rawi")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForMaskedLM tokenizer = AutoTokenizer.from_pretrained("Al-Sila/Rawi") model = AutoModelForMaskedLM.from_pretrained("Al-Sila/Rawi", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Rawi
Rawi is a Classical Arabic encoder adapted to pre-modern biographical and historical prose, with a focus on the Islamic West (the Maghrib and al-Andalus). It is built for token- and span-level work on these texts: names, nisbas, places, offices and chains of transmission. On held-out western biographical entries it matches CAMeLBERT-CA with an 8,192-token context, and halves the pseudo-perplexity of the model it was adapted from. Rāwī is Arabic for the transmitter in an isnād.
Model details
- Developed by: Abderahmane Ainouche and Riadh Moulahcene, École Nationale Supérieure d'Intelligence Artificielle (ENSIA), Algiers, as part of the SILA project
- Model type: ModernBERT encoder, masked language model, 149,587,560 parameters
- Language: Classical Arabic
- Base model:
NAMAA-Space/AraModernBert-Base-V1.0at revision5b76dac145f7d4b71844c685c0627df6fc5e6fd9 - Context length: 8,192 tokens
- Tokenizer: the base model's, unchanged (50,280 entries)
- Licence: CC BY-NC-SA 4.0
How to use
Rawi needs transformers 5 (tested on 5.16.1) and no custom code. FlashAttention-2 is optional.
import torch
from transformers import AutoModelForMaskedLM, AutoTokenizer
repo = "Al-Sila/Rawi"
tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModelForMaskedLM.from_pretrained(repo).eval()
text = f"حدثنا محمد {tokenizer.mask_token} عبد الله"
inputs = tokenizer(text, return_tensors="pt")
position = (inputs["input_ids"][0] == tokenizer.mask_token_id).nonzero().item()
with torch.inference_mode():
logits = model(**inputs).logits[0, position]
print([tokenizer.decode([int(i)]).strip() for i in logits.topk(5).indices])
# ['بن', 'حدثنا', 'أبو', 'ثنا', 'قال']
For entity recognition, onomastic parsing or relation extraction, load it with
AutoModelForTokenClassification or as the encoder of your own head, and fine-tune.
Uses
Intended use. Research on pre-modern Arabic texts: a starting encoder for named-entity recognition, onomastic parsing, relation extraction and retrieval over biographical dictionaries, chronicles and geographical works, and fill-mask over historical prose.
Out of scope. Modern Standard Arabic and dialectal Arabic, which it was not adapted to. Commercial use, which the licence does not allow.
Training data
The adapt selection of OpenITI release 2025.1.9: 5,620 texts and
985,096,880 tokens of pre-modern Arabic prose. Texts flagged in the release as uncorrected OCR are
excluded. Duplicate segments are removed. Text is normalised to the base tokenizer's orthography:
presentation forms, zero-width characters, no-break spaces and tatwil.
The mixture is weighted towards the Islamic West. Western means written by an author attributed to the Maghrib or al-Andalus.
| Region | Tokens read | Segments | Share |
|---|---|---|---|
| Western (Maghrib and al-Andalus) | 150,863,465 | 81,006 | 48.0% |
| Rest of the corpus | 163,709,335 | 98,776 | 52.0% |
Held out from training:
- a fixed random 20% of the entries of the western biographical dictionaries used for annotation
- three books in full: Ibn ʿAbd al-Barr's al-Istīʿāb, Ibn Farḥūn's al-Dībāj al-mudhahhab and Ibn ʿAbd al-Malik al-Marrākushī's al-Dhayl wa-l-takmila
- evaluation segments of western, eastern and eastern biographical text
Training procedure
| Objective | Masked language modelling over whole-word spans (geometric, p = 0.2, up to 10 words) |
| Mask rate | 30% falling linearly to 15% |
| Steps | 1,200 steps of 262,144 tokens, in packed rows of 16,384 tokens |
| Tokens read | 311,397,956 |
| Schedule | The first 15% of steps at the corpus mixture, the next 65% at 50% western text, the last 20% at 80% western text |
| Optimizer | AdamW, learning rate 1e-4 after a 3% warmup, 1−√ decay to 1e-6 over the last 20% |
| Precision | bf16 |
| Compute | One NVIDIA A100 40GB on Modal: 1.5 GPU-hours, 81,648 tokens/s |
| Seed | 20260902 (single run) |
Evaluation
The base model and Rawi are scored with the same instruments on held-out text. Loss is masked language-model cross-entropy per token at a fixed 15% mask. Intervals are 95%, from 10,000 bootstrap resamples over books.
Pseudo-perplexity per word on 200 held-out western biographical entries of any length, read whole (lower is better):
| Base model | Rawi | Change |
|---|---|---|
| 51.54 | 19.79 | −62% |
Held-out loss by population (lower is better):
| Population | Books | Base model | Rawi |
|---|---|---|---|
| Western | 119 | 5.639 [5.472, 5.816] | 4.899 [4.737, 5.065] |
| Eastern | 208 | 5.364 [5.226, 5.501] | 4.617 [4.481, 4.749] |
| Western biography, held-out entries | 29 | 6.109 [5.804, 6.443] | 5.303 [4.940, 5.623] |
| Western biography, books never seen | 3 | 5.424 [4.917, 6.191] | 4.672 [4.292, 5.299] |
| Eastern biography | 118 | 4.817 [4.606, 5.030] | 4.127 [3.903, 4.348] |
Loss falls in every population, including the three books held out in full.
Onomastic cloze. The nisba or place of origin (min ahl …) in held-out western entries is masked and predicted, top-1 accuracy:
| Category | Items | Books | Base model | Rawi |
|---|---|---|---|---|
| Nisba | 2,723 | 30 | 0.248 [0.136, 0.322] | 0.265 [0.168, 0.329] |
| Place of origin | 293 | 10 | 0.512 [0.385, 0.597] | 0.570 [0.479, 0.679] |
Comparison with other Arabic encoders
Every model reads the same held-out western biographical entries: 200 drawn across 32 books, at most 7 from any one, 20 to 130 words each, with hamza, tāʾ marbūṭa and alif maqṣūra kept as written. The 3 entries that exceed one model's 256-token scoring window are set aside for all of them, leaving 197. Tokens per word is measured on 40 western texts. Intervals are 95%, from 10,000 bootstrap resamples over entries. The difference column is paired on the same entries (lower is better).
| Model | Context | Tokens per western word | Pseudo-perplexity / word | Difference from CAMeLBERT-CA |
|---|---|---|---|---|
| Rawi | 8,192 | 1.358 | 13.68 [11.76, 15.95] | −1.22 to +1.25 |
| CAMeL-Lab/bert-base-arabic-camelbert-ca | 512 | 1.450 | 13.63 [11.64, 15.91] | reference |
| NAMAA-Space/AraModernBert-Base-V1.0 (base) | 8,192 | 1.358 | 26.83 [22.33, 32.27] | +10.03 to +17.11 |
| aubmindlab/bert-base-arabertv02 | 512 | 1.381 | 32.34 [26.54, 39.42] | +14.40 to +24.13 |
| jhu-clsp/mmBERT-base | 8,192 | 1.941 | 42.59 [35.71, 51.08] | +23.47 to +35.80 |
Rawi halves its base model's pseudo-perplexity and matches CAMeLBERT-CA, the established Classical Arabic encoder, with sixteen times its context. CAMeLBERT-CA's classical training corpus is an earlier OpenITI release, so it may have read these entries. Rawi's adaptation corpus excludes them.
What was not measured
- Downstream tasks. Fine-tuned results for entity recognition, onomastic parsing and relation extraction will be reported when the annotated gold data is complete.
- Behaviour on Modern Standard Arabic, dialectal Arabic and manuscript transcriptions.
Limitations
- Western text is still harder for the model than eastern text. The western minus eastern loss gap is 0.275 for the base model and 0.282 for Rawi.
- Some western nisbas are weak. In fill-mask, al-Maghribī and al-Ifrīqī still rank below common eastern nisbas such as al-Baṣrī.
- The source texts are printed critical editions, mostly unvocalised. Input with full diacritics is outside what the model was trained on.
- Results are from one training run.
Glossary
- Isnād: the chain of transmitters through whom a report or a text was received, written as a sequence such as ḥaddathanā X ʿan Y.
- Nasab: the line of descent in a name, joined by ibn (son of), as in Muḥammad ibn ʿAbd Allāh.
- Kunya: a name formed with Abū or Umm (father or mother of), as in Abū Bakr.
- Nisba: an adjective of origin, school, tribe or trade, as in al-Andalusī, al-Mālikī or al-Qurṭubī.
- Min ahl …: "from the people of …", the formula that gives a scholar's town or region in a biographical entry.
- Pseudo-perplexity: a masked language model's perplexity computed by masking each token in turn and scoring the model's prediction for it. Per word, so models with different tokenizers are comparable.
Licence
CC BY-NC-SA 4.0. The base model is Apache-2.0. The training corpus, OpenITI, is CC BY-NC-SA 4.0, and its share-alike terms are applied to the weights. Models built on Rawi carry the same licence.
Citation
@misc{sila_rawi_2026,
title = {{Rawi}: A Classical Arabic Encoder for the Biographical Literature of the Islamic West},
author = {Ainouche, Abderahmane and Moulahcene, Riadh},
year = {2026},
url = {https://huggingface.co/Al-Sila/Rawi},
note = {Base model NAMAA-Space/AraModernBert-Base-V1.0; trained on OpenITI 2025.1.9 (CC BY-NC-SA 4.0)}
}
- Downloads last month
- 19
Model tree for Al-Sila/Rawi
Base model
NAMAA-Space/AraModernBert-Base-V1.0