Polish ModernBERT 512 Large

Polish ModernBERT is a family of monolingual ModernBERT-based encoders pretrained for Polish. The family covers two model scales (Base and Large) and two context lengths (512 and 8,192 tokens). The models are general-purpose pretrained encoders intended to serve as strong foundations for a broad range of Polish NLP applications. They can be fine-tuned or adapted for downstream tasks including, but not limited to, text and document classification, sequence labeling, regression, semantic similarity, retrieval, reranking, and representation learning.

📄 Paper: Polish ModernBERT: The Long and Short of Polish Language Understanding
🤗 Model collection: Polish ModernBERT
📚 LongContext benchmark: Dataset

Model family

Model Parameters Context Vocabulary Tokenizer
pl-ModernBERT-512-base 149M 512 50,008 SentencePiece Unigram
pl-ModernBERT-512-large 475M 512 128,256 SentencePiece Unigram + byte fallback
pl-ModernBERT-base 149M 8,192 50,008 SentencePiece Unigram
pl-ModernBERT-large 475M 8,192 128,256 SentencePiece Unigram + byte fallback

All variants use the core ModernBERT architecture with RoPE, GeGLU feed-forward layers, pre-normalization, and alternating global/local attention. The Base models use 22 layers with hidden size 768, while the Large models use 28 layers with hidden size 1,024.

Training

The models were pretrained on approximately 44.5B Polish tokens (197 GB after preprocessing) from a curated Polish corpus, Common Crawl (CC-MAIN-2019-43), and the Polish pol_Latn subset of FineTranslations.

Pretraining follows a staged recipe selected through downstream validation: four 512-token stages progressively transition from token-level MLM to whole-word masking, increased emphasis on curated data, and final annealing, followed by long-context continuation to 8,192 tokens. During context extension, the global RoPE theta is increased from 10,000 to 160,000.

Pretraining schedule used for Polish ModernBERT. The final 512-token checkpoints initialize the corresponding 8K variants.

Training was performed with Composer in BF16 using StableAdamW and distributed data parallelism across 8 NVIDIA GH200 GPUs.

Evaluation

The models were evaluated on 30 Polish NLU tasks spanning KLEJ, FinBench, the five-task LongContext benchmark, and a range of additional Polish NLP tasks. Each model-task configuration was fine-tuned using five random seeds; scores below are means on a 0–100 scale.

Evaluation Results

The tables below summarize average performance across the main evaluation groups. Scores are reported on a 0–100 scale. Overall Avg. is the macro-average across all 30 individual tasks rather than the average of the four group-level scores.

Base models

Model Context KLEJ Avg. FinBench Avg. Other Tasks Avg. LongContext Avg. Overall Avg.
XLM-R Base 512 84.46 83.13 77.10 50.75 76.12
HerBERT Base 512 85.83 83.65 79.28 61.93 79.23
Polish RoBERTa-v2 Base 512 86.75 85.14 80.26 57.66 79.42
Polish ModernBERT 512 Base 512 86.84 86.95 82.07 71.59 82.73
EuroBERT-210M 8K 77.16 82.38 73.71 72.30 76.24
mmBERT Small 8K 80.11 83.07 78.01 69.02 78.15
Polish RoBERTa-8K Base 8K 86.86 85.98 82.36 67.47 81.95
Polish ModernBERT Base 8K 87.00 86.69 83.08 77.15 83.99

Large models

Model Context KLEJ Avg. FinBench Avg. Other Tasks Avg. LongContext Avg. Overall Avg.
XLM-R Large 512 87.29 85.44 82.64 67.87 82.13
HerBERT Large 512 87.83 87.33 83.42 70.26 83.33
Polish RoBERTa-v2 Large 512 88.69 87.90 83.58 70.49 83.80
Polish ModernBERT 512 Large 512 88.48 88.18 83.60 73.58 84.31
EuroBERT-610M 8K 80.10 86.11 78.95 77.36 80.46
mmBERT Base 8K 83.17 85.69 80.80 73.44 81.27
Polish RoBERTa-8K Large 8K 88.52 87.98 84.27 75.88 84.89
Polish ModernBERT Large 8K 88.31 87.83 83.90 78.49 85.11

KLEJ

Polish ModernBERT is competitive with the strongest Polish BERT/RoBERTa encoders on KLEJ. The Base variants obtain the highest KLEJ average in both context settings, while the Large variants remain within 0.21 points of the corresponding Polish RoBERTa models.

Detailed KLEJ results - Base models
Task pl-RoBERTa-v2-base pl-ModernBERT-512-base pl-RoBERTa-8K-base pl-ModernBERT-base
NKJP-NER 94.32 94.38 94.16 94.49
CDSC-E 94.05 94.66 94.54 94.46
CDSC-R 94.64 94.11 94.90 94.06
CBD 70.57 68.56 69.35 71.40
POLEMO-IN 90.97 92.88 91.27 92.14
POLEMO-OUT 79.11 83.77 81.26 83.04
DYK 70.38 66.90 69.35 67.28
PSC 98.88 97.79 98.90 97.68
AR 87.83 88.53 88.05 88.46
Average 86.75 86.84 86.86 87.00
Detailed KLEJ results - Large models
Task pl-RoBERTa-v2-large pl-ModernBERT-512-large pl-RoBERTa-8K-large pl-ModernBERT-large
NKJP-NER 95.75 95.05 95.64 94.38
CDSC-E 94.16 94.60 94.28 94.68
CDSC-R 95.25 95.14 95.33 94.47
CBD 73.10 72.60 73.23 71.47
POLEMO-IN 93.55 93.38 93.05 93.05
POLEMO-OUT 83.81 84.41 83.64 84.78
DYK 74.87 73.28 74.05 74.63
PSC 98.37 98.81 98.56 98.47
AR 89.36 89.07 88.91 88.88
Average 88.69 88.48 88.52 88.31

FinBench

Polish ModernBERT shows particularly strong performance on FinBench. The Base variants achieve the highest FinBench average in both context settings. At Large scale, the 512-token Polish ModernBERT obtains the highest average, while the 8K variant remains close to the corresponding Polish RoBERTa model.

Detailed FinBench results - Base models
Task pl-RoBERTa-v2-base pl-ModernBERT-512-base pl-RoBERTa-8K-base pl-ModernBERT-base
Banking-Short 78.75 80.41 79.79 80.08
Banking-Long 85.03 87.29 86.99 87.16
Banking77 88.26 91.85 89.27 91.66
FPB 83.55 83.20 83.63 83.40
GCN 95.02 94.87 94.87 94.83
Stooq 80.25 84.08 81.32 83.03
Average 85.14 86.95 85.98 86.69
Detailed FinBench results - Large models
Task pl-RoBERTa-v2-large pl-ModernBERT-512-large pl-RoBERTa-8K-large pl-ModernBERT-large
Banking-Short 81.69 82.07 81.99 81.94
Banking-Long 87.89 88.40 88.35 88.89
Banking77 92.45 92.96 92.74 92.62
FPB 85.26 84.80 85.42 84.60
GCN 95.04 95.08 94.97 94.88
Stooq 85.07 85.77 84.41 84.02
Average 87.90 88.18 87.98 87.83

Other Tasks

Across the additional Polish NLP tasks, the Base variants of Polish ModernBERT obtain the highest average in both context settings. At Large scale, the 512-token model achieves the highest average, while the 8K variant remains competitive with the corresponding Polish RoBERTa model.

Detailed Other Tasks results - Base models
Task pl-RoBERTa-v2-base pl-ModernBERT-512-base pl-RoBERTa-8K-base pl-ModernBERT-base
8TAGS 78.03 80.69 79.21 80.86
BAN-PL 92.19 93.10 92.62 93.08
MIPD 58.58 67.11 64.39 68.03
PPC 87.05 84.40 86.02 85.70
SICK-E 86.71 86.31 86.31 86.61
SICK-R 82.58 83.16 83.16 83.77
TwitterEMO 66.46 69.02 68.75 69.52
IMDB 91.06 92.05 95.02 94.40
EURLEX 74.51 79.31 79.12 79.61
NKJP-NER* 85.41 85.54 88.97 89.21
Average 80.26 82.07 82.36 83.08
Detailed Other Tasks results - Large models
Task pl-RoBERTa-v2-large pl-ModernBERT-512-large pl-RoBERTa-8K-large pl-ModernBERT-large
8TAGS 81.64 82.50 81.44 82.24
BAN-PL 93.80 94.00 93.99 93.51
MIPD 67.27 68.28 68.50 68.99
PPC 89.96 88.04 89.48 87.20
SICK-E 88.33 87.88 88.96 87.47
SICK-R 85.93 84.69 86.54 84.91
TwitterEMO 70.70 70.20 70.60 70.35
IMDB 94.36 93.77 96.03 95.93
EURLEX 79.19 79.84 79.77 79.76
NKJP-NER* 84.62 86.84 87.36 88.66
Average 83.58 83.60 84.27 83.90

Long-context performance

Polish ModernBERT shows its largest gains on the LongContext benchmark, which consists of five tasks designed specifically to evaluate long-document understanding. The 8K variants achieve the highest LongContext average at both model scales.

At Base scale, pl-ModernBERT-base improves over pl-RoBERTa-8K-base by 9.68 points (77.15 vs. 67.47) while using 22% fewer parameters (149M vs. 190M). At Large scale, pl-ModernBERT-large improves over the corresponding Polish RoBERTa-8K baseline by 2.61 points (78.49 vs. 75.88).

Detailed LongContext results - Base models
Task pl-RoBERTa-v2-base pl-ModernBERT-512-base pl-RoBERTa-8K-base pl-ModernBERT-base
SCOTUS-Dom 79.12 82.81 79.26 84.48
SCOTUS-Dec 69.86 70.74 63.20 77.79
BookSummary 85.02 83.71 88.96 90.22
ECtHR-PL-AVA 33.65 61.42 64.66 68.01
ECtHR-PL-VA 20.65 59.29 41.28 65.27
Average 57.66 71.59 67.47 77.15
Detailed LongContext results - Large models
Task pl-RoBERTa-v2-large pl-ModernBERT-512-large pl-RoBERTa-8K-large pl-ModernBERT-large
SCOTUS-Dom 83.01 83.83 83.99 85.78
SCOTUS-Dec 72.46 71.39 68.73 78.21
BookSummary 87.47 86.54 93.11 91.74
ECtHR-PL-AVA 58.33 64.62 68.29 69.48
ECtHR-PL-VA 51.17 61.50 65.27 67.24
Average 70.49 73.58 75.88 78.49

The 512-token variants are evaluated on LongContext with truncation and should be treated as practical short-context baselines rather than controlled context-length ablations.

Inference efficiency

In a common BF16 inference setup on a single NVIDIA H100, Polish ModernBERT provides favorable quality–efficiency trade-offs relative to the corresponding Polish RoBERTa baselines.

For 512-token inference, the Base and Large models reduce latency by approximately 26% and 47%, respectively, while reducing peak GPU memory usage by 54% and 18%. In the 8K setting, the corresponding latency reductions are 6% and 25%, with peak memory reductions of 24% and 21%.

Quality–efficiency trade-offs for Polish ModernBERT and the evaluated encoder baselines. The x-axis shows inference latency per sample, the y-axis shows downstream performance, and marker size represents peak GPU memory usage. Measurements were performed in BF16 on a single NVIDIA H100 GPU.

Usage

from transformers import AutoTokenizer, AutoModel

model_id = "OPI-PIB/pl-ModernBERT-512-large"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModel.from_pretrained(model_id)

text = "W białodrzewiu jaśnie dźni słoneczno, miodzie złoci białopałem żyśnie, drzewia pełni pszczelą i pasieczną, a przez liście kraśnie pęk słowiśnie."

inputs = tokenizer(
    text,
    return_tensors="pt",
    truncation=True,
    max_length=512,
)

outputs = model(**inputs)
token_embeddings = outputs.last_hidden_state

For downstream tasks such as classification, regression, sequence labeling, retrieval, or reranking, the pretrained encoder can be adapted to the target task.

The raw checkpoints are not retrieval-specific sentence-embedding models. PIRB retrieval results reported in the paper were obtained after contrastive fine-tuning.

For longer inputs, use one of the corresponding 8K checkpoints from the Polish ModernBERT family.

Flash Attention 2

🏎️ For the highest training and inference efficiency, we recommend using Polish ModernBERT with Flash Attention 2 when supported by your GPU and environment.

pip install flash-attn --no-build-isolation

The model can then be loaded with:

from transformers import AutoModel

model = AutoModel.from_pretrained(
    model_id,
    attn_implementation="flash_attention_2",
    torch_dtype="auto",
)

Authors

Michał Perełkiewicz, Sławomir Dadas, Rafał Poświata, Małgorzata Grębowiec

AI Lab, National Information Processing Institute
(Ośrodek Przetwarzania Informacji – Państwowy Instytut Badawczy, OPI PIB)

Warsaw, Poland

Corresponding author: mperelkiewicz@opi.org.pl

Acknowledgments

This work was supported by the Gaia AI Factory project, funded by the European Union under Grant Agreement No. 101314359 through the EuroHPC Joint Undertaking (EuroHPC JU).

We gratefully acknowledge the Polish high-performance computing infrastructure PLGrid (HPC Center: ACK Cyfronet AGH) for providing computational resources and support within computational grant PLG/2025/018315.

Citation

If you use Polish ModernBERT in your work, please cite:

@misc{perełkiewicz2026polishmodernbertlongshort,
      title={Polish ModernBERT: The Long and Short of Polish Language Understanding}, 
      author={Michał Perełkiewicz and Sławomir Dadas and Rafał Poświata and Małgorzata Grębowiec},
      year={2026},
      eprint={2609.01379},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2609.01379}, 
}
Downloads last month
12
Safetensors
Model size
0.5B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including OPI-PIB/pl-ModernBERT-512-large

Paper for OPI-PIB/pl-ModernBERT-512-large