metadata
license: cc-by-nc-nd-4.0
datasets:
- taln-ls2n/Adminset
language:
- fr
library_name: transformers
tags:
- camembert
- BERT
- Administrative documents
AdminBERT 4GB: A Small French Language model adapted to Administrative documents
AdminBERT-4GB is a French language model adapted on a large corpus of 10 millions French administrative texts. It is a derivative of CamemBERT model, which is based on the RoBERTa architecture. AdminBERT-4GB is trained using the Masked Language Modeling (MLM) objective with 30% mask rate for 2 epochs on 8 V100 GPUs. The dataset used for training is a sample of Adminset.