RefalMachine
/

ruadapt_qwen2.5_3B_ext_u48_instruct_v4

Text Generation

Safetensors

Russian

qwen2

conversational

Model card Files Files and versions Community

RefalMachine commited on about 1 month ago

Commit

ec34b94

•

1 Parent(s): ee1059f

Update README.md

Browse files

Files changed (1) hide show

README.md +14 -4

README.md CHANGED Viewed

@@ -10,19 +10,21 @@ base_model:
 - RefalMachine/ruadapt_qwen2.5_3B_ext_u48_full_lr5e4_peft_mlp_32_32_bs256
 ---
-### Model description
 Instruction-tuned version of RefalMachine/ruadapt_qwen2.5_3B_ext_u48_full_lr5e4_peft_mlp_32_32_bs256 with extended tokenizer after LEP (Learned Embedding Propagation, paper will be soon) procedure.
 Thanks to the extended tokenizer, the model works more efficiently with the Russian language (up to 60% speed up compared to Qwen-2.5-3B-Instruct in terms of characters)
-### Метрики и оценка качества
 #### Результаты на Ru-Arena-General
 В качестве референсых ответов, с которыми сравниваются модели выступают ответы от gpt-3.5-turbo-0125, поэтому она имеет винрейт 50%.
-Здесь приведена лишь часть лидерборда, подробнее смотрите в репозитории бенчмарка.
 | Model Name                                       | Winrate  | 95% CI             | Average # Tokens |
 |--------------------------------------------------|--------|--------------------|------------------|
@@ -41,7 +43,15 @@ Thanks to the extended tokenizer, the model works more efficiently with the Russ
 | c4ai-command-r-v01                               | 49.0   | (-1.7, 2.2)        | 529              |
 | meta-llama-3.1-8b-instruct                       | 43.1   | (-2.8, 2.3)        | 628              |
-### How to cite:
 Tikhomirov M., Chernyshev D. Facilitating large language model Russian adaptation with Learned Embedding Propagation // 2024 (will be soon)

 - RefalMachine/ruadapt_qwen2.5_3B_ext_u48_full_lr5e4_peft_mlp_32_32_bs256
 ---
+## Model description
 Instruction-tuned version of RefalMachine/ruadapt_qwen2.5_3B_ext_u48_full_lr5e4_peft_mlp_32_32_bs256 with extended tokenizer after LEP (Learned Embedding Propagation, paper will be soon) procedure.
 Thanks to the extended tokenizer, the model works more efficiently with the Russian language (up to 60% speed up compared to Qwen-2.5-3B-Instruct in terms of characters)
+## Метрики и оценка качества
+Модель была оценена на Ru-Arena-General, MERA, llmtf_open
 #### Результаты на Ru-Arena-General
 В качестве референсых ответов, с которыми сравниваются модели выступают ответы от gpt-3.5-turbo-0125, поэтому она имеет винрейт 50%.
+Приведена лишь часть лидерборда, подробнее смотрите в репозитории бенчмарка (https://huggingface.co/spaces/Vikhrmodels/arenahardlb).
 | Model Name                                       | Winrate  | 95% CI             | Average # Tokens |
 |--------------------------------------------------|--------|--------------------|------------------|
 | c4ai-command-r-v01                               | 49.0   | (-1.7, 2.2)        | 529              |
 | meta-llama-3.1-8b-instruct                       | 43.1   | (-2.8, 2.3)        | 628              |
+#### Результаты на MERA
+TODO
+#### Результаты на llmtf_open
+TODO
+## How to cite:
 Tikhomirov M., Chernyshev D. Facilitating large language model Russian adaptation with Learned Embedding Propagation // 2024 (will be soon)