File size: 883 Bytes
0d8f644
 
 
 
 
 
 
306c059
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
0d8f644
1812076
 
2198ac3
 
1c13844
a1b0465
2198ac3
 
0d8f644
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
---
license: gpl-3.0
language:
- en
- zh
- ja
- de
datasets:
- JosephusCheung/GuanacoDataset
- meta-math/MetaMathQA
- jondurbin/airoboros-3.1
- WizardLM/WizardLM_evol_instruct_V2_196k
- RyokoAI/ShareGPT52K
- RyokoAI/Fandom23K
- milashkaarshif/MoeGirlPedia_wikitext_raw_archive
- wikipedia
- wiki_lingua
- garage-bAInd/Open-Platypus
- LDJnr/Puffin
- BAAI/COIG
- TigerResearch/tigerbot-zhihu-zh-10k
- liwu/MNBVC
- teknium/openhermes
- CausalLM/Refined-Anime-Text
- microsoft/orca-math-word-problems-200k
- m-a-p/CodeFeedback-Filtered-Instruction
---
## TBA

Tokenizer is different from cohere - and chat template is ChatML - fully fine-tuned at 128K+

No loras, no quants, no tricks, 30M+ sft data.

Pressure Testing from: https://github.com/LeonEricsson/llmcontext

![image/png](https://cdn-uploads.huggingface.co/production/uploads/63468a143ea42ee2cb49ddd1/2XbONpyTeMH1qWCtE9ziH.png)