Text Generation
Transformers
PyTorch
English
llama
Eval Results
Inference Endpoints
text-generation-inference
Commit
2c6ddf2
1 Parent(s): e63a535

Adding Evaluation Results (#4)

Browse files

- Adding Evaluation Results (69c9834c8d6ff51179344da3a52b4d4dd7e2438a)


Co-authored-by: Open LLM Leaderboard PR Bot <leaderboard-pr-bot@users.noreply.huggingface.co>

Files changed (1) hide show
  1. README.md +120 -3
README.md CHANGED
@@ -1,5 +1,8 @@
1
  ---
 
 
2
  license: llama2
 
3
  datasets:
4
  - garage-bAInd/Open-Platypus
5
  - ehartford/dolphin
@@ -10,10 +13,110 @@ datasets:
10
  - tatsu-lab/alpaca
11
  - databricks/databricks-dolly-15k
12
  - WizardLM/WizardLM_evol_instruct_V2_196k
13
- language:
14
- - en
15
- library_name: transformers
16
  pipeline_tag: text-generation
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
17
  ---
18
 
19
  # model_007_13b_v2
@@ -210,3 +313,17 @@ Detailed results can be found [here](https://huggingface.co/datasets/open-llm-le
210
  | Winogrande (5-shot) | 75.85 |
211
  | GSM8K (5-shot) | 1.36 |
212
  | DROP (3-shot) | 43.97 |
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  ---
2
+ language:
3
+ - en
4
  license: llama2
5
+ library_name: transformers
6
  datasets:
7
  - garage-bAInd/Open-Platypus
8
  - ehartford/dolphin
 
13
  - tatsu-lab/alpaca
14
  - databricks/databricks-dolly-15k
15
  - WizardLM/WizardLM_evol_instruct_V2_196k
 
 
 
16
  pipeline_tag: text-generation
17
+ model-index:
18
+ - name: model_007_13b_v2
19
+ results:
20
+ - task:
21
+ type: text-generation
22
+ name: Text Generation
23
+ dataset:
24
+ name: AI2 Reasoning Challenge (25-Shot)
25
+ type: ai2_arc
26
+ config: ARC-Challenge
27
+ split: test
28
+ args:
29
+ num_few_shot: 25
30
+ metrics:
31
+ - type: acc_norm
32
+ value: 61.95
33
+ name: normalized accuracy
34
+ source:
35
+ url: https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard?query=psmathur/model_007_13b_v2
36
+ name: Open LLM Leaderboard
37
+ - task:
38
+ type: text-generation
39
+ name: Text Generation
40
+ dataset:
41
+ name: HellaSwag (10-Shot)
42
+ type: hellaswag
43
+ split: validation
44
+ args:
45
+ num_few_shot: 10
46
+ metrics:
47
+ - type: acc_norm
48
+ value: 82.48
49
+ name: normalized accuracy
50
+ source:
51
+ url: https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard?query=psmathur/model_007_13b_v2
52
+ name: Open LLM Leaderboard
53
+ - task:
54
+ type: text-generation
55
+ name: Text Generation
56
+ dataset:
57
+ name: MMLU (5-Shot)
58
+ type: cais/mmlu
59
+ config: all
60
+ split: test
61
+ args:
62
+ num_few_shot: 5
63
+ metrics:
64
+ - type: acc
65
+ value: 57.32
66
+ name: accuracy
67
+ source:
68
+ url: https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard?query=psmathur/model_007_13b_v2
69
+ name: Open LLM Leaderboard
70
+ - task:
71
+ type: text-generation
72
+ name: Text Generation
73
+ dataset:
74
+ name: TruthfulQA (0-shot)
75
+ type: truthful_qa
76
+ config: multiple_choice
77
+ split: validation
78
+ args:
79
+ num_few_shot: 0
80
+ metrics:
81
+ - type: mc2
82
+ value: 53.5
83
+ source:
84
+ url: https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard?query=psmathur/model_007_13b_v2
85
+ name: Open LLM Leaderboard
86
+ - task:
87
+ type: text-generation
88
+ name: Text Generation
89
+ dataset:
90
+ name: Winogrande (5-shot)
91
+ type: winogrande
92
+ config: winogrande_xl
93
+ split: validation
94
+ args:
95
+ num_few_shot: 5
96
+ metrics:
97
+ - type: acc
98
+ value: 75.85
99
+ name: accuracy
100
+ source:
101
+ url: https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard?query=psmathur/model_007_13b_v2
102
+ name: Open LLM Leaderboard
103
+ - task:
104
+ type: text-generation
105
+ name: Text Generation
106
+ dataset:
107
+ name: GSM8k (5-shot)
108
+ type: gsm8k
109
+ config: main
110
+ split: test
111
+ args:
112
+ num_few_shot: 5
113
+ metrics:
114
+ - type: acc
115
+ value: 1.36
116
+ name: accuracy
117
+ source:
118
+ url: https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard?query=psmathur/model_007_13b_v2
119
+ name: Open LLM Leaderboard
120
  ---
121
 
122
  # model_007_13b_v2
 
313
  | Winogrande (5-shot) | 75.85 |
314
  | GSM8K (5-shot) | 1.36 |
315
  | DROP (3-shot) | 43.97 |
316
+
317
+ # [Open LLM Leaderboard Evaluation Results](https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard)
318
+ Detailed results can be found [here](https://huggingface.co/datasets/open-llm-leaderboard/details_psmathur__model_007_13b_v2)
319
+
320
+ | Metric |Value|
321
+ |---------------------------------|----:|
322
+ |Avg. |55.41|
323
+ |AI2 Reasoning Challenge (25-Shot)|61.95|
324
+ |HellaSwag (10-Shot) |82.48|
325
+ |MMLU (5-Shot) |57.32|
326
+ |TruthfulQA (0-shot) |53.50|
327
+ |Winogrande (5-shot) |75.85|
328
+ |GSM8k (5-shot) | 1.36|
329
+