BEE-spoke-data
/

TinyLlama-3T-1.1bee

Text Generation

Transformers

Safetensors

text-generation-inference

Model card Files Files and versions Community

pszemraj

leaderboard-pr-bot commited on Mar 4

Commit

f5478a0

•

1 Parent(s): bba037e

Adding Evaluation Results (#1)

Browse files

- Adding Evaluation Results (f74de0fad37df5fb0839f07f3593a248afbe31cb)

Co-authored-by: Open LLM Leaderboard PR Bot <leaderboard-pr-bot@users.noreply.huggingface.co>

Files changed (1) hide show

README.md +129 -15

README.md CHANGED Viewed

@@ -1,13 +1,17 @@
 ---
 license: apache-2.0
-base_model: TinyLlama/TinyLlama-1.1B-intermediate-step-1431k-3T
 tags:
 - bees
 - bzz
 - honey
 - oprah winfrey
 metrics:
 - accuracy
 inference:
   parameters:
     max_new_tokens: 64
@@ -31,27 +35,124 @@ widget:
   example_title: Beekeeping PPE
 - text: The term "robbing" in beekeeping refers to the act of
   example_title: Robbing in Beekeeping
-- text: |-
-    Question: What's the primary function of drone bees in a hive?
-    Answer:
   example_title: Role of Drone Bees
 - text: To harvest honey from a hive, beekeepers often use a device known as a
   example_title: Honey Harvesting Device
-- text: >-
-    Problem: You have a hive that produces 60 pounds of honey per year. You
-    decide to split the hive into two. Assuming each hive now produces at a 70%
-    rate compared to before, how much honey will you get from both hives next
-    year?
-    To calculate
   example_title: Beekeeping Math Problem
 - text: In beekeeping, "swarming" is the process where
   example_title: Swarming
 pipeline_tag: text-generation
-datasets:
-- BEE-spoke-data/bees-internal
-language:
-- en
 ---
 <!-- This model card has been generated automatically according to the information the Trainer had access to. You
@@ -111,4 +212,17 @@ The following hyperparameters were used during training:
 - Transformers 4.36.2
 - Pytorch 2.1.0
 - Datasets 2.16.1
-- Tokenizers 0.15.0

 ---
+language:
+- en
 license: apache-2.0
 tags:
 - bees
 - bzz
 - honey
 - oprah winfrey
+datasets:
+- BEE-spoke-data/bees-internal
 metrics:
 - accuracy
+base_model: TinyLlama/TinyLlama-1.1B-intermediate-step-1431k-3T
 inference:
   parameters:
     max_new_tokens: 64
   example_title: Beekeeping PPE
 - text: The term "robbing" in beekeeping refers to the act of
   example_title: Robbing in Beekeeping
+- text: 'Question: What''s the primary function of drone bees in a hive?
+    Answer:'
   example_title: Role of Drone Bees
 - text: To harvest honey from a hive, beekeepers often use a device known as a
   example_title: Honey Harvesting Device
+- text: 'Problem: You have a hive that produces 60 pounds of honey per year. You decide
+    to split the hive into two. Assuming each hive now produces at a 70% rate compared
+    to before, how much honey will you get from both hives next year?
+    To calculate'
   example_title: Beekeeping Math Problem
 - text: In beekeeping, "swarming" is the process where
   example_title: Swarming
 pipeline_tag: text-generation
+model-index:
+- name: TinyLlama-3T-1.1bee
+  results:
+  - task:
+      type: text-generation
+      name: Text Generation
+    dataset:
+      name: AI2 Reasoning Challenge (25-Shot)
+      type: ai2_arc
+      config: ARC-Challenge
+      split: test
+      args:
+        num_few_shot: 25
+    metrics:
+    - type: acc_norm
+      value: 33.79
+      name: normalized accuracy
+    source:
+      url: https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard?query=BEE-spoke-data/TinyLlama-3T-1.1bee
+      name: Open LLM Leaderboard
+  - task:
+      type: text-generation
+      name: Text Generation
+    dataset:
+      name: HellaSwag (10-Shot)
+      type: hellaswag
+      split: validation
+      args:
+        num_few_shot: 10
+    metrics:
+    - type: acc_norm
+      value: 60.29
+      name: normalized accuracy
+    source:
+      url: https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard?query=BEE-spoke-data/TinyLlama-3T-1.1bee
+      name: Open LLM Leaderboard
+  - task:
+      type: text-generation
+      name: Text Generation
+    dataset:
+      name: MMLU (5-Shot)
+      type: cais/mmlu
+      config: all
+      split: test
+      args:
+        num_few_shot: 5
+    metrics:
+    - type: acc
+      value: 25.86
+      name: accuracy
+    source:
+      url: https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard?query=BEE-spoke-data/TinyLlama-3T-1.1bee
+      name: Open LLM Leaderboard
+  - task:
+      type: text-generation
+      name: Text Generation
+    dataset:
+      name: TruthfulQA (0-shot)
+      type: truthful_qa
+      config: multiple_choice
+      split: validation
+      args:
+        num_few_shot: 0
+    metrics:
+    - type: mc2
+      value: 38.13
+    source:
+      url: https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard?query=BEE-spoke-data/TinyLlama-3T-1.1bee
+      name: Open LLM Leaderboard
+  - task:
+      type: text-generation
+      name: Text Generation
+    dataset:
+      name: Winogrande (5-shot)
+      type: winogrande
+      config: winogrande_xl
+      split: validation
+      args:
+        num_few_shot: 5
+    metrics:
+    - type: acc
+      value: 60.22
+      name: accuracy
+    source:
+      url: https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard?query=BEE-spoke-data/TinyLlama-3T-1.1bee
+      name: Open LLM Leaderboard
+  - task:
+      type: text-generation
+      name: Text Generation
+    dataset:
+      name: GSM8k (5-shot)
+      type: gsm8k
+      config: main
+      split: test
+      args:
+        num_few_shot: 5
+    metrics:
+    - type: acc
+      value: 0.45
+      name: accuracy
+    source:
+      url: https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard?query=BEE-spoke-data/TinyLlama-3T-1.1bee
+      name: Open LLM Leaderboard
 ---
 <!-- This model card has been generated automatically according to the information the Trainer had access to. You
 - Transformers 4.36.2
 - Pytorch 2.1.0
 - Datasets 2.16.1
+- Tokenizers 0.15.0
+# [Open LLM Leaderboard Evaluation Results](https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard)
+Detailed results can be found [here](https://huggingface.co/datasets/open-llm-leaderboard/details_BEE-spoke-data__TinyLlama-3T-1.1bee)
+|             Metric              |Value|
+|---------------------------------|----:|
+|Avg.                             |36.46|
+|AI2 Reasoning Challenge (25-Shot)|33.79|
+|HellaSwag (10-Shot)              |60.29|
+|MMLU (5-Shot)                    |25.86|
+|TruthfulQA (0-shot)              |38.13|
+|Winogrande (5-shot)              |60.22|
+|GSM8k (5-shot)                   | 0.45|