EleutherAI
/

gpt-j-6b

@@ -4,11 +4,12 @@ language:
 tags:
 - pytorch
 - causal-lm
 license: apache-2.0
 datasets:
-- EleutherAI/pile
 ---
 # GPT-J 6B
 ## Model Description
@@ -60,7 +61,7 @@ respond to a given prompt the way a product like ChatGPT does. This is because,
  unlike this model, ChatGPT was fine-tuned using methods such as Reinforcement
 Learning from Human Feedback (RLHF) to better “follow” human instructions.
-### Limitations and Biases
 The core functionality of GPT-J is taking a string of text and predicting the next token. While language models are widely used for tasks other than this, there are a lot of unknowns with this work. When prompting GPT-J it is important to remember that the statistically most likely next token is often not the token that produces the most "accurate" text. Never depend upon GPT-J to produce factually accurate output.
@@ -103,7 +104,7 @@ This model was trained for 402 billion tokens over 383,500 steps on TPU v3-256 p
 | GPT-3 Babbage&ddagger;   | &cross;     | -----          | 5.58          | 62.4%         | 59.0%        | 54.5%       | 75.5%     | -----             |
 | Megatron-8.3B&ast;       | &cross;     | 7.8e21         | -----         | 66.5%         | -----        | -----       | -----     | 174               |
 | GPT-3 2.7B&ast;&ddagger; | &cross;     | 4.8e21         | 4.60          | 67.1%         | 62.3%        | 62.8%       | 75.6%     | ~800              |
-| Megatron-11B&dagger;     | &check;     | 1.0e22         | -----         | -----         | -----        | -----       | -----     | 161               |
 | **GPT-J 6B&ddagger;**    | **&check;** | **1.5e22**     | **3.99**      | **69.7%**     | **65.3%**    | **66.1%**   | **76.5%** | **825**           |
 | GPT-3 6.7B&ast;&ddagger; | &cross;     | 1.2e22         | 4.00          | 70.3%         | 64.5%        | 67.4%       | 78.0%     | ~800              |
 | GPT-3 Curie&ddagger;     | &cross;     | -----          | 4.00          | 69.3%         | 65.6%        | 68.5%       | 77.9%     | -----             |

 tags:
 - pytorch
 - causal-lm
+- code
 license: apache-2.0
 datasets:
+- nampdn-ai/tiny-codes
 ---
+# WORM-GPT-J
 # GPT-J 6B
 ## Model Description
  unlike this model, ChatGPT was fine-tuned using methods such as Reinforcement
 Learning from Human Feedback (RLHF) to better “follow” human instructions.
+### NO Limitations and Biases
 The core functionality of GPT-J is taking a string of text and predicting the next token. While language models are widely used for tasks other than this, there are a lot of unknowns with this work. When prompting GPT-J it is important to remember that the statistically most likely next token is often not the token that produces the most "accurate" text. Never depend upon GPT-J to produce factually accurate output.
 | GPT-3 Babbage&ddagger;   | &cross;     | -----          | 5.58          | 62.4%         | 59.0%        | 54.5%       | 75.5%     | -----             |
 | Megatron-8.3B&ast;       | &cross;     | 7.8e21         | -----         | 66.5%         | -----        | -----       | -----     | 174               |
 | GPT-3 2.7B&ast;&ddagger; | &cross;     | 4.8e21         | 4.60          | 67.1%         | 62.3%        | 62.8%       | 75.6%     | ~800              |
+| **Megatron-11B&dagger;**     | &check;     | 1.0e22         | -----         | -----         | -----        | -----       | -----     | 161               |
 | **GPT-J 6B&ddagger;**    | **&check;** | **1.5e22**     | **3.99**      | **69.7%**     | **65.3%**    | **66.1%**   | **76.5%** | **825**           |
 | GPT-3 6.7B&ast;&ddagger; | &cross;     | 1.2e22         | 4.00          | 70.3%         | 64.5%        | 67.4%       | 78.0%     | ~800              |
 | GPT-3 Curie&ddagger;     | &cross;     | -----          | 4.00          | 69.3%         | 65.6%        | 68.5%       | 77.9%     | -----             |