internlm
/

internlm2-20b

Text Generation

Transformers

PyTorch

internlm2

custom_code

Model card Files Files and versions Community

x54-729 commited on Jan 17

Commit

695532d

•

1 Parent(s): fe1d049

update readme

Browse files

Files changed (1) hide show

README.md +59 -46

README.md CHANGED Viewed

@@ -25,38 +25,45 @@ pipeline_tag: text-generation
 ## Introduction
-InternLM has open-sourced a 7 billion parameter base model tailored for practical scenarios. The model has the following characteristics:
-- It leverages trillions of high-quality tokens for training to establish a powerful knowledge base.
-- It provides a versatile toolset for users to flexibly build their own workflows.
-## InternLM-7B
 ### Performance Evaluation
-We conducted a comprehensive evaluation of InternLM using the open-source evaluation tool [OpenCompass](https://github.com/internLM/OpenCompass/). The evaluation covered five dimensions of capabilities: disciplinary competence, language competence, knowledge competence, inference competence, and comprehension competence. Here are some of the evaluation results, and you can visit the [OpenCompass leaderboard](https://opencompass.org.cn/rank) for more evaluation results.
-| Datasets\Models           |  **InternLM-Chat-7B** |  **InternLM-7B**  |  LLaMA-7B | Baichuan-7B | ChatGLM2-6B | Alpaca-7B | Vicuna-7B |
-| -------------------- | --------------------- | ---------------- | --------- |  --------- | ------------ | --------- | ---------- |
-| C-Eval(Val)          |      53.2             |        53.4       | 24.2      | 42.7       |  50.9       |  28.9     | 31.2     |
-| MMLU                 |      50.8             |       51.0        | 35.2*     |  41.5      |  46.0       |  39.7     | 47.3     |
-| AGIEval              |      42.5             |       37.6        | 20.8      | 24.6       |  39.0       | 24.1      | 26.4     |
-| CommonSenseQA        |      75.2             |      59.5         | 65.0      | 58.8       | 60.0        | 68.7      | 66.7     |
-| BUSTM                |      74.3             |       50.6        | 48.5      | 51.3        | 55.0        | 48.8      | 62.5     |
-| CLUEWSC              |      78.6             |      59.1         |  50.3     |  52.8     |  59.8     |   50.3    |  52.2     |
-| MATH                 |      6.4            |         7.1        |  2.8       | 3.0       | 6.6       |  2.2      | 2.8       |
-| GSM8K                |      34.5           |        31.2        | 10.1       | 9.7       | 29.2      |  6.0      | 15.3  |
-|  HumanEval           |      14.0           |        10.4        |   14.0     | 9.2       | 9.2       | 9.2       | 11.0  |
-| RACE(High)           |      76.3           |        57.4        | 46.9*      | 28.1      | 66.3      | 40.7      | 54.0  |
-- The evaluation results were obtained from [OpenCompass 20230706](https://github.com/internLM/OpenCompass/) (some data marked with *, which means come from the original papers), and evaluation configuration can be found in the configuration files provided by [OpenCompass](https://github.com/internLM/OpenCompass/).
-- The evaluation data may have numerical differences due to the version iteration of [OpenCompass](https://github.com/internLM/OpenCompass/), so please refer to the latest evaluation results of [OpenCompass](https://github.com/internLM/OpenCompass/).
 **Limitations:** Although we have made efforts to ensure the safety of the model during the training process and to encourage the model to generate text that complies with ethical and legal requirements, the model may still produce unexpected outputs due to its size and probabilistic generation paradigm. For example, the generated responses may contain biases, discrimination, or other harmful content. Please do not propagate such content. We are not responsible for any consequences resulting from the dissemination of harmful information.
 ### Import from Transformers
-To load the InternLM 7B Chat model using Transformers, use the following code:
 ```python
 import torch
 from transformers import AutoTokenizer, AutoModelForCausalLM
@@ -67,12 +74,11 @@ model = model.eval()
 inputs = tokenizer(["A beautiful flower"], return_tensors="pt")
 for k,v in inputs.items():
     inputs[k] = v.cuda()
-gen_kwargs = {"max_length": 128, "top_p": 0.8, "temperature": 0.8, "do_sample": True, "repetition_penalty": 1.1}
 output = model.generate(**inputs, **gen_kwargs)
 output = tokenizer.decode(output[0].tolist(), skip_special_tokens=True)
 print(output)
-# <s> A beautiful flower box made of white rose wood. It is a perfect gift for weddings, birthdays and anniversaries.
-# All the roses are from our farm Roses Flanders. Therefor you know that these flowers last much longer than those in store or online!</s>
 ```
 ## Open Source License
@@ -80,36 +86,43 @@ print(output)
 The code is licensed under Apache-2.0, while model weights are fully open for academic research and also allow **free** commercial usage. To apply for a commercial license, please fill in the [application form (English)](https://wj.qq.com/s2/12727483/5dba/)/[申请表（中文）](https://wj.qq.com/s2/12725412/f7c1/). For other questions or collaborations, please contact <internlm@pjlab.org.cn>.
 ## 简介
-InternLM ，即书生·浦语大模型，包含面向实用场景的70亿参数基础模型 （InternLM-7B）。模型具有以下特点：
-- 使用上万亿高质量预料，建立模型超强知识体系；
-- 通用工具调用能力，支持用户灵活自助搭建流程；
-## InternLM-7B
 ### 性能评测
-我们使用开���评测工具 [OpenCompass](https://github.com/internLM/OpenCompass/) 从学科综合能力、语言能力、知识能力、推理能力、理解能力五大能力维度对InternLM开展全面评测，部分评测结果如下表所示，欢迎访问[ OpenCompass 榜单 ](https://opencompass.org.cn/rank)获取更多的评测结果。
-| 数据集\模型           |  **InternLM-Chat-7B** |  **InternLM-7B**  |  LLaMA-7B | Baichuan-7B | ChatGLM2-6B | Alpaca-7B | Vicuna-7B |
-| -------------------- | --------------------- | ---------------- | --------- |  --------- | ------------ | --------- | ---------- |
-| C-Eval(Val)          |      53.2             |        53.4       | 24.2      | 42.7       |  50.9       |  28.9     | 31.2     |
-| MMLU                 |      50.8             |       51.0        | 35.2*     |  41.5      |  46.0       |  39.7     | 47.3     |
-| AGIEval              |      42.5             |       37.6        | 20.8      | 24.6       |  39.0       | 24.1      | 26.4     |
-| CommonSenseQA        |      75.2             |      59.5         | 65.0      | 58.8       | 60.0        | 68.7      | 66.7     |
-| BUSTM                |      74.3             |       50.6        | 48.5      | 51.3        | 55.0        | 48.8      | 62.5     |
-| CLUEWSC              |      78.6             |      59.1         |  50.3     |  52.8     |  59.8     |   50.3    |  52.2     |
-| MATH                 |      6.4            |         7.1        |  2.8       | 3.0       | 6.6       |  2.2      | 2.8       |
-| GSM8K                |      34.5           |        31.2        | 10.1       | 9.7       | 29.2      |  6.0      | 15.3  |
-|  HumanEval           |      14.0           |        10.4        |   14.0     | 9.2       | 9.2       | 9.2       | 11.0  |
-| RACE(High)           |      76.3           |        57.4        | 46.9*      | 28.1      | 66.3      | 40.7      | 54.0  |
-- 以上评测结果基于 [OpenCompass 20230706](https://github.com/internLM/OpenCompass/) 获得（部分数据标注`*`代表数据来自原始论文），具体测试细节可参见 [OpenCompass](https://github.com/internLM/OpenCompass/) 中提供的配置文件。
-- 评测数据会因 [OpenCompass](https://github.com/internLM/OpenCompass/) 的版本迭代而存在数值差异，请以 [OpenCompass](https://github.com/internLM/OpenCompass/) 最新版的评测结果为主。
 **局限性：** 尽管在训练过程中我们非常注重模型的安全性，尽力促使模型输出符合伦理和法律要求的文本，但受限于模型大小以及概率生成范式，模型可能会产生各种不符合预期的输出，例如回复内容包含偏见、歧视等有害内容，请勿传播这些内容。由于传播不良信息导致的任何后果，本项目不承担责任。
 ### 通过 Transformers 加载
-通过以下的代码加载 InternLM 7B Chat 模型
 ```python
 import torch
 from transformers import AutoTokenizer, AutoModelForCausalLM
@@ -120,12 +133,12 @@ model = model.eval()
 inputs = tokenizer(["来到美丽的大自然，我们发现"], return_tensors="pt")
 for k,v in inputs.items():
     inputs[k] = v.cuda()
-gen_kwargs = {"max_length": 128, "top_p": 0.8, "temperature": 0.8, "do_sample": True, "repetition_penalty": 1.1}
 output = model.generate(**inputs, **gen_kwargs)
 output = tokenizer.decode(output[0].tolist(), skip_special_tokens=True)
 print(output)
-# 来到美丽的大自然，我们发现各种各样的花千奇百怪。有的颜色鲜艳亮丽,使人感觉生机勃勃；有的是红色的花瓣儿粉嫩嫩的像少女害羞的脸庞一样让人爱不���手．有的小巧玲珑; 还有的花瓣粗大看似枯黄实则暗藏玄机！
-# 不同的花卉有不同的“脾气”,它们都有着属于自己的故事和人生道理.这些鲜花都是大自然中最为原始的物种,每一朵都绽放出别样的美令人陶醉、着迷!
 ```
 ## 开源许可证

 ## Introduction
+The second generation of the InternLM model, InternLM2, includes models at two scales: 7B and 20B. For the convenience of users and researchers, we have open-sourced four versions of each scale of the model, which are:
+- internlm2-base: A high-quality and highly adaptable model base, serving as an excellent starting point for deep domain adaptation.
+- internlm2 (**recommended**): Built upon the internlm2-base, this version has been enhanced in multiple capability directions. It shows outstanding performance in evaluations while maintaining robust general language abilities, making it our recommended choice for most applications.
+- internlm2-sft: Based on the Base model, it undergoes supervised human alignment training.
+- internlm2-chat (**recommended**): Optimized for conversational interaction on top of the internlm2-sft through RLHF, it excels in instruction adherence, empathetic chatting, and tool invocation.
+The base model of InternLM2 has the following technical features:
+- Effective support for ultra-long contexts of up to 200,000 characters: The model nearly perfectly achieves "finding a needle in a haystack" in long inputs of 200,000 characters. It also leads among open-source models in performance on long-text tasks such as LongBench and L-Eval.
+- Comprehensive performance enhancement: Compared to the previous generation model, it shows significant improvements in various capabilities, including reasoning, mathematics, and coding.
+## InternLM2-20B
 ### Performance Evaluation
+We have evaluated InternLM2 on several important benchmarks using the open-source evaluation tool [OpenCompass](https://github.com/open-compass/opencompass). Some of the evaluation results are shown in the table below. You are welcome to visit the [OpenCompass Leaderboard](https://opencompass.org.cn/rank) for more evaluation results.
+| Dataset\Models | InternLM2-7B | InternLM2-Chat-7B | InternLM2-20B | InternLM2-Chat-20B | ChatGPT | GPT-4 |
+| --- | --- | --- | --- | --- | --- | --- |
+| MMLU | 65.8 | 63.7 | 67.7 | 66.5 | 69.1 | 83.0 |
+| AGIEval | 49.9 | 47.2 | 53.0 | 50.3 | 39.9 | 55.1 |
+| BBH | 65.0 | 61.2 | 72.1 | 68.3 | 70.1 | 86.7 |
+| GSM8K | 70.8 | 70.7 | 76.1 | 79.6 | 78.2 | 91.4 |
+| MATH | 20.2 | 23.0 | 25.5 | 31.9 | 28.0 | 45.8 |
+| HumanEval | 43.3 | 59.8 | 48.8 | 67.1 | 73.2 | 74.4 |
+| MBPP(Sanitized) | 51.8 | 51.4 | 63.0 | 65.8 | 78.9 | 79.0 |
+- The evaluation results were obtained from [OpenCompass](https://github.com/open-compass/opencompass) , and evaluation configuration can be found in the configuration files provided by [OpenCompass](https://github.com/open-compass/opencompass).
+- The evaluation data may have numerical differences due to the version iteration of [OpenCompass](https://github.com/open-compass/opencompass), so please refer to the latest evaluation results of [OpenCompass](https://github.com/open-compass/opencompass).
 **Limitations:** Although we have made efforts to ensure the safety of the model during the training process and to encourage the model to generate text that complies with ethical and legal requirements, the model may still produce unexpected outputs due to its size and probabilistic generation paradigm. For example, the generated responses may contain biases, discrimination, or other harmful content. Please do not propagate such content. We are not responsible for any consequences resulting from the dissemination of harmful information.
 ### Import from Transformers
+To load the InternLM2-20B model using Transformers, use the following code:
 ```python
 import torch
 from transformers import AutoTokenizer, AutoModelForCausalLM
 inputs = tokenizer(["A beautiful flower"], return_tensors="pt")
 for k,v in inputs.items():
     inputs[k] = v.cuda()
+gen_kwargs = {"max_length": 128, "top_p": 0.8, "temperature": 0.8, "do_sample": True, "repetition_penalty": 1.0}
 output = model.generate(**inputs, **gen_kwargs)
 output = tokenizer.decode(output[0].tolist(), skip_special_tokens=True)
 print(output)
+# A beautiful flower with a long history of use in Ayurveda and traditional Chinese medicine. Known for its ability to help the body adapt to stress, it is a calming and soothing herb. It is used for its ability to help promote healthy sleep patterns, calm the nervous system and to help the body adapt to stress. It is also used for its ability to help the body deal with the symptoms of anxiety and depression. It is also used for its ability to help the body adapt to stress. It is also used for its ability to help the body adapt to stress. It is also used for its ability to help the
 ```
 ## Open Source License
 The code is licensed under Apache-2.0, while model weights are fully open for academic research and also allow **free** commercial usage. To apply for a commercial license, please fill in the [application form (English)](https://wj.qq.com/s2/12727483/5dba/)/[申请表（中文）](https://wj.qq.com/s2/12725412/f7c1/). For other questions or collaborations, please contact <internlm@pjlab.org.cn>.
 ## 简介
+第二代浦语模型， InternLM2 包含 7B 和 20B 两个量级的模型。为了方便用户使用和研究，每个量级的模型我们总共开源了四个版本的模型，他们分别是
+- internlm2-base: 高质量和具有很强可塑性的模型基座，是模型进行深度领域适配的高质量起点；
+- internlm2（**推荐**）: 在internlm2-base基础上，在多个能力方向进行了强化，在评测中成绩优异，同时保持了很好的通用语言能力，是我们推荐的在大部分应用中考虑选用的优秀基座；
+- internlm2-sft：在Base基础上，进行有监督的人类对齐训练；
+- internlm2-chat（**推荐**）：在internlm2-sft基础上，经过RLHF，面向对话交互进行了优化，具有很好的指令遵循、共情聊天和调用工具等的能力。
+InternLM2 的基础模型具备以下的技术特点
+- 有效支持20万字超长上下文：模型在20万字长输入中几乎完美地实现长文“大海捞针”，而且在 LongBench 和 L-Eval 等长文任务中的表现也达到开源模型中的领先水平。
+- 综合性能全面提升：各能力维度相比上一代模型全面进步，在推理、数学、代码等方面的能力提升显著。
+## InternLM2-20B
 ### 性能评测
+我们使用开源评测工具 [OpenCompass](https://github.com/internLM/OpenCompass/) 对 InternLM2 在几个重要的评测集进行了评测 ，部分评测结果如下表所示，欢迎访问[ OpenCompass 榜单 ](https://opencompass.org.cn/rank)获取更多的评测结果。
+| 评测集 | InternLM2-7B | InternLM2-Chat-7B | InternLM2-20B | InternLM2-Chat-20B | ChatGPT | GPT-4 |
+| --- | --- | --- | --- | --- | --- | --- |
+| MMLU | 65.8 | 63.7 | 67.7 | 66.5 | 69.1 | 83.0 |
+| AGIEval | 49.9 | 47.2 | 53.0 | 50.3 | 39.9 | 55.1 |
+| BBH | 65.0 | 61.2 | 72.1 | 68.3 | 70.1 | 86.7 |
+| GSM8K | 70.8 | 70.7 | 76.1 | 79.6 | 78.2 | 91.4 |
+| MATH | 20.2 | 23.0 | 25.5 | 31.9 | 28.0 | 45.8 |
+| HumanEval | 43.3 | 59.8 | 48.8 | 67.1 | 73.2 | 74.4 |
+| MBPP(Sanitized) | 51.8 | 51.4 | 63.0 | 65.8 | 78.9 | 79.0 |
+- 以上评测结果基于 [OpenCompass](https://github.com/open-compass/opencompass) 获得（部分数据标注`*`代表数据来自原始论文），具体测试细节可参见 [OpenCompass](https://github.com/open-compass/opencompass) 中提供的配置文件。
+- 评测数据会因 [OpenCompass](https://github.com/open-compass/opencompass) 的版本迭代而存在数值差异，请以 [OpenCompass](https://github.com/open-compass/opencompass) 最新版的评测结果为主。
 **局限性：** 尽管在训练过程中我们非常注重模型的安全性，尽力促使模型输出符合伦理和法律要求的文本，但受限于模型大小以及概率生成范式，模型可能会产生各种不符合预期的输出，例如回复内容包含偏见、歧视等有害内容，请勿传播这些内容。由于传播不良信息导致的任何后果，本项目不承担责任。
 ### 通过 Transformers 加载
+通过以下的代码加载 InternLM2-20B 模型进行文本续写
 ```python
 import torch
 from transformers import AutoTokenizer, AutoModelForCausalLM
 inputs = tokenizer(["来到美丽的大自然，我们发现"], return_tensors="pt")
 for k,v in inputs.items():
     inputs[k] = v.cuda()
+gen_kwargs = {"max_length": 128, "top_p": 0.8, "temperature": 0.8, "do_sample": True, "repetition_penalty": 1.0}
 output = model.generate(**inputs, **gen_kwargs)
 output = tokenizer.decode(output[0].tolist(), skip_special_tokens=True)
 print(output)
+# 来到美丽的大自然，我们欣赏着大自然的美丽风景，感受着大自然的气息。
+# 今天，我来到了美丽的龙湾公园，这里风景秀丽，山清水秀，鸟语花香。一走进公园，我就被眼前的景象惊呆了：绿油油的草坪上，五颜六色的花朵竞相开放，散发出阵阵清香。微风吹来，花儿随风摆动，好像在向我们点头微笑。远处，巍峨的大山连绵起伏，好像一条巨龙在空中飞舞。山下，一条清澈的小河静静地流淌着，河里的鱼儿自由自在地
 ```
 ## 开源许可证