upload

Files changed (10) hide show

README.md +135 -2
config.json +48 -0
configuration_telechat.py +210 -0
generation_config.json +14 -0
modeling_telechat.py +1105 -0
pytorch_model.bin.index.json +458 -0
special_tokens_map.json +30 -0
tokenization_telechat.py +403 -0
tokenizer.model +3 -0
tokenizer_config.json +54 -0

README.md CHANGED Viewed

@@ -1,3 +1,136 @@
 ---
-license: apache-2.0
----

+<div align="center">
+<h1>
+  星辰大模型(52B)
+</h1>
+</div>
+# 目录
+- [模型介绍](#模型介绍)
+- [数据开源](#数据开源)
+- [效果评测](#效果评测)
+- [模型推理和部署](#模型推理和部署)
+- [模型微调](#模型微调)
+- [声明、协议、引用](#声明协议引用)
+# 最新动态
+- 5.16 开源52B版本chat模型
+# 模型介绍
+### 星辰大模型（52B）
+- 星辰大模型（52B）是一款开源多语言大模型，其模型基座使用更优数据配比，采用课程学习方式，训练了2万亿tokens的中英文高质量数据。
+- 我们开源了使用星辰语义大模型52B基座微调的对话模型，以及基于Deepspeed的微调代码和huggingface推理代码。
+- 星辰大模型（52B）在模型评测中取得了领先的效果，在榜单评测上超过LLaMA-2-70B-Chat，与Qwen-72B-chat可比；通用对话性能已经超过GPT-3.5-Turbo。
+### 模型结构
+我们采用标准的 `Decoder-only` 结构设计了 **TeleChat** 模型，并在模型维度做了如下的一些改进：
+- **位置编码**：我们使用 [Rotary Embedding](https://arxiv.org/pdf/2104.09864.pdf) 的位置编码方法，该方法将相对位置信息依赖集成到 self-attention 中，并且具有较好的位置外推性。Rotary Embedding还可以较好地与Flash-Attention v2 配合使用，将模型的训练速度提升约20%。
+- **激活函数**：我们使用 [SwiGLU](https://arxiv.org/pdf/2002.05202.pdf) 激活函数来替代GELU激活函数。
+- **层标准化**: 基于 [RMSNorm](https://arxiv.org/abs/1910.07467) 的 Pre-Normalization。
+- **词嵌入层与输出层解耦**：我们将星辰52B的词嵌入层和输出lm head层参数分开，有助于增强训练稳定性和收敛性。
+|         | layer_num | hidden_size | ffn_hidden_size | head_num | tie_word_embeddings |
+| ------- | --------- | ----------- | --------------- | -------- | ------------------- |
+| 星辰52B | 64        | 8192        | 21824           | 64       | 否                  |
 ---
+我们开源的星辰52B模型：
+- 支持deepspeed微调，开源了基于deepspeed的训练代码，支持Zero并行显存优化，同时集成了FlashAttention2
+- 多轮能力支持。开源了多轮数据构建方式，针对多轮模型训练集成了针对多轮的mask loss训练方式，更好的聚焦多轮答案，提升问答效果。
+# 效果评测
+星辰52B模型相比同规模模型在评测效果方面也有较好的表现，我们的评测集涵盖了包括MMLU、AGIEval、CMMLU、 GSM8K、MATH、HumanEval 等数据集，评测能力包括了自然语言理解、知识、数学计算和推理、代码生成等
+## 评测集介绍
+### 通用能力
+- MMLU 数据集是一个全面的英文评测数据集，涵盖了 57 个学科，包括人文学科、社会科学、自然科学、初等数学、美国历史、计算机科学、法律等等。
+- CMMLU 数据集同样是一个全面的中文评估测试集，涵盖了从基础学科到高级专业水平的67个主题。
+- AGIEval 数据集是一个专门为评估基础模型在难度较高的标准化考试（如大学入学考试、法学院入学考试、数学竞赛和律师资格考试）的语境中而设计的基准测试，包括中文试题和英文试题。
+### 推理和代码能力
+- GSM8K 数据集包含了8.5K高质量的小学数学题，能够评估语言模型在数学推理能力上的表现，我们利用[官方](https://github.com/openai/grade-school-math)的评测方案在test集上进行了4-shot测试。
+- MATH 数据集包含了12.5K具有挑战性的高中数学竞赛题，难度较大，对语言模型的推理能力要求较高，基于[官方](https://github.com/hendrycks/math)的评测方案，我们在test集上进行了4-shot测试。
+- HumanEval 数据集是一个由openai提供的代码能力测试数据集，它由 164 个编程问题组成，要求根据给定的问题和代码模板，生成正确的代码片段，我们利用[官方](https://github.com/openai/human-eval)评测方案在test集上进行了zero-shot测试。
+## 评测结果如下
+| Model            |   MMLU   |   CMMLU   |  AGIEval  |  GSM8K   |   MATH   | HumanEval |   BBH    | HellaSwag |
+| :--------------- | :------: | :-------: | :-------: | :------: | :------: | :-------: | :------: | :-------: |
+|                  |  5-shot  |  5-shot   | zero-shot |  4-shot  |  4-shot  | zero-shot |  3-shot  | zero-shot |
+| LLaMA-2-70B-Chat |   63.8   |   43.3    |   37.9    |   59.3   |   10.4   |   32.3    |   60.8   |   80.6    |
+| Qwen-72B-chat    |    74    |   81.4    |   58.5    |   67.4   |   31.8   |   49.4    |    68    |   84.7    |
+| 星辰7B-chat      |   60.5   |   64.3    |   46.8    |   36.7   |   10.3   |   20.1    |   19.5   |   36.7    |
+| 星辰12B-chat     |   73.3   |   74.2    |   51.7    |   57.2   |   16.0   |   22.0    |   52.2   |   71.5    |
+| **星辰52B-chat** | **76.6** | **73.79** | **61.1**  | **63.5** | **13.5** | **36.6**  | **60.3** | **86.3**  |
+说明：榜单均基于[OpenCompass](https://github.com/open-compass/OpenCompass/)平台提供的评测方法进行评估，而对于对比模型，我们同时参考了官方汇报结果和OpenCompass结果。
+### 对话能力评测
+为了评价模型的对话能力，研发团队建立了包含2500+单轮、多轮对话交互的内部评测系统，涵盖闲聊问答、专业知识、翻译、逻辑思维、长文写作、幻觉测试、安全测试、角色扮演、任务执行、数学能力等多个维度，并使用Judge模型基于详细的评价指标文档进行自动打分。在当前评测数据上，星辰52B模型的综合平均得分为83.8，高于GPT-3.5-Turbo的82.3。这一结果表明，星辰52B模型能较好地支持下游任务应用。
+# 模型推理和部署
+### 模型推理
+当前模型支持fp16精度推理，适配4卡40G A100进行推理。具体推理操作请参考`infer.py`文件，该文件中有单轮和多轮的推理示例。推理结果示例见`infer_result.txt`。
+**模型推理方法示范**
+```python
+import os
+import torch
+from transformers import AutoModelForCausalLM, AutoTokenizer
+from transformers import GenerationConfig
+PATH = "/path/to/TeleChat-52B-chat"
+tokenizer = AutoTokenizer.from_pretrained(PATH, use_fast=False, trust_remote_code=True)
+model = AutoModelForCausalLM.from_pretrained(PATH,torch_dtype=torch.bfloat16,device_map='auto',trust_remote_code=True)
+question = "你作为一名气候保护协会的会员，你准备写一篇全球气候变化的新闻报告，要求体现出全球气候变化以前与现在情况的对比，字数要求1000字。"
+generate_config = GenerationConfig.from_pretrained(PATH)
+answer = model.chat(tokenizer,question, history_input_list = [], history_output_list = [],generation_config = generate_config)
+print("machine:",answer)
+```
+# 声明、协议、引用
+### 声明
+我们在此声明，不要使用TeleChat模型及其衍生模型进行任何危害国家社会安全或违法的活动。同时，我们也要求使用者不要将TeleChat模型用于没有安全审查和备案的互联网服务。我们希望所有使用者遵守上述原则，确保科技发展在合法合规的环境下进行。
+我们已经尽我们所能，来确保模型训练过程中使用的数据的合规性。然而，尽管我们已经做出了巨大的努力，但由于模型和数据的复杂性，仍有可能存在一些无法预见的问题。因此，如果由于使用TeleChat开源模型而导致的任何问题，包括但不限于数据安全问题、公共舆论风险，或模型被误导、滥用、传播或不当利用所带来的任何风险和问题，我们将不承担任何责任。
+### 协议
+社区使用 TeleChat 模型需要遵循《[TeleChat模型社区许可协议](./TeleChat模型社区许可协议.pdf)》。TeleChat模型支持商业用途，如果您计划将 TeleChat 模型或其衍生品用于商业目的，您需要通过以下联系邮箱 tele_ai@chinatelecom.cn，提交《TeleChat模型社区许可协议》要求的申请材料。审核通过后，将特此授予您一个非排他性、全球性、不可转让、不可再许可、可撤销的商用版权许可。
+### 引用
+如需引用我们的工作，请使用如下 reference:
+```
+@misc{wang2024telechat,
+      title={TeleChat Technical Report},
+      author={Zihan Wang and Xinzhang Liu and Shixuan Liu and Yitong Yao and Yuyao Huang and Zhongjiang He and Xuelong Li and Yongxiang Li and Zhonghao Che and Zhaoxi Zhang and Yan Wang and Xin Wang and Luwen Pu and Huihan Xu and Ruiyu Fang and Yu Zhao and Jie Zhang and Xiaomeng Huang and Zhilong Lu and Jiaxin Peng and Wenjun Zheng and Shiquan Wang and Bingkai Yang and Xuewei he and Zhuoru Jiang and Qiyi Xie and Yanhan Zhang and Zhongqiu Li and Lingling Shi and Weiwei Fu and Yin Zhang and Zilu Huang and Sishi Xiong and Yuxiang Zhang and Chao Wang and Shuangyong Song},
+      year={2024},
+      eprint={2401.03804},
+      archivePrefix={arXiv},
+      primaryClass={cs.CL}
+}
+@misc{li2024teleflm,
+      title={Tele-FLM Technical Report},
+      author={Xiang Li and Yiqun Yao and Xin Jiang and Xuezhi Fang and Chao Wang and Xinzhang Liu and Zihan Wang and Yu Zhao and Xin Wang and Yuyao Huang and Shuangyong Song and Yongxiang Li and Zheng Zhang and Bo Zhao and Aixin Sun and Yequan Wang and Zhongjiang He and Zhongyuan Wang and Xuelong Li and Tiejun Huang},
+      year={2024},
+      eprint={2404.16645},
+      archivePrefix={arXiv},
+      primaryClass={cs.CL}
+}
+```

config.json ADDED Viewed

	@@ -0,0 +1,48 @@

+{
+  "activation_function": "silu",
+  "add_bias_linear": false,
+  "attn_pdrop": 0.0,
+  "auto_map": {
+    "AutoConfig": "configuration_telechat.TELECHATConfig",
+    "AutoModel": "modeling_telechat.TELECHAT",
+    "AutoModelForCausalLM": "modeling_telechat.TELECHAT"
+  },
+  "bos_token_id": 1,
+  "embd_pdrop": 0.0,
+  "enable_flash_attn": true,
+  "eos_token_id": 2,
+  "initializer_range": 0.02,
+  "input_mult": 1.0,
+  "layer_norm_epsilon": 1e-05,
+  "model_type": "telechat",
+  "mup_base_width": 256,
+  "mup_scale_factor": 32.0,
+  "n_embd": 8192,
+  "n_head": 64,
+  "n_inner": 21824,
+  "n_layer": 64,
+  "n_positions": 4096,
+  "output_mult": 1.0,
+  "pad_token_id": 3,
+  "relative_encoding": "rotary",
+  "reorder_and_upcast_attn": true,
+  "resid_pdrop": 0.0,
+  "rotary_theta": 10000,
+  "rotary_use_xpos": false,
+  "rotary_xpos_scale_base": 512,
+  "scale_attn_by_inverse_layer_idx": true,
+  "scale_attn_weights": true,
+  "summary_activation": null,
+  "summary_first_dropout": 0.1,
+  "summary_proj_to_labels": true,
+  "summary_type": "cls_index",
+  "summary_use_proj": true,
+  "tie_word_embeddings": false,
+  "tokenizer_class": "TELECHATTokenizer",
+  "transformers_version": "4.34.1",
+  "unk_token_id": 0,
+  "use_RMSNorm": true,
+  "use_cache": true,
+  "use_mup": true,
+  "vocab_size": 80896
+}

configuration_telechat.py ADDED Viewed

	@@ -0,0 +1,210 @@

+# coding=utf-8
+# Copyright 2018 The OpenAI Team Authors and HuggingFace Inc. team.
+# Copyright (c) 2018, NVIDIA CORPORATION.  All rights reserved.
+#
+# Licensed under the Apache License, Version 2.0 (the "License");
+# you may not use this file except in compliance with the License.
+# You may obtain a copy of the License at
+#
+#     http://www.apache.org/licenses/LICENSE-2.0
+#
+# Unless required by applicable law or agreed to in writing, software
+# distributed under the License is distributed on an "AS IS" BASIS,
+# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
+# See the License for the specific language governing permissions and
+# limitations under the License.
+""" TELECHAT configuration"""
+from transformers.configuration_utils import PretrainedConfig
+from transformers.utils import logging
+logger = logging.get_logger(__name__)
+TELECHAT_PRETRAINED_CONFIG_ARCHIVE_MAP = {
+}
+class TELECHATConfig(PretrainedConfig):
+    """
+    xxxxxx
+    Configuration objects inherit from [`PretrainedConfig`] and can be used to control the model outputs. Read the
+    documentation from [`PretrainedConfig`] for more information.
+    Args:
+        vocab_size (`int`, *optional*, defaults to 50257):
+            Vocabulary size of the GPT-2 model. Defines the number of different tokens that can be represented by the
+            `inputs_ids` passed when calling [`GPT2Model`] or [`TFGPT2Model`].
+        n_positions (`int`, *optional*, defaults to 1024):
+            The maximum sequence length that this model might ever be used with. Typically set this to something large
+            just in case (e.g., 512 or 1024 or 2048).
+        n_embd (`int`, *optional*, defaults to 768):
+            Dimensionality of the embeddings and hidden states.
+        n_layer (`int`, *optional*, defaults to 12):
+            Number of hidden layers in the Transformer encoder.
+        n_head (`int`, *optional*, defaults to 12):
+            Number of attention heads for each attention layer in the Transformer encoder.
+        n_inner (`int`, *optional*, defaults to None):
+            Dimensionality of the inner feed-forward layers. `None` will set it to 4 times n_embd
+        activation_function (`str`, *optional*, defaults to `"gelu"`):
+            Activation function, to be selected in the list `["relu", "silu", "gelu", "tanh", "gelu_new"]`.
+        resid_pdrop (`float`, *optional*, defaults to 0.1):
+            The dropout probability for all fully connected layers in the embeddings, encoder, and pooler.
+        embd_pdrop (`int`, *optional*, defaults to 0.1):
+            The dropout ratio for the embeddings.
+        attn_pdrop (`float`, *optional*, defaults to 0.1):
+            The dropout ratio for the attention.
+        layer_norm_epsilon (`float`, *optional*, defaults to 1e-5):
+            The epsilon to use in the layer normalization layers.
+        initializer_range (`float`, *optional*, defaults to 0.02):
+            The standard deviation of the truncated_normal_initializer for initializing all weight matrices.
+        summary_type (`string`, *optional*, defaults to `"cls_index"`):
+            Argument used when doing sequence summary, used in the models [`GPT2DoubleHeadsModel`] and
+            [`TFGPT2DoubleHeadsModel`].
+            Has to be one of the following options:
+                - `"last"`: Take the last token hidden state (like XLNet).
+                - `"first"`: Take the first token hidden state (like BERT).
+                - `"mean"`: Take the mean of all tokens hidden states.
+                - `"cls_index"`: Supply a Tensor of classification token position (like GPT/GPT-2).
+                - `"attn"`: Not implemented now, use multi-head attention.
+        summary_use_proj (`bool`, *optional*, defaults to `True`):
+            Argument used when doing sequence summary, used in the models [`GPT2DoubleHeadsModel`] and
+            [`TFGPT2DoubleHeadsModel`].
+            Whether or not to add a projection after the vector extraction.
+        summary_activation (`str`, *optional*):
+            Argument used when doing sequence summary. Used in for the multiple choice head in
+            [`GPT2DoubleHeadsModel`].
+            Pass `"tanh"` for a tanh activation to the output, any other value will result in no activation.
+        summary_proj_to_labels (`bool`, *optional*, defaults to `True`):
+            Argument used when doing sequence summary, used in the models [`GPT2DoubleHeadsModel`] and
+            [`TFGPT2DoubleHeadsModel`].
+            Whether the projection outputs should have `config.num_labels` or `config.hidden_size` classes.
+        summary_first_dropout (`float`, *optional*, defaults to 0.1):
+            Argument used when doing sequence summary, used in the models [`GPT2DoubleHeadsModel`] and
+            [`TFGPT2DoubleHeadsModel`].
+            The dropout ratio to be used after the projection and activation.
+        scale_attn_weights (`bool`, *optional*, defaults to `True`):
+            Scale attention weights by dividing by sqrt(hidden_size)..
+        use_cache (`bool`, *optional*, defaults to `True`):
+            Whether or not the model should return the last key/values attentions (not used by all models).
+        scale_attn_by_inverse_layer_idx (`bool`, *optional*, defaults to `False`):
+            Whether to additionally scale attention weights by `1 / layer_idx + 1`.
+        reorder_and_upcast_attn (`bool`, *optional*, defaults to `False`):
+            Whether to scale keys (K) prior to computing attention (dot-product) and upcast attention
+            dot-product/softmax to float() when training with mixed precision.
+    Example:
+    ```python
+    >>> from transformers import GPT2Config, GPT2Model
+    >>> # Initializing a GPT2 configuration
+    >>> configuration = GPT2Config()
+    >>> # Initializing a model (with random weights) from the configuration
+    >>> model = GPT2Model(configuration)
+    >>> # Accessing the model configuration
+    >>> configuration = model.config
+    ```"""
+    model_type = "telechat"
+    keys_to_ignore_at_inference = ["past_key_values"]
+    attribute_map = {
+        "hidden_size": "n_embd",
+        "max_position_embeddings": "n_positions",
+        "num_attention_heads": "n_head",
+        "num_hidden_layers": "n_layer",
+    }
+    def __init__(
+            self,
+            vocab_size=80000,
+            n_positions=1024,
+            n_embd=768,
+            n_layer=12,
+            n_head=12,
+            n_inner=None,
+            activation_function="gelu_new",
+            resid_pdrop=0.1,
+            embd_pdrop=0.1,
+            attn_pdrop=0.1,
+            layer_norm_epsilon=1e-5,
+            initializer_range=0.02,
+            summary_type="cls_index",
+            summary_use_proj=True,
+            summary_activation=None,
+            summary_proj_to_labels=True,
+            summary_first_dropout=0.1,
+            scale_attn_weights=True,
+            use_cache=True,
+            bos_token_id=None,
+            eos_token_id=None,
+            sep_token_id=None,
+            pad_token_id=None,
+            unk_token_id=None,
+            scale_attn_by_inverse_layer_idx=False,
+            reorder_and_upcast_attn=False,
+            relative_encoding=None,
+            rotary_theta=10000,
+            rotary_use_xpos=True,
+            rotary_xpos_scale_base=512,
+            use_mup=False,
+            mup_scale_factor=1.0,
+            output_mult=1.0,
+            input_mult=1.0,
+            mup_base_width=256,
+            enable_flash_attn=True,
+            use_RMSNorm=False,
+            add_bias_linear=True,
+            **kwargs,
+    ):
+        self.vocab_size = vocab_size
+        self.n_positions = n_positions
+        self.n_embd = n_embd
+        self.n_layer = n_layer
+        self.n_head = n_head
+        self.n_inner = n_inner
+        self.activation_function = activation_function
+        self.resid_pdrop = resid_pdrop
+        self.embd_pdrop = embd_pdrop
+        self.attn_pdrop = attn_pdrop
+        self.layer_norm_epsilon = layer_norm_epsilon
+        self.initializer_range = initializer_range
+        self.summary_type = summary_type
+        self.summary_use_proj = summary_use_proj
+        self.summary_activation = summary_activation
+        self.summary_first_dropout = summary_first_dropout
+        self.summary_proj_to_labels = summary_proj_to_labels
+        self.scale_attn_weights = scale_attn_weights
+        self.use_cache = use_cache
+        self.scale_attn_by_inverse_layer_idx = scale_attn_by_inverse_layer_idx
+        self.reorder_and_upcast_attn = reorder_and_upcast_attn
+        self.relative_encoding = relative_encoding
+        self.use_RMSNorm = use_RMSNorm
+        self.add_bias_linear = add_bias_linear
+        # for rotary
+        self.rotary_theta = rotary_theta
+        self.rotary_use_xpos = rotary_use_xpos
+        self.rotary_xpos_scale_base = rotary_xpos_scale_base
+        # for mup
+        self.use_mup = use_mup
+        self.mup_scale_factor = mup_scale_factor
+        self.output_mult = output_mult
+        self.input_mult = input_mult
+        self.mup_base_width = mup_base_width
+        self.bos_token_id = bos_token_id
+        self.eos_token_id = eos_token_id
+        self.unk_token_id = unk_token_id
+        self.sep_token_id = sep_token_id
+        self.pad_token_id = pad_token_id
+        self.enable_flash_attn = enable_flash_attn
+        self.architectures = ["TELECHAT"]
+        self.auto_map = {
+            "AutoConfig": "configuration_telechat.TELECHATConfig",
+            "AutoModel": "modeling_telechat.TELECHAT",
+            "AutoModelForCausalLM": "modeling_telechat.TELECHAT"
+        }
+        super().__init__(bos_token_id=bos_token_id, eos_token_id=eos_token_id, sep_token_id = sep_token_id, pad_token_id = pad_token_id, **kwargs)

generation_config.json ADDED Viewed

	@@ -0,0 +1,14 @@

+{
+  "max_length": 2048,
+  "do_sample": false,
+  "use_cache": true,
+  "temperature": 0.3,
+  "top_k": 5,
+  "top_p": 0.85,
+  "repetition_penalty": 1.02,
+  "pad_token_id": 3,
+  "bos_token_id": 1,
+  "eos_token_id": 2,
+  "user_token_id": 20,
+  "bot_token_id": 21
+}

modeling_telechat.py ADDED Viewed

	@@ -0,0 +1,1105 @@

+# coding=utf-8
+# Copyright 2018 The OpenAI Team Authors and HuggingFace Inc. team.
+# Copyright (c) 2018, NVIDIA CORPORATION.  All rights reserved.
+#
+# This code is based on OpenAI's GPT-2 library. It has been modified from its
+# original forms to accommodate architectural differences compared to GPT-2.
+#
+# Licensed under the Apache License, Version 2.0 (the "License");
+# you may not use this file except in compliance with the License.
+# You may obtain a copy of the License at
+#
+#     http://www.apache.org/licenses/LICENSE-2.0
+#
+# Unless required by applicable law or agreed to in writing, software
+# distributed under the License is distributed on an "AS IS" BASIS,
+# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
+# See the License for the specific language governing permissions and
+# limitations under the License.
+"""PyTorch TELECHAT model."""
+from typing import Optional, Tuple, Union
+import math
+import torch
+from einops import rearrange
+from torch import einsum, nn
+from torch.cuda.amp import autocast
+import torch.nn.functional as F
+from transformers.activations import ACT2FN
+from transformers.modeling_outputs import (
+    BaseModelOutputWithPastAndCrossAttentions,
+    CausalLMOutputWithCrossAttentions,
+    SequenceClassifierOutputWithPast,
+)
+from transformers.modeling_utils import PreTrainedModel
+from transformers.pytorch_utils import find_pruneable_heads_and_indices, prune_conv1d_layer
+from transformers.utils import logging
+from transformers.utils.model_parallel_utils import assert_device_map, get_device_map
+try:
+    from flash_attn.flash_attn_interface import flash_attn_unpadded_func # flashattn1
+    print("# FLASH ATTENTION 1 DETECTED #")
+except ImportError:
+    try:
+        from flash_attn.flash_attn_interface import flash_attn_varlen_func as flash_attn_unpadded_func # flashattn2
+        print("# FLASH ATTENTION 2 DETECTED #")
+    except ImportError:
+        print("# NO FLASH ATTENTION DETECTED #")
+        flash_attn_unpadded_func = None
+from .configuration_telechat import TELECHATConfig
+def debug_print_tensor(t, name, title='', show_dim=10):
+    # return
+    prefix = f'{title} -> '
+    if isinstance(t, torch.Tensor):
+        if len(t.shape) == 1:
+            output = f"{name}[{t.shape}]: {t[:show_dim]}"
+        elif len(t.shape) == 2:
+            output = f"{name}[{t.shape}]: {t[-1, :show_dim]}"
+        elif len(t.shape) == 3:
+            output = f" {name}[{t.shape}]: {t[-1, -1, :show_dim]}"
+        elif len(t.shape) == 4:
+            output = f"{name}[{t.shape}]: {t[-1, -1, -1, :show_dim]}"
+        else:
+            output = f"{name}[{t.shape}]"
+    elif isinstance(t, list):
+        output = f"{name} [{len(t)}]: {t[:show_dim]}"
+    else:
+        output = f"{name} 未知类型: {type(t)}"
+    print(prefix + output)
+class Conv1D(nn.Module):
+    def __init__(self, nf, nx, bias=True):
+        super().__init__()
+        self.nf = nf
+        self.weight = nn.Parameter(torch.empty(nx, nf))
+        self.bias = None
+        if bias:
+            self.bias = nn.Parameter(torch.zeros(nf))
+        nn.init.normal_(self.weight, std=0.02)
+    def forward(self, x):
+        if self.bias is not None:
+            return torch.matmul(x, self.weight) + self.bias
+        else:
+            return torch.matmul(x, self.weight)
+class RMSNorm(nn.Module):
+    def __init__(self, hidden_size, eps=1e-5):
+        super().__init__()
+        self.weight = nn.Parameter(torch.ones(hidden_size))
+        self.eps = eps
+    def forward(self, hidden_states):
+        input_dtype = hidden_states.dtype
+        hidden_states = hidden_states.to(torch.float32)
+        variance = hidden_states.pow(2).mean(-1, keepdim=True)
+        hidden_states = hidden_states * torch.rsqrt(variance + self.eps)
+        return self.weight * hidden_states.to(input_dtype)
+logger = logging.get_logger(__name__)
+def exists(v):
+    return v is not None
+class RotaryEmbedding(nn.Module):
+    def __init__(self, dim, use_xpos=False, xpos_scale_base=512, theta=10000):
+        super().__init__()
+        inv_freq = 1.0 / (theta ** (torch.arange(0, dim, 2).float() / dim))
+        self.register_buffer('inv_freq', inv_freq)
+        self.cache = dict()
+        self.cache_scale = dict()
+        self.use_xpos = use_xpos
+        if not use_xpos:
+            self.register_buffer('scale', None)
+            return
+        scale = (torch.arange(0, dim, 2) + 0.4 * dim) / (1.4 * dim)
+        self.register_buffer('scale', scale)
+        self.scale_base = xpos_scale_base
+    def forward(self, seq, cache_key=None):
+        if cache_key is not None and cache_key in self.cache:
+            return self.cache[cache_key]
+        inv_freq = self.inv_freq.to(device=seq.device)
+        freqs = einsum('i , j -> i j', seq, inv_freq)
+        # first part even vector components, second part odd vector components,
+        #  2 * dim in dimension size
+        scale = torch.cat((freqs, freqs), dim=-1)
+        if exists(cache_key):
+            self.cache[cache_key] = scale
+        return scale
+    def rotate_queries_and_keys(self, q, k, seq_dim=-2):
+        """
+        use this only when xpos is activated.
+        """
+        assert self.use_xpos and q.device == k.device
+        device, seq_len_k, seq_len_q = k.device, k.shape[seq_dim], q.shape[seq_dim]
+        pos_seq_k = torch.arange(seq_len_k, device=device, dtype=torch.float32)
+        pos_seq_q = torch.arange(seq_len_k - seq_len_q, seq_len_k, device=device, dtype=torch.float32)
+        freqs_k = self.forward(pos_seq_k, cache_key=f"{0}:{seq_len_k}")
+        freqs_q = self.forward(pos_seq_q, cache_key=f"{seq_len_k - seq_len_q}:{seq_len_k}")
+        scale_k = self.get_scale(pos_seq_k)
+        scale_q = self.get_scale(pos_seq_q, offset=seq_len_k - seq_len_q)  # 这里的offset是Q相对于K的offset
+        rotated_q = apply_rotary_emb(freqs_q, q, scale=scale_q)
+        rotated_k = apply_rotary_emb(freqs_k, k, scale=scale_k ** -1)
+        return rotated_q, rotated_k
+    def rotate_queries_or_keys(self, t, seq_dim=-2, offset=0):
+        """
+        use this only when xpos is NOT activated.
+        """
+        # t's shape e.g.  -> (batchsize, headnum, seqlen, dimofhead)
+        assert not self.use_xpos, 'you must use `.rotate_queries_and_keys` method instead and pass in both queries and keys, for length extrapolatable rotary embeddings'
+        device, seq_len = t.device, t.shape[seq_dim]
+        pos_seq_t = torch.arange(offset, offset + seq_len, device=device, dtype=torch.float32)
+        freqs = self.forward(pos_seq_t, cache_key=f"{offset}:{offset+seq_len}")
+        # freqs   seqlen  x  dim
+        return apply_rotary_emb(freqs, t)
+    def get_scale(self, t, cache_key=None, offset=0, ):
+        assert self.use_xpos, 'This function is only useful for xpos.'
+        if exists(cache_key) and cache_key in self.cache_scale:
+            return self.cache_scale[cache_key]
+        if callable(t):
+            t = t()
+        length = len(t)
+        min_pos = -(length + offset) // 2
+        max_pos = length + offset + min_pos
+        power = torch.arange(min_pos, max_pos, 1).to(device=self.scale.device) / self.scale_base
+        scale = self.scale ** rearrange(power, 'n -> n 1')
+        scale = scale[-length:, :]
+        scale = torch.cat((scale, scale), dim=-1)
+        if exists(cache_key):
+            self.cache_scale[cache_key] = scale
+        return scale
+def rotate_half(x):
+    """
+    change sign so the last dimension becomes [-odd, +even]
+    """
+    x1, x2 = torch.chunk(x, 2, dim=-1)
+    return torch.cat((-x2, x1), dim=-1)
+def apply_rotary_emb(freqs, t, start_index=0, scale=1.):
+    """
+    freq: seqlen  x  dim
+       t: [batchsize  *  headnum  ,  seqlen  , dim (dim_of_head actually)]
+    """
+    dtype_t = t.dtype
+    freqs = freqs.to(device=t.device)
+    if isinstance(scale, torch.Tensor):
+        scale = scale.to(device=t.device)
+    rot_dim = freqs.shape[-1]
+    end_index = start_index + rot_dim
+    t_left, t, t_right = t[..., :start_index], t[..., start_index:end_index], t[..., end_index:]
+    t = (t * freqs.cos() + rotate_half(t) * freqs.sin()) * scale
+    rotated = torch.cat((t_left, t, t_right), dim=-1)
+    rotated = rotated.to(dtype=dtype_t)
+    return rotated
+class TELECHATAttention(nn.Module):
+    def __init__(self, config, layer_idx=None):
+        super().__init__()
+        max_positions = config.max_position_embeddings
+        self.register_buffer(
+            "bias",
+            torch.tril(torch.ones((max_positions, max_positions), dtype=torch.bool)).view(
+                1, 1, max_positions, max_positions
+            ),
+        )
+        self.register_buffer("masked_bias", torch.tensor(-1e4))
+        self.embed_dim = config.hidden_size
+        self.num_heads = config.num_attention_heads
+        self.head_dim = self.embed_dim // self.num_heads
+        self.split_size = self.embed_dim
+        if self.head_dim * self.num_heads != self.embed_dim:
+            raise ValueError(
+                f"`embed_dim` must be divisible by num_heads (got `embed_dim`: {self.embed_dim} and `num_heads`:"
+                f" {self.num_heads})."
+            )
+        self.scale_attn_weights = config.scale_attn_weights
+        # Layer-wise attention scaling, reordering, and upcasting
+        self.scale_attn_by_inverse_layer_idx = config.scale_attn_by_inverse_layer_idx
+        # for alignment with megatron-lm in softmax scale
+        self.layer_idx = max(1, layer_idx)
+        self.reorder_and_upcast_attn = config.reorder_and_upcast_attn
+        self.relative_encoding = config.relative_encoding
+        self.rotary_use_xpos = config.rotary_use_xpos
+        self.use_mup = config.use_mup
+        self.c_attn = Conv1D(3 * self.embed_dim, self.embed_dim, bias=config.add_bias_linear)
+        self.c_proj = Conv1D(self.embed_dim, self.embed_dim, bias=config.add_bias_linear)
+        self.attn_dropout = nn.Dropout(config.attn_pdrop)
+        self.resid_dropout = nn.Dropout(config.resid_pdrop)
+        self.pruned_heads = set()
+        self.use_flash_attn = False
+    def set_max_positions(self, max_positions, device='cuda'):
+        self.max_positions = max_positions
+        self.register_buffer(
+            "bias",
+            torch.tril(torch.ones((self.max_positions, self.max_positions), dtype=torch.bool)).view(
+                1, 1, self.max_positions, self.max_positions
+            ).to(device=device)
+        )
+    def prune_heads(self, heads):
+        if len(heads) == 0:
+            return
+        heads, index = find_pruneable_heads_and_indices(heads, self.num_heads, self.head_dim, self.pruned_heads)
+        index_attn = torch.cat([index, index + self.split_size, index + (2 * self.split_size)])
+        # Prune conv1d layers
+        self.c_attn = prune_conv1d_layer(self.c_attn, index_attn, dim=1)
+        self.c_proj = prune_conv1d_layer(self.c_proj, index, dim=0)
+        # Update hyper params
+        self.split_size = (self.split_size // self.num_heads) * (self.num_heads - len(heads))
+        self.num_heads = self.num_heads - len(heads)
+        self.pruned_heads = self.pruned_heads.union(heads)
+    def _attn(self, query, key, value, attention_mask=None, head_mask=None):
+        # (batch, head, seq_length, head_features)
+        # batch_size, head_num, k_seq_len(q_seq_len), head_features
+        batch_size, head_num, k_seq_len, head_features = key.shape
+        _, _, q_seq_len, _ = query.shape
+        if self.use_flash_attn:
+            # print("*")
+            # attn_output = torch.nn.functional._scaled_dot_product_attention(query, key, value, is_causal=True)
+            # attn_weights = None
+            # return attn_output, attn_weights
+            batch_size, seqlen_q = query.shape[0], query.shape[2]
+            seqlen_k = key.shape[2]
+            query, key, value = [rearrange(x, 'b h s ... -> (b s) h ...') for x in [query, key, value]]
+            cu_seqlens_q = torch.arange(0, (batch_size + 1) * seqlen_q, step=seqlen_q, dtype=torch.int32,
+                                        device=query.device)
+            is_causal = seqlen_q == seqlen_k
+            cu_seqlens_k = torch.arange(0, (batch_size + 1) * seqlen_k, step=seqlen_k, dtype=torch.int32,
+                                        device=query.device)
+            dropout_p = 0
+            softmax_scale = 1/torch.full([], (value.size(-1) ** 0.5), dtype=value.dtype, device=value.device) if self.scale_attn_weights else 1
+            attn_output = flash_attn_unpadded_func(
+                query, key, value, cu_seqlens_q, cu_seqlens_k, seqlen_q, seqlen_k,
+                dropout_p,
+                softmax_scale=softmax_scale, causal=is_causal
+            )
+            attn_output = rearrange(attn_output, '(b s) h ... -> b h s ...', b=batch_size)
+            attn_weights = None
+            return attn_output, attn_weights
+        attn_weights = torch.matmul(query, key.transpose(-1, -2))
+        if self.scale_attn_weights:
+            if self.use_mup:
+                attn_weights = attn_weights / torch.full(
+                    [], value.size(-1) / (value.size(-1) ** 0.5), dtype=attn_weights.dtype,
+                    device=attn_weights.device
+                )
+            else:
+                attn_weights = attn_weights / torch.full(
+                    [], value.size(-1) ** 0.5, dtype=attn_weights.dtype, device=attn_weights.device
+                )
+        if not self.is_cross_attention:
+            # if only "normal" attention layer implements causal mask
+            query_length, key_length = query.size(-2), key.size(-2)
+            causal_mask = self.bias[:, :, key_length - query_length: key_length, :key_length]
+            mask_value = torch.finfo(attn_weights.dtype).min
+            # Need to be a tensor, otherwise we get error: `RuntimeError: expected scalar type float but found double`.
+            # Need to be on the same device, otherwise `RuntimeError: ..., x and y to be on the same device`
+            mask_value = torch.full([], mask_value, dtype=attn_weights.dtype).to(attn_weights.device)
+            attn_weights = torch.where(causal_mask, attn_weights.to(attn_weights.dtype), mask_value)
+        if attention_mask is not None:
+            # Apply the attention mask
+            attn_weights = attn_weights + attention_mask
+        attn_weights = nn.functional.softmax(attn_weights, dim=-1)
+        # Downcast (if necessary) back to V's dtype (if in mixed-precision) -- No-Op otherwise
+        attn_weights = attn_weights.type(value.dtype)
+        attn_weights = self.attn_dropout(attn_weights)
+        # Mask heads if we want to
+        if head_mask is not None:
+            attn_weights = attn_weights * head_mask
+        attn_output = torch.matmul(attn_weights, value)
+        return attn_output, attn_weights
+    def _upcast_and_reordered_attn(self, query, key, value, attention_mask=None, head_mask=None):
+        # Use `torch.baddbmm` (a bit more efficient w/ alpha param for scaling -- from Megatron-LM)
+        bsz, num_heads, q_seq_len, dk = query.size()
+        _, _, k_seq_len, _ = key.size()
+        # Preallocate attn_weights for `baddbmm`
+        attn_weights = torch.empty(bsz * num_heads, q_seq_len, k_seq_len, dtype=query.dtype, device=query.device)
+        # Compute Scale Factor
+        scale_factor = 1.0
+        if self.scale_attn_weights:
+            scale_factor /= float(value.size(-1)) ** 0.5
+        if self.scale_attn_by_inverse_layer_idx:
+            scale_factor /= float(self.layer_idx)
+        # Upcast (turn off autocast) and reorder (Scale K by 1 / root(dk))
+        with autocast(enabled=False):
+            q, k = query.reshape(-1, q_seq_len, dk), key.transpose(-1, -2).reshape(-1, dk, k_seq_len)
+            attn_weights = torch.baddbmm(attn_weights, q, k, beta=0, alpha=scale_factor)
+            attn_weights = attn_weights.reshape(bsz, num_heads, q_seq_len, k_seq_len)
+        if not self.is_cross_attention:
+            attn_weights = attn_weights.float()
+            if self.scale_attn_by_inverse_layer_idx:
+                attn_weights *= self.layer_idx
+            # if only "normal" attention layer implements causal mask
+            query_length, key_length = query.size(-2), key.size(-2)
+            causal_mask = self.bias[:, :, key_length - query_length: key_length, :key_length]
+            mask_value = -10000.0  # align with megatron-lm
+            # Need to be a tensor, otherwise we get error: `RuntimeError: expected scalar type float but found double`.
+            # Need to be on the same device, otherwise `RuntimeError: ..., x and y to be on the same device`
+            mask_value = torch.tensor(mask_value, dtype=attn_weights.dtype).to(attn_weights.device)
+            attn_weights = torch.where(causal_mask, attn_weights, mask_value)
+        if attention_mask is not None:
+            # Apply the attention mask
+            attn_weights = attn_weights + attention_mask
+        attn_weights = nn.functional.softmax(attn_weights, dim=-1)
+        # Downcast (if necessary) back to V's dtype (if in mixed-precision) -- No-Op if otherwise
+        if attn_weights.dtype != torch.float32:
+            raise RuntimeError("Error with upcasting, attn_weights does not have dtype torch.float32")
+        attn_weights = attn_weights.type(value.dtype)
+        attn_weights = self.attn_dropout(attn_weights)
+        # Mask heads if we want to
+        if head_mask is not None:
+            attn_weights = attn_weights * head_mask
+        attn_output = torch.matmul(attn_weights, value)
+        return attn_output, attn_weights
+    def _split_heads(self, tensor, num_heads, attn_head_size):
+        """
+        Splits hidden_size dim into attn_head_size and num_heads
+        """
+        new_shape = tensor.size()[:-1] + (num_heads, attn_head_size)
+        tensor = tensor.view(new_shape)
+        return tensor.permute(0, 2, 1, 3)  # (batch, head, seq_length, head_features)
+    def _merge_heads(self, tensor, num_heads, attn_head_size):
+        """
+        Merges attn_head_size dim and num_attn_heads dim into hidden_size
+        """
+        tensor = tensor.permute(0, 2, 1, 3).contiguous()
+        new_shape = tensor.size()[:-2] + (num_heads * attn_head_size,)
+        return tensor.view(new_shape)
+    def forward(
+            self,
+            hidden_states: Optional[Tuple[torch.FloatTensor]],
+            layer_past: Optional[Tuple[torch.Tensor]] = None,
+            attention_mask: Optional[torch.FloatTensor] = None,
+            head_mask: Optional[torch.FloatTensor] = None,
+            encoder_hidden_states: Optional[torch.Tensor] = None,
+            encoder_attention_mask: Optional[torch.FloatTensor] = None,
+            rotary_embedding: Optional[RotaryEmbedding] = None,
+            use_cache: Optional[bool] = False,
+            output_attentions: Optional[bool] = False,
+    ) -> Tuple[Union[torch.Tensor, Tuple[torch.Tensor]], ...]:
+        if encoder_hidden_states is not None:
+            if not hasattr(self, "q_attn"):
+                raise ValueError(
+                    "If class is used as cross attention, the weights `q_attn` have to be defined. "
+                    "Please make sure to instantiate class with `GPT2Attention(..., is_cross_attention=True)`."
+                )
+            query = self.q_attn(hidden_states)
+            key, value = self.c_attn(encoder_hidden_states).split(self.split_size, dim=2)
+            attention_mask = encoder_attention_mask
+        else:
+            query, key, value = self.c_attn(hidden_states).split(self.split_size, dim=2)
+        query = self._split_heads(query, self.num_heads, self.head_dim)
+        key = self._split_heads(key, self.num_heads, self.head_dim)
+        value = self._split_heads(value, self.num_heads, self.head_dim)
+        if layer_past is not None:
+            past_key, past_value = layer_past
+            key = torch.cat((past_key, key), dim=-2)
+            value = torch.cat((past_value, value), dim=-2)
+        if use_cache is True:
+            present = (key, value)
+        else:
+            present = None
+        batch_size, head_num, k_seq_len, head_features = key.shape
+        _, _, q_seq_len, _ = query.shape
+        query_offset = k_seq_len - q_seq_len
+        if rotary_embedding is not None:
+            query = query.contiguous().view(batch_size * head_num, q_seq_len, head_features)
+            key = key.contiguous().view(batch_size * head_num, k_seq_len, head_features)
+            # batch_size * head_num,  k_seq_len(q_seq_len), head_features
+            if self.rotary_use_xpos:
+                # query: [batch_size * head_num, seqlen, hn]
+                query, key = rotary_embedding.rotate_queries_and_keys(query, key)
+            else:
+                query = rotary_embedding.rotate_queries_or_keys(query, offset=query_offset)
+                key = rotary_embedding.rotate_queries_or_keys(key)
+            # batch_size * head_num, k_seq_len(q_seq_len), head_features
+            query = query.view(batch_size, head_num, q_seq_len, head_features)
+            key = key.view(batch_size, head_num, k_seq_len, head_features)
+        if self.reorder_and_upcast_attn and not self.use_flash_attn:
+            attn_output, attn_weights = self._upcast_and_reordered_attn(query, key, value, attention_mask, head_mask)
+        else:
+            attn_output, attn_weights = self._attn(query, key, value, attention_mask, head_mask)
+        attn_output = self._merge_heads(attn_output, self.num_heads, self.head_dim)
+        attn_output = self.c_proj(attn_output)
+        attn_output = self.resid_dropout(attn_output)
+        outputs = (attn_output, present)
+        if output_attentions:
+            outputs += (attn_weights,)
+        return outputs
+class TELECHATMLP(nn.Module):
+    def __init__(self, intermediate_size, config):
+        super().__init__()
+        embed_dim = config.hidden_size
+        if config.activation_function=='silu':
+            up_intermediate_size = 2 * intermediate_size
+        else:
+            up_intermediate_size = intermediate_size
+        self.c_fc = Conv1D(up_intermediate_size, embed_dim, bias=config.add_bias_linear)
+        self.c_proj = Conv1D(embed_dim, intermediate_size, bias=config.add_bias_linear)
+        if config.activation_function=='silu':
+            def swiglu(x):
+                x = torch.chunk(x, 2, dim=-1)
+                return F.silu(x[0]) * x[1]
+            self.act = swiglu
+        else:
+            self.act = ACT2FN[config.activation_function]
+        self.dropout = nn.Dropout(config.resid_pdrop)
+    def forward(self, hidden_states: Optional[Tuple[torch.FloatTensor]]) -> torch.FloatTensor:
+        hidden_states = self.c_fc(hidden_states)
+        # print(f'activation func: {self.act}')
+        # print(f'before act: hidden_states {hidden_states.shape}')
+        hidden_states = self.act(hidden_states)
+        # print(f'after  act: hidden_states {hidden_states.shape}')
+        hidden_states = self.c_proj(hidden_states)
+        hidden_states = self.dropout(hidden_states)
+        return hidden_states
+class TELECHATBlock(nn.Module):
+    def __init__(self, config, layer_idx=None):
+        super().__init__()
+        LayerNorm = nn.LayerNorm if not config.use_RMSNorm else RMSNorm
+        hidden_size = config.hidden_size
+        inner_dim = config.n_inner if config.n_inner is not None else 4 * hidden_size
+        self.layer_idx = layer_idx
+        self.ln_1 = LayerNorm(hidden_size, eps=config.layer_norm_epsilon)
+        self.attn = TELECHATAttention(config, layer_idx=layer_idx)
+        self.ln_2 = LayerNorm(hidden_size, eps=config.layer_norm_epsilon)
+        self.mlp = TELECHATMLP(inner_dim, config)
+    def forward(
+            self,
+            hidden_states: Optional[Tuple[torch.FloatTensor]],
+            layer_past: Optional[Tuple[torch.Tensor]] = None,
+            attention_mask: Optional[torch.FloatTensor] = None,
+            head_mask: Optional[torch.FloatTensor] = None,
+            encoder_hidden_states: Optional[torch.Tensor] = None,
+            encoder_attention_mask: Optional[torch.FloatTensor] = None,
+            rotary_embedding: Optional[RotaryEmbedding] = None,
+            use_cache: Optional[bool] = False,
+            output_attentions: Optional[bool] = False,
+    ) -> Union[Tuple[torch.Tensor], Optional[Tuple[torch.Tensor, Tuple[torch.FloatTensor, ...]]]]:
+        residual = hidden_states
+        hidden_states = self.ln_1(hidden_states)
+        # debug_print_tensor(hidden_states, 'after ln_1')
+        attn_outputs = self.attn(
+            hidden_states,
+            layer_past=layer_past,
+            attention_mask=attention_mask,
+            head_mask=head_mask,
+            rotary_embedding=rotary_embedding,
+            use_cache=use_cache,
+            output_attentions=output_attentions
+        )
+        attn_output = attn_outputs[0]  # output_attn: a, present, (attentions)
+        outputs = attn_outputs[1:]
+        # residual connection
+        hidden_states = attn_output + residual
+        residual = hidden_states
+        hidden_states = self.ln_2(hidden_states)
+        feed_forward_hidden_states = self.mlp(hidden_states)
+        # residual connection
+        hidden_states = residual + feed_forward_hidden_states
+        if use_cache:
+            outputs = (hidden_states,) + outputs
+        else:
+            outputs = (hidden_states,) + outputs[1:]
+        # debug_print_tensor(hidden_states, 'block output')
+        return outputs
+class TELECHATPretrainedModel(PreTrainedModel):
+    """
+    An abstract class to handle weights initialization and a simple interface for downloading and loading pretrained
+    models.
+    """
+    config_class = TELECHATConfig
+    load_tf_weights = None
+    base_model_prefix = "transformer"
+    is_parallelizable = True
+    supports_gradient_checkpointing = True
+    _no_split_modules = ["TELECHATBlock"]
+    def __init__(self, *inputs, **kwargs):
+        super().__init__(*inputs, **kwargs)
+    def _init_weights(self, module):
+        """Initialize the weights."""
+        if isinstance(module, (nn.Linear, Conv1D)):
+            # Slightly different from the TF version which uses truncated_normal for initialization
+            # cf https://github.com/pytorch/pytorch/pull/5617
+            module.weight.data.normal_(mean=0.0, std=self.config.initializer_range)
+            if module.bias is not None:
+                module.bias.data.zero_()
+        elif isinstance(module, nn.Embedding):
+            module.weight.data.normal_(mean=0.0, std=self.config.initializer_range)
+            if module.padding_idx is not None:
+                module.weight.data[module.padding_idx].zero_()
+        elif isinstance(module, nn.LayerNorm) or isinstance(module, RMSNorm):
+            module.bias.data.zero_()
+            module.weight.data.fill_(1.0)
+        # Reinitialize selected weights subject to the OpenAI GPT-2 Paper Scheme:
+        #   > A modified initialization which accounts for the accumulation on the residual path with model depth. Scale
+        #   > the weights of residual layers at initialization by a factor of 1/√N where N is the # of residual layers.
+        #   >   -- GPT-2 :: https://openai.com/blog/better-language-models/
+        #
+        # Reference (Megatron-LM): https://github.com/NVIDIA/Megatron-LM/blob/main/megatron/model/gpt_model.py
+        for name, p in module.named_parameters():
+            if name == "c_proj.weight":
+                # Special Scaled Initialization --> There are 2 Layer Norms per Transformer Block
+                p.data.normal_(mean=0.0, std=(self.config.initializer_range / math.sqrt(2 * self.config.n_layer)))
+    def _set_gradient_checkpointing(self, module, value=False):
+        if isinstance(module, TELECHATTransformer):
+            module.gradient_checkpointing = value
+class TELECHATTransformer(TELECHATPretrainedModel):
+    _keys_to_ignore_on_load_missing = ["attn.masked_bias"]
+    def __init__(self, config):
+        super().__init__(config)
+        self.embed_dim = config.hidden_size
+        self.relative_encoding = config.relative_encoding
+        self.wte = nn.Embedding(config.vocab_size, self.embed_dim)
+        self.use_mup = config.use_mup
+        if self.use_mup:
+            self.input_mult = config.input_mult
+        if self.relative_encoding is None:
+            self.wpe = nn.Embedding(config.max_position_embeddings, self.embed_dim)
+        elif self.relative_encoding == 'rotary':
+            pe_dim = config.n_embd // config.n_head
+            self.wpe = RotaryEmbedding(pe_dim,
+                                       use_xpos=config.rotary_use_xpos,
+                                       xpos_scale_base=config.rotary_xpos_scale_base,
+                                       theta=config.rotary_theta
+                                       )
+        else:
+            raise RuntimeError(
+                f'Unknown relative positional encoding type: `relative_encoding`={self.relative_encoding}')
+        self.drop = nn.Dropout(config.embd_pdrop)
+        self.h = nn.ModuleList([TELECHATBlock(config, layer_idx=i + 1) for i in range(config.num_hidden_layers)])
+        LayerNorm = nn.LayerNorm if not config.use_RMSNorm else RMSNorm
+        self.ln_f = LayerNorm(self.embed_dim, eps=config.layer_norm_epsilon)
+        # Model parallel
+        self.model_parallel = False
+        self.device_map = None
+        self.gradient_checkpointing = False
+        # Initialize weights and apply final processing
+        self.post_init()
+    # @add_start_docstrings(PARALLELIZE_DOCSTRING)
+    def parallelize(self, device_map=None):
+        # Check validity of device_map
+        self.device_map = (
+            get_device_map(len(self.h), range(torch.cuda.device_count())) if device_map is None else device_map
+        )
+        assert_device_map(self.device_map, len(self.h))
+        self.model_parallel = True
+        self.first_device = "cpu" if "cpu" in self.device_map.keys() else "cuda:" + str(min(self.device_map.keys()))
+        self.last_device = "cuda:" + str(max(self.device_map.keys()))
+        self.wte = self.wte.to(self.first_device)
+        self.wpe = self.wpe.to(self.first_device)
+        # Load onto devices
+        for k, v in self.device_map.items():
+            for block in v:
+                cuda_device = "cuda:" + str(k)
+                self.h[block] = self.h[block].to(cuda_device)
+        # ln_f to last
+        self.ln_f = self.ln_f.to(self.last_device)
+    def deparallelize(self):
+        self.model_parallel = False
+        self.device_map = None
+        self.first_device = "cpu"
+        self.last_device = "cpu"
+        self.wte = self.wte.to("cpu")
+        self.wpe = self.wpe.to("cpu")
+        for index in range(len(self.h)):
+            self.h[index] = self.h[index].to("cpu")
+        self.ln_f = self.ln_f.to("cpu")
+        torch.cuda.empty_cache()
+    def get_input_embeddings(self):
+        return self.wte
+    def set_input_embeddings(self, new_embeddings):
+        self.wte = new_embeddings
+    def _prune_heads(self, heads_to_prune):
+        """
+        Prunes heads of the model. heads_to_prune: dict of {layer_num: list of heads to prune in this layer}
+        """
+        for layer, heads in heads_to_prune.items():
+            self.h[layer].attn.prune_heads(heads)
+    def forward(
+            self,
+            input_ids: Optional[torch.LongTensor] = None,
+            past_key_values: Optional[Tuple[Tuple[torch.Tensor]]] = None,
+            attention_mask: Optional[torch.FloatTensor] = None,
+            token_type_ids: Optional[torch.LongTensor] = None,
+            position_ids: Optional[torch.LongTensor] = None,
+            head_mask: Optional[torch.FloatTensor] = None,
+            inputs_embeds: Optional[torch.FloatTensor] = None,
+            encoder_hidden_states: Optional[torch.Tensor] = None,
+            encoder_attention_mask: Optional[torch.FloatTensor] = None,
+            use_cache: Optional[bool] = None,
+            output_attentions: Optional[bool] = None,
+            output_hidden_states: Optional[bool] = None,
+            return_dict: Optional[bool] = None,
+    ) -> Union[Tuple, BaseModelOutputWithPastAndCrossAttentions]:
+        output_attentions = output_attentions if output_attentions is not None else self.config.output_attentions
+        output_hidden_states = (
+            output_hidden_states if output_hidden_states is not None else self.config.output_hidden_states
+        )
+        use_cache = use_cache if use_cache is not None else self.config.use_cache
+        return_dict = return_dict if return_dict is not None else self.config.use_return_dict
+        if input_ids is not None and inputs_embeds is not None:
+            raise ValueError("You cannot specify both input_ids and inputs_embeds at the same time")
+        elif input_ids is not None:
+            input_shape = input_ids.size()
+            input_ids = input_ids.view(-1, input_shape[-1])
+            batch_size = input_ids.shape[0]
+        elif inputs_embeds is not None:
+            input_shape = inputs_embeds.size()[:-1]
+            batch_size = inputs_embeds.shape[0]
+        else:
+            raise ValueError("You have to specify either input_ids or inputs_embeds")
+        device = input_ids.device if input_ids is not None else inputs_embeds.device
+        if token_type_ids is not None:
+            token_type_ids = token_type_ids.view(-1, input_shape[-1])
+        if position_ids is not None:
+            position_ids = position_ids.view(-1, input_shape[-1])
+        if past_key_values is None:
+            past_length = 0
+            past_key_values = tuple([None] * len(self.h))
+        else:
+            past_length = past_key_values[0][0].size(-2)
+        if position_ids is None:
+            position_ids = torch.arange(past_length, input_shape[-1] + past_length, dtype=torch.long, device=device)
+            position_ids = position_ids.unsqueeze(0).view(-1, input_shape[-1])
+        # GPT2Attention mask.
+        if attention_mask is not None:
+            if batch_size <= 0:
+                raise ValueError("batch_size has to be defined and > 0")
+            attention_mask = attention_mask.view(batch_size, -1)
+            # We create a 3D attention mask from a 2D tensor mask.
+            # Sizes are [batch_size, 1, 1, to_seq_length]
+            # So we can broadcast to [batch_size, num_heads, from_seq_length, to_seq_length]
+            # this attention mask is more simple than the triangular masking of causal attention
+            # used in OpenAI GPT, we just need to prepare the broadcast dimension here.
+            attention_mask = attention_mask[:, None, None, :]
+            # Since attention_mask is 1.0 for positions we want to attend and 0.0 for
+            # masked positions, this operation will create a tensor which is 0.0 for
+            # positions we want to attend and the dtype's smallest value for masked positions.
+            # Since we are adding it to the raw scores before the softmax, this is
+            # effectively the same as removing these entirely.
+            attention_mask = attention_mask.to(dtype=self.dtype)  # fp16 compatibility
+            attention_mask = (1.0 - attention_mask) * torch.finfo(self.dtype).min
+        # If a 2D or 3D attention mask is provided for the cross-attention
+        # we need to make broadcastable to [batch_size, num_heads, seq_length, seq_length]
+        # if self.config.add_cross_attention and encoder_hidden_states is not None:
+        #     encoder_batch_size, encoder_sequence_length, _ = encoder_hidden_states.size()
+        #     encoder_hidden_shape = (encoder_batch_size, encoder_sequence_length)
+        #     if encoder_attention_mask is None:
+        #         encoder_attention_mask = torch.ones(encoder_hidden_shape, device=device)
+        #     encoder_attention_mask = self.invert_attention_mask(encoder_attention_mask)
+        # else:
+        #     encoder_attention_mask = None
+        encoder_attention_mask = None
+        # Prepare head mask if needed
+        # 1.0 in head_mask indicate we keep the head
+        # attention_probs has shape bsz x n_heads x N x N
+        # head_mask has shape n_layer x batch x n_heads x N x N
+        head_mask = self.get_head_mask(head_mask, self.config.n_layer)
+        if inputs_embeds is None:
+            inputs_embeds = self.wte(input_ids)
+        # Mup
+        if self.use_mup:
+            inputs_embeds = inputs_embeds * self.input_mult
+        if self.relative_encoding is None:
+            position_embeds = self.wpe(position_ids)
+            hidden_states = inputs_embeds + position_embeds
+        elif self.relative_encoding == 'rotary':
+            hidden_states = inputs_embeds
+        if token_type_ids is not None:
+            token_type_embeds = self.wte(token_type_ids)
+            hidden_states = hidden_states + token_type_embeds
+        hidden_states = self.drop(hidden_states)
+        output_shape = input_shape + (hidden_states.size(-1),)
+        presents = () if use_cache else None
+        all_self_attentions = () if output_attentions else None
+        all_cross_attentions = () if output_attentions and self.config.add_cross_attention else None
+        all_hidden_states = () if output_hidden_states else None
+        # debug_print_tensor(hidden_states, 'after embedding')
+        for i, (block, layer_past) in enumerate(zip(self.h, past_key_values)):
+            # Model parallel
+            if self.model_parallel:
+                torch.cuda.set_device(hidden_states.device)
+                # Ensure layer_past is on same device as hidden_states (might not be correct)
+                if layer_past is not None:
+                    layer_past = tuple(past_state.to(hidden_states.device) for past_state in layer_past)
+                # Ensure that attention_mask is always on the same device as hidden_states
+                if attention_mask is not None:
+                    attention_mask = attention_mask.to(hidden_states.device)
+                if isinstance(head_mask, torch.Tensor):
+                    head_mask = head_mask.to(hidden_states.device)
+            if output_hidden_states:
+                all_hidden_states = all_hidden_states + (hidden_states,)
+            if self.gradient_checkpointing and self.training:
+                if use_cache:
+                    logger.warning(
+                        "`use_cache=True` is incompatible with gradient checkpointing. Setting `use_cache=False`..."
+                    )
+                    use_cache = False
+                def create_custom_forward(module):
+                    def custom_forward(*inputs):
+                        # None for past_key_value
+                        return module(*inputs, use_cache, output_attentions)
+                    return custom_forward
+                outputs = torch.utils.checkpoint.checkpoint(
+                    create_custom_forward(block),
+                    hidden_states,
+                    None,
+                    attention_mask,
+                    head_mask[i],
+                    encoder_hidden_states,
+                    encoder_attention_mask,
+                )
+            else:
+                outputs = block(
+                    hidden_states,
+                    layer_past=layer_past,
+                    attention_mask=attention_mask,
+                    head_mask=head_mask[i],
+                    encoder_hidden_states=encoder_hidden_states,
+                    encoder_attention_mask=encoder_attention_mask,
+                    rotary_embedding=self.wpe if self.relative_encoding == 'rotary' else None,
+                    use_cache=use_cache,
+                    output_attentions=output_attentions
+                )
+            hidden_states = outputs[0]
+            if use_cache is True:
+                presents = presents + (outputs[1],)
+            if output_attentions:
+                all_self_attentions = all_self_attentions + (outputs[2 if use_cache else 1],)
+                # if self.config.add_cross_attention:
+                #     all_cross_attentions = all_cross_attentions + (outputs[3 if use_cache else 2],)
+            # Model Parallel: If it's the last layer for that device, put things on the next device
+            if self.model_parallel:
+                for k, v in self.device_map.items():
+                    if i == v[-1] and "cuda:" + str(k) != self.last_device:
+                        hidden_states = hidden_states.to("cuda:" + str(k + 1))
+        hidden_states = self.ln_f(hidden_states)
+        hidden_states = hidden_states.view(output_shape)
+        # Add last hidden state
+        if output_hidden_states:
+            all_hidden_states = all_hidden_states + (hidden_states,)
+        if not return_dict:
+            return tuple(
+                v
+                for v in [hidden_states, presents, all_hidden_states, all_self_attentions, all_cross_attentions]
+                if v is not None
+            )
+        return BaseModelOutputWithPastAndCrossAttentions(
+            last_hidden_state=hidden_states,
+            past_key_values=presents,
+            hidden_states=all_hidden_states,
+            attentions=all_self_attentions,
+            cross_attentions=all_cross_attentions,
+        )
+class TELECHAT(TELECHATPretrainedModel):
+    _keys_to_ignore_on_load_missing = [r"attn.masked_bias", r"attn.bias", r"lm_head.weight"]
+    def __init__(self, config):
+        super().__init__(config)
+        self.transformer = TELECHATTransformer(config)
+        self.lm_head = nn.Linear(config.n_embd, config.vocab_size, bias=False)
+        self.use_mup = config.use_mup
+        if self.use_mup:
+            self.mup_scale_factor = config.mup_scale_factor
+            self.output_mult = config.output_mult / self.mup_scale_factor
+        # 初始化时先根据config里的开关决定是否开启flashattn, 用户可以通过修改config或者model.enable_flash_attn修改flashattn的开关
+        self.enable_flash_attn(config.enable_flash_attn)
+        # Model parallel
+        self.model_parallel = False
+        self.device_map = None
+        # Initialize weights and apply final processing
+        self.post_init()
+    def enable_flash_attn(self, enabled: bool):
+        for block in self.transformer.h:
+            block.attn.use_flash_attn = enabled
+        print(f"TELECHAT flash attention {'enabled' if enabled else 'disabled'}")
+        # torch.backends.cuda.enable_flash_sdp(enabled)
+    def set_max_positions(self, max_positions):
+        for layer in self.transformer.h:
+            device = layer.ln_1.weight.device
+            layer.attn.set_max_positions(max_positions, device=device)
+    def parallelize(self, device_map=None):
+        self.device_map = (
+            get_device_map(len(self.transformer.h), range(torch.cuda.device_count()))
+            if device_map is None
+            else device_map
+        )
+        assert_device_map(self.device_map, len(self.transformer.h))
+        self.transformer.parallelize(self.device_map)
+        self.lm_head = self.lm_head.to(self.transformer.first_device)
+        self.model_parallel = True
+    def deparallelize(self):
+        self.transformer.deparallelize()
+        self.transformer = self.transformer.to("cpu")
+        self.lm_head = self.lm_head.to("cpu")
+        self.model_parallel = False
+        torch.cuda.empty_cache()
+    def get_output_embeddings(self):
+        return self.lm_head
+    def set_output_embeddings(self, new_embeddings):
+        self.lm_head = new_embeddings
+    def prepare_inputs_for_generation(self, input_ids, past_key_values=None, **kwargs):
+        token_type_ids = kwargs.get("token_type_ids", None)
+        # only last token for inputs_ids if past is defined in kwargs
+        if past_key_values:
+            input_ids = input_ids[:, -1].unsqueeze(-1)
+            if token_type_ids is not None:
+                token_type_ids = token_type_ids[:, -1].unsqueeze(-1)
+        attention_mask = kwargs.get("attention_mask", None)
+        position_ids = kwargs.get("position_ids", None)
+        if attention_mask is not None and position_ids is None:
+            # create position_ids on the fly for batch generation
+            position_ids = attention_mask.long().cumsum(-1) - 1
+            position_ids.masked_fill_(attention_mask == 0, 1)
+            if past_key_values:
+                position_ids = position_ids[:, -1].unsqueeze(-1)
+        else:
+            position_ids = None
+        return {
+            "input_ids": input_ids,
+            "past_key_values": past_key_values,
+            "use_cache": kwargs.get("use_cache"),
+            "position_ids": position_ids,
+            "attention_mask": attention_mask,
+            "token_type_ids": token_type_ids,
+        }
+    def forward(
+            self,
+            input_ids: Optional[torch.LongTensor] = None,
+            past_key_values: Optional[Tuple[Tuple[torch.Tensor]]] = None,
+            attention_mask: Optional[torch.FloatTensor] = None,
+            token_type_ids: Optional[torch.LongTensor] = None,
+            position_ids: Optional[torch.LongTensor] = None,
+            head_mask: Optional[torch.FloatTensor] = None,
+            inputs_embeds: Optional[torch.FloatTensor] = None,
+            encoder_hidden_states: Optional[torch.Tensor] = None,
+            encoder_attention_mask: Optional[torch.FloatTensor] = None,
+            labels: Optional[torch.LongTensor] = None,
+            use_cache: Optional[bool] = None,
+            output_attentions: Optional[bool] = None,
+            output_hidden_states: Optional[bool] = None,
+            return_dict: Optional[bool] = None,
+    ) -> Union[Tuple, CausalLMOutputWithCrossAttentions, SequenceClassifierOutputWithPast]:
+        return_dict = return_dict if return_dict is not None else self.config.use_return_dict
+        transformer_outputs = self.transformer(
+            input_ids,
+            past_key_values=past_key_values,
+            attention_mask=attention_mask,
+            token_type_ids=token_type_ids,
+            position_ids=position_ids,
+            head_mask=head_mask,
+            inputs_embeds=inputs_embeds,
+            encoder_hidden_states=encoder_hidden_states,
+            encoder_attention_mask=encoder_attention_mask,
+            use_cache=use_cache,
+            output_attentions=output_attentions,
+            output_hidden_states=output_hidden_states,
+            return_dict=return_dict
+        )
+        hidden_states = transformer_outputs[0]
+        # Set device for model parallelism
+        if self.model_parallel:
+            torch.cuda.set_device(self.transformer.first_device)
+            hidden_states = hidden_states.to(self.lm_head.weight.device)
+        lm_logits = self.lm_head(hidden_states)
+        # Mup
+        if self.use_mup:
+            lm_logits = lm_logits * self.output_mult
+        loss = None
+        if labels is not None:
+            # Shift so that tokens < n predict n
+            shift_logits = lm_logits[..., :-1, :].contiguous()
+            shift_labels = labels[..., 1:].contiguous()
+            # Flatten the tokens
+            loss_fct = nn.CrossEntropyLoss()
+            loss = loss_fct(shift_logits.view(-1, shift_logits.size(-1)), shift_labels.view(-1))
+        if not return_dict:
+            output = (lm_logits,) + transformer_outputs[1:]
+            return ((loss,) + output) if loss is not None else output
+        return CausalLMOutputWithCrossAttentions(
+            loss=loss,
+            logits=lm_logits,
+            past_key_values=transformer_outputs.past_key_values,
+            hidden_states=transformer_outputs.hidden_states,
+            attentions=transformer_outputs.attentions,
+            cross_attentions=transformer_outputs.cross_attentions,
+        )
+    def chat(self,tokenizer, question, history_input_list, history_output_list,generation_config):
+            '''
+            :param question: 当前问题
+            :param history_input_list: 历史问题列表, list of strings
+            :param history_output_list: 历史回答列表, list of string
+            :return: response
+            '''
+            inputs = ""
+            assert len(history_output_list) == len(history_output_list)
+            for i in range(len(history_input_list)):
+                inputs += "<_user>" + history_input_list[i] + "<_bot>" + history_output_list[i] + "<_end>"
+            inputs += "<_user>" + question + "<_bot>"
+            print("input:", inputs)
+            input_ids = tokenizer.encode(inputs,
+                                         return_tensors="pt"
+                                         )
+            if len(input_ids[0]) >= 2000:
+                input_ids = input_ids[:, -2000:]
+            input_ids = input_ids.to(0)
+            output = self.generate(input_ids,generation_config)
+            response = tokenizer.decode(output[0].cpu().numpy().tolist()).split('<_bot>')[-1].split('</s>')[0]
+            return response
+    @staticmethod
+    def _reorder_cache(past: Tuple[Tuple[torch.Tensor]], beam_idx: torch.Tensor) -> Tuple[Tuple[torch.Tensor]]:
+        """
+        This function is used to re-order the `past_key_values` cache if [`~PreTrainedModel.beam_search`] or
+        [`~PreTrainedModel.beam_sample`] is called. This is required to match `past_key_values` with the correct
+        beam_idx at every generation step.
+        """
+        return tuple(
+            tuple(past_state.index_select(0, beam_idx.to(past_state.device)) for past_state in layer_past)
+            for layer_past in past
+        )

pytorch_model.bin.index.json ADDED Viewed

	@@ -0,0 +1,458 @@

+{
+  "metadata": {
+    "total_size": 105665020032
+  },
+  "weight_map": {
+    "lm_head.weight": "pytorch_model-00011-of-00011.bin",
+    "transformer.h.0.attn.c_attn.weight": "pytorch_model-00001-of-00011.bin",
+    "transformer.h.0.attn.c_proj.weight": "pytorch_model-00001-of-00011.bin",
+    "transformer.h.0.attn.masked_bias": "pytorch_model-00001-of-00011.bin",
+    "transformer.h.0.ln_1.weight": "pytorch_model-00001-of-00011.bin",
+    "transformer.h.0.ln_2.weight": "pytorch_model-00001-of-00011.bin",
+    "transformer.h.0.mlp.c_fc.weight": "pytorch_model-00001-of-00011.bin",
+    "transformer.h.0.mlp.c_proj.weight": "pytorch_model-00001-of-00011.bin",
+    "transformer.h.1.attn.c_attn.weight": "pytorch_model-00001-of-00011.bin",
+    "transformer.h.1.attn.c_proj.weight": "pytorch_model-00001-of-00011.bin",
+    "transformer.h.1.attn.masked_bias": "pytorch_model-00001-of-00011.bin",
+    "transformer.h.1.ln_1.weight": "pytorch_model-00001-of-00011.bin",
+    "transformer.h.1.ln_2.weight": "pytorch_model-00001-of-00011.bin",
+    "transformer.h.1.mlp.c_fc.weight": "pytorch_model-00001-of-00011.bin",
+    "transformer.h.1.mlp.c_proj.weight": "pytorch_model-00001-of-00011.bin",
+    "transformer.h.10.attn.c_attn.weight": "pytorch_model-00002-of-00011.bin",
+    "transformer.h.10.attn.c_proj.weight": "pytorch_model-00002-of-00011.bin",
+    "transformer.h.10.attn.masked_bias": "pytorch_model-00002-of-00011.bin",
+    "transformer.h.10.ln_1.weight": "pytorch_model-00002-of-00011.bin",
+    "transformer.h.10.ln_2.weight": "pytorch_model-00002-of-00011.bin",
+    "transformer.h.10.mlp.c_fc.weight": "pytorch_model-00002-of-00011.bin",
+    "transformer.h.10.mlp.c_proj.weight": "pytorch_model-00002-of-00011.bin",
+    "transformer.h.11.attn.c_attn.weight": "pytorch_model-00002-of-00011.bin",
+    "transformer.h.11.attn.c_proj.weight": "pytorch_model-00002-of-00011.bin",
+    "transformer.h.11.attn.masked_bias": "pytorch_model-00002-of-00011.bin",
+    "transformer.h.11.ln_1.weight": "pytorch_model-00002-of-00011.bin",
+    "transformer.h.11.ln_2.weight": "pytorch_model-00002-of-00011.bin",
+    "transformer.h.11.mlp.c_fc.weight": "pytorch_model-00003-of-00011.bin",
+    "transformer.h.11.mlp.c_proj.weight": "pytorch_model-00003-of-00011.bin",
+    "transformer.h.12.attn.c_attn.weight": "pytorch_model-00003-of-00011.bin",
+    "transformer.h.12.attn.c_proj.weight": "pytorch_model-00003-of-00011.bin",
+    "transformer.h.12.attn.masked_bias": "pytorch_model-00003-of-00011.bin",
+    "transformer.h.12.ln_1.weight": "pytorch_model-00003-of-00011.bin",
+    "transformer.h.12.ln_2.weight": "pytorch_model-00003-of-00011.bin",
+    "transformer.h.12.mlp.c_fc.weight": "pytorch_model-00003-of-00011.bin",
+    "transformer.h.12.mlp.c_proj.weight": "pytorch_model-00003-of-00011.bin",
+    "transformer.h.13.attn.c_attn.weight": "pytorch_model-00003-of-00011.bin",
+    "transformer.h.13.attn.c_proj.weight": "pytorch_model-00003-of-00011.bin",
+    "transformer.h.13.attn.masked_bias": "pytorch_model-00003-of-00011.bin",
+    "transformer.h.13.ln_1.weight": "pytorch_model-00003-of-00011.bin",
+    "transformer.h.13.ln_2.weight": "pytorch_model-00003-of-00011.bin",
+    "transformer.h.13.mlp.c_fc.weight": "pytorch_model-00003-of-00011.bin",
+    "transformer.h.13.mlp.c_proj.weight": "pytorch_model-00003-of-00011.bin",
+    "transformer.h.14.attn.c_attn.weight": "pytorch_model-00003-of-00011.bin",
+    "transformer.h.14.attn.c_proj.weight": "pytorch_model-00003-of-00011.bin",
+    "transformer.h.14.attn.masked_bias": "pytorch_model-00003-of-00011.bin",
+    "transformer.h.14.ln_1.weight": "pytorch_model-00003-of-00011.bin",
+    "transformer.h.14.ln_2.weight": "pytorch_model-00003-of-00011.bin",
+    "transformer.h.14.mlp.c_fc.weight": "pytorch_model-00003-of-00011.bin",
+    "transformer.h.14.mlp.c_proj.weight": "pytorch_model-00003-of-00011.bin",
+    "transformer.h.15.attn.c_attn.weight": "pytorch_model-00003-of-00011.bin",
+    "transformer.h.15.attn.c_proj.weight": "pytorch_model-00003-of-00011.bin",
+    "transformer.h.15.attn.masked_bias": "pytorch_model-00003-of-00011.bin",
+    "transformer.h.15.ln_1.weight": "pytorch_model-00003-of-00011.bin",
+    "transformer.h.15.ln_2.weight": "pytorch_model-00003-of-00011.bin",
+    "transformer.h.15.mlp.c_fc.weight": "pytorch_model-00003-of-00011.bin",
+    "transformer.h.15.mlp.c_proj.weight": "pytorch_model-00003-of-00011.bin",
+    "transformer.h.16.attn.c_attn.weight": "pytorch_model-00003-of-00011.bin",
+    "transformer.h.16.attn.c_proj.weight": "pytorch_model-00003-of-00011.bin",
+    "transformer.h.16.attn.masked_bias": "pytorch_model-00003-of-00011.bin",
+    "transformer.h.16.ln_1.weight": "pytorch_model-00003-of-00011.bin",
+    "transformer.h.16.ln_2.weight": "pytorch_model-00003-of-00011.bin",
+    "transformer.h.16.mlp.c_fc.weight": "pytorch_model-00003-of-00011.bin",
+    "transformer.h.16.mlp.c_proj.weight": "pytorch_model-00003-of-00011.bin",
+    "transformer.h.17.attn.c_attn.weight": "pytorch_model-00003-of-00011.bin",
+    "transformer.h.17.attn.c_proj.weight": "pytorch_model-00003-of-00011.bin",
+    "transformer.h.17.attn.masked_bias": "pytorch_model-00003-of-00011.bin",
+    "transformer.h.17.ln_1.weight": "pytorch_model-00003-of-00011.bin",
+    "transformer.h.17.ln_2.weight": "pytorch_model-00003-of-00011.bin",
+    "transformer.h.17.mlp.c_fc.weight": "pytorch_model-00004-of-00011.bin",
+    "transformer.h.17.mlp.c_proj.weight": "pytorch_model-00004-of-00011.bin",
+    "transformer.h.18.attn.c_attn.weight": "pytorch_model-00004-of-00011.bin",
+    "transformer.h.18.attn.c_proj.weight": "pytorch_model-00004-of-00011.bin",
+    "transformer.h.18.attn.masked_bias": "pytorch_model-00004-of-00011.bin",
+    "transformer.h.18.ln_1.weight": "pytorch_model-00004-of-00011.bin",
+    "transformer.h.18.ln_2.weight": "pytorch_model-00004-of-00011.bin",
+    "transformer.h.18.mlp.c_fc.weight": "pytorch_model-00004-of-00011.bin",
+    "transformer.h.18.mlp.c_proj.weight": "pytorch_model-00004-of-00011.bin",
+    "transformer.h.19.attn.c_attn.weight": "pytorch_model-00004-of-00011.bin",
+    "transformer.h.19.attn.c_proj.weight": "pytorch_model-00004-of-00011.bin",
+    "transformer.h.19.attn.masked_bias": "pytorch_model-00004-of-00011.bin",
+    "transformer.h.19.ln_1.weight": "pytorch_model-00004-of-00011.bin",
+    "transformer.h.19.ln_2.weight": "pytorch_model-00004-of-00011.bin",
+    "transformer.h.19.mlp.c_fc.weight": "pytorch_model-00004-of-00011.bin",
+    "transformer.h.19.mlp.c_proj.weight": "pytorch_model-00004-of-00011.bin",
+    "transformer.h.2.attn.c_attn.weight": "pytorch_model-00001-of-00011.bin",
+    "transformer.h.2.attn.c_proj.weight": "pytorch_model-00001-of-00011.bin",
+    "transformer.h.2.attn.masked_bias": "pytorch_model-00001-of-00011.bin",
+    "transformer.h.2.ln_1.weight": "pytorch_model-00001-of-00011.bin",
+    "transformer.h.2.ln_2.weight": "pytorch_model-00001-of-00011.bin",
+    "transformer.h.2.mlp.c_fc.weight": "pytorch_model-00001-of-00011.bin",
+    "transformer.h.2.mlp.c_proj.weight": "pytorch_model-00001-of-00011.bin",
+    "transformer.h.20.attn.c_attn.weight": "pytorch_model-00004-of-00011.bin",
+    "transformer.h.20.attn.c_proj.weight": "pytorch_model-00004-of-00011.bin",
+    "transformer.h.20.attn.masked_bias": "pytorch_model-00004-of-00011.bin",
+    "transformer.h.20.ln_1.weight": "pytorch_model-00004-of-00011.bin",
+    "transformer.h.20.ln_2.weight": "pytorch_model-00004-of-00011.bin",
+    "transformer.h.20.mlp.c_fc.weight": "pytorch_model-00004-of-00011.bin",
+    "transformer.h.20.mlp.c_proj.weight": "pytorch_model-00004-of-00011.bin",
+    "transformer.h.21.attn.c_attn.weight": "pytorch_model-00004-of-00011.bin",
+    "transformer.h.21.attn.c_proj.weight": "pytorch_model-00004-of-00011.bin",
+    "transformer.h.21.attn.masked_bias": "pytorch_model-00004-of-00011.bin",
+    "transformer.h.21.ln_1.weight": "pytorch_model-00004-of-00011.bin",
+    "transformer.h.21.ln_2.weight": "pytorch_model-00004-of-00011.bin",
+    "transformer.h.21.mlp.c_fc.weight": "pytorch_model-00004-of-00011.bin",
+    "transformer.h.21.mlp.c_proj.weight": "pytorch_model-00004-of-00011.bin",
+    "transformer.h.22.attn.c_attn.weight": "pytorch_model-00004-of-00011.bin",
+    "transformer.h.22.attn.c_proj.weight": "pytorch_model-00004-of-00011.bin",
+    "transformer.h.22.attn.masked_bias": "pytorch_model-00004-of-00011.bin",
+    "transformer.h.22.ln_1.weight": "pytorch_model-00004-of-00011.bin",
+    "transformer.h.22.ln_2.weight": "pytorch_model-00004-of-00011.bin",
+    "transformer.h.22.mlp.c_fc.weight": "pytorch_model-00004-of-00011.bin",
+    "transformer.h.22.mlp.c_proj.weight": "pytorch_model-00004-of-00011.bin",
+    "transformer.h.23.attn.c_attn.weight": "pytorch_model-00004-of-00011.bin",
+    "transformer.h.23.attn.c_proj.weight": "pytorch_model-00004-of-00011.bin",
+    "transformer.h.23.attn.masked_bias": "pytorch_model-00004-of-00011.bin",
+    "transformer.h.23.ln_1.weight": "pytorch_model-00004-of-00011.bin",
+    "transformer.h.23.ln_2.weight": "pytorch_model-00004-of-00011.bin",
+    "transformer.h.23.mlp.c_fc.weight": "pytorch_model-00005-of-00011.bin",
+    "transformer.h.23.mlp.c_proj.weight": "pytorch_model-00005-of-00011.bin",
+    "transformer.h.24.attn.c_attn.weight": "pytorch_model-00005-of-00011.bin",
+    "transformer.h.24.attn.c_proj.weight": "pytorch_model-00005-of-00011.bin",
+    "transformer.h.24.attn.masked_bias": "pytorch_model-00005-of-00011.bin",
+    "transformer.h.24.ln_1.weight": "pytorch_model-00005-of-00011.bin",
+    "transformer.h.24.ln_2.weight": "pytorch_model-00005-of-00011.bin",
+    "transformer.h.24.mlp.c_fc.weight": "pytorch_model-00005-of-00011.bin",
+    "transformer.h.24.mlp.c_proj.weight": "pytorch_model-00005-of-00011.bin",
+    "transformer.h.25.attn.c_attn.weight": "pytorch_model-00005-of-00011.bin",
+    "transformer.h.25.attn.c_proj.weight": "pytorch_model-00005-of-00011.bin",
+    "transformer.h.25.attn.masked_bias": "pytorch_model-00005-of-00011.bin",
+    "transformer.h.25.ln_1.weight": "pytorch_model-00005-of-00011.bin",
+    "transformer.h.25.ln_2.weight": "pytorch_model-00005-of-00011.bin",
+    "transformer.h.25.mlp.c_fc.weight": "pytorch_model-00005-of-00011.bin",
+    "transformer.h.25.mlp.c_proj.weight": "pytorch_model-00005-of-00011.bin",
+    "transformer.h.26.attn.c_attn.weight": "pytorch_model-00005-of-00011.bin",
+    "transformer.h.26.attn.c_proj.weight": "pytorch_model-00005-of-00011.bin",
+    "transformer.h.26.attn.masked_bias": "pytorch_model-00005-of-00011.bin",
+    "transformer.h.26.ln_1.weight": "pytorch_model-00005-of-00011.bin",
+    "transformer.h.26.ln_2.weight": "pytorch_model-00005-of-00011.bin",
+    "transformer.h.26.mlp.c_fc.weight": "pytorch_model-00005-of-00011.bin",
+    "transformer.h.26.mlp.c_proj.weight": "pytorch_model-00005-of-00011.bin",
+    "transformer.h.27.attn.c_attn.weight": "pytorch_model-00005-of-00011.bin",
+    "transformer.h.27.attn.c_proj.weight": "pytorch_model-00005-of-00011.bin",
+    "transformer.h.27.attn.masked_bias": "pytorch_model-00005-of-00011.bin",
+    "transformer.h.27.ln_1.weight": "pytorch_model-00005-of-00011.bin",
+    "transformer.h.27.ln_2.weight": "pytorch_model-00005-of-00011.bin",
+    "transformer.h.27.mlp.c_fc.weight": "pytorch_model-00005-of-00011.bin",
+    "transformer.h.27.mlp.c_proj.weight": "pytorch_model-00005-of-00011.bin",
+    "transformer.h.28.attn.c_attn.weight": "pytorch_model-00005-of-00011.bin",
+    "transformer.h.28.attn.c_proj.weight": "pytorch_model-00005-of-00011.bin",
+    "transformer.h.28.attn.masked_bias": "pytorch_model-00005-of-00011.bin",
+    "transformer.h.28.ln_1.weight": "pytorch_model-00005-of-00011.bin",
+    "transformer.h.28.ln_2.weight": "pytorch_model-00005-of-00011.bin",
+    "transformer.h.28.mlp.c_fc.weight": "pytorch_model-00005-of-00011.bin",
+    "transformer.h.28.mlp.c_proj.weight": "pytorch_model-00005-of-00011.bin",
+    "transformer.h.29.attn.c_attn.weight": "pytorch_model-00005-of-00011.bin",
+    "transformer.h.29.attn.c_proj.weight": "pytorch_model-00005-of-00011.bin",
+    "transformer.h.29.attn.masked_bias": "pytorch_model-00005-of-00011.bin",
+    "transformer.h.29.ln_1.weight": "pytorch_model-00005-of-00011.bin",
+    "transformer.h.29.ln_2.weight": "pytorch_model-00005-of-00011.bin",
+    "transformer.h.29.mlp.c_fc.weight": "pytorch_model-00006-of-00011.bin",
+    "transformer.h.29.mlp.c_proj.weight": "pytorch_model-00006-of-00011.bin",
+    "transformer.h.3.attn.c_attn.weight": "pytorch_model-00001-of-00011.bin",
+    "transformer.h.3.attn.c_proj.weight": "pytorch_model-00001-of-00011.bin",
+    "transformer.h.3.attn.masked_bias": "pytorch_model-00001-of-00011.bin",
+    "transformer.h.3.ln_1.weight": "pytorch_model-00001-of-00011.bin",
+    "transformer.h.3.ln_2.weight": "pytorch_model-00001-of-00011.bin",
+    "transformer.h.3.mlp.c_fc.weight": "pytorch_model-00001-of-00011.bin",
+    "transformer.h.3.mlp.c_proj.weight": "pytorch_model-00001-of-00011.bin",
+    "transformer.h.30.attn.c_attn.weight": "pytorch_model-00006-of-00011.bin",
+    "transformer.h.30.attn.c_proj.weight": "pytorch_model-00006-of-00011.bin",
+    "transformer.h.30.attn.masked_bias": "pytorch_model-00006-of-00011.bin",
+    "transformer.h.30.ln_1.weight": "pytorch_model-00006-of-00011.bin",
+    "transformer.h.30.ln_2.weight": "pytorch_model-00006-of-00011.bin",
+    "transformer.h.30.mlp.c_fc.weight": "pytorch_model-00006-of-00011.bin",
+    "transformer.h.30.mlp.c_proj.weight": "pytorch_model-00006-of-00011.bin",
+    "transformer.h.31.attn.c_attn.weight": "pytorch_model-00006-of-00011.bin",
+    "transformer.h.31.attn.c_proj.weight": "pytorch_model-00006-of-00011.bin",
+    "transformer.h.31.attn.masked_bias": "pytorch_model-00006-of-00011.bin",
+    "transformer.h.31.ln_1.weight": "pytorch_model-00006-of-00011.bin",
+    "transformer.h.31.ln_2.weight": "pytorch_model-00006-of-00011.bin",
+    "transformer.h.31.mlp.c_fc.weight": "pytorch_model-00006-of-00011.bin",
+    "transformer.h.31.mlp.c_proj.weight": "pytorch_model-00006-of-00011.bin",
+    "transformer.h.32.attn.c_attn.weight": "pytorch_model-00006-of-00011.bin",
+    "transformer.h.32.attn.c_proj.weight": "pytorch_model-00006-of-00011.bin",
+    "transformer.h.32.attn.masked_bias": "pytorch_model-00006-of-00011.bin",
+    "transformer.h.32.ln_1.weight": "pytorch_model-00006-of-00011.bin",
+    "transformer.h.32.ln_2.weight": "pytorch_model-00006-of-00011.bin",
+    "transformer.h.32.mlp.c_fc.weight": "pytorch_model-00006-of-00011.bin",
+    "transformer.h.32.mlp.c_proj.weight": "pytorch_model-00006-of-00011.bin",
+    "transformer.h.33.attn.c_attn.weight": "pytorch_model-00006-of-00011.bin",
+    "transformer.h.33.attn.c_proj.weight": "pytorch_model-00006-of-00011.bin",
+    "transformer.h.33.attn.masked_bias": "pytorch_model-00006-of-00011.bin",
+    "transformer.h.33.ln_1.weight": "pytorch_model-00006-of-00011.bin",
+    "transformer.h.33.ln_2.weight": "pytorch_model-00006-of-00011.bin",
+    "transformer.h.33.mlp.c_fc.weight": "pytorch_model-00006-of-00011.bin",
+    "transformer.h.33.mlp.c_proj.weight": "pytorch_model-00006-of-00011.bin",
+    "transformer.h.34.attn.c_attn.weight": "pytorch_model-00006-of-00011.bin",
+    "transformer.h.34.attn.c_proj.weight": "pytorch_model-00006-of-00011.bin",
+    "transformer.h.34.attn.masked_bias": "pytorch_model-00006-of-00011.bin",
+    "transformer.h.34.ln_1.weight": "pytorch_model-00006-of-00011.bin",
+    "transformer.h.34.ln_2.weight": "pytorch_model-00006-of-00011.bin",
+    "transformer.h.34.mlp.c_fc.weight": "pytorch_model-00006-of-00011.bin",
+    "transformer.h.34.mlp.c_proj.weight": "pytorch_model-00006-of-00011.bin",
+    "transformer.h.35.attn.c_attn.weight": "pytorch_model-00006-of-00011.bin",
+    "transformer.h.35.attn.c_proj.weight": "pytorch_model-00006-of-00011.bin",
+    "transformer.h.35.attn.masked_bias": "pytorch_model-00006-of-00011.bin",
+    "transformer.h.35.ln_1.weight": "pytorch_model-00006-of-00011.bin",
+    "transformer.h.35.ln_2.weight": "pytorch_model-00006-of-00011.bin",
+    "transformer.h.35.mlp.c_fc.weight": "pytorch_model-00007-of-00011.bin",
+    "transformer.h.35.mlp.c_proj.weight": "pytorch_model-00007-of-00011.bin",
+    "transformer.h.36.attn.c_attn.weight": "pytorch_model-00007-of-00011.bin",
+    "transformer.h.36.attn.c_proj.weight": "pytorch_model-00007-of-00011.bin",
+    "transformer.h.36.attn.masked_bias": "pytorch_model-00007-of-00011.bin",
+    "transformer.h.36.ln_1.weight": "pytorch_model-00007-of-00011.bin",
+    "transformer.h.36.ln_2.weight": "pytorch_model-00007-of-00011.bin",
+    "transformer.h.36.mlp.c_fc.weight": "pytorch_model-00007-of-00011.bin",
+    "transformer.h.36.mlp.c_proj.weight": "pytorch_model-00007-of-00011.bin",
+    "transformer.h.37.attn.c_attn.weight": "pytorch_model-00007-of-00011.bin",
+    "transformer.h.37.attn.c_proj.weight": "pytorch_model-00007-of-00011.bin",
+    "transformer.h.37.attn.masked_bias": "pytorch_model-00007-of-00011.bin",
+    "transformer.h.37.ln_1.weight": "pytorch_model-00007-of-00011.bin",
+    "transformer.h.37.ln_2.weight": "pytorch_model-00007-of-00011.bin",
+    "transformer.h.37.mlp.c_fc.weight": "pytorch_model-00007-of-00011.bin",
+    "transformer.h.37.mlp.c_proj.weight": "pytorch_model-00007-of-00011.bin",
+    "transformer.h.38.attn.c_attn.weight": "pytorch_model-00007-of-00011.bin",
+    "transformer.h.38.attn.c_proj.weight": "pytorch_model-00007-of-00011.bin",
+    "transformer.h.38.attn.masked_bias": "pytorch_model-00007-of-00011.bin",
+    "transformer.h.38.ln_1.weight": "pytorch_model-00007-of-00011.bin",
+    "transformer.h.38.ln_2.weight": "pytorch_model-00007-of-00011.bin",
+    "transformer.h.38.mlp.c_fc.weight": "pytorch_model-00007-of-00011.bin",
+    "transformer.h.38.mlp.c_proj.weight": "pytorch_model-00007-of-00011.bin",
+    "transformer.h.39.attn.c_attn.weight": "pytorch_model-00007-of-00011.bin",
+    "transformer.h.39.attn.c_proj.weight": "pytorch_model-00007-of-00011.bin",
+    "transformer.h.39.attn.masked_bias": "pytorch_model-00007-of-00011.bin",
+    "transformer.h.39.ln_1.weight": "pytorch_model-00007-of-00011.bin",
+    "transformer.h.39.ln_2.weight": "pytorch_model-00007-of-00011.bin",
+    "transformer.h.39.mlp.c_fc.weight": "pytorch_model-00007-of-00011.bin",
+    "transformer.h.39.mlp.c_proj.weight": "pytorch_model-00007-of-00011.bin",
+    "transformer.h.4.attn.c_attn.weight": "pytorch_model-00001-of-00011.bin",
+    "transformer.h.4.attn.c_proj.weight": "pytorch_model-00001-of-00011.bin",
+    "transformer.h.4.attn.masked_bias": "pytorch_model-00001-of-00011.bin",
+    "transformer.h.4.ln_1.weight": "pytorch_model-00001-of-00011.bin",
+    "transformer.h.4.ln_2.weight": "pytorch_model-00001-of-00011.bin",
+    "transformer.h.4.mlp.c_fc.weight": "pytorch_model-00001-of-00011.bin",
+    "transformer.h.4.mlp.c_proj.weight": "pytorch_model-00001-of-00011.bin",
+    "transformer.h.40.attn.c_attn.weight": "pytorch_model-00007-of-00011.bin",
+    "transformer.h.40.attn.c_proj.weight": "pytorch_model-00007-of-00011.bin",
+    "transformer.h.40.attn.masked_bias": "pytorch_model-00007-of-00011.bin",
+    "transformer.h.40.ln_1.weight": "pytorch_model-00007-of-00011.bin",
+    "transformer.h.40.ln_2.weight": "pytorch_model-00007-of-00011.bin",
+    "transformer.h.40.mlp.c_fc.weight": "pytorch_model-00007-of-00011.bin",
+    "transformer.h.40.mlp.c_proj.weight": "pytorch_model-00007-of-00011.bin",
+    "transformer.h.41.attn.c_attn.weight": "pytorch_model-00007-of-00011.bin",
+    "transformer.h.41.attn.c_proj.weight": "pytorch_model-00007-of-00011.bin",
+    "transformer.h.41.attn.masked_bias": "pytorch_model-00007-of-00011.bin",
+    "transformer.h.41.ln_1.weight": "pytorch_model-00007-of-00011.bin",
+    "transformer.h.41.ln_2.weight": "pytorch_model-00007-of-00011.bin",
+    "transformer.h.41.mlp.c_fc.weight": "pytorch_model-00008-of-00011.bin",
+    "transformer.h.41.mlp.c_proj.weight": "pytorch_model-00008-of-00011.bin",
+    "transformer.h.42.attn.c_attn.weight": "pytorch_model-00008-of-00011.bin",
+    "transformer.h.42.attn.c_proj.weight": "pytorch_model-00008-of-00011.bin",
+    "transformer.h.42.attn.masked_bias": "pytorch_model-00008-of-00011.bin",
+    "transformer.h.42.ln_1.weight": "pytorch_model-00008-of-00011.bin",
+    "transformer.h.42.ln_2.weight": "pytorch_model-00008-of-00011.bin",
+    "transformer.h.42.mlp.c_fc.weight": "pytorch_model-00008-of-00011.bin",
+    "transformer.h.42.mlp.c_proj.weight": "pytorch_model-00008-of-00011.bin",
+    "transformer.h.43.attn.c_attn.weight": "pytorch_model-00008-of-00011.bin",
+    "transformer.h.43.attn.c_proj.weight": "pytorch_model-00008-of-00011.bin",
+    "transformer.h.43.attn.masked_bias": "pytorch_model-00008-of-00011.bin",
+    "transformer.h.43.ln_1.weight": "pytorch_model-00008-of-00011.bin",
+    "transformer.h.43.ln_2.weight": "pytorch_model-00008-of-00011.bin",
+    "transformer.h.43.mlp.c_fc.weight": "pytorch_model-00008-of-00011.bin",
+    "transformer.h.43.mlp.c_proj.weight": "pytorch_model-00008-of-00011.bin",
+    "transformer.h.44.attn.c_attn.weight": "pytorch_model-00008-of-00011.bin",
+    "transformer.h.44.attn.c_proj.weight": "pytorch_model-00008-of-00011.bin",
+    "transformer.h.44.attn.masked_bias": "pytorch_model-00008-of-00011.bin",
+    "transformer.h.44.ln_1.weight": "pytorch_model-00008-of-00011.bin",
+    "transformer.h.44.ln_2.weight": "pytorch_model-00008-of-00011.bin",
+    "transformer.h.44.mlp.c_fc.weight": "pytorch_model-00008-of-00011.bin",
+    "transformer.h.44.mlp.c_proj.weight": "pytorch_model-00008-of-00011.bin",
+    "transformer.h.45.attn.c_attn.weight": "pytorch_model-00008-of-00011.bin",
+    "transformer.h.45.attn.c_proj.weight": "pytorch_model-00008-of-00011.bin",
+    "transformer.h.45.attn.masked_bias": "pytorch_model-00008-of-00011.bin",
+    "transformer.h.45.ln_1.weight": "pytorch_model-00008-of-00011.bin",
+    "transformer.h.45.ln_2.weight": "pytorch_model-00008-of-00011.bin",
+    "transformer.h.45.mlp.c_fc.weight": "pytorch_model-00008-of-00011.bin",
+    "transformer.h.45.mlp.c_proj.weight": "pytorch_model-00008-of-00011.bin",
+    "transformer.h.46.attn.c_attn.weight": "pytorch_model-00008-of-00011.bin",
+    "transformer.h.46.attn.c_proj.weight": "pytorch_model-00008-of-00011.bin",
+    "transformer.h.46.attn.masked_bias": "pytorch_model-00008-of-00011.bin",
+    "transformer.h.46.ln_1.weight": "pytorch_model-00008-of-00011.bin",
+    "transformer.h.46.ln_2.weight": "pytorch_model-00008-of-00011.bin",
+    "transformer.h.46.mlp.c_fc.weight": "pytorch_model-00008-of-00011.bin",
+    "transformer.h.46.mlp.c_proj.weight": "pytorch_model-00008-of-00011.bin",
+    "transformer.h.47.attn.c_attn.weight": "pytorch_model-00008-of-00011.bin",
+    "transformer.h.47.attn.c_proj.weight": "pytorch_model-00008-of-00011.bin",
+    "transformer.h.47.attn.masked_bias": "pytorch_model-00008-of-00011.bin",
+    "transformer.h.47.ln_1.weight": "pytorch_model-00008-of-00011.bin",
+    "transformer.h.47.ln_2.weight": "pytorch_model-00008-of-00011.bin",
+    "transformer.h.47.mlp.c_fc.weight": "pytorch_model-00009-of-00011.bin",
+    "transformer.h.47.mlp.c_proj.weight": "pytorch_model-00009-of-00011.bin",
+    "transformer.h.48.attn.c_attn.weight": "pytorch_model-00009-of-00011.bin",
+    "transformer.h.48.attn.c_proj.weight": "pytorch_model-00009-of-00011.bin",
+    "transformer.h.48.attn.masked_bias": "pytorch_model-00009-of-00011.bin",
+    "transformer.h.48.ln_1.weight": "pytorch_model-00009-of-00011.bin",
+    "transformer.h.48.ln_2.weight": "pytorch_model-00009-of-00011.bin",
+    "transformer.h.48.mlp.c_fc.weight": "pytorch_model-00009-of-00011.bin",
+    "transformer.h.48.mlp.c_proj.weight": "pytorch_model-00009-of-00011.bin",
+    "transformer.h.49.attn.c_attn.weight": "pytorch_model-00009-of-00011.bin",
+    "transformer.h.49.attn.c_proj.weight": "pytorch_model-00009-of-00011.bin",
+    "transformer.h.49.attn.masked_bias": "pytorch_model-00009-of-00011.bin",
+    "transformer.h.49.ln_1.weight": "pytorch_model-00009-of-00011.bin",
+    "transformer.h.49.ln_2.weight": "pytorch_model-00009-of-00011.bin",
+    "transformer.h.49.mlp.c_fc.weight": "pytorch_model-00009-of-00011.bin",
+    "transformer.h.49.mlp.c_proj.weight": "pytorch_model-00009-of-00011.bin",
+    "transformer.h.5.attn.c_attn.weight": "pytorch_model-00001-of-00011.bin",
+    "transformer.h.5.attn.c_proj.weight": "pytorch_model-00001-of-00011.bin",
+    "transformer.h.5.attn.masked_bias": "pytorch_model-00001-of-00011.bin",
+    "transformer.h.5.ln_1.weight": "pytorch_model-00001-of-00011.bin",
+    "transformer.h.5.ln_2.weight": "pytorch_model-00001-of-00011.bin",
+    "transformer.h.5.mlp.c_fc.weight": "pytorch_model-00002-of-00011.bin",
+    "transformer.h.5.mlp.c_proj.weight": "pytorch_model-00002-of-00011.bin",
+    "transformer.h.50.attn.c_attn.weight": "pytorch_model-00009-of-00011.bin",
+    "transformer.h.50.attn.c_proj.weight": "pytorch_model-00009-of-00011.bin",
+    "transformer.h.50.attn.masked_bias": "pytorch_model-00009-of-00011.bin",
+    "transformer.h.50.ln_1.weight": "pytorch_model-00009-of-00011.bin",
+    "transformer.h.50.ln_2.weight": "pytorch_model-00009-of-00011.bin",
+    "transformer.h.50.mlp.c_fc.weight": "pytorch_model-00009-of-00011.bin",
+    "transformer.h.50.mlp.c_proj.weight": "pytorch_model-00009-of-00011.bin",
+    "transformer.h.51.attn.c_attn.weight": "pytorch_model-00009-of-00011.bin",
+    "transformer.h.51.attn.c_proj.weight": "pytorch_model-00009-of-00011.bin",
+    "transformer.h.51.attn.masked_bias": "pytorch_model-00009-of-00011.bin",
+    "transformer.h.51.ln_1.weight": "pytorch_model-00009-of-00011.bin",
+    "transformer.h.51.ln_2.weight": "pytorch_model-00009-of-00011.bin",
+    "transformer.h.51.mlp.c_fc.weight": "pytorch_model-00009-of-00011.bin",
+    "transformer.h.51.mlp.c_proj.weight": "pytorch_model-00009-of-00011.bin",
+    "transformer.h.52.attn.c_attn.weight": "pytorch_model-00009-of-00011.bin",
+    "transformer.h.52.attn.c_proj.weight": "pytorch_model-00009-of-00011.bin",
+    "transformer.h.52.attn.masked_bias": "pytorch_model-00009-of-00011.bin",
+    "transformer.h.52.ln_1.weight": "pytorch_model-00009-of-00011.bin",
+    "transformer.h.52.ln_2.weight": "pytorch_model-00009-of-00011.bin",
+    "transformer.h.52.mlp.c_fc.weight": "pytorch_model-00009-of-00011.bin",
+    "transformer.h.52.mlp.c_proj.weight": "pytorch_model-00009-of-00011.bin",
+    "transformer.h.53.attn.c_attn.weight": "pytorch_model-00009-of-00011.bin",
+    "transformer.h.53.attn.c_proj.weight": "pytorch_model-00009-of-00011.bin",
+    "transformer.h.53.attn.masked_bias": "pytorch_model-00009-of-00011.bin",
+    "transformer.h.53.ln_1.weight": "pytorch_model-00009-of-00011.bin",
+    "transformer.h.53.ln_2.weight": "pytorch_model-00009-of-00011.bin",
+    "transformer.h.53.mlp.c_fc.weight": "pytorch_model-00010-of-00011.bin",
+    "transformer.h.53.mlp.c_proj.weight": "pytorch_model-00010-of-00011.bin",
+    "transformer.h.54.attn.c_attn.weight": "pytorch_model-00010-of-00011.bin",
+    "transformer.h.54.attn.c_proj.weight": "pytorch_model-00010-of-00011.bin",
+    "transformer.h.54.attn.masked_bias": "pytorch_model-00010-of-00011.bin",
+    "transformer.h.54.ln_1.weight": "pytorch_model-00010-of-00011.bin",
+    "transformer.h.54.ln_2.weight": "pytorch_model-00010-of-00011.bin",
+    "transformer.h.54.mlp.c_fc.weight": "pytorch_model-00010-of-00011.bin",
+    "transformer.h.54.mlp.c_proj.weight": "pytorch_model-00010-of-00011.bin",
+    "transformer.h.55.attn.c_attn.weight": "pytorch_model-00010-of-00011.bin",
+    "transformer.h.55.attn.c_proj.weight": "pytorch_model-00010-of-00011.bin",
+    "transformer.h.55.attn.masked_bias": "pytorch_model-00010-of-00011.bin",
+    "transformer.h.55.ln_1.weight": "pytorch_model-00010-of-00011.bin",
+    "transformer.h.55.ln_2.weight": "pytorch_model-00010-of-00011.bin",
+    "transformer.h.55.mlp.c_fc.weight": "pytorch_model-00010-of-00011.bin",
+    "transformer.h.55.mlp.c_proj.weight": "pytorch_model-00010-of-00011.bin",
+    "transformer.h.56.attn.c_attn.weight": "pytorch_model-00010-of-00011.bin",
+    "transformer.h.56.attn.c_proj.weight": "pytorch_model-00010-of-00011.bin",
+    "transformer.h.56.attn.masked_bias": "pytorch_model-00010-of-00011.bin",
+    "transformer.h.56.ln_1.weight": "pytorch_model-00010-of-00011.bin",
+    "transformer.h.56.ln_2.weight": "pytorch_model-00010-of-00011.bin",
+    "transformer.h.56.mlp.c_fc.weight": "pytorch_model-00010-of-00011.bin",
+    "transformer.h.56.mlp.c_proj.weight": "pytorch_model-00010-of-00011.bin",
+    "transformer.h.57.attn.c_attn.weight": "pytorch_model-00010-of-00011.bin",
+    "transformer.h.57.attn.c_proj.weight": "pytorch_model-00010-of-00011.bin",
+    "transformer.h.57.attn.masked_bias": "pytorch_model-00010-of-00011.bin",
+    "transformer.h.57.ln_1.weight": "pytorch_model-00010-of-00011.bin",
+    "transformer.h.57.ln_2.weight": "pytorch_model-00010-of-00011.bin",
+    "transformer.h.57.mlp.c_fc.weight": "pytorch_model-00010-of-00011.bin",
+    "transformer.h.57.mlp.c_proj.weight": "pytorch_model-00010-of-00011.bin",
+    "transformer.h.58.attn.c_attn.weight": "pytorch_model-00010-of-00011.bin",
+    "transformer.h.58.attn.c_proj.weight": "pytorch_model-00010-of-00011.bin",
+    "transformer.h.58.attn.masked_bias": "pytorch_model-00010-of-00011.bin",
+    "transformer.h.58.ln_1.weight": "pytorch_model-00010-of-00011.bin",
+    "transformer.h.58.ln_2.weight": "pytorch_model-00010-of-00011.bin",
+    "transformer.h.58.mlp.c_fc.weight": "pytorch_model-00010-of-00011.bin",
+    "transformer.h.58.mlp.c_proj.weight": "pytorch_model-00010-of-00011.bin",
+    "transformer.h.59.attn.c_attn.weight": "pytorch_model-00010-of-00011.bin",
+    "transformer.h.59.attn.c_proj.weight": "pytorch_model-00010-of-00011.bin",
+    "transformer.h.59.attn.masked_bias": "pytorch_model-00010-of-00011.bin",
+    "transformer.h.59.ln_1.weight": "pytorch_model-00010-of-00011.bin",
+    "transformer.h.59.ln_2.weight": "pytorch_model-00010-of-00011.bin",
+    "transformer.h.59.mlp.c_fc.weight": "pytorch_model-00011-of-00011.bin",
+    "transformer.h.59.mlp.c_proj.weight": "pytorch_model-00011-of-00011.bin",
+    "transformer.h.6.attn.c_attn.weight": "pytorch_model-00002-of-00011.bin",
+    "transformer.h.6.attn.c_proj.weight": "pytorch_model-00002-of-00011.bin",
+    "transformer.h.6.attn.masked_bias": "pytorch_model-00002-of-00011.bin",
+    "transformer.h.6.ln_1.weight": "pytorch_model-00002-of-00011.bin",
+    "transformer.h.6.ln_2.weight": "pytorch_model-00002-of-00011.bin",
+    "transformer.h.6.mlp.c_fc.weight": "pytorch_model-00002-of-00011.bin",
+    "transformer.h.6.mlp.c_proj.weight": "pytorch_model-00002-of-00011.bin",
+    "transformer.h.60.attn.c_attn.weight": "pytorch_model-00011-of-00011.bin",
+    "transformer.h.60.attn.c_proj.weight": "pytorch_model-00011-of-00011.bin",
+    "transformer.h.60.attn.masked_bias": "pytorch_model-00011-of-00011.bin",
+    "transformer.h.60.ln_1.weight": "pytorch_model-00011-of-00011.bin",
+    "transformer.h.60.ln_2.weight": "pytorch_model-00011-of-00011.bin",
+    "transformer.h.60.mlp.c_fc.weight": "pytorch_model-00011-of-00011.bin",
+    "transformer.h.60.mlp.c_proj.weight": "pytorch_model-00011-of-00011.bin",
+    "transformer.h.61.attn.c_attn.weight": "pytorch_model-00011-of-00011.bin",
+    "transformer.h.61.attn.c_proj.weight": "pytorch_model-00011-of-00011.bin",
+    "transformer.h.61.attn.masked_bias": "pytorch_model-00011-of-00011.bin",
+    "transformer.h.61.ln_1.weight": "pytorch_model-00011-of-00011.bin",
+    "transformer.h.61.ln_2.weight": "pytorch_model-00011-of-00011.bin",
+    "transformer.h.61.mlp.c_fc.weight": "pytorch_model-00011-of-00011.bin",
+    "transformer.h.61.mlp.c_proj.weight": "pytorch_model-00011-of-00011.bin",
+    "transformer.h.62.attn.c_attn.weight": "pytorch_model-00011-of-00011.bin",
+    "transformer.h.62.attn.c_proj.weight": "pytorch_model-00011-of-00011.bin",
+    "transformer.h.62.attn.masked_bias": "pytorch_model-00011-of-00011.bin",
+    "transformer.h.62.ln_1.weight": "pytorch_model-00011-of-00011.bin",
+    "transformer.h.62.ln_2.weight": "pytorch_model-00011-of-00011.bin",
+    "transformer.h.62.mlp.c_fc.weight": "pytorch_model-00011-of-00011.bin",
+    "transformer.h.62.mlp.c_proj.weight": "pytorch_model-00011-of-00011.bin",
+    "transformer.h.63.attn.c_attn.weight": "pytorch_model-00011-of-00011.bin",
+    "transformer.h.63.attn.c_proj.weight": "pytorch_model-00011-of-00011.bin",
+    "transformer.h.63.attn.masked_bias": "pytorch_model-00011-of-00011.bin",
+    "transformer.h.63.ln_1.weight": "pytorch_model-00011-of-00011.bin",
+    "transformer.h.63.ln_2.weight": "pytorch_model-00011-of-00011.bin",
+    "transformer.h.63.mlp.c_fc.weight": "pytorch_model-00011-of-00011.bin",
+    "transformer.h.63.mlp.c_proj.weight": "pytorch_model-00011-of-00011.bin",
+    "transformer.h.7.attn.c_attn.weight": "pytorch_model-00002-of-00011.bin",
+    "transformer.h.7.attn.c_proj.weight": "pytorch_model-00002-of-00011.bin",
+    "transformer.h.7.attn.masked_bias": "pytorch_model-00002-of-00011.bin",
+    "transformer.h.7.ln_1.weight": "pytorch_model-00002-of-00011.bin",
+    "transformer.h.7.ln_2.weight": "pytorch_model-00002-of-00011.bin",
+    "transformer.h.7.mlp.c_fc.weight": "pytorch_model-00002-of-00011.bin",
+    "transformer.h.7.mlp.c_proj.weight": "pytorch_model-00002-of-00011.bin",
+    "transformer.h.8.attn.c_attn.weight": "pytorch_model-00002-of-00011.bin",
+    "transformer.h.8.attn.c_proj.weight": "pytorch_model-00002-of-00011.bin",
+    "transformer.h.8.attn.masked_bias": "pytorch_model-00002-of-00011.bin",
+    "transformer.h.8.ln_1.weight": "pytorch_model-00002-of-00011.bin",
+    "transformer.h.8.ln_2.weight": "pytorch_model-00002-of-00011.bin",
+    "transformer.h.8.mlp.c_fc.weight": "pytorch_model-00002-of-00011.bin",
+    "transformer.h.8.mlp.c_proj.weight": "pytorch_model-00002-of-00011.bin",
+    "transformer.h.9.attn.c_attn.weight": "pytorch_model-00002-of-00011.bin",
+    "transformer.h.9.attn.c_proj.weight": "pytorch_model-00002-of-00011.bin",
+    "transformer.h.9.attn.masked_bias": "pytorch_model-00002-of-00011.bin",
+    "transformer.h.9.ln_1.weight": "pytorch_model-00002-of-00011.bin",
+    "transformer.h.9.ln_2.weight": "pytorch_model-00002-of-00011.bin",
+    "transformer.h.9.mlp.c_fc.weight": "pytorch_model-00002-of-00011.bin",
+    "transformer.h.9.mlp.c_proj.weight": "pytorch_model-00002-of-00011.bin",
+    "transformer.ln_f.weight": "pytorch_model-00011-of-00011.bin",
+    "transformer.wte.weight": "pytorch_model-00001-of-00011.bin"
+  }
+}

special_tokens_map.json ADDED Viewed

	@@ -0,0 +1,30 @@

+{
+  "bos_token": {
+    "content": "<s>",
+    "lstrip": false,
+    "normalized": true,
+    "rstrip": false,
+    "single_word": false
+  },
+  "eos_token": {
+    "content": "</s>",
+    "lstrip": false,
+    "normalized": true,
+    "rstrip": false,
+    "single_word": false
+  },
+  "pad_token": {
+    "content": "<pad>",
+    "lstrip": false,
+    "normalized": true,
+    "rstrip": false,
+    "single_word": false
+  },
+  "unk_token": {
+    "content": "<unk>",
+    "lstrip": false,
+    "normalized": true,
+    "rstrip": false,
+    "single_word": false
+  }
+}

tokenization_telechat.py ADDED Viewed

	@@ -0,0 +1,403 @@

+# coding=utf-8
+# Copyright 2022 EleutherAI and the HuggingFace Inc. team. All rights reserved.
+#
+# This code is based on EleutherAI's GPT-NeoX library and the GPT-NeoX
+# and OPT implementations in this library. It has been modified from its
+# original forms to accommodate minor architectural differences compared
+# to GPT-NeoX and OPT used by the Meta AI team that trained the model.
+#
+# Licensed under the Apache License, Version 2.0 (the "License");
+# you may not use this file except in compliance with the License.
+# You may obtain a copy of the License at
+#
+#     http://www.apache.org/licenses/LICENSE-2.0
+#
+# Unless required by applicable law or agreed to in writing, software
+# distributed under the License is distributed on an "AS IS" BASIS,
+# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
+# See the License for the specific language governing permissions and
+# limitations under the License.
+"""Tokenization classes for TELECHAT."""
+import os
+from shutil import copyfile
+from typing import Any, Dict, List, Optional, Tuple
+import sentencepiece as spm
+import re
+from transformers.convert_slow_tokenizer import import_protobuf
+from transformers import AddedToken, PreTrainedTokenizer
+from transformers.utils import logging
+from transformers.tokenization_utils_base import TextInput
+logger = logging.get_logger(__name__)
+VOCAB_FILES_NAMES = {"vocab_file": "tokenizer.model"}
+PRETRAINED_VOCAB_FILES_MAP = {
+    "vocab_file": {},
+    "tokenizer_file": {},
+}
+PRETRAINED_POSITIONAL_EMBEDDINGS_SIZES = {
+    "telechat-tokenizer": 8192,
+}
+SPIECE_UNDERLINE = "▁"
+class TELECHATTokenizer(PreTrainedTokenizer):
+    """
+    Construct a TELECHAT tokenizer. Based on byte-level Byte-Pair-Encoding. The default padding token is unset as there is
+    no padding token in the original model.
+    Args:
+        vocab_file (`str`):
+            Path to the vocabulary file.
+        unk_token (`str` or `tokenizers.AddedToken`, *optional*, defaults to `"<unk>"`):
+            The unknown token. A token that is not in the vocabulary cannot be converted to an ID and is set to be this
+            token instead.
+        bos_token (`str` or `tokenizers.AddedToken`, *optional*, defaults to `"<s>"`):
+            The beginning of sequence token that was used during pretraining. Can be used a sequence classifier token.
+        eos_token (`str` or `tokenizers.AddedToken`, *optional*, defaults to `"</s>"`):
+            The end of sequence token.
+        pad_token (`str` or `tokenizers.AddedToken`, *optional*):
+            A special token used to make arrays of tokens the same size for batching purpose. Will then be ignored by
+            attention mechanisms or loss computation.
+        sp_model_kwargs (`Dict[str, Any]`, `Optional`, *optional*):
+            Will be passed to the `SentencePieceProcessor.__init__()` method. The [Python wrapper for
+            SentencePiece](https://github.com/google/sentencepiece/tree/master/python) can be used, among other things,
+            to set:
+            - `enable_sampling`: Enable subword regularization.
+            - `nbest_size`: Sampling parameters for unigram. Invalid for BPE-Dropout.
+              - `nbest_size = {0,1}`: No sampling is performed.
+              - `nbest_size > 1`: samples from the nbest_size results.
+              - `nbest_size < 0`: assuming that nbest_size is infinite and samples from the all hypothesis (lattice)
+                using forward-filtering-and-backward-sampling algorithm.
+            - `alpha`: Smoothing parameter for unigram sampling, and dropout probability of merge operations for
+              BPE-dropout.
+        add_bos_token (`bool`, *optional*, defaults to `True`):
+            Whether or not to add an `bos_token` at the start of sequences.
+        add_eos_token (`bool`, *optional*, defaults to `False`):
+            Whether or not to add an `eos_token` at the end of sequences.
+        clean_up_tokenization_spaces (`bool`, *optional*, defaults to `False`):
+            Whether or not to cleanup spaces after decoding, cleanup consists in removing potential artifacts like
+            extra spaces.
+        spaces_between_special_tokens (`bool`, *optional*, defaults to `False`):
+            Whether or not to add spaces between special tokens.
+    """
+    vocab_files_names = VOCAB_FILES_NAMES
+    pretrained_vocab_files_map = PRETRAINED_VOCAB_FILES_MAP
+    max_model_input_sizes = PRETRAINED_POSITIONAL_EMBEDDINGS_SIZES
+    model_input_names = ["input_ids", "attention_mask"]
+    def __init__(
+        self,
+        vocab_file,
+        bos_token="<s>",
+        eos_token="</s>",
+        unk_token="<unk>",
+        pad_token=None,
+        sp_model_kwargs: Optional[Dict[str, Any]] = None,
+        add_bos_token=False,
+        add_eos_token=False,
+        clean_up_tokenization_spaces=False,
+        spaces_between_special_tokens=False,
+        **kwargs,
+    ):
+        self.sp_model_kwargs = {} if sp_model_kwargs is None else sp_model_kwargs
+        bos_token = AddedToken(bos_token, normalized=False, special=True) if isinstance(bos_token, str) else bos_token
+        eos_token = AddedToken(eos_token, normalized=False, special=True) if isinstance(eos_token, str) else eos_token
+        pad_token = AddedToken(pad_token, normalized=False, special=True) if isinstance(pad_token, str) else pad_token
+        self.vocab_file = vocab_file
+        self.add_bos_token = add_bos_token
+        self.add_eos_token = add_eos_token
+        self.sp_model = self.get_spm_processor(kwargs.pop("from_slow", False))
+        super().__init__(
+            bos_token=bos_token,
+            eos_token=eos_token,
+            unk_token=unk_token,
+            pad_token=pad_token,
+            add_bos_token=add_bos_token,
+            add_eos_token=add_eos_token,
+            sp_model_kwargs=self.sp_model_kwargs,
+            clean_up_tokenization_spaces=clean_up_tokenization_spaces,
+            spaces_between_special_tokens=spaces_between_special_tokens,
+            **kwargs,
+        )
+    @property
+    def unk_token_length(self):
+        return len(self.sp_model.encode(str(self.unk_token)))
+    # Copied from transformers.models.t5.tokenization_t5.T5Tokenizer.get_spm_processor
+    def get_spm_processor(self, from_slow=False):
+        tokenizer = spm.SentencePieceProcessor(**self.sp_model_kwargs)
+        with open(self.vocab_file, "rb") as f:
+            sp_model = f.read()
+            model_pb2 = import_protobuf(f"The new behaviour of {self.__class__.__name__} (with `self.legacy = False`)")
+            model = model_pb2.ModelProto.FromString(sp_model)
+            normalizer_spec = model_pb2.NormalizerSpec()
+            normalizer_spec.add_dummy_prefix = True
+            model.normalizer_spec.MergeFrom(normalizer_spec)
+            sp_model = model.SerializeToString()
+            tokenizer.LoadFromSerializedProto(sp_model)
+        return tokenizer
+    def __getstate__(self):
+        state = self.__dict__.copy()
+        state["sp_model"] = None
+        state["sp_model_proto"] = self.sp_model.serialized_model_proto()
+        return state
+    def __setstate__(self, d):
+        self.__dict__ = d
+        self.sp_model = spm.SentencePieceProcessor(**self.sp_model_kwargs)
+        self.sp_model.LoadFromSerializedProto(self.sp_model_proto)
+    @property
+    def vocab_size(self):
+        """Returns vocab size"""
+        return self.sp_model.get_piece_size()
+    def get_vocab(self):
+        """Returns vocab as a dict"""
+        vocab = {self.convert_ids_to_tokens(i): i for i in range(self.vocab_size)}
+        vocab.update(self.added_tokens_encoder)
+        return vocab
+    def tokenize(self, text: TextInput, **kwargs) -> List[str]:
+        """
+        Converts a string in a sequence of tokens, using the tokenizer.
+        Split in words for word-based vocabulary or sub-words for sub-word-based vocabularies
+        (BPE/SentencePieces/WordPieces). Takes care of added tokens.
+        Args:
+            text (`str`):
+                The sequence to be encoded.
+            **kwargs (additional keyword arguments):
+                Passed along to the model-specific `prepare_for_tokenization` preprocessing method.
+        Returns:
+            `List[str]`: The list of tokens.
+        """
+        split_special_tokens = kwargs.pop("split_special_tokens", self.split_special_tokens)
+        remove_dummy_prefix = kwargs.pop("remove_dummy_prefix", False)
+        text, kwargs = self.prepare_for_tokenization(text, **kwargs)
+        if kwargs:
+            logger.warning(f"Keyword arguments {kwargs} not recognized.")
+        if hasattr(self, "do_lower_case") and self.do_lower_case:
+            # convert non-special tokens to lowercase. Might be super slow as well?
+            escaped_special_toks = [re.escape(s_tok) for s_tok in (self.all_special_tokens)]
+            escaped_special_toks += [
+                re.escape(s_tok.content)
+                for s_tok in (self._added_tokens_decoder.values())
+                if not s_tok.special and s_tok.normalized
+            ]
+            pattern = r"(" + r"|".join(escaped_special_toks) + r")|" + r"(.+?)"
+            text = re.sub(pattern, lambda m: m.groups()[0] or m.groups()[1].lower(), text)
+        if split_special_tokens:
+            no_split_token = []
+            tokens = [text]
+        else:
+            no_split_token = self._added_tokens_encoder.keys()  # don't split on any of the added tokens
+            # "This is something<special_token_1>  else"
+            tokens = self.tokens_trie.split(text)
+        # ["This is something", "<special_token_1>", "  else"]
+        for i, token in enumerate(tokens):
+            if token in no_split_token:
+                tok_extended = self._added_tokens_decoder.get(self._added_tokens_encoder[token], None)
+                left = tokens[i - 1] if i > 0 else None
+                right = tokens[i + 1] if i < len(tokens) - 1 else None
+                if isinstance(tok_extended, AddedToken):
+                    if tok_extended.rstrip and right:
+                        # A bit counter-intuitive but we strip the left of the string
+                        # since tok_extended.rstrip means the special token is eating all white spaces on its right
+                        tokens[i + 1] = right.lstrip()
+                    # Strip white spaces on the left
+                    if tok_extended.lstrip and left:
+                        tokens[i - 1] = left.rstrip()  # Opposite here
+                    if tok_extended.single_word and left and left[-1] != " ":
+                        tokens[i - 1] += token
+                        tokens[i] = ""
+                    elif tok_extended.single_word and right and right[0] != " ":
+                        tokens[i + 1] = token + tokens[i + 1]
+                        tokens[i] = ""
+                else:
+                    raise ValueError(
+                        f"{tok_extended} cannot be tokenized because it was not properly added"
+                        f" to the tokenizer. This means that it is not an `AddedToken` but a {type(tok_extended)}"
+                    )
+        # ["This is something", "<special_token_1>", "else"]
+        tokenized_text = []
+        for token in tokens:
+            # Need to skip eventual empty (fully stripped) tokens
+            if not token:
+                continue
+            if token in no_split_token:
+                tokenized_text.append(token)
+            else:
+                tokenized_text.extend(self._tokenize(token, remove_dummy_prefix=remove_dummy_prefix))
+        # ["This", " is", " something", "<special_token_1>", "else"]
+        return tokenized_text
+    def _tokenize(self, text, **kwargs):
+        """
+        Returns a tokenized string.
+        We add a option to remove dummpy prefix during tokenization instead of changing the default behaviour of the sentencepiece tokenizer.
+        This is useful when there're two tokenized sentences to be merged into one as the last one will have an extra dummy prefix which results in a
+        inconsistant pattern.
+        """
+        tokens = self.sp_model.encode(text, out_type=str)
+        if text.startswith((SPIECE_UNDERLINE, " ")):
+            return tokens
+        if len(tokens) > 0 and kwargs.get("remove_dummy_prefix") is True:
+            tokens[0] = tokens[0].replace(SPIECE_UNDERLINE, "", 1)
+        return tokens
+    def _convert_token_to_id(self, token):
+        """Converts a token (str) in an id using the vocab."""
+        return self.sp_model.piece_to_id(token)
+    def _convert_id_to_token(self, index):
+        """Converts an index (integer) in a token (str) using the vocab."""
+        token = self.sp_model.IdToPiece(index)
+        return token
+    def convert_tokens_to_string(self, tokens):
+        """Converts a sequence of tokens (string) in a single string."""
+        current_sub_tokens = []
+        out_string = ""
+        # prev_is_special = False
+        for i, token in enumerate(tokens):
+            # make sure that special tokens are not decoded using sentencepiece model
+            if token in self.all_special_tokens:
+                # if not prev_is_special and i != 0 and self.legacy:
+                #     out_string += " "
+                out_string += self.sp_model.decode(current_sub_tokens) + token
+                # prev_is_special = True
+                current_sub_tokens = []
+            else:
+                current_sub_tokens.append(token)
+                # prev_is_special = False
+        out_string += self.sp_model.decode(current_sub_tokens)
+        return out_string
+    def save_vocabulary(self, save_directory, filename_prefix: Optional[str] = None) -> Tuple[str]:
+        """
+        Save the vocabulary and special tokens file to a directory.
+        Args:
+            save_directory (`str`):
+                The directory in which to save the vocabulary.
+        Returns:
+            `Tuple(str)`: Paths to the files saved.
+        """
+        if not os.path.isdir(save_directory):
+            logger.error(f"Vocabulary path ({save_directory}) should be a directory")
+            return
+        out_vocab_file = os.path.join(
+            save_directory, (filename_prefix + "-" if filename_prefix else "") + VOCAB_FILES_NAMES["vocab_file"]
+        )
+        if os.path.abspath(self.vocab_file) != os.path.abspath(out_vocab_file) and os.path.isfile(self.vocab_file):
+            copyfile(self.vocab_file, out_vocab_file)
+        elif not os.path.isfile(self.vocab_file):
+            with open(out_vocab_file, "wb") as fi:
+                content_spiece_model = self.sp_model.serialized_model_proto()
+                fi.write(content_spiece_model)
+        return (out_vocab_file,)
+    def build_inputs_with_special_tokens(self, token_ids_0, token_ids_1=None):
+        bos_token_id = [self.bos_token_id] if self.add_bos_token else []
+        eos_token_id = [self.eos_token_id] if self.add_eos_token else []
+        output = bos_token_id + token_ids_0 + eos_token_id
+        if token_ids_1 is not None:
+            output = output + bos_token_id + token_ids_1 + eos_token_id
+        return output
+    def get_special_tokens_mask(
+        self, token_ids_0: List[int], token_ids_1: Optional[List[int]] = None, already_has_special_tokens: bool = False
+    ) -> List[int]:
+        """
+        Retrieve sequence ids from a token list that has no special tokens added. This method is called when adding
+        special tokens using the tokenizer `prepare_for_model` method.
+        Args:
+            token_ids_0 (`List[int]`):
+                List of IDs.
+            token_ids_1 (`List[int]`, *optional*):
+                Optional second list of IDs for sequence pairs.
+            already_has_special_tokens (`bool`, *optional*, defaults to `False`):
+                Whether or not the token list is already formatted with special tokens for the model.
+        Returns:
+            `List[int]`: A list of integers in the range [0, 1]: 1 for a special token, 0 for a sequence token.
+        """
+        if already_has_special_tokens:
+            return super().get_special_tokens_mask(
+                token_ids_0=token_ids_0, token_ids_1=token_ids_1, already_has_special_tokens=True
+            )
+        bos_token_id = [1] if self.add_bos_token else []
+        eos_token_id = [1] if self.add_eos_token else []
+        if token_ids_1 is None:
+            return bos_token_id + ([0] * len(token_ids_0)) + eos_token_id
+        return (
+            bos_token_id
+            + ([0] * len(token_ids_0))
+            + eos_token_id
+            + bos_token_id
+            + ([0] * len(token_ids_1))
+            + eos_token_id
+        )
+    def create_token_type_ids_from_sequences(
+        self, token_ids_0: List[int], token_ids_1: Optional[List[int]] = None
+    ) -> List[int]:
+        """
+        Creates a mask from the two sequences passed to be used in a sequence-pair classification task. An ALBERT
+        sequence pair mask has the following format:
+        ```
+        0 0 0 0 0 0 0 0 0 0 0 1 1 1 1 1 1 1 1 1
+        | first sequence    | second sequence |
+        ```
+        if token_ids_1 is None, only returns the first portion of the mask (0s).
+        Args:
+            token_ids_0 (`List[int]`):
+                List of ids.
+            token_ids_1 (`List[int]`, *optional*):
+                Optional second list of IDs for sequence pairs.
+        Returns:
+            `List[int]`: List of [token type IDs](../glossary#token-type-ids) according to the given sequence(s).
+        """
+        bos_token_id = [self.bos_token_id] if self.add_bos_token else []
+        eos_token_id = [self.eos_token_id] if self.add_eos_token else []
+        output = [0] * len(bos_token_id + token_ids_0 + eos_token_id)
+        if token_ids_1 is not None:
+            output += [1] * len(bos_token_id + token_ids_1 + eos_token_id)
+        return output

tokenizer.model ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:1e2bf2c2d38bab8a4d7107e36073be27be40a625b2f4e57f5a0609bdb70deed8
+size 1159468

tokenizer_config.json ADDED Viewed

	@@ -0,0 +1,54 @@

+{
+  "add_bos_token": false,
+  "add_eos_token": false,
+  "added_tokens_decoder": {
+    "0": {
+      "content": "<unk>",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "1": {
+      "content": "<s>",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "2": {
+      "content": "</s>",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "3": {
+      "content": "<pad>",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    }
+  },
+  "auto_map": {
+    "AutoTokenizer": [
+      "tokenization_telechat.TELECHATTokenizer",
+      null
+    ]
+  },
+  "bos_token": "<s>",
+  "clean_up_tokenization_spaces": false,
+  "eos_token": "</s>",
+  "model_max_length": 8192,
+  "pad_token": "<pad>",
+  "sp_model_kwargs": {},
+  "spaces_between_special_tokens": false,
+  "tokenizer_class": "TELECHATTokenizer",
+  "unk_token": "<unk>",
+  "use_fast": false
+}