Quant for 4.25

Browse files

Files changed (9) hide show

README.md +79 -51
config.json +40 -0
generation_config.json +6 -0
model.safetensors.index.json +0 -0
original_repo_url.txt +1 -0
output.safetensors +3 -0
special_tokens_map.json +23 -0
tokenizer.json +0 -0
tokenizer_config.json +43 -0

README.md CHANGED Viewed

@@ -7,63 +7,91 @@ datasets:
 - argilla/dpo-mix-7k
 language:
 - en
-quantized_by: bartowski
-pipeline_tag: text-generation
 ---
-## Exllama v2 Quantizations of sparsetral-16x7B-v2-SPIN_iter0
-Using <a href="https://github.com/turboderp/exllamav2/releases/tag/v0.0.13">turboderp's ExLlamaV2 v0.0.13</a> for quantization.
-<b>The "main" branch only contains the measurement.json, download one of the other branches for the model (see below)</b>
-Each branch contains an individual bits per weight, with the main one containing only the meaurement.json for further conversions.
-Original model: https://huggingface.co/serpdotai/sparsetral-16x7B-v2-SPIN_iter0
-| Branch | Bits | lm_head bits | VRAM (4k) | VRAM (16k) | VRAM (32k) | Description |
-| ----- | ---- | ------- | ------ | ------ | ------ | ------------ |
-| [8_0](https://huggingface.co/bartowski/sparsetral-16x7B-v2-SPIN_iter0-exl2/tree/8_0) | 8.0 | 8.0 | 8.3 GB | 9.7 GB | 11.8 GB | Maximum quality that ExLlamaV2 can produce, near unquantized performance. |
-| [6_5](https://huggingface.co/bartowski/sparsetral-16x7B-v2-SPIN_iter0-exl2/tree/6_5) | 6.5 | 8.0 | 7.1 GB | 8.5 GB | 10.6 GB | Very similar to 8.0, good tradeoff of size vs performance, **recommended**. |
-| [5_0](https://huggingface.co/bartowski/sparsetral-16x7B-v2-SPIN_iter0-exl2/tree/5_0) | 5.0 | 6.0 | 5.7 GB | 7.1 GB | 9.2 GB | Slightly lower quality vs 6.5, but usable on 8GB cards. |
-| [4_25](https://huggingface.co/bartowski/sparsetral-16x7B-v2-SPIN_iter0-exl2/tree/4_25) | 4.25 | 6.0 | 5.1 GB | 6.5 GB | 8.6 GB | GPTQ equivalent bits per weight, slightly higher quality. |
-| [3_5](https://huggingface.co/bartowski/sparsetral-16x7B-v2-SPIN_iter0-exl2/tree/3_5) | 3.5 | 6.0 | 4.4 GB | 5.8 GB | 7.9 GB | Lower quality, only use if you have to. |
-## Download instructions
-With git:
-```shell
-git clone --single-branch --branch 6_5 https://huggingface.co/bartowski/sparsetral-16x7B-v2-SPIN_iter0-exl2 sparsetral-16x7B-v2-SPIN_iter0-exl2-6_5
 ```
-With huggingface hub (credit to TheBloke for instructions):
-```shell
-pip3 install huggingface-hub
 ```
-To download the `main` (only useful if you only care about measurement.json) branch to a folder called `sparsetral-16x7B-v2-SPIN_iter0-exl2`:
-```shell
-mkdir sparsetral-16x7B-v2-SPIN_iter0-exl2
-huggingface-cli download bartowski/sparsetral-16x7B-v2-SPIN_iter0-exl2 --local-dir sparsetral-16x7B-v2-SPIN_iter0-exl2 --local-dir-use-symlinks False
 ```
-To download from a different branch, add the `--revision` parameter:
-Linux:
-```shell
-mkdir sparsetral-16x7B-v2-SPIN_iter0-exl2-6_5
-huggingface-cli download bartowski/sparsetral-16x7B-v2-SPIN_iter0-exl2 --revision 6_5 --local-dir sparsetral-16x7B-v2-SPIN_iter0-exl2-6_5 --local-dir-use-symlinks False
-```
-Windows (which apparently doesn't like _ in folders sometimes?):
-```shell
-mkdir sparsetral-16x7B-v2-SPIN_iter0-exl2-6.5
-huggingface-cli download bartowski/sparsetral-16x7B-v2-SPIN_iter0-exl2 --revision 6_5 --local-dir sparsetral-16x7B-v2-SPIN_iter0-exl2-6.5 --local-dir-use-symlinks False
-```
-Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

 - argilla/dpo-mix-7k
 language:
 - en
 ---
+This model is [sparsetral-16x7B-v2](https://huggingface.co/serpdotai/sparsetral-16x7B-v2) further tuned utilizing [SPIN](https://arxiv.org/abs/2401.01335) on [OpenHermes-2.5](https://huggingface.co/datasets/teknium/OpenHermes-2.5) mixed with traditional DPO samples. This is iteration_0, plan to keep making iterations until improvements stop.
+## Training
+- 8x A6000s
+- Base model is [sparsetral-16x7B-v2](https://huggingface.co/serpdotai/sparsetral-16x7B-v2)
+- [Forked version of unsloth](https://github.com/serp-ai/unsloth) for efficient training
+- Sequence Length: 4096
+- Effective batch size: 64
+- Learning Rate: 5e-7 with linear decay (0.1 warmup ratio)
+- Epochs: 2
+- 50k samples (~15k traditional dpo samples, rest SPIN)
+- QLoRA:
+  - 256 r and 256 alpha
+  - ```python
+    target_modules=[
+        "q_proj",
+        "k_proj",
+        "v_proj",
+        "o_proj",
+        "gate_proj",
+        "up_proj",
+        "down_proj",
+        "adapter_down",
+        "adapter_up",
+    ]
+    ```
+## Prompt Format
 ```
+<|im_start|>system\n{message}<|im_end|>\n<|im_start|>user\n{message}<|im_end|>\n<|im_start|>assistant\n
 ```
+## Usage
+```python
+from transformers import AutoModelForCausalLM, AutoTokenizer
+tokenizer = AutoTokenizer.from_pretrained("serpdotai/sparsetral-16x7B-v2-SPIN_iter0", trust_remote_code=True)
+model = AutoModelForCausalLM.from_pretrained("serpdotai/sparsetral-16x7B-v2-SPIN_iter0", device_map="auto", trust_remote_code=True).eval()
+system_str = "<|im_start|>system\n{message}<|im_end|>\n"
+user_str = "<|im_start|>user\n{message}<|im_end|>\n"
+assistant_str = "<|im_start|>assistant\n{message}<|im_end|>\n"
+def construct_prompt(messages):
+    prompt = ""
+    for message in messages:
+        if message["from"] in ["human", "user"]:
+            prompt += user_str.format(
+                message=message["value"]
+            )
+        elif message["from"] in ["gpt", "assistant"]:
+            prompt += assistant_str.format(
+                message=message["value"]
+            )
+        elif message["from"] in ["system", "instruction"]:
+            prompt += system_str.format(
+                message=message["value"]
+            )
+        else:
+            raise ValueError(
+                f"Unknown message type: {message['from']}"
+            )
+    return prompt + "<|im_start|>assistant\n"
+system = "You are a helpful assistant who will help the user to the best of their ability. If you don't know something, say \"I don't know\""
+user = "Are you sentient?"
+messages = [
+    {"from": "system", "value": system},
+    {"from": "user", "value": user},
+]
+prompt = construct_prompt(messages)
+inputs = tokenizer(prompt, return_tensors="pt")
+inputs = inputs.to(model.device)
+pred = model.generate(**inputs, max_length=4096, do_sample=True, top_k=50, top_p=0.99, temperature=0.9, num_return_sequences=1)
+print(tokenizer.decode(pred.cpu()[0], skip_special_tokens=True))
 ```
+## Other Information
+Paper reference: [Parameter-Efficient Sparsity Crafting from Dense to Mixture-of-Experts for Instruction Tuning on General Tasks](https://arxiv.org/abs/2401.02731)
+[Original Paper repo](https://github.com/wuhy68/Parameter-Efficient-MoE)
+[Forked repo with mistral support (sparsetral)](https://github.com/serp-ai/Parameter-Efficient-MoE)
+If you are interested in faster inferencing, check out our [fork of vLLM](https://github.com/serp-ai/vllm) that adds sparsetral support

config.json ADDED Viewed

	@@ -0,0 +1,40 @@

+{
+  "_name_or_path": "serpdotai/sparsetral-16x7B-v2",
+  "adapter_dim": 512,
+  "adapter_dropout": 0.0,
+  "architectures": [
+    "MistralForCausalLM"
+  ],
+  "attention_dropout": 0.0,
+  "auto_map": {
+    "AutoConfig": "serpdotai/sparsetral-16x7B-v2--configuration_sparsetral.SparsetralConfig",
+    "AutoModel": "serpdotai/sparsetral-16x7B-v2--modeling_sparsetral.MistralModel",
+    "AutoModelForCausalLM": "serpdotai/sparsetral-16x7B-v2--modeling_sparsetral.MistralForCausalLM"
+  },
+  "bos_token_id": 1,
+  "eos_token_id": 2,
+  "hidden_act": "silu",
+  "hidden_size": 4096,
+  "initializer_range": 0.02,
+  "intermediate_size": 14336,
+  "max_position_embeddings": 32768,
+  "model_type": "mistral",
+  "moe_dtype": "bfloat16",
+  "moe_scaling": 1,
+  "num_attention_heads": 32,
+  "num_experts": 16,
+  "num_hidden_layers": 32,
+  "num_key_value_heads": 8,
+  "output_router_logits": false,
+  "pretraining_tp": 1,
+  "rms_norm_eps": 1e-05,
+  "rope_theta": 1000000.0,
+  "router_aux_loss_coef": 0.01,
+  "sliding_window": null,
+  "tie_word_embeddings": false,
+  "topk": 4,
+  "torch_dtype": "bfloat16",
+  "transformers_version": "4.37.2",
+  "use_cache": true,
+  "vocab_size": 32000
+}

generation_config.json ADDED Viewed

	@@ -0,0 +1,6 @@

+{
+  "_from_model_config": true,
+  "bos_token_id": 1,
+  "eos_token_id": 2,
+  "transformers_version": "4.37.2"
+}

model.safetensors.index.json ADDED Viewed

The diff for this file is too large to render. See raw diff

original_repo_url.txt ADDED Viewed

	@@ -0,0 +1 @@


1	+ https://huggingface.co/serpdotai/sparsetral-16x7B-v2-SPIN_iter0

output.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:948bc95778d72124823a490e30b46e015d416de8da01192f10e5ba3cc548b8a8
+size 4074366700

special_tokens_map.json ADDED Viewed

	@@ -0,0 +1,23 @@

+{
+  "bos_token": {
+    "content": "<s>",
+    "lstrip": false,
+    "normalized": false,
+    "rstrip": false,
+    "single_word": false
+  },
+  "eos_token": {
+    "content": "</s>",
+    "lstrip": false,
+    "normalized": false,
+    "rstrip": false,
+    "single_word": false
+  },
+  "unk_token": {
+    "content": "<unk>",
+    "lstrip": false,
+    "normalized": false,
+    "rstrip": false,
+    "single_word": false
+  }
+}

tokenizer.json ADDED Viewed

The diff for this file is too large to render. See raw diff

tokenizer_config.json ADDED Viewed

	@@ -0,0 +1,43 @@

+{
+  "add_bos_token": true,
+  "add_eos_token": false,
+  "added_tokens_decoder": {
+    "0": {
+      "content": "<unk>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "1": {
+      "content": "<s>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "2": {
+      "content": "</s>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    }
+  },
+  "additional_special_tokens": [],
+  "bos_token": "<s>",
+  "chat_template": "{{ bos_token }}{% for message in messages %}{% if (message['role'] == 'user') != (loop.index0 % 2 == 0) %}{{ raise_exception('Conversation roles must alternate user/assistant/user/assistant/...') }}{% endif %}{% if message['role'] == 'user' %}{{ '[INST] ' + message['content'] + ' [/INST]' }}{% elif message['role'] == 'assistant' %}{{ message['content'] + eos_token}}{% else %}{{ raise_exception('Only user and assistant roles are supported!') }}{% endif %}{% endfor %}",
+  "clean_up_tokenization_spaces": false,
+  "eos_token": "</s>",
+  "legacy": true,
+  "model_max_length": 1000000000000000019884624838656,
+  "pad_token": null,
+  "sp_model_kwargs": {},
+  "spaces_between_special_tokens": false,
+  "tokenizer_class": "LlamaTokenizer",
+  "unk_token": "<unk>",
+  "use_default_system_prompt": false
+}