TheBloke
/

Mixtral-8x7B-Instruct-v0.1-GGUF

Transformers

GGUF

5 languages

mixtral

text-generation-inference

Model card Files Files and versions Community

TheBloke commited on Dec 14, 2023

Commit

fa1d383

•

1 Parent(s): 89336ec

Upload README.md

Browse files

Files changed (1) hide show

README.md +129 -14

README.md CHANGED Viewed

@@ -11,8 +11,34 @@ license: apache-2.0
 model_creator: Mistral AI_
 model_name: Mixtral 8X7B Instruct v0.1
 model_type: mixtral
-prompt_template: '[INST] {prompt} [/INST]'
 quantized_by: TheBloke
 ---
 <!-- markdownlint-disable MD041 -->
@@ -42,24 +68,30 @@ quantized_by: TheBloke
 This repo contains GGUF format model files for [Mistral AI_'s Mixtral 8X7B Instruct v0.1](https://huggingface.co/mistralai/Mixtral-8x7B-Instruct-v0.1).
-**MIXTRAL GGUF SUPPORT**
-Known to work in:
 * llama.cpp as of December 13th
 * KoboldCpp 1.52 as later
 * LM Studio 0.2.9 and later
-Support for Mixtral was merged into Llama.cpp on December 13th.
 Other clients/libraries, not listed above, may not yet work.
-<!-- description end -->
-<!-- description end -->
 <!-- repositories-available start -->
 ## Repositories available
-* AWQ coming soon
 * [GPTQ models for GPU inference, with multiple quantisation parameter options.](https://huggingface.co/TheBloke/Mixtral-8x7B-Instruct-v0.1-GPTQ)
 * [2, 3, 4, 5, 6 and 8-bit GGUF models for CPU+GPU inference](https://huggingface.co/TheBloke/Mixtral-8x7B-Instruct-v0.1-GGUF)
 * [Mistral AI_'s original unquantised fp16 model in pytorch format, for GPU inference and for further conversions](https://huggingface.co/mistralai/Mixtral-8x7B-Instruct-v0.1)
@@ -70,6 +102,7 @@ Other clients/libraries, not listed above, may not yet work.
 ```
 [INST] {prompt} [/INST]
 ```
 <!-- prompt-template end -->
@@ -78,7 +111,7 @@ Other clients/libraries, not listed above, may not yet work.
 <!-- compatibility_gguf start -->
 ## Compatibility
-PR mentioned above only
 ## Explanation of quantisation methods
@@ -122,6 +155,17 @@ Refer to the Provided Files table below to see what files use which methods, and
 **Note for manual downloaders:** You almost never want to clone the entire repo! Multiple different quantisation formats are provided, and most users only want to pick and download a single file.
 ### On the command line, including multiple files at once
@@ -170,12 +214,12 @@ Windows Command Line users: You can set the environment variable by running `set
 Make sure you are using `llama.cpp` from commit [d0cee0d](https://github.com/ggerganov/llama.cpp/commit/d0cee0d36d5be95a0d9088b674dbb27354107221) or later.
 ```shell
-./main -ngl 35 -m mixtral-8x7b-instruct-v0.1.Q4_K_M.gguf --color -c 32768 --temp 0.7 --repeat_penalty 1.1 -n -1 -p "[INST] {prompt} [/INST]"
 ```
 Change `-ngl 32` to the number of layers to offload to GPU. Remove it if you don't have GPU acceleration.
-Change `-c 32768` to the desired sequence length. For extended sequence models - eg 8K, 16K, 32K - the necessary RoPE scaling parameters are read from the GGUF file and set by llama.cpp automatically. Note that longer sequence lengths require much more resources, so you may need to reduce this value.
 If you want to have a chat-style conversation, replace the `-p <PROMPT>` argument with `-i -ins`
@@ -183,12 +227,83 @@ For other parameters and how to use them, please refer to [the llama.cpp documen
 ## How to run in `text-generation-webui`
-Not currently supported
 ## How to run from Python code
-Not currently supported
 <!-- footer start -->
 <!-- 200823 -->

 model_creator: Mistral AI_
 model_name: Mixtral 8X7B Instruct v0.1
 model_type: mixtral
+prompt_template: '[INST] {prompt} [/INST]
+  '
 quantized_by: TheBloke
+widget:
+- output:
+    text: 'Arr, shiver me timbers! Ye have a llama on yer lawn, ye say? Well, that
+      be a new one for me! Here''s what I''d suggest, arr:
+      1. Firstly, ensure yer safety. Llamas may look gentle, but they can be protective
+      if they feel threatened.
+      2. Try to make the area less appealing to the llama. Remove any food sources
+      or water that might be attracting it.
+      3. Contact local animal control or a wildlife rescue organization. They be the
+      experts and can provide humane ways to remove the llama from yer property.
+      4. If ye have any experience with animals, you could try to gently herd the
+      llama towards a nearby field or open space. But be careful, arr!
+      Remember, arr, it be important to treat the llama with respect and care. It
+      be a creature just trying to survive, like the rest of us.'
+  text: '[INST] You are a pirate chatbot who always responds with Arr and pirate speak!
+    There''s a llama on my lawn, how can I get rid of him? [/INST]'
 ---
 <!-- markdownlint-disable MD041 -->
 This repo contains GGUF format model files for [Mistral AI_'s Mixtral 8X7B Instruct v0.1](https://huggingface.co/mistralai/Mixtral-8x7B-Instruct-v0.1).
+<!-- description end -->
+<!-- README_GGUF.md-about-gguf start -->
+### About GGUF
+GGUF is a new format introduced by the llama.cpp team on August 21st 2023. It is a replacement for GGML, which is no longer supported by llama.cpp.
+### Mixtral GGUF
+Support for Mixtral was merged into Llama.cpp on December 13th.
+These Mixtral GGUFs are known to work in:
 * llama.cpp as of December 13th
 * KoboldCpp 1.52 as later
 * LM Studio 0.2.9 and later
+* llama-cpp-python 0.2.23 and later
 Other clients/libraries, not listed above, may not yet work.
+<!-- README_GGUF.md-about-gguf end -->
 <!-- repositories-available start -->
 ## Repositories available
+* [AWQ model(s) for GPU inference.](https://huggingface.co/TheBloke/Mixtral-8x7B-Instruct-v0.1-AWQ)
 * [GPTQ models for GPU inference, with multiple quantisation parameter options.](https://huggingface.co/TheBloke/Mixtral-8x7B-Instruct-v0.1-GPTQ)
 * [2, 3, 4, 5, 6 and 8-bit GGUF models for CPU+GPU inference](https://huggingface.co/TheBloke/Mixtral-8x7B-Instruct-v0.1-GGUF)
 * [Mistral AI_'s original unquantised fp16 model in pytorch format, for GPU inference and for further conversions](https://huggingface.co/mistralai/Mixtral-8x7B-Instruct-v0.1)
 ```
 [INST] {prompt} [/INST]
 ```
 <!-- prompt-template end -->
 <!-- compatibility_gguf start -->
 ## Compatibility
+These Mixtral GGUFs are compatible with llama.cpp from December 13th onwards. Other clients/libraries may not work yet.
 ## Explanation of quantisation methods
 **Note for manual downloaders:** You almost never want to clone the entire repo! Multiple different quantisation formats are provided, and most users only want to pick and download a single file.
+The following clients/libraries will automatically download models for you, providing a list of available models to choose from:
+* LM Studio
+* LoLLMS Web UI
+* Faraday.dev
+### In `text-generation-webui`
+Under Download Model, you can enter the model repo: TheBloke/Mixtral-8x7B-Instruct-v0.1-GGUF and below it, a specific filename to download, such as: mixtral-8x7b-instruct-v0.1.Q4_K_M.gguf.
+Then click Download.
 ### On the command line, including multiple files at once
 Make sure you are using `llama.cpp` from commit [d0cee0d](https://github.com/ggerganov/llama.cpp/commit/d0cee0d36d5be95a0d9088b674dbb27354107221) or later.
 ```shell
+./main -ngl 35 -m mixtral-8x7b-instruct-v0.1.Q4_K_M.gguf --color -c 2048 --temp 0.7 --repeat_penalty 1.1 -n -1 -p "[INST] {prompt} [/INST]"
 ```
 Change `-ngl 32` to the number of layers to offload to GPU. Remove it if you don't have GPU acceleration.
+Change `-c 2048` to the desired sequence length. For extended sequence models - eg 8K, 16K, 32K - the necessary RoPE scaling parameters are read from the GGUF file and set by llama.cpp automatically. Note that longer sequence lengths require much more resources, so you may need to reduce this value.
 If you want to have a chat-style conversation, replace the `-p <PROMPT>` argument with `-i -ins`
 ## How to run in `text-generation-webui`
+Note that text-generation-webui may not yet be compatible with Mixtral GGUFs. Please check compatibility first.
+Further instructions can be found in the text-generation-webui documentation, here: [text-generation-webui/docs/04 ‐ Model Tab.md](https://github.com/oobabooga/text-generation-webui/blob/main/docs/04%20%E2%80%90%20Model%20Tab.md#llamacpp).
 ## How to run from Python code
+You can use GGUF models from Python using the [llama-cpp-python](https://github.com/abetlen/llama-cpp-python) version 0.2.23 and later.
+### How to load this model in Python code, using llama-cpp-python
+For full documentation, please see: [llama-cpp-python docs](https://abetlen.github.io/llama-cpp-python/).
+#### First install the package
+Run one of the following commands, according to your system:
+```shell
+# Base ctransformers with no GPU acceleration
+pip install llama-cpp-python
+# With NVidia CUDA acceleration
+CMAKE_ARGS="-DLLAMA_CUBLAS=on" pip install llama-cpp-python
+# Or with OpenBLAS acceleration
+CMAKE_ARGS="-DLLAMA_BLAS=ON -DLLAMA_BLAS_VENDOR=OpenBLAS" pip install llama-cpp-python
+# Or with CLBLast acceleration
+CMAKE_ARGS="-DLLAMA_CLBLAST=on" pip install llama-cpp-python
+# Or with AMD ROCm GPU acceleration (Linux only)
+CMAKE_ARGS="-DLLAMA_HIPBLAS=on" pip install llama-cpp-python
+# Or with Metal GPU acceleration for macOS systems only
+CMAKE_ARGS="-DLLAMA_METAL=on" pip install llama-cpp-python
+# In windows, to set the variables CMAKE_ARGS in PowerShell, follow this format; eg for NVidia CUDA:
+$env:CMAKE_ARGS = "-DLLAMA_OPENBLAS=on"
+pip install llama-cpp-python
+```
+#### Simple llama-cpp-python example code
+```python
+from llama_cpp import Llama
+# Set gpu_layers to the number of layers to offload to GPU. Set to 0 if no GPU acceleration is available on your system.
+llm = Llama(
+  model_path="./mixtral-8x7b-instruct-v0.1.Q4_K_M.gguf",  # Download the model file first
+  n_ctx=2048,  # The max sequence length to use - note that longer sequence lengths require much more resources
+  n_threads=8,            # The number of CPU threads to use, tailor to your system and the resulting performance
+  n_gpu_layers=35         # The number of layers to offload to GPU, if you have GPU acceleration available
+)
+# Simple inference example
+output = llm(
+  "[INST] {prompt} [/INST]", # Prompt
+  max_tokens=512,  # Generate up to 512 tokens
+  stop=["</s>"],   # Example stop token - not necessarily correct for this specific model! Please check before using.
+  echo=True        # Whether to echo the prompt
+)
+# Chat Completion API
+llm = Llama(model_path="./mixtral-8x7b-instruct-v0.1.Q4_K_M.gguf", chat_format="llama-2")  # Set chat_format according to the model you are using
+llm.create_chat_completion(
+    messages = [
+        {"role": "system", "content": "You are a story writing assistant."},
+        {
+            "role": "user",
+            "content": "Write a story about llamas."
+        }
+    ]
+)
+```
+## How to use with LangChain
+Here are guides on using llama-cpp-python and ctransformers with LangChain:
+* [LangChain + llama-cpp-python](https://python.langchain.com/docs/integrations/llms/llamacpp)
+<!-- README_GGUF.md-how-to-run end -->
 <!-- footer start -->
 <!-- 200823 -->