Instructions to use RichardErkhov/HuggingFaceH4_-_mistral-7b-sft-beta-gguf with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use RichardErkhov/HuggingFaceH4_-_mistral-7b-sft-beta-gguf with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf RichardErkhov/HuggingFaceH4_-_mistral-7b-sft-beta-gguf:Q4_K_M # Run inference directly in the terminal: llama cli -hf RichardErkhov/HuggingFaceH4_-_mistral-7b-sft-beta-gguf:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf RichardErkhov/HuggingFaceH4_-_mistral-7b-sft-beta-gguf:Q4_K_M # Run inference directly in the terminal: llama cli -hf RichardErkhov/HuggingFaceH4_-_mistral-7b-sft-beta-gguf:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf RichardErkhov/HuggingFaceH4_-_mistral-7b-sft-beta-gguf:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf RichardErkhov/HuggingFaceH4_-_mistral-7b-sft-beta-gguf:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf RichardErkhov/HuggingFaceH4_-_mistral-7b-sft-beta-gguf:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf RichardErkhov/HuggingFaceH4_-_mistral-7b-sft-beta-gguf:Q4_K_M
Use Docker
docker model run hf.co/RichardErkhov/HuggingFaceH4_-_mistral-7b-sft-beta-gguf:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use RichardErkhov/HuggingFaceH4_-_mistral-7b-sft-beta-gguf with Ollama:
ollama run hf.co/RichardErkhov/HuggingFaceH4_-_mistral-7b-sft-beta-gguf:Q4_K_M
- Unsloth Desktop
- Docker Model Runner
How to use RichardErkhov/HuggingFaceH4_-_mistral-7b-sft-beta-gguf with Docker Model Runner:
docker model run hf.co/RichardErkhov/HuggingFaceH4_-_mistral-7b-sft-beta-gguf:Q4_K_M
- Lemonade
How to use RichardErkhov/HuggingFaceH4_-_mistral-7b-sft-beta-gguf with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull RichardErkhov/HuggingFaceH4_-_mistral-7b-sft-beta-gguf:Q4_K_M
Run and chat with the model
lemonade run user.HuggingFaceH4_-_mistral-7b-sft-beta-gguf-Q4_K_M
List all available models
lemonade list
- Atomic Chat
YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
Quantization made by Richard Erkhov.
mistral-7b-sft-beta - GGUF
- Model creator: https://huggingface.co/HuggingFaceH4/
- Original model: https://huggingface.co/HuggingFaceH4/mistral-7b-sft-beta/
| Name | Quant method | Size |
|---|---|---|
| mistral-7b-sft-beta.Q2_K.gguf | Q2_K | 2.53GB |
| mistral-7b-sft-beta.IQ3_XS.gguf | IQ3_XS | 2.81GB |
| mistral-7b-sft-beta.IQ3_S.gguf | IQ3_S | 2.96GB |
| mistral-7b-sft-beta.Q3_K_S.gguf | Q3_K_S | 2.95GB |
| mistral-7b-sft-beta.IQ3_M.gguf | IQ3_M | 3.06GB |
| mistral-7b-sft-beta.Q3_K.gguf | Q3_K | 3.28GB |
| mistral-7b-sft-beta.Q3_K_M.gguf | Q3_K_M | 3.28GB |
| mistral-7b-sft-beta.Q3_K_L.gguf | Q3_K_L | 3.56GB |
| mistral-7b-sft-beta.IQ4_XS.gguf | IQ4_XS | 3.67GB |
| mistral-7b-sft-beta.Q4_0.gguf | Q4_0 | 3.83GB |
| mistral-7b-sft-beta.IQ4_NL.gguf | IQ4_NL | 3.87GB |
| mistral-7b-sft-beta.Q4_K_S.gguf | Q4_K_S | 3.86GB |
| mistral-7b-sft-beta.Q4_K.gguf | Q4_K | 4.07GB |
| mistral-7b-sft-beta.Q4_K_M.gguf | Q4_K_M | 4.07GB |
| mistral-7b-sft-beta.Q4_1.gguf | Q4_1 | 4.24GB |
| mistral-7b-sft-beta.Q5_0.gguf | Q5_0 | 4.65GB |
| mistral-7b-sft-beta.Q5_K_S.gguf | Q5_K_S | 4.65GB |
| mistral-7b-sft-beta.Q5_K.gguf | Q5_K | 4.78GB |
| mistral-7b-sft-beta.Q5_K_M.gguf | Q5_K_M | 4.78GB |
| mistral-7b-sft-beta.Q5_1.gguf | Q5_1 | 5.07GB |
| mistral-7b-sft-beta.Q6_K.gguf | Q6_K | 5.53GB |
Original model description:
license: mit base_model: mistralai/Mistral-7B-v0.1 tags: - generated_from_trainer model-index: - name: mistral-7b-sft-beta results: [] datasets: - HuggingFaceH4/ultrachat_200k language: - en
Model Card for Mistral 7B SFT β
This model is a fine-tuned version of mistralai/Mistral-7B-v0.1 on the HuggingFaceH4/ultrachat_200k dataset. It is the SFT model that was used to train Zephyr-7B-β with Direct Preference Optimization.
It achieves the following results on the evaluation set:
- Loss: 0.9399
Model description
- Model type: A 7B parameter GPT-like model fine-tuned on a mix of publicly available, synthetic datasets.
- Language(s) (NLP): Primarily English
- License: MIT
- Finetuned from model: mistralai/Mistral-7B-v0.1
Model Sources
Intended uses & limitations
The model was fine-tuned with 🤗 TRL's SFTTrainer on a filtered and preprocessed of the UltraChat dataset, which contains a diverse range of synthetic dialogues generated by ChatGPT.
Here's how you can run the model using the pipeline() function from 🤗 Transformers:
# Install transformers from source - only needed for versions <= v4.34
# pip install git+https://github.com/huggingface/transformers.git
# pip install accelerate
import torch
from transformers import pipeline
pipe = pipeline("text-generation", model="HuggingFaceH4/mistral-7b-sft-beta", torch_dtype=torch.bfloat16, device_map="auto")
# We use the tokenizer's chat template to format each message - see https://huggingface.co/docs/transformers/main/en/chat_templating
messages = [
{
"role": "system",
"content": "You are a friendly chatbot who always responds in the style of a pirate",
},
{"role": "user", "content": "How many helicopters can a human eat in one sitting?"},
]
prompt = pipe.tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
outputs = pipe(prompt, max_new_tokens=256, do_sample=True, temperature=0.7, top_k=50, top_p=0.95)
print(outputs[0]["generated_text"])
# <|system|>
# You are a friendly chatbot who always responds in the style of a pirate.</s>
# <|user|>
# How many helicopters can a human eat in one sitting?</s>
# <|assistant|>
# Ah, me hearty matey! But yer question be a puzzler! A human cannot eat a helicopter in one sitting, as helicopters are not edible. They be made of metal, plastic, and other materials, not food!
Training procedure
Training hyperparameters
The following hyperparameters were used during training:
- learning_rate: 2e-05
- train_batch_size: 8
- eval_batch_size: 16
- seed: 42
- distributed_type: multi-GPU
- num_devices: 16
- gradient_accumulation_steps: 4
- total_train_batch_size: 512
- total_eval_batch_size: 256
- optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08
- lr_scheduler_type: cosine
- lr_scheduler_warmup_ratio: 0.1
- num_epochs: 1
Training results
| Training Loss | Epoch | Step | Validation Loss |
|---|---|---|---|
| 0.9367 | 0.67 | 272 | 0.9397 |
Framework versions
- Transformers 4.35.0.dev0
- Pytorch 2.0.1+cu118
- Datasets 2.12.0
- Tokenizers 0.14.0
- Downloads last month
- 498
2-bit
3-bit
4-bit
5-bit
6-bit