YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Quantization made by Richard Erkhov.

Github

Discord

Request more models

Mistral-7B-Technical-Tutorial-Summarization-QLoRA - GGUF

Name Quant method Size
Mistral-7B-Technical-Tutorial-Summarization-QLoRA.Q2_K.gguf Q2_K 2.53GB
Mistral-7B-Technical-Tutorial-Summarization-QLoRA.IQ3_XS.gguf IQ3_XS 2.81GB
Mistral-7B-Technical-Tutorial-Summarization-QLoRA.IQ3_S.gguf IQ3_S 2.96GB
Mistral-7B-Technical-Tutorial-Summarization-QLoRA.Q3_K_S.gguf Q3_K_S 2.95GB
Mistral-7B-Technical-Tutorial-Summarization-QLoRA.IQ3_M.gguf IQ3_M 3.06GB
Mistral-7B-Technical-Tutorial-Summarization-QLoRA.Q3_K.gguf Q3_K 3.28GB
Mistral-7B-Technical-Tutorial-Summarization-QLoRA.Q3_K_M.gguf Q3_K_M 3.28GB
Mistral-7B-Technical-Tutorial-Summarization-QLoRA.Q3_K_L.gguf Q3_K_L 3.56GB
Mistral-7B-Technical-Tutorial-Summarization-QLoRA.IQ4_XS.gguf IQ4_XS 3.67GB
Mistral-7B-Technical-Tutorial-Summarization-QLoRA.Q4_0.gguf Q4_0 3.83GB
Mistral-7B-Technical-Tutorial-Summarization-QLoRA.IQ4_NL.gguf IQ4_NL 3.87GB
Mistral-7B-Technical-Tutorial-Summarization-QLoRA.Q4_K_S.gguf Q4_K_S 3.86GB
Mistral-7B-Technical-Tutorial-Summarization-QLoRA.Q4_K.gguf Q4_K 4.07GB
Mistral-7B-Technical-Tutorial-Summarization-QLoRA.Q4_K_M.gguf Q4_K_M 4.07GB
Mistral-7B-Technical-Tutorial-Summarization-QLoRA.Q4_1.gguf Q4_1 4.24GB
Mistral-7B-Technical-Tutorial-Summarization-QLoRA.Q5_0.gguf Q5_0 4.65GB
Mistral-7B-Technical-Tutorial-Summarization-QLoRA.Q5_K_S.gguf Q5_K_S 4.65GB
Mistral-7B-Technical-Tutorial-Summarization-QLoRA.Q5_K.gguf Q5_K 4.78GB
Mistral-7B-Technical-Tutorial-Summarization-QLoRA.Q5_K_M.gguf Q5_K_M 4.78GB
Mistral-7B-Technical-Tutorial-Summarization-QLoRA.Q5_1.gguf Q5_1 5.07GB
Mistral-7B-Technical-Tutorial-Summarization-QLoRA.Q6_K.gguf Q6_K 5.53GB
Mistral-7B-Technical-Tutorial-Summarization-QLoRA.Q8_0.gguf Q8_0 7.17GB

Original model description:

language: - en license: apache-2.0 tags: - text-generation-inference - transformers - unsloth - mistral - trl - sft base_model: unsloth/mistral-7b-instruct-v0.2-bnb-4bit

Model Specifications

  • Max Sequence Length: 16384 (with auto support for RoPE Scaling)
  • Data Type: Auto detection, with options for Float16 and Bfloat16
  • Quantization: 4bit, to reduce memory usage

Training Data

Used a private dataset with hundreds of technical tutorials and associated summaries.

Implementation Highlights

  • Efficiency: Emphasis on reducing memory usage and accelerating download speeds through 4bit quantization.
  • Adaptability: Auto detection of data types and support for advanced configuration options like RoPE scaling, LoRA, and gradient checkpointing.

Uploaded Model

  • Developed by: ndebuhr
  • License: apache-2.0
  • Finetuned from model : unsloth/mistral-7b-instruct-v0.2-bnb-4bit

Configuration and Usage

from transformers import AutoModelForCausalLM, AutoTokenizer, pipeline
import torch

input_text = ""

# Set device based on CUDA availability
device = "cuda" if torch.cuda.is_available() else "cpu"

# Load the model and tokenizer
model_name = "ndebuhr/Mistral-7B-Technical-Tutorial-Summarization-QLoRA"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(model_name).to(device)

instruction = "Clarify and summarize this tutorial transcript"
prompt = """{}

### Raw Transcript:
{}

### Summary:
"""

# Tokenize the input text
inputs = tokenizer(
    prompt.format(instruction, input_text),
    return_tensors="pt",
    truncation=True,
    max_length=16384
).to(device)

# Generate outputs
outputs = model.generate(
    **inputs,
    max_length=16384,
    num_return_sequences=1,
    use_cache=True
)

# Decode the generated text
generated_text = tokenizer.batch_decode(outputs, skip_special_tokens=True)

Compute Infrastructure

  • Fine-tuning: used 1xA100 (40GB)
  • Inference: recommend 1xL4 (24GB)

This mistral model was trained 2x faster with Unsloth and Huggingface's TRL library.

Downloads last month
537
GGUF
Model size
7B params
Architecture
llama
Hardware compatibility
Log In to add your hardware

2-bit

3-bit

4-bit

5-bit

6-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support