You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

Ri-Gemma-E2B-it-QAT-GGUF

Ri-Gemma-E2B-it-QAT-GGUF is the GGUF release of Ri-Gemma-E2B-IT-QAT-Khasi-Chatbot, optimized for efficient local inference with llama.cpp, Ollama, LM Studio, Jan, Open WebUI, and other GGUF-compatible applications.

The model provides fluent Khasi conversations, instruction following, translation, and reasoning while enabling deployment on consumer hardware through multiple quantization formats.

Quantizations

Quantization Size Recommended Use
Q4_K_M 3.43 GB Best balance between quality and speed. Recommended for most users.
Q6_K 3.85 GB Higher quality with moderate memory requirements.
Q8_0 4.97 GB Near-lossless quantization with excellent response quality.
F16 9.31 GB Full-precision model for maximum quality and research purposes.

Features

  • Native Khasi conversational AI
  • English ↔ Khasi translation
  • Instruction following
  • General knowledge
  • Logical reasoning
  • Writing assistance
  • Local offline inference
  • Compatible with popular GGUF runtimes

Recommended Software

This model can be used with:

  • llama.cpp
  • Ollama
  • LM Studio
  • Jan
  • GPT4All
  • Open WebUI
  • KoboldCpp

Training

This GGUF release is converted from:

Ri-Gemma-E2B-IT-QAT-Khasi-Chatbot

which was instruction-tuned using:

Dataset: toiar/khasi-instruction-response-v2

Dataset Statistics

Category Count
Total Instruction-Response Pairs 77,810
Khasi-Specific Data 72,810
External English Reasoning & General Data 5,000

The dataset combines:

  • Khasi conversations
  • Instruction following
  • Translation
  • Cultural knowledge
  • Educational content
  • General knowledge
  • High-quality reasoning examples

Intended Uses

This model is suitable for:

  • Local AI assistants
  • Offline chatbots
  • Language learning
  • Translation
  • Educational applications
  • Research on Khasi NLP
  • Low-resource language development

Limitations

This model may:

  • Produce incorrect factual information.
  • Make reasoning mistakes on complex tasks.
  • Be sensitive to ambiguous prompts.
  • Reflect biases present in the training data.

Human verification is recommended for important decisions.

Citation

If you use this model in your work, please cite:

@misc{ri_gemma_e2b_it_qat_gguf,
  title        = {Ri-Gemma-E2B-it-QAT-GGUF},
  author       = {Toiar},
  year         = {2026},
  publisher    = {Hugging Face},
  howpublished = {\url{https://huggingface.co/toiar/Ri-Gemma-E2B-it-QAT-GGUF}}
}
Downloads last month
-
GGUF
Model size
5B params
Architecture
gemma4
Hardware compatibility
Log In to add your hardware

4-bit

6-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for toiar/Ri-Gemma-E2B-it-QAT-GGUF