Instructions to use tahsinahsen/birag-gemma4-e2b-response-only-lora with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use tahsinahsen/birag-gemma4-e2b-response-only-lora with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("tahsinahsen/birag-gemma4-e2b-response-only-lora") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Unsloth Studio
How to use tahsinahsen/birag-gemma4-e2b-response-only-lora with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for tahsinahsen/birag-gemma4-e2b-response-only-lora to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for tahsinahsen/birag-gemma4-e2b-response-only-lora to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for tahsinahsen/birag-gemma4-e2b-response-only-lora to start chatting
Load model with FastModel
pip install unsloth from unsloth import FastModel model, tokenizer = FastModel.from_pretrained( model_name="tahsinahsen/birag-gemma4-e2b-response-only-lora", max_seq_length=2048, ) - MLX LM
How to use tahsinahsen/birag-gemma4-e2b-response-only-lora with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "tahsinahsen/birag-gemma4-e2b-response-only-lora"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "tahsinahsen/birag-gemma4-e2b-response-only-lora" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "tahsinahsen/birag-gemma4-e2b-response-only-lora", "messages": [ {"role": "user", "content": "Hello"} ] }'
BIRAG Gemma 4 E2B โ Response-only LoRA
This repository contains an Unsloth MLX LoRA adapter trained on top of
unsloth/gemma-4-E2B-it.
Training configuration
- Target: response_only
- Maximum sequence length: 8096
- LoRA rank: 16
- LoRA alpha: 16
- Epochs: 3
- Training records: 1790
- Validation records: 225
- Test records: 225
- Seed: 3407
Usage
This is a LoRA adapter, not a standalone merged model. It requires the
unsloth/gemma-4-E2B-it base model and an Unsloth-compatible MLX runtime.
Limitations
The model is intended for research and supportive guidance. It must not be treated as medical, psychological, or clinical advice.
Hardware compatibility
Log In to add your hardware
Quantized