Instructions to use gepardzik/G4-Exp1-26B-A4B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use gepardzik/G4-Exp1-26B-A4B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="gepardzik/G4-Exp1-26B-A4B")# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("gepardzik/G4-Exp1-26B-A4B") model = AutoModelForMultimodalLM.from_pretrained("gepardzik/G4-Exp1-26B-A4B", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use gepardzik/G4-Exp1-26B-A4B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "gepardzik/G4-Exp1-26B-A4B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "gepardzik/G4-Exp1-26B-A4B", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/gepardzik/G4-Exp1-26B-A4B
- SGLang
How to use gepardzik/G4-Exp1-26B-A4B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "gepardzik/G4-Exp1-26B-A4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "gepardzik/G4-Exp1-26B-A4B", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "gepardzik/G4-Exp1-26B-A4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "gepardzik/G4-Exp1-26B-A4B", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use gepardzik/G4-Exp1-26B-A4B with Docker Model Runner:
docker model run hf.co/gepardzik/G4-Exp1-26B-A4B
SFT of google/gemma-4-26B-A4B on 16252928 tokens of pre-2023 texts and merged with llmfan46/G4-MeroMero-26B-A4B-it-uncensored-heretic using task arithmetic.
Observations:
- Still feels like regular gemma
- More lexical flexibility, won't fail when tasked to use cuss words
- Slightly less bullet points and BOLD TEXT
- Unfortunately still prone to use not simple language, but annoying rhetorical figures and “slop punctuation.”
- Due to nature of merged models, it's uncensored by default
- Small drop in MTP acceptance rate on writing due to mentioned style changes
- Tool calling capability is good (at least in Q6_K and better quants)
Use at your own responsibility.
- Downloads last month
- 6
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support