Instructions to use 4cee/raze-v2-gemma3n-e4b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- llama-cpp-python
How to use 4cee/raze-v2-gemma3n-e4b with llama-cpp-python:
# !pip install llama-cpp-python from llama_cpp import Llama llm = Llama.from_pretrained( repo_id="4cee/raze-v2-gemma3n-e4b", filename="raze_v2_gemma3n_e4b.gguf", )
llm.create_chat_completion( messages = "No input example has been defined for this model task." )
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use 4cee/raze-v2-gemma3n-e4b with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf 4cee/raze-v2-gemma3n-e4b # Run inference directly in the terminal: llama cli -hf 4cee/raze-v2-gemma3n-e4b
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf 4cee/raze-v2-gemma3n-e4b # Run inference directly in the terminal: llama cli -hf 4cee/raze-v2-gemma3n-e4b
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf 4cee/raze-v2-gemma3n-e4b # Run inference directly in the terminal: ./llama-cli -hf 4cee/raze-v2-gemma3n-e4b
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf 4cee/raze-v2-gemma3n-e4b # Run inference directly in the terminal: ./build/bin/llama-cli -hf 4cee/raze-v2-gemma3n-e4b
Use Docker
docker model run hf.co/4cee/raze-v2-gemma3n-e4b
- LM Studio
- Jan
- Ollama
How to use 4cee/raze-v2-gemma3n-e4b with Ollama:
ollama run hf.co/4cee/raze-v2-gemma3n-e4b
- Unsloth Studio
How to use 4cee/raze-v2-gemma3n-e4b with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for 4cee/raze-v2-gemma3n-e4b to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for 4cee/raze-v2-gemma3n-e4b to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for 4cee/raze-v2-gemma3n-e4b to start chatting
- Atomic Chat new
- Docker Model Runner
How to use 4cee/raze-v2-gemma3n-e4b with Docker Model Runner:
docker model run hf.co/4cee/raze-v2-gemma3n-e4b
- Lemonade
How to use 4cee/raze-v2-gemma3n-e4b with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull 4cee/raze-v2-gemma3n-e4b
Run and chat with the model
lemonade run user.raze-v2-gemma3n-e4b-{{QUANT_TAG}}List all available models
lemonade list
- Disclaimer 2: I have no idea how to use HuggingFace, Github, or literally anything like that. This was a minor project that I did for fun. All of this was vibe coded with the help of Gemini-3-preview. Please forgive me if anything goes wrong.
This is a custom QLoRA fine-tune of Gemma-3n-E4B-it. It's trained on online conversations of my own friend group, with consent.
Disclaimer: this model is STILL VERY UNSTABLE. Most times it generates half-legible nonsense. Be weary!
On a related note; it will just hallucinate usernames. Or respond as multiple users.
This is the second model I've trained on this dataset. Due to the unfortunate nature of the dataset, it's still weird and stupid. If you want a more stable model still with a distinct personality, use raze-v3-hybrid, or raze-v3-calcium. Unfortunately it does not have the visual capabilities of the base model. I don't know how to keep them and it would require a lot of difficulty trying to make it like that.
Half of the training data was formatted as such:
{"messages": [{"role": "user", "content": ""Below is a chat log. Continue the conversation as [username1]. \n\n[username1]: [message1]\n[username2]:[message2]\n\n"}, {"role": "assistant", "content": "[response message]"}]}
And the other half was formatted like this:
{"messages": [{"role": "user", "content": "Roleplay as [username]. Reply to the following message.\n\n[message]\n\n"}, {"role": "assistant", "content": "[response]"}]}
This way, it can handle both one-on-one conversation, and conversation as a group. If you want the most accurate responses (why would you, it's funnier without it), then use something like that. I think it's best suited as an automated application.
Not much is actually different in terms of the model itself compared to v1. However I think the splitting of the dataset helped to mellow it out a little bit. It should be more legible.
Sorry if this doesn't fit whatever huggingface standards stuff. I see a lot of models split into the 4 different safetensors files and I deleted those. You get ggufs, deal with it i guess???
Disclaimer 2: I have no idea how to use HuggingFace, Github, or literally anything like that. This was a minor project that I did for fun. All of this was vibe coded with the help of Gemini-3-preview. Please forgive me if anything goes wrong.
Disclaimer 3: This model was not trained for the explicit purpose of generating anything harmful or against the Gemma Prohibited Use Policy. I did my best to filter out such content, however it may still be present and/or randomly generated by the model. Please don't sue me Google, oh god.
License and Terms
This model is a derivative of Gemma 3n E4B by Google.
Gemma is provided under and subject to the Gemma Terms of Use found at https://ai.google.dev/gemma/terms.
By using this model, you agree to the Gemma Terms of Use and the Prohibited Use Policy. (https://ai.google.dev/gemma/prohibited_use_policy)
- Downloads last month
- 6
We're not able to determine the quantization variants.