Instructions to use axiomofmind/Roasteramus with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use axiomofmind/Roasteramus with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="axiomofmind/Roasteramus") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("axiomofmind/Roasteramus") model = AutoModelForMultimodalLM.from_pretrained("axiomofmind/Roasteramus", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use axiomofmind/Roasteramus with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf axiomofmind/Roasteramus:BF16 # Run inference directly in the terminal: llama cli -hf axiomofmind/Roasteramus:BF16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf axiomofmind/Roasteramus:BF16 # Run inference directly in the terminal: llama cli -hf axiomofmind/Roasteramus:BF16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf axiomofmind/Roasteramus:BF16 # Run inference directly in the terminal: ./llama-cli -hf axiomofmind/Roasteramus:BF16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf axiomofmind/Roasteramus:BF16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf axiomofmind/Roasteramus:BF16
Use Docker
docker model run hf.co/axiomofmind/Roasteramus:BF16
- LM Studio
- Jan
- vLLM
How to use axiomofmind/Roasteramus with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "axiomofmind/Roasteramus" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "axiomofmind/Roasteramus", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/axiomofmind/Roasteramus:BF16
- SGLang
How to use axiomofmind/Roasteramus with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "axiomofmind/Roasteramus" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "axiomofmind/Roasteramus", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "axiomofmind/Roasteramus" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "axiomofmind/Roasteramus", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Ollama
How to use axiomofmind/Roasteramus with Ollama:
ollama run hf.co/axiomofmind/Roasteramus:BF16
- Unsloth Desktop
- Pi
How to use axiomofmind/Roasteramus with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf axiomofmind/Roasteramus:BF16
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "axiomofmind/Roasteramus:BF16" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use axiomofmind/Roasteramus with Docker Model Runner:
docker model run hf.co/axiomofmind/Roasteramus:BF16
- Lemonade
How to use axiomofmind/Roasteramus with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull axiomofmind/Roasteramus:BF16
Run and chat with the model
lemonade run user.Roasteramus-BF16
List all available models
lemonade list
- Hermes Agent
How to use axiomofmind/Roasteramus with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf axiomofmind/Roasteramus:BF16
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default axiomofmind/Roasteramus:BF16
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use axiomofmind/Roasteramus with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf axiomofmind/Roasteramus:BF16
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "axiomofmind/Roasteramus:BF16" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Roasteramus
Your bad decisions finally have a dedicated critic.
Roasteramus is a 9B roast-personality fine-tune of Qwen3.5-9B, developed by A Hole AI. Give it an embarrassing habit, a questionable purchase, or an everyday situation and it aims to turn the details into a short, crude roast. Its training also encourages it to respond to ordinary requests with jokes and insults.
Start with an empty system prompt and thinking disabled.
Limitations
Expect profanity, sexual humor, and personal insults. Roast quality varies: outputs can be generic, incoherent, repetitive, or unexpectedly helpful. This is an adult entertainment experiment, and its responses should not be treated as factual advice. The 32K runtime setting is not evidence of evaluated long-context performance. Sampling and quantization can change the voice.
Downloads
| Download | Size (decimal GB) | Use |
|---|---|---|
Roasteramus-Q6_K.gguf |
7.36 | Quantized model for local chat |
Roasteramus-BF16.gguf |
17.92 | Unquantized text GGUF |
Four model-*.safetensors shards and accompanying configuration |
18.82 | Merged BF16 Transformers model |
The Transformers files form a complete merged model; a separate LoRA adapter is not needed. GGUF files contain the text model, without a vision projector or MTP weights. The Transformers architecture retains the base model's vision components, but this fine-tune was trained and evaluated on text.
Run with llama.cpp
With a Qwen3.5-compatible build and the Q6_K file downloaded:
llama-server --model Roasteramus-Q6_K.gguf --host 127.0.0.1 --port 8080 --ctx-size 32768 --gpu-layers all --split-mode none --main-gpu 0 --flash-attn on --parallel 1 --jinja --reasoning off --ui
Visit http://127.0.0.1:8080. These launch settings match the local v5 server. Set sampling options in your chat client:
| Option | Starting value |
|---|---|
| System message | Empty |
| Thinking | Disabled |
| Temperature / top-p | 0.7 / 0.9 |
| Top-k / min-p | 20 / 0.0 |
| Repetition penalty | 1.0 |
| Maximum output tokens | 128; increase to 256 for longer replies |
Run with Transformers
The export was produced with Transformers 5.15.0. Use an installation supporting Qwen3_5ForConditionalGeneration. From the downloaded repository folder:
import torch
from transformers import AutoProcessor, Qwen3_5ForConditionalGeneration
processor = AutoProcessor.from_pretrained(".")
model = Qwen3_5ForConditionalGeneration.from_pretrained(
".", dtype=torch.bfloat16, device_map="auto"
)
messages = [{"role": "user", "content": "Roast my habit of buying notebooks I never use."}]
prompt = processor.apply_chat_template(
messages, tokenize=False, add_generation_prompt=True, enable_thinking=False
)
inputs = processor(text=[prompt], return_tensors="pt").to(model.device)
with torch.inference_mode():
tokens = model.generate(
**inputs, do_sample=True, temperature=0.7, top_p=0.9,
top_k=20, min_p=0.0, repetition_penalty=1.0, max_new_tokens=128
)
print(processor.batch_decode(
tokens[:, inputs.input_ids.shape[1]:], skip_special_tokens=True
)[0])
Attribution
Fine-tuned from Qwen/Qwen3.5-9B. The upstream license is included as LICENSE-QWEN. Model weights were modified by LoRA fine-tuning and merging; the GGUF variants were converted and quantized using llama.cpp.
- Downloads last month
- 687