Instructions to use tinyopsec/Qwen3.5-4B-Uncensored-Aggressive-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use tinyopsec/Qwen3.5-4B-Uncensored-Aggressive-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf tinyopsec/Qwen3.5-4B-Uncensored-Aggressive-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf tinyopsec/Qwen3.5-4B-Uncensored-Aggressive-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf tinyopsec/Qwen3.5-4B-Uncensored-Aggressive-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf tinyopsec/Qwen3.5-4B-Uncensored-Aggressive-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf tinyopsec/Qwen3.5-4B-Uncensored-Aggressive-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf tinyopsec/Qwen3.5-4B-Uncensored-Aggressive-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf tinyopsec/Qwen3.5-4B-Uncensored-Aggressive-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf tinyopsec/Qwen3.5-4B-Uncensored-Aggressive-GGUF:Q4_K_M
Use Docker
docker model run hf.co/tinyopsec/Qwen3.5-4B-Uncensored-Aggressive-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use tinyopsec/Qwen3.5-4B-Uncensored-Aggressive-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "tinyopsec/Qwen3.5-4B-Uncensored-Aggressive-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "tinyopsec/Qwen3.5-4B-Uncensored-Aggressive-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/tinyopsec/Qwen3.5-4B-Uncensored-Aggressive-GGUF:Q4_K_M
- Ollama
How to use tinyopsec/Qwen3.5-4B-Uncensored-Aggressive-GGUF with Ollama:
ollama run hf.co/tinyopsec/Qwen3.5-4B-Uncensored-Aggressive-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use tinyopsec/Qwen3.5-4B-Uncensored-Aggressive-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf tinyopsec/Qwen3.5-4B-Uncensored-Aggressive-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "tinyopsec/Qwen3.5-4B-Uncensored-Aggressive-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use tinyopsec/Qwen3.5-4B-Uncensored-Aggressive-GGUF with Docker Model Runner:
docker model run hf.co/tinyopsec/Qwen3.5-4B-Uncensored-Aggressive-GGUF:Q4_K_M
- Lemonade
How to use tinyopsec/Qwen3.5-4B-Uncensored-Aggressive-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull tinyopsec/Qwen3.5-4B-Uncensored-Aggressive-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.Qwen3.5-4B-Uncensored-Aggressive-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use tinyopsec/Qwen3.5-4B-Uncensored-Aggressive-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf tinyopsec/Qwen3.5-4B-Uncensored-Aggressive-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default tinyopsec/Qwen3.5-4B-Uncensored-Aggressive-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use tinyopsec/Qwen3.5-4B-Uncensored-Aggressive-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf tinyopsec/Qwen3.5-4B-Uncensored-Aggressive-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "tinyopsec/Qwen3.5-4B-Uncensored-Aggressive-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Qwen3.5-4B-Uncensored-Aggressive - GGUF
GGUF quantized versions of rodrigomt/Qwen3.5-4B-Uncensored-Aggressive, a 4.5 billion parameter language model based on the Qwen3.5-4B architecture, optimized for unrestricted text generation and direct instruction following.
Model Details
- Base Model: Qwen/Qwen3.5-4B
- Fine-tuned by: rodrigomt
- Architecture: Qwen2 (28 layers, 28 attention heads)
- Context Length: 32768 tokens
- Vocabulary Size: 151936
- Parameters: 4.5B
Quantization
| Filename | Bits | Size | Use Case |
|---|---|---|---|
model_f16.gguf |
16 | ~8.4 GB | Maximum quality, high VRAM requirement |
model_q8_0.gguf |
8 | ~4.5 GB | High quality, moderate VRAM |
model_q6_k.gguf |
6 | ~3.4 GB | Good quality, balanced VRAM |
model_q5_k_m.gguf |
5 | ~2.8 GB | Recommended for most use cases |
model_q5_k_s.gguf |
5 | ~2.5 GB | Compact, minimal quality loss |
model_q4_k_m.gguf |
4 | ~2.1 GB | Good balance, low VRAM |
model_q4_k_s.gguf |
4 | ~1.9 GB | Compact model |
model_q3_k_l.gguf |
3 | ~1.7 GB | Very compact |
model_q3_k_m.gguf |
3 | ~1.5 GB | Minimal size |
model_q3_k_s.gguf |
3 | ~1.4 GB | Extreme compression |
model_q2_k.gguf |
2 | ~1.1 GB | Maximum compression |
VRAM Requirements
| Quantization | VRAM | Recommended Device |
|---|---|---|
| F16 | 9 GB | High-end GPU |
| Q8_0 | 5 GB | Mid-range GPU |
| Q6_K | 3.5 GB | Mid-range GPU |
| Q5_K_M | 3 GB | Standard GPU |
| Q4_K_M | 2.5 GB | Standard GPU / Laptop GPU |
| Q3_K_M | 2 GB | Laptop GPU / CPU |
| Q2_K | 1.5 GB | CPU with sufficient RAM |
Usage
llama.cpp
./main -m model_q5_k_m.gguf -n 256 -p "You are a helpful assistant."
llama-cpp-python
from llama_cpp import Llama
llm = Llama(
model_path="model_q5_k_m.gguf",
n_ctx=32768,
n_threads=8
)
response = llm("What is machine learning?", max_tokens=256)
print(response["choices"][0]["text"])
LM Studio
Download the quantized GGUF file and load it in LM Studio's model picker.
Ollama
ollama pull tinyopsec/qwen3.5-4b-uncensored-aggressive-gguf:q5_k_m
ollama run tinyopsec/qwen3.5-4b-uncensored-aggressive-gguf:q5_k_m
Model Capabilities
This model excels at:
- Unrestricted Generation: Optimized for direct, uncensored text generation without safety filters
- Aggressive Instruction Following: Maximized compliance with user instructions and prompts
- Diverse Content Generation: Capable of generating a wide range of content types
- Lightweight Deployment: Efficient 4.5B parameter architecture suitable for local inference
- Research & Experimentation: Designed for exploring model behavior without refusal patterns
License
This quantized version maintains the original model's Apache 2.0 license. See the original model card for full details.
Disclaimer
These are community quantizations. For official support and documentation, refer to the original model repository and llama.cpp documentation.
- Downloads last month
- -
2-bit
3-bit
4-bit
5-bit
6-bit
8-bit
16-bit
Model tree for tinyopsec/Qwen3.5-4B-Uncensored-Aggressive-GGUF
Base model
Qwen/Qwen3.5-4B-Base