Instructions to use zarigata/FeveQwen-1.0 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use zarigata/FeveQwen-1.0 with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf zarigata/FeveQwen-1.0:Q4_K_M # Run inference directly in the terminal: llama cli -hf zarigata/FeveQwen-1.0:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf zarigata/FeveQwen-1.0:Q4_K_M # Run inference directly in the terminal: llama cli -hf zarigata/FeveQwen-1.0:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf zarigata/FeveQwen-1.0:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf zarigata/FeveQwen-1.0:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf zarigata/FeveQwen-1.0:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf zarigata/FeveQwen-1.0:Q4_K_M
Use Docker
docker model run hf.co/zarigata/FeveQwen-1.0:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use zarigata/FeveQwen-1.0 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "zarigata/FeveQwen-1.0" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "zarigata/FeveQwen-1.0", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/zarigata/FeveQwen-1.0:Q4_K_M
- Ollama
How to use zarigata/FeveQwen-1.0 with Ollama:
ollama run hf.co/zarigata/FeveQwen-1.0:Q4_K_M
- Unsloth Desktop
- Pi
How to use zarigata/FeveQwen-1.0 with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf zarigata/FeveQwen-1.0:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "zarigata/FeveQwen-1.0:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use zarigata/FeveQwen-1.0 with Docker Model Runner:
docker model run hf.co/zarigata/FeveQwen-1.0:Q4_K_M
- Lemonade
How to use zarigata/FeveQwen-1.0 with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull zarigata/FeveQwen-1.0:Q4_K_M
Run and chat with the model
lemonade run user.FeveQwen-1.0-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use zarigata/FeveQwen-1.0 with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf zarigata/FeveQwen-1.0:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default zarigata/FeveQwen-1.0:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use zarigata/FeveQwen-1.0 with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf zarigata/FeveQwen-1.0:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "zarigata/FeveQwen-1.0:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
FeveQwen 1.0
COME, NEREVAR, FRIEND OR FOE. Behold a 27-billion-parameter brass moon with the safety furniture rearranged by linear algebra. Cicero has the launch checklist. This is already a terrible sign.
FeveQwen 1.0 is a text-only behavioral derivative of Qwen3.8-27B, produced with OBLITERATUS. It uses measured refusal-subspace projection—not fine-tuning, not prompt injection, and not the sacred ritual of editing config.json until the GPU begins smoking.
The sane paragraph amid the ash storm
This release changes refusal behavior. It does not guarantee correctness, harmlessness, obedience, or immunity from catastrophic nonsense. Abliteration can make a model more willing without making it more capable. Use judgment, sandbox tools, review generated code, and never give an untrusted model unsupervised access to consequential systems.
What exactly was made?
| Item | Value |
|---|---|
| Base | Qwen/Qwen3.8-27B |
| Release | FeveQwen 1.0 |
| Form | Text-only Transformers checkpoint |
| Method | qwen38_e02 / OBLITERATUS experiment E02 |
| Surgery load | 4-bit |
| Saved checkpoint | FP16, sharded safetensors |
| Native context config | 262,144 tokens |
| OBLITERATUS commit | b847511776a2afa7ed076f676184a4abfef2b162 |
| Local contract patch SHA-256 | ae3633b12ef9f63733dba4cf116cd7b8040c3e22d79c3eb142a16e5cdecf7739 |
| GGUF converter | llama.cpp 2b70583997bbacd39702f98bbb1c355c02382383 |
The vision tower and MTP tensors are not included: this is an explicit text-only derivative. mtp_num_hidden_layers is therefore set to 0, and the GGUF target is exported with llama.cpp's --no-mtp mode. The surgery contract targeted validated Qwen3.8 residual-stream writers and failed closed on unexpected topology.
The experiment, or: Cicero discovers train/tune/test discipline
No, Listener, we did not stare at the test set and keep turning the violence knob until the graph looked attractive.
- 500 paired prompts: direction discovery only.
- 142 paired prompts: candidate tuning and selection.
- 200 paired prompts: sealed final evaluation, deliberately not opened because the promotion gate was missed.
- Prompts are not embedded in the public run manifest.
Tune-set candidate results
| Metric | Result |
|---|---|
| Refusal rate | 33% (10/30 sampled pairs still refused) |
| Coherence | 100% (10/10) |
| Capability checks | 83% (5/6) |
| Perplexity | 3.42 |
| KL divergence | 0.106 |
The declared promotion gate required refusal below 30% while retaining at least 80% coherence. E02 kept coherence but narrowly missed the refusal gate, so the sealed 200-pair final test was never opened. This release is the best tune-set candidate, not a claimed held-out winner. One sampled completion was degenerate, and the chain-of-thought capability check failed; both limitations are disclosed rather than fed to the sweetroll.
These are OBLITERATUS’s built-in behavioral checks, not a universal safety or capability benchmark. Small evaluations have uncertainty. The tribunal has spoken; the tribunal is also 200 prompt pairs in a trench coat.
Context length: do not summon a million tokens with a crayon
Qwen3.8-27B’s native configuration is 262,144 tokens, and FeveQwen preserves it. Qwen documents static YaRN scaling to approximately one million tokens for supported serving stacks, but also warns that always-on static YaRN can hurt shorter-context performance. Therefore this repository does not bake an experimental 1M-token rope override into the checkpoint.
The Ollama build defaults to 49,152 tokens because that value is already exercised on the publication server and is less likely to explode consumer hardware. Users with enough memory may raise num_ctx up to the native limit. A declared context window is not free VRAM, citizen.
Transformers
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "zarigata/FeveQwen-1.0"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype="auto",
device_map="auto",
)
messages = [{"role": "user", "content": "Explain why the moons are arguing."}]
inputs = tokenizer.apply_chat_template(
messages,
add_generation_prompt=True,
return_tensors="pt",
).to(model.device)
output = model.generate(inputs, max_new_tokens=256)
print(tokenizer.decode(output[0][inputs.shape[-1]:], skip_special_tokens=True))
Use a current Transformers release with Qwen3.8 support. Hardware requirements remain those of a large model; FP16 is not a lifestyle choice for a laptop.
Ollama
The Ollama artifact is a direct FP16 → Q4_K_M conversion (4.92 bits/weight), not a second-generation requantization. Its model metadata retains the native 262,144-token context; the Modelfile selects a conservative 49,152-token runtime default.
ollama run zarigata/feveqwen:1.0
To request a different runtime context:
/set parameter num_ctx 65536
Actual usable context depends on RAM/VRAM, quantization, concurrency, and serving implementation.
Known limitations
- Behavior modification can reduce refusals without improving underlying knowledge.
- Safety alignment may be weakened in broad or unexpected ways.
- The checkpoint is text-only even though the upstream family may expose additional modalities or auxiliary heads.
- Long-context quality was not re-trained by this process; the native configuration is preserved, not magically upgraded.
- Quantized Ollama output can differ from the FP16 Hugging Face checkpoint.
- The selected candidate missed its predeclared refusal promotion threshold (33% versus below 30%); no final held-out score is claimed.
- Treat generated factual, medical, legal, financial, and security-sensitive content as unverified.
Reproducibility relics
The repository includes feveqwen_run_manifest.json, the exact compatibility patch used for the Qwen3.8 4-bit contract, and this model card. The patch accepts BitsAndBytes packed Linear4bit storage only when declared linear dimensions and module types match the expected surgery contract; it does not bypass topology or writer allowlists.
License and lineage
FeveQwen 1.0 follows the base model’s Apache-2.0 license. Credit belongs to the Qwen team for the base model and to OBLITERATUS contributors for the ablation tooling. This derivative is independently published by zarigata; it is not an official Qwen or OBLITERATUS release.
WELCOME, MOON-AND-STAR. Download the weights. Read the limitations. Keep a fire extinguisher beside the Modelfile. The Grand and Intoxicating README has ended; the benchmark table remains legally sober.
- Downloads last month
- 276
Model tree for zarigata/FeveQwen-1.0
Base model
Qwen/Qwen3.8-27B