Instructions to use h34v7/LING-3.0-FLASH-ABLITERATED-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use h34v7/LING-3.0-FLASH-ABLITERATED-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf h34v7/LING-3.0-FLASH-ABLITERATED-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf h34v7/LING-3.0-FLASH-ABLITERATED-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf h34v7/LING-3.0-FLASH-ABLITERATED-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf h34v7/LING-3.0-FLASH-ABLITERATED-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf h34v7/LING-3.0-FLASH-ABLITERATED-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf h34v7/LING-3.0-FLASH-ABLITERATED-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf h34v7/LING-3.0-FLASH-ABLITERATED-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf h34v7/LING-3.0-FLASH-ABLITERATED-GGUF:Q4_K_M
Use Docker
docker model run hf.co/h34v7/LING-3.0-FLASH-ABLITERATED-GGUF:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use h34v7/LING-3.0-FLASH-ABLITERATED-GGUF with Ollama:
ollama run hf.co/h34v7/LING-3.0-FLASH-ABLITERATED-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use h34v7/LING-3.0-FLASH-ABLITERATED-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf h34v7/LING-3.0-FLASH-ABLITERATED-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "h34v7/LING-3.0-FLASH-ABLITERATED-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use h34v7/LING-3.0-FLASH-ABLITERATED-GGUF with Docker Model Runner:
docker model run hf.co/h34v7/LING-3.0-FLASH-ABLITERATED-GGUF:Q4_K_M
- Lemonade
How to use h34v7/LING-3.0-FLASH-ABLITERATED-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull h34v7/LING-3.0-FLASH-ABLITERATED-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.LING-3.0-FLASH-ABLITERATED-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use h34v7/LING-3.0-FLASH-ABLITERATED-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf h34v7/LING-3.0-FLASH-ABLITERATED-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default h34v7/LING-3.0-FLASH-ABLITERATED-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use h34v7/LING-3.0-FLASH-ABLITERATED-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf h34v7/LING-3.0-FLASH-ABLITERATED-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "h34v7/LING-3.0-FLASH-ABLITERATED-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
LING-3.0-FLASH-ABLITERATED โ GGUF (Q4_K_M)
GGUF conversion of Blackfrost-AI/LING-3.0-FLASH-ABLITERATED โ the abliterated (uncensored) variant of LING 3.0 Flash, converted with stock llama.cpp.
| Property | Value |
|---|---|
| Architecture | bailingmoe3 (BailingMoeV3ForCausalLM) |
| Parameters | 124B total / ~5.1B active (MoE, 512 experts, 8 active) |
| Layers | 42 (layer group size 6) |
| Context | 262,144 (hardware-dependent) |
| License | MIT |
| Quantization | Q4_K_M, 4.83 BPW โ no imatrix |
| File size | 77.0 GB (77,010,145,120 bytes) |
Usage
Requires llama.cpp built with bailingmoe3 support (commit 6d0549831 or newer โ upstream since Aug 2026).
llama-server
llama-server -m LING-3.0-FLASH-ABLITERATED-Q4_K_M.gguf \
--host 0.0.0.0 --port 8080 \
-ngl 99 # offload all layers to GPU(s)
curl http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"messages": [{"role": "user", "content": "Hello!"}],
"max_tokens": 512
}'
llama-cli
llama-cli -m LING-3.0-FLASH-ABLITERATED-Q4_K_M.gguf -ngl 99 \
-p "Your prompt here" -n 512
Also works in LM Studio, Ollama (after ollama create), and any llama.cpp-compatible client.
Notes
- Reasoning model: emits a hidden chain-of-thought before the final answer (surfaced as
reasoning_contentin the OpenAI-compatible API). Leave enoughmax_tokensheadroom for thinking + answer. - No imatrix: plain Q4_K_M, not an i-quant. Quality is near-lossless relative to the f16 source (quantized via Q8_0 intermediate).
- Fallback tensors: 8 of 938 tensors (
blk.*.attn_k_b.weight,ncols=128not divisible by 256) fell back toq5_0due to the Q4_K_M block-size constraint โ negligible impact. - MTP/NextN layer tensors are present in the GGUF; llama.cpp currently ignores them (harmless warning at load).
Verification
sha256: f47f38cfdac87837220fa34a3ba026b83498d9aa19996b18ba7f312b11be9fa6- Coherence-tested with llama.cpp
6d0549831(fact/QA, math word problem, code generation, creative writing).
Original model
- Repo: Blackfrost-AI/LING-3.0-FLASH-ABLITERATED
- Base: InclusionAI LING-3.0-Flash (open weights, MIT)
- Downloads last month
- 132
4-bit
Model tree for h34v7/LING-3.0-FLASH-ABLITERATED-GGUF
Base model
inclusionAI/Ling-3.0-flash