LING-3.0-FLASH-ABLITERATED โ€” GGUF (Q4_K_M)

GGUF conversion of Blackfrost-AI/LING-3.0-FLASH-ABLITERATED โ€” the abliterated (uncensored) variant of LING 3.0 Flash, converted with stock llama.cpp.

Property Value
Architecture bailingmoe3 (BailingMoeV3ForCausalLM)
Parameters 124B total / ~5.1B active (MoE, 512 experts, 8 active)
Layers 42 (layer group size 6)
Context 262,144 (hardware-dependent)
License MIT
Quantization Q4_K_M, 4.83 BPW โ€” no imatrix
File size 77.0 GB (77,010,145,120 bytes)

Usage

Requires llama.cpp built with bailingmoe3 support (commit 6d0549831 or newer โ€” upstream since Aug 2026).

llama-server

llama-server -m LING-3.0-FLASH-ABLITERATED-Q4_K_M.gguf \
  --host 0.0.0.0 --port 8080 \
  -ngl 99          # offload all layers to GPU(s)
curl http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "messages": [{"role": "user", "content": "Hello!"}],
    "max_tokens": 512
  }'

llama-cli

llama-cli -m LING-3.0-FLASH-ABLITERATED-Q4_K_M.gguf -ngl 99 \
  -p "Your prompt here" -n 512

Also works in LM Studio, Ollama (after ollama create), and any llama.cpp-compatible client.

Notes

  • Reasoning model: emits a hidden chain-of-thought before the final answer (surfaced as reasoning_content in the OpenAI-compatible API). Leave enough max_tokens headroom for thinking + answer.
  • No imatrix: plain Q4_K_M, not an i-quant. Quality is near-lossless relative to the f16 source (quantized via Q8_0 intermediate).
  • Fallback tensors: 8 of 938 tensors (blk.*.attn_k_b.weight, ncols=128 not divisible by 256) fell back to q5_0 due to the Q4_K_M block-size constraint โ€” negligible impact.
  • MTP/NextN layer tensors are present in the GGUF; llama.cpp currently ignores them (harmless warning at load).

Verification

  • sha256: f47f38cfdac87837220fa34a3ba026b83498d9aa19996b18ba7f312b11be9fa6
  • Coherence-tested with llama.cpp 6d0549831 (fact/QA, math word problem, code generation, creative writing).

Original model

Downloads last month
132
GGUF
Model size
127B params
Architecture
bailingmoe3
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for h34v7/LING-3.0-FLASH-ABLITERATED-GGUF

Quantized
(3)
this model