Webcoda AI โ€” 35B-A3B (Q6_K GGUF)

Webcoda AI is a knowledge assistant fine-tuned from unsloth/Qwen3.6-35B-A3B (a Mixture-of-Experts model, 256 experts / ~3B active) to answer questions about Webcoda, a Sydney-based digital agency. This is the highest-quality model in the Webcoda AI family โ€” it passed 6/6 on the factual + identity validation gate.

Run it (llama.cpp / LM Studio / Ollama)

One universal file โ€” runs on Mac (Metal), AMD (Vulkan/ROCm), and NVIDIA (CUDA), or CPU. Requires a recent llama.cpp build (linear-attention qwen3_5_moe support).

hf download hardin/Webcoda-AI-35B-A3B-GGUF Webcoda-AI-35B-A3B-Q6_K.gguf --local-dir .
llama-server --model Webcoda-AI-35B-A3B-Q6_K.gguf --jinja --ctx-size 8192 --n-gpu-layers 999

Identity is baked in โ€” a bare "What is your name?" answers "Webcoda AI" with no system prompt needed.

Model details

  • Base: unsloth/Qwen3.6-35B-A3B (Qwen3_5MoeForConditionalGeneration)
  • Fine-tune: bf16 LoRA, r=128, on the attention + MoE expert projections (17.5% of params trainable), merged to 16-bit then quantized to Q6_K (~28.5 GB).
  • Text-only: the base model's vision tower and MTP (nextn) speculative head were removed during GGUF export; standard autoregressive inference is unaffected.
  • Tokenizer: PreTrainedTokenizerFast (repacked from the base's newer backend).

Family

  • hardin/Webcoda-AI-14B-GGUF โ€” smaller/faster (dense)
  • hardin/Webcoda-AI-27B-GGUF โ€” dense mid-size
  • hardin/Webcoda-AI-35B-A3B-GGUF โ€” this (MoE, best accuracy, ~3B active params โ†’ fast)
Downloads last month
139
GGUF
Model size
35B params
Architecture
qwen35moe
Hardware compatibility
Log In to add your hardware

6-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for hardin/Webcoda-AI-35B-A3B-GGUF

Quantized
(3)
this model