ARustyCoder93

A 93-of-256 routed-expert cut of DeepSeek-V4-Flash in GGUF, made for local agentic coding on one 96 GB GPU. This is not a DeepSeek release. It is a derived model produced by expert pruning; DeepSeek did not make or endorse it. It inherits DeepSeek-V4-Flash's MIT license.

Full recipe, scripts, llama.cpp patch and every checksum: https://github.com/apchat-agent/arustycoder

file arustycoder93.gguf — 62,213,283,424 bytes
sha256 90a4c8a8a32fc664d822f1df093d5e228111c94eebf9398ae5c5737391ecb058
source antirez/deepseek-v4-ggufDeepSeek-V4-Flash-MXFP4Experts-F16HC-F16Compressor-F16Indexer-Q8Attn-Q8Shared-Q8Out-chat-v2-mxfp4-0731.gguf (MXFP4 experts, Q8 attention/shared/output, F16 compressor/indexer); kept tensors are byte-identical to it
experts 93 per layer kept out of 256 (top-6 routing unchanged); 43 MoE layers
runs on stock llama.cpp at or after commit 22b8e31 (2026-08-15). Our serving build for the 2026-08-16 smoke tests was that commit plus the census patch from the recipe repo, which is inert for this file (no full-router tensors). llama-server -m arustycoder93.gguf -ngl 99 -c 262144 -fa on
footprint ~67 GB resident at 256K context on an RTX PRO 6000; ~46 tok/s decode, ~470 tok/s prompt

How the 93 experts were chosen

Two "grow-from-nothing" censuses (start from a random 8 experts per layer with the full 256-wide router kept as side tensors, record which experts the router wants, rebuild at 16, 32, 64) — one on a code/systems corpus, one on a general corpus. ARustyCoder93 is the union of the two keep-64 sets: 78..93 experts per layer (mean 87.3), padded to 93 in the smaller layers by merged rank because a GGUF needs a single expert count. Each keep-64 alone predicts held-out text well but cannot finish generating; the union does. Method, numbers and the merge rule are in the repository README.

What it does and does not do

  • 6/6 coding asks (write Rust, fix Python, explain C, refactor JS, a tool call, a multi-file plan) finished cleanly, 86-1993 completion tokens, with the server settings above. Prompts and raw responses are in the recipe repo under results/.
  • Known limit: a harder single-shot prompt (a generic LRU cache) fell into a verbatim repetition loop in its reasoning to a 12k-token cap. Expect this on some hard reasoning; a DRY sampler stops the loop but corrupts identifiers in code.
  • No standard benchmark has been run on this cut. It is an experiment that turned out useful for interactive coding. Use the unpruned model when correctness matters more than footprint.

Files

  • arustycoder93.gguf — the model
  • config.txt, prune.log — build record of this exact file
  • merged-wants.tsv, coder-census-64.wants.tsv, general-census-64.wants.tsv — the expert rankings it was pruned from
  • PROVENANCE.md — checksums, versions, commits
  • LICENSE-DeepSeek-V4-Flash — the upstream MIT notice (DeepSeek); LICENSE — MIT for the recipe

HF metadata note: base_model points at DeepSeek-V4-Flash; there is no "pruned" relation tag on the Hub, so none is set.

Provenance / credit

Base model: DeepSeek-V4-Flash by DeepSeek (MIT). GGUF source and its quantization: antirez/deepseek-v4-gguf (tagged MIT on the Hub). Pruning method, patch and scripts: this project (MIT), 2026-08-15/17.

Downloads last month
111
GGUF
Model size
108B params
Architecture
deepseek4
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ayourtch/ARustyCoder93

Quantized
(128)
this model