qwen3-coder-next.knot
Sovereign .knot encoding of Qwen3-Coder-Next (arch=qwen3next) for the
Gnosis fat-station runtime. Converted from GGUF with gguf-to-knot.py
(K-quant block passthrough, arch-aware config).
Why a knot
.knot is a range-addressable weight container: a station fetches only the
tensors for the layers it owns, over HTTP, instead of loading a whole file.
That is what lets an A3B MoE this size run sharded across Cloudflare
Containers (12 GiB each) rather than needing one large host.
Architecture
Hybrid Mamba-2 gated-DeltaNet SSM with periodic full attention, and a sparse MoE FFN on every layer. Roughly 3B parameters are active per token. Long context is cheap because only the periodic attention layers carry a KV cache.
Serving
Pipeline-parallel across N stations, each owning a contiguous layer range:
FAT_STATION_KNOT=<this knot url>
FAT_STATION_SHARDED=1
FAT_STATION_ROLE=entry|mid|exit
FAT_STATION_LAYERS=0..6
MoE shards cleanly along layers: a layer owns its entire expert set, so top-k routing never needs a tensor held by another station.
Published by AFFECTIVELY. https://huggingface.co/forkjoin-ai
Model tree for forkjoin-ai/qwen3-coder-next-knot
Base model
Qwen/Qwen3-Coder-Next