Instructions to use akoumpa/Moonlight-V4-1B-h16d256 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use akoumpa/Moonlight-V4-1B-h16d256 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="akoumpa/Moonlight-V4-1B-h16d256")# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("akoumpa/Moonlight-V4-1B-h16d256") model = AutoModelForCausalLM.from_pretrained("akoumpa/Moonlight-V4-1B-h16d256", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use akoumpa/Moonlight-V4-1B-h16d256 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "akoumpa/Moonlight-V4-1B-h16d256" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "akoumpa/Moonlight-V4-1B-h16d256", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/akoumpa/Moonlight-V4-1B-h16d256
- SGLang
How to use akoumpa/Moonlight-V4-1B-h16d256 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "akoumpa/Moonlight-V4-1B-h16d256" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "akoumpa/Moonlight-V4-1B-h16d256", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "akoumpa/Moonlight-V4-1B-h16d256" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "akoumpa/Moonlight-V4-1B-h16d256", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use akoumpa/Moonlight-V4-1B-h16d256 with Docker Model Runner:
docker model run hf.co/akoumpa/Moonlight-V4-1B-h16d256
Moonlight-V4-1B-h16d256
An untrained, ~1B-parameter DeepSeek-V4-architecture configuration (no weights): DeepSeek-V4-architecture 1B config with 16 x 256 attention heads (CSA ratio 4 + lightning indexer). It is the small end of a Moonlight-style down-scaling of DeepSeek-V4, sized to pre-train on two 48 GB GPUs, with attention dimensions chosen so that the TileLang sparse-attention kernel fits GPUs with 99 KB of shared memory (Ada / consumer class). This is the V4-faithful variant: 7 Compressed Sparse Attention (CSA) layers with ratio-4 overlapped compression and a lightning indexer, 6 Heavily Compressed Attention (HCA) layers (ratio 128) and 2 pure sliding-window layers.
The sibling repo replaces the ratio-4 CSA layers by ratio-8 compression without an indexer, which is cheaper to train at short context and avoids the indexer kernel entirely. Sibling: akoumpa/Moonlight-V4-1B-h16d256-r8.
| total | non-embedding | activated / token | activated non-embedding | |
|---|---|---|---|---|
| parameters | 999,670,879 (999.7M) | 734.9M | 539.6M | 274.8M |
Lineage
| model | architecture | size | status |
|---|---|---|---|
| Moonlight-16B-A3B (Moonshot) | DeepSeek-V3 | 16B / 3B active | released, trained with Muon |
| Moonlight-V4-16B-A3B | DeepSeek-V4 at Moonlight's width/depth/experts | 16.5B / 3.0B active | proposed config |
| Moonlight-V4-1B (8 heads x 512) | same, shrunk | 1.0B / 0.54B active | config; eager attention only |
| Moonlight-V4-1B-h16d256 | same, attention re-shaped for the kernels | 999.7M / 539.6M active | this repo |
DeepSeek-V4 (Flash: 284B/13B, Pro: 1.6T/49B) keeps head_dim = 512 and 64 or 128 query heads; Moonlight's 16 heads
were kept for the 16B analogue, and here 16 heads x 256 dims give the same 4096-wide query space as the 8 x 512
1B config with half the shared-KV width.
Architecture
| component | setting |
|---|---|
| hidden size / layers | 1024 / 15 |
attention schedule (compress_ratios) |
[0, 0, 4, 128, 4, 128, 4, 128, 4, 128, 4, 128, 4, 128, 4]: 2 sliding-window, 7 compressed (ratio 4, overlapped, with indexer), 6 HCA (ratio 128) |
| attention | shared-KV MQA (num_key_value_heads 1): 16 query heads x head_dim 256 (last 64 dims RoPE), the same 256-dim entry is key and value |
| query path | q_lora_rank 256 -> 16 x 256; per-head RMSNorm before RoPE |
| output projection | grouped low-rank: o_groups 2 (8 heads per group) x o_lora_rank 1024 -> hidden |
| sliding window / attention sinks | 128 tokens on every layer / one learnable sink logit per head |
| Lightning indexer | 64 heads x 128 dims, index_topk 1024 (a power of two, as the indexer kernels require); at <= 4K tokens this is >= the 1024 pooled entries, i.e. dense |
| MoE (every layer) | 32 routed experts x 384 (SwiGLU, clamp 10.0), top-6, 1 shared expert; sqrtsoftplus scoring, aux-loss-free bias (noaux_tc), routed_scaling_factor 2.436 |
| hash-routed layers | first 1 MoE layer(s) route by a fixed token-id table (tid2eid) |
| residual stream | manifold-constrained hyper-connections, hc_mult 4, 20 Sinkhorn iterations |
| positions | 4096 tokens, plain RoPE (theta 10000 on sliding layers, 160000 on compressed layers, no YaRN) |
| MTP | none (num_nextn_predict_layers 0) |
| vocabulary | 129280 (DeepSeek-V4 tokenizer, BOS 0, EOS 1) |
| norm eps / init | 1e-06 / 0.02 |
Per-layer parameters: attention 7.9M (sliding) / 11.6M (ratio 4) / 8.4M (HCA); MoE 39.0M
total, 8.3M activated (each expert 1.2M); mHC mixers 196,662. Embedding and
head are 132.4M each. KV cache per sequence at 4K / 32K tokens (FP8 non-RoPE dims, bf16 RoPE dims): 3.3 MiB / 22.1 MiB.
Core-attention FLOPs per generated token at 4K context: 0.27 GF, against 0.81 GF of linear layers.
python count_params.py config.json reproduces these numbers (the script also reproduces the published sizes of
Moonlight, DeepSeek-V3, V4-Flash and V4-Pro).
Design rules: head_dim and index_topk are powers of two and index_n_heads divides 128 (kernel requirements);
o_lora_rank 1024 keeps DeepSeek-V4's per-group output projection shape; routed_scaling_factor follows Moonlight's
recipe (expected 1/||p||_2 of renormalised top-k scores) applied to sqrt-softplus with 32 experts / top-6.
Files
| file | purpose |
|---|---|
config.json |
Hugging Face config (same key set as deepseek-ai/DeepSeek-V4-Flash) |
inference_config.json |
the same model in the key format of DeepSeek's reference inference/model.py |
tokenizer.json, tokenizer_config.json |
DeepSeek-V4 tokenizer (MIT), copied from deepseek-ai/DeepSeek-V4-Flash |
count_params.py |
parameter / KV-cache / FLOP calculator for DeepSeek-V3- and V4-style configs |
training/pretrain.yaml, training/train.py |
NeMo Automodel recipe (FSDP2, 2 GPUs) and launcher that seeds the hash table and mHC mixers |
training/prepare_data.py, training/finite_nanogpt.py, training/init_utils.py |
data shards from parquet text, bounded validation dataset, from-scratch initialisers |
Loading
transformers (>= 5.8, native deepseek_v4)
from transformers import AutoTokenizer, DeepseekV4Config, DeepseekV4ForCausalLM
cfg = DeepseekV4Config.from_pretrained("akoumpa/Moonlight-V4-1B-h16d256") # legacy keys (compress_ratios, num_hash_layers, ...) are folded in
tok = AutoTokenizer.from_pretrained("akoumpa/Moonlight-V4-1B-h16d256") # or PreTrainedTokenizerFast.from_pretrained
model = DeepseekV4ForCausalLM(cfg) # random init, 999.7M parameters
The transformers implementation is inference-oriented: with no KV cache it appends compressed entries to the
key axis without a causal mask and gathers per-query top-k entries into an S x k key axis, so do not train with it.
Use it for architecture inspection and, once you have weights, for generation.
NeMo Automodel (native training implementation)
from nemo_automodel.components.models.common import BackendConfig
from nemo_automodel.components.models.deepseek_v4.config import DeepseekV4Config
from nemo_automodel.components.models.deepseek_v4.model import DeepseekV4ForCausalLM
cfg = DeepseekV4Config.from_pretrained("akoumpa/Moonlight-V4-1B-h16d256")
backend = BackendConfig(attn="eager", linear="torch", rms_norm="torch_fp32", rope_fusion=False,
dispatcher="torch", experts="torch_mm", enable_hf_state_dict_adapter=False)
model = DeepseekV4ForCausalLM(cfg, backend=backend)
model.initialize_weights(dtype=torch.bfloat16)
Or in a recipe: NeMoAutoModelForCausalLM.from_config with config: DeepseekV4Config.from_pretrained("akoumpa/Moonlight-V4-1B-h16d256")
(see training/pretrain.yaml).
DeepSeek reference code
inference_config.json drops into the inference/ folder of the DeepSeek-V4 release (ModelArgs keys, n_mtp_layers 0).
Training from scratch
The released implementations only ever load checkpoints, so two things must be initialised by hand; training/init_utils.py
does both and training/train.py calls it after the recipe's own weight init on a fresh start:
- Hash-routing table.
tid2eidis created as zeros (every token to expert 0).fill_hash_tableswrites a balanced token-id hash: 32 experts, 6 distinct experts per token, equal load over the vocabulary. - mHC mixers. NeMo Automodel leaves the
fn/base/scaletensors uninitialised;init_hyper_connectionsapplies transformers' rule (normal(0, 0.02) projection, zero bias, unit gates).
Recipe notes that cost time to find (all encoded in training/):
- Automodel wraps DeepSeek-V4's fp32 tensors (attention sinks, compressor position biases, mHC mixers,
lm_head) as their own FSDP2 units whose forward returns the parameter; if they reshard after forward, attention reads a freed tensor.train.pycallsset_reshard_after_forward(False)on those units. - Use the logits-based
MaskedCrossEntropy: the fused linear cross-entropy rejects the fp32lm_headx bf16 hidden states. - For iterable datasets the recipe passes no batch size to the DataLoader; set
dataloader.batch_sizeexplicitly. NanogptDatasetis an infinite stream; validation usesFiniteNanogptDataset.- The indexer's top-k has no gradient path (Automodel freezes its parameters). With
index_topk1024 every query sees all causal compressed entries up to 4K tokens, i.e. DeepSeek's own dense warm-up regime; sparse training at longer contexts needs an indexer distillation loss. - To train beyond 4K, add DeepSeek-V4's YaRN block (
rope_scaling: factor 16,original_max_position_embeddings65536) and raisemax_position_embeddingsandindex_topk.
Measured on two RTX 5880 Ada (48 GB, sm_89) GPUs, bf16, torch_mm experts, chunked cross-entropy, single-GPU forward+backward
(TileLang sparse attention with Sinkhorn and indexer on torch; the eager path for the same model reaches 4.2k tok/s at
B=2 x 2048 and runs out of memory at B=4):
| micro-batch | fwd+bwd time, throughput, peak memory |
|---|---|
| 2 x 2048 | 740 ms, 5.5k tok/s, 15.1 GiB |
| 4 x 2048 | 1227 ms, 6.7k tok/s, 27.9 GiB |
| 2 x 4096 | 1421 ms, 5.8k tok/s, 33.1 GiB |
Forward-time breakdown at B=2 x 2048: attention 59 ms, indexer 82 ms, MoE 39 ms, mHC mixers 35 ms, compressor 12 ms (247 ms forward). With the recipe's FSDP2 data parallelism over 2 GPUs and
AdamW, a 32-sequence x 2048-token global batch is a reasonable starting point (training/pretrain.yaml).
TileLang kernels
Automodel's vendored Miles/TileLang kernels (sparse attention, indexer) and DeepSeek's TileKernels Sinkhorn were written for Hopper's 227 KB of shared memory. On a 99 KB-per-block GPU:
| kernel | shape rule | shared memory | Ada (99 KB) |
|---|---|---|---|
| sparse attention fwd/bwd | head_dim power of two; heads padded to 16, chunked by 16, multiples of 64 above 64 |
147 KB at head_dim 512, fits at 256 |
runs with this config; parity with the torch reference verified |
| lightning indexer fwd | index_n_heads <= 64, multiple of 8, divides 128; index_topk power of two (bwd) |
224 KB (block_N 256 x 128 fp32) |
does not fit; run the indexer on torch |
| Sinkhorn (mHC) | needs the tile_kernels package |
small | needs TileKernels installed or the torch fallback |
Automodel currently switches all three together (backend.attn: tilelang); the measurements above used per-kernel
selection (only sparse attention on TileLang). On Hopper GPUs the default head_dim 512 kernels fit and the 16 x 256
choice is optional.
Limitations and notes
- No trained weights are provided; all numbers are architecture-derived or short from-scratch measurements.
- The name follows the Moonlight / DeepSeek-V4 lineage for clarity; this repository is not affiliated with Moonshot AI or DeepSeek.
- The tokenizer files are DeepSeek's (MIT licence,
deepseek-ai/DeepSeek-V4-Flash).
References
- Liu et al., Muon is Scalable for LLM Training (Moonlight), arXiv:2502.16982
- DeepSeek-AI, DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence, arXiv:2606.19348
- Xie et al., Manifold-Constrained Hyper-Connections (mHC), 2026
- Roller et al., Hash Layers for Large Sparse Models, NeurIPS 2021
- DeepSeek-AI, DeepSeek-V3.2 (DeepSeek Sparse Attention, lightning indexer), 2025
- Kernels: Miles (sparse attention / indexer, vendored in NeMo Automodel), TileKernels (Sinkhorn), TileLang
- Downloads last month
- 246