LongCat-Flash-Lite-Sparse-4bit (MLX)

4-bit MLX quantization of meituan-longcat/LongCat-Flash-Lite-Sparse (69B-A3B, LongcatCausalLM).

To our knowledge this is the first working implementation of LongCat-Flash-Lite-Sparse in any framework — no upstream serving stack (mlx-lm, vLLM, SGLang, llama.cpp) supports the oe_embed_* variant yet.

4-bit (36 GB of weights) is the smallest and fastest variant, for a 64 GB Mac. Also available: 6-bit (52 GB, 96 GB Macs) and 8-bit (~68 GB, 128 GB Macs, near-lossless).

What's in this checkpoint

LongCat-Flash-Lite-Sparse adds three things vanilla LongCat-Flash lacks:

  • LongCat Sparse Attention (LSA) — a DeepSeek-style lightning indexer over MLA, with streaming-aware indexing (fixed sink + local window) and cross-layer index reuse (cli_factor). Native long context.
  • Zero-computation (identity) experts in the ScMoE decoder (256 routed + 128 identity, top-12).
  • N-gram ("oe") input embedding — ~46% of the parameters, fused into the token embedding.

The n-gram fix

The oe embedding hash and tables are identical to the published n-gram references (the Scaling Embeddings paper, mlx-lm, SGLang, llama.cpp, Meituan's dense modeling). The one difference in LongcatCausalLM is the fusion: it keeps the word embedding at full scaleword + Σ projections / (1 + num_embedders) — rather than the dense form (word + Σ projections) / (1 + num_embedders). Dividing the word by 13 garbles generation; this build applies the correct fusion.

Usage

Requires mlx-vlm with longcat_flash_sparse support (PR #2063):

pip install git+https://github.com/Lazarus-931/mlx-vlm@add-longcat-flash
from mlx_vlm import load, generate
model, processor = load("AlazarM/LongCat-Flash-Lite-Sparse-4bit", trust_remote_code=True)
tok = processor.tokenizer
text = tok.apply_chat_template(
    [{"role": "user", "content": "What is the capital of France?"}],
    tokenize=False, add_generation_prompt=True,
)
print(generate(model, processor, text, max_tokens=64, temperature=0.0))
# -> The capital of France is Paris.

Throughput (M5 Max, 128 GB, batch 1, greedy)

Same methodology across quantizations (chunk-512 prefill, warmed kernels).

Decode tok/s

ctx 4-bit 6-bit 8-bit
512 112 87 80
1024 101 83 75
2048 85 72 65
4096 84 72 65
8192 83 71 65
16384 79 66 64
32768 73 64 60

Prefill tok/s

ctx 4-bit 6-bit 8-bit
512 3142 2581 2421
2048 2386 2366 1923
8192 1756 1627 1312
32768 623 504 492

Footprint — peak memory across 512→32k: 4-bit ~39–45 GB · 6-bit ~56–63 GB · 8-bit ~74–80 GB.

LSA's dynamic sparse selection activates once the KV length exceeds index_topk (2048), keeping decode nearly flat (4-bit 112→73 tok/s to 32k). Batch-1 decode is partly weight-bandwidth-bound, so lower precision is faster; higher precision trades that for quality — only ~3B params are active per token, so quant error has little room to hide.

License

MIT, inherited from the base model.

Downloads last month
83
Safetensors
Model size
11B params
Tensor type
BF16
·
U32
·
F32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for AlazarM/LongCat-Flash-Lite-Sparse-4bit

Quantized
(3)
this model