LongCat-Flash-Lite-Sparse-8bit (MLX)

8-bit MLX quantization of meituan-longcat/LongCat-Flash-Lite-Sparse (69B-A3B, LongcatCausalLM).

To our knowledge this is the first working implementation of LongCat-Flash-Lite-Sparse in any framework — no upstream serving stack (mlx-lm, vLLM, SGLang, llama.cpp) supports the oe_embed_* variant yet.

Near-lossless 8-bit (68 GB of weights), for a 128 GB Mac. Smaller-footprint variants: 6-bit (54 GB, 96 GB Macs) and 4-bit (~36 GB, 64 GB Macs).

What's in this checkpoint

LongCat-Flash-Lite-Sparse adds three things vanilla LongCat-Flash lacks:

  • LongCat Sparse Attention (LSA) — a DeepSeek-style lightning indexer over MLA, with streaming-aware indexing (fixed sink + local window) and cross-layer index reuse. Native long context.
  • Zero-computation (identity) experts in the ScMoE decoder (256 routed + 128 identity, top-12).
  • N-gram ("oe") input embedding — ~46% of the parameters, fused into the token embedding.

The n-gram fix

The oe embedding hash and tables are identical to the published n-gram references (the Scaling Embeddings paper, mlx-lm, SGLang, llama.cpp, Meituan's dense modeling). The one difference in LongcatCausalLM is the fusion: it keeps the word embedding at full scaleword + Σ projections / (1 + num_embedders) — rather than the dense form (word + Σ projections) / (1 + num_embedders). Dividing the word by 1 + num_embedders garbles generation; this build applies the correct fusion.

Usage

Requires mlx-vlm with longcat_flash_sparse support (PR #2063):

pip install git+https://github.com/Lazarus-931/mlx-vlm@add-longcat-flash
from mlx_vlm import load, generate
model, processor = load("AlazarM/LongCat-Flash-Lite-Sparse-8bit", trust_remote_code=True)
tok = processor.tokenizer
text = tok.apply_chat_template(
    [{"role": "user", "content": "What is the capital of France?"}],
    tokenize=False, add_generation_prompt=True,
)
print(generate(model, processor, text, max_tokens=64, temperature=0.0))
# -> The capital of France is Paris.

Throughput (M5 Max, 128 GB, batch 1, greedy)

Decode tok/s across the published quantizations:

ctx 4-bit 6-bit 8-bit
512 112 87 80
2048 85 72 65
8192 83 71 65
32768 73 64 60

Batch-1 decode is partly weight-bandwidth-bound, so lower precision is faster (~30% spread 4→8-bit); LSA keeps all three nearly flat as context grows. Peak memory across 512→32k: 4-bit ~39–45 GB, 6-bit ~56–63 GB, 8-bit ~74–80 GB. The extra precision trades speed + memory for quality — since only ~3B params are active per token, quant error has little room to hide, so the 8-bit quality gain is meaningful.

License

MIT, inherited from the base model.

Downloads last month
24
Safetensors
Model size
19B params
Tensor type
BF16
·
U32
·
F32
·
MLX
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for AlazarM/LongCat-Flash-Lite-Sparse-8bit

Quantized
(3)
this model