Agents-A1-Flash (Tencent STQ1_0 + FP8 Mixed)

Compressed release of InternScience/Agents-A1 (35B MoE, 256 experts, 8 active per token).

Compression Specification

  • Routed MoE Experts: Tencent STQ1_0 (Sherry 3:4 fine-grained ternary sparsity, 1.3125 bpw).
  • Attention & Shared Experts: FP8 (float8_e4m3fn).
  • Layer Norms & Vectors: FP16.
  • Original Size: ~70 GB BF16
  • Compressed Size: 8.19 GB (88.3% reduction)
  • Serving Budget: Runs in <3.5 GB VRAM using on-demand storage caching or ~8.2 GB full VRAM on an NVIDIA T4 GPU.

Shard Distribution

  • model-00001 to model-00011 (~441 MB each): STQ1_0 routed MoE expert layers.
  • model-00000, model-00012, model-00013: FP8 embeddings, vision/dense layers, and attention blocks.
Downloads last month
259
Safetensors
Model size
8B params
Tensor type
F16
F8_E4M3
U8
Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support

Model tree for SofiTesfay2010/Agents-A1-Flash

Finetuned
(5)
this model