Agents-A1-Flash (Tencent STQ1_0 + FP8 Mixed)
Compressed release of InternScience/Agents-A1 (35B MoE, 256 experts, 8 active per token).
Compression Specification
- Routed MoE Experts: Tencent STQ1_0 (Sherry 3:4 fine-grained ternary sparsity, 1.3125 bpw).
- Attention & Shared Experts: FP8 (
float8_e4m3fn). - Layer Norms & Vectors: FP16.
- Original Size: ~70 GB BF16
- Compressed Size:
8.19 GB (88.3% reduction) - Serving Budget: Runs in <3.5 GB VRAM using on-demand storage caching or ~8.2 GB full VRAM on an NVIDIA T4 GPU.
Shard Distribution
model-00001tomodel-00011(~441 MB each): STQ1_0 routed MoE expert layers.model-00000,model-00012,model-00013: FP8 embeddings, vision/dense layers, and attention blocks.
- Downloads last month
- 259
Model tree for SofiTesfay2010/Agents-A1-Flash
Base model
InternScience/Agents-A1