Use on OpenRouter 路 Use on Vercel AI Gateway 路 Release blog post
poolside/Laguna-S-2.1-DFlash-INT4
DFlash speculator for the INT4 target poolside/Laguna-S-2.1-INT4. The speculator is a 6-layer Laguna-style draft model (BF16); pair it with the INT4 base for lower-latency serving via speculative decoding.
Trained: e0630_rhiemann_baseline SFT, DFlash Stage-2, 15k steps. Recommended
serving setting: num_speculative_tokens=7.
DFlash upstream support is in progress (vLLM #46853, SGLang #29446, TRT-LLM #15666). Use
poolside/Laguna-S-2.1-INT4 as the target model.
Benchmarks
Measured with TP=2, temperature=0, and num_speculative_tokens=15.
Throughput speedup
| Concurrency | GSM8K | MATH-500 | HumanEval | MBPP | MT-Bench |
|---|---|---|---|---|---|
| 1 | 3.697x | 3.216x | 3.776x | 2.704x | 2.525x |
| 4 | 2.815x | 2.521x | 2.995x | 2.145x | 1.980x |
| 8 | 2.541x | 2.247x | 2.865x | 1.970x | 1.953x |
| 16 | 2.426x | 2.161x | 2.629x | 1.895x | 1.935x |
Acceptance length
| Concurrency | GSM8K | MATH-500 | HumanEval | MBPP | MT-Bench |
|---|---|---|---|---|---|
| 1 | 6.273 | 5.363 | 6.416 | 4.532 | 4.217 |
| 4 | 6.111 | 5.336 | 6.416 | 4.498 | 4.154 |
| 8 | 6.133 | 5.356 | 6.777 | 4.565 | 4.383 |
| 16 | 6.145 | 5.393 | 6.314 | 4.507 | 4.636 |
- Downloads last month
- 16
Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support