poolside-banner

Use on OpenRouterUse on Vercel AI GatewayRelease blog post


poolside/Laguna-S-2.1-DFlash-INT4

DFlash speculator for the INT4 target poolside/Laguna-S-2.1-INT4. The speculator is a 6-layer Laguna-style draft model (BF16); pair it with the INT4 base for lower-latency serving via speculative decoding.

Trained: e0630_rhiemann_baseline SFT, DFlash Stage-2, 15k steps. Recommended serving setting: num_speculative_tokens=7. DFlash upstream support is in progress (vLLM #46853, SGLang #29446, TRT-LLM #15666). Use poolside/Laguna-S-2.1-INT4 as the target model.

Benchmarks

Measured with TP=2, temperature=0, and num_speculative_tokens=15.

Throughput speedup

Concurrency GSM8K MATH-500 HumanEval MBPP MT-Bench
1 3.697x 3.216x 3.776x 2.704x 2.525x
4 2.815x 2.521x 2.995x 2.145x 1.980x
8 2.541x 2.247x 2.865x 1.970x 1.953x
16 2.426x 2.161x 2.629x 1.895x 1.935x

Acceptance length

Concurrency GSM8K MATH-500 HumanEval MBPP MT-Bench
1 6.273 5.363 6.416 4.532 4.217
4 6.111 5.336 6.416 4.498 4.154
8 6.133 5.356 6.777 4.565 4.383
16 6.145 5.393 6.314 4.507 4.636
Downloads last month
16
Safetensors
Model size
1B params
Tensor type
BF16
Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support

Model tree for EazyDuzIt/Laguna-S-2.1-DFlash-INT4

Finetuned
(2)
this model