Image-Text-to-Text
NInfer
qwen3.8
nvfp4
w4a4
blackwell
multimodal
conversational
cuda
rtx-5090
Eval Results (legacy)
Instructions to use cometkim/Qwen3.8-27B-nvfp4full-NInfer with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- NInfer
How to use cometkim/Qwen3.8-27B-nvfp4full-NInfer with NInfer:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
Blackwell Pro 4000 SFF / non SFF Prefill and MTP 3-5 Results Text / Code
#1
by Schestex - opened
llama.cpp vs Ninfer
| Context | SFF 70 W | non-SFF 145 W | non-SFF Vorteil |
|---|---|---|---|
| 512 | 2.188 tok/s | 3.575 tok/s | +63 % |
| 2K | 2.635 | 4.571 | +73 % |
| 4K | 2.556 | 4.447 | +74 % |
| 7.680 | 2.397 | 4.157 | +73 % |
| 8K | 2.381 | 4.169 | +75 % |
| 16K | 2.165 | 3.784 | +75 % |
| 32K | 1.836 | 3.221 | +75 % |
| 64K | 1.452 | 2.430 | +67 % |
MTP4 Sweetspot for text
| SFF | MTP3 | MTP4 | MTP5 |
|---|---|---|---|
| Decode | 31,44 | 33,85 tok/s | 30,48 |
| Decode-time | 32,54 s | 30,22 s | 33,56 s |
| Acceptance rate | 38,54 % | 30,03 % | 22,62 % |
| Accepted length | 2,15 | 2,20 | 2,13 |
| Rounds | 475 | 465 | 480 |
| Accepted tokens | 548 | 558 | 542 |
| Fallbacks | 0 | 0 | 1 |
MTP4 Sweetspot for code
| Code-Output | SFF | non-SFF | SSF vs non-SFF |
|---|---|---|---|
| MTP3 | 52,25 | 88,63 tok/s | +69,6 % |
| MTP4 | 64,16 | 105,24 tok/s | +64,0 % |
| MTP5 | 64,61 | 107,96 tok/s | +67,1 % |