gpt-oss-120b-rngd-tp2-3.0
GPT-OSS-120B MXFP4 compiled for two Furiosa RNGD cards (16 PEs), using modified two-chip kernels.
Compilation
furiosa-llm 2026.4.0b7, compiler 0.11.0-dev (1f95359ba3).
-tp 16 -pp 1 --max-model-len 131072 -O O0 --concurrency 2
--tokenwise-buckets 1,128
--attention-buckets '(1,128,128),(1,4096,128),(1,4096,1),(1,32768,128),(1,32768,1),(1,131072,128),(1,131072,1)'
Serving
Tested with Furiosa SDK 2026.3.0, driver 2026.3.1, and firmware 2026.3.0.
furiosa-llm serve iAcloud/gpt-oss-120b-rngd-tp2-3.0 --devices npu:0,npu:1
Source and license
Original model: openai/gpt-oss-120b.
Model weights are unchanged. See LICENSE for the Apache 2.0 license.
- Downloads last month
- 145
Model tree for iAcloud/gpt-oss-120b-rngd-tp2-3.0
Base model
openai/gpt-oss-120b