gpt-oss-120b-rngd-tp2-3.0

GPT-OSS-120B MXFP4 compiled for two Furiosa RNGD cards (16 PEs), using modified two-chip kernels.

Compilation

furiosa-llm 2026.4.0b7, compiler 0.11.0-dev (1f95359ba3).

-tp 16 -pp 1 --max-model-len 131072 -O O0 --concurrency 2
--tokenwise-buckets 1,128
--attention-buckets '(1,128,128),(1,4096,128),(1,4096,1),(1,32768,128),(1,32768,1),(1,131072,128),(1,131072,1)'

Serving

Tested with Furiosa SDK 2026.3.0, driver 2026.3.1, and firmware 2026.3.0.

furiosa-llm serve iAcloud/gpt-oss-120b-rngd-tp2-3.0 --devices npu:0,npu:1

Source and license

Original model: openai/gpt-oss-120b.

Model weights are unchanged. See LICENSE for the Apache 2.0 license.

Downloads last month
145
Safetensors
Model size
117B params
Tensor type
BF16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for iAcloud/gpt-oss-120b-rngd-tp2-3.0

Quantized
(131)
this model