Qwen/Qwen3-Coder-Next-Int8-ov

Base model converted using NoLlama project.

Model conversion details:

    Name in API/UI : Qwen3-Coder-Next      (from directory name)
    Kind           : LLM (text)
    Architecture   : Qwen3NextForCausalLM / qwen3_next
    Weights        : INT8 (asymmetric, channel-wise)   74.4 GB on disk
    MoE            : 512 experts, 10 active per token
    Geometry       : 48 layers, 262,144-token context, 96 KB/token KV
    Exported with  : OpenVINO 2026.2.1-21919-ede283a88e3-releases/2026/2, optimum-intel 2.0.0, transformers 4.57.6
    Agent mode     : tool calling on GPU/CPU; never on NPU (hard prompt cap)
    Integrity      : weights complete

HW details:

  • Intel Core Ultra i9 285H
  • 128GB RAM DDR5
  • Intel ARC T140 shared VRAM - 102GB (80%)
  • Windows drive C pagefile.sys settings: 400GB <-- this only was used for coversion from base model to int8
  • Free space on disk: 400GB <-- this only was used for coversion from base model to int8

Using details:

You need to set shared VRAM size 114GB (90%) to set KV cache 24GB.

Downloads last month
8
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for dmitriyteteruk/Qwen3-Coder-Next-int8-ov

Finetuned
(35)
this model