Granite-4.2-3B-NPU2

IBM Granite 4.2 3B quantised to the q4nx container that FastFlowLM loads, for AMD Ryzen AI NPUs (XDNA2 / NPU2).

The base model is IBM's, licensed Apache-2.0. This repository redistributes a quantised conversion of it under the same licence, with attribution. Nothing here is a new model; the weights are IBM's, in a different container.

Use

Requires a FastFlowLM build that includes the granite family. Support is not in an upstream release yet — see the pull request linked below.

flm run granite:3b

or place this directory as Granite-4.2-3B-NPU2 where FastFlowLM looks for models.

What is in here

file
model.q4nx the weights, Q4_1 semantics in the q4nx tile layout (2.6 GB)
config.json granite geometry, plus a record of the folded multipliers
tokenizer.json, tokenizer_config.json as shipped with the source package
chat_template.jinja as shipped with the source package

The tokenizer and chat template are deliberately unmodified. They look mismatched — the tokenizer carries <|start_of_role|> as a single token while the template emits ChatML markers that cost six tokens each — and replacing either makes the model's output worse, which was measured rather than assumed:

template prompt result
as shipped (ChatML) 44 tokens thinking trace, </think>, correct answer
role markers 14 tokens echoes the question first
role markers + <think> 15 tokens URL-encoded garbage

Swapping in upstream granite-4.2's own tokenizer is worse still: it and this one differ on exactly 17 control-token ids out of 100352, so the model receives <|pad|> where a turn marker should be, and the output degrades in a way that stays fluent and is easy to misread as a model problem.

Geometry

Hidden 2560 · 40 layers · 40 query heads over 8 kv heads · head_dim 64 · intermediate 8192 · vocab 100352 · RoPE theta 1e7, half-split · RMSNorm eps 1e-5.

Granite needs head_dim 64 at hidden 2560, and every design FastFlowLM ships at hidden ≥ 2560 is head_dim 128, so it needed an engine of its own.

Provenance

Converted from IBM's published weights with the q4nx-build tooling. The conversion was verified tensor-by-tensor against the GGUF it was fed, and the resulting forward pass was checked against an independent numpy implementation written from config.json — cosine agreement at prefill and at decode steps 1 and 50.

Speed

FastFlowLM's granite engine currently runs on the CPU: about 8.7 tok/s decode on a Ryzen AI 9 HX 370. NPU kernels for this geometry exist and measure 13.6 tok/s of device time for the whole layer stack, but are not yet wired into the C++ dispatch path.

Links

Downloads last month
27
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for OpenFlowLM/Granite-4.2-3B-NPU2

Quantized
(54)
this model