Nemotron-3-Nano-4B β€” LiteRT-LM

nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16 converted to the LiteRT-LM (.litertlm) format for on-device inference with Google's LiteRT-LM runtime. Requires litert-lm β‰₯ 0.15. To our knowledge this is the first Nemotron-3-Nano in LiteRT form.

A reasoning model (<think>, ChatML turns) on a three-kind hybrid stack: 21 Mamba2 selective-scan layers + 17 plain MLP layers + 4 grouped-query attention layers (42 total). Only the 4 attention layers keep KV (4096-token budget here), the mamba layers carry constant-size conv + SSM state, and the MLP layers carry no state at all β€” 50 state buffers in total (42 mamba, 8 KV), so memory stays nearly flat with context length.

File Recipe Size
Nemotron-3-Nano-4B_int8.litertlm int8 dynamic on linears + embedding (convs and the scan stay float); fp32 activations declared 4.13 GB

Geometry: hidden 3136, 40 query / 8 KV heads, mamba 96 heads Γ— 80 dim (state 128, conv 4, 8 groups), vocab 131,072, untied embeddings.

Correctness

All numbers below were measured on this exact file (litert-lm 0.16.0, Apple M4 Max).

  • 8-question sanity gate: 7/8 on CPU, 8/8 on GPU, non-degenerate on both. The single CPU miss is the rhyme-completion item β€” it answers "violets are purple" where the gate wants "blue"; the GPU run answers "blue". Every arithmetic, factual, and translation item is correct on both backends.
  • Chat template is byte-equal to the source: the embedded Jinja matches the repo's chat_template.jinja exactly (10,504 / 10,504 bytes). Note the source repo's tokenizer_config.json carries a different 10,497-byte copy; the bundle embeds the one AutoTokenizer actually resolves.
  • Turn-end stop tokens are <|im_end|> (id 11) alongside the exported id 2.
  • No spurious start token. The source tokenizer sets add_bos_token: False and the template never renders a leading BOS, so the <s> the bundler would otherwise prepend is dropped β€” the on-device token stream matches the training stream. Honest note: at this scale the model is robust either way (the gate scores 7/8 with the token and 7/8 without, and greedy decoding in PyTorch is byte-identical on 2 of 3 probes), so this is a correctness-of-convention fix rather than a rescue.

Usage

# CPU
litert-lm run ./Nemotron-3-Nano-4B_int8.litertlm --prompt "What is the capital of France? Answer in one word."

# GPU β€” pass --cache no (see the honest note below)
litert-lm run ./Nemotron-3-Nano-4B_int8.litertlm --backend gpu --cache no --prompt "..."

The bundle carries the tokenizer and the stock Nemotron-3-Nano chat template. Seven prefill signatures (1024/256/64/16/4/1 + decode) are exported so the runtime picks tight chunks.

Performance

litert-lm benchmark <file> -p 256 -d 256 --runs 3 --cache no, litert-lm 0.16.0, Apple M4 Max. Two independent runs:

Backend Prefill (256) Decode TTFT
GPU (--cache no) 803 / 792 tok/s 83.3 / 82.4 tok/s 0.33 s
CPU 99.6 / 113.3 tok/s 22.7 / 22.5 tok/s 2.64 / 2.30 s

Both figures per cell are the two runs, not a range estimate. GPU repeats within ~1.4%; CPU prefill spreads ~13%, partly because the host was not idle during these runs (another export was using the machine) β€” read the CPU column as an order of magnitude, not a precise figure. --cache no matters for more than tidiness here: with the compiled-graph cache the benchmark reports a much faster CPU prefill because it is not doing the same work.

Honest notes

  • GPU requires --cache no on this bundle. With the compiled-graph cache enabled, litert-lm run --backend gpu fails with WebGPU Invalid BindGroup validation errors, and an 8-question sweep through the Mac verify harness returns token soup (0/8). The same file with --cache no answers 8/8. Measured as a one-variable comparison β€” same runner, same file, cache flag flipped β€” so the cache path is where it goes wrong; the root cause is not isolated further here.
  • Not measured on a phone yet. The desktop numbers above are Mac-only. A 4B of this shape did not fit an 8 GB Android phone when the sibling Nemotron-H-4B was measured, so expect to need a higher-RAM device; that is an expectation carried over from a different bundle, not a measurement of this one.
  • It is a reasoning model. Answers arrive after a <think> block, so give it a token budget that fits the thought (the gate above used 3200).
  • int8 is applied to linears and the embedding only; the convolutions and the selective scan stay float, which is what keeps the hybrid state numerically sane.

Conversion notes

Converted with litert-torch plus a hybrid-cache patch. One command, no per-model work β€” the reproduction script, the patch, and the full measurement record are in hf-to-litertlm:

python scripts/convert.py nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16

Two things that route this model correctly and are worth knowing if you convert your own:

  • auto_map in a config is not proof of remote code. This repo declares auto_map, but transformers registers nemotron_h natively, so without trust_remote_code the library implementation loads and the repo's Python is never imported. A converter that refuses on auto_map alone will refuse this model for no reason.
  • β‰₯3B exports use a reduced 7-signature prefill ladder. Every exported signature costs engine RAM whether or not it is called, and a 4B hybrid with the full 11-signature ladder is exactly the shape that trips memory limits at GPU program init.

Conversion took 1645 s on an M4 Max. See REPRODUCE.md for the Nemotron-H family recipe and the measurements behind every claim on this card.

Downloads last month
11
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for litert-community/Nemotron-3-Nano-4B