Instructions to use litert-community/North-Micro-Vision-Instruct with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- LiteRT-LM
How to use litert-community/North-Micro-Vision-Instruct with LiteRT-LM:
# LiteRT-LM runs on various platforms (Android, iOS, Windows, Linux, macOS, IoT, Web/WASM) # and supports many APIs (C++, Python, Kotlin, Swift, JavaScript, Flutter). # For platform-specific integration guides, please refer to the official developer website: # https://ai.google.dev/edge/litert-lm # To try LiteRT-LM, the easiest way is to use our CLI tool. # 1. Install the LiteRT-LM CLI tool: pip install -U litert-lm # 2. Download and run this model locally: # See: https://ai.google.dev/edge/litert-lm/cli litert-lm run \ --from-huggingface-repo=litert-community/North-Micro-Vision-Instruct \ --prompt="Write me a poem"
- LiteRT
How to use litert-community/North-Micro-Vision-Instruct with LiteRT:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
North-Micro-Vision-Instruct β LiteRT-LM (on-device Vision-Language Model)
CohereLabs/North-Micro-Vision-Instruct converted to the LiteRT-LM (.litertlm) format for on-device image+text inference with Google's LiteRT-LM runtime β the first Cohere-family model in this format.
North-Micro-Vision is Cohere's 2.48B VLM: a 400M SigLIP2-SO400M-scale vision tower (Qwen3-VL-style encoder with DeepStack mergers) feeding a 2B Cohere decoder, trained for 11 languages. This bundle runs it through LiteRT-LM's fast_vlm path β attach an image, ask a question, get a grounded answer fully on-device.
| Files | North-Micro-Vision-Instruct_wi8.litertlm (3.07 GB, primary) Β· North-Micro-Vision-Instruct_int4.litertlm (2.19 GB, size-constrained variant) |
| Vision | 27-block ViT (hidden 1152, 16 heads, patch 16) made static 512Γ512 β 1024 patches β 2Γ2 merge β 256 image tokens; DeepStack (3 extra vision embeddings) folded into the single image embedding; int8 weights |
| Adapter | Patch merger + the three DeepStack mergers, summed; int8; output at the 2048 text hidden size |
| Decoder | 2B Cohere decoder (28L, hidden 2048, GQA kv8, parallel attn+MLP blocks, sliding/full 3:1, tied 262k-vocab embedding) β int8 dynamic weights (wi8) or int4 blockwise-32; int8 externalized embedder; mixed-precision activations (fp32_fp16) declared in-bundle |
| Context (KV cache) | 4096 |
| Image input | resized to 512Γ512 (normalization (x/255β0.5)/0.5 baked into the encoder) |
| Chat format | Cohere turn tokens (`< |
| Base model | CohereLabs/North-Micro-Vision-Instruct (Apache-2.0) |
Performance (measured)
Apple M4 Max, litert-lm 0.16.0, litert-lm benchmark -p 256 -d 256 --runs 3 --cache no, wi8 bundle:
| Backend | Prefill (tok/s) | Decode (tok/s) | Init (s) |
|---|---|---|---|
| CPU | 671.8 | 27.3 | 6.7 |
| GPU | 1236.6 | 80.6 | 1.8 |
(GPU rows verified with a real image generation on --backend gpu --vision-backend gpu, not benchmark-only.)
Pixel 8a (litert-lm v0.16.1 CLI, --disable_cache), wi8 bundle β the decoder fully delegates to the OpenCL GPU (1380/1380 prefill, 1204/1204 decode nodes):
| Config | Prefill (tok/s) | Decode (tok/s) |
|---|---|---|
| decoder GPU + vision CPU, image turn (271-token prefill) | 124.5 | 4.3β4.4 |
| decoder GPU, text-only (255-token prefill) | 181.1 | 6.4 |
Run the decoder on the GPU on 8 GB-class phones β the CPU backend pages against the 2.5 GB decoder (0.6 tok/s).
iPhone 17 Pro (Metal decoder + Metal vision, cold, this exact 3.07 GB file): peak resident 2.9β3.0 GB with the vision executor loaded (2.1 GB text-only); decode 13.6 tok/s on a 39-token text turn; vision time-to-first-token 2.3β5.0 s cold (302-token image prompt, prefill 322β362 tok/s). Same phone, same model through Apple's Core AI runtime (own measurement, 2026-08-14, int8 decoder + fp16 tower): 18.2 tok/s decode, 21.5 tok/s prefill, image oracle 24/24. Read it as "same class, Core AI ahead on decode (~1.3Γ), LiteRT-LM ahead on prefill and vision TTFT"; both numbers are cold single runs on the same device and the same day-class build, not a tuned benchmark.
On-device demo (Pixel 8a, this exact file, decoder on the OpenCL GPU + vision on CPU)
![]() |
![]() |
| Q: Describe this image. A: "The image shows two cats lying on a pink surface, which appears to be a couch or a cushion. Both cats are in a relaxed position, with their bodies stretched out and their heads resting on the surface. They are both wearing collars, indicating they are pets. The cats are positioned close to each other, suggesting a sense of companionshipβ¦" |
Q: What is in this image? Mention the colors. A: "The image shows a vintage Volkswagen Beetle car parked in front of a building with a yellow wall and two wooden doors. The car is painted in a light teal or turquoise color, and it has a classic design with round wheels and a rounded body shape. The building behind the car has a weathered appearance⦠The ground is paved with rectangular tiles." |
| Q: How many cats are there, and what is between them? A: "There are two cats in the image, and they are both lying on a pink surface. Between them, there are two remote controls." |
Q: Where is the car parked, and what is behind it? A: "The car is parked on a paved street, and behind it is a yellow wall with two wooden doors. The setting appears to be an urban or semi-urban area." |
Verbatim outputs from litert_lm_advanced_main (litert-lm v0.16.1, --backend=gpu --vision_backend=cpu --disable_cache, greedy, 80-token cap) on a Pixel 8a; per turn: TTFT 2.5β2.8 s (β270-token image prompt, prefill 110β128 tok/s), decode 2.9β3.6 tok/s. The photos are the Hugging Face documentation sample images (COCO cats / Beetle), resized by the runtime to the bundle's 512Γ512.
Quality
- 9-case COCO suite (3 images Γ 3 questions, 48-token greedy) against a fp32 PyTorch oracle running the same single-embedding / 1-D-position contract: content-correct and image-grounded on 9/9 (cats on a pink couch with remotes, the two kitchens, colour palettes, "where is this scene"); the int8 vision encoder shifts token choices, so token-exact is 1/9 (the desktop fp16-vision build keeps 5/9 token-exact).
- 8-question text gate: 7/8, non-degenerate. The one miss ("17 + 25" read as "1.7 + 2.5") reproduces token-for-token on the HF fp32 model β Cohere's per-digit pre-tokenizer, not a conversion artifact.
- Device: Pixel 8a Ask-Image 2/2 grounded ("two cats lying on a pink surface, possibly a couch or a bedβ¦", "warm brown wooden counter, black stove, white apron, hanging pots and pans"); iPhone 17 Pro vision-grounding probes 2/2 ("Does this image contain visible written text?" β No on the no-text fractal, Yes on the text probe), 8-question text gate 6/8 on-device (the same digit quirk plus one rhyme miss).
- int4 variant: coherent and content-correct on all 9 suite cases but terser and further from the fp32 wording; prefer wi8 where storage allows.
What the fast_vlm contract changes, and what it costs. The released model injects three DeepStack vision embeddings after decoder layers 0/1/2 and uses interleaved M-RoPE. This bundle folds the DeepStack embeddings into the single image embedding (exactly representable; teacher-forced top-1 vs the released model 0.96 fold-only, 0.93 with the runtime's 1-D positions) and the runtime supplies plain sequential positions in place of M-RoPE. Measured effect on probe prompts: describe / VQA / spatial relations / single-cell lookup preserved; 2-D table cross-cell questions and digit-dense OCR degrade (row count off-by-one, "$652,000" read as "$652,000,000", a duplicated word in a dense paragraph). Same class of trade as the Qwen2-VL-2B bundle. Use it for reading and describing; don't rely on it to rank table cells.
One image per chat. Send each image in a fresh conversation, as with the other fast_vlm bundles.
Run on Android β Google AI Edge Gallery
Install a recent Google AI Edge Gallery, download North-Micro-Vision-Instruct_wi8.litertlm, import it (tap +, enable "Support image"), attach an image and ask. Choose the GPU accelerator for the decoder.
CLI:
litert-lm run North-Micro-Vision-Instruct_wi8.litertlm \
--prompt "What is in this image?" --attachment photo.jpg \
--backend gpu --vision-backend cpu
Run on iPhone / macOS
Use the LiteRT-LM Swift runtime (swift-litert-lm). Load the bundle with the vision tower enabled (Modality.textImage), attach a photo, and ask. Vision-only bundle (no audio tower): bring the engine up with the vision modality only.
Conversion notes
- LiteRT-LM
fast_vlmbundle: VISION_ENCODER ([1,512,512,3]β[1,1024,4608]β the final block plus the three DeepStack taps, concatenated) + VISION_ADAPTER ([1,1024,4608]β[1,256,2048]) + single-token EMBEDDER + PREFILL_DECODE (embeddings-input, cache 4096). - DeepStack fold. The three DeepStack mergers' outputs are added to the main merger output inside the adapter, so the decoder receives one embedding per image token that carries all four vision taps.
- Static rewrite of the dynamic-res tower: fixed 32Γ32 grid, precomputed 2-D rope and resampled learned position embedding, full attention;
Conv3d(temporal 2) folded toConv2dwith the summed temporal kernel; patches kept in raster order through the encoder with the 2Γ2 merge done by strided slices + concat in the adapter (no GATHER_ND, which the mobile GPU delegate cannot compile). Every activation keeps a leading batch dim (rank β₯3) β required for correct results on the Metal GPU delegate. LayerNorm inputs are pre-scaled by calibrated powers of two so the tower survives fp16 GPU precision. - Decoder re-host. The Cohere decoder is re-hosted as a standalone
Cohere2ForCausalLM(parallel block, mean-subtracting LayerNorm, NoPE full-attention layers, logit scale 0.25, tied head) with its rotary layout patched to the checkpoint's half-split convention β text-only logits identical to the original (max |Ξ| = 0.0). The decoder section declaresprefer_activation_type = fp32_fp16: some mobile GPUs accumulate fp16 and overflow at the image-token positions, which blanks the vision conditioning (the model answers as if it saw nothing); the in-bundle declaration selects mixed precision automatically. - Tokenizer: the HF
tokenizer.jsonis bundled as-is (byte-level BPE, 255k + 38 special tokens). - Reproduce: hf-to-litertlm β
bash scripts/reproduce_vlm.sh north-micro-vision(REPRODUCE.mdhas the full recipe and gotchas).
License
Apache-2.0, inherited from the base model CohereLabs/North-Micro-Vision-Instruct.
- Downloads last month
- 26
Model tree for litert-community/North-Micro-Vision-Instruct
Base model
CohereLabs/North-Micro-Vision-Instruct
