Qwen3-VL-4B-Instruct INT4 OpenVINO

OpenVINO IR export of Qwen/Qwen3-VL-4B-Instruct for the Piccolo AI OpenVINO Model Server provider.

Provenance

  • Source revision: ebb281ec70b05090aa6165b016eac8ec08e71b17
  • OVMS export requirements: tag v2026.2, commit 9b795c5fad08cc4abf06a8f751b80a5fc5ae1001
  • Optimum Intel export commit: d4dd21a3aa89c0671d85b704847ac06a378e761c
  • OpenVINO: 2026.2.0rc2
  • OpenVINO Tokenizers: 2026.2.0.0rc2
  • Transformers: 5.0.0
  • NNCF: 3.2.0

Compression

The language model was exported with symmetric INT4 weight compression:

  • ratio: 1.0
  • group size: 128
  • all 252 ratio-defining language layers: INT4_SYM

Auxiliary vision/exporter components use the exporter's per-channel INT8 compression where applicable.

Validation

The exported artifact was validated with the Piccolo AI OVMS provider 0.1.3 and OVMS 2026.2 on an Intel GPU target:

  • all seven OpenVINO IR graphs load with the OpenVINO Tokenizers extension;
  • the language graph contains 36 ScaledDotProductAttention operations;
  • the vision-merger graph contains 24 ScaledDotProductAttention operations;
  • strict Intel GPU startup reached AVAILABLE;
  • streamed text generation and synthetic-image understanding passed;
  • concurrent text and vision requests completed without backend crashes, GPU resets, or OOM kills;
  • text aggregate generation throughput rose from 5.65 tok/s at concurrency 1 to 11.52 tok/s at concurrency 8 for a fixed 128-token completion test;
  • the deployment cgroup peaked at 8.02 GiB, with 0 max-limit events, 0 OOM events, and 0 swap use.

These are bounded validation measurements from one Piccolo appliance and test workload. They are not universal performance claims. Higher concurrency improved aggregate throughput but increased queue latency and crossed the deployment's 8.005 GiB soft memory boundary.

Serving contract

Piccolo mounts this repository read-only at /models/model. The Piccolo AI provider exposes it through the OpenAI-compatible /v3 API under the stable model identifier piccolo-chat.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for boris93/Qwen3-VL-4B-Instruct-int4-ov

Finetuned
(361)
this model