Instructions to use hsmin92/internvl35-fp8 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use hsmin92/internvl35-fp8 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="hsmin92/internvl35-fp8") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("hsmin92/internvl35-fp8") model = AutoModelForMultimodalLM.from_pretrained("hsmin92/internvl35-fp8", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use hsmin92/internvl35-fp8 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "hsmin92/internvl35-fp8" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "hsmin92/internvl35-fp8", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/hsmin92/internvl35-fp8
- SGLang
How to use hsmin92/internvl35-fp8 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "hsmin92/internvl35-fp8" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "hsmin92/internvl35-fp8", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "hsmin92/internvl35-fp8" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "hsmin92/internvl35-fp8", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use hsmin92/internvl35-fp8 with Docker Model Runner:
docker model run hf.co/hsmin92/internvl35-fp8
- InternVL3.5-4B-HF FP8 Dynamic
InternVL3.5-4B-HF FP8 Dynamic
A compressed-tensors FP8 Dynamic quantization of OpenGVLab/InternVL3_5-4B-HF, prepared for memory-conscious vLLM serving.
A companion AWQ W4A16 build of the same base model is at hsmin92/internvl35-4b-awq-w4a16-g128. On an 8 GB card the AWQ build is the better fit — see Choosing between the FP8 and AWQ builds.
⚠️ Do not use --dtype half
--dtype half produces a server that starts normally, returns HTTP 200 from
/health, reports a healthy container health check — and answers every
request with !!!!!!!!. Dequantizing FP8 E4M3 weights into FP16 overflows
the FP16 range, the logits become NaN, and argmax selects token 0.
It fails silently, so a liveness probe will not catch it. Always send a real request after a configuration change, not just a health check.
Use --dtype bfloat16. The checkpoint declares bfloat16 in config.json,
Ampere supports BF16 natively, and the memory cost is identical to FP16.
The bundled
scripts/start_vllm_wsl_rtx3070.shin this repository still passes--dtype halfand is affected. Change that line to--dtype bfloat16before using the script, or use thevllm servecommand in Serving with vLLM instead.
Important runtime note
The checkpoint stores the language decoder Linear weights in FP8 E4M3 format. On an NVIDIA Ampere GPU such as the RTX 3070, vLLM serves these weights through its W8A16 FP8 path: weights stay compressed to 8-bit while activations run in 16-bit. This is a VRAM-saving configuration; a speedup is not guaranteed on Ampere, and in practice this build is 2–4× slower than the AWQ build on the same card.
Quantization scope
Quantized:
- 252 language decoder
Linearmodules - Scheme:
FP8_DYNAMIC - Weight format: FP8 E4M3
- Activation scaling: dynamic per token at runtime
- Calibration dataset: not required — unlike AWQ, this build carries no calibration-domain risk
Kept in BF16:
- Vision tower
- Multimodal projector
- Input embeddings
lm_head- Normalization layers and other protected parameters
Generated from the base-model revision:
6bd4487402110ef9889ba50eb7aefeb302526fed
See quantization/recipe.py for the compression recipe.
Verified environment
| Component | Version / value |
|---|---|
| GPU | NVIDIA GeForce RTX 3070 8 GB (Ampere, SM 8.6) |
| Host | Windows 11 + WSL2 (Ubuntu 24.04), Docker Desktop |
| NVIDIA driver | 591.86 |
| Serving image | vllm/vllm-openai:v0.26.0 |
| vLLM | 0.26.0 (V1 engine) |
| PyTorch | 2.11.0+cu130 |
| Transformers | 5.14.1 |
| Quantization backend | compressed-tensors 0.17.0 |
| Attention backend | FlashInfer with fp8 KV, FlashAttention with fp16 KV |
Startup with the command below:
Model loading took 5.51 GiB memory and 7.46 seconds
GPU KV cache size: 9,088 tokens
Maximum concurrency for 8,192 tokens per request: 1.11x
Total process footprint is about 7.8 GB on an 8 GB card that is also driving a desktop. This build uses essentially the whole card; see VRAM budget.
Serving with vLLM
vllm serve hsmin92/internvl35-fp8 \
--served-model-name internvl35-fp8 \
--dtype bfloat16 \
--max-model-len 8192 \
--max-num-seqs 8 \
--max-num-batched-tokens 4096 \
--limit-mm-per-prompt '{"image":2,"video":0}' \
--gpu-memory-utilization 0.85 \
--kv-cache-memory-bytes 640M \
--kv-cache-dtype fp8 \
--enforce-eager \
--trust-remote-code \
--enable-auto-tool-choice \
--tool-call-parser hermes
Then use the OpenAI-compatible endpoint at http://127.0.0.1:8000/v1.
Each non-obvious flag is explained below.
VRAM budget
Measured, all figures reported by vLLM itself:
weights 5,642 MiB
context + activation peak 604 MiB
KV cache (pinned, fp8) 640 MiB -> 9,088 tokens
────────────────────────────────────────
vLLM process ~6,886 MiB -> --gpu-memory-utilization 0.841
There is almost no slack. Two consequences:
--kv-cache-dtype fp8 is required to reach an 8,192-token context. With
FP16 KV, vLLM reports that only 5,680 tokens fit:
To serve at least one request with the model's max seq len (8192),
1.12 GiB KV cache is needed, which is larger than the available
KV cache memory (0.78 GiB). Based on the available memory, the
estimated maximum model length is 5680.
FP8 KV halves the per-token cost from 144 KiB to 72 KiB. It was validated on Ampere through the auto-selected FlashInfer backend; text, a 7k-token context, and two-image requests all produce correct output.
--enforce-eager is required. CUDA graphs need roughly another 0.5–1 GiB
that this build does not have. The AWQ build does capture them.
--gpu-memory-utilization is a fraction of the free memory the process can
see, which is not what host nvidia-smi reports. Under WSL2 the container
saw 6.93 GiB while the Windows host reported 7.48 GiB free, so the effective
ceiling was 0.866. Setting 0.87 was rejected outright:
Free memory on device cuda:0 (6.93/8.0 GiB) on startup is less than
desired GPU memory utilization (0.87, 6.96 GiB)
Pin the KV cache rather than letting utilization decide
Without --kv-cache-memory-bytes, vLLM expands the KV cache to fill whatever
--gpu-memory-utilization allows, so the cache size drifts between restarts
with whatever the desktop happens to be using — observed swings of several
thousand tokens on an otherwise identical configuration. Pinning it makes
startup deterministic, which matters here because the margin above the
8,192-token requirement is thin.
--max-num-seqs: leaving it at the default costs throughput
The vLLM default is 128. When more concurrent requests arrive than the KV cache can hold, vLLM preempts and recomputes, and throughput collapses below what a smaller limit achieves. With a 640 MiB fp8 cache (9,088 tokens) and typical requests of ~1,084 prompt tokens, about 8 requests fit; use that as the limit.
Do not derive the limit from the Maximum concurrency for 8,192 tokens per request line in the startup log — that figure assumes every request fills the
full context window, which real requests rarely do.
--max-num-batched-tokens also sizes the multimodal encoder cache
This is not a VRAM tuning knob. The vLLM default of 2048 rejects large images outright, on any card:
image item with 2816 embedding tokens, which exceeds the
pre-allocated encoder cache size 2048
InternVL dynamic tiling allows up to 12 patches plus a thumbnail, so a single image can reach 3,329 embedding tokens. Use at least 4096.
Tool-call flags are needed even without tool calling
--enable-auto-tool-choice and --tool-call-parser are required for the
server to accept requests that carry a tool_choice field. Open WebUI sends
tool_choice: "auto" by default, so omitting the flags makes it fail with:
"auto" tool choice requires --enable-auto-tool-choice and
--tool-call-parser to be set
Actual tool calling still does not work: the bundled chat_template.jinja has
no tool rendering, so tools=[...] is silently ignored.
Image token cost is driven by aspect ratio, not size
Tile count is chosen from the aspect ratio, so token cost is not monotonic in resolution. Measured on this model family:
| Input size | Image tokens | Tiles |
|---|---|---|
| 448×448 | 257 | 1 |
| 896×448 | 769 | 3 |
| 960×544 | 769 | 3 |
| 800×450 | 769 | 3 |
| 2048×2048 | 2,561 | 10 |
| 200×150 | 3,329 | 13 |
| 4000×300 | 3,329 | 13 |
A 200×150 crop costs 4.3× more than an 800×450 image, because 4:3 maps exactly onto a 4×3 tile grid and the crop is upscaled to 1792×1344 for no added information. Across randomly sized crops of a 960×544 frame, total prompt cost swung between 1,040 and 4,112 tokens.
When sending region crops, letterbox them into a fixed 448×448 canvas (preserving aspect ratio, no upscaling) to make cost constant. On a small-text reading task this was several times faster than sending the raw crop, with no loss of accuracy down to 14 px glyphs.
Quick start: WSL2 + RTX 3070
The bundled startup script disables the V2 model runner, FlashInfer sampling JIT, and DeepGEMM for Ampere/WSL2 compatibility. It also passes
--dtype half, which must be changed to--dtype bfloat16— see the warning at the top of this card.
1. Install system build requirements
sudo apt update
sudo apt install -y build-essential python3.12-dev
Do not install a Linux NVIDIA display driver inside WSL2. The Windows NVIDIA
driver exposes libcuda.so under /usr/lib/wsl/lib.
2. Create the Python environment
Install uv first when it is not already available.
uv venv --python 3.12 .venv-vllm
source .venv-vllm/bin/activate
uv pip install "vllm==0.26.0" hf_xet
3. Validate the environment
./scripts/check_wsl_runtime.sh
4. Start the server
Prefer the vllm serve command in Serving with vLLM.
The bundled script is kept for the originally validated conservative profile
(2,048-token context, single sequence) and still needs the --dtype fix.
Successful startup ends with:
Application startup complete.
Running in Docker instead
The native script exists to work around WSL2 CUDA/NVRTC linking. Those
workarounds are unnecessary in a container: the official image ships CUDA and
NVRTC, and the NVIDIA container runtime injects libcuda.so.
docker run --rm --gpus all --ipc host -p 8000:8000 \
vllm/vllm-openai:v0.26.0 \
--model hsmin92/internvl35-fp8 \
--dtype bfloat16 --max-model-len 8192 --max-num-seqs 8 \
--max-num-batched-tokens 4096 --gpu-memory-utilization 0.85 \
--kv-cache-memory-bytes 640M --kv-cache-dtype fp8 --enforce-eager \
--limit-mm-per-prompt '{"image":2,"video":0}' --trust-remote-code \
--enable-auto-tool-choice --tool-call-parser hermes
API tests
curl -s http://127.0.0.1:8000/health
curl -s http://127.0.0.1:8000/v1/models | python3 -m json.tool
./examples/chat_text.sh
python examples/chat_image.py /path/to/image.jpg "Describe the scene."
The image client sends the local image as a base64 data URL and uses only the Python standard library.
Choosing between the FP8 and AWQ builds
Measured on the same RTX 3070, same vLLM version, each tuned to its own best settings on this card:
| FP8 Dynamic | AWQ W4A16 G128 | |
|---|---|---|
| Weights on GPU | 5.51 GiB | 3.84 GiB |
| KV cache | 9,088 tokens (fp8 KV) | 12,288 tokens (fp16 KV) |
| KV dtype needed for 8192 ctx | fp8 (compromise) | fp16 |
| CUDA graphs on 8 GB | not possible | yes |
--gpu-memory-utilization needed |
0.85 (near the 0.866 ceiling) | 0.79 |
| Process VRAM | 7.8 GB | 6.6 GB |
| Text generation, single request | 12–18 tok/s | ~2–4× faster |
| Calibration data required | no | yes |
| Weight precision | 8-bit | 4-bit |
On an 8 GB card the AWQ build is the better choice — it is faster, leaves 1.2 GB free, and needs no precision compromises on the KV cache or CUDA graphs.
This FP8 build remains preferable when weight precision matters more than memory: it is 8-bit rather than 4-bit, and it required no calibration data, so it carries no risk of a calibration set that does not match your domain. It is also the better starting point on a larger card where VRAM is not the constraint.
Accuracy has not been compared between the two builds, or against the BF16 base model. Benchmark both on your own data before choosing.
Native Linux and other GPUs
The WSL2 script applies compatibility settings required by the validated RTX 3070 environment:
VLLM_USE_V2_MODEL_RUNNER=0because the V2 runner required UVA in this WSL setup.VLLM_USE_FLASHINFER_SAMPLER=0because FlashInfer sampling JIT requirednvcc.VLLM_USE_DEEP_GEMM=0because DeepGEMM targets Hopper./usr/lib/wsl/libis added to the compile and runtime linker paths.- pip-installed CUDA NVRTC libraries are added to
LD_LIBRARY_PATH.
Native Linux systems with a full CUDA Toolkit, or newer GPUs, may not need these. Start from the documented configuration, then remove compatibility flags one at a time and validate output quality, memory, and stability after each change. On a card with native FP8 Tensor Cores (Ada/Hopper) this checkpoint runs on the true FP8 path rather than the W8A16 compatibility path, and the performance picture changes completely.
Intended use
- Image understanding
- Multi-image comparison (for example a full frame plus a zoomed crop)
- CCTV frame summarization
- Visual question answering
- OpenAI-compatible multimodal API integration
For video on an 8 GB GPU, sample a small number of frames externally, resize them appropriately, and send them as multiple images rather than passing a video.
Limitations
- A quantized derivative, not an independently trained model.
- No accuracy benchmark has been published for this checkpoint. Functional validation only: text generation, a 7k-token context, and multi-image requests all produce coherent, correct answers on simple synthetic probes.
- The vision tower and output head remain BF16 and are a meaningful share of the loaded weights.
- FP8 on Ampere is served through a W8A16 compatibility path, not native FP8 Tensor Core execution.
- The bundled
chat_template.jinjahas no tool-calling support. - VRAM figures depend on driver, desktop applications, context length, multimodal limits, and vLLM version.
Attribution and license
A quantized derivative of OpenGVLab/InternVL3_5-4B-HF. The original project and this derivative are distributed under the Apache-2.0 license. Review the upstream model card for original training details, limitations, and citation information.
Citation
@article{wang2025internvl3_5,
title={InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency},
author={Wang, Weiyun and Gao, Zhangwei and Gu, Lixin and Pu, Hengjun and Cui, Long and Wei, Xingguang and Liu, Zhaoyang and Jing, Linglin and Ye, Shenglong and Shao, Jie and others},
journal={arXiv preprint arXiv:2508.18265},
year={2025}
}
- Downloads last month
- -
Model tree for hsmin92/internvl35-fp8
Base model
OpenGVLab/InternVL3_5-4B-Pretrained