Instructions to use Mantisec/Qwen3.5-4B-FP16 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Mantisec/Qwen3.5-4B-FP16 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="Mantisec/Qwen3.5-4B-FP16") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("Mantisec/Qwen3.5-4B-FP16") model = AutoModelForMultimodalLM.from_pretrained("Mantisec/Qwen3.5-4B-FP16", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Mantisec/Qwen3.5-4B-FP16 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Mantisec/Qwen3.5-4B-FP16" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Mantisec/Qwen3.5-4B-FP16", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/Mantisec/Qwen3.5-4B-FP16
- SGLang
How to use Mantisec/Qwen3.5-4B-FP16 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Mantisec/Qwen3.5-4B-FP16" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Mantisec/Qwen3.5-4B-FP16", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Mantisec/Qwen3.5-4B-FP16" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Mantisec/Qwen3.5-4B-FP16", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use Mantisec/Qwen3.5-4B-FP16 with Docker Model Runner:
docker model run hf.co/Mantisec/Qwen3.5-4B-FP16
Qwen3.5-4B-fp16
FP16 conversion of Qwen/Qwen3.5-4B, produced by bfsquish v0.1.0.
Intended use
Optimized for FP16 inference and fine-tuning on NVIDIA V100 (Volta) GPUs, which lack native BF16 Tensor Core support. The weights have been converted from BF16 to FP16 using the effective clamp_only strategy (see below).
Conversion details
| Field | Value |
|---|---|
| Source model | Qwen/Qwen3.5-4B |
| Source revision | latest |
| Requested strategy | auto |
| Effective strategy | clamp_only |
| Tool | bfsquish v0.1.0 |
| Converted at (UTC) | 2026-09-05T13:51:06.034581+00:00 |
| Target hardware | NVIDIA V100 (Volta, sm_70) |
| Target runtime | NVIDIA V100 (Volta sm_70), FP16 Tensor Cores |
Conversion strategy
Direct upcast to FP32, clamp to FP16 range (+/-65504), downcast to FP16. The simplest strategy, near-lossless for well-behaved trained weights whose values are concentrated near zero.
Transform plan
v1 strategy applied: clamp_only (no v2 plan bundle was produced for this conversion).
Enhanced chat template
Prompt rendering uses the selected third-party enhanced template; model weights are unchanged by this step. The standalone and embedded forms were smoke-rendered to the same prompt before publication.
| Field | Value |
|---|---|
| Selection | latest |
| Source | peculiar-ragdoll/Qwen-Sharp-Chat-Templates |
| Resolved revision | fa3a1295882d31132770c156fced4e616b5db25d |
| Template version | qwen3.8-froggeric-v22.4.0 |
| Template license | Apache-2.0 |
| Published forms | chat_template.jinja and tokenizer_config.json#chat_template |
Reproduce the prompt-format step with --chat-template latest --chat-template-repo peculiar-ragdoll/Qwen-Sharp-Chat-Templates --chat-template-revision fa3a1295882d31132770c156fced4e616b5db25d.
Numerical quality
| Metric | Value |
|---|---|
| Validation verdict | PASS |
| Validation method | generate |
| Max abs logit diff (vs. source) | 0.215126 |
| Min cosine similarity (vs. source) | 0.999978 |
| Token agreement rate | 100.00% |
| Inf/NaN scan | passed (no inf/nan) |
Validation notes
- generation agreement passed despite logit drift: token_agreement=100.00% (pass≥98.00%), min_cos=0.999978, max_diff=0.2151
Reproducing this conversion
bfsquish run \
--model Qwen/Qwen3.5-4B \
--output-dir ./out \
--strategy auto
License
Inherited from the source model. Refer to the source model's license for terms of use.
- Downloads last month
- 30