Instructions to use ewin-reg/MiniCPM5-V-1B-unofficial with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ewin-reg/MiniCPM5-V-1B-unofficial with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="ewin-reg/MiniCPM5-V-1B-unofficial")# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("ewin-reg/MiniCPM5-V-1B-unofficial", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use ewin-reg/MiniCPM5-V-1B-unofficial with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ewin-reg/MiniCPM5-V-1B-unofficial" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ewin-reg/MiniCPM5-V-1B-unofficial", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/ewin-reg/MiniCPM5-V-1B-unofficial
- SGLang
How to use ewin-reg/MiniCPM5-V-1B-unofficial with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "ewin-reg/MiniCPM5-V-1B-unofficial" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ewin-reg/MiniCPM5-V-1B-unofficial", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "ewin-reg/MiniCPM5-V-1B-unofficial" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ewin-reg/MiniCPM5-V-1B-unofficial", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use ewin-reg/MiniCPM5-V-1B-unofficial with Docker Model Runner:
docker model run hf.co/ewin-reg/MiniCPM5-V-1B-unofficial
MiniCPM5-V-1B (unofficial, experimental)
This is an unofficial, community/personal experiment. It is not affiliated with, endorsed by, or
produced by OpenBMB. OpenBMB has never released a vision-capable version of MiniCPM5-1B (their
MiniCPM-V vision line is built on a different backbone, Qwen3.5, not MiniCPM5). This repo grafts
a vision encoder onto MiniCPM5-1B and trains a connector from scratch, because no such model
existed anywhere.
What this is
- Base LLM:
openbmb/MiniCPM5-1B(apache-2.0), weights unmodified. - Vision encoder:
google/siglip2-base-patch16-512, weights unmodified, frozen throughout training. - Connector: a single linear projector (pixel-shuffle x4 +
Linear(12288, 1536), ~18.9M params) trained from random initialization. This is the only thing actually trained here. - Code: the vision encoder loading/wiring and pixel-shuffle projector structure are adapted
from nanoVLM (MIT license) โ reimplemented, not copied
verbatim, to plug into MiniCPM5's
transformersLlamaForCausalLMinterface instead of nanoVLM's own SmolLM2-based decoder.
What this is NOT
- Not a finished VLM. This is a single stage-1 (LLaVA-style) alignment run: 10,000 image-caption pairs from Flickr30k, 1 epoch (1,250 steps), batch size 8, projector-only gradient (both the vision encoder and the LLM stay frozen). Real VLM releases (including OpenBMB's own MiniCPM-V line) train on millions of image-text pairs across multiple stages. This model has seen roughly 0.1-1% as much data.
- Not benchmarked. No BFCL/MME/OCRBench/etc. numbers are reported here, because none were run. Anyone citing accuracy numbers for this model is making them up โ don't.
- Not reliably accurate. See examples below: it correctly identifies broad scene content (people, objects, rough activity) at a noticeably-better-than-random rate, but regularly gets specific details wrong (color, exact object, count).
Training details
- Data: 10,000 (image, caption) pairs,
nlphuji/flickr30ktest split, no data augmentation. - Optimizer: AdamW, lr 1e-3, linear warmup + linear decay, batch size 8, 1 epoch (1,250 steps).
- Only the projector's
18.9M parameters received gradient updates. SigLIP2 (86.4M params) and MiniCPM5-1B (~1.1B params) were frozen for the entire run. - Loss (next-token cross-entropy on caption tokens only): started at 7.95 (first 20-step avg), ended at 2.78 (last 20-step avg) โ a real, monotonic-ish decrease, not noise. Still slowly decreasing at the end of the run, meaning more training would likely help further; this is an early checkpoint, not a converged one.
- Hardware: single Colab T4 GPU, ~48 minutes wall-clock.
Qualitative examples (held-out images, not seen during training)
| Ground-truth caption | Model output |
|---|---|
| "A young woman lifts her voice in song in a field near an old wooden structure." | "A girl wearing a red shirt and jeans is playing a guitar." |
| "A tan dog climbing over a bare tree trunk with a man standing in the background" | "A man in a red jacket and brown hair is standing on a tree with a dog." |
| "This is a lady in a blue outfit... holding a tennis racket in her hand." | "A woman wearing a red dress... is holding a tennis ball." |
| "Young boy in a white striped shirt and headband holding a tennis racket." | "A young boy wearing a white shirt and shorts is playing a game with a ball." |
The model picks up real scene elements (dog + tree + man; tennis-adjacent activity + ball/racket) but gets specifics wrong (colors, exact object). This is the expected shape of output for a lightly-trained stage-1 projector, not a bug.
Usage
Requires the base MiniCPM5-1B and SigLIP2 weights (downloaded automatically via transformers/
huggingface_hub) plus the small model.py / vision_encoder.py wrapper and this repo's trained
projector.safetensors.
import torch
from huggingface_hub import hf_hub_download
from safetensors.torch import load_file
# model.py / vision_encoder.py from this repo define MiniCPM5V
from model import MiniCPM5V
model = MiniCPM5V().to("cuda")
proj_path = hf_hub_download("ewinregirgojr/MiniCPM5-V-1B-unofficial", "projector.safetensors")
model.projector.load_state_dict(load_file(proj_path))
model.eval()
from PIL import Image
import numpy as np
def img_transform(img):
img = img.resize((512, 512), Image.BICUBIC)
arr = (np.asarray(img, dtype=np.float32) / 255.0 - 0.5) / 0.5
return torch.from_numpy(arr).permute(2, 0, 1).contiguous()
image = Image.open("your_image.jpg").convert("RGB")
pixel_values = img_transform(image).unsqueeze(0).to("cuda", dtype=torch.bfloat16)
ids = model.make_prompt_ids("Describe the image.\n", "cuda")
attn = torch.ones_like(ids)
out = model.generate(input_ids=ids, attention_mask=attn, pixel_values=pixel_values, max_new_tokens=40)
print(model.tokenizer.decode(out[0], skip_special_tokens=True))
Limitations
- English captioning only tested; no multilingual, multi-image, video, or OCR capability trained.
- Single-image only.
- No safety/alignment tuning specific to vision inputs.
- Small stage-1 run โ expect hallucination and detail errors, not a reliable VQA/captioning tool.
License
MiniCPM5-1B and SigLIP2 weights are used unmodified under their original licenses (both
apache-2.0). This repo's own contribution (the projector weights and wrapper code) is released
under apache-2.0. Code structure for the vision encoder/projector is adapted from nanoVLM
(MIT license, https://github.com/huggingface/nanoVLM) โ changes: reimplemented to target
MiniCPM5-1B's transformers interface instead of nanoVLM's native decoder, image-token handling
adjusted to MiniCPM5's tokenizer/vocab.
Not affiliated with or endorsed by OpenBMB, Google, or the nanoVLM authors.
Model tree for ewin-reg/MiniCPM5-V-1B-unofficial
Base model
google/siglip2-base-patch16-512