Instructions to use HugC/VisME-Qwen25VL-7B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use HugC/VisME-Qwen25VL-7B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="HugC/VisME-Qwen25VL-7B")# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("HugC/VisME-Qwen25VL-7B") model = AutoModelForMultimodalLM.from_pretrained("HugC/VisME-Qwen25VL-7B", device_map="auto") - Notebooks
- Google Colab
- Kaggle
VisME (Qwen2.5-VL-7B)
English | ç®€ä½“ä¸æ–‡
VisME is a universal multimodal embedding model from our CVPR 2026 paper Illuminating Visual Identity in Universal Multimodal Embeddings. Built on Qwen2.5-VL-7B-Instruct, VisME produces dense embeddings for images, text, and image–text pairs, with strong performance on both general multimodal benchmarks (MMEB) and identity-centric retrieval (MVEB).
Highlights
- Identity-aware retrieval — trained with visual-identity supervision via the MVEB benchmark (4 meta-tasks, 28 test subsets).
- Strong general capability — jointly trained on MMEB-v1 and MVEB, maintaining competitive performance on standard MMEB tasks.
Performance
Results on MMEB (Cls / VQA / Ret / Grd) and MVEB (ID-Rec / Re-ID / ID-Grd / ID-Edit). Metrics are reported as in the paper (average Precision@1 across datasets within each task group).
Training Data
VisME is trained on a mixture of:
| Benchmark | Role |
|---|---|
| MMEB-v1 | General multimodal embedding tasks (classification, VQA, retrieval, grounding) |
| MVEB | Visual-identity retrieval across recognition, re-ID, grounding, and editing |
The MVEB datasets are now open-sourced. Training code will be released soon.
Usage
import torch
import torch.nn.functional as F
from io import BytesIO
from urllib.request import Request, urlopen
from PIL import Image
from visme import VisME
model = VisME("path/to/this/repo")
model = model.cuda().eval()
instruction = "Represent the face with the following text."
text = "Retrieve all images with the same cartoon character."
samples = {
"SpiderMan_comic_E616": "https://static.wikia.nocookie.net/spiderman/images/a/ad/Peter_Parker_%28Earth-616%29_017.png/revision/latest?cb=20210807043502",
"SpiderMan_promo_E199999": "https://static.wikia.nocookie.net/marveldatabase/images/2/28/Peter_Parker_%28Earth-199999%29_from_Spider-Man_No_Way_Home_promotional_art_002.jpg/revision/latest/scale-to-width-down/1000?cb=20230730084204",
"Toxin": "https://static.wikia.nocookie.net/superheroes/images/5/53/Toxin.jpg/revision/latest?cb=20240813171932",
}
def load_image(url: str) -> Image.Image:
req = Request(url, headers={"User-Agent": "Mozilla/5.0"})
with urlopen(req, timeout=30) as resp:
return Image.open(BytesIO(resp.read())).convert("RGB")
names = list(samples.keys())
batch = [
{"image": load_image(samples[name]), "text": text, "instruction": instruction}
for name in names
]
with torch.no_grad():
embeddings = model.encode_input(batch) # shape: (3, 3584), L2-normalized
# Pairwise cosine similarity
# SpiderMan_comic_E616 <-> SpiderMan_promo_E199999
# SpiderMan_comic_E616 <-> Toxin
# SpiderMan_promo_E199999 <-> Toxin
for i in range(len(names)):
for j in range(i + 1, len(names)):
sim = F.cosine_similarity(embeddings[i : i + 1], embeddings[j : j + 1]).item()
print(f"cosine_similarity({names[i]}, {names[j]}) = {sim:.4f}")
Evaluation
Benchmark evaluation code for MMEB and MVEB is coming soon. In the meantime, please refer to the paper for full experimental settings and baselines.
Citation
If you find VisME or MVEB useful, please cite:
@inproceedings{cao2026illuminating,
title={Illuminating Visual Identity in Universal Multimodal Embeddings},
author={Cao, Jiawei and Feng, Junyi and Hua, Jiashen and Huang, Ziheng and Deng, Bing and Wu, Kaijie and Gu, Chaochen and Ye, Jieping},
booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition},
pages={8737--8748},
year={2026}
}
License
This model is built upon Qwen2.5-VL-7B-Instruct. Please follow the license terms of the base model and the respective datasets used during training.
- Downloads last month
- 19
Model tree for HugC/VisME-Qwen25VL-7B
Base model
Qwen/Qwen2.5-VL-7B-Instruct