Instructions to use openbmb/MiniCPM-V-4.6-gguf with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use openbmb/MiniCPM-V-4.6-gguf with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="openbmb/MiniCPM-V-4.6-gguf") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("openbmb/MiniCPM-V-4.6-gguf", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use openbmb/MiniCPM-V-4.6-gguf with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf openbmb/MiniCPM-V-4.6-gguf:Q4_K_M # Run inference directly in the terminal: llama cli -hf openbmb/MiniCPM-V-4.6-gguf:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf openbmb/MiniCPM-V-4.6-gguf:Q4_K_M # Run inference directly in the terminal: llama cli -hf openbmb/MiniCPM-V-4.6-gguf:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf openbmb/MiniCPM-V-4.6-gguf:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf openbmb/MiniCPM-V-4.6-gguf:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf openbmb/MiniCPM-V-4.6-gguf:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf openbmb/MiniCPM-V-4.6-gguf:Q4_K_M
Use Docker
docker model run hf.co/openbmb/MiniCPM-V-4.6-gguf:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use openbmb/MiniCPM-V-4.6-gguf with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "openbmb/MiniCPM-V-4.6-gguf" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "openbmb/MiniCPM-V-4.6-gguf", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/openbmb/MiniCPM-V-4.6-gguf:Q4_K_M
- SGLang
How to use openbmb/MiniCPM-V-4.6-gguf with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "openbmb/MiniCPM-V-4.6-gguf" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "openbmb/MiniCPM-V-4.6-gguf", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "openbmb/MiniCPM-V-4.6-gguf" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "openbmb/MiniCPM-V-4.6-gguf", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Ollama
How to use openbmb/MiniCPM-V-4.6-gguf with Ollama:
ollama run hf.co/openbmb/MiniCPM-V-4.6-gguf:Q4_K_M
- Unsloth Studio
How to use openbmb/MiniCPM-V-4.6-gguf with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for openbmb/MiniCPM-V-4.6-gguf to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for openbmb/MiniCPM-V-4.6-gguf to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for openbmb/MiniCPM-V-4.6-gguf to start chatting
- Pi
How to use openbmb/MiniCPM-V-4.6-gguf with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf openbmb/MiniCPM-V-4.6-gguf:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "openbmb/MiniCPM-V-4.6-gguf:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- OpenClaw new
How to use openbmb/MiniCPM-V-4.6-gguf with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf openbmb/MiniCPM-V-4.6-gguf:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "openbmb/MiniCPM-V-4.6-gguf:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Docker Model Runner
How to use openbmb/MiniCPM-V-4.6-gguf with Docker Model Runner:
docker model run hf.co/openbmb/MiniCPM-V-4.6-gguf:Q4_K_M
- Lemonade
How to use openbmb/MiniCPM-V-4.6-gguf with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull openbmb/MiniCPM-V-4.6-gguf:Q4_K_M
Run and chat with the model
lemonade run user.MiniCPM-V-4.6-gguf-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use openbmb/MiniCPM-V-4.6-gguf with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf openbmb/MiniCPM-V-4.6-gguf:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default openbmb/MiniCPM-V-4.6-gguf:Q4_K_M
Run Hermes
hermes
- Atomic Chat
not working with llama.cpp
When I run this model and upload an image, the llama.cpp throws:
In the UI: Error: "This model supports: text files, PDFs"
through the local API: "image input is not supported - hint: if this is unexpected, you may need to provide the mmproj"
{
"messages": [
{
"role": "user",
"content": [
{
"type": "text",
"text": "describe this image"
},
{
"type": "image_url",
"image_url": {
"url": "data:image/png;base64,/9j/4AAQ...
}
}
]
}
],
"model": "MiniCPM-V-4_6-F16",
"mmproj":"mmproj-model-f16"
}
logs:
[51243] 0.01.938.469 I srv init: init: chat template, thinking = 1
[51243] 0.01.938.493 I srv main: model loaded
[51243] 0.01.938.495 I srv main: server is listening on http://127.0.0.1:51243
[51243] cmd_child_to_router:ready
[51243] cmd_child_to_router:info:{"id":"MiniCPM-V-4_6-Q4_K_M","aliases":["MiniCPM-V-4_6-Q4_K_M"],"tags":[],"object":"model","created":1786180099,"owned_by":"llamacpp","meta":{"vocab_type":2,"n_vocab":248094,"n_ctx":262144,"n_ctx_train":262144,"n_embd":1024,"n_params":752161600,"size":518145904}}
0.32.966.450 I srv proxy_reques: proxying request to model MiniCPM-V-4_6-Q4_K_M on port 51243
[51243] 0.01.938.689 I srv update_slots: all slots are idle
[51243] 0.01.938.694 I srv operator(): child server monitoring thread started, waiting for EOF on stdin...
[51243] 0.01.939.675 W srv operator(): got exception: {"error":{"code":500,"message":"image input is not supported - hint: if this is unexpected, you may need to provide the mmproj","type":"server_error"}}
Tested with Q4, then F16, with mmproj:
@fawogin598
This is a usage/configuration issue rather than a model support issue.
The "mmproj" field in the OpenAI request body is not supported and will not load the projector.
The projector must be specified when starting llama.cpp:
llama-server \
-m MiniCPM-V-4_6-Q4_K_M.gguf \
--mmproj mmproj-model-f16.gguf
You can refer to this document for guidance. It might be helpful.
https://github.com/OpenSQZ/MiniCPM-V-CookBook/blob/main/deployment/llama.cpp/minicpm-v4_6_llamacpp.md
