Instructions to use VertexAGI/prism-caption-3-micro with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use VertexAGI/prism-caption-3-micro with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("VertexAGI/prism-caption-3-micro") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use VertexAGI/prism-caption-3-micro with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf VertexAGI/prism-caption-3-micro:Q4_K_M # Run inference directly in the terminal: llama cli -hf VertexAGI/prism-caption-3-micro:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf VertexAGI/prism-caption-3-micro:Q4_K_M # Run inference directly in the terminal: llama cli -hf VertexAGI/prism-caption-3-micro:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf VertexAGI/prism-caption-3-micro:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf VertexAGI/prism-caption-3-micro:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf VertexAGI/prism-caption-3-micro:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf VertexAGI/prism-caption-3-micro:Q4_K_M
Use Docker
docker model run hf.co/VertexAGI/prism-caption-3-micro:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use VertexAGI/prism-caption-3-micro with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "VertexAGI/prism-caption-3-micro" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "VertexAGI/prism-caption-3-micro", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/VertexAGI/prism-caption-3-micro:Q4_K_M
- Ollama
How to use VertexAGI/prism-caption-3-micro with Ollama:
ollama run hf.co/VertexAGI/prism-caption-3-micro:Q4_K_M
- Unsloth Desktop
- MLX LM
How to use VertexAGI/prism-caption-3-micro with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "VertexAGI/prism-caption-3-micro"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "VertexAGI/prism-caption-3-micro" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "VertexAGI/prism-caption-3-micro", "messages": [ {"role": "user", "content": "Hello"} ] }' - Docker Model Runner
How to use VertexAGI/prism-caption-3-micro with Docker Model Runner:
docker model run hf.co/VertexAGI/prism-caption-3-micro:Q4_K_M
- Lemonade
How to use VertexAGI/prism-caption-3-micro with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull VertexAGI/prism-caption-3-micro:Q4_K_M
Run and chat with the model
lemonade run user.prism-caption-3-micro-Q4_K_M
List all available models
lemonade list
- Atomic Chat
Prism Caption 3 Micro
Prism Caption 3 Micro is a chat-titling model: given the first user message of a conversation, it writes a short, specific, correctly formatted title (4-6 words, title case, naming the actual subject). It is the same ~350M model as Prism Caption 2.5 Micro, retrained on a 30,000-example dataset, fine-tuned with LoRA on LFM2-350M. Part of the Prism family of small, single-purpose models.
Evaluation
Prism Caption 3 comes in two sizes, Micro (354M) and Pico (135M). Both are compared below with their untuned base models and the previous generation, on the same inputs: 275 held-out topics that never appear in any training bank, and 30 hand-written, realistic multi-sentence chat openers (the training prompts are short templated phrasings, so the second set checks that the model generalizes beyond the template).
275 held-out topics
| System | Format issues | Relevant | 3-6 words | Avg words | Speed (per title) |
|---|---|---|---|---|---|
| Base LFM2-350M | 231/275 | 247/275 | 90/275 | 10.2 | 69 ms |
| Prism Caption 2.5 Micro (published) | 27/275 | 274/275 | 222/275 | 5.2 | 52 ms |
| Prism Caption 3 Micro | 0/275 | 275/275 | 264/275 | 4.9 | 63 ms |
| Base SmolLM2-135M | 273/275 | 255/275 | 4/275 | 20.8 | 121 ms |
| Prism Caption 3 Pico | 2/275 | 275/275 | 219/275 | 5.3 | 48 ms |
30 realistic multi-sentence openers
| System | Format issues | Relevant | 3-6 words | Avg words | Speed (per title) |
|---|---|---|---|---|---|
| Base LFM2-350M | 25/30 | 27/30 | 9/30 | 14.2 | 85 ms |
| Prism Caption 2.5 Micro (published) | 9/30 | 29/30 | 19/30 | 6.8 | 60 ms |
| Prism Caption 3 Micro | 0/30 | 30/30 | 27/30 | 4.8 | 64 ms |
| Base SmolLM2-135M | 29/30 | 29/30 | 1/30 | 20.3 | 126 ms |
| Prism Caption 3 Pico | 2/30 | 29/30 | 25/30 | 4.7 | 47 ms |
"Format issues" = formatting problems (too long, too terse, leaked preamble, trailing punctuation, multiline). "Relevant" = the title shares a content word with the message. "3-6 words" = title length within the target range. "Speed" = average wall-clock time to generate one title (up to 28 new tokens, greedy decoding, single request) with mlx-lm on an Apple M4 (16 GB); speeds from separate runs differ by roughly +/-15 ms, so treat gaps smaller than that as noise. The Caption 3 rows above are the published default build (MLX 6-bit); every build is in the table further down.
The relevance check is a word-overlap rule and the format check is rule-based: this is a regression-style check, not a human or model judge, and both Caption 3 sizes are near its ceiling. The Caption 2.5 Micro row was measured on the published (fused, 4-bit) weights; the 0/275 issues figure on its own model card was measured earlier with the unfused adapter loaded.
Formats
This repo holds both an MLX and a GGUF build, plus extra MLX versions:
| Format | Location | Notes |
|---|---|---|
| MLX 6-bit (default) | repo root (model.safetensors + config) |
mlx-lm on Apple Silicon. Best balance: no measurable quality loss versus fp16 at lower latency |
| MLX fp16 | mlx-fp16/ |
Full-precision reference |
| MLX 4-bit, group size 32 | mlx-4bit-g32/ |
Fastest MLX option; small quality trade-off versus 6-bit (see the table below) |
| GGUF Q4_K_M | prism_caption_3_micro_Q4_K_M.gguf |
llama.cpp and compatible runtimes (LM Studio, Ollama, ...) |
| GGUF Q8_0 | prism_caption_3_micro_Q8_0.gguf |
Higher-fidelity GGUF |
Why not plain 4-bit
The model was fine-tuned on a 4-bit base and then fused. Fusing a LoRA into 4-bit weights and re-quantizing with the default settings loses part of the fine-tune: the plain 4-bit MLX build measured 16/275 formatting issues versus 0/275 for fp16 (same model, same prompts), so it is not published. The builds here fuse the LoRA at full precision first and quantize afterwards. The 6-bit build holds quality; the group-32 4-bit build is the faster option and trades a little quality for it (see the table: lengths loosen a bit, and Pico picks up a few format issues). GGUF K-quants held quality too.
All builds measured
275 held-out topics
| Build | Format issues | Relevant | 3-6 words | Avg words | Speed (per title) |
|---|---|---|---|---|---|
| MLX 6-bit (repo root, default) | 0/275 | 275/275 | 264/275 | 4.9 | 63 ms |
MLX fp16 (mlx-fp16/) |
0/275 | 275/275 | 267/275 | 4.9 | 105 ms |
MLX 4-bit group-32 (mlx-4bit-g32/) |
0/275 | 275/275 | 240/275 | 5.3 | 56 ms |
| GGUF Q8_0 | 0/275 | 275/275 | 267/275 | 4.9 | 53 ms |
| GGUF Q4_K_M | 0/275 | 275/275 | 236/275 | 5.3 | 77 ms |
| MLX 8-bit (measured, not published) | 0/275 | 275/275 | 267/275 | 4.8 | 69 ms |
| MLX plain 4-bit (measured, not published) | 16/275 | 275/275 | 205/275 | 5.6 | 54 ms |
30 realistic multi-sentence openers
| Build | Format issues | Relevant | 3-6 words | Avg words | Speed (per title) |
|---|---|---|---|---|---|
| MLX 6-bit (repo root, default) | 0/30 | 30/30 | 27/30 | 4.8 | 64 ms |
MLX fp16 (mlx-fp16/) |
1/30 | 30/30 | 27/30 | 5.0 | 106 ms |
MLX 4-bit group-32 (mlx-4bit-g32/) |
0/30 | 30/30 | 24/30 | 5.3 | 57 ms |
| GGUF Q8_0 | 1/30 | 30/30 | 27/30 | 5.0 | 77 ms |
| GGUF Q4_K_M | 0/30 | 30/30 | 27/30 | 5.0 | 59 ms |
| MLX 8-bit (measured, not published) | 1/30 | 30/30 | 27/30 | 4.9 | 74 ms |
| MLX plain 4-bit (measured, not published) | 8/30 | 30/30 | 20/30 | 6.9 | 60 ms |
GGUF rows were measured through llama-server (Metal), which adds a few ms of local HTTP overhead; MLX rows through mlx-lm.
Usage -- MLX
from mlx_lm import load, generate
model, tokenizer = load("VertexAGI/prism-caption-3-micro") # 6-bit default. For mlx-fp16/ or mlx-4bit-g32/, download the repo and pass that subfolder's local path
messages = [{"role": "system", "content": (
"You name chat conversations. Given the user's first message, reply with ONLY a short, specific chat title (4-6 words, title case, no quotes, no punctuation at the end, no preamble). The title MUST name the main subject of the message -- do not over-abbreviate into something vague. Nothing else -- just the title."
)}, {"role": "user", "content": "Any advice on how to fix a leaking kitchen faucet?"}]
text = tokenizer.apply_chat_template(messages, add_generation_prompt=True, tokenize=False)
print(generate(model, tokenizer, prompt=text, max_tokens=24))
Usage -- GGUF (llama.cpp)
hf download VertexAGI/prism-caption-3-micro prism_caption_3_micro_Q4_K_M.gguf --local-dir .
llama-cli -m prism_caption_3_micro_Q4_K_M.gguf -st -n 24 --temp 0 \
-sys "You name chat conversations. Given the user's first message, reply with ONLY a short, specific chat title (4-6 words, title case, no quotes, no punctuation at the end, no preamble). The title MUST name the main subject of the message -- do not over-abbreviate into something vague. Nothing else -- just the title." \
-p "Any advice on how to fix a leaking kitchen faucet?"
Model Details
| Base model | LiquidAI/LFM2-350M |
| Fine-tuning base checkpoint | the 4-bit MLX checkpoint mlx-community/LFM2-350M-4bit |
| Architecture | LFM2: hybrid short-convolution / attention |
| Fine-tuning method | LoRA (rank 8, scale 20.0, 16 (full depth) layers; 2.998M (0.846%) trainable parameters) |
| Framework | MLX / mlx-lm, on Apple Silicon |
| License | LFM Open License v1.0, inherited from the LFM2-350M base model. Free for research/non-commercial use and for commercial use under $10M annual revenue. |
Training Data
A 30,000-example chat-titling dataset (27,000 train / 3,000 validation), the 13,000-example set behind Prism Caption 2.5 Micro extended with 17,000 new examples over a much larger topic bank: 9,207 unique topics (up from 1,207), 28,367 unique prompts, 23,241 unique titles. Each example is a first user message paired with a teacher-written title. Teachers (fast models, cycled/switched adaptively by recent success rate):
| Teacher | Examples | Share |
|---|---|---|
openai/gpt-oss-20b (NIM) |
23,690 | 79.0% |
nvidia/nemotron-3.5-lightning-30b-a3b (NIM) |
4,250 | 14.2% |
poolside/laguna-s-2.1:free (OpenRouter) |
1,027 | 3.4% |
nvidia/nemotron-3-nano-omni-30b-a3b-reasoning (NIM) |
600 | 2.0% |
openai/gpt-oss-120b (NIM, retired 2026-09-03) |
433 | 1.4% |
New examples were filtered for format (3-8 words, no preamble) and relevance (the title must share a content word with the message). The 275 evaluation topics are excluded from every training bank.
Training Procedure
- Method: LoRA, rank 8, scale 20.0, dropout 0.0, 16 (full depth) layers; Adam, learning rate 1e-5, batch size 4, sequence length 256
- Steps: 6,750 iterations (exactly one epoch of the 27,000 training examples), validation every 250 steps
- Validation loss: 7.122 at initialization, best 0.333 at iteration 6,500 (final 0.345) -- the best checkpoint was used
- Throughput: ~2.35 it/s, ~1,050 tokens/s, peak memory ~1.2 GB (Apple M4)
- Hyperparameters are identical to Prism Caption 2.5 Micro, so the comparison with it isolates the data (13k to 30k examples, larger topic bank) rather than the recipe.
- Release builds: the LoRA adapter was fused into the de-quantized base at full precision, then quantized (MLX 6-bit / group-32 4-bit) or converted to GGUF (Q8_0 / Q4_K_M) from that fp16 model.
Eval scripts and raw results for every build are in eval/.
Limitations
Trained on synthetic titles distilled from a mix of teacher models, so some stylistic inconsistency between teachers may remain. Validated only on English, conversational, everyday-topic inputs; highly technical or non-English inputs are untested. Titles are short and extractive-leaning; a message with several unrelated requests will get a title for only one of them. Evaluation uses rule-based checks on 275 + 30 prompts, which are near ceiling for both sizes, so differences between Micro and Pico on this eval are small and should not be over-read.
License
LFM Open License v1.0, inherited from the LFM2-350M base model. Free for research/non-commercial use and for commercial use under $10M annual revenue.
- Downloads last month
- 202
6-bit
Model tree for VertexAGI/prism-caption-3-micro
Base model
LiquidAI/LFM2-350M