Instructions to use objectai/obj_v1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use objectai/obj_v1 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="objectai/obj_v1") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("objectai/obj_v1") model = AutoModelForMultimodalLM.from_pretrained("objectai/obj_v1", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use objectai/obj_v1 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "objectai/obj_v1" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "objectai/obj_v1", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/objectai/obj_v1
- SGLang
How to use objectai/obj_v1 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "objectai/obj_v1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "objectai/obj_v1", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "objectai/obj_v1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "objectai/obj_v1", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use objectai/obj_v1 with Docker Model Runner:
docker model run hf.co/objectai/obj_v1
obj_v1
Vision-language model fine-tuned for structured data extraction from Indian financial documents. Give it a page image and a JSON schema; it returns the schema filled in from what is on the page.
A 4B vision-language model, LoRA fine-tuned and merged. Nothing extra is needed at load time -- it is a plain bf16 checkpoint.
Serving with vLLM
vllm serve objectai/obj_v1 \
--served-model-name obj_v1 \
--max-model-len 16384 \
--limit-mm-per-prompt '{"image":1}' \
--mm-processor-kwargs '{"max_pixels":1003520}' \
--trust-remote-code
max_pixels is 1280x28x28, the resolution the model was trained at. Raising it
wastes KV cache; lowering it makes small print unreadable.
Calling it
The server is OpenAI-compatible, so an ordinary chat completion works:
import base64, json, openai
client = openai.OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
image = base64.b64encode(open("cheque.jpg", "rb").read()).decode()
schema = {"cheque_details": {"amount": "number", "payee": "string",
"date": "string", "cheque_number": "string"}}
response = client.chat.completions.create(
model="obj_v1",
temperature=0.0,
max_tokens=8192,
messages=[
{"role": "system", "content":
"You are a document data extraction model. "
"Extract only values present in the document. "
"Use null for fields that are absent or illegible. "
"Output a single compact JSON object matching the requested schema. "
"No prose, no markdown, no explanation."},
{"role": "user", "content": [
{"type": "image_url",
"image_url": {"url": f"data:image/jpeg;base64,{image}"}},
{"type": "text",
"text": f"document_type: cheque\nschema: {json.dumps(schema)}"},
]},
],
)
print(response.choices[0].message.content)
Prompt format
Match training or accuracy drops. The system prompt above is verbatim, and the user turn is the image followed by exactly two lines:
document_type: <type>
schema: <compact json>
Set temperature=0.0 so the same page yields the same answer.
Requirements
| Weights | 8.9 GB (bf16) |
| VRAM | 16 GB minimum, 24 GB comfortable |
| Precision | bf16 (Ampere or newer; use fp16 below that) |
| Context | 16384 covers the longest documents |
Runs on an L4, A10G, L40S, A100 or RTX 4090. On a T4 add --dtype float16.
Long documents matter: bank_statement and form16 answers run to ~2500
tokens, so max_tokens below 4096 truncates them mid-JSON.
Output
Compact JSON matching the requested schema. Fields absent from the page come
back null rather than guessed. Values found on the page that the schema did
not ask for are placed under extras when that key is included in the schema.
Limitations
- Trained on Indian financial documents; other domains and layouts are untested.
- Handwriting is the weakest case, particularly digits at low resolution.
- The model does not verify its own arithmetic. Totals that must reconcile should be checked by the caller.
License
Apache 2.0. Fine-tuned from Qwen3-VL-4B-Instruct, which is Apache 2.0.
- Downloads last month
- 2
Model tree for objectai/obj_v1
Base model
Qwen/Qwen3-VL-4B-Instruct