Instructions to use SamuelBang/LeRF-9B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use SamuelBang/LeRF-9B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="SamuelBang/LeRF-9B") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("SamuelBang/LeRF-9B") model = AutoModelForMultimodalLM.from_pretrained("SamuelBang/LeRF-9B", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use SamuelBang/LeRF-9B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "SamuelBang/LeRF-9B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "SamuelBang/LeRF-9B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/SamuelBang/LeRF-9B
- SGLang
How to use SamuelBang/LeRF-9B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "SamuelBang/LeRF-9B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "SamuelBang/LeRF-9B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "SamuelBang/LeRF-9B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "SamuelBang/LeRF-9B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use SamuelBang/LeRF-9B with Docker Model Runner:
docker model run hf.co/SamuelBang/LeRF-9B
LeRF-9B
LeRF: Learning Reference Coordinate Frames for Perspective Taking Reasoning
Bang Xiao1,2 · Wenqi Jia1 · Ozgur Kara1 · Tiancheng Shen3 · Yibo Yang4
Bolin Lai1,5,^ · Junho Kim1,^ · James Matthew Rehg1,^
1University of Illinois Urbana-Champaign · 2Zhiyuan College, Shanghai Jiao Tong University
3University of California, Merced · 4Shanghai Jiao Tong University · 5Amazon AGI
^Corresponding authors
Model Summary
Vision-language models (VLMs) struggle with perspective taking: when a question asks about spatial relations from another entity's or an imagined observer's viewpoint, they often fall back to the camera view. LeRF trains a VLM to build and use an explicit entity-centered reference frame. Given an image and a question, the model decides whether a frame is needed. If it is, it grounds the reference entity and predicts the frame's projected origin and its front / left / up axes. A lightweight renderer draws the frame on the image, and the model reasons over this visual cue. No external perception models or 3D reconstruction are used.
How It Works
Two-turn inference.
- Turn 1 (thinking off). The model outputs either a single
draw_reference_framecall —reference_object(text), plusorigin,axis_front,axis_left,axis_upas[x, y]points in[0, 1000]normalized image coordinates — or the literal stringNO_TOOL_CALLwhen the question can be answered from the camera view. - Tool. The frame is rendered on the image (red = front, green = left, blue = up; a
⊙/⊗depth marker replaces a strongly foreshortened axis). - Turn 2 (thinking on). The model sees the rendered image (not the numeric coordinates),
reasons over it, and answers in
\boxed{}.
Training.
- Stage I: SFT. Learns reference-entity grounding and projected frame prediction from pose datasets (ImageNet3D, Omni6DPose-SOPE, BEDLAM). About 20% of the data are no-tool-call examples that teach selective tool use.
- Stage II: RL (GRPO). Trained on MultihopSpatial and SpatialReasoner-RL with a binary final-answer reward. The frame-prediction turn is excluded from the policy gradient, so only the reasoning/answer turn receives the advantage.
Usage
The model is meant to be used together with the draw_reference_frame tool and the two-turn
client from the GitHub repository. The system prompt and
tool schema live in tool/system_prompt.txt and tool/prompts.py.
1. Serve with vLLM (OpenAI-compatible API):
vllm serve SamuelBang/LeRF-9B --served-model-name lerf --max-model-len 32768 \
--limit-mm-per-prompt '{"image": 2}' --trust-remote-code
2. Run the client:
git clone https://github.com/bangx7/LeRF.git && cd LeRF/tool
pip install pillow requests
# single question
python inference.py --model lerf --image example.jpg \
--question "From the man's perspective, which object is on his left?" \
--options cup bed laptop phone --save_renders renders/
# batch: one JSON per line with image / question / options (and optionally answer, id)
python inference.py --model lerf --input questions.jsonl --output results.jsonl
Recommended settings (matching training): temperature 0.6, images downscaled to at most 1M pixels, a 10,240-token thinking budget within a 12,288-token answer turn.
Tested with Python 3.12, CUDA 13.0, PyTorch 2.11, vLLM 0.24.0 and transformers 5.10.4.
Evaluation
Accuracy (%) on OmniSpatial perspective taking (OmniSpatial-PT: Ego / Allo / Hypo), 3DSRBench (Orientation / Multi-Object) and ViewSpatial-Bench (person-perspective Object View Orientation / Relative Direction). Bold marks the best open-source result.
| Method | Ego | Allo | Hypo | 3DSR Ori | 3DSR M-Obj | VS P-Obj | VS P-Rel |
|---|---|---|---|---|---|---|---|
| Proprietary | |||||||
| GPT-5.6-Luna (medium) | 83.33 | 49.73 | 45.78 | 60.04 | 55.34 | 46.99 | 70.07 |
| GPT-5.6-Terra (medium) | 81.37 | 55.85 | 53.01 | 63.32 | 56.39 | 45.08 | 77.20 |
| Claude Sonnet 5 (medium) | 80.39 | 42.55 | 49.40 | 34.94 | 44.77 | 51.31 | 51.43 |
| Claude Sonnet 5 (high) | 84.31 | 48.14 | 45.78 | 43.15 | 46.81 | 51.51 | 60.10 |
| Open-source | |||||||
| Qwen3.5-4B | 74.71 | 42.55 | 44.34 | 42.28 | 43.26 | 51.01 | 57.43 |
| Qwen3.5-9B | 80.20 | 47.13 | 44.58 | 48.17 | 48.66 | 56.23 | 65.51 |
| Qwen3.5-9B + APC | 42.16 | 27.66 | 30.12 | 44.98 | 33.22 | 59.34 | 37.53 |
| SpatialReasoner | 40.39 | 35.11 | 35.66 | 52.05 | 50.64 | 42.37 | 45.61 |
| Ours | |||||||
| LeRF-4B | 72.35 | 49.36 | 46.75 | 45.88 | 44.55 | 56.26 | 67.85 |
| LeRF-9B (this model) | 74.31 | 54.04 | 55.66 | 53.76 | 50.29 | 61.91 | 74.23 |
See the paper for more baselines, reference-frame estimation accuracy and ablations.
Acknowledgements
Built on Qwen3.5, LLaMA-Factory and verl.
Citation
Coming soon.
- Downloads last month
- 13