bartender-2

Qwen3-VL-4B-Instruct, LoRA-merged, fine-tuned on simulated bar-scene frames (Gazebo, bartender_robot_sim) to describe a bartending robot's camera view as JSON: bottles, glasses, gripper contents, obstruction.

Known limitation: trained only on the old 3-class sim vocabulary (whiskey / cola / beer). It has not seen the real bar's other bottles (vodka, liqueur, gin, wine) and will misname them. A 6-class retrain is in progress.

Prompt used at training time is in prompt.txt in this repo.

Downloads last month
14
Safetensors
Model size
4B params
Tensor type
BF16
·
Video Preview
loading

Model tree for cheeselover69/bartender-2

Adapter
(214)
this model