Configuration Parsing Warning:In adapter_config.json: "peft.task_type" must be a string

SO-101 VLM agent (LoRA v4, "V4f"), trained only in simulation

A LoRA adapter for Qwen3.8-27B that turns the model into a skill-level agent for tabletop pick-and-place with the SO-101 arm. At each step the model reads three camera images (top, side, gripper) plus the command and chooses one of nine skills: move to an object, grasp, lift, move to a place, rotate, release, open, done, give up. The choice is a single-token readout. The base model points at objects in the top image; code executes bounded motion primitives.

Trained only in simulation. No teleoperation and no real-world training data were used.

Adapter

Base model Qwen/Qwen3.8-27B, revision 1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0
Type LoRA, rank 16, alpha 32, dropout 0.05, on the language model's attention (q/k/v/o) and MLP projections
Trainable parameters 79.7M
File adapter_model.safetensors, 318.8 MB, SHA-256 ab6678ec7cb849a3fd8ce4a63ea4d2c9e270c6d49a4ec8c7f9790c0d098436fe

Training data

Simulated SO-101 episodes in MuJoCo. The final round continued from an earlier adapter for one epoch on 5,347 examples (4,085 s on one H100). Its states came from 200 of the agent's own simulated rollouts (DAgger-style); each state was labelled from simulated outcomes: restore the state, try every candidate step, finish with a scripted policy, and score success minus 0.02 per step. Trained task families: put in a container, stack, sort, bar in a tray, take out. Held out: gather, next to, in then out, and all test seeds.

Results

Simulation (600 photorealistic held-out test episodes, 544 solvable): 440 / 544 successes (81%, 95% interval 77–84%); 199 / 240 clean episodes and 220 / 269 episodes with an injected fault; correct give-up on 54 / 56 unreachable ones.

Real SO-101 (field report, 28 runs over two days, the adapter changing between runs): 6 of 9 single-ball "put the ball in the container" runs, one two-ball command in 104 s, and 0 of 10 "take the ball out" runs. Most failures were in calibration, motion primitives and layout rather than in the model's decisions. See the paper and the code repository for the full ledger and videos.

Use

Serve the base model with the adapter, for example with vLLM (--enable-lora --lora-modules run7-v4=<path>), and run the agent and real-arm adapter from the code repository. The adapter expects the agent's prompts and single-token readouts; it is not a general chat model.

License

Apache-2.0, the same as the base model.

Limitations

Evaluated on one simulated arm family and one real desk. Real-world success depends heavily on camera calibration, the motion primitives and the scene layout. Take-out tasks, balls against container walls and poses at joint limits failed on the real arm. Not for use where a failure could hurt people or property.

Links

Downloads last month
10
Video Preview
loading

Model tree for squiredaniiar/so101-vlm-agent

Base model

Qwen/Qwen3.8-27B
Adapter
(136)
this model