Papers
arxiv:2605.19410

Vision Harnessing Agent for Open Ad-hoc Segmentation

Published on Sep 29
Authors:
,

Abstract

Segmentation has become easy when the concept is known, requiring retrieval of a learned visual grounding from text. It remains hard for open ad-hoc concepts, where the grounding may not exist as one learned mask and must often be constructed from image evidence through parts, relations, exclusions, and collections. We propose a Vision-guided Ad-hoc Segmentation Agent (VASA), the first vision harnessing agent for open ad-hoc segmentation. VASA is training-free and couples a VLM agent, a segmentation foundation model, and a visual harness that maintains a working mask to make visual progress persistent, inspectable, and editable. Rather than revising text prompts alone, it plans visual operations, invokes segmentation tools, inspects results, edits the mask, and recovers from errors. We construct PARS, a new benchmark that turns part-level labels into open ad-hoc concepts through long-form definition queries. We show that VASA is consistently effective across six VLMs with varying capabilities. Using Qwen3-VL 32B Thinking as the VLM, VASA outperforms various baselines on PARS, surpassing SAM3 Agent by 13.5%-25.3%. On RefCOCOm, VASA improves over SAM3 Agent by 4.8%-8.8% and over other agentic baselines by more. VASA also remains competitive with SAM3 Agent on ReasonSeg for common, named concepts. These results validate VASA's agentic visual construction for open ad-hoc segmentation.

Community

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2605.19410
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2605.19410 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2605.19410 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2605.19410 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.