Visual grounding is the task of localising the region of an image or scene that corresponds to a natural-language expression, linking words to specific visual entities. It connects language understanding to perception, enabling models to point at, select or act on the object a user refers to. Visual grounding is foundational for vision-language models and for agents that operate graphical interfaces.
Content
- It requires joint reasoning over language and visual features to resolve referring expressions, often outputting bounding boxes, masks or coordinates. Grounding lets agents identify clickable UI elements or scene objects, bridging instruction following and concrete action.