Gaze control regulates robot eye and head movement to establish, maintain, and redirect visual attention toward objects and people, conveying robot intent and facilitating natural Human-Robot Interaction.

Semantic Classification

Content

Gaze control in robots encompasses two distinct challenges: technical control of eye and head actuators to point visual sensors toward desired targets, and cognitive modelling of where gaze should be directed based on task context and social interaction norms. Technical implementation typically employs Pan-Tilt Unit mechanisms for head orientation and motorised eye rotation if articulated eyes are present, controlled through Servo Motors and coordinated motion planning. Visual servoing approaches use image feedback to maintain gaze on tracked features despite disturbances.

The social dimension emerges from psychological evidence showing humans interpret robot gaze direction as indicating attention and intention; robots that gaze toward task objects enhance human understanding of robot plans. During collaboration, robots that gaze at human interaction partners strengthen social presence and reduce task execution time. Implementing effective social gaze requires predicting human gaze direction through Computer Vision, estimating human interest from body language, and coordinating robot gaze to maintain Joint Attention—the shared focus of two agents on a common object or location.

Contemporary gaze control systems integrate perception, planning, and social reasoning: Convolutional Neural Network models predict salient regions humans find interesting, Reinforcement Learning agents learn gaze policies that maximise human satisfaction, and hierarchical controllers balance smooth continuous gaze tracking with attention-shifting to newly detected task-relevant objects. Research explores gaze aversion mechanisms enabling robots to look away at socially appropriate moments (avoiding staring), gaze-based communication enabling robots to express uncertainty or confusion through eye movement patterns, and group interaction scenarios where multiple robots coordinate gaze to enhance collective understanding.

Provenance