Grounded language understanding is the capacity of an agent to connect linguistic meaning to perception and action in a physical or simulated environment, rather than treating language as a purely symbolic, text-only system. It requires linking words and phrases to sensory referents, spatial relations and executable actions so that an instruction such as ‘pick up the red block’ resolves to concrete perception and motor commands. It is a prerequisite for embodied AI and embodied-minds research, where an agent must interpret language in the context of its own body and surroundings.