Screen capture is the acquisition of the visual contents of a display, window, or application as still images or a stream of frames. It underpins screenshots, screen recording, and remote presentation, and is increasingly used to provide perceptual input to AI agents that operate graphical interfaces. Capture is mediated by operating-system or browser APIs subject to user permission.

Content

  • Capture pipelines read framebuffers or composited surfaces through OS or browser APIs, optionally constrained to a single window or region. For AI agents, captured frames are paired with accessibility metadata to ground grounding and action selection, while permission prompts and privacy indicators mitigate misuse.