Cross-domain marker for metaverse components combining artificial intelligence with human interaction systems including conversational AI, gesture recognition, emotion detection, and intelligent user experience adaptation.
Bridge-To
Semantic Classification
Content
Compositional Relationships (Components)
SubClassOf(ai:ETSIDomainAIHumanInterface
ObjectSomeValuesFrom(ai:hasPart ai:ConversationalAIClassification))
SubClassOf(ai:ETSIDomainAIHumanInterface
ObjectSomeValuesFrom(ai:hasPart ai:IntelligentUXCategorization))
SubClassOf(ai:ETSIDomainAIHumanInterface
ObjectSomeValuesFrom(ai:hasPart ai:SentimentAnalysis))
SubClassOf(ai:ETSIDomainAIHumanInterface
ObjectSomeValuesFrom(ai:hasPart ai:EmotionRecognition))
SubClassOf(ai:ETSIDomainAIHumanInterface
ObjectSomeValuesFrom(ai:hasPart ai:GestureRecognition))
SubClassOf(ai:ETSIDomainAIHumanInterface
ObjectSomeValuesFrom(ai:hasPart ai:AffectiveComputing))
SubClassOf(ai:ETSIDomainAIHumanInterface
ObjectSomeValuesFrom(ai:hasPart ai:SpeechRecognition))
SubClassOf(ai:ETSIDomainAIHumanInterface
ObjectSomeValuesFrom(ai:hasPart ai:GazeTracking))
Dependency Relationships
SubClassOf(ai:ETSIDomainAIHumanInterface
ObjectSomeValuesFrom(ai:requires ai:NaturalLanguageProcessing))
SubClassOf(ai:ETSIDomainAIHumanInterface
ObjectSomeValuesFrom(ai:requires ai:MachineLearning))
SubClassOf(ai:ETSIDomainAIHumanInterface
ObjectSomeValuesFrom(ai:requires ai:ComputerVision))
SubClassOf(ai:ETSIDomainAIHumanInterface
ObjectSomeValuesFrom(ai:requires ai:SpeechRecognition))
SubClassOf(ai:ETSIDomainAIHumanInterface
ObjectSomeValuesFrom(ai:requires ai:MultimodalAI))
SubClassOf(ai:ETSIDomainAIHumanInterface
ObjectSomeValuesFrom(ai:requires ai:AffectiveComputing))
SubClassOf(ai:ETSIDomainAIHumanInterface
ObjectSomeValuesFrom(ai:requires ai:UserModelling))
SubClassOf(ai:ETSIDomainAIHumanInterface
ObjectSomeValuesFrom(ai:requires ai:ContextAwareness))
Capability Relationships
SubClassOf(ai:ETSIDomainAIHumanInterface
ObjectSomeValuesFrom(ai:enables ai:ConversationalAI))
SubClassOf(ai:ETSIDomainAIHumanInterface
ObjectSomeValuesFrom(ai:enables ai:VirtualAssistant))
SubClassOf(ai:ETSIDomainAIHumanInterface
ObjectSomeValuesFrom(ai:enables ai:HumanRobotInteraction))
SubClassOf(ai:ETSIDomainAIHumanInterface
ObjectSomeValuesFrom(ai:enables ai:AccessibleXR))
SubClassOf(ai:ETSIDomainAIHumanInterface
ObjectSomeValuesFrom(ai:enables ai:IntelligentTutoringSystem))
SubClassOf(ai:ETSIDomainAIHumanInterface
ObjectSomeValuesFrom(ai:enables ai:Telecollaboration))
Implementation Relationships
SubClassOf(ai:ETSIDomainAIHumanInterface
ObjectSomeValuesFrom(ai:uses ai:LargeLanguageModels))
SubClassOf(ai:ETSIDomainAIHumanInterface
ObjectSomeValuesFrom(ai:uses ai:TransformerArchitecture))
SubClassOf(ai:ETSIDomainAIHumanInterface
ObjectSomeValuesFrom(ai:uses ai:ConvolutionalNeuralNetwork))
SubClassOf(ai:ETSIDomainAIHumanInterface
ObjectSomeValuesFrom(ai:uses ai:ReinforcementLearningFromHumanFeedback))
SubClassOf(ai:ETSIDomainAIHumanInterface
ObjectSomeValuesFrom(ai:uses ai:AttentionMechanism))
SubClassOf(ai:ETSIDomainAIHumanInterface
ObjectSomeValuesFrom(ai:uses ai:TextToSpeech))
SubClassOf(ai:ETSIDomainAIHumanInterface
ObjectSomeValuesFrom(ai:implements ai:IntentRecognition))
SubClassOf(ai:ETSIDomainAIHumanInterface
ObjectSomeValuesFrom(ai:implements ai:DialogueStateTracking))
Reduction Relationships
SubClassOf(ai:ETSIDomainAIHumanInterface
ObjectSomeValuesFrom(ai:reducesTo ai:ETSIDomainAI))
SubClassOf(ai:ETSIDomainAIHumanInterface
ObjectSomeValuesFrom(ai:reducesTo ai:InteractionDomain))
SubClassOf(ai:ETSIDomainAIHumanInterface
ObjectSomeValuesFrom(ai:reducesTo ai:NaturalLanguageProcessing))
SubClassOf(ai:ETSIDomainAIHumanInterface
ObjectSomeValuesFrom(ai:reducesTo ai:HumanComputerInteraction))
About
ETSI Domain AI + Human Interface occupies a strategically critical position within the ETSI metaverse taxonomy. While the other AI sub-domains — ETSI Domain AI + Creative Media, ETSI Domain AI + Data Mgmt, and ETSI Domain AI + Governance — address specific capability categories or compliance obligations, the Human Interface domain addresses the most fundamental challenge of the metaverse paradigm: how do humans and intelligent systems come to understand each other across the gap of embodiment, intent, and context? The domain’s scope spans from the millisecond timescale of gesture detection and speech recognition through to the long-arc timescale of longitudinal user modelling that personalises metaverse environments over months of use. The taxonomic role of this cross-domain node is to flag that a given ETSI MEC-deployed component makes AI-mediated claims about, or exerts AI-mediated influence upon, the human user’s experience of the metaverse environment — triggering downstream design obligations for naturalness, accessibility, latency tolerance, cultural sensitivity, and user-centred evaluation that do not apply to purely back-end AI data processing or content generation components.
The intellectual genealogy of this domain traces to the Human-Computer Interaction (HCI) research tradition established by pioneers including Douglas Engelbart (the Augmentation Research Center, 1960s), Alan Kay (Dynabook concept, Xerox PARC, 1970s), and Ben Shneiderman (direct manipulation interfaces, 1983). The application of machine learning to HCI began in earnest with statistical gesture recognition (CMU LISTEN project, 1990s), early affective computing work (Rosalind Picard, MIT Media Lab, 1997), and multimodal interaction research through the W3C Multimodal Interaction Working Group (established 2002). The contemporary ETSI domain represents the convergence of these traditions with the representational power of deep learning and the deployment vehicle of multi-access edge computing. Where earlier HCI research focused on the ergonomics of fixed interfaces — keyboard layouts, menu structures, icon design — ETSI Domain AI + Human Interface is fundamentally concerned with the design of interfaces that do not have a fixed form but instead emerge from continuous AI inference about what the user needs, knows, and can do at any given moment.
What distinguishes ETSI Domain AI + Human Interface from generic HCI is the explicit requirement for AI-mediated adaptivity — the capacity of the system to continuously update its model of the user, the context, and the interaction history, and to use that model to reshape the interface in real time. This distinguishes static GUI design from intelligent adaptive interfaces, rule-based dialogue trees from Conversational AI, and scripted avatar behaviour from AI-driven Embodied Conversational Agent design. The ETSI GS MEC infrastructure provides the latency-sensitive edge computing substrate that makes real-time human-AI interface adaptation feasible at scale: by bringing AI inference within 10-20ms of the user (at the network edge rather than a distant cloud data centre), MEC enables the sub-perceptual-latency loop closures that naturalistic multimodal interaction requires. The bridge to Telecollaboration is architecturally significant — in shared XR spaces, the AI human interface layer must coordinate across multiple users simultaneously, managing attention, presence, and communication bandwidth in multi-party AI-mediated collaboration. Multi-access edge computing is the enabling infrastructure that makes this possible: cloud-hosted AI inference for gesture recognition would add 80-200ms round-trip latency, producing visible lag between hand movement and virtual response that breaks the perceptual illusion of natural interaction. MEC brings inference to within 5-20ms, enabling the sub-50ms total system latency required for natural-feeling gesture interaction in XR.
The domain is experiencing rapid capability expansion driven by three concurrent developments as of 2026: (1) the emergence of natively multimodal large language models (GPT-4o, Gemini 2.0 Flash, Claude 4) that process and generate speech, image, and text within a unified architecture, eliminating the pipeline fragmentation that constrained prior-generation voice and multimodal interfaces; (2) advances in real-time gesture and body pose estimation enabled by lightweight vision transformer models deployable on edge hardware; (3) growing sophistication of emotion and affect recognition systems, including a UK-based project using supercomputer-generated synthetic facial expression data to train more generalisable emotion recognition models (BiometricUpdate.com, 2025). A fourth dimension — AI-mediated accessibility — is gaining particular prominence as governments mandate inclusive design: EU Web Accessibility Directive extensions to XR environments, UK Equality Act digital accessibility interpretations, and WCAG 3.0’s emerging XR guidelines all create positive obligations for AI human interface components to serve users with diverse ability profiles without requiring separate “accessibility modes” that create second-class user experiences.
The relationship between ETSI Domain AI + Human Interface and ETSI Domain AI + Governance is particularly intricate and deserves explicit treatment. Many of the AI capabilities in the Human Interface domain — emotion recognition, sentiment analysis, personalisation based on inferred psychological state, biometric identification from voice or gaze — are precisely the capabilities that the EU AI Act subjects to the most stringent governance requirements. Emotion recognition used in workplace or educational settings is in the EU AI Act Annex III high-risk category; biometric categorisation from facial images is similarly high-risk. This means that metaverse components bearing the ETSI Domain AI + Human Interface classification will frequently also require ETSI Domain AI + Governance classification, and system architects must reason about the joint obligations that arise. The ETSI taxonomy’s bridging architecture, where domain nodes can be combined, is precisely designed to capture this co-occurrence pattern.
The history of multimodal AI interaction in the human-computer interface domain also illustrates how standardisation efforts like ETSI Domain AI + Human Interface can pre-empt proprietary fragmentation. In the early 2010s, before the deep learning era, multiple competing proprietary gesture vocabularies, voice command protocols, and affective computing APIs existed without common interfaces, making cross-platform XR human interface development extremely difficult. ETSI’s domain classification approach creates the shared vocabulary that allows interface capability descriptions to be exchanged between systems, enabling MEC platforms from different vendors to negotiate about what human interface AI capabilities are available and required, and service orchestrators to dynamically allocate appropriate inference resources when a user session’s interaction modality requires gesture recognition, emotion sensing, or conversational AI.
Components and Architecture
-
Conversational AI Classification Layer: Natural-language dialogue management for metaverse agents and avatars. Implements Intent Recognition, Natural Language Understanding, Dialogue State Tracking, and response generation using Large Language Models backed by Retrieval-Augmented Generation. Connects users to Digital Twin environments, navigation systems, and service interfaces through natural speech and text. The ETSI MEC Phase 4 architecture (GS MEC 003 V4.1.1, 2025) specifies edge-hosted inference services for low-latency conversational AI in multi-access XR environments.
-
Gesture Recognition Pipeline: Computer-vision-based detection and classification of hand gestures, body poses, and deictic pointing acts from RGB or depth camera streams. Uses Convolutional Neural Network and vision transformer architectures for real-time landmark detection (MediaPipe Holistic, HRNet-class models). Translates gesture vocabularies into metaverse control signals — navigation, object manipulation, avatar expression, and interface interaction. Critical for Accessible XR use cases where traditional controllers are unavailable or inappropriate.
-
Emotion Recognition and Affective Computing Module: Multimodal affect sensing from facial expression analysis (action unit detection, valence-arousal mapping), speech prosody (pitch, energy, speech rate), physiological signals (heart rate variability, electrodermal activity via wearables), and behavioural cues. Outputs are used to adapt interface pacing, content difficulty, avatar emotionality, and support escalation (e.g., detecting user frustration to trigger intervention). UK supercomputing resources (ARCHER2) are being applied to train emotion recognition systems on large synthetic datasets to improve generalisation beyond biased training corpora (2025).
-
Intelligent UX Categorization and Adaptive Personalisation: Long-term user modelling that maintains representations of user preferences, expertise level, interaction history, accessibility needs, and contextual factors. Used to dynamically adapt metaverse environment layout, information density, interaction modality, content recommendations, and agent behaviour. Integrates with ETSI Domain AI + Governance obligations — personalisation engines using inferred sensitive characteristics (health status, emotional state) must satisfy Regulatory Compliance requirements under EU AI Act Article 50 transparency obligations.
-
Speech Recognition and Voice Interface Stack: Automatic Speech Recognition (ASR) for real-time voice input, speaker identification, voice activity detection, and noise robustness in XR acoustic environments. Modern deployments use Whisper-class transformer ASR models achieving under 5% WER on clean speech; edge-optimised variants (Whisper Tiny, Moonshine) enable on-device inference for privacy-sensitive deployments. End-to-end speech-to-speech models (GPT-4o voice mode, Gemini 2.0 Flash Live) eliminate the ASR/NLU boundary for ultra-low-latency voice interaction.
-
Gaze Tracking and Attention Modelling: Eye-tracking hardware (infrared corneal reflection, video oculography) integrated with AI-based gaze estimation models provides foveal attention signals used for: foveated rendering (rendering at high resolution only where the user is looking, reducing GPU load by 5-10x); attentional state monitoring (detecting distraction, cognitive load overload); social gaze in multi-user telecollaboration; and gaze-directed interface selection as an input modality for users with motor impairments.
-
Accessibility Technology Framework: AI-mediated adaptations ensuring XR environments are navigable and usable by users with visual, auditory, motor, and cognitive impairments. Includes real-time captioning and sign-language interpretation, gesture-based navigation alternatives, cognitive load adaptation (simplifying complex interfaces when user difficulty signals are detected), and voice-only navigation modes. Aligned with W3C Multimodal Interaction Working Group specifications and UK Equality Act 2010 digital accessibility obligations.
-
Telecollaboration AI Mediation Layer: Multi-user shared XR space management using AI to coordinate presence signalling, turn-taking in multi-party conversation, attention direction (highlighting speakers, flagging relevant events), avatar proxemics, and collaborative object manipulation. The AI layer abstracts over heterogeneous user hardware (VR headsets, AR glasses, mobile devices, flat-screen telepresence) to create coherent shared interaction spaces.
Use Cases and Major Families
-
AI-Driven Metaverse Avatars and Agents: Intelligent non-player characters (NPCs) and AI avatars that engage in natural-language conversation, respond to gesture and gaze, express procedurally-generated emotional affect, and adapt their behaviour to user context. Deployed in virtual commerce, entertainment, training simulation, and social XR. ETSI Domain AI + Human Interface classifies the AI components enabling this behaviour, triggering downstream ETSI Domain AI + Governance obligations where the AI agent acts on behalf of or influences vulnerable users.
-
Multi-modal XR Navigation: Natural-language, gesture, and gaze-directed navigation in complex 3D environments — replacing traditional game-controller or keyboard-based navigation with AI-interpreted natural interaction. Particularly transformative for professional metaverse applications (architectural walkthroughs, medical simulation, industrial training) where domain experts need to interact naturally without mastering VR controllers.
-
Intelligent Tutoring and Training Simulation: AI human interface components enabling adaptive educational experiences in XR — dialogue-based instruction by AI tutors, gesture-assessed practical skill training, emotion-aware pacing that responds to frustration or disengagement, and personalised difficulty adaptation. UK universities and the NHS are evaluating XR training platforms with AI human interface layers for surgical training and clinical skills assessment.
-
Healthcare and Assistive Applications: AI-mediated interfaces for users with communication difficulties — augmentative and alternative communication (AAC) systems using AI prediction, brain-computer interface (BCI) integration for users with motor impairments, emotion recognition for non-verbal communication support, and gaze-directed smart home control. The domain classification signals that these deployments require medical-grade reliability and governance compliance.
-
Enterprise Collaboration and Remote Work: AI-enhanced Telecollaboration platforms where intelligent systems manage meeting facilitation, real-time translation, action item extraction, attention monitoring, and collaborative document editing through natural multimodal interaction. The AI human interface layer sits between human participants and the shared digital workspace, providing contextual assistance without disrupting conversational flow.
-
Customer Service and Retail in XR: AI virtual assistants embodied as avatars in virtual retail environments — providing product information, navigating virtual showrooms, assisting with configuration decisions, and escalating to human agents when AI confidence is insufficient. Conversational AI Classification components handle dialogue; Sentiment Analysis detects customer dissatisfaction; adaptive personalisation tailors the retail experience to inferred preferences.
-
AI Companion Systems: Long-term AI companions in metaverse environments that build and maintain persistent models of individual users, adapting personality, communication style, and proactive assistance patterns over extended interaction histories. Intersection with ETSI Domain AI + Governance is significant — companion AI that infers emotional vulnerability requires governance controls preventing exploitation.
Academic Context
The academic foundations of ETSI Domain AI + Human Interface span three research traditions that have increasingly converged. Human-Computer Interaction (HCI) provided foundational theory: Shneiderman’s (1983) direct manipulation principles, Weiser’s (1991) ubiquitous computing vision, Norman’s (1988) affordance theory, and Grudin’s (1994) CSCW framework for computer-supported cooperative work that directly underpins the Telecollaboration bridge. Affective Computing emerged from Picard’s (1997) foundational monograph at MIT Media Lab, establishing emotion recognition from physiological signals as a computer science research area. Multimodal HCI research formalised through W3C and research groups studying combined speech-gesture interaction (Bolt, 1980; Oviatt, 2003).
The deep learning revolution transformed the field from 2012. Krizhevsky et al.’s ImageNet breakthrough (2012) enabled practical facial expression recognition from monocular video. Vaswani et al.’s transformer architecture (2017) provided the unified sequence processing capability that powers modern conversational AI and multimodal understanding. The multimodal large language model era (GPT-4V, 2023; Gemini, 2023; Claude 3, 2024) enabled genuine cross-modal understanding — describing images in natural language, reasoning about video content, and processing speech, text, and vision within unified architectures.
Human-AI Interaction (HAII) has emerged as a distinct research sub-field addressing the specific challenges of AI-driven interfaces: explainability in conversational systems (Gunning et al., 2019), calibrated trust in AI recommendations, mixed-initiative interaction where AI and human share control, and the ethics of affective AI. The ACM CHI and UIST conference communities are the primary venues, with CSCW covering the telecollaboration dimension. A 2025 survey (arXiv:2503.16472) provides a comprehensive mapping of human-AI interaction design standards, identifying ETSI MEC, ISO/IEC SC 42, and W3C multimodal specifications as the primary standardisation anchors for the domain.
Current Landscape (2026)
The ETSI MEC Phase 4 work programme (GS MEC 003 V4.1.1, May 2025) has expanded the MEC framework’s scope to explicitly encompass AI inference services at the network edge, with AI human interface workloads — conversational AI inference, gesture recognition, emotion detection — identified as primary edge-latency-sensitive use cases. Phase 4 specifications developed in collaboration with the Linux Foundation CAMARA project and TM Forum aim to harmonise MEC with open-source edge-cloud implementations, broadening the ecosystem for deploying AI human interface components at scale.
The multimodal AI inflection point reached in 2024-2025 is transformative for this domain. GPT-4o (OpenAI, May 2024) demonstrated end-to-end voice interaction with under 320ms average latency and real-time vision understanding — enabling conversational interfaces that respond naturally to visual context in XR environments without pipeline-based ASR/NLU separation. Gemini 2.0 Flash Live (Google, December 2024) extended this to streaming video input, enabling AI systems to perceive and respond to the user’s physical environment in real time. These capabilities unlock genuinely natural multimodal interaction at the edge, provided sufficient compute can be positioned close to users via ETSI MEC infrastructure.
Gesture recognition has advanced to consumer deployment maturity with Apple Vision Pro (2024) achieving sub-millimetre hand tracking via Apple’s R1 chip, Meta Quest 3’s improved passthrough hand tracking, and Google’s Project Starline (now deployed in limited enterprise settings) enabling life-sized 3D video communication. These hardware advances combined with lightweight AI pose estimation models (MediaPipe, RTMPose, RTMO for real-time multi-person pose estimation) create the production substrate for ETSI Domain AI + Human Interface gesture components.
Emotion recognition faces significant governance pressure as of 2026. The EU AI Act includes emotion recognition systems used in workplace and educational settings in its high-risk annex, requiring conformity assessment and explicit transparency obligations. UK DSIT has flagged emotion recognition as an area requiring careful consideration in the UK AI Assurance Roadmap. Despite this regulatory complexity, research investment continues: a 2025 UK project using ARCHER2 supercomputer resources applied synthetic facial expression data to improve model generalisation, addressing the known demographic bias problems in existing training datasets.
UK Context
The United Kingdom has significant academic depth across the components that constitute ETSI Domain AI + Human Interface. At the University of Cambridge, the Dialogue Systems Group (Steve Young, Milica Gasic) pioneered the POMDP-based spoken dialogue framework foundational to task-oriented conversational AI; subsequent Cambridge research has extended this to deep reinforcement learning for dialogue management and transformer-based conversation modelling. The Cambridge Computer Laboratory’s Human-Inspired Computing group works on affective computing and physiological sensing for HCI. At the University of Edinburgh, the Institute for Language, Cognition and Computation contributes to speech recognition, spoken dialogue, and natural language generation; the Centre for Speech Technology Research has produced foundational work on low-latency TTS and voice conversion relevant to avatar voice synthesis. Imperial College London’s HCI group and the Adaptive Systems group work on user modelling, context-aware computing, and AI-mediated personalisation.
Bath University’s Centre for Digital Entertainment and Goldsmiths’ Department of Computing have active XR and AI creative interaction research programmes. The UK Research and Innovation (UKRI) Future Leaders Fellowships have funded multiple researchers working on AI-mediated interaction, affective computing, and accessible XR. The EPSRC’s Human-Robot Interaction programme and the UKRI TrustWorthy Autonomous Systems hub fund research directly relevant to this domain.
In Northern England, the University of Sheffield’s Speech and Language Technologies Group (co-founder of the Natural Language Processing research tradition in the UK) contributes to speech recognition and dialogue research. The University of Leeds has HCI and AI research groups, including work on tangible interaction and accessible computing. Manchester Metropolitan University and the wider Greater Manchester innovation ecosystem (including the NOMA innovation district and Manchester’s AI cluster around MediaCityUK and Bruntwood SciTech’s Manchester Science Park) house companies deploying conversational AI and intelligent interface systems for retail, customer service, and public sector applications. The BBC and ITV at MediaCityUK are active adopters of AI human interface technologies for broadcast production, interactive content, and audience engagement.
The UK AI Council’s 2025 report highlighted AI-enabled interaction design as a priority capability for the UK’s digital economy, with conversational AI and multimodal interfaces among the top five AI application categories by economic impact. The NHS Long Term Plan and NHSX AI Lab have funded evaluations of AI-mediated clinical interfaces — including AI consultation assistants, symptom checker conversational AI, and emotion-aware mental health support chatbots — creating a substantial UK public-sector testbed for ETSI Domain AI + Human Interface components.
Future Directions (2026-2030)
-
End-to-End Multimodal Native Interfaces: The full integration of speech, vision, gesture, and biosignal perception into unified multimodal transformers — eliminating pipeline seams, enabling cross-modal reasoning (understanding that a user pointing at an object while saying “that one” requires joint resolution of deictic gesture and spoken reference), and approaching the richness of human-to-human communication bandwidth in human-AI interaction.
-
Persistent Long-Term User Models: AI interface components that maintain rich, longitudinal models of individual users across sessions, devices, and metaverse environments — capturing expertise evolution, preference drift, social graph, emotional patterns, and life context — enabling genuinely personalised AI companions that know their users deeply without requiring repeated context provision.
-
On-Device and Privacy-First Interaction AI: Compressed, distilled versions of conversational and multimodal AI models running entirely on-device (smartphone, XR headset, edge node) — enabling private interaction AI that never transmits voice, video, or behavioural data to cloud servers, critical for healthcare, legal, and personal assistant use cases where data sovereignty is non-negotiable.
-
Brain-Computer Interface Integration: Non-invasive BCI sensors (EEG, fNIRS, dry-contact arrays in XR headsets) feeding AI decoders that interpret cognitive intent, attention state, and workload from neural signals — enabling direct neural interaction modalities that augment gesture and voice for users who need or prefer alternative input channels.
-
Theory of Mind and Social Intelligence: AI interface components with genuine models of user beliefs, intentions, knowledge states, and social relationships — enabling systems that understand what a user knows and doesn’t know, what they’re trying to achieve beyond their immediate utterance, and how to adapt communication style to social context and relationship type.
-
Regulatory-Compliant Emotion AI: Emotion recognition systems designed from the ground up to satisfy EU AI Act high-risk requirements — with bias-audited training data, demographic parity across protected characteristics, explicit consent mechanisms, and mandatory transparency disclosures — creating a legitimate market for affect-aware AI interfaces that currently face regulatory uncertainty.
-
AI-Mediated Accessibility at Scale: Universal design of AI human interfaces that adapt to any user’s ability profile in real time — semantic subtitles, gesture alternatives for voice, voice alternatives for gesture, cognitive simplification modes, and cross-modal re-rendering of information — making metaverse XR environments genuinely inclusive as AI perception and generation capabilities improve.
Research and Literature
- Picard, R. W. (1997). Affective Computing. Cambridge, MA: MIT Press.
- Norman, D. A. (1988). The Design of Everyday Things. New York: Basic Books.
- Weiser, M. (1991). The computer for the 21st century. Scientific American, 265(3), 94-104.
- Bolt, R. A. (1980). “Put-That-There”: Voice and gesture at the graphics interface. SIGGRAPH 1980, Computer Graphics, 14(3), 262-270.
- Oviatt, S. (2003). Multimodal interfaces. In A. Sears & J. Jacko (Eds.), The Human-Computer Interaction Handbook (pp. 413-432). Lawrence Erlbaum.
- Shneiderman, B. (1983). Direct manipulation: A step beyond programming languages. Computer, 16(8), 57-69.
- Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., & Polosukhin, I. (2017). Attention is all you need. NeurIPS 2017.
- Young, S., Gasic, M., Thomson, B., & Williams, J. D. (2013). POMDP-based statistical spoken dialogue systems: A review. Proceedings of the IEEE, 101(5), 1160-1179.
- Grudin, J. (1994). Computer-supported cooperative work: History and focus. Computer, 27(5), 19-26.
- Krizhevsky, A., Sutskever, I., & Hinton, G. E. (2012). ImageNet classification with deep convolutional neural networks. NeurIPS 2012.
- Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. NAACL 2019.
- Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., … & Lowe, R. (2022). Training language models to follow instructions with human feedback. NeurIPS 2022.
- Selvaraju, R. R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., & Batra, D. (2017). Grad-CAM: Visual explanations from deep networks via gradient-based localization. ICCV 2017.
- Gunning, D., Stefik, M., Choi, J., Miller, T., Stumpf, S., & Yang, G.-Z. (2019). XAI — Explainable artificial intelligence. Science Robotics, 4(37), eaay7120.
- Ekman, P., & Friesen, W. V. (1978). Facial Action Coding System: A Technique for the Measurement of Facial Movement. Palo Alto, CA: Consulting Psychologists Press.
- Schuller, B. W., & Batliner, A. (2014). Computational Paralinguistics: Emotion, Affect and Personality in Speech and Language Processing. Chichester: Wiley.
- Zhang, K., Zhang, Z., Li, Z., & Qiao, Y. (2016). Joint face detection and alignment using multitask cascaded convolutional networks. IEEE Signal Processing Letters, 23(10), 1499-1503.
- Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., … & Kiela, D. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. NeurIPS 2020.
- Cao, Z., Simon, T., Wei, S.-E., & Sheikh, Y. (2017). Realtime multi-person 2D pose estimation using part affinity fields. CVPR 2017.
- ETSI (2025). ETSI GS MEC 003 V4.1.1 — Multi-access Edge Computing (MEC); Framework and Reference Architecture. May 2025.
- ETSI (2025). ETSI GR MEC 043 V4.1.1 — Multi-access Edge Computing (MEC); Study on MEC Support for Advanced XR and Media Services. August 2025.
- W3C (2024). Multimodal Architecture and Interfaces (MMI). W3C Recommendation, Multimodal Interaction Working Group.
- European Commission (2024). Regulation (EU) 2024/1689 — EU Artificial Intelligence Act. Official Journal of the European Union, 12 July 2024.
- arXiv (2025). Human-AI Interaction Design Standards: A Survey. arXiv:2503.16472.
- Biometric Update (2025). UK project uses supercomputers, synthetic data to improve emotion recognition. BiometricUpdate.com, May 2025.
- OpenAI (2024). GPT-4o System Card. OpenAI Technical Report, May 2024.
- Oviatt, S., Coulston, R., & Lunsford, R. (2004). When do we interact multimodally? Cognitive load and multimodal communication patterns. ICMI 2004.
- NHS-AI Lab (2024). AI in NHS Interaction Design: Evaluation of Conversational AI Triage Pilots. NHSX Publication.
Formal Interaction Architecture Specification
The ETSI Domain AI + Human Interface domain node encodes a layered interaction architecture that can be decomposed into a formal stack of processing stages. The following specification, synthesised from ETSI MEC Phase 4 architecture documents, W3C Multimodal Interaction Framework, and ISO/IEC JTC 1/SC 42 AI-HCI guidance, defines the canonical processing pipeline for components classified under this domain:
Layer 1 — Multimodal Perception The perception layer ingests raw sensor streams and transforms them into structured symbolic representations of user behaviour. Constituent processing pipelines include: (a) Audio perception: microphone array capture → noise suppression (spectral subtraction, RNNoise-class models) → voice activity detection (LSTM-based VAD, 95%+ precision at SNR >5dB) → speaker identification (x-vector or ECAPA-TDNN speaker embeddings) → automatic speech recognition (Whisper Large v3 or domain-adapted variants, target WER <5% on clean speech, <15% in XR ambient noise) → text normalisation; (b) Visual perception: RGB-D or stereo camera capture → face detection (RetinaFace, SSD-class models) → facial landmark detection (MediaPipe Face Mesh, 468 landmarks at <5ms) → facial expression classification (action unit detection or direct emotion category prediction) → body pose estimation (MediaPipe Holistic, RTMPose, OpenPose — 33 body landmarks + 21 hand landmarks per hand) → gaze estimation (appearance-based regression, EfficientGaze-class models achieving <3° mean angular error); (c) Biosignal perception (wearable-enabled deployments): PPG heart rate variability → autonomic arousal estimation → EDA skin conductance → valence arousal mapping; (d) Behavioural context: interaction history → session state tracker → long-term user model retrieval.
Layer 2 — Intent and State Inference The inference layer combines perception outputs to construct a unified model of user intent, cognitive state, and emotional valence. Processing stages: (a) Natural Language Understanding: tokenisation → embedding (BERT/RoBERTa encoder or LLM context encoding) → Intent Recognition (multi-class classification or generative intent prediction) → slot filling (named entity recognition, constraint satisfaction) → dialogue act classification; (b) Gesture semantics: gesture trajectory segmentation → vocabulary matching (DTW or gesture CNN classifier) → deictic reference resolution (linking pointing gestures to visual objects in scene graph); (c) Affect fusion: late fusion of facial expression, prosodic, and behavioural affect signals using learned fusion weights or transformer-based cross-modal attention; (d) Context integration: world state query (scene graph, user location in XR environment, active task context) → pragmatic inference (combining intent, affect, and context into disambiguated user goal representation).
Layer 3 — Dialogue Management and Response Planning The management layer maintains conversational state and plans system responses. In LLM-based architectures, the LLM itself serves as an implicit dialogue state tracker and response planner, operating over the full conversational history encoded in the context window. In hybrid architectures, an explicit Dialogue State Tracking module maintains a structured belief state that feeds policy-based action selection: (a) Belief state update: Bayesian or neural update of dialogue state from new user turn; (b) Policy execution: learned dialogue policy (RL-trained or rule-based) selects next system action from action space (inform, request, confirm, offer, reject, escalate); (c) Response generation: template-based, retrieval-based, or generative (LLM) surface realisation converting abstract dialogue act to natural language text; (d) Retrieval-Augmented Generation grounding: vector similarity search over domain knowledge base prepends relevant facts to LLM generation context, reducing hallucination and enabling currency of information.
Layer 4 — Output Rendering The rendering layer transforms planned responses into multimodal system outputs. Constituent rendering pipelines: (a) Speech synthesis: Text-to-Speech (neural TTS: VITS, FastSpeech2, XTTS2 for voice cloning in avatar applications; streaming first-packet latency target <100ms); (b) Avatar animation: phoneme-driven lip synchronisation (wav2lip-class models), facial expression synthesis (blend shape or neural renderer), body gesture generation (co-speech gesture synthesis from speech prosody features); (c) Visual overlay: AR annotation rendering keyed to detected objects, attention-directed information overlays following gaze position; (d) Haptic feedback: tactile signal encoding of interaction events for haptic gloves or suit actuators; (e) Spatial audio: binaural rendering of synthesised voice output localised to avatar position in 3D space.
Layer 5 — Adaptation and Personalisation The adaptation layer continuously updates the user model and applies learned preferences to modulate all upstream layers. Components: (a) Preference learning: implicit feedback signals (dwell time, interaction depth, abandonment rate, emotional response) update Bayesian or neural user preference model; (b) Expertise modelling: Bayesian knowledge tracing or item response theory models update user knowledge state from interaction outcomes; (c) Interface adaptation: personalisation signals routed to Layer 1 (adapting input sensitivity thresholds for the user’s voice characteristics), Layer 2 (adapting intent priors toward the user’s observed goal patterns), Layer 3 (adapting response verbosity and technicality to user expertise level), and Layer 4 (adapting voice prosody, avatar personality, and visual complexity to user preference history).
Interaction Metrics and Evaluation Criteria
The evaluation of ETSI Domain AI + Human Interface components requires interaction-centric metrics that go beyond accuracy on held-out test sets, capturing the quality of the human-AI exchange as experienced by real users in XR environments:
-
Conversational AI Quality Metrics: Task completion rate (fraction of user goals successfully achieved without human agent escalation); conversational turns-to-completion (average number of dialogue turns required to complete a user task, with lower values indicating better intent understanding and response efficiency); first response latency (time from end-of-utterance detection to first token of AI response, target <400ms for natural-feeling dialogue in XR); BLEU/ROUGE scores for reference response alignment in task-oriented dialogues; MOS (Mean Opinion Score) for TTS voice naturalness (5-point scale, target >4.0 for consumer applications).
-
Gesture Recognition Metrics: Top-1 and Top-5 accuracy on standardised gesture vocabulary test sets (e.g., ChaLearn gesture datasets); recognition latency (time from gesture completion to label assignment, target <50ms for natural interaction); false positive rate per hour (inadvertent gesture activations, target <1/hour for consumer XR); cross-user generalisation accuracy on users not in training set (out-of-distribution robustness).
-
Emotion Recognition Metrics: Per-class F1 score across Ekman basic emotion categories (happiness, sadness, anger, fear, disgust, surprise, neutral); valence-arousal RMSE in continuous affect space; demographic parity across age, gender, and ethnicity protected groups (subject to EU AI Act high-risk assessment obligations); cross-corpus generalisation accuracy (model trained on one dataset evaluated on another, testing domain generalisation).
-
Adaptive Interface Quality Metrics: User satisfaction rating (NASA-TLX cognitive load subscales; SUS System Usability Scale, target >85 for accessible XR); personalisation uplift (difference in task completion rate between personalised and non-personalised interface conditions); accessibility equivalence metric (ratio of task completion rate for users with disabilities to users without, target >0.90 under universal design obligations).
-
Latency Budget Breakdown (MEC Context): End-to-end interaction loop latency decomposed as: sensor capture latency (<1ms with integrated headset sensors) + edge inference latency (target <20ms for gesture/emotion, <50ms for NLU, <100ms for LLM response generation at edge) + rendering latency (<11ms for 90Hz displays) + network round-trip to edge node (<5ms for co-located MEC) = total budget target <200ms for natural-feeling interaction across all modalities.
Benchmark Datasets and Evaluation Resources
ETSI Domain AI + Human Interface components are evaluated against a diverse set of benchmark datasets spanning its constituent component areas:
-
Conversational AI Benchmarks: MultiWOZ 2.4 (Ye et al., 2021) — the standard task-oriented dialogue benchmark covering hotel/taxi/restaurant/attraction/train domains with 10,000 dialogues; DSTC (Dialogue State Tracking Challenges, 1-12) benchmark series; Persona-Chat (Zhang et al., 2018) for open-domain personality-consistent dialogue; SuperGLUE (Wang et al., 2019) for underlying NLU evaluation; the Human Evaluation Framework for Conversational AI (Deriu et al., 2021).
-
Gesture Recognition Benchmarks: ChaLearn Looking at People gesture benchmark (Escalera et al., 2014, 20 Italian gesture classes); EgoGesture (Zhang et al., 2018, 83 gesture classes from egocentric video); NTU RGB+D 120 (Liu et al., 2019) for skeleton-based action and gesture recognition; Microsoft MSRC-12 Kinect gesture dataset.
-
Emotion Recognition Benchmarks: AffectNet (Mollahosseini et al., 2019, 0.4M facial images with valence/arousal/expression labels); RAVDESS (Livingstone & Russo, 2018, audio-visual acted emotion dataset); SEMAINE (McKeown et al., 2012, naturalistic audiovisual affect corpus); AVEC emotion recognition challenge series (2011-2019).
-
Multimodal Interaction Benchmarks: CMU-MOSI and CMU-MOSEI (sentiment and emotion from speech-video-text); AVA-ActiveSpeaker (Roth et al., 2020) for active speaker detection in video; FaceForensics++ (Rossler et al., 2019) for evaluating deepfake detection in human interface security.
-
Accessibility Evaluation: WCAG 2.2 / WCAG 3.0 conformance test suites; WebAIM contrast ratio calculators; screen reader compatibility evaluation frameworks (NVDA, JAWS, VoiceOver); cognitive load assessment protocols from the Universal Design for Learning (UDL) framework.
Key Terminology
-
Affective Computing: The field concerned with designing systems that recognise, interpret, simulate, and respond to human emotion and affect, through analysis of facial expression, speech prosody, physiological signals, and behavioural cues.
-
Embodied Conversational Agent (ECA): An AI agent with an animated avatar body that uses synthesised speech, facial expression, gesture, and gaze alongside verbal dialogue to communicate with human users — combining conversational AI with graphical avatar animation.
-
Foveated Rendering: A technique that renders metaverse scenes at full resolution only in the central visual area (the fovea) where gaze is directed, and at lower resolution in the periphery — requiring real-time gaze tracking and AI-based gaze prediction to be effective.
-
Gesture Vocabulary: The defined set of recognised gestures within a specific AI human interface system — typically including pointing (deictic), manipulation (pinch, grab, push), navigation (swipe, scroll), and communicative (thumbs-up, wave) gesture categories.
-
Multimodal Fusion: The combination of information from multiple input modalities (speech, gesture, gaze, text, image) to produce a unified interpretation of user intent — early fusion operates on raw signals, late fusion operates on modality-specific interpretations, and hybrid fusion combines both strategies.
-
Turn-Taking: In conversational AI embedded in multi-user XR, the mechanism by which the system detects end-of-utterance boundaries, determines when the AI versus a human should speak, and manages overlapping contributions in multi-party dialogue.
-
User Modelling: The AI process of constructing and maintaining a computational representation of an individual user’s knowledge, preferences, expertise level, goals, emotional state, and interaction history, used to personalise interface behaviour and content presentation.
-
Voice Activity Detection (VAD): The AI component that distinguishes speech from background noise and silence in audio streams — essential for initiating ASR processing only when a user is speaking, reducing false activations and computational load.