WO2026005102A1 - Methods and systems for identifying intentional and unintentional user gesture - Google Patents

Methods and systems for identifying intentional and unintentional user gesture

Info

Publication number
WO2026005102A1
WO2026005102A1 PCT/KR2024/011466 KR2024011466W WO2026005102A1 WO 2026005102 A1 WO2026005102 A1 WO 2026005102A1 KR 2024011466 W KR2024011466 W KR 2024011466W WO 2026005102 A1 WO2026005102 A1 WO 2026005102A1
Authority
WO
WIPO (PCT)
Prior art keywords
user
data
real
gesture
feature
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
PCT/KR2024/011466
Other languages
French (fr)
Inventor
Ankur Agrawal
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Samsung Electronics Co Ltd
Original Assignee
Samsung Electronics Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Samsung Electronics Co Ltd filed Critical Samsung Electronics Co Ltd
Publication of WO2026005102A1 publication Critical patent/WO2026005102A1/en
Pending legal-status Critical Current
Anticipated expiration legal-status Critical

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F3/00Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
    • G06F3/01Input arrangements or combined input and output arrangements for interaction between user and computer
    • G06F3/017Gesture based interaction, e.g. based on a set of recognized hand gestures
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F3/00Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
    • G06F3/01Input arrangements or combined input and output arrangements for interaction between user and computer
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F3/00Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
    • G06F3/01Input arrangements or combined input and output arrangements for interaction between user and computer
    • G06F3/011Arrangements for interaction with the human body, e.g. for user immersion in virtual reality
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F3/00Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
    • G06F3/01Input arrangements or combined input and output arrangements for interaction between user and computer
    • G06F3/011Arrangements for interaction with the human body, e.g. for user immersion in virtual reality
    • G06F3/013Eye tracking input arrangements
    • GPHYSICS
    • G02OPTICS
    • G02BOPTICAL ELEMENTS, SYSTEMS OR APPARATUS
    • G02B27/00Optical systems or apparatus not provided for by any of the groups G02B1/00 - G02B26/00, G02B30/00
    • G02B27/01Head-up displays
    • G02B27/0179Display position adjusting means not related to the information to be displayed
    • G02B2027/0187Display position adjusting means not related to the information to be displayed slaved to motion of at least a part of the body of the user, e.g. head, eye

Definitions

  • the present disclosure relates to extended reality and more particularly relates to a method and system for identifying intentional and unintentional user gestures while a user is immersed in an extended reality.
  • XR extended reality
  • VR virtual reality
  • AR augmented reality
  • the user freely looks around a virtual scene and is able to interact with objects using different eye gestures.
  • the user also interacts with objects in the virtual scene by different hand gestures.
  • any factor outside a Field of View (FOV) of the user immersed in the XR and which may cause unwanted user gestures is not tracked.
  • FOV Field of View
  • the existing techniques fail to provide solution for predicting any upcoming real event which is known in the real world, while the user is immersed in the extended reality.
  • a relative motion of the user which may result in an incorrect head position is not tracked.
  • unintentional gestures may lead to exceptional and harmful results.
  • the VR/AR environments are functionally guided by gestures and any unwanted gesture can lead to undesirable outcomes.
  • real-time motion tracking for any AR/VR system includes multiple approaches, for example, a marker-based motion tracking.
  • the marker-based motion tracking detects and identifies a marker and calculates a relative pose of an object, for an observer. Also, there needs to be marker positioned in advance and if the marker goes out of view, motion tracking stops.
  • Another approach is model-based tracking which uses a prior model of an environment to be tracked. The method fails in unorganized natural scenes.
  • a third approach is marker less approach which depends on natural features. This approach is effective tracking performance in unprepared environments. However, performance is degraded when facing motion blur and fast motion. Also, real-time performance for mobile AR/VR is beyond the traditional video frame-rate.
  • the method includes detecting a real-world gesture performed by a user immersed in a virtual reality (VR) environment and an augmented reality (AR) environment. Further, the method includes receiving data associated with the user and corresponding surroundings. Further, the method includes determining a plurality of data features based on the received data and the real-world gesture. The method further includes categorizing the plurality of data features into at least one user feature and at least one event feature. The method further includes estimating an intent probability of the gesture based on an analysis of the gesture and the categorized plurality of data features. Further, the method includes identifying that the real-world gesture is intentional for the VR/AR environment based on determining that the intent probability is above a predefined threshold.
  • VR virtual reality
  • AR augmented reality
  • the method includes identifying that the gesture is unintentional when the intent probability is below the predefined threshold. Further, the method includes continuing with a scene of the AR/VR environment based on the identification that the gesture is unintentional.
  • a system of identifying an intentional gesture comprises a processor and a memory communicatively coupled to the processor.
  • the memory is configured to store instructions which when executed by the processor causes the processor to detect a real-world gesture performed by a user immersed in a virtual reality (VR) environment and an augmented reality (AR) environment.
  • the processor is configured to receive data associated with the user and corresponding surroundings.
  • the processor is configured to determine a plurality of data features based on the received data and the real-world gesture. Further, the processor is configured to categorize the plurality of data features into at least one user feature and at least one event feature.
  • the processor is configured to estimate an intent probability of the gesture based on an analysis of the gesture and the categorized plurality of data features. Further, the processor is configured to identify that the real-world gesture is intentional for the VR/AR environment based on determining that the intent probability is above a predefined threshold.
  • a method of identifying an intentional gesture may comprise: detecting a real-world gesture performed by a user immersed in a virtual reality or an augmented reality (VR/AR) environment.
  • the method may comprise receiving 1504 data associated with the user and corresponding surroundings.
  • the method may comprise determining a plurality of data features based on the received data and the real-world gesture.
  • the method may comprise categorizing the plurality of data features into at least one user feature and at least one event feature.
  • the method may comprise estimating an intent probability of the real-world gesture based on the real-world gesture and the categorized plurality of data features.
  • the method may comprise identifying that the real-world gesture is intentional for the VR/AR environment based on the intent probability eing above a first threshold.
  • a computer-readable storage medium may store instructions.
  • the instructions when executed by at least one processor, may cause the at least one processor to perform the method.
  • a system (100) of identifying an intentional gesture may comprise at least one processor (104) comprising processing circuitry.
  • the system may comprise memory (102) comprising one or more storage media, the memory communicatively coupled to the processor (104), and configured to store instructions which, when executed by the at least one processor (104) individually or collectively, causes the system (100) to: detect a real-world gesture performed by a user immersed in a virtual reality (VR) or an augmented reality (VR/AR) environment.
  • the instructions may be further configured to, when executed by the at least one processor (104) individually or collectively, cause the system (100) to receive data associated with the user and corresponding surroundings.
  • the instructions may be further configured to, when executed by the at least one processor (104) individually or collectively, cause the system (100) to determine a plurality of data features based on the received data and the real-world gesture.
  • the instructions may be further configured to, when executed by the at least one processor (104) individually or collectively, cause the system (100) to categorize the plurality of data features into at least one user feature and at least one event feature.
  • the instructions may be further configured to, when executed by the at least one processor (104) individually or collectively, cause the system (100) to estimate an intent probability of the real-world gesture based on the real-world gesture and the categorized plurality of data features.
  • the instructions may be further configured to, when executed by the at least one processor (104) individually or collectively, cause the system (100) to identify that the real-world gesture is intentional for the VR/AR environment based on the intent probability being above a first threshold.
  • FIG. 1 illustrates a block diagram of a system for identifying intentional gestures, according to one or more embodiments of the present disclosure
  • FIG. 2 illustrates a block diagram depicting modules of the system, according to one or more embodiments of the present disclosure
  • FIG. 3a illustrates an exemplary scenario of the system depicting a user immersed in a virtual reality environment, according to one or more embodiments of the present disclosure
  • FIG. 3b illustrates a plurality of devices communicatively coupled to the user immersed in the virtual reality environment, according to one or more embodiments of the present disclosure
  • FIG. 4 illustrates a block diagram of a smart data processing engine of the system, according to one or more embodiments of the present disclosure
  • FIG. 5 illustrates a block diagram of a hybrid semantic fusion unit of the system, according to one or more embodiments of the present disclosure
  • FIG. 6 illustrates a block diagram of an intelligent unintentional event estimator operationally coupled with a virtual reality (VR) scene correction unit of the system, according to one or more embodiments of the present disclosure
  • FIG. 7 illustrates a flowchart depicting an operational flow of the system, according to one or more embodiments of the present disclosure
  • FIG. 8 illustrates an exemplary scenario depicting different mechanisms to perform gesture detection when the user is immersed in the VR scene, according to one or more embodiments of the present disclosure
  • FIG. 9 illustrates an exemplary scenario depicting the user's out of sync emotion detection by a visual semantic gathering unit, according to one or more embodiments of the present disclosure
  • FIG. 10 illustrates an exemplary scenario of a convolutional neural network (CNN) model depicting the user's real emotion detection by the visual semantic gathering unit, according to one or more embodiments of the present disclosure
  • FIG. 11 illustrates an exemplary scenario depicting user expected emotion as per content by the visual semantic gathering unit, according to one or more embodiments of the present disclosure
  • FIG. 12 illustrates a flowchart of the visual semantic gathering unit depicting a field of view (FOV) scene user attention detection, according to one or more embodiments of the present disclosure
  • FIG. 13 illustrates a table depicting event feature and user feature calculation by the hybrid semantic fusion unit, according to one or more embodiments of the present disclosure
  • FIG. 14 illustrates a table having a plurality of features as a training data for the intelligent unintentional event estimator, according to one or more embodiments of the present disclosure.
  • FIG. 15 illustrates a flowchart depicting a method of identifying an intentional gesture, according to one or more embodiments of the present disclosure.
  • any terms used herein such as but not limited to “includes,” “comprises”, “has”, “have”, and grammatical variants thereof do not specify an exact limitation or restriction and certainly do not exclude the possible addition of one or more features or elements, unless otherwise stated, and must not be taken to exclude the possible removal of one or more of the listed features and elements, unless otherwise stated with the limiting language “must comprise” or “needs to include.”
  • a or B at least one of A or/and B, or “one or more of A or/and B” used in the various embodiments of the present disclosure include any and all combinations of words enumerated with it.
  • “A or B,” “at least one of A and B,” or “at least one of A or B” means (1) including at least one A, (2) including at least one B, or (3) including both at least one A and at least one B.
  • first and second used in various embodiments of the present disclosure may modify various elements of various embodiments, these terms do not limit the corresponding elements. For example, these terms do not limit an order and/or importance of the corresponding elements. These terms may be used for the purpose of distinguishing one element from another element.
  • a first user device and a second user device all indicate user devices and may indicate different user devices.
  • a first element may be named a second element without departing from the scope of right of various embodiments of the present disclosure, and similarly, a second element may be named a first element.
  • a processor configured to (set to) perform A, B, and C may be a dedicated processor, for example, an embedded processor, for performing a corresponding operation, or a generic-purpose processor, for example, a Central Processing Unit (CPU) or an application processor (AP), capable of performing a corresponding operation by executing one or more software programs stored in a memory device.
  • a dedicated processor for example, an embedded processor, for performing a corresponding operation
  • a generic-purpose processor for example, a Central Processing Unit (CPU) or an application processor (AP), capable of performing a corresponding operation by executing one or more software programs stored in a memory device.
  • CPU Central Processing Unit
  • AP application processor
  • a term "module” used in the present document may imply a unit including, for example, one of hardware, software, and firmware or a combination of two or more of them.
  • the “module” may be interchangeably used with a term such as a unit, a logic, a logical block, a component, a circuit, and the like.
  • the “module” may be a minimum unit of an integrally constituted component or may be a part thereof.
  • the “module” may be a minimum unit for performing one or more functions or may be a part thereof.
  • the “module” may be mechanically or electrically implemented.
  • the "module” of the present disclosure may include at least one of an Application-Specific Integrated Circuit (ASIC) chip, a Field-Programmable Gate Arrays (FPGAs), and a programmable-logic device, which are known or will be developed, and which perform certain operations.
  • ASIC Application-Specific Integrated Circuit
  • FPGAs Field-Programmable Gate Arrays
  • programmable-logic device which are known or will be developed, and which perform certain operations.
  • modules that carry out a described function or functions.
  • modules which may be referred to herein as units or blocks or the like, or may include blocks or units, are physically implemented by analog or digital circuits such as logic gates, integrated circuits, microprocessors, microcontrollers, memory circuits, passive electronic components, active electronic components, optical components, hardwired circuits, or the like, and may optionally be driven by firmware and software.
  • the circuits may, for example, be embodied in one or more semiconductor chips, or on substrate supports such as printed circuit boards and the like.
  • circuits constituting a block may be implemented by dedicated hardware, by a processor (e.g., one or more programmed microprocessors and associated circuitry), or by a combination of dedicated hardware to perform some functions of the block and a processor to perform other functions of the block.
  • a processor e.g., one or more programmed microprocessors and associated circuitry
  • Each block of the embodiments may be physically separated into two or more interacting and discrete blocks without departing from the scope of the disclosure.
  • the blocks of the embodiments may be physically combined into more complex blocks without departing from the scope of the disclosure.
  • the objective of the present disclosure is to provide users with a more streamlined and effortless approach to identify intentional and unintentional gestures performed by a user or happening within a vicinity of the user, when the user is immersed in a virtual reality (VR) or augmented reality (AR) environment.
  • the present disclosure detects a real-world gesture performed by the user immersed in the VR/AR environment.
  • the real-world gesture may include the movement of user by way of blinking of eyes, left or right, or both, winking of eyes, hand movement, while being immersed in the VR/AR environment.
  • the present disclosure gathers data and groups features within the data into categories. Further, the categories and grouped features are used to train a convolutional neural network (CNN) model such that the model is trained based on the grouped features.
  • CNN convolutional neural network
  • the present disclosure further implements the trained model to determine or estimate a probability of the intention of the gesture to be intentional or unintentional and therefore, based on the estimation of a scene within the AR/VR environment is adjusted.
  • the system and the method disclosed by the present disclosure are described in greater detail in conjunction with FIGS. 1-15.
  • FIG. 1 illustrates a block diagram of a system 100 for identifying intentional gestures, according to one or more embodiments of the present disclosure.
  • Figure 1 is described in conjunction with Figures 2-15.
  • the system 100 may include a memory 102 and a processor 104 communicatively coupled to the memory 102.
  • the system 100 may include one or more processors or at least one processor.
  • the one or more processors or at least one processor may individually or collectively execute the computer-readable instructions.
  • the processor 104 may be implemented as one or more microprocessors, microcomputers, microcontrollers, digital signal processors, central processing units, state machines, logic circuitries, processing circuitries, and/or any devices that manipulate signals based on operational instructions.
  • the processor 104 may be configured to fetch and execute computer-readable instructions (e.g., instructions 106) and data stored in the memory 102.
  • the processor 104 may be a general-purpose processor, such as a central processing unit (CPU), an application processor (AP), or the like, and an AI-dedicated processor such as a neural processing unit (NPU).
  • the processor 104 may control the processing of input data in accordance with a predefined operating rule or artificial intelligence (AI) model stored in the non-volatile memory and the volatile memory, i.e., the memory 102.
  • the predefined operating rule or artificial intelligence model is provided through training or learning.
  • the processor 104 may be operatively coupled to each of the memory, the I/O Interface.
  • the processor 104 may be configured to process, execute, or perform a plurality of operations described herein.
  • the memory 102 may include any non-transitory computer-readable medium known in the art including, for example, volatile memory, such as static random-access memory (SRAM) and dynamic random-access memory (DRAM), and/or non-volatile memory, such as read-only memory (ROM), erasable programmable ROM, flash memories, hard disks, optical disks, and magnetic tapes.
  • volatile memory such as static random-access memory (SRAM) and dynamic random-access memory (DRAM)
  • non-volatile memory such as read-only memory (ROM), erasable programmable ROM, flash memories, hard disks, optical disks, and magnetic tapes.
  • the memory 102 is communicatively coupled with the processor to store processing instructions for completing the process. Further, the memory may include an operating system for performing one or more tasks of the system, as performed by a generic operating system in a computing domain.
  • the memory 102 is operable to store instructions executable by the processor 104.
  • system 100 may include a set of instructions that can be executed to cause the system 100 to perform any one or more of the methods disclosed.
  • the system 100 may operate as a standalone device or may be connected, e.g., using a network, to other computer systems or peripheral devices.
  • the system 100 may operate in the capacity of a server or as a client user computer in a server-client user network environment, or as a peer system in a peer-to-peer (or distributed) network environment.
  • the system 100 can also be implemented as or incorporated across various devices, such as a personal computer (PC), a tablet PC, a personal digital assistant (PDA), a mobile device, a palmtop computer, a laptop computer, a desktop computer, a communications device, a wireless telephone, a land-line telephone, a web appliance, a network router, switch or bridge, or any other machine capable of executing a set of instructions (sequential or otherwise) that specify actions to be taken by that machine.
  • PC personal computer
  • PDA personal digital assistant
  • a mobile device a palmtop computer
  • laptop computer a laptop computer
  • a desktop computer a communications device
  • a wireless telephone a land-line telephone
  • web appliance a web appliance
  • network router switch or bridge
  • the system 100 may include the processor 104 e.g., a central processing unit (CPU), a graphics processing unit (GPU), or both.
  • the processor 104 may be a component in a variety of systems.
  • the processor 104 may be part of a standard personal computer or a workstation.
  • the processor 104 may be one or more general processors, digital signal processors, application-specific integrated circuits, field-programmable gate arrays, servers, networks, digital circuits, analog circuits, combinations thereof, or other now known or later developed devices for analyzing and processing data.
  • the processor 104 may implement a software program, such as code generated manually (i.e., programmed).
  • the system 100 may include the memory 102, such as a memory 102 that can communicate via a bus 122.
  • the memory 102 may include but is not limited to computer-readable storage media such as various types of volatile and non-volatile storage media, including but not limited to random access memory, read-only memory, programmable read-only memory, electrically programmable read-only memory, electrically erasable read-only memory, flash memory, magnetic tape or disk, optical media and the like.
  • memory 102 includes a cache or random-access memory for the processor 104.
  • the memory 102 is separate from the processor 104, such as a cache memory of a processor, the system memory, or other memory.
  • the memory 102 may be an external storage device or database for storing data.
  • the memory 102 is operable to store instructions 106 executable by the processor 104.
  • the functions, acts or tasks illustrated in the figures or described may be performed by the programmed processor 104 for executing the instructions 106 stored in the memory 102.
  • the functions, acts or tasks are independent of the particular type of instructions set, storage media, processor or processing strategy and may be performed by software, hardware, integrated circuits, firmware, micro-code and the like, operating alone or in combination.
  • processing strategies may include multiprocessing, multitasking, parallel processing and the like.
  • the system 100 may or may not further include a display unit 114, such as a liquid crystal display (LCD), an organic light-emitting diode (OLED), a flat panel display, a solid-state display, a cathode ray tube (CRT), a projector, a printer or other now known or later developed display device for outputting determined information.
  • a display unit 114 such as a liquid crystal display (LCD), an organic light-emitting diode (OLED), a flat panel display, a solid-state display, a cathode ray tube (CRT), a projector, a printer or other now known or later developed display device for outputting determined information.
  • the display 114 may act as an interface for the user to see the functioning of the processor 104, or specifically as an interface with the software stored in the memory 102 or a drive unit 108.
  • the system 100 may include an input device 116 configured to allow the user to interact with any of the components of system 100.
  • the system 100 may also include the drive unit 108.
  • the drive unit 108 may include a computer-readable medium 110 in which one or more sets of instructions 112, e.g., software, can be embedded.
  • the instructions 112 may embody one or more of the methods or logic as described. In an example, the instructions 112 may reside completely, or at least partially, within the memory 102 or within the processor 104 during execution by the system 100.
  • the present disclosure contemplates a computer-readable medium that includes instructions 112 or receives and executes instructions 112 responsive to a propagated signal so that a device connected to a network 120 can communicate voice, video, audio, images, or any other data over the network 120. Further, the instructions 112 may be transmitted or received over the network 120 via a communication port or interface 118 or using a bus.
  • the communication port or interface 118 may be a part of the processor 104 or maybe a separate component.
  • the communication port or interface 118 may be created in software or maybe a physical connection in hardware.
  • the communication port or interface 118 may be configured to connect with a network 120, external media, the display 114, or any other components in system 100, or combinations thereof.
  • connection with the network 120 may be a physical connection, such as a wired Ethernet connection or may be established wirelessly as discussed later.
  • additional connections with other components of the system 100 may be physical or may be established wirelessly.
  • the network 120 may alternatively be directly connected to the bus 122.
  • the network 120 may include wired networks, wireless networks, Ethernet AVB networks, or combinations thereof.
  • the wireless network may be a cellular telephone network, an 802.11, 802.16, 802.20, 802.1Q or WiMax network.
  • the network 120 may be a public network, such as the Internet, a private network, such as an intranet, or combinations thereof, and may utilize a variety of networking protocols now available or later developed including, but not limited to TCP/IP based networking protocols.
  • the system 100 may not be limited to operation with any particular standards and protocols. For example, standards for Internet and other packet-switched network transmissions (e.g., TCP/IP, UDP/IP, HTML, and HTTP) may be used.
  • the processor 104 may be configured to detect the real-world gesture performed by the user immersed in the VR environment. Further, the processor 104 may be configured to receive data associated with the user and corresponding surroundings. In an embodiment, the data corresponds to one or more of visual data, inertial measurement unit (IMU) sensors-based data, ambient sensors-based data, audio data, and multimedia communication data received from the plurality of devices in the vicinity of the user. In an embodiment, the real-world gesture performed by the user may be detected by a plurality of devices within the vicinity of the user. In an embodiment, the plurality of devices may include a smartphone, a smartwatch, a VR headset, and cameras disposed within the vicinity of the user.
  • IMU inertial measurement unit
  • the processor 104 may be configured to determine a plurality of data features based on the received data and the real-world gesture.
  • the plurality of data features may include a visual user emotion feature, a visual user attention feature, an inertial measurement unit (IMU) event prediction feature, an IMU motion feature, and an ambient event prediction feature.
  • IMU inertial measurement unit
  • the processor 104 may be configured to categorize the plurality of data features into at least one user feature and at least one event feature.
  • the at least one user feature may be indicative of visual user emotion by analyzing visual data received via the plurality of devices in the vicinity of the user to detect whether an emotion of the user matches an expected emotion for currently displayed content in the VR/AR environment.
  • the processor 104 may be configured to create a visual map of a field of view (FOV) of the user and track one or more objects within the surroundings of the user. Further, the processor 104 may be configured to determine the at least one user feature that is indicative of visual user attention based on the visual map to detect the attention of the user with respect to each of the one or more objects.
  • FOV field of view
  • the processor 104 may be configured to determine the at least one event feature indicative of a motion or an ambient event occurring in the vicinity of the user based on analysing data associated with one or more IMU sensors located within the VR headset, and one or more ambient sensors, audio, and multimedia communication from the plurality of devices in the vicinity of the user.
  • the processor 104 may be configured to collect data related to the one or more events outside and inside the FOV of the user by the plurality of devices outside the FOV of the user. In an embodiment, the processor 104 may be configured to determine the at least one event feature indicating a probable ambient event occurring in an area outside the FOV of the user based on analysis of the collected data. In an embodiment, the processor 104 may be configured to track changes in user facial expressions and eye gaze for object tracking within the FOV of the user. In an embodiment, the processor 104 may be configured to collect data from the one or more ambient sensors to anticipate events outside the FOV of the user.
  • the processor 104 may be configured to estimate an intent probability of the gesture based on an analysis of the gesture and the categorized plurality of data features.
  • the processor 104 may be configured to identify that the real-world gesture is intentional for a VR/AR environment 602, as shown in Figure 6, based on determining that the intent probability is above a predefined threshold.
  • the processor 104 may be configured to identify that the gesture is unintentional when the intent probability is below the predefined threshold. Further, the processor 104 may be configured to continue with the scene of the AR/VR environment based on the identification that the gesture is unintentional.
  • FIG. 2 illustrates a block diagram depicting modules of the system 100, according to one or more embodiments of the present disclosure.
  • the system 100 may include the following modules: an input device data observer 202, a smart data processing engine 204, a hybrid semantic fusion unit 206, an intelligent unintentional event estimator 208, and a VR scene gesture executor 210.
  • the input device data observer 202 may be configured to fetch data relative to the user and surroundings. All the devices in vicinity of the user of the system (200) may be used to fetch data.
  • the input device data observer 202 may be configured to fetch data using mobile devices, watch or wearable devices, and/or Internet-of-Things (IoT) devices in vicinity of the user (e.g. a user 212) of the system (200).
  • IoT Internet-of-Things
  • AR/VR headset may be configured to capture information of surroundings of the user.
  • the smart data processing engine 204 may be adapted to receive the real-world gesture.
  • the real-world gesture may be received from a user 212.
  • the user 212 may be a VR user a VR/AR user.
  • the smart data processing engine 204 may be adapted to process the received real-world gesture into a plurality of data features.
  • the smart data processing engine 204 may fetch all raw data observed from input devices.
  • the raw data may be obtained by the input devices data observer 202.
  • the smart data processing engine 204 may learn and analyse the raw data.
  • the smart data processing engine 204 may process the raw data into one or more key data features.
  • the hybrid semantic fusion unit 206 may be adapted to categorise the plurality of data features.
  • the hybrid semantic fusion unit 206 may fuse the plurality of data features into the at least one user feature and the at least one event feature, as mentioned above.
  • the hybrid semantic fusion unit 206 may intelligently fuse received data at any particular time.
  • the hybrid semantic fusion unit 206 may receive information related to any gesture performed at any time.
  • the hybrid semantic fusion unit 206 may fuse received information related to gestures performed and user-generated (or VR/AR system-generated) pose data in real time.
  • the intelligent unintentional event estimator 208 may be adapted to estimate the intent probability of the gesture based on the analysis of the gesture and the at least one user feature and the at least one event feature.
  • the intelligent unintentional event estimator 208 may employ one or more machine learning model configured to mark a gesture to be intentional or unintentional.
  • the intelligent unintentional event estimator 208 may evaluate the gesture using one or more machine learning models.
  • the intelligent unintentional event estimator 208 may mark the gesture to be intentional or unintentional.
  • the VR scene gesture executor 210 may dynamically adjust the VR/AR environment based on user pose modifications.
  • the VR scene gesture executor 210 may provide valid intentional gestures to the user 212 or the VR/AR headset.
  • the VR Scene gesture executor 210 may transmit valid intentional gestures to the VR/AR headset.
  • the intelligent unintentional event estimator 208 may determine that the intent probability is above or equal to a predefined threshold (or a first threshold). In this case, the real-world gesture is intentional for the VR/AR environment. In this case, the VR scene gesture executor 210 may adjust the scene of the AR/VR environment. In an embodiment, the predefined threshold may be 40%, but not limited thereto. In one case, the intelligent unintentional event estimator 208 may determine that the intent probability is below the predefined threshold. In this case, the gesture is unintentional. In this case, the VR scene gesture executor 210 may continue with the scene within the VR/AR environment.
  • a predefined threshold or a first threshold
  • FIG. 3a illustrates an exemplary scenario of the system 100 depicting the user 212 immersed in the VR/AR environment, according to one or more embodiments of the present disclosure.
  • FIG. 3b illustrates a plurality of devices 302 communicatively coupled to the user 212 immersed in the VR environment, according to one or more embodiments of the present disclosure. Referring to Fig. 3b, while three devices 302a, 302b, and 302c are depicted, it may be apparent that only one or more than one device may be deployed within the scope of the disclosure.
  • the plurality of devices 302 may include a phone 302a, a smartwatch (or a wearable device) 302b, an AR/VR headset 302c, etc.
  • the AR/VR headset 302c may be referred to as VR headset 302c and may be used interchangeably throughout the detailed description.
  • the plurality of devices 302 may be communicatively coupled with the processor 104 via the network 120.
  • the input device data observer 202 may be adapted to monitor the plurality of devices 302 and capture (or fetch) raw data from the plurality of devices 302.
  • the raw data may include a phone data, a wearable data, and/or a VR/AR headset data.
  • the phone data may include but not limited to audio, location, call/message, alarm, and emergency reminder.
  • the wearable data may include but not limited to heart rate, motion tracking, skin temperature, sleep cycle, and proximity.
  • the VR/AR headset data may include IMU (e.g., sensor information such as information obtained from a gyroscope sensor and/or an accelerometer sensor), face and face tracker unit (e.g., facial expressions and/or eye focus), object tracking in user FOV, and/or head orientations.
  • the plurality of devices 302 may be referred to as user connected devices.
  • an AR/VR system may detect a real-world gesture of a user.
  • the user in an AR/VR world may use the AR/VR system.
  • the AR/VR system may detect the real-world gesture of the user at time t+1.
  • the AR/VR system may obtain data from input devices around the user.
  • the AR/VR system may process the obtained data.
  • FIG. 4 illustrates a block diagram of the smart data processing engine 204 of the system 100, according to one or more embodiments of the present disclosure.
  • the smart data processing engine 204 may include a visual semantic gathering unit 402, an IMU attribute tracker 404, and an ambient information converging unit 406 communicatively coupled to the smart data processing engine 204.
  • the visual semantic gathering unit 402 may be adapted to track visual based data using a camera. Further, the visual semantic gathering unit 402 may be adapted to create the visual map of the FOV and track events within the FOV of the user 212. Further, the visual semantic gathering unit 402 may be adapted to capture the visual scene data of the user 212.
  • the visual semantic gathering unit 402 may be adapted to detect gesture by various mechanisms, as illustrated in FIG. 8.
  • the user 212 wearing the VR headset 302c may perform the gesture, for example by hand.
  • the visual semantic gathering unit 402 may be adapted to extract gesture image and a Red-Green-Blue (RGB) image of the gesture image is captured. Further, a heat map may be generated and a gesture type may be predicted. The gesture type may be a feature in the AR/VR environment.
  • the visual semantic gathering unit 402 may be adapted to analyse user visual information and detect changes in facial expressions and eye gaze object tracking of the user 212.
  • the visual semantic gathering unit 402 may be adapted for the VR user 212 out of synchronization emotion detection and FOV scene user attention detection, as illustrated in FIG. 12.
  • the VR user 212 out of synchronization emotion detection of the visual semantic gathering unit 402 may be adapted to match expected emotion as per the VR scene and emotion at any time (t).
  • the VR scene may be referred to as a VR/AR scene and may be used interchangeably throughout the detailed description.
  • the VR user 212 out of synchronization emotion detection may be adapted to predict whether the emotion might be due to real-world impact.
  • the VR user 212 may express emotion as per VR content/scene and may express emotion out of synchronization as per the VR content.
  • the FOV scene user attention detection may be adapted to create the visual map of the scene and track all objects in the vicinity of the VR user 212. Further, the FOV scene user attention detection may be adapted to calculate the attention of the VR user 212 within an object.
  • the visual semantic gathering unit 402 will be described in greater detail in conjunction with FIGS. 9-11 in the later part of the detailed description.
  • the IMU attribute tracker 404 may be adapted to capture sensor based data for the user 212 from the VR/AR headset 302c and the connected devices.
  • the IMU sensors may include a gyroscope, an accelerometer, a position sensor, etc.
  • the sensor based data may include an angular velocity, orientation, centre of mass of location, acceleration, and velocity.
  • the IMU attribute tracker 404 may be adapted to integrate data based on time.
  • the IMU sensor information may be monitored. The sudden movement and relative motion probability may be detected, for example, using Kalman Filter and artificial intelligence (AI) expert system.
  • the IMU attribute tracker 404 may have a monitoring service running that constantly monitors all the parameters.
  • the IMU attribute tracker 404 may be adapted to generate output features based on the sensor based data, including an IMU event prediction feature and an IMU motion feature.
  • the IMU event prediction feature may depict movement at any given time (t) from rest or constant motion position. Further, the IMU motion feature may depict relative motion probability at any given time (t).
  • the IMU sensors may comprise a gyroscope sensor, an accelerometer sensor, and/or a position sensor.
  • the sensor data obtained from the IMU sensors may be calibrated. Calibrated sensor data may be applied to the Kalman Filter.
  • calibrated sensor data obtained from the gyroscope sensor may be applied to an orientation Kalman filter. Filtered results from the orientation Kalman filter may comprise angular velocity information and orientation information.
  • calibrated sensor data obtained from the accelerometer sensor may be applied to a position Kalman filter.
  • Calibrated sensor data obtained from the position sensor may be applied to the position Kalman filter.
  • Time derivative may be also applied to the Calibrated sensor data obtained from the position sensor.
  • the results of the time derivative may comprise velocity information.
  • the results of the orientation Kalman Filter, the position Kalman filter, and the time derivative may be applied to the AI expert system.
  • the AI expert system may comprise an expert knowledge database, rules for steady state detection, and rules for relative motion detections. Based on the database and the rules included in the AI expert system. the AI expert system may generate one or more IMU event prediction features and one or more IMU motion features.
  • the ambient information converging unit 406 may be adapted to capture and assemble different real-world ambient parameters.
  • the real-world ambient parameters may include audio, phone information, and the one or more parameters obtained from ambient sensors.
  • the audio may comprise from a phone ringing, a voice of mother, mother shouting from another room.
  • the phone information may comprise an urgent call situation, dad calling, etc.
  • the one or more parameters obtained from ambient sensors may comprise a phone located near the user. Based on the real-world ambient parameters an ambient activity may be detected.
  • the ambient information converging unit 406 may be adapted to gather data related to events outside/inside FOV of the user 212, such as audio, relative motion, and visual scene data using other devices outside FOV of the user 212. Based on the gathered data the ambient information converging unit 406 may be adapted to track probable events to happen in the area outside FOV of the user 212.
  • the ambient information converging unit 406 may be adapted to feed the sensor-based data and the real-world ambient parameters to an ambient information collector. Further, the ambient information converging unit 406 may generate a predicted output, such as an ambient activity feature. In an embodiment, the ambient information converging unit 406 may be adapted to generate a pre-trained dataset for a supervised machine learning engine (for example, a support vector machine (SVM) ML) model 702, as shown in FIG. 7. Further, the SVM ML model 702 may be adapted to learn about the information being captured from the ambient surroundings, i.e., the real-world ambient parameters. In an embodiment, an activity happening outside the FOV of the user 212 may be detected.
  • a supervised machine learning engine for example, a support vector machine (SVM) ML
  • SVM ML model 702 may be adapted to learn about the information being captured from the ambient surroundings, i.e., the real-world ambient parameters. In an embodiment, an activity happening outside the FOV of the user 212 may
  • the SVM ML model 702 may be adapted to be fed with input parameters, for example, the ambient sensors, audio data, and multimedia/communication on a user device.
  • the ambient sensor may be the IMU sensor and the input parameters from the ambient sensors may include an ultrasound sensor to locate a person, a vibration sensor to detect a sitting or standing state, infrared sensor to detect entry, a microwave sensor to detect small movements, and/or thermal sensor to detect physical activity.
  • the input parameters from the audio data may be employed to identify a speaker, deduce an emergency content, and/or detect queries to the VR user 212.
  • the input parameters from the multimedia may be utilized to detect urgent calls/messages, and categorize urgent multimedia.
  • the SVM ML model 702 may be adapted to generate output parameters based on the received/fed input parameters.
  • the output parameters may include the IMU event prediction feature to depict the probability of distraction of the VR user 212.
  • the AV/VR system may determine one or more IMU prediction event features and one or more IMU motion features.
  • the AV/VR system may determine an IMU prediction event feature and an IMU motion feature at time t+1 based on head orientation data, eye focus changing data, acceleration data, and/or a heart rate data, etc. For example, at time t+1, the IMU prediction event feature (or its value) may be determined as '0.7' and the IMU motion feature (or its value) may be determined as '0'.
  • the AV/VR system may determine one or more ambient event prediction features.
  • the AV/VR system may determine an ambient event prediction feature at time t+1 based on head orientation data, eye focus changing data, etc. For example, at time t+1, the ambient event prediction feature (or its value) may be determined as 'event: emergency, loud object damage, a weight of 0.7'.
  • FIG. 5 illustrates a block diagram 500 of the hybrid semantic fusion unit 206 of the system 100, according to one or more embodiments of the present disclosure.
  • the hybrid semantic fusion unit 206 may also be referred to as a hybrid content synthesizer.
  • the hybrid semantic fusion unit 206 may be adapted to receive the plurality of data features.
  • the plurality of data features may be referred to as processed data features.
  • the plurality of data features may comprise, for example, the visual user emotion feature, the visual user attention feature, the inertial measurement unit (IMU) event prediction feature, the IMU motion feature, and/or the ambient event prediction feature, as mentioned earlier.
  • the hybrid semantic fusion unit 206 may be adapted to categorize the plurality of data features into at least one user feature 502 and at least one event feature 504, as discussed earlier.
  • the at least one user feature 502 and at least one event feature 504 may be fed to the intelligent unintentional event estimator 208.
  • the intelligent unintentional event estimator 208 is described in conjunction with FIG. 6.
  • the hybrid semantic fusion unit 206 may receive inputs from the smart data processing engine 204, for example, the visual semantic gathering unit 402, the IMU attribute tracker 404, the ambient information converging unit 406, and fuse the received inputs into broader categories, the at least one user feature 502 and the at least one event feature 504, for eliminating redundant data.
  • the redundant data may be eliminated by considering features from multiple input sources, e.g., same features may be considered in visual, sensor, and ambient sources. Further, during classification or categorization, any feature may only be considered once, thereby eliminating redundant data. In an example, similar attributes from all different data processing units are merged for better score data.
  • a user feature may be determined as a combination of emotion feature and an attention feature.
  • an event feature may be determined as a combination of a user event, an ambient event, and a relative motion.
  • the hybrid semantic fusion unit 206 may receive a first visual user emotion feature determined as 'shock, 3', a first visual user attention feature determined as 'loss in attention', a first IMU even prediction feature determined as '0.7', a first IMU motion feature determined as '0', and a first ambient event prediction feature determined as 'event: emergency, loud damage, a weight of 0.7'.
  • the hybrid semantic fusion unit 206 may calculate a first user feature based on the first visual user emotion feature and the first visual user attention feature.
  • the first user feature may be determined as (1, 1).
  • the hybrid semantic fusion unit 206 may calculate a first event feature based on the first IMU event prediction feature, the first IMU motion feature, and the first ambient event prediction feature.
  • the first event feature may be determined as (1, 0, 1).
  • FIG. 6 illustrates a block diagram 600 of the intelligent unintentional event estimator 208 operationally coupled with a VR scene correction unit 210 of the system 100, according to one or more embodiments of the present disclosure.
  • the intelligent unintentional event estimator 208 may be adapted to estimate the intent probability of the gesture based on an analysis of the gesture and the at least one user feature 502 and the at least one event feature 504. In an example embodiment, the intelligent unintentional event estimator 208 may estimate that an intention probability is 30% and a relative motion probability is 55%. The intelligent unintentional event estimator 208 may identify that the real-world gesture is intentional for the VR/AR environment 602 or the VR scene 602 based on determining that the intent probability is above the predefined threshold, as discussed above.
  • the VR scene correction unit 210 may be adapted to adjust the VR scene 602 for the VR/AR environment, as mentioned earlier.
  • the VR scene correction unit 210 may include a gesture analyzer unit 604, a final user pose evaluator unit 606 and a VR scene updater unit 608.
  • the gesture analyzer unit 604 may be adapted to atomically analyze gesture input and prepare to update the event to the VR scene 602.
  • the final user pose evaluator unit 606 may be adapted to finalize the relative pose of the user 212 based on different calculated attributes.
  • the VR scene updater unit 608 may be adapted to update the VR scene for both inputs relative pose and gesture applicability.
  • the intelligent unintentional event estimator 208 may determine that (a) the gesture of the user at corresponding time is unintentional and should be rejected and (b)the motion of the user at corresponding time is not affected.
  • the VR scene updater unit 608 may adjust based on the determination (a) and (b). For example, based on that the gesture is determined as unintentional and the motion is determined as not affected, the VR scene updater unit 608 may continue a current VR scene of VR system.
  • FIG. 7 illustrates a flowchart 700 depicting an operational flow of the system 100, according to one or more embodiments of the present disclosure.
  • the system 100 may include the following modules such as the input device data observer 202, the smart data processing engine 204, the hybrid semantic fusion unit 206, the intelligent unintentional event estimator 208, and the VR scene gesture executor 210 or the VR scene correction unit 210.
  • the intelligent unintentional event estimator 208 may be referred to as the SVM ML model 702 as described above. The operation of each module or element of the system 100 depicted by the flowchart is already explained above in greater detail in conjunction with Figure 1-6.
  • FIG. 8 illustrates an exemplary scenario 800 depicting different mechanisms to perform gesture detection when the user 212 is immersed in the VR scene 602, according to one or more embodiments of the present disclosure.
  • objects present with the FOV of the user 212 may be tracked and analyzed using the camera installed within the VR headset 302c.
  • the IMU sensors may be inbuilt within the VR headset 302c.
  • the objects present outside the FOV of the user 212 that is, ambient information, may be collected using a camera arrangement and the phone, watch content, etc.
  • the user 212 immersed within the VR scene 602 may perform the gesture and the performed gesture may be analyzed and based on the gesture performed, whether intentional or unintentional, the feature in the AR/VR environment may be detected.
  • Alex wearing the VR headset 302c and immersed in the VR scene performs a hand gesture like a gun
  • the visual semantic gathering unit 402 of the smart data processing engine 204 detects a heat map of the hand and predicts the type of gesture based on the heat map.
  • FIG. 9 illustrates an exemplary scenario 900 depicting the user 212 out-of-sync emotion detection by the visual semantic gathering unit 402, according to one or more embodiments of the present disclosure.
  • the visual semantic gathering unit 402 may be adapted to determine the out-of-synchronization emotion detection of the VR user 212.
  • the VR user 212 may be immersed in the VR scene.
  • the VR user 212 based on the VR scene may perform an emotion.
  • the emotion may be categorized as a real emotion and an expected emotion.
  • the real emotion of the VR/AR user 212 may be detected and at 906, the expected emotion as per content by the VR/AR user may be detected.
  • the real emotion may be categorized as an emotion 1 model and the expected emotion may be categorized as an emotion 2 model.
  • an emotion contradiction calculator may be adapted to calculate emotion deviation strength of the emotion 1 model and the emotion 2 model.
  • a visual user emotion feature may be detected.
  • the real emotion and the expected emotion of the user 212 as per the content within the VR scene may be further explained in conjunction with FIGS. 10-11.
  • the AR/VR system may determine one or more visual user emotion features.
  • the AR/VR system may determine a areal emotion of the user and a AR/VR content emotion (e.g., one or more expected emotions as per AR/VR content).
  • the AR/VR system may match the expected emotion for the AR/VR content and the real emotion of the user at the time t+1.
  • the AR/VR system may compare a strength (or a probability or a deviation between the expected emotion and the real emotion) of each emotion with a predefined threshold. Based on the comparison, the AR/VR system may determine whether a visual user emotion feature is intentional or unintentional.
  • the emotion 'shocking' is determined as intentional (and/or the visual user emotion feature corresponding to the emotion 'shocking' is determined as intentional).
  • FIG. 10 illustrates an exemplary scenario 1000 of training and verifying a convolutional neural network (CNN) model depicting the user's real emotion detection by the visual semantic gathering unit 402, according to one or more embodiments of the present disclosure.
  • CNN convolutional neural network
  • the CNN model may be adapted to detect the real emotion (emotion 1 model) of the VR user 212.
  • the exemplary scenario 1000 may include a training stage 1002 and a verification state 1004.
  • the training stage 1002 may be adapted to train a deep CNN model 1006 using a dataset, e.g., an image dataset to generate a pre-trained model 1008.
  • the pre-trained model 1008 may be adapted to feed with emotion recognition adaption to form the pre-trained model 1010 with a new dense layer.
  • the pre-trained model 1010 may be fine-tuned using a cleansed dataset 1012 from a facial expression dataset 1014.
  • the facial expression dataset 1014 may include multiple facial expressions.
  • the cleansed dataset 1012 may be obtained from the facial expression dataset 1014 by cropping facial parts of expressions. Further, by fine-tuning the pre-trained model 1010, an emotion recognition model 1016 may be generated (or obtained).
  • the emotion recognition model 1016 may be verified or tested at the verification stage 1006.
  • the verification stage 1016 may be performed by feeding a cropped picture of a face to the emotion recognition model 1016.
  • the output of the emotion recognition model 1016 may be estimated based on corresponding gesture of the face.
  • the corresponding gesture of the face may be predefined for the verification.
  • the emotion recognition model 1016 may be adapted to predict output probability of each emotion. Therefore, the real emotion of the user 212 may be detected from the emotion recognition model 1016.
  • FIG. 11 illustrates an exemplary scenario 1100 depicting user expected emotion as per content by the visual semantic gathering unit 402, according to one or more embodiments of the present disclosure.
  • the expected emotion of the user 212 as per content may be determined by training a single task CNN model from the VR scene and by one or more feature extractors from an audio of the VR scene and surroundings.
  • the one or more feature extractors may be adapted to extract features from the VR/AR content.
  • a clip from the VR scene 602 may be input to the pre-trained single task CNN model.
  • the pre-trained single task CNN model may extract one or more features from the clip.
  • the feature extractor may comprise a 3x3 convolution layer and a max pooling unit (or a max pooling layer).
  • the audio may include a zeta function corresponding to singularities in space and time.
  • the zeta function may be generated from the audio by the system 100.
  • the zeta function may be input to the one or more feature extractors.
  • the one or more feature extractors may extract corresponding feature(s) from the zeta function.
  • the one or more feature extractors and the single task CNN model may be configured to generate a feature fused layer.
  • the feature fused layer may include one or more fully convolutional layers to generate an output.
  • the output may include a predicted user emotion.
  • FIG. 12 illustrates a flowchart of the visual semantic gathering unit 402 depicting the FOV scene user attention detection 1200, according to one or more embodiments of the present disclosure.
  • the FOV scene user attention detection 1200 may be determined by the visual semantic gathering unit 402.
  • the VR user 212 attentional information of user 212 in real-world scenes may be identified by an eye-tracking paradigm.
  • the eye-tracking paradigm may directly compare eye movements while the VR user 212 may be immersed in the VR/AR environment.
  • a meaning map of a scene may be created using a spatial distribution algorithm for high-level semantic features in an environment, e.g., faces, objects, doors, etc.
  • the meaning map may be relevant to understanding the semantic content and affordances available to the VR user 212 in the VR/AR scene.
  • a gaze map may be captured and compared with the meaning map to calculate attention parameters.
  • the attention features may include an average fixation duration (AFD), an average fixation number (AFN), gaze shifts (GS), and a gaze spherical distribution (GSD).
  • the attention parameters may include an active attention and a passive attention. The active attention and the passive attention may be calculated by superimposing the meaning map into the gaze map.
  • the user's attention on semantically meaning regions may be analyzed as a whole for the VR scene. For example, for any scene for a given time (t) the attention parameters for the VR user 212 may be calculated.
  • the AFD may be an average of user fixation duration across scenes.
  • the user fixation duration is inversely proportional to the attention, for example, less is the user fixation duration, and more is the attention.
  • the AFN may be an average of total fixation across scenes. For example, less is the AFN, more is the attention.
  • GS may be larger in the active attention.
  • the GSD may be less centrally trending in the active attention.
  • a change in the attention strength of user 212 for the FOV scene while performing the gesture may need to be considered to validate the intention of the gesture.
  • the AR/VR system may determine one or more visual user attention features. For example, the AR/VR system may calculate corresponding attention of the user with each object. Accordingly, the AR/VR system may determine that the user attention has been changed regarding a first object and a second object. In this case, a count of visual user attention features may be calculated as 2; that is, the count of visual user attention features may correspond to a number of objects of which the user attention has been changed.
  • FIG. 13 illustrates an exemplary table depicting an event feature calculation 1300 and a user feature calculation 1302 by the hybrid semantic fusion unit 206, according to one or more embodiments of the present disclosure.
  • the event feature calculation 1300 may be segregated into three parts, for example, the IMU event prediction feature, the ambient event prediction feature and the IMU motion feature.
  • the event feature calculation 1300 may be evaluated as the at least one event feature with features as, a user event, an ambient event, and a relative motion.
  • the IMU event prediction feature may generate the user event
  • the ambient event prediction feature may generate the ambient event
  • the IMU motion feature may generate the relative motion.
  • the user feature calculation 1302 may be grouped into two parts, for example, a visual user emotion feature, and a visual user attention feature.
  • the user feature calculation 1302 may determine the at least one user feature with features such as the emotion and the attention.
  • a visual emotion feature (VEF) is greater than or equal to the predefined threshold (or a first threshold) (e.g., 0.5)
  • the emotion may be intentional.
  • VEF is less than 0.5
  • the emotion may be unintentional.
  • a count (or a number) of user attention feature(s) (UAF) is greater than or equal to another predetermined threshold (e.g., 2)
  • the attention may be intentional.
  • UAF is less than 2
  • the attention may be unintentional.
  • the predefined threshold of the VEF is directly dependent on expressions of the user and may vary among users based on expressions.
  • FIG. 14 illustrates a table having a plurality of features as a training data for the intelligent unintentional event estimator 208, according to one or more embodiments of the present disclosure.
  • the intelligent unintentional event estimator 208 may correspond to the SVM ML model 702 that may be adapted to determine a multi-layer perceptron model (MLP) by training on a dataset of gestures, user features, event features, user intentions and a relative motion.
  • the features of all feeder machine learning engines may be both static and/or dynamic.
  • the fused feature scores of all machine learning engines are fed to the intelligent unintentional event estimator 208.
  • the VR user 212 may perform a left hand gesture, a user feature 1 and an event feature 3 may be detected based on the left hand gesture and the user intention may be intentional and the relative motion may be positive. These feature scores of the left hand gesture may be used as training data for training the MLP.
  • content is passed as training data to a supervised machine learning engine, i.e., MLP.
  • FIG. 15 illustrates a flowchart depicting a method of identifying an intentional gesture, according to one or more embodiments of the present disclosure.
  • the method 1500 may include detecting a real-world gesture performed by the user 212 immersed in a virtual reality (VR) environment and/or an augmented reality (AR) environment.
  • the method 1500 may include receiving data associated with the user 212 and corresponding surroundings.
  • the method 1500 may include determining a plurality of data features based on the received data and the real-world gesture. Further, at step 1508, the method 1500 may include categorizing the plurality of data features into at least one user feature and at least one event feature.
  • the method 1500 may include determining the at least one user feature indicative of visual user emotion by analyzing visual data received via one or more devices in the vicinity of the user 212 to detect whether an emotion of the user matches an expected emotion for currently displayed content in the VR/AR environment.
  • the method 1500 may include analyzing visual data received via one or more devices in the vicinity of the user 212, detecting whether an emotion of the user matches an expected emotion for currently displayed content in the VR/AR environment based on the analyzed visual data, and determining a first user feature indicative of visual user emotion based on detecting whether the emotion of the user matches the expected emotion.
  • the method 1500 may include creating the visual map of the FOV of the user 212 and tracking one or more objects within the surroundings of the user 212.
  • the method 1500 may include determining the at least one user feature indicative of visual user attention based on the visual map to detect attention of the user with respect to each of the one or more objects. For example, the method 1500 may include detecting an attention of the user with respect to each of the one or more objects and determining a second user feature indicative of visual user attention based on the visual map and the detected attention of the user.
  • the method 1500 may include analysing data associated with at least one of: one or more inertial measurement unit (IMU) sensors located within a VR headset, one or more ambient sensors, audio, or multimedia communication from one or more devices in vicinity of the user; and determining the at least one event feature indicative of a motion or an ambient event occurring in vicinity of the user based on the analyzed data.
  • IMU inertial measurement unit
  • the method 1500 may include collecting data related to the one or more events outside and inside a field of view (FOV) of the user by one or more devices outside the FOV of the user; and determining the at least one event feature indicating a probable ambient event occurring in an area outside the FOV of the user based on the collected data.
  • the method 1500 may include analyzing the collected data and determining the at least one event feature indicating the probable ambient event based on the analyzed collected data.
  • the method 1500 may include estimating an intent probability of the gesture based on an analysis of the gesture and the categorized plurality of data features.
  • the method 1500 may include identifying that the real-world gesture is intentional for the VR/AR environment based on the intent probability being above the predefined (or a first) threshold. Alternatively or additionally, the method 1500 may include comparing the intent probability with the first threshold. In an embodiment, the method 1500 may include adjusting a scene within the VR/AR environment based on identifying whether the real-world gesture is intentional. Alternatively or additionally, the method 1500 may include identifying that the gesture is unintentional based on the intent probability being below the predefined threshold (or the first threshold). Alternatively or additionally, the method 1500 may include continuing with a scene of the AR/VR environment based on identifying that the gesture is unintentional.
  • the present disclosure may be adapted to ensure a smooth streamlined experience of the AR/VR environment for the user.
  • the present disclosure may be adapted to identify gestures performed by the user while he/she is immersed in the VR/AR environment.
  • the present disclosure classifies them into different categories and estimates a probability that a certain gesture performed by the user is intentionally done by the user or is unintentional.
  • a certain gesture is an act by the user or any other act/event happening within the vicinity of the user affecting the user. In this manner, the user experiences an AR/VR experience without any disturbance and experiences enhanced or controlled VR scenes.
  • the user is not affected by any factor outside the FOV of the user which may cause unwanted user gesture.

Landscapes

  • Engineering & Computer Science (AREA)
  • General Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Human Computer Interaction (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • User Interface Of Digital Computer (AREA)

Abstract

Disclosed herein is a method (1500) of identifying an intentional gesture. The method (1500) includes detecting (1502) a real-world gesture performed by a user immersed in a virtual reality or an augmented reality (VR/AR) environment and receiving (1504) data associated with the user and corresponding surroundings. Further, the method (1500) includes determining (1506) a plurality of data features based on the received data and the real-world gesture. Further, the method (1500) includes categorizing (1508) the plurality of data features into at least one user feature and at least one event feature. Further, the method (1500) includes estimating (1510) an intent probability of the real-world gesture based on the real-world gesture and the categorized plurality of data features. Further, the method (1500) includes identifying (1512) that the real-world gesture is intentional for the VR/AR environment based on the intent probability being above a first threshold.

Description

METHODS AND SYSTEMS FOR IDENTIFYING INTENTIONAL AND UNINTENTIONAL USER GESTURE
The present disclosure relates to extended reality and more particularly relates to a method and system for identifying intentional and unintentional user gestures while a user is immersed in an extended reality.
With a rise of virtual reality technology, different types of interaction methods are introduced for a user in an extended reality (XR) including both virtual reality (VR) environment and augmented reality (AR) environment. The user freely looks around a virtual scene and is able to interact with objects using different eye gestures. The user also interacts with objects in the virtual scene by different hand gestures. In existing techniques, any factor outside a Field of View (FOV) of the user immersed in the XR and which may cause unwanted user gestures is not tracked. Also, the existing techniques fail to provide solution for predicting any upcoming real event which is known in the real world, while the user is immersed in the extended reality. Additionally, a relative motion of the user which may result in an incorrect head position is not tracked. In certain critical situations, unintentional gestures may lead to exceptional and harmful results. The VR/AR environments are functionally guided by gestures and any unwanted gesture can lead to undesirable outcomes.
Currently, real-time motion tracking for any AR/VR system includes multiple approaches, for example, a marker-based motion tracking. The marker-based motion tracking detects and identifies a marker and calculates a relative pose of an object, for an observer. Also, there needs to be marker positioned in advance and if the marker goes out of view, motion tracking stops. Another approach is model-based tracking which uses a prior model of an environment to be tracked. The method fails in unorganized natural scenes. A third approach is marker less approach which depends on natural features. This approach is effective tracking performance in unprepared environments. However, performance is degraded when facing motion blur and fast motion. Also, real-time performance for mobile AR/VR is beyond the traditional video frame-rate.
There is a need to provide a fast and robust system that efficiently maps real world state of user with the immersed virtual world and a tracks user pose to analyze different real world possible situations and distinguish intentional and unintentional gestures performed by the user. The system is required for avoiding unintentional user gesture while the user is immersed in the virtual scene and therefore, avoiding unwanted event in the AR/VR environment.
Therefore, in light of the above-mentioned challenges, a solution is required to overcome the above-mentioned challenges associated with identifying the intentional and unintentional user gesture when the user is immersed in the virtual reality or augmented reality environments.
This summary is provided to introduce a selection of concepts in a simplified format that is further described in the detailed description of the disclosure. This summary is not intended to identify key or essential inventive concepts of the disclosure, nor is it intended to determine the scope of the disclosure.
Disclosed herein is a method for identifying an intentional gesture. The method includes detecting a real-world gesture performed by a user immersed in a virtual reality (VR) environment and an augmented reality (AR) environment. Further, the method includes receiving data associated with the user and corresponding surroundings. Further, the method includes determining a plurality of data features based on the received data and the real-world gesture. The method further includes categorizing the plurality of data features into at least one user feature and at least one event feature. The method further includes estimating an intent probability of the gesture based on an analysis of the gesture and the categorized plurality of data features. Further, the method includes identifying that the real-world gesture is intentional for the VR/AR environment based on determining that the intent probability is above a predefined threshold.
In an embodiment, the method includes identifying that the gesture is unintentional when the intent probability is below the predefined threshold. Further, the method includes continuing with a scene of the AR/VR environment based on the identification that the gesture is unintentional.
In one or more embodiments, a system of identifying an intentional gesture. The system comprises a processor and a memory communicatively coupled to the processor. The memory is configured to store instructions which when executed by the processor causes the processor to detect a real-world gesture performed by a user immersed in a virtual reality (VR) environment and an augmented reality (AR) environment. Further, the processor is configured to receive data associated with the user and corresponding surroundings. Further, the processor is configured to determine a plurality of data features based on the received data and the real-world gesture. Further, the processor is configured to categorize the plurality of data features into at least one user feature and at least one event feature. Further, the processor is configured to estimate an intent probability of the gesture based on an analysis of the gesture and the categorized plurality of data features. Further, the processor is configured to identify that the real-world gesture is intentional for the VR/AR environment based on determining that the intent probability is above a predefined threshold.
According to an embodiment of the present disclosure, a method of identifying an intentional gesture may comprise: detecting a real-world gesture performed by a user immersed in a virtual reality or an augmented reality (VR/AR) environment. The method may comprise receiving 1504 data associated with the user and corresponding surroundings. The method may comprise determining a plurality of data features based on the received data and the real-world gesture. The method may comprise categorizing the plurality of data features into at least one user feature and at least one event feature. The method may comprise estimating an intent probability of the real-world gesture based on the real-world gesture and the categorized plurality of data features. The method may comprise identifying that the real-world gesture is intentional for the VR/AR environment based on the intent probability eing above a first threshold.
According to an embodiment of the present disclosure, a computer-readable storage medium may store instructions. The instructions, when executed by at least one processor, may cause the at least one processor to perform the method.
According to an embodiment of the present disclosure, a system (100) of identifying an intentional gesture may comprise at least one processor (104) comprising processing circuitry. The system may comprise memory (102) comprising one or more storage media, the memory communicatively coupled to the processor (104), and configured to store instructions which, when executed by the at least one processor (104) individually or collectively, causes the system (100) to: detect a real-world gesture performed by a user immersed in a virtual reality (VR) or an augmented reality (VR/AR) environment. The instructions may be further configured to, when executed by the at least one processor (104) individually or collectively, cause the system (100) to receive data associated with the user and corresponding surroundings. The instructions may be further configured to, when executed by the at least one processor (104) individually or collectively, cause the system (100) to determine a plurality of data features based on the received data and the real-world gesture. The instructions may be further configured to, when executed by the at least one processor (104) individually or collectively, cause the system (100) to categorize the plurality of data features into at least one user feature and at least one event feature. The instructions may be further configured to, when executed by the at least one processor (104) individually or collectively, cause the system (100) to estimate an intent probability of the real-world gesture based on the real-world gesture and the categorized plurality of data features. The instructions may be further configured to, when executed by the at least one processor (104) individually or collectively, cause the system (100) to identify that the real-world gesture is intentional for the VR/AR environment based on the intent probability being above a first threshold.
To further clarify the advantages and features of the present disclosure, a more particular description of the disclosure will be rendered by reference to specific embodiments thereof, which is illustrated in the appended drawing. It is appreciated that these drawings depict only typical embodiments of the disclosure and are therefore not to be considered limiting its scope. The disclosure will be described and explained with additional specificity and detail with the accompanying drawings.
These and other features, aspects, and advantages of the present disclosure will become better understood when the following detailed description is read with reference to the accompanying drawings in which like characters represent like parts throughout the drawings, wherein:
FIG. 1 illustrates a block diagram of a system for identifying intentional gestures, according to one or more embodiments of the present disclosure;
FIG. 2 illustrates a block diagram depicting modules of the system, according to one or more embodiments of the present disclosure;
FIG. 3a illustrates an exemplary scenario of the system depicting a user immersed in a virtual reality environment, according to one or more embodiments of the present disclosure;
FIG. 3b illustrates a plurality of devices communicatively coupled to the user immersed in the virtual reality environment, according to one or more embodiments of the present disclosure;
FIG. 4 illustrates a block diagram of a smart data processing engine of the system, according to one or more embodiments of the present disclosure;
FIG. 5 illustrates a block diagram of a hybrid semantic fusion unit of the system, according to one or more embodiments of the present disclosure;
FIG. 6 illustrates a block diagram of an intelligent unintentional event estimator operationally coupled with a virtual reality (VR) scene correction unit of the system, according to one or more embodiments of the present disclosure;
FIG. 7 illustrates a flowchart depicting an operational flow of the system, according to one or more embodiments of the present disclosure;
FIG. 8 illustrates an exemplary scenario depicting different mechanisms to perform gesture detection when the user is immersed in the VR scene, according to one or more embodiments of the present disclosure;
FIG. 9 illustrates an exemplary scenario depicting the user's out of sync emotion detection by a visual semantic gathering unit, according to one or more embodiments of the present disclosure;
FIG. 10 illustrates an exemplary scenario of a convolutional neural network (CNN) model depicting the user's real emotion detection by the visual semantic gathering unit, according to one or more embodiments of the present disclosure;
FIG. 11 illustrates an exemplary scenario depicting user expected emotion as per content by the visual semantic gathering unit, according to one or more embodiments of the present disclosure;
FIG. 12 illustrates a flowchart of the visual semantic gathering unit depicting a field of view (FOV) scene user attention detection, according to one or more embodiments of the present disclosure;
FIG. 13 illustrates a table depicting event feature and user feature calculation by the hybrid semantic fusion unit, according to one or more embodiments of the present disclosure;
FIG. 14 illustrates a table having a plurality of features as a training data for the intelligent unintentional event estimator, according to one or more embodiments of the present disclosure; and
FIG. 15 illustrates a flowchart depicting a method of identifying an intentional gesture, according to one or more embodiments of the present disclosure.
Further, skilled artisans will appreciate those elements in the drawings are illustrated for simplicity and may not have necessarily been drawn to scale. For example, the flow charts illustrate the method in terms of the most prominent steps involved to help and improve understanding of aspects of the present disclosure. Furthermore, in terms of the construction of the device, one or more components of the device may have been represented in the drawings by conventional symbols, and the drawings may show only those specific details that are pertinent to understanding the embodiments of the present disclosure so as not to obscure the drawings with details that will be readily apparent to those of ordinary skill in the art having the benefit of the description herein.
It should be understood at the outset that although illustrative implementations of the embodiments of the present disclosure are illustrated below, the present disclosure may be implemented using any number of techniques, whether currently known or in existence. The present disclosure should in no way be limited to the illustrative implementations, drawings, and techniques illustrated below, including the exemplary design and implementation illustrated and described herein, but may be modified within the scope of the appended claims along with their full scope of equivalents.
The term "some", "one or more embodiment", "one or more example embodiments", as used herein is defined as "one, or more than one, or all." Accordingly, the terms "one," "more than one," "more than one, but not all" or "all" would all fall under the definition of "some." The term "some embodiments" may refer to one embodiment, several embodiments, or to all embodiments. Accordingly, the term "some embodiments" is defined as meaning "one embodiment, or more than one embodiment, or all embodiments."
The terminology and structure employed herein are for describing, teaching, and illuminating some embodiments and their specific features and elements and do not limit, restrict, or reduce the spirit and scope of the claims or their equivalents.
More specifically, any terms used herein such as but not limited to "includes," "comprises", "has", "have", and grammatical variants thereof do not specify an exact limitation or restriction and certainly do not exclude the possible addition of one or more features or elements, unless otherwise stated, and must not be taken to exclude the possible removal of one or more of the listed features and elements, unless otherwise stated with the limiting language "must comprise" or "needs to include."
Whether or not a certain feature or element was limited to being used only once, either way, it may still be referred to as "one or more features", "one or more elements", "at least one feature" or "at least one element." Furthermore, the use of the terms "one or more" or "at least one" feature or element does not preclude there being none of that feature or element unless otherwise specified by limiting language such as "there needs to be one or more 쪋" or "one or more element is required."
The terms "A or B," "at least one of A or/and B," or "one or more of A or/and B" used in the various embodiments of the present disclosure include any and all combinations of words enumerated with it. For example, "A or B," "at least one of A and B,” or "at least one of A or B" means (1) including at least one A, (2) including at least one B, or (3) including both at least one A and at least one B.
Although the term such as "first" and "second" used in various embodiments of the present disclosure may modify various elements of various embodiments, these terms do not limit the corresponding elements. For example, these terms do not limit an order and/or importance of the corresponding elements. These terms may be used for the purpose of distinguishing one element from another element. For example, a first user device and a second user device all indicate user devices and may indicate different user devices. For example, a first element may be named a second element without departing from the scope of right of various embodiments of the present disclosure, and similarly, a second element may be named a first element.
The expression "configured to (or set to)" used in various embodiments of the present disclosure may be replaced with "suitable for," "having the capacity to," "designed to," "adapted to," "made to," or "capable of" according to the situation. The term "configured to (set to)" does not necessarily mean "specifically designed to" as hardware. Instead, the expression "apparatus configured to . . . " may mean that the apparatus is "capable of . . . " along with other devices or parts in a certain situation. For example, "a processor configured to (set to) perform A, B, and C" may be a dedicated processor, for example, an embedded processor, for performing a corresponding operation, or a generic-purpose processor, for example, a Central Processing Unit (CPU) or an application processor (AP), capable of performing a corresponding operation by executing one or more software programs stored in a memory device.
A term "module" used in the present document may imply a unit including, for example, one of hardware, software, and firmware or a combination of two or more of them. The "module" may be interchangeably used with a term such as a unit, a logic, a logical block, a component, a circuit, and the like. The "module" may be a minimum unit of an integrally constituted component or may be a part thereof. The "module" may be a minimum unit for performing one or more functions or may be a part thereof. The "module" may be mechanically or electrically implemented. For example, the "module" of the present disclosure may include at least one of an Application-Specific Integrated Circuit (ASIC) chip, a Field-Programmable Gate Arrays (FPGAs), and a programmable-logic device, which are known or will be developed, and which perform certain operations.
Unless otherwise defined, all terms, and especially any technical and/or scientific terms, used herein may be taken to have the same meaning as commonly understood by one having ordinary skill in the art.
The embodiments herein and the various features and advantageous details thereof are explained more fully with reference to the non-limiting embodiments that are illustrated in the accompanying drawings and detailed in the following description. Descriptions of well-known components and processing techniques are omitted so as to not unnecessarily obscure the embodiments herein.
As is traditional in the field, embodiments may be described and illustrated in terms of modules that carry out a described function or functions. These modules, which may be referred to herein as units or blocks or the like, or may include blocks or units, are physically implemented by analog or digital circuits such as logic gates, integrated circuits, microprocessors, microcontrollers, memory circuits, passive electronic components, active electronic components, optical components, hardwired circuits, or the like, and may optionally be driven by firmware and software. The circuits may, for example, be embodied in one or more semiconductor chips, or on substrate supports such as printed circuit boards and the like. The circuits constituting a block may be implemented by dedicated hardware, by a processor (e.g., one or more programmed microprocessors and associated circuitry), or by a combination of dedicated hardware to perform some functions of the block and a processor to perform other functions of the block. Each block of the embodiments may be physically separated into two or more interacting and discrete blocks without departing from the scope of the disclosure. Likewise, the blocks of the embodiments may be physically combined into more complex blocks without departing from the scope of the disclosure.
The accompanying drawings are used to help easily understand various technical features and it should be understood that the embodiments presented herein are not limited by the accompanying drawings. As such, the present disclosure should be construed to extend to any alterations, equivalents, and substitutes in addition to those which are particularly set out in the accompanying drawings. Although the terms first, second, third, etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are generally only used to distinguish one element from another.
Embodiments of the present disclosure will be described below in detail with reference to the accompanying drawings.
The objective of the present disclosure is to provide users with a more streamlined and effortless approach to identify intentional and unintentional gestures performed by a user or happening within a vicinity of the user, when the user is immersed in a virtual reality (VR) or augmented reality (AR) environment. The present disclosure detects a real-world gesture performed by the user immersed in the VR/AR environment. The real-world gesture may include the movement of user by way of blinking of eyes, left or right, or both, winking of eyes, hand movement, while being immersed in the VR/AR environment. The present disclosure then gathers data and groups features within the data into categories. Further, the categories and grouped features are used to train a convolutional neural network (CNN) model such that the model is trained based on the grouped features. The present disclosure further implements the trained model to determine or estimate a probability of the intention of the gesture to be intentional or unintentional and therefore, based on the estimation of a scene within the AR/VR environment is adjusted. The system and the method disclosed by the present disclosure are described in greater detail in conjunction with FIGS. 1-15.
FIG. 1 illustrates a block diagram of a system 100 for identifying intentional gestures, according to one or more embodiments of the present disclosure. Figure 1 is described in conjunction with Figures 2-15.
The system 100 may include a memory 102 and a processor 104 communicatively coupled to the memory 102. The system 100 may include one or more processors or at least one processor. According to an embodiment, the one or more processors or at least one processor may individually or collectively execute the computer-readable instructions. The processor 104 may be implemented as one or more microprocessors, microcomputers, microcontrollers, digital signal processors, central processing units, state machines, logic circuitries, processing circuitries, and/or any devices that manipulate signals based on operational instructions. Among other capabilities, the processor 104 may be configured to fetch and execute computer-readable instructions (e.g., instructions 106) and data stored in the memory 102. At this time, the processor 104 may be a general-purpose processor, such as a central processing unit (CPU), an application processor (AP), or the like, and an AI-dedicated processor such as a neural processing unit (NPU). The processor 104 may control the processing of input data in accordance with a predefined operating rule or artificial intelligence (AI) model stored in the non-volatile memory and the volatile memory, i.e., the memory 102. The predefined operating rule or artificial intelligence model is provided through training or learning. Further, the processor 104 may be operatively coupled to each of the memory, the I/O Interface. The processor 104 may be configured to process, execute, or perform a plurality of operations described herein.
The memory 102 may include any non-transitory computer-readable medium known in the art including, for example, volatile memory, such as static random-access memory (SRAM) and dynamic random-access memory (DRAM), and/or non-volatile memory, such as read-only memory (ROM), erasable programmable ROM, flash memories, hard disks, optical disks, and magnetic tapes. The memory 102 is communicatively coupled with the processor to store processing instructions for completing the process. Further, the memory may include an operating system for performing one or more tasks of the system, as performed by a generic operating system in a computing domain. The memory 102 is operable to store instructions executable by the processor 104.
In some embodiments, the system 100 may include a set of instructions that can be executed to cause the system 100 to perform any one or more of the methods disclosed. The system 100 may operate as a standalone device or may be connected, e.g., using a network, to other computer systems or peripheral devices.
In a networked deployment, the system 100 may operate in the capacity of a server or as a client user computer in a server-client user network environment, or as a peer system in a peer-to-peer (or distributed) network environment. The system 100 can also be implemented as or incorporated across various devices, such as a personal computer (PC), a tablet PC, a personal digital assistant (PDA), a mobile device, a palmtop computer, a laptop computer, a desktop computer, a communications device, a wireless telephone, a land-line telephone, a web appliance, a network router, switch or bridge, or any other machine capable of executing a set of instructions (sequential or otherwise) that specify actions to be taken by that machine. Further, while a single system 100 is illustrated, the term "system" shall also be taken to include any collection of systems or sub-systems that individually or jointly execute a set, or multiple sets, of instructions to perform one or more computer functions.
As discussed, the system 100 may include the processor 104 e.g., a central processing unit (CPU), a graphics processing unit (GPU), or both. The processor 104 may be a component in a variety of systems. For example, the processor 104 may be part of a standard personal computer or a workstation. The processor 104 may be one or more general processors, digital signal processors, application-specific integrated circuits, field-programmable gate arrays, servers, networks, digital circuits, analog circuits, combinations thereof, or other now known or later developed devices for analyzing and processing data. The processor 104 may implement a software program, such as code generated manually (i.e., programmed).
As mentioned above, the system 100 may include the memory 102, such as a memory 102 that can communicate via a bus 122. The memory 102 may include but is not limited to computer-readable storage media such as various types of volatile and non-volatile storage media, including but not limited to random access memory, read-only memory, programmable read-only memory, electrically programmable read-only memory, electrically erasable read-only memory, flash memory, magnetic tape or disk, optical media and the like. In one example, memory 102 includes a cache or random-access memory for the processor 104. In alternative examples, the memory 102 is separate from the processor 104, such as a cache memory of a processor, the system memory, or other memory. The memory 102 may be an external storage device or database for storing data. The memory 102 is operable to store instructions 106 executable by the processor 104. The functions, acts or tasks illustrated in the figures or described may be performed by the programmed processor 104 for executing the instructions 106 stored in the memory 102. The functions, acts or tasks are independent of the particular type of instructions set, storage media, processor or processing strategy and may be performed by software, hardware, integrated circuits, firmware, micro-code and the like, operating alone or in combination. Likewise, processing strategies may include multiprocessing, multitasking, parallel processing and the like.
As shown, the system 100 may or may not further include a display unit 114, such as a liquid crystal display (LCD), an organic light-emitting diode (OLED), a flat panel display, a solid-state display, a cathode ray tube (CRT), a projector, a printer or other now known or later developed display device for outputting determined information. The display 114 may act as an interface for the user to see the functioning of the processor 104, or specifically as an interface with the software stored in the memory 102 or a drive unit 108.
Additionally, the system 100 may include an input device 116 configured to allow the user to interact with any of the components of system 100. The system 100 may also include the drive unit 108. The drive unit 108 may include a computer-readable medium 110 in which one or more sets of instructions 112, e.g., software, can be embedded. Further, the instructions 112 may embody one or more of the methods or logic as described. In an example, the instructions 112 may reside completely, or at least partially, within the memory 102 or within the processor 104 during execution by the system 100.
The present disclosure contemplates a computer-readable medium that includes instructions 112 or receives and executes instructions 112 responsive to a propagated signal so that a device connected to a network 120 can communicate voice, video, audio, images, or any other data over the network 120. Further, the instructions 112 may be transmitted or received over the network 120 via a communication port or interface 118 or using a bus. The communication port or interface 118 may be a part of the processor 104 or maybe a separate component. The communication port or interface 118 may be created in software or maybe a physical connection in hardware. The communication port or interface 118 may be configured to connect with a network 120, external media, the display 114, or any other components in system 100, or combinations thereof. The connection with the network 120 may be a physical connection, such as a wired Ethernet connection or may be established wirelessly as discussed later. Likewise, the additional connections with other components of the system 100 may be physical or may be established wirelessly. The network 120 may alternatively be directly connected to the bus 122.
The network 120 may include wired networks, wireless networks, Ethernet AVB networks, or combinations thereof. The wireless network may be a cellular telephone network, an 802.11, 802.16, 802.20, 802.1Q or WiMax network. Further, the network 120 may be a public network, such as the Internet, a private network, such as an intranet, or combinations thereof, and may utilize a variety of networking protocols now available or later developed including, but not limited to TCP/IP based networking protocols. The system 100 may not be limited to operation with any particular standards and protocols. For example, standards for Internet and other packet-switched network transmissions (e.g., TCP/IP, UDP/IP, HTML, and HTTP) may be used.
The processor 104 may be configured to detect the real-world gesture performed by the user immersed in the VR environment. Further, the processor 104 may be configured to receive data associated with the user and corresponding surroundings. In an embodiment, the data corresponds to one or more of visual data, inertial measurement unit (IMU) sensors-based data, ambient sensors-based data, audio data, and multimedia communication data received from the plurality of devices in the vicinity of the user. In an embodiment, the real-world gesture performed by the user may be detected by a plurality of devices within the vicinity of the user. In an embodiment, the plurality of devices may include a smartphone, a smartwatch, a VR headset, and cameras disposed within the vicinity of the user. Further, the processor 104 may be configured to determine a plurality of data features based on the received data and the real-world gesture. In an embodiment, the plurality of data features may include a visual user emotion feature, a visual user attention feature, an inertial measurement unit (IMU) event prediction feature, an IMU motion feature, and an ambient event prediction feature.
Further, the processor 104 may be configured to categorize the plurality of data features into at least one user feature and at least one event feature. In some embodiments, the at least one user feature may be indicative of visual user emotion by analyzing visual data received via the plurality of devices in the vicinity of the user to detect whether an emotion of the user matches an expected emotion for currently displayed content in the VR/AR environment. In an embodiment, the processor 104 may be configured to create a visual map of a field of view (FOV) of the user and track one or more objects within the surroundings of the user. Further, the processor 104 may be configured to determine the at least one user feature that is indicative of visual user attention based on the visual map to detect the attention of the user with respect to each of the one or more objects.
Further, the processor 104 may be configured to determine the at least one event feature indicative of a motion or an ambient event occurring in the vicinity of the user based on analysing data associated with one or more IMU sensors located within the VR headset, and one or more ambient sensors, audio, and multimedia communication from the plurality of devices in the vicinity of the user.
In an embodiment, the processor 104 may be configured to collect data related to the one or more events outside and inside the FOV of the user by the plurality of devices outside the FOV of the user. In an embodiment, the processor 104 may be configured to determine the at least one event feature indicating a probable ambient event occurring in an area outside the FOV of the user based on analysis of the collected data. In an embodiment, the processor 104 may be configured to track changes in user facial expressions and eye gaze for object tracking within the FOV of the user. In an embodiment, the processor 104 may be configured to collect data from the one or more ambient sensors to anticipate events outside the FOV of the user.
Further, the processor 104 may be configured to estimate an intent probability of the gesture based on an analysis of the gesture and the categorized plurality of data features. In an embodiment, the processor 104 may be configured to identify that the real-world gesture is intentional for a VR/AR environment 602, as shown in Figure 6, based on determining that the intent probability is above a predefined threshold. In an embodiment, the processor 104 may be configured to identify that the gesture is unintentional when the intent probability is below the predefined threshold. Further, the processor 104 may be configured to continue with the scene of the AR/VR environment based on the identification that the gesture is unintentional.
FIG. 2 illustrates a block diagram depicting modules of the system 100, according to one or more embodiments of the present disclosure.
The system 100 may include the following modules: an input device data observer 202, a smart data processing engine 204, a hybrid semantic fusion unit 206, an intelligent unintentional event estimator 208, and a VR scene gesture executor 210.
Further, the input device data observer 202 may be configured to fetch data relative to the user and surroundings. All the devices in vicinity of the user of the system (200) may be used to fetch data. For example, the input device data observer 202 may be configured to fetch data using mobile devices, watch or wearable devices, and/or Internet-of-Things (IoT) devices in vicinity of the user (e.g. a user 212) of the system (200). For example, AR/VR headset may be configured to capture information of surroundings of the user.
Further, the smart data processing engine 204 may be adapted to receive the real-world gesture. The real-world gesture may be received from a user 212. The user 212 may be a VR user a VR/AR user. Further, the smart data processing engine 204 may be adapted to process the received real-world gesture into a plurality of data features. For example, the smart data processing engine 204 may fetch all raw data observed from input devices. The raw data may be obtained by the input devices data observer 202. The smart data processing engine 204 may learn and analyse the raw data. The smart data processing engine 204 may process the raw data into one or more key data features.
Further, the hybrid semantic fusion unit 206 may be adapted to categorise the plurality of data features. The hybrid semantic fusion unit 206 may fuse the plurality of data features into the at least one user feature and the at least one event feature, as mentioned above. For example, the hybrid semantic fusion unit 206 may intelligently fuse received data at any particular time. The hybrid semantic fusion unit 206 may receive information related to any gesture performed at any time. The hybrid semantic fusion unit 206 may fuse received information related to gestures performed and user-generated (or VR/AR system-generated) pose data in real time.
Further, the intelligent unintentional event estimator 208 may be adapted to estimate the intent probability of the gesture based on the analysis of the gesture and the at least one user feature and the at least one event feature. For example, the intelligent unintentional event estimator 208 may employ one or more machine learning model configured to mark a gesture to be intentional or unintentional. The intelligent unintentional event estimator 208 may evaluate the gesture using one or more machine learning models. The intelligent unintentional event estimator 208 may mark the gesture to be intentional or unintentional.
Further, the VR scene gesture executor 210 may dynamically adjust the VR/AR environment based on user pose modifications. The VR scene gesture executor 210 may provide valid intentional gestures to the user 212 or the VR/AR headset. For example, the VR Scene gesture executor 210 may transmit valid intentional gestures to the VR/AR headset.
In one case, the intelligent unintentional event estimator 208 may determine that the intent probability is above or equal to a predefined threshold (or a first threshold). In this case, the real-world gesture is intentional for the VR/AR environment. In this case, the VR scene gesture executor 210 may adjust the scene of the AR/VR environment. In an embodiment, the predefined threshold may be 40%, but not limited thereto. In one case, the intelligent unintentional event estimator 208 may determine that the intent probability is below the predefined threshold. In this case, the gesture is unintentional. In this case, the VR scene gesture executor 210 may continue with the scene within the VR/AR environment.
FIG. 3a illustrates an exemplary scenario of the system 100 depicting the user 212 immersed in the VR/AR environment, according to one or more embodiments of the present disclosure. FIG. 3b illustrates a plurality of devices 302 communicatively coupled to the user 212 immersed in the VR environment, according to one or more embodiments of the present disclosure. Referring to Fig. 3b, while three devices 302a, 302b, and 302c are depicted, it may be apparent that only one or more than one device may be deployed within the scope of the disclosure.
The plurality of devices 302 may include a phone 302a, a smartwatch (or a wearable device) 302b, an AR/VR headset 302c, etc. In an embodiment, the AR/VR headset 302c may be referred to as VR headset 302c and may be used interchangeably throughout the detailed description. In an embodiment, the plurality of devices 302 may be communicatively coupled with the processor 104 via the network 120. In an embodiment, the input device data observer 202 may be adapted to monitor the plurality of devices 302 and capture (or fetch) raw data from the plurality of devices 302. In an exemplary embodiment, the raw data may include a phone data, a wearable data, and/or a VR/AR headset data. For example, the phone data may include but not limited to audio, location, call/message, alarm, and emergency reminder. The wearable data may include but not limited to heart rate, motion tracking, skin temperature, sleep cycle, and proximity. The VR/AR headset data may include IMU (e.g., sensor information such as information obtained from a gyroscope sensor and/or an accelerometer sensor), face and face tracker unit (e.g., facial expressions and/or eye focus), object tracking in user FOV, and/or head orientations. In an embodiment, the plurality of devices 302 may be referred to as user connected devices.
In an embodiment, an AR/VR system may detect a real-world gesture of a user. At time t, the user in an AR/VR world may use the AR/VR system. At time t, the AR/VR system may detect the real-world gesture of the user at time t+1. The AR/VR system may obtain data from input devices around the user. The AR/VR system may process the obtained data.
FIG. 4 illustrates a block diagram of the smart data processing engine 204 of the system 100, according to one or more embodiments of the present disclosure.
The smart data processing engine 204 may include a visual semantic gathering unit 402, an IMU attribute tracker 404, and an ambient information converging unit 406 communicatively coupled to the smart data processing engine 204. The visual semantic gathering unit 402 may be adapted to track visual based data using a camera. Further, the visual semantic gathering unit 402 may be adapted to create the visual map of the FOV and track events within the FOV of the user 212. Further, the visual semantic gathering unit 402 may be adapted to capture the visual scene data of the user 212.
The visual semantic gathering unit 402 may be adapted to detect gesture by various mechanisms, as illustrated in FIG. 8. The user 212 wearing the VR headset 302c may perform the gesture, for example by hand. Further, the visual semantic gathering unit 402 may be adapted to extract gesture image and a Red-Green-Blue (RGB) image of the gesture image is captured. Further, a heat map may be generated and a gesture type may be predicted. The gesture type may be a feature in the AR/VR environment. The visual semantic gathering unit 402 may be adapted to analyse user visual information and detect changes in facial expressions and eye gaze object tracking of the user 212. The visual semantic gathering unit 402 may be adapted for the VR user 212 out of synchronization emotion detection and FOV scene user attention detection, as illustrated in FIG. 12. The VR user 212 out of synchronization emotion detection of the visual semantic gathering unit 402 may be adapted to match expected emotion as per the VR scene and emotion at any time (t). In an embodiment, the VR scene may be referred to as a VR/AR scene and may be used interchangeably throughout the detailed description. Further, the VR user 212 out of synchronization emotion detection may be adapted to predict whether the emotion might be due to real-world impact. In an embodiment, the VR user 212 may express emotion as per VR content/scene and may express emotion out of synchronization as per the VR content. The FOV scene user attention detection may be adapted to create the visual map of the scene and track all objects in the vicinity of the VR user 212. Further, the FOV scene user attention detection may be adapted to calculate the attention of the VR user 212 within an object. The visual semantic gathering unit 402 will be described in greater detail in conjunction with FIGS. 9-11 in the later part of the detailed description.
Further, the IMU attribute tracker 404 may be adapted to capture sensor based data for the user 212 from the VR/AR headset 302c and the connected devices. In an example, the IMU sensors may include a gyroscope, an accelerometer, a position sensor, etc. In an embodiment, the sensor based data may include an angular velocity, orientation, centre of mass of location, acceleration, and velocity. Further, the IMU attribute tracker 404 may be adapted to integrate data based on time. In an embodiment, the IMU sensor information may be monitored. The sudden movement and relative motion probability may be detected, for example, using Kalman Filter and artificial intelligence (AI) expert system. The IMU attribute tracker 404 may have a monitoring service running that constantly monitors all the parameters. In an embodiment, the IMU attribute tracker 404 may be adapted to generate output features based on the sensor based data, including an IMU event prediction feature and an IMU motion feature. The IMU event prediction feature may depict movement at any given time (t) from rest or constant motion position. Further, the IMU motion feature may depict relative motion probability at any given time (t).
In an embodiment, the IMU sensors may comprise a gyroscope sensor, an accelerometer sensor, and/or a position sensor. The sensor data obtained from the IMU sensors may be calibrated. Calibrated sensor data may be applied to the Kalman Filter. For example, calibrated sensor data obtained from the gyroscope sensor may be applied to an orientation Kalman filter. Filtered results from the orientation Kalman filter may comprise angular velocity information and orientation information. For example, calibrated sensor data obtained from the accelerometer sensor may be applied to a position Kalman filter. Calibrated sensor data obtained from the position sensor may be applied to the position Kalman filter. The orientation information from the orientation Kalman filter may be applied to the position Kalman filter. Filtered results from the position Kalman filter may comprise centre of mass location information and acceleration information. Time derivative may be also applied to the Calibrated sensor data obtained from the position sensor. The results of the time derivative may comprise velocity information. The results of the orientation Kalman Filter, the position Kalman filter, and the time derivative may be applied to the AI expert system. The AI expert system may comprise an expert knowledge database, rules for steady state detection, and rules for relative motion detections. Based on the database and the rules included in the AI expert system. the AI expert system may generate one or more IMU event prediction features and one or more IMU motion features.
Further, the ambient information converging unit 406 may be adapted to capture and assemble different real-world ambient parameters. In an embodiment, the real-world ambient parameters may include audio, phone information, and the one or more parameters obtained from ambient sensors. In an example, the audio may comprise from a phone ringing, a voice of mother, mother shouting from another room. The phone information may comprise an urgent call situation, dad calling, etc. The one or more parameters obtained from ambient sensors may comprise a phone located near the user. Based on the real-world ambient parameters an ambient activity may be detected.
Further, the ambient information converging unit 406 may be adapted to gather data related to events outside/inside FOV of the user 212, such as audio, relative motion, and visual scene data using other devices outside FOV of the user 212. Based on the gathered data the ambient information converging unit 406 may be adapted to track probable events to happen in the area outside FOV of the user 212.
Further, the ambient information converging unit 406 may be adapted to feed the sensor-based data and the real-world ambient parameters to an ambient information collector. Further, the ambient information converging unit 406 may generate a predicted output, such as an ambient activity feature. In an embodiment, the ambient information converging unit 406 may be adapted to generate a pre-trained dataset for a supervised machine learning engine (for example, a support vector machine (SVM) ML) model 702, as shown in FIG. 7. Further, the SVM ML model 702 may be adapted to learn about the information being captured from the ambient surroundings, i.e., the real-world ambient parameters. In an embodiment, an activity happening outside the FOV of the user 212 may be detected. In an embodiment, the SVM ML model 702 may be adapted to be fed with input parameters, for example, the ambient sensors, audio data, and multimedia/communication on a user device. The ambient sensor may be the IMU sensor and the input parameters from the ambient sensors may include an ultrasound sensor to locate a person, a vibration sensor to detect a sitting or standing state, infrared sensor to detect entry, a microwave sensor to detect small movements, and/or thermal sensor to detect physical activity. The input parameters from the audio data may be employed to identify a speaker, deduce an emergency content, and/or detect queries to the VR user 212. The input parameters from the multimedia may be utilized to detect urgent calls/messages, and categorize urgent multimedia. Further, the SVM ML model 702 may be adapted to generate output parameters based on the received/fed input parameters. The output parameters may include the IMU event prediction feature to depict the probability of distraction of the VR user 212.
In an embodiment, the AV/VR system may determine one or more IMU prediction event features and one or more IMU motion features. The AV/VR system may determine an IMU prediction event feature and an IMU motion feature at time t+1 based on head orientation data, eye focus changing data, acceleration data, and/or a heart rate data, etc. For example, at time t+1, the IMU prediction event feature (or its value) may be determined as '0.7' and the IMU motion feature (or its value) may be determined as '0'.
In an embodiment, the AV/VR system may determine one or more ambient event prediction features. The AV/VR system may determine an ambient event prediction feature at time t+1 based on head orientation data, eye focus changing data, etc. For example, at time t+1, the ambient event prediction feature (or its value) may be determined as 'event: emergency, loud object damage, a weight of 0.7'.
FIG. 5 illustrates a block diagram 500 of the hybrid semantic fusion unit 206 of the system 100, according to one or more embodiments of the present disclosure.
The hybrid semantic fusion unit 206 may also be referred to as a hybrid content synthesizer. The hybrid semantic fusion unit 206 may be adapted to receive the plurality of data features. The plurality of data features may be referred to as processed data features. The plurality of data features may comprise, for example, the visual user emotion feature, the visual user attention feature, the inertial measurement unit (IMU) event prediction feature, the IMU motion feature, and/or the ambient event prediction feature, as mentioned earlier. Further, the hybrid semantic fusion unit 206 may be adapted to categorize the plurality of data features into at least one user feature 502 and at least one event feature 504, as discussed earlier. The at least one user feature 502 and at least one event feature 504 may be fed to the intelligent unintentional event estimator 208. The intelligent unintentional event estimator 208 is described in conjunction with FIG. 6.
The hybrid semantic fusion unit 206 may receive inputs from the smart data processing engine 204, for example, the visual semantic gathering unit 402, the IMU attribute tracker 404, the ambient information converging unit 406, and fuse the received inputs into broader categories, the at least one user feature 502 and the at least one event feature 504, for eliminating redundant data. In an embodiment, the redundant data may be eliminated by considering features from multiple input sources, e.g., same features may be considered in visual, sensor, and ambient sources. Further, during classification or categorization, any feature may only be considered once, thereby eliminating redundant data. In an example, similar attributes from all different data processing units are merged for better score data.
In an embodiment, a user feature may be determined as a combination of emotion feature and an attention feature. In an embodiment, an event feature may be determined as a combination of a user event, an ambient event, and a relative motion.
In an embodiment, the hybrid semantic fusion unit 206 may receive a first visual user emotion feature determined as 'shock, 3', a first visual user attention feature determined as 'loss in attention', a first IMU even prediction feature determined as '0.7', a first IMU motion feature determined as '0', and a first ambient event prediction feature determined as 'event: emergency, loud damage, a weight of 0.7'. The hybrid semantic fusion unit 206 may calculate a first user feature based on the first visual user emotion feature and the first visual user attention feature. For example, based on the strength (or the value) of the first visual user emotion feature (i.e., 3) being greater than a first threshold and based on the first visual user attention feature determined as 'loss in attention', the first user feature may be determined as (1, 1). The hybrid semantic fusion unit 206 may calculate a first event feature based on the first IMU event prediction feature, the first IMU motion feature, and the first ambient event prediction feature. For example, based on the strength (or the value) of first IMU event prediction feature (i.e., 0.7) being greater than a second threshold, based on the weight of the first ambient event prediction feature being greater than a third threshold, and based on the IMU motion feature being less than a fourth threshold, the first event feature may be determined as (1, 0, 1).
FIG. 6 illustrates a block diagram 600 of the intelligent unintentional event estimator 208 operationally coupled with a VR scene correction unit 210 of the system 100, according to one or more embodiments of the present disclosure.
The intelligent unintentional event estimator 208 may be adapted to estimate the intent probability of the gesture based on an analysis of the gesture and the at least one user feature 502 and the at least one event feature 504. In an example embodiment, the intelligent unintentional event estimator 208 may estimate that an intention probability is 30% and a relative motion probability is 55%. The intelligent unintentional event estimator 208 may identify that the real-world gesture is intentional for the VR/AR environment 602 or the VR scene 602 based on determining that the intent probability is above the predefined threshold, as discussed above. Further, the intelligent unintentional event estimator 208, based on the predefined threshold of the intent probability, the VR scene correction unit 210 may be adapted to adjust the VR scene 602 for the VR/AR environment, as mentioned earlier. Further, the VR scene correction unit 210 may include a gesture analyzer unit 604, a final user pose evaluator unit 606 and a VR scene updater unit 608. The gesture analyzer unit 604 may be adapted to atomically analyze gesture input and prepare to update the event to the VR scene 602. Further, the final user pose evaluator unit 606 may be adapted to finalize the relative pose of the user 212 based on different calculated attributes. Further, the VR scene updater unit 608 may be adapted to update the VR scene for both inputs relative pose and gesture applicability.
In an embodiment, based on the determined intention probability of 30% and on the determined relative motion probability of 55%, the intelligent unintentional event estimator 208 may determine that (a) the gesture of the user at corresponding time is unintentional and should be rejected and (b)the motion of the user at corresponding time is not affected. The VR scene updater unit 608 may adjust based on the determination (a) and (b). For example, based on that the gesture is determined as unintentional and the motion is determined as not affected, the VR scene updater unit 608 may continue a current VR scene of VR system.
FIG. 7 illustrates a flowchart 700 depicting an operational flow of the system 100, according to one or more embodiments of the present disclosure.
As mentioned in the preceding paragraphs, the system 100 may include the following modules such as the input device data observer 202, the smart data processing engine 204, the hybrid semantic fusion unit 206, the intelligent unintentional event estimator 208, and the VR scene gesture executor 210 or the VR scene correction unit 210. The intelligent unintentional event estimator 208 may be referred to as the SVM ML model 702 as described above. The operation of each module or element of the system 100 depicted by the flowchart is already explained above in greater detail in conjunction with Figure 1-6.
FIG. 8 illustrates an exemplary scenario 800 depicting different mechanisms to perform gesture detection when the user 212 is immersed in the VR scene 602, according to one or more embodiments of the present disclosure.
As discussed above, objects present with the FOV of the user 212 may be tracked and analyzed using the camera installed within the VR headset 302c. In an embodiment, the IMU sensors may be inbuilt within the VR headset 302c. Further, the objects present outside the FOV of the user 212, that is, ambient information, may be collected using a camera arrangement and the phone, watch content, etc. The user 212 immersed within the VR scene 602 may perform the gesture and the performed gesture may be analyzed and based on the gesture performed, whether intentional or unintentional, the feature in the AR/VR environment may be detected. For example, Alex wearing the VR headset 302c and immersed in the VR scene performs a hand gesture like a gun, and the visual semantic gathering unit 402 of the smart data processing engine 204 detects a heat map of the hand and predicts the type of gesture based on the heat map.
FIG. 9 illustrates an exemplary scenario 900 depicting the user 212 out-of-sync emotion detection by the visual semantic gathering unit 402, according to one or more embodiments of the present disclosure.
The visual semantic gathering unit 402 may be adapted to determine the out-of-synchronization emotion detection of the VR user 212. At 902, the VR user 212 may be immersed in the VR scene. In an embodiment, the VR user 212, based on the VR scene may perform an emotion. The emotion may be categorized as a real emotion and an expected emotion. At 904, the real emotion of the VR/AR user 212 may be detected and at 906, the expected emotion as per content by the VR/AR user may be detected. The real emotion may be categorized as an emotion 1 model and the expected emotion may be categorized as an emotion 2 model. At 908, an emotion contradiction calculator may be adapted to calculate emotion deviation strength of the emotion 1 model and the emotion 2 model. The emotion contradiction calculator may take input of emotion features from the emotion 1 model and the emotion 2 model. Further, the emotion contradiction calculator may determine whether the emotion deviation strength exists. In one case, the emotion contradiction calculator may determine that the emotion deviation strength exists. In this case, the emotion contradiction calculator may calculate the deviation strength value. For example, E1=4, E2=1, therefore, dE=3 wherein dE is a deviation between the E1 and the E2. At 910, a visual user emotion feature may be detected. The real emotion and the expected emotion of the user 212 as per the content within the VR scene may be further explained in conjunction with FIGS. 10-11.
In an embodiment, The AR/VR system may determine one or more visual user emotion features. The AR/VR system may determine a areal emotion of the user and a AR/VR content emotion (e.g., one or more expected emotions as per AR/VR content). The AR/VR system may match the expected emotion for the AR/VR content and the real emotion of the user at the time t+1. The AR/VR system may compare a strength (or a probability or a deviation between the expected emotion and the real emotion) of each emotion with a predefined threshold. Based on the comparison, the AR/VR system may determine whether a visual user emotion feature is intentional or unintentional. For example, if a determined strength of an emotion 'shocking' is greater than or equal to the predefined threshold, the emotion 'shocking' is determined as intentional (and/or the visual user emotion feature corresponding to the emotion 'shocking' is determined as intentional).
FIG. 10 illustrates an exemplary scenario 1000 of training and verifying a convolutional neural network (CNN) model depicting the user's real emotion detection by the visual semantic gathering unit 402, according to one or more embodiments of the present disclosure.
The CNN model may be adapted to detect the real emotion (emotion 1 model) of the VR user 212. The exemplary scenario 1000 may include a training stage 1002 and a verification state 1004. The training stage 1002 may be adapted to train a deep CNN model 1006 using a dataset, e.g., an image dataset to generate a pre-trained model 1008. Further, the pre-trained model 1008 may be adapted to feed with emotion recognition adaption to form the pre-trained model 1010 with a new dense layer. Further, the pre-trained model 1010 may be fine-tuned using a cleansed dataset 1012 from a facial expression dataset 1014. In an embodiment, the facial expression dataset 1014 may include multiple facial expressions. The cleansed dataset 1012 may be obtained from the facial expression dataset 1014 by cropping facial parts of expressions. Further, by fine-tuning the pre-trained model 1010, an emotion recognition model 1016 may be generated (or obtained).
Further, the emotion recognition model 1016 may be verified or tested at the verification stage 1006. The verification stage 1016 may be performed by feeding a cropped picture of a face to the emotion recognition model 1016. The output of the emotion recognition model 1016 may be estimated based on corresponding gesture of the face. For example, the corresponding gesture of the face may be predefined for the verification. Further, the emotion recognition model 1016 may be adapted to predict output probability of each emotion. Therefore, the real emotion of the user 212 may be detected from the emotion recognition model 1016.
FIG. 11 illustrates an exemplary scenario 1100 depicting user expected emotion as per content by the visual semantic gathering unit 402, according to one or more embodiments of the present disclosure.
The expected emotion of the user 212 as per content may be determined by training a single task CNN model from the VR scene and by one or more feature extractors from an audio of the VR scene and surroundings. In an embodiment, the one or more feature extractors may be adapted to extract features from the VR/AR content. In an embodiment, a clip from the VR scene 602 may be input to the pre-trained single task CNN model. The pre-trained single task CNN model may extract one or more features from the clip. In an embodiment, the feature extractor may comprise a 3x3 convolution layer and a max pooling unit (or a max pooling layer). In an embodiment, the audio may include a zeta function corresponding to singularities in space and time. In an embodiment, the zeta function may be generated from the audio by the system 100. In an embodiment, the zeta function may be input to the one or more feature extractors. The one or more feature extractors may extract corresponding feature(s) from the zeta function. Further, the one or more feature extractors and the single task CNN model may be configured to generate a feature fused layer. Further, the feature fused layer may include one or more fully convolutional layers to generate an output. The output may include a predicted user emotion.
FIG. 12 illustrates a flowchart of the visual semantic gathering unit 402 depicting the FOV scene user attention detection 1200, according to one or more embodiments of the present disclosure.
The FOV scene user attention detection 1200 may be determined by the visual semantic gathering unit 402. The VR user 212 attentional information of user 212 in real-world scenes may be identified by an eye-tracking paradigm. The eye-tracking paradigm may directly compare eye movements while the VR user 212 may be immersed in the VR/AR environment.
A meaning map of a scene may be created using a spatial distribution algorithm for high-level semantic features in an environment, e.g., faces, objects, doors, etc. In an embodiment, the meaning map may be relevant to understanding the semantic content and affordances available to the VR user 212 in the VR/AR scene. Further, a gaze map may be captured and compared with the meaning map to calculate attention parameters. The attention features may include an average fixation duration (AFD), an average fixation number (AFN), gaze shifts (GS), and a gaze spherical distribution (GSD). In an embodiment, the attention parameters may include an active attention and a passive attention. The active attention and the passive attention may be calculated by superimposing the meaning map into the gaze map. In an embodiment, the user's attention on semantically meaning regions may be analyzed as a whole for the VR scene. For example, for any scene for a given time (t) the attention parameters for the VR user 212 may be calculated.
In an embodiment, the AFD may be an average of user fixation duration across scenes. In an embodiment, the user fixation duration is inversely proportional to the attention, for example, less is the user fixation duration, and more is the attention. Further, the AFN may be an average of total fixation across scenes. For example, less is the AFN, more is the attention. Further, GS may be larger in the active attention. Further, the GSD may be less centrally trending in the active attention. In an embodiment, a change in the attention strength of user 212 for the FOV scene while performing the gesture may need to be considered to validate the intention of the gesture. In an example, for loss in attention, AFD2>AFD1, AFN2>AFN1, GS2<GS1, and GSD2<GSD1. According to this example, when any three of the conditions are met, the user attention may change negatively from the previous instance.
In an embodiment, The AR/VR system may determine one or more visual user attention features. For example, the AR/VR system may calculate corresponding attention of the user with each object. Accordingly, the AR/VR system may determine that the user attention has been changed regarding a first object and a second object. In this case, a count of visual user attention features may be calculated as 2; that is, the count of visual user attention features may correspond to a number of objects of which the user attention has been changed.
FIG. 13 illustrates an exemplary table depicting an event feature calculation 1300 and a user feature calculation 1302 by the hybrid semantic fusion unit 206, according to one or more embodiments of the present disclosure.
The event feature calculation 1300 may be segregated into three parts, for example, the IMU event prediction feature, the ambient event prediction feature and the IMU motion feature. The event feature calculation 1300 may be evaluated as the at least one event feature with features as, a user event, an ambient event, and a relative motion. In an embodiment, the IMU event prediction feature may generate the user event, the ambient event prediction feature may generate the ambient event, and the IMU motion feature may generate the relative motion.
The user feature calculation 1302 may be grouped into two parts, for example, a visual user emotion feature, and a visual user attention feature. The user feature calculation 1302 may determine the at least one user feature with features such as the emotion and the attention. In one example, if a visual emotion feature (VEF) is greater than or equal to the predefined threshold (or a first threshold) (e.g., 0.5), then the emotion may be intentional. On the other hand, if VEF is less than 0.5, the emotion may be unintentional. In an example, if a count (or a number) of user attention feature(s) (UAF) is greater than or equal to another predetermined threshold (e.g., 2), the attention may be intentional. On the other hand, if UAF is less than 2, the attention may be unintentional. In an embodiment, the predefined threshold of the VEF is directly dependent on expressions of the user and may vary among users based on expressions.
FIG. 14 illustrates a table having a plurality of features as a training data for the intelligent unintentional event estimator 208, according to one or more embodiments of the present disclosure.
The intelligent unintentional event estimator 208 may correspond to the SVM ML model 702 that may be adapted to determine a multi-layer perceptron model (MLP) by training on a dataset of gestures, user features, event features, user intentions and a relative motion. In an embodiment, the features of all feeder machine learning engines may be both static and/or dynamic. In an embodiment, the fused feature scores of all machine learning engines are fed to the intelligent unintentional event estimator 208. For example, the VR user 212 may perform a left hand gesture, a user feature 1 and an event feature 3 may be detected based on the left hand gesture and the user intention may be intentional and the relative motion may be positive. These feature scores of the left hand gesture may be used as training data for training the MLP. In an embodiment, for every gesture, on the basis of feature scores, content is passed as training data to a supervised machine learning engine, i.e., MLP.
FIG. 15 illustrates a flowchart depicting a method of identifying an intentional gesture, according to one or more embodiments of the present disclosure.
At step 1502, the method 1500 may include detecting a real-world gesture performed by the user 212 immersed in a virtual reality (VR) environment and/or an augmented reality (AR) environment. At step 1504, the method 1500 may include receiving data associated with the user 212 and corresponding surroundings. At step 1506, the method 1500 may include determining a plurality of data features based on the received data and the real-world gesture. Further, at step 1508, the method 1500 may include categorizing the plurality of data features into at least one user feature and at least one event feature.
Alternatively or additionally, the method 1500 may include determining the at least one user feature indicative of visual user emotion by analyzing visual data received via one or more devices in the vicinity of the user 212 to detect whether an emotion of the user matches an expected emotion for currently displayed content in the VR/AR environment. For example, the method 1500 may include analyzing visual data received via one or more devices in the vicinity of the user 212, detecting whether an emotion of the user matches an expected emotion for currently displayed content in the VR/AR environment based on the analyzed visual data, and determining a first user feature indicative of visual user emotion based on detecting whether the emotion of the user matches the expected emotion. Alternatively or additionally, The method 1500 may include creating the visual map of the FOV of the user 212 and tracking one or more objects within the surroundings of the user 212. The method 1500 may include determining the at least one user feature indicative of visual user attention based on the visual map to detect attention of the user with respect to each of the one or more objects. For example, the method 1500 may include detecting an attention of the user with respect to each of the one or more objects and determining a second user feature indicative of visual user attention based on the visual map and the detected attention of the user.
Alternatively or additionally, the method 1500 may include analysing data associated with at least one of: one or more inertial measurement unit (IMU) sensors located within a VR headset, one or more ambient sensors, audio, or multimedia communication from one or more devices in vicinity of the user; and determining the at least one event feature indicative of a motion or an ambient event occurring in vicinity of the user based on the analyzed data.
Alternatively or additionally, the method 1500 may include collecting data related to the one or more events outside and inside a field of view (FOV) of the user by one or more devices outside the FOV of the user; and determining the at least one event feature indicating a probable ambient event occurring in an area outside the FOV of the user based on the collected data. Alternatively or additionally, the method 1500 may include analyzing the collected data and determining the at least one event feature indicating the probable ambient event based on the analyzed collected data.
At step 1510, the method 1500 may include estimating an intent probability of the gesture based on an analysis of the gesture and the categorized plurality of data features. At step 1512, the method 1500 may include identifying that the real-world gesture is intentional for the VR/AR environment based on the intent probability being above the predefined (or a first) threshold. Alternatively or additionally, the method 1500 may include comparing the intent probability with the first threshold. In an embodiment, the method 1500 may include adjusting a scene within the VR/AR environment based on identifying whether the real-world gesture is intentional. Alternatively or additionally, the method 1500 may include identifying that the gesture is unintentional based on the intent probability being below the predefined threshold (or the first threshold). Alternatively or additionally, the method 1500 may include continuing with a scene of the AR/VR environment based on identifying that the gesture is unintentional.
The present disclosure may be adapted to ensure a smooth streamlined experience of the AR/VR environment for the user. The present disclosure may be adapted to identify gestures performed by the user while he/she is immersed in the VR/AR environment. The present disclosure classifies them into different categories and estimates a probability that a certain gesture performed by the user is intentionally done by the user or is unintentional. In an embodiment, a certain gesture is an act by the user or any other act/event happening within the vicinity of the user affecting the user. In this manner, the user experiences an AR/VR experience without any disturbance and experiences enhanced or controlled VR scenes. The user is not affected by any factor outside the FOV of the user which may cause unwanted user gesture.
Although specific units/modules have been illustrated in the figure and described above, it should be understood that the system may include other hardware modules or software modules or combinations as may be required for performing various functions.
The various embodiments described above are provided by way of illustration only and should not be construed to limit the scope of the disclosure. Various modifications and changes may be made to the principles described herein without following the example embodiments and applications illustrated and described herein, and without departing from the spirit and scope of the disclosure.
Those skilled in the art will appreciate that the operations described herein in the present disclosure may be carried out in other specific ways than those set forth herein without departing from essential characteristics of the present disclosure. The above-described embodiments are therefore to be construed in all aspects as illustrative and not restrictive. The scope of the disclosure should be determined by the appended claims, not by the above description, and all changes coming within the meaning of the appended claims are intended to be embraced therein.
The drawings and the forgoing description give examples of embodiments. Those skilled in the art will appreciate that one or more of the described elements may well be combined into a single functional element. Alternatively, certain elements may be split into multiple functional elements. Elements from one embodiment may be added to another embodiment. For example, orders of processes described herein may be changed and are not limited to the manner described herein.
Moreover, the actions of any flow diagram need not be implemented in the order shown; nor do all of the acts necessarily need to be performed. Also, those acts that are not dependent on other acts may be performed in parallel with the other acts. The scope of embodiments is by no means limited by these specific examples. Numerous variations, whether explicitly given in the specification or not, such as differences in structure, dimension, and use of material, are possible. The scope of embodiments is at least as broad as given by the following claims.
Benefits, other advantages, and solutions to problems have been described above with regard to specific embodiments. However, the benefits, advantages, solutions to problems, and any component(s) that may cause any benefit, advantage, or solution to occur or become more pronounced are not to be construed as a critical, required, or essential feature or component of any or all the claims.

Claims (15)

  1. A method (1500) of identifying an intentional gesture, the method (1500) comprising:
    detecting (1502) a real-world gesture performed by a user immersed in a virtual reality or an augmented reality (VR/AR) environment;
    receiving (1504) data associated with the user and corresponding surroundings;
    determining (1506) a plurality of data features based on the received data and the real-world gesture;
    categorizing (1508) the plurality of data features into at least one user feature and at least one event feature;
    estimating (1510) an intent probability of the real-world gesture based on the real-world gesture and the categorized plurality of data features; and
    identifying (1512) that the real-world gesture is intentional for the VR/AR environment based on the intent probability being above a first threshold.
  2. The method (1500) as claimed in claim 1 further comprising:
    adjusting a scene within the VR/AR environment based on identifying whether the real-world gesture is intentional.
  3. The method (1500) as claimed in claim 1 or 2 further comprising:
    identifying that the real-world gesture is unintentional based on the intent probability being below the first threshold; and
    continuing with a scene within the AR/VR environment based on identifying that the real-world gesture is unintentional.
  4. The method (1500) as claimed in any one of claims 1 to 3 further comprising:
    estimating a relative motion probability based on the real-world gesture and the categorized plurality of data features; and
    identifying that the relative motion is intentional for the VR/AR environment based on the relative motion probability being above a second threshold.
  5. The method (1500) as claimed in claim 4 further comprising:
    adjusting a scene within the VR/AR environment based on identifying whether the real-world gesture is intentional and based on identifying whether the relative motion is intentional for the VR/AR environment.
  6. The method (1500) as claimed in any one of claims 1 to 5, wherein the data corresponds to one or more of visual data, inertial measurement unit (IMU) sensors-based data, ambient sensors-based data, audio data, multimedia communication data received from one or more devices in vicinity of the user.
  7. The method (1500) as claimed in any one of claims 1 to 6, wherein determining the plurality of data features comprises:
    analyzing visual data received via one or more devices in vicinity of the user;
    detecting whether an emotion of the user matches an expected emotion for currently displayed content in the VR/AR environment based on the analyzed visual data;
    determining a first user feature indicative of visual user emotion based on detecting whether the emotion of the user matches the expected emotion;
    creating the visual map of a field of view (FOV) of the user;
    tracking one or more objects within the surroundings of the user;
    detecting an attention of the user with respect to each of the one or more objects; and
    determining a second user feature indicative of visual user attention based on the visual map and the detected attention of the user.
  8. The method (1500) as claimed in any one of claims 1 to 7, wherein determining the plurality of data features comprises:
    analyzing data associated with at least one of: one or more inertial measurement unit (IMU) sensors located within a VR headset, one or more ambient sensors, audio, or multimedia communication from one or more devices in vicinity of the user; and
    determining a first event feature indicative of a motion or an ambient event occurring in vicinity of the user based on the analyzed data.
  9. The method (1500) as claimed in any one of claims 1 to 8, wherein determining the plurality of data features comprises:
    collecting data related to the one or more events outside and inside a field of view (FOV) of the user by one or more devices outside the FOV of the user; and
    determining the at least one event feature indicating a probable ambient event occurring in an area outside the FOV of the user based on the collected data.
  10. The method (1500) as claimed in any one of claims 1 to 9, wherein receiving the data comprises:
    tracking changes in user facial expressions and eye gaze for object tracking within a field of view (FOV) of the user.
  11. The method (1500) as claimed in claim 1, wherein determining the plurality of data features comprises:
    collecting data from one or more ambient sensors to anticipate events outside a field of view (FOV) of the user.
  12. A computer-readable storage medium storing one or more instructions, wherein the one or more instructions, when executed by at least one processor, cause the at least one processor to perform the method of any one of claims 1 to 11.
  13. A system (100) of identifying an intentional gesture, the system (100) comprising:
    at least one processor (104) comprising processing circuitry;
    memory (102) comprising one or more storage media, the memory communicatively coupled to the processor (104), and configured to store instructions which, when executed by the at least one processor (104) individually or collectively, causes the system (100) to:
    detect a real-world gesture performed by a user immersed in a virtual reality (VR) or an augmented reality (VR/AR) environment;
    receive data associated with the user and corresponding surroundings;
    determine a plurality of data features based on the received data and the real-world gesture;
    categorize the plurality of data features into at least one user feature and at least one event feature;
    estimate an intent probability of the real-world gesture based on the real-world gesture and the categorized plurality of data features; and
    identify that the real-world gesture is intentional for the VR/AR environment based on the intent probability being above a first threshold.
  14. The system (100) as claimed in claim 13, wherein the instructions are further configured to, when executed by the at least one processor (104) individually or collectively, cause the system (100) to:
    identify that the real-world gesture is unintentional based on the intent probability being below the first threshold; and
    continue with a scene of the AR/VR environment based on identifying that the real-world gesture is unintentional.
  15. The system (100) as claimed in claim 13 or 14, wherein the instructions are further configured to, when executed by the at least one processor (104) individually or collectively, cause the system (100) to:
    estimate a relative motion probability based on the real-world gesture and the categorized plurality of data features; and
    identify that the relative motion is intentional for the VR/AR environment based on the relative motion probability being above a second threshold.
PCT/KR2024/011466 2024-06-28 2024-08-05 Methods and systems for identifying intentional and unintentional user gesture Pending WO2026005102A1 (en)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
IN202411049554 2024-06-28
IN202411049554 2024-06-28

Publications (1)

Publication Number Publication Date
WO2026005102A1 true WO2026005102A1 (en) 2026-01-02

Family

ID=98222181

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/KR2024/011466 Pending WO2026005102A1 (en) 2024-06-28 2024-08-05 Methods and systems for identifying intentional and unintentional user gesture

Country Status (1)

Country Link
WO (1) WO2026005102A1 (en)

Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20190188895A1 (en) * 2017-12-14 2019-06-20 Magic Leap, Inc. Contextual-based rendering of virtual avatars
US20210248835A1 (en) * 2015-12-11 2021-08-12 Google Llc Context sensitive user interface activation in an augmented and/or virtual reality environment
US20220342481A1 (en) * 2021-04-22 2022-10-27 Coapt Llc Biometric enabled virtual reality systems and methods for detecting user intentions and manipulating virtual avatar control based on user intentions for providing kinematic awareness in holographic space, two-dimensional (2d), or three-dimensional (3d) virtual space
US20230154175A1 (en) * 2018-04-20 2023-05-18 Meta Platforms Technologies, Llc Auto-completion for Gesture-input in Assistant Systems
US20240077986A1 (en) * 2022-09-07 2024-03-07 Meta Platforms Technologies, Llc Opportunistic adaptive tangible user interfaces for use in extended reality environments

Patent Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20210248835A1 (en) * 2015-12-11 2021-08-12 Google Llc Context sensitive user interface activation in an augmented and/or virtual reality environment
US20190188895A1 (en) * 2017-12-14 2019-06-20 Magic Leap, Inc. Contextual-based rendering of virtual avatars
US20230154175A1 (en) * 2018-04-20 2023-05-18 Meta Platforms Technologies, Llc Auto-completion for Gesture-input in Assistant Systems
US20220342481A1 (en) * 2021-04-22 2022-10-27 Coapt Llc Biometric enabled virtual reality systems and methods for detecting user intentions and manipulating virtual avatar control based on user intentions for providing kinematic awareness in holographic space, two-dimensional (2d), or three-dimensional (3d) virtual space
US20240077986A1 (en) * 2022-09-07 2024-03-07 Meta Platforms Technologies, Llc Opportunistic adaptive tangible user interfaces for use in extended reality environments

Similar Documents

Publication Publication Date Title
WO2018212494A1 (en) Method and device for identifying object
WO2021100919A1 (en) Method, program, and system for determining whether abnormal behavior occurs, on basis of behavior sequence
WO2018212617A1 (en) Method for providing 360-degree video and device for supporting the same
WO2019093599A1 (en) Apparatus for generating user interest information and method therefor
WO2022158819A1 (en) Method and electronic device for determining motion saliency and video playback style in video
WO2023120840A1 (en) Method for providing cognitive rehabilitation content to elderly by using virtual reality technology, and computing device using same
WO2023182796A1 (en) Artificial intelligence device for sensing defective products on basis of product images and method therefor
EP3635627A1 (en) System and method for optical tracking
WO2021107734A1 (en) Method and device for recommending golf-related contents, and non-transitory computer-readable recording medium
WO2026005102A1 (en) Methods and systems for identifying intentional and unintentional user gesture
WO2024048944A1 (en) Apparatus and method for detecting a user intent for image capturing or video recording
WO2022097805A1 (en) Method, device, and system for detecting abnormal event
WO2023163515A1 (en) Interactive display system for dogs, operating method thereof, and interactive display device for dogs
WO2023080667A1 (en) Surveillance camera wdr image processing through ai-based object recognition
WO2021251733A1 (en) Display device and control method therefor
WO2025230303A1 (en) Object tracking method and apparatus for tracking same object within store
WO2021040317A1 (en) Apparatus, method and computer program for determining configuration settings for a display apparatus
WO2021199311A1 (en) Monitoring device, monitoring method, and recording medium
WO2023200038A1 (en) Eye tracking system and method for monitoring golf hitting, and non-transitory computer-readable recording medium
WO2024010133A1 (en) Image noise learning server and image noise reduction device using machine learning
JP2024075108A (en) Information processing device, information processing method, and program
WO2021251780A1 (en) Systems and methods for live conversation using hearing devices
WO2022050622A1 (en) Display device and method of controlling same
JP2023074793A (en) Image processing system and image processing method
WO2024043752A1 (en) Method and electronic device for motion-based image enhancement

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 24944844

Country of ref document: EP

Kind code of ref document: A1