EP4662543A1 - Apparatus and method for guiding a viewing direction of a user, and user equipment - Google Patents
Apparatus and method for guiding a viewing direction of a user, and user equipmentInfo
- Publication number
- EP4662543A1 EP4662543A1 EP24703165.1A EP24703165A EP4662543A1 EP 4662543 A1 EP4662543 A1 EP 4662543A1 EP 24703165 A EP24703165 A EP 24703165A EP 4662543 A1 EP4662543 A1 EP 4662543A1
- Authority
- EP
- European Patent Office
- Prior art keywords
- user
- people
- scene
- data
- viewing
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F3/00—Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
- G06F3/01—Input arrangements or combined input and output arrangements for interaction between user and computer
- G06F3/011—Arrangements for interaction with the human body, e.g. for user immersion in virtual reality
- G06F3/013—Eye tracking input arrangements
-
- A—HUMAN NECESSITIES
- A61—MEDICAL OR VETERINARY SCIENCE; HYGIENE
- A61B—DIAGNOSIS; SURGERY; IDENTIFICATION
- A61B5/00—Measuring for diagnostic purposes; Identification of persons
- A61B5/16—Devices for psychotechnics; Testing reaction times ; Devices for evaluating the psychological state
- A61B5/163—Devices for psychotechnics; Testing reaction times ; Devices for evaluating the psychological state by tracking eye movement, gaze, or pupil change
-
- A—HUMAN NECESSITIES
- A61—MEDICAL OR VETERINARY SCIENCE; HYGIENE
- A61B—DIAGNOSIS; SURGERY; IDENTIFICATION
- A61B5/00—Measuring for diagnostic purposes; Identification of persons
- A61B5/16—Devices for psychotechnics; Testing reaction times ; Devices for evaluating the psychological state
- A61B5/165—Evaluating the state of mind, e.g. depression, anxiety
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F3/00—Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
- G06F3/01—Input arrangements or combined input and output arrangements for interaction between user and computer
- G06F3/011—Arrangements for interaction with the human body, e.g. for user immersion in virtual reality
- G06F3/015—Input arrangements based on nervous system activity detection, e.g. brain waves [EEG] detection, electromyograms [EMG] detection, electrodermal response detection
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V20/00—Scenes; Scene-specific elements
- G06V20/20—Scenes; Scene-specific elements in augmented reality scenes
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16H—HEALTHCARE INFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR THE HANDLING OR PROCESSING OF MEDICAL OR HEALTHCARE DATA
- G16H20/00—ICT specially adapted for therapies or health-improving plans, e.g. for handling prescriptions, for steering therapy or for monitoring patient compliance
- G16H20/70—ICT specially adapted for therapies or health-improving plans, e.g. for handling prescriptions, for steering therapy or for monitoring patient compliance relating to mental therapies, e.g. psychological therapy or autogenous training
Definitions
- the present disclosure relates to changing a user’s perception.
- examples of the present disclosure relate to an apparatus and a method for guiding a viewing direction of a user.
- Further examples of the present disclosure relate to a user equipment, a non-transitory machine-readable medium and a program.
- the present disclosure provides an apparatus for guiding a viewing direction of a user.
- the apparatus comprises interface circuitry configured to receive first data indicative of one or more physiological property of the user.
- the interface circuitry is further configured to receive second data indicative of viewing directions of a plurality of people.
- the plurality of people are interacting in the same scene as the user.
- the apparatus comprises processing circuitry configured to determine at least one viewing focus of the plurality of people in the scene based on the second data. Further, the processing circuitry is configured to determine whether the user experiences negative emotions based on the first data. If it is determined that the user experiences negative emotions, the processing circuitry is configured to cause an output for guiding the viewing direction of the user toward one of the at least one viewing focus of the plurality of people in the scene by a user equipment of the user.
- the present disclosure provides a user equipment for a user.
- the user equipment comprises one or more sensor configured to measure one or more physiological property of the user while the user is interacting in the same scene as a plurality of people.
- the user equipment comprises interface circuitry configured to send data indicative of the one or more physiological property of the user to an external device.
- the interface circuitry is additionally configured to receive, from the external device, control data encoded with one or more commands for controlling the user equipment to perform output for guiding the viewing direction of the user toward a determined viewing focus of the plurality of people in the scene.
- the user equipment comprises a humanmachine interface configured to perform output for guiding the viewing direction of the user toward the viewing focus based on the control data.
- the present disclosure provides a non-transitory machine- readable medium having stored thereon a program having a program code for performing the method according to the second aspect, when the program is executed on a processor or a programmable hardware.
- the present disclosure provides a program having a program code for performing the method according to the second aspect, when the program is executed on a processor or a programmable hardware.
- Fig. 1 illustrates an example of an apparatus for guiding a viewing direction of a user
- Fig. 2 illustrates a first exemplary scene
- Fig. 6 illustrates a flowchart of an example of a method for guiding a viewing direction of a user.
- Fig- 1 illustrates an exemplary apparatus 100 for guiding a viewing (view) direction of a user.
- the viewing direction of the user is the sight direction of the user along which the user is viewing to perceive his/her environment.
- the viewing direction may be a spatial direction in the real world, a virtual world, or a mixture thereof.
- the apparatus 100 comprises at least interface circuitry 110 and processing circuitry 120.
- the processing circuitry 120 is coupled to the interface circuitry 110.
- the interface circuitry 110 is configured to receive first data 101 indicative of (representing, encoded with) one or more physiological property of the user.
- a physiological property of the user is a property (quantity, characteristic) describing the physiology of the user.
- the physiological property is a property describing one or more function, behavior and/or mechanism in the user’s body.
- the one or more physiological property may be one or more of measured eye-tracking data, measured gaze-tracking data, a walking pattern, a heart (pulse) rate, a heart rate variability of the user, a respiration rate, a blood pressure and an electrodermal activity of the user.
- the interface circuitry 110 is further configured to receive second data 102 indicative of (representing, encoded with) viewing directions of a plurality of people (i.e., N > 2 people).
- the viewing directions of the people are the sight directions of the people along which the people are viewing to perceive their environment.
- the viewing directions may be spatial directions in the real world, a virtual world, or a mixture thereof.
- the second data 102 may comprise at least one of measured eye-tracking data and measured gazetracking data of the people.
- the plurality of people are interacting in the same scene as the user.
- the scene is a place or sphere of an occurrence or action.
- the scene may cover the entire environment in a field of view of the user, but is not limited thereto.
- the scene may, according to examples, additionally cover the environment of the user outside the user’s field of view (such as a region behind or above the user’s face or the face of an avatar representing the user).
- the environment may be a real world environment, a virtual world environment, or a mixture thereof.
- the scene may be a virtual reality scene.
- a virtual reality scene is a scene in a virtual world, which is simulated to give the user an immersive feel of the virtual world.
- the virtual reality scene may be scene in a video game or a metaverse.
- the processing circuitry 120 is configured to receive and process the first data 101 and the second data 102.
- the processing circuitry 120 may be a single dedicated processor, a single shared processor, or a plurality of individual processors, some of which or all of which may be shared, a digital signal processor (DSP) hardware, an application specific integrated circuit (ASIC), a neuromorphic processor or a field programmable gate array (FPGA).
- DSP digital signal processor
- ASIC application specific integrated circuit
- FPGA field programmable gate array
- the processing circuitry 120 may optionally be coupled to, e.g., read only memory (ROM) for storing software, random access memory (RAM) and/or non-volatile memory.
- the apparatus 100 may comprise further circuitry.
- the processing circuitry 120 is configured to determine at least one viewing (view) focus of the plurality of people in the scene based on the second data 102.
- a viewing focus is a point or region in the scene to which the people direct their attention at a particular time instant by looking at or toward the viewing focus.
- the viewing focus may be an object in the scene such as an animal at which at least part of the people are looking at a particular time.
- the second data 102 is indicative of the viewing directions of a plurality of people
- analysis of the second data 102 allows to determine the one or more viewing focus of the plurality of people for a particular time instant.
- the processing circuitry 120 is configured to determine whether the user experiences negative emotions (e.g., now or in the future) based on the first data 101.
- a negative emotion is any feeling which causes the user to feel uncomfortable or sad. For example, discomfort, stress or anxiety may be negative emotions of the user.
- the human body reacts physiologically to emotions. For example, the eye-movement behavior of a stressed or anxious human being is different from the eye-movement behavior of a non-stressed or relaxed human being. Further, the pulse rate as well as the electrodermal activity of a stressed or anx- ious human is increased compared to a non-stressed or relaxed human being.
- the processing circuitry 120 may be configured to determine whether the user is currently experiencing negative emotions or will (is likely to) experience negative emotions in the future (e.g., the near future such as in a few seconds). In other words, the processing circuitry 120 may be configured to determine the current emotional status of the user and/or predict the future emotional status of the user.
- Various techniques for determining whether a user experiences negative emotions based on one or more physiological property exist. Some exemplary techniques will be described later with more details.
- the processing circuitry 120 is configured to cause an output for guiding the viewing direction of the user toward one of the at least one viewing focus of the plurality of people in the scene by a user equipment 130 of the user.
- the user equipment 130 is a device directly used by the user such as an optical, acoustical and/or haptic human-machine interface.
- the user equipment 130 may comprise at least one of optical output means (structure, device) such as a display or a projector, an acoustical output means (structure, device) such as a loudspeaker or an earphone, and a haptic output means (structure, device) such as a vibrating device.
- the graphical output may be one or more symbol or graphic element for drawing the user’s attention to the at least one viewing focus (e.g., a text, an arrow or a frame highlighting the at least one viewing focus).
- the audio output may be one or more sounds or voice commands for drawing the user’s attention to the at least one viewing focus.
- the haptic output may be one or more mechanical outputs by the user equipment 130 perceivable by the user (e.g., motion, vibration, pressure) for drawing the user’s attention to the at least one viewing focus.
- the processing circuitry 120 may, e.g., be configured to generate control data 104 for the user equipment 130.
- the control data 104 may be encoded with (be indicative of, represent) one or more commands for controlling the user equipment 130 to perform the output. Accordingly, the interface circuitry 110 may be configured to send the control data 104 to the user equipment 130 such that the user equipment 130 is able to cause the output for guiding the viewing direction of the user based on the control data 104. In other examples, interface circuitry of the apparatus 100 other than the interface circuitry 110 may be used for sending the control data 104 to the user equipment 130.
- the user’s attention may be directed to the actual object or region in the scene at which the other people in the scene are actually looking.
- the apparatus 100 enables the user to understand that the people are not focusing on him/her but on one or more other object or region in the scene.
- the apparatus 100 may allow to reduce the experience of negative emotions.
- the apparatus 100 may be used for Cognitive Behavioral Therapy (CBT) as the apparatus 100 may allow to make the user aware that his/her perception of the scene, which is based his/her inner beliefs, is erroneous (negatively biased). As the at least one viewing focus of the other people in the scene is determined and output to the user is caused, the user is enabled to understand the actual actions or behaviors of other people in the scene. As a consequence, the apparatus 100 may support the user in understanding his/her reactions and behavior to the other people in the scene. For example, the scene and the behavior of the people in the scene may be scripted to confront the user with predefined situations. The scene and the behavior of the people may, e.g., be selected by a coach or therapist as part of a CBT.
- CBT Cognitive Behavioral Therapy
- the processing circuitry 120 may be configured to not cause an output for guiding the viewing direction of the user as the user’s perception need not be changed. Accordingly, confusion or unnecessary disturbance of the user may be avoided.
- the apparatus 100 allows the user to recognize that his/her perception of the people’s behavior is not correct. Based on the second data 102 indicative of the viewing directions of the plurality of people 220, the apparatus 100 determines the viewing focus of the plurality of people 220 in the scene 200. In the example of Fig. 2, the apparatus 100 hence determines that the dog 230 is the viewing focus of the people 220.
- the apparatus 100 determines, based on the first data 101 indicative of one or more physiological property of the user 210, that the user 210 experiences negative emotions, the apparatus 100 causes an output for guiding the viewing direction of the user 210 toward the viewing focus of the plurality of people 220 in the scene 200 by a user equipment of the user 210. In other words, the apparatus 100 causes an output for guiding the viewing direction of the user 210 toward the dog 230 by a user equipment of the user 210. Accordingly, the apparatus 100 enables the user 210 to recognize that the people 220 look at the dog 230 and not him/her. As described above, this may, e.g., allow to reduce negative emotions of the user 210 and to make the user 210 aware that his/her perception of the scene 200 is not correct.
- eye-tracking data and/or gaze-tracking data of the user may be analyzed.
- the eye movement and eye gaze may be exemplary physiological properties of the user that may be monitored and analyzed for determining whether the user experiences negative emotions.
- the first data 101 may comprise at least one of measured eye-tracking data and measured gaze-tracking data of the user.
- the processing circuitry 120 may, e.g., be configured to extract one or more feature from the at least one of the measured eye-tracking and the measured gaze-tracking data of the user. For example, the blink rate or the pupil width may be extracted from the measured eye-tracking and/or the measured gaze-tracking data.
- the processing circuitry 120 may be further configured to determine whether the user experiences negative emotions based on the one or more extracted feature.
- Document D. Venugopal, J. Amudha and C. Jyotsna "Developing an application using eye tracker," 2016 IEEE International Conference on Recent Trends in Electronics, Information & Communication Technology (RTEICT), Bangalore, India, 2016, pp.
- the first data 101 may comprises at least measured sensor data indicative of one or more of a heart rate, a respiration rate, a blood pressure and an electrodermal activity of the user.
- these parameters change depending on the emotional status.
- the heart rate, the respiration rate, the blood pressure and the electrodermal activity of a user increases when a user is stressed or anxious compared to a time instant at which the user is relaxed.
- the processing circuitry 120 may, e.g., be configured to extract one or more feature from the measured sensor data and determine whether the user experiences negative emotions based on the one or more extracted feature.
- the contextual information may indicate information about the user (e.g., information about known stressors such as certain animals or other objects or situations causing stress).
- the contextual information may indicated presence of objects such as specific animals in the scene that frighten the user. The presence of such objects in the scene makes it more likely that the user experiences negative emotions.
- the day time and the setting may influence the emotional status of the user. For example, many human beings feel uncomfortable in dark streets. In case the scene plays in a scene at night time, it is more likely that the user experiences negative emotions.
- the above pieces of information are exemplary pieces of contextual information of the scene. Analysis of the third data 103 in addition to the first data 101 may allow to improve the accuracy of determining whether the user experiences negative emotions.
- a trained machine-learning model For analyzing the first data 101 and optionally the third data 103, a trained machine-learning model may be used.
- the processing circuitry 120 may be configured to determine whether the user experiences negative emotions using a trained machine-learning model.
- the machine-learning model is a data structure and/or set of rules representing a statistical model that the processing circuitry 120 uses to determine whether the user experiences negative emotions without using explicit instructions, instead relying on models and inference.
- the data structure and/or set of rules represents learned knowledge (e.g. based on training performed by a machine-learning algorithm as described above and below).
- machinelearning instead of a rule-based transformation of data, a transformation of data may be used, that is inferred from an analysis of training data.
- the machine-learning model is trained by a machine-learning algorithm.
- the term "machine-learning algorithm” denotes a set of instructions that are used to create, train or use a machine-learning model.
- the machine-learning model may be trained using training data such as known physical properties of the user as input and predefined emotional stati of the user (e.g., relaxed, stressed, anxious, user experiences negative emotions, user does not experience negative emotions) as target output. This may allow to train a personalized machine-learning model, which is personalized to the user.
- training data of other users may be used in addition to or instead of the training data of the user together with emotional stati of the other users (people) to train a (more) generic machinelearning model, which may allow earlier predictions.
- the machine-learning model By training the machine-learning model with a large set of training data and associated training content information, the machine-learning model "learns" to determine whether the user experiences negative emotions in the training data, so that a target determination whether the user experiences negative emotions is obtained using the machine-learning model.
- the machine-learning model By training the machine-learning model using training physical properties of the user and desired emotional stati of the user, the machine-learning model "learns" a transformation between the physical properties of the user and the desired output, which can be used to provide an output based on non-training physical properties of the user provided to the machine-learning model.
- the machine-learning model may be trained using training input data (e.g. known physical properties of the user).
- the machine-learning model may be trained using a training method called "supervised learning".
- supervised learning the machine-learning model is trained using a plurality of training samples, wherein each sample may comprise a plurality of input data values, and a plurality of desired output values, i.e., each training sample is associated with a desired output value.
- the machine-learning model "learns" which output value to provide based on an input sample that is similar to the samples provided during the training.
- a training sample may comprise one or more physical property of the user as input data and a desired emotional status of the user (i.e., either user experiences negative emotions or user does not experience negative emotions) as desired output data.
- semi-supervised learning may be used.
- semi-supervised learning some of the training samples lack a corresponding desired output value.
- Supervised learning may be based on a supervised learning algorithm (e.g. a classification algorithm or a similarity learning algorithm).
- Classification algorithms may be used as the desired outputs of the trained machine-learning model are restricted to a limited set of values (categorical variables), i.e., the input is classified to one of the limited set of values (e.g., user experiences negative emotions or user does not experience negative emotions).
- Similarity learning algorithms are similar to classification algorithms but are based on learning from examples using a similarity function that measures how similar or related two objects are.
- Reinforcement learning is a third group of machine-learning algorithms.
- reinforcement learning may be used to train the machine-learning model.
- one or more software actors (called “software agents") are trained to take actions in an environment. Based on the taken actions, a reward is calculated.
- Reinforcement learning is based on training the one or more software agents to choose the actions such that the cumulative reward is increased, leading to software agents that become better at the task they are given (as evidenced by increasing rewards).
- Feature learning may be used.
- the machine-learning model may at least partially be trained using feature learning, and/or the machine-learning algorithm may comprise a feature learning component.
- Feature learning algorithms which may be called representation learning algorithms, may preserve the information in their input but also transform it in a way that makes it useful, often as a pre-processing step before performing classification or predictions.
- Feature learning may be based on principal components analysis or cluster analysis, for example.
- the machine-learning model may be an Artificial Neural Network (ANN).
- ANNs are systems that are inspired by biological neural networks, such as can be found in a retina or a brain.
- ANNs comprise a plurality of interconnected nodes and a plurality of connections, so-called edges, between the nodes.
- input nodes that receiving input values (e.g., the respective physical properties of the user), hidden nodes that are (only) connected to other nodes, and output nodes that provide output values (e.g., emotional status of the user such as user experiences negative emotions or user does not experience negative emotions).
- Each node may represent an artificial neuron.
- Each edge may transmit information from one node to another.
- the output of a node may be defined as a (non-linear) function of its inputs (e.g. of the sum of its inputs).
- the inputs of a node may be used in the function based on a "weight" of the edge or of the node that provides the input.
- the weight of nodes and/or of edges may be adjusted in the learning process.
- the training of an ANN may comprise adjusting the weights of the nodes and/or edges of the ANN, i.e., to achieve a desired output for a given input.
- the machine-learning model may be a support vector machine, a random forest model or a gradient boosting model. Support vector machines (i.e.
- support vector networks are supervised learning models with associated learning algorithms that may be used to analyze data (e.g. in classification or regression analysis).
- Support vector machines may be trained by providing an input with a plurality of training input values (e.g., physical properties of the user) that belong to one of two categories (e.g., emotional status of the user such as user experiences negative emotions or user does not experience negative emotions).
- the support vector machine may be trained to assign a new input value to one of the two categories.
- the machine-learning model may be a Bayesian network, which is a probabilistic directed acyclic graphical model.
- a Bayesian network may represent a set of random variables and their conditional dependencies using a directed acyclic graph.
- the machine-learning model may be based on a genetic algorithm, which is a search algorithm and heuristic technique that mimics the process of natural selection.
- the machine-learning model may be a combination of the above examples.
- features of various physiological parameters such as pupil, gaze, heart rate or electrodermal activity may be taken (extracted) from the first data 101 and input into the trained machine-learning model to determine whether the user experiences negative emotions with high accuracy.
- the trained machine-learning model may be understood as a negative emotions detection algorithm, in particular as a stress detection or anxiety detection algorithm.
- the detection that the user experiences the negative emotions may optionally be used to further train the trained machine-learning model.
- the processing circuitry 120 may be configured to determine at least one of contextual information about a region and/or object in the scene causing the negative emotions of the user and information about one or more physiological property of the user while the user experiences the negative emotions.
- the processing circuitry 120 may be configured to determine a viewing focus of the user while the user experiences the negative emotions to determine the region and/or object in the scene causing the negative emotions of the user.
- contextual information such as a type of the object, a content of the region, a distance to the object or region, etc.
- the processing circuitry 120 may be determined contextual information about the region and/or object in the scene causing the negative emotions of the user.
- the processing circuitry 120 may be configured to determine information such as a setting or a time of the scene, a number of people present in the scene or the region, distances of the user to the people present in the scene or the region, etc. as contextual information about the region and/or object in the scene causing the negative emotions of the user.
- the processing circuitry 120 may additionally be configured to further train the trained machine-learning model based on the at least one of the contextual information about the region and/or object in the scene causing the negative emotions of the user and the information about the one or more physiological property of the user while the user experiences the negative emotions.
- the trained machinelearning model may “learn” further correlations between objects, etc. causing negative emotions of the user and the corresponding effects/patterns in the one or more physiological property of the user. In this manner, the accuracy of the trained machine-learning model may be further improved.
- the evaluation of contextual information of the scene such as the contextual information included in the third data 103 may be improved by the further training of the trained machine-learning model.
- the processing circuitry 120 may be configured to continuously determine whether the user experiences negative emotions. For example, the processing circuitry 120 may be configured to determine whether the user experiences negative emotions at regular time intervals. In other example, the processing circuitry 120 may be configured to sporadically determine whether the user experiences negative emotions. For example, the processing circuitry 120 may be configured to determine whether the user experiences negative emotions if certain patterns or features are detected in the one or more physiological property by the processing circuitry 120 (e.g., in case the processing circuitry 120 monitors the respective characteristic of the one or more physiological property indicated by the first data 101).
- Fig- 3 illustrates an exemplary scene 300 to highlight how a viewing focus may be determined.
- Five people 301, 305 are illustrated in Fig. 3.
- the viewing directions of the people 301, ..., 305 at a first time instant ti are indicated by the arrows 361-1, ..., 361-5.
- the viewing directions of the people 301, ..., 305 at a second time instant t2 are indicated by the arrows 362-1, ..., 362-5.
- the viewing directions of the people 301, ..., 305 for a third time instant t3 are indicated by the arrows 363-1, ..., 363-5.
- the second time instant t2 succeeds first time instant ti.
- the third time instant t3 succeeds second time instant t2.
- a plurality of trees 310 are present in the scene as well as a bird 320, some flowers 330 and a fountain 340.
- the first person 301 is looking at a dog 350 playing at the fountain 340 as indicated by arrow 361-1.
- the second person 302, the third person 303, the fourth person 304 and the fifth person 305 are looking at a jogger 360 as indicated by arrows 361- 2, 361-3, 361-4 and 361-5.
- the first person 301, the second person 302 and the fifth person 305 are looking at the dog 350 as indicated by arrows 362-1, 32-2 and 362-5.
- the third person 303 is looking at the fourth person 304 as indicated by arrow 362-3.
- the fourth person 304 is looking at the jogger 360 as indicated by arrow 362-4.
- the first person 301 is looking at a dog 340 as indicated by arrow 363-1.
- the second person 302 is looking at the first person 301 as indicated by arrow 363-2.
- the third person 303 and the fourth person 304 are looking at each other as indicated by arrows 363-3 and 363-4.
- the fifth person 361-5 is again looking at the jogger 360 as indicated by arrow 363-5.
- the processing circuitry 120 receives the second data 102 which are indicative of the viewing directions of the people 301, ..., 305.
- the second data 102 indicate the viewing directions for the time instants ti, t2 and ti as indicated by the arrows in Fig. 3.
- the processing circuitry 120 may be configured to determine one or more region and/or object in the scene at which the people are looking for the respective time instant. For example, for the first time instant ti, the dog 350 and the jogger 360 are determined as objects at which the people 301, . . ., 305 are looking.
- the processing circuitry 120 may be further configured to determine a respective attention rating for the determined one or more region and/or object in the scene.
- the attention rating is a classification according to the attention grade of the people 301, ..., 305.
- a value or score is determined for the one or more region and/or object in the scene which describes the level of attention the one or more region and/or object in the scene gets from the people 301, ..., 305. That is, the attention rating is a measure of how much attention the determined one or more region and/or object in the scene gets from the people 301, ..., 305.
- a high attention rating may indicate that the determined one or more region and/or object in the scene gets a lot of attention from the people 301, . . ., 305
- a low attention rating may indicate that the determined one or more region and/or object in the scene gets only little attention from the people 301, . . ., 305.
- the processing circuitry 120 may be configured to determine one or more of the one or more region and/or object having the highest attention rating as the one or more viewing focus of the plurality of people.
- Various criteria for selecting one or more of the one or more region and/or object as the one or more viewing focus of the plurality of people may be used. For example, a predetermined number of the one or more region and/or object having the highest attention may be determined as the one or more viewing focus of the plurality of people.
- the region(s) and/or object(s) having the highest attention rating, wherein the attention ratio is above a threshold may be determined as the one or more viewing focus of the plurality of people.
- the present disclosure is not limited thereto.
- the processing circuitry 120 may be configured to determine, based on the second data 102, at least one of a respective number of times and a respective accumulated time duration the plurality of people look at a respective one of the one or more region and/or object. For the first time instant ti, the processing circuitry 120 may, e.g., determine how many times the people 301, ..., 305 look at the dog and how long the people 301, . . ., 305 look at dog. As in indicated in Fig.
- the first person 301 is looking at the dog 350 at the first time instant ti such that it is determined that the plurality of people 301, . . 305 look one time at the dog 350 for an accumulated time tdog, ti.
- the second person 302, the third person 303, the fourth person 304 and the fifth person 305 are looking at the jogger 360 at the first time instant ti such that it is determined that the plurality of people 301, . . ., 305 look four times at the jogger 360 for an accumulated time tj ogg er, ti.
- the processing circuitry 120 may, e.g., determine how many times the people 301, . .
- the first person 301, the second person 302 and the fifth person 305 are looking at the dog 350 at the second time instant t2 such that it is determined that the plurality of people 301, . . ., 305 look three times at the dog 350 for an accumulated time tdo g ,t2.
- the third person 303 is looking at the fourth person 304 at the second time instant t2 such that it is determined that the plurality of people 301, . . ., 305 look one time at the fourth person 304 for an accumulated time t pe rson4, t2.
- the fourth person 304 is looking at the jogger 360 at the second time instant t2 such that it is determined that the plurality of people 301, . . ., 305 look one time at the jogger 360 for an accumulated time tj 0gg er,t2. The same may be performed for the third time instant i by the processing circuitry 120.
- the processing circuitry 120 may be further configured to determine the respective attention rating for the one or more region based on the at least one of the respective number of times and the respective accumulated time duration the plurality of people look at a respective one of the one or more region and/or object. For example, for the second time instant t2, the processing circuitry 120 may determine a respective attention rating for the dog 350, the fourth person 304 and the jogger 360 as the people look at these elements in the scene. The respective attention may be determined based on the number of times and the respective accumulated time duration the plurality of people look at the respective element in the scene. For example, the more often the people in the scene look at an element in the scene, the higher may be attention rating for the element.
- processing circuitry 120 may, e.g., determine that the plurality of people 301, ..., 305 look one time at the dog 350 for an accumulated time tdog, ti and that the plurality of people 301, ..., 305 look four times at the jogger 360 for an accumulated time tjogger, ti. Accordingly, the processing circuitry 120 may further determine an attention rating for the dog 350 and the jogger 360.
- the processing circuitry 120 may, e.g., determine an attention rating for the dog 350, the fourth person 304 and the jogger 360. As most people look at the dog 350, the attention rating for the dog 350 may be high, whereas the attention ratings for the fourth person 304 and the jogger 360 may be lower. The same may be performed for the third time instant t3 by the processing circuitry 120.
- the jogger 360 may be determined as viewing focus of the plurality of people 301, ..., 305 for the first time instant ti by the processing circuitry 120 and the dog 350 may be determined as viewing focus of the plurality of people 301, . . ., 305 for the second time instant t2.
- the processing circuitry 120 may create a heat map highlighting the one or more region and/or object in the scene that attract the attention of the plurality of people 301, . . ., 305.
- the viewing directions of the plurality of people in the scene indicated by the second data 102 as well the analysis results may further be used for an enhanced interpretation of the scene. This is exemplary illustrated in Fig. 4 for the scene 300.
- the first person 301 is looking at the dog 350 at all three instants ti, t2 and t3. Due to the first person 301’s constant attention on the dog 350, the processing circuitry 120 may determine that the first person 301 is probably the owner of the dog 350 and hence better understand the scene 300.
- the processing circuitry 120 may take into account further data for determining the respective attention rating for the determined one or more region and/or object in the scene.
- the interface circuitry 110 may be further configured to determine the respective attention rating for the one or more region and/or object based on the third data 103, i.e., based on the contextual information of the scene.
- the contextual information may, e.g., indicate that the jogger 360 is running in the scene. Moving objects attract more attention than static objects.
- the determination of the respective attention rating may be more accurate.
- the attention rating for the jogger 360 may be increased for the first time instant ti as the contextual data indicate that the jogger is running (moving) in the scene 300.
- the processing circuitry 120 may be configured to continuously determine the respective attention rating for the one or more region and/or object. Continuously determining the respective attention rating allows to adapt the respective attention rating to the current state of the scene. According to examples, the processing circuitry 120 may be configured to determine the respective attention rating for the one or more region and/or object based on the attention rating for the one or more region and/or object determined for one or more previous time instant. Similarly, the processing circuitry 120 may be configured to determine the respective attention rating for the one or more region and/or object based on the current viewing directions of the plurality of people and viewing directions of the plurality of people for one or more previous time instant.
- the processing circuitry 120 may determine the attention ratio for the dog 350 for the second time instant t2 based on the attention ratio for the dog 350 for the first time instant ti. Accordingly, the development of the scene may be taken into account to increase the accuracy of the attention ratio determination.
- the processing circuitry 120 may determine that the second person 302 is interested in a communication with the first person 301 about the dog 350 and hence better understand the scene 300.
- the third person 303 and the fourth person 304 look at each other after both looked at the jogger 360 at the first time instant ti.
- the third person 303 already looks at the fourth person 304 at the second time instant t2, while the fourth person 304 still looks at the jogger 360 at the second time instant t2.
- the processing circuitry 120 may determine that the third person 303 and the fourth person 304 are interested in a communication about the jogger 360 and hence better understand the scene 300.
- the processing circuitry 120 may hence be configured to determine intentions of the people in the scene and determine, e.g., a respective probability for the determined intention.
- the determined intentions as well as the probabilities for the determined intentions may be used for determining the respective attention ratio for the one or more object and/or region in the scene 300.
- determined intentions as well as the probabilities for the determined intentions may further be used as contextual information of the scene for determining whether the user experiences negative emotions.
- Adding contextual information to the heat map with the attention rating may allow to create a storyline (e.g., people’s attention drifts from the sunset to a swarm of birds because moving objects cause instantaneous attention and then back to the sunset) which can help in a CBT to explain why people focus attention on specific sequences of objects. It might also learn the input from the attention model to predict which scene will cause the most attention in the future in healthy user and in users with mental diseases.
- a storyline e.g., people’s attention drifts from the sunset to a swarm of birds because moving objects cause instantaneous attention and then back to the sunset
- the processing circuitry 120 is configured to determine the at least one viewing focus of the plurality of people using a trained machine-learning model.
- the trained machine-learning model for determining the at least one viewing focus of the plurality of people in the scene may be integrated into the trained machine-learning model for determining whether the user experiences negative emotions or be a separate machine-learning model.
- the machine learning model may be trained analogously to what is described above. For example, for the machine-learning model to determine the at least one viewing focus, the machine-learning model may be trained using training data indicative of viewing directions of a plurality of people in a scene as input and one or more predefined viewing focus as target output. However, as described also other training methods such as semi-supervised learning, unsupervised learning or reinforcement learning may be used.
- the viewing directions of the plurality of people in the scene may be taken (extracted) from the second data 102 and input into the trained machine-learning model to determine the at least one viewing focus of the plurality of people.
- the trained machine-learning model may be understood as an attention model.
- the trained machine-learning model may be used analogously to what is described above to create a heatmap where objects or regions on the screen/scene get rated with different attention inputs.
- the rating may depend on how often and how long people look at the specific object or region.
- further features such as determined intentions of the people in the scene and probabilities for the determined intentions may be used.
- the attention rating value may change over time such that a continuous update of the rating allows to determine at which time instant the specific object or region caused attraction.
- the first data 101, the second data 102 and the third data 103 described above may be provided from various sources.
- the first data 101 may at least in part be received from the user equipment of the user.
- the second data 102 may, e.g., at least in part be received from one or more user equipment of the plurality of people.
- at least part of the first data 101 and/or at least part of the second data 102 may be received from a computing hardware (e.g., a server or a computing cloud) hosting the virtual world that includes the virtual reality scene.
- the scene is an augmented reality scene
- at least part of the first data 101 and/or at least part of the second data 102 may be received from a computing hardware (e.g., a server or a computing cloud) providing the virtual objects augmented to the real world.
- the third data 103 may at least in part be received from the respective computing platform.
- at least part of the third data 103 may be generated by the processing circuitry 120 itself while performing viewing direction guidance according to the present disclosure.
- the apparatus 100 may, e.g., be part of or be coupled to a computing hardware (e.g., a server or a computing cloud) hosting the virtual world that includes the virtual reality scene.
- a computing hardware e.g., a server or a computing cloud
- the apparatus 100 may, e.g., be part of or be coupled to a computing hardware (e.g., a server or a computing cloud) providing the virtual objects augmented to the real world.
- Fig. 5 schematically illustrates an exemplary user equipment 500 for the user.
- the user equipment 500 comprises one or more sensor 510 configured to measure one or more physiological property of the user while the user is interacting in the same scene as a plurality of people.
- the one or more physiological property may be one or more of measured eye-tracking data, measured gaze-tracking data of the user, a heart (pulse) rate, a heart rate variability of the user, a respiration rate, a blood pressure and an electrodermal activity of the user.
- the one or more sensor 510 may be or be configured to perform the functionalities of one or more of an eye-tracking sensor, a gazetracking sensor, a Galvanic Skin Response (GSR) sensor, a PhotoPlethysmoGraphy (PPG) sensor, a Laser Doppler Flowmetry (LDF) sensor, and an ElectroMyoGraphy (EMG) sensor.
- GSR Galvanic Skin Response
- PPG PhotoPlethysmoGraphy
- LDF Laser Doppler Flowmetry
- EMG ElectroMyoGraphy
- the user equipment 500 comprises interface circuitry 520 coupled to the one or more sensor 510.
- the interface circuitry 520 is configured to send data 501 indicative of (representing, encoded with) the measured one or more physiological property of the user to an external device 550 (i.e., a device external to / separate from the user equipment 500).
- the external device 550 may be the apparatus 100 described above. If the scene is a virtual reality scene, the external device 550 may, e.g., be a computing hardware hosting the virtual world that includes the virtual reality scene. If the scene is an augmented reality scene, the external device 550 may, e.g., be a computing hardware providing the virtual objects augmented to the real world. The respective computing hardware may forward the data 501 to, e.g., the apparatus 100.
- the interface circuitry 520 is further configured to receive, from the external device 550, control data 502 encoded with one or more commands for controlling the user equipment 500 to perform output for guiding the viewing direction of the user toward a determined viewing focus of the plurality of people in the scene.
- the control data 502 may, e.g., be determined by the apparatus 100 as described above.
- the user equipment 500 further comprises a human-machine interface 530 configured to perform output for guiding the viewing direction of the user toward the viewing focus based on the control data 502.
- the human-machine interface 530 may comprise at least one of optical output means (structure, device) such as a display or a projector, an acoustical output means (structure, device) such as a loudspeaker or an earphone, and an haptic output means (structure, device) such as a vibrating device.
- the graphical output may be one or more symbol or graphic element for drawing the user’s attention to the viewing focus (e.g., a text, an arrow or a frame highlighting the at least one viewing focus).
- the audio output may be one or more sounds or voice commands for drawing the user’s attention to the viewing focus.
- the haptic output may be one or more mechanical outputs by the human-machine interface 530 perceivable by the user (e.g., motion, vibration, pressure) for drawing the user’s attention to the viewing focus.
- the user equipment 500 may allow to guide the user’s viewing direction toward the viewing focus based on the control data 502. In particular, the user equipment 500 may be used together with the apparatus 100 to make a user understand that the people in the scene are not focusing on him/her but on one or more other object or region in the scene.
- the user equipment 500 may, e.g., be a head-mounted equipment such as a virtual reality headset or smart glasses.
- the user equipment 500 may optionally further comprise one or more further sensor 540 configured to measure viewing directions of (at least part of) the plurality of people in the scene.
- the one or more further sensor 540 may be or be configured to perform the functionalities of one or more of a camera, an eye-tracking sensor and a gazetracking sensor.
- the interface circuitry 520 may be further configured to send data 503 indicative of the measured viewing directions of the plurality of people to the external device 550.
- the user equipment 500 may allow to collect viewing directions of the people in the vicinity of the user such that the proposed viewing direction guidance may be performed even in case the other people in the scene do not wear or use a respective user equipment that allows to measure the viewing direction of the person.
- the interface circuitry 520 may be configured to receive from user equipments 560 of at least part of the other people in the scene data 504 indicative of measured viewing directions of the plurality of people in the scene.
- the interface circuitry may receive data 504 indicative of measured viewing directions from plural other user equipments 560.
- the interface circuitry 520 may be further configured to send data 503 indicative of the measured viewing directions of the plurality of people to the external device 550.
- the interface circuitry 520 may be configured to send data 505 indicative of the one or more physiological property of the user to the user equipments 560 of at least part of the people in the scene.
- the method 600 comprises determining 608 whether the user experiences negative emotions based on the first data. If it is determined that the user experiences negative emotions, the method 600 comprises causing 610 an output for guiding the viewing direction of the user toward one of the at least one viewing focus of the plurality of people in the scene by a user equipment of the user.
- the method 600 may, e.g., enable the user to understand that the people are not focusing on him but on one or more other object or region in the scene. As the user is enabled to recognize that he/she is not the other people’s object or region of attention in the scene, the method 600 may allow to reduce the experience of negative emotions.
- the method 600 may be used for CBT as the method 600 may allow to make the user aware that his/her perception of the scene, which is based his/her inner beliefs, is erroneous.
- the method 600 may comprise one or more additional optional features corresponding to one or more aspects of the proposed technique or one or more examples described above. For example, if it is not determined that the user experiences negative emotions, the method 600 may comprise not causing 610 an output for guiding the viewing direction of the user as the user’s perception need not be changed. Accordingly, confusion or unnecessary disturbance of the user may be avoided.
- An apparatus for guiding a viewing direction of a user comprising: interface circuitry configured to: receive first data indicative of one or more physiological property of the user; and receive second data indicative of viewing directions of a plurality of people, the plurality of people interacting in the same scene as the user; and processing circuitry configured to: determine at least one viewing focus of the plurality of people in the scene based on the second data; determine whether the user experiences negative emotions based on the first data; and if it is determined that the user experiences negative emotions, cause an output for guiding the viewing direction of the user toward one of the at least one viewing focus of the plurality of people in the scene by a user equipment of the user.
- the processing circuitry is configured to: extract one or more feature from the at least one of the measured eye-tracking and the measured gaze-tracking data of the user; and determine whether the user experiences negative emotions based on the one or more extracted feature.
- the processing circuitry is configured to: extract one or more feature from the measured sensor data; and determine whether the user experiences negative emotions based on the one or more extracted feature.
- the processing circuitry is further configured to: determine at least one of contextual information about a region and/or object in the scene causing the negative emotions of the user and information about one or more physiological property of the user while the user experiences the negative emotions; and further train the trained machine-learning model based on the at least one of the contextual information about the region and/or object in the scene causing the negative emotions of the user and the information about the one or more physiological property of the user while the user experiences the negative emotions.
- the processing circuitry is, based on the second data, configured to: determine one or more region and/or object in the scene at which the people are looking; determine a respective attention rating for the one or more region and/or object; and determine one or more of the one or more region and/or object having the highest attention rating as the one or more viewing focus of the plurality of people.
- the processing circuitry is configured to: determine, based on the second data, at least one of a respective number of times and a respective accumulated time duration the plurality of people look at a respective one of the one or more region and/or object; and determine the respective attention rating for the one or more region based on the at least one of the respective number of times and the respective accumulated time duration the plurality of people look at a respective one of the one or more region and/or object.
- the processing circuitry is configured to generate control data for the user equipment, the control data being encoded with one or more commands for controlling the user equipment to perform the output, and wherein the interface circuitry is configured to send the control data to the user equipment.
- a user equipment for a user comprising: one or more sensor configured to measure one or more physiological property of the user while the user is interacting in the same scene as a plurality of people; interface circuitry configured to: send data indicative of the one or more physiological property of the user to an external device; and receive, from the external device, control data encoded with one or more commands for controlling the user equipment to perform output for guiding the viewing direction of the user toward a determined viewing focus of the plurality of people in the scene; and a human-machine interface configured to perform output for guiding the viewing direction of the user toward the viewing focus based on the control data.
- a method for guiding a viewing direction of a user comprising: receiving first data indicative of a physiological property of the user; receiving second data indicative of viewing directions of a plurality of people, the plurality of people interacting in the same scene as the user; determining at least one viewing focus of the plurality of people in the scene based on the second data; determining whether the user experiences negative emotions based on the first data; and if it is determined that the user experiences negative emotions, causing an output for guiding the viewing direction of the user toward one of the at least one viewing focus of the plurality of people in the scene by a user equipment of the user.
- Examples may further be or relate to a (computer) program including a program code to execute one or more of the above methods when the program is executed on a computer, processor or other programmable hardware component.
- steps, operations or processes of different ones of the methods described above may also be executed by programmed computers, processors or other programmable hardware components.
- Examples may also cover program storage devices, such as digital data storage media, which are machine-, processor- or computer-readable and encode and/or contain machine-executable, processorexecutable or computer-executable programs and instructions.
- Program storage devices may include or be digital storage devices, magnetic storage media such as magnetic disks and magnetic tapes, hard disk drives, or optically readable digital data storage media, for example.
- Other examples may also include computers, processors, control units, (field) programmable logic arrays ((F)PLAs), (field) programmable gate arrays ((F)PGAs), graphics processor units (GPU), ASICs, integrated circuits (ICs) or system-on-a-chip (SoCs) systems programmed to execute the steps of the methods described above.
- FPLAs field programmable logic arrays
- FPGAs field programmable gate arrays
- GPU graphics processor units
- ASICs integrated circuits
- ICs integrated circuits
- SoCs system-on-a-chip
- aspects described in relation to a device or system should also be understood as a description of the corresponding method.
- a block, device or functional aspect of the device or system may correspond to a feature, such as a method step, of the corresponding method.
- aspects described in relation to a method shall also be understood as a description of a corresponding block, a corresponding element, a property or a functional feature of a corresponding device or a corresponding system.
Landscapes
- Engineering & Computer Science (AREA)
- Health & Medical Sciences (AREA)
- Life Sciences & Earth Sciences (AREA)
- Physics & Mathematics (AREA)
- Theoretical Computer Science (AREA)
- General Engineering & Computer Science (AREA)
- General Health & Medical Sciences (AREA)
- Psychiatry (AREA)
- General Physics & Mathematics (AREA)
- Hospice & Palliative Care (AREA)
- Psychology (AREA)
- Child & Adolescent Psychology (AREA)
- Developmental Disabilities (AREA)
- Biomedical Technology (AREA)
- Social Psychology (AREA)
- Medical Informatics (AREA)
- Public Health (AREA)
- Veterinary Medicine (AREA)
- Educational Technology (AREA)
- Surgery (AREA)
- Molecular Biology (AREA)
- Heart & Thoracic Surgery (AREA)
- Human Computer Interaction (AREA)
- Biophysics (AREA)
- Animal Behavior & Ethology (AREA)
- Pathology (AREA)
- Multimedia (AREA)
- Dermatology (AREA)
- Neurology (AREA)
- Neurosurgery (AREA)
- Primary Health Care (AREA)
- Epidemiology (AREA)
- User Interface Of Digital Computer (AREA)
Abstract
An apparatus for guiding a viewing direction of a user is provided. The apparatus includes interface circuitry configured to receive first data indicative of one or more physiological property of the user. The interface circuitry is further configured to receive second data in- dicative of viewing directions of a plurality of people. The plurality of people are interacting in the same scene as the user. Additionally, the apparatus includes processing circuitry con- figured to determine at least one viewing focus of the plurality of people in the scene based on the second data. Further, the processing circuitry is configured to determine whether the user experiences negative emotions based on the first data. If it is determined that the user experiences negative emotions, the processing circuitry is configured to cause an output for guiding the viewing direction of the user toward one of the at least one viewing focus of the plurality of people in the scene by a user equipment of the user.
Description
APPARATUS AND METHOD FOR GUIDING A VIEWING DIRECTION OF A USER, AND USER EQUIPMENT
Field
The present disclosure relates to changing a user’s perception. In particular, examples of the present disclosure relate to an apparatus and a method for guiding a viewing direction of a user. Further examples of the present disclosure relate to a user equipment, a non-transitory machine-readable medium and a program.
Background
Human beings perceive their environment visually. The way a human being experiences the world is based on underlying beliefs and the self-perception of the user. As a consequence, some human beings tend to interpret actions or behaviors, in particular the viewing behavior, of other people incorrectly and experience feelings such as stress, discomfort or anxiety.
Hence, there may be a demand for (objectively) informing a user about the viewing behavior of other people.
Summary
This demand is met by an apparatus for guiding a viewing direction of a user, a method for guiding a viewing direction of a user, a user equipment, a non-transitory machine-readable medium and a program in accordance with the independent claims. Advantageous embodiments are defined the dependent claims.
According to a first aspect, the present disclosure provides an apparatus for guiding a viewing direction of a user. The apparatus comprises interface circuitry configured to receive first data indicative of one or more physiological property of the user. The interface circuitry is further configured to receive second data indicative of viewing directions of a plurality of people. The plurality of people are interacting in the same scene as the user. Additionally, the apparatus comprises processing circuitry configured to determine at least one viewing
focus of the plurality of people in the scene based on the second data. Further, the processing circuitry is configured to determine whether the user experiences negative emotions based on the first data. If it is determined that the user experiences negative emotions, the processing circuitry is configured to cause an output for guiding the viewing direction of the user toward one of the at least one viewing focus of the plurality of people in the scene by a user equipment of the user.
According to a second aspect, the present disclosure provides a method for guiding a viewing direction of a user. The method comprises receiving first data indicative of a physiological property of the user. In addition, the method comprises receiving second data indicative of viewing directions of a plurality of people. The plurality of people are interacting in the same scene as the user. The method further comprises determining at least one viewing focus of the plurality of people in the scene based on the second data. Additionally, the method comprises determining whether the user experiences negative emotions based on the first data. If it is determined that the user experiences negative emotions, the method comprises causing an output for guiding the viewing direction of the user toward one of the at least one viewing focus of the plurality of people in the scene by a user equipment of the user.
According to a third aspect, the present disclosure provides a user equipment for a user. The user equipment comprises one or more sensor configured to measure one or more physiological property of the user while the user is interacting in the same scene as a plurality of people. Further, the user equipment comprises interface circuitry configured to send data indicative of the one or more physiological property of the user to an external device. The interface circuitry is additionally configured to receive, from the external device, control data encoded with one or more commands for controlling the user equipment to perform output for guiding the viewing direction of the user toward a determined viewing focus of the plurality of people in the scene. In addition, the user equipment comprises a humanmachine interface configured to perform output for guiding the viewing direction of the user toward the viewing focus based on the control data.
According to a fourth aspect, the present disclosure provides a non-transitory machine- readable medium having stored thereon a program having a program code for performing the method according to the second aspect, when the program is executed on a processor or a programmable hardware.
According to a fifth aspect, the present disclosure provides a program having a program code for performing the method according to the second aspect, when the program is executed on a processor or a programmable hardware.
Brief description of the Figures
Some examples of apparatuses and/or methods will be described in the following by way of example only, and with reference to the accompanying figures, in which
Fig. 1 illustrates an example of an apparatus for guiding a viewing direction of a user;
Fig. 2 illustrates a first exemplary scene;
Fig. 3 illustrates a second exemplary scene;
Fig. 4 illustrates an enhanced interpretation of the second exemplary scene;
Fig. 5 illustrates an example of a user equipment; and
Fig. 6 illustrates a flowchart of an example of a method for guiding a viewing direction of a user.
Detailed Description
Some examples are now described in more detail with reference to the enclosed figures. However, other possible examples are not limited to the features of these embodiments described in detail. Other examples may include modifications of the features as well as equivalents and alternatives to the features. Furthermore, the terminology used herein to describe certain examples should not be restrictive of further possible examples.
Throughout the description of the figures same or similar reference numerals refer to same or similar elements and/or features, which may be identical or implemented in a modified
form while providing the same or a similar function. The thickness of lines, layers and/or areas in the figures may also be exaggerated for clarification.
When two elements A and B are combined using an “or”, this is to be understood as disclosing all possible combinations, i.e. only A, only B as well as A and B, unless expressly defined otherwise in the individual case. As an alternative wording for the same combinations, "at least one of A and B" or "A and/or B" may be used. This applies equivalently to combinations of more than two elements.
If a singular form, such as “a”, “an” and “the” is used and the use of only a single element is not defined as mandatory either explicitly or implicitly, further examples may also use several elements to implement the same function. If a function is described below as implemented using multiple elements, further examples may implement the same function using a single element or a single processing entity. It is further understood that the terms "include", "including", "comprise" and/or "comprising", when used, describe the presence of the specified features, integers, steps, operations, processes, elements, components and/or a group thereof, but do not exclude the presence or addition of one or more other features, integers, steps, operations, processes, elements, components and/or a group thereof.
Fig- 1 illustrates an exemplary apparatus 100 for guiding a viewing (view) direction of a user. The viewing direction of the user is the sight direction of the user along which the user is viewing to perceive his/her environment. The viewing direction may be a spatial direction in the real world, a virtual world, or a mixture thereof.
The apparatus 100 comprises at least interface circuitry 110 and processing circuitry 120. The processing circuitry 120 is coupled to the interface circuitry 110. The interface circuitry 110 is configured to receive first data 101 indicative of (representing, encoded with) one or more physiological property of the user. A physiological property of the user is a property (quantity, characteristic) describing the physiology of the user. In other words, the physiological property is a property describing one or more function, behavior and/or mechanism in the user’s body. For example, the one or more physiological property may be one or more of measured eye-tracking data, measured gaze-tracking data, a walking pattern, a heart (pulse) rate, a heart rate variability of the user, a respiration rate, a blood pressure and an electrodermal activity of the user.
The interface circuitry 110 is further configured to receive second data 102 indicative of (representing, encoded with) viewing directions of a plurality of people (i.e., N > 2 people). The viewing directions of the people are the sight directions of the people along which the people are viewing to perceive their environment. The viewing directions may be spatial directions in the real world, a virtual world, or a mixture thereof. For example, the second data 102 may comprise at least one of measured eye-tracking data and measured gazetracking data of the people.
The plurality of people are interacting in the same scene as the user. The scene is a place or sphere of an occurrence or action. The scene may cover the entire environment in a field of view of the user, but is not limited thereto. The scene may, according to examples, additionally cover the environment of the user outside the user’s field of view (such as a region behind or above the user’s face or the face of an avatar representing the user). The environment may be a real world environment, a virtual world environment, or a mixture thereof. In particular, the scene may be a virtual reality scene. A virtual reality scene is a scene in a virtual world, which is simulated to give the user an immersive feel of the virtual world. For example, the virtual reality scene may be scene in a video game or a metaverse. In other examples, the scene may be an augmented (mixed) reality scene. An augmented reality scene is a scene in a representation comprising real and virtual objects. For example, real world content may be combined with computer(artificially)-generated content such that it is perceived by the user as an immersive aspect of the real world.
As described above, the plurality of people as well as the user are interacting in the same scene. The interaction of the people and the user in the scene may be manifold. In particular, the plurality of people and the user may interact with each other, but need not. In some examples, the plurality of people and the user may act independently from each other in the scene. For example, there may be direct interaction between the people and the user, i.e., the user and at least one of the people may mutually or reciprocally act with or influence each other (e.g., via verbal communication, actions or gestures). In other example, there may only be direction between two or more of the people, but not between the people and the user. In case of a virtual reality scene, the interaction may be performed by the respective avatar of the user and the people in the virtual world. In still other examples, there may be no direc-
tion interaction between the plurality of people and also no direct interaction between the user and the plurality of people. In these examples, the plurality of people and also the user may individually interact in the scene (e.g., with one or more respective object in the scene).
The processing circuitry 120 is configured to receive and process the first data 101 and the second data 102. For example, the processing circuitry 120 may be a single dedicated processor, a single shared processor, or a plurality of individual processors, some of which or all of which may be shared, a digital signal processor (DSP) hardware, an application specific integrated circuit (ASIC), a neuromorphic processor or a field programmable gate array (FPGA). The processing circuitry 120 may optionally be coupled to, e.g., read only memory (ROM) for storing software, random access memory (RAM) and/or non-volatile memory. Optionally, the apparatus 100 may comprise further circuitry.
In particular, the processing circuitry 120 is configured to determine at least one viewing (view) focus of the plurality of people in the scene based on the second data 102. A viewing focus is a point or region in the scene to which the people direct their attention at a particular time instant by looking at or toward the viewing focus. For example, the viewing focus may be an object in the scene such as an animal at which at least part of the people are looking at a particular time. There may be more than one viewing focus of the plurality of people at a particular time instant as different ones of the plurality of people may focus their attention to different objections or regions in/of the scene at a particular time instant. As the second data 102 is indicative of the viewing directions of a plurality of people, analysis of the second data 102 allows to determine the one or more viewing focus of the plurality of people for a particular time instant. Various techniques for determining a viewing focus based on viewing directions of people exist. Some exemplary techniques will be described later with more details.
Furthermore, the processing circuitry 120 is configured to determine whether the user experiences negative emotions (e.g., now or in the future) based on the first data 101. A negative emotion is any feeling which causes the user to feel miserable or sad. For example, discomfort, stress or anxiety may be negative emotions of the user. The human body reacts physiologically to emotions. For example, the eye-movement behavior of a stressed or anxious human being is different from the eye-movement behavior of a non-stressed or relaxed human being. Further, the pulse rate as well as the electrodermal activity of a stressed or anx-
ious human is increased compared to a non-stressed or relaxed human being. As the first data 101 is indicative of one or more physiological property of the user, analysis of the first data 101 allows to determine whether the user experiences negative emotions. In particular, the processing circuitry 120 may be configured to determine whether the user is currently experiencing negative emotions or will (is likely to) experience negative emotions in the future (e.g., the near future such as in a few seconds). In other words, the processing circuitry 120 may be configured to determine the current emotional status of the user and/or predict the future emotional status of the user. Various techniques for determining whether a user experiences negative emotions based on one or more physiological property exist. Some exemplary techniques will be described later with more details.
If it is determined that the user experiences negative emotions, the processing circuitry 120 is configured to cause an output for guiding the viewing direction of the user toward one of the at least one viewing focus of the plurality of people in the scene by a user equipment 130 of the user. The user equipment 130 is a device directly used by the user such as an optical, acoustical and/or haptic human-machine interface. For example, the user equipment 130 may comprise at least one of optical output means (structure, device) such as a display or a projector, an acoustical output means (structure, device) such as a loudspeaker or an earphone, and a haptic output means (structure, device) such as a vibrating device. The user equipment 130 may, e.g., be a head-mounted equipment such as a virtual reality headset or smart glasses. The output for guiding the viewing direction of the user may be manifold. For example, the processing circuitry 120 may be configured to cause at least one of a graphic output, an audio output and a haptic output by the user equipment 130 if it is determined that the user experiences negative emotions. The type(s) of output may, e.g., be adapted to one or more of preferences of the user, capabilities of the user equipment 130 and a determined level of negative emotions experienced by the user (e.g., the number and/or the intensity of the outputs may correlate to the determined level of negative emotions). For example, the graphical output may be one or more symbol or graphic element for drawing the user’s attention to the at least one viewing focus (e.g., a text, an arrow or a frame highlighting the at least one viewing focus). The audio output may be one or more sounds or voice commands for drawing the user’s attention to the at least one viewing focus. The haptic output may be one or more mechanical outputs by the user equipment 130 perceivable by the user (e.g., motion, vibration, pressure) for drawing the user’s attention to the at least one viewing focus.
For causing the output by the user equipment 130, the processing circuitry 120 may, e.g., be configured to generate control data 104 for the user equipment 130. The control data 104 may be encoded with (be indicative of, represent) one or more commands for controlling the user equipment 130 to perform the output. Accordingly, the interface circuitry 110 may be configured to send the control data 104 to the user equipment 130 such that the user equipment 130 is able to cause the output for guiding the viewing direction of the user based on the control data 104. In other examples, interface circuitry of the apparatus 100 other than the interface circuitry 110 may be used for sending the control data 104 to the user equipment 130.
Many human beings such as the user feel uncomfortable, stressed or even anxious when being in the vicinity of other people. Due their underlying beliefs and their self-perception they erroneously interpret the actions or behaviors of other people such as their viewing behavior. Very often these human beings are afraid that the other people are staring at them, making fun of them or do not accept them. In other words, these human beings erroneously interpret the actions or behaviors of other people as reactions to their own presence or behavior such that negative emotions are caused. Therefore, such human beings perceive other people in a scene as potential threat. By determining the at least one viewing focus of the other people in the scene and causing the output for guiding the viewing direction of the user toward one of the at least one viewing focus of the plurality of people in the scene, the user’s attention may be directed to the actual object or region in the scene at which the other people in the scene are actually looking. By drawing the user’s attention to the object or region in the scene at which the other people in the scene are actually looking, the apparatus 100 enables the user to understand that the people are not focusing on him/her but on one or more other object or region in the scene. As the user is enabled to recognize that he/she is not the other people’s object or region of attention in the scene, the apparatus 100 may allow to reduce the experience of negative emotions.
Furthermore, the apparatus 100 may be used for Cognitive Behavioral Therapy (CBT) as the apparatus 100 may allow to make the user aware that his/her perception of the scene, which is based his/her inner beliefs, is erroneous (negatively biased). As the at least one viewing focus of the other people in the scene is determined and output to the user is caused, the user is enabled to understand the actual actions or behaviors of other people in the scene. As a
consequence, the apparatus 100 may support the user in understanding his/her reactions and behavior to the other people in the scene. For example, the scene and the behavior of the people in the scene may be scripted to confront the user with predefined situations. The scene and the behavior of the people may, e.g., be selected by a coach or therapist as part of a CBT. However, it is to be noted that the present disclosure is not limited thereto. In other words, the scene and the behavior of the people in the scene need not be scripted. For example, if the scene is a virtual reality scene in a virtual world such as a metaverse, the other people may be random real people interacting in the metaverse via their respective avatar. Similarly, if the scene is an augmented reality scene, the other people may be random real people interacting in the real world.
On the other hand, if it is not determined that the user experiences negative emotions the processing circuitry 120 may be configured to not cause an output for guiding the viewing direction of the user as the user’s perception need not be changed. Accordingly, confusion or unnecessary disturbance of the user may be avoided.
Fig- 2 illustrates an exemplary scene 200. A user 210 as well as a plurality of other people 220 are interacting in the scene 200. In particular, the plurality of people 220 looking in the direction of the user 210. Due to his/her inner beliefs, the user 210 may feel that the plurality of people 220 look/stare at him/her as indicated by the arrows 221 in Fig. 2. Accordingly, the user 210 may experience negative emotions such as discomfort, stress or anxiety. However, as indicated by the arrows 222 in Fig. 2, the plurality of people 220 actually look at the dog 230 which is located near the user 210 in the scene 200.
The apparatus 100 allows the user to recognize that his/her perception of the people’s behavior is not correct. Based on the second data 102 indicative of the viewing directions of the plurality of people 220, the apparatus 100 determines the viewing focus of the plurality of people 220 in the scene 200. In the example of Fig. 2, the apparatus 100 hence determines that the dog 230 is the viewing focus of the people 220.
In case the apparatus 100 determines, based on the first data 101 indicative of one or more physiological property of the user 210, that the user 210 experiences negative emotions, the apparatus 100 causes an output for guiding the viewing direction of the user 210 toward the viewing focus of the plurality of people 220 in the scene 200 by a user equipment of the
user 210. In other words, the apparatus 100 causes an output for guiding the viewing direction of the user 210 toward the dog 230 by a user equipment of the user 210. Accordingly, the apparatus 100 enables the user 210 to recognize that the people 220 look at the dog 230 and not him/her. As described above, this may, e.g., allow to reduce negative emotions of the user 210 and to make the user 210 aware that his/her perception of the scene 200 is not correct.
As described above, various techniques for determining whether a user experiences negative emotions based on one or more physiological property exist. Some exemplary techniques will be described in the following. However, it is to be noted that the present disclosure is not limited thereto.
For example, eye-tracking data and/or gaze-tracking data of the user may be analyzed. There is a direct innervation of the autonomous nervous system of a human being to the muscles controlling the eyes such that various features related to the eyes such as the blinking rate or the pupil width allow to estimate the emotional status of the user. Therefore, the eye movement and eye gaze may be exemplary physiological properties of the user that may be monitored and analyzed for determining whether the user experiences negative emotions. For example, the first data 101 may comprise at least one of measured eye-tracking data and measured gaze-tracking data of the user. For determining whether the user experiences negative emotions, the processing circuitry 120 may, e.g., be configured to extract one or more feature from the at least one of the measured eye-tracking and the measured gaze-tracking data of the user. For example, the blink rate or the pupil width may be extracted from the measured eye-tracking and/or the measured gaze-tracking data. The processing circuitry 120 may be further configured to determine whether the user experiences negative emotions based on the one or more extracted feature. Document D. Venugopal, J. Amudha and C. Jyotsna, "Developing an application using eye tracker," 2016 IEEE International Conference on Recent Trends in Electronics, Information & Communication Technology (RTEICT), Bangalore, India, 2016, pp. 1518-1522, doi: 10.1109/RTEICT.2016.7808086, document M. S. Yousefi, F. Reisi, M. R. Daliri and V. Shalchyan, "Stress Detection Using Eye Tracking Data: An Evaluation of Full Parameters," in IEEE Access, vol. 10, pp. 118941-118952, 2022, doi: 10.1109/ ACCESS.2022.3221179 or document C. Jyotsna and J. Amudha, "Eye Gaze as an Indicator for Stress Level Analysis in Students," 2018 International Conference on Advances in Computing, Communications and Informatics (ICACCI), Bangalore, India,
2018, pp. 1588-1593, doi: 10.1109/ICACCI.2018.8554715 describe exemplary techniques for determining the emotional status of a user in accordance with the above general description.
Similarly to the above, other physiological properties may be monitored and analyzed. For example, the first data 101 may comprises at least measured sensor data indicative of one or more of a heart rate, a respiration rate, a blood pressure and an electrodermal activity of the user. As described above, these parameters change depending on the emotional status. For example, the heart rate, the respiration rate, the blood pressure and the electrodermal activity of a user increases when a user is stressed or anxious compared to a time instant at which the user is relaxed. For determining whether the user experiences negative emotions, the processing circuitry 120 may, e.g., be configured to extract one or more feature from the measured sensor data and determine whether the user experiences negative emotions based on the one or more extracted feature.
However, it is to be noted that the determination of the emotional status need not be based on physiological properties of the user only. According to examples of the present disclosure, contextual information may be used in addition. For example, the interface circuitry 110 may be configured to receive third data 103 indicative of contextual information about the scene in which the user and the plurality of people interact. The processing circuitry 120 may be configured to determine whether the user experiences negative emotions further based on the third data 103, i.e., based on the contextual information about the scene. The contextual information is any information indicating (describing) the context of the scene and hence supporting the understanding of the scene. For example, the contextual information may indicate a setting, a time, objects (items), actions, etc. of or in the scene. Alternatively or additionally, the contextual information may indicate information about the user (e.g., information about known stressors such as certain animals or other objects or situations causing stress). For example, the contextual information may indicated presence of objects such as specific animals in the scene that frighten the user. The presence of such objects in the scene makes it more likely that the user experiences negative emotions. Similarly, the day time and the setting may influence the emotional status of the user. For example, many human beings feel uncomfortable in dark streets. In case the scene plays in a scene at night time, it is more likely that the user experiences negative emotions. The above pieces of information are exemplary pieces of contextual information of the scene. Analysis
of the third data 103 in addition to the first data 101 may allow to improve the accuracy of determining whether the user experiences negative emotions.
For analyzing the first data 101 and optionally the third data 103, a trained machine-learning model may be used. In other words, the processing circuitry 120 may be configured to determine whether the user experiences negative emotions using a trained machine-learning model.
The machine-learning model is a data structure and/or set of rules representing a statistical model that the processing circuitry 120 uses to determine whether the user experiences negative emotions without using explicit instructions, instead relying on models and inference. The data structure and/or set of rules represents learned knowledge (e.g. based on training performed by a machine-learning algorithm as described above and below). In machinelearning, instead of a rule-based transformation of data, a transformation of data may be used, that is inferred from an analysis of training data.
The machine-learning model is trained by a machine-learning algorithm. The term "machine-learning algorithm" denotes a set of instructions that are used to create, train or use a machine-learning model. For the machine-learning model to determine whether the user experiences negative emotions, the machine-learning model may be trained using training data such as known physical properties of the user as input and predefined emotional stati of the user (e.g., relaxed, stressed, anxious, user experiences negative emotions, user does not experience negative emotions) as target output. This may allow to train a personalized machine-learning model, which is personalized to the user. In other example, training data of other users (people) may be used in addition to or instead of the training data of the user together with emotional stati of the other users (people) to train a (more) generic machinelearning model, which may allow earlier predictions. By training the machine-learning model with a large set of training data and associated training content information, the machine-learning model "learns" to determine whether the user experiences negative emotions in the training data, so that a target determination whether the user experiences negative emotions is obtained using the machine-learning model. By training the machine-learning model using training physical properties of the user and desired emotional stati of the user, the machine-learning model "learns" a transformation between the physical properties of the
user and the desired output, which can be used to provide an output based on non-training physical properties of the user provided to the machine-learning model.
The machine-learning model may be trained using training input data (e.g. known physical properties of the user). For example, the machine-learning model may be trained using a training method called "supervised learning". In supervised learning, the machine-learning model is trained using a plurality of training samples, wherein each sample may comprise a plurality of input data values, and a plurality of desired output values, i.e., each training sample is associated with a desired output value. By specifying both training samples and desired output values, the machine-learning model "learns" which output value to provide based on an input sample that is similar to the samples provided during the training. For example, a training sample may comprise one or more physical property of the user as input data and a desired emotional status of the user (i.e., either user experiences negative emotions or user does not experience negative emotions) as desired output data.
Apart from supervised learning, semi-supervised learning may be used. In semi-supervised learning, some of the training samples lack a corresponding desired output value. Supervised learning may be based on a supervised learning algorithm (e.g. a classification algorithm or a similarity learning algorithm). Classification algorithms may be used as the desired outputs of the trained machine-learning model are restricted to a limited set of values (categorical variables), i.e., the input is classified to one of the limited set of values (e.g., user experiences negative emotions or user does not experience negative emotions). Similarity learning algorithms are similar to classification algorithms but are based on learning from examples using a similarity function that measures how similar or related two objects are.
Apart from supervised or semi-supervised learning, unsupervised learning may be used to train the machine-learning model. In unsupervised learning, (only) input data are supplied and an unsupervised learning algorithm is used to find structure in the input data such as training physical properties of the user (e.g. by grouping or clustering the input data, finding commonalities in the data). Clustering is the assignment of input data comprising a plurality of input values into subsets (clusters) so that input values within the same cluster are similar according to one or more (pre-defined) similarity criteria, while being dissimilar to input
values that are included in other clusters. The input data for the unsupervised learning may be training physical properties of the user.
Reinforcement learning is a third group of machine-learning algorithms. In other words, reinforcement learning may be used to train the machine-learning model. In reinforcement learning, one or more software actors (called "software agents") are trained to take actions in an environment. Based on the taken actions, a reward is calculated. Reinforcement learning is based on training the one or more software agents to choose the actions such that the cumulative reward is increased, leading to software agents that become better at the task they are given (as evidenced by increasing rewards).
Furthermore, additional techniques may be applied to some of the machine-learning algorithms. For example, feature learning may be used. In other words, the machine-learning model may at least partially be trained using feature learning, and/or the machine-learning algorithm may comprise a feature learning component. Feature learning algorithms, which may be called representation learning algorithms, may preserve the information in their input but also transform it in a way that makes it useful, often as a pre-processing step before performing classification or predictions. Feature learning may be based on principal components analysis or cluster analysis, for example.
For example, the machine-learning model may be an Artificial Neural Network (ANN). ANNs are systems that are inspired by biological neural networks, such as can be found in a retina or a brain. ANNs comprise a plurality of interconnected nodes and a plurality of connections, so-called edges, between the nodes. There are usually three types of nodes, input nodes that receiving input values (e.g., the respective physical properties of the user), hidden nodes that are (only) connected to other nodes, and output nodes that provide output values (e.g., emotional status of the user such as user experiences negative emotions or user does not experience negative emotions). Each node may represent an artificial neuron. Each edge may transmit information from one node to another. The output of a node may be defined as a (non-linear) function of its inputs (e.g. of the sum of its inputs). The inputs of a node may be used in the function based on a "weight" of the edge or of the node that provides the input. The weight of nodes and/or of edges may be adjusted in the learning process. In other words, the training of an ANN may comprise adjusting the weights of the nodes and/or edges of the ANN, i.e., to achieve a desired output for a given input.
Alternatively, the machine-learning model may be a support vector machine, a random forest model or a gradient boosting model. Support vector machines (i.e. support vector networks) are supervised learning models with associated learning algorithms that may be used to analyze data (e.g. in classification or regression analysis). Support vector machines may be trained by providing an input with a plurality of training input values (e.g., physical properties of the user) that belong to one of two categories (e.g., emotional status of the user such as user experiences negative emotions or user does not experience negative emotions). The support vector machine may be trained to assign a new input value to one of the two categories. Alternatively, the machine-learning model may be a Bayesian network, which is a probabilistic directed acyclic graphical model. A Bayesian network may represent a set of random variables and their conditional dependencies using a directed acyclic graph. Alternatively, the machine-learning model may be based on a genetic algorithm, which is a search algorithm and heuristic technique that mimics the process of natural selection.
In some examples, the machine-learning model may be a combination of the above examples.
For example, features of various physiological parameters such as pupil, gaze, heart rate or electrodermal activity may be taken (extracted) from the first data 101 and input into the trained machine-learning model to determine whether the user experiences negative emotions with high accuracy. In some example, the trained machine-learning model may be understood as a negative emotions detection algorithm, in particular as a stress detection or anxiety detection algorithm.
The detection that the user experiences the negative emotions may optionally be used to further train the trained machine-learning model. For example, if it is determined that the user experiences negative emotions, the processing circuitry 120 may be configured to determine at least one of contextual information about a region and/or object in the scene causing the negative emotions of the user and information about one or more physiological property of the user while the user experiences the negative emotions. For example, the processing circuitry 120 may be configured to determine a viewing focus of the user while the user experiences the negative emotions to determine the region and/or object in the scene causing the negative emotions of the user. Then contextual information such as a type of the
object, a content of the region, a distance to the object or region, etc. may be determined contextual information about the region and/or object in the scene causing the negative emotions of the user. Alternatively or additionally, the processing circuitry 120 may be configured to determine information such as a setting or a time of the scene, a number of people present in the scene or the region, distances of the user to the people present in the scene or the region, etc. as contextual information about the region and/or object in the scene causing the negative emotions of the user. The processing circuitry 120 may additionally be configured to further train the trained machine-learning model based on the at least one of the contextual information about the region and/or object in the scene causing the negative emotions of the user and the information about the one or more physiological property of the user while the user experiences the negative emotions. Accordingly, the trained machinelearning model may “learn” further correlations between objects, etc. causing negative emotions of the user and the corresponding effects/patterns in the one or more physiological property of the user. In this manner, the accuracy of the trained machine-learning model may be further improved. In particular, the evaluation of contextual information of the scene such as the contextual information included in the third data 103 may be improved by the further training of the trained machine-learning model.
It is to be noted that the processing circuitry 120 may be configured to continuously determine whether the user experiences negative emotions. For example, the processing circuitry 120 may be configured to determine whether the user experiences negative emotions at regular time intervals. In other example, the processing circuitry 120 may be configured to sporadically determine whether the user experiences negative emotions. For example, the processing circuitry 120 may be configured to determine whether the user experiences negative emotions if certain patterns or features are detected in the one or more physiological property by the processing circuitry 120 (e.g., in case the processing circuitry 120 monitors the respective characteristic of the one or more physiological property indicated by the first data 101).
As described above, various techniques for determining a viewing focus based on viewing directions of people exist. Some exemplary techniques will be described in the following with reference to Fig. 3 and Fig. 4. However, it is to be noted that the present disclosure is not limited thereto.
Fig- 3 illustrates an exemplary scene 300 to highlight how a viewing focus may be determined. Five people 301, 305 are illustrated in Fig. 3. However, it is to be noted that the technique works for any other plurality of persons as well. The user is not illustrated in Fig. 3 for reasons of simplicity. The viewing directions of the people 301, ..., 305 at a first time instant ti are indicated by the arrows 361-1, ..., 361-5. The viewing directions of the people 301, ..., 305 at a second time instant t2 are indicated by the arrows 362-1, ..., 362-5. The viewing directions of the people 301, ..., 305 for a third time instant t3 are indicated by the arrows 363-1, ..., 363-5. The second time instant t2 succeeds first time instant ti. The third time instant t3 succeeds second time instant t2.
A plurality of trees 310 are present in the scene as well as a bird 320, some flowers 330 and a fountain 340.
At the first time instant ti, the first person 301 is looking at a dog 350 playing at the fountain 340 as indicated by arrow 361-1. The second person 302, the third person 303, the fourth person 304 and the fifth person 305 are looking at a jogger 360 as indicated by arrows 361- 2, 361-3, 361-4 and 361-5.
At the second time instant t2, the first person 301, the second person 302 and the fifth person 305 are looking at the dog 350 as indicated by arrows 362-1, 32-2 and 362-5. The third person 303 is looking at the fourth person 304 as indicated by arrow 362-3. The fourth person 304 is looking at the jogger 360 as indicated by arrow 362-4.
At the third time instant t3, the first person 301 is looking at a dog 340 as indicated by arrow 363-1. The second person 302 is looking at the first person 301 as indicated by arrow 363-2. The third person 303 and the fourth person 304 are looking at each other as indicated by arrows 363-3 and 363-4. The fifth person 361-5 is again looking at the jogger 360 as indicated by arrow 363-5.
As described above, the processing circuitry 120 receives the second data 102 which are indicative of the viewing directions of the people 301, ..., 305. In other words, the second data 102 indicate the viewing directions for the time instants ti, t2 and ti as indicated by the arrows in Fig. 3. Based on the second data 102, the processing circuitry 120 may be configured to determine one or more region and/or object in the scene at which the people are
looking for the respective time instant. For example, for the first time instant ti, the dog 350 and the jogger 360 are determined as objects at which the people 301, . . ., 305 are looking.
The processing circuitry 120 may be further configured to determine a respective attention rating for the determined one or more region and/or object in the scene. The attention rating is a classification according to the attention grade of the people 301, ..., 305. In other words, a value or score is determined for the one or more region and/or object in the scene which describes the level of attention the one or more region and/or object in the scene gets from the people 301, ..., 305. That is, the attention rating is a measure of how much attention the determined one or more region and/or object in the scene gets from the people 301, ..., 305. For example, a high attention rating may indicate that the determined one or more region and/or object in the scene gets a lot of attention from the people 301, . . ., 305, whereas a low attention rating may indicate that the determined one or more region and/or object in the scene gets only little attention from the people 301, . . ., 305.
Furthermore, the processing circuitry 120 may be configured to determine one or more of the one or more region and/or object having the highest attention rating as the one or more viewing focus of the plurality of people. Various criteria for selecting one or more of the one or more region and/or object as the one or more viewing focus of the plurality of people may be used. For example, a predetermined number of the one or more region and/or object having the highest attention may be determined as the one or more viewing focus of the plurality of people. Alternatively or additionally, the region(s) and/or object(s) having the highest attention rating, wherein the attention ratio is above a threshold, may be determined as the one or more viewing focus of the plurality of people. However, the present disclosure is not limited thereto.
Various approaches or techniques may be used for determining a respective attention rating for the determined one or more region and/or object in the scene at which the people are looking. For example, the processing circuitry 120 may be configured to determine, based on the second data 102, at least one of a respective number of times and a respective accumulated time duration the plurality of people look at a respective one of the one or more region and/or object. For the first time instant ti, the processing circuitry 120 may, e.g., determine how many times the people 301, ..., 305 look at the dog and how long the people 301, . . ., 305 look at dog. As in indicated in Fig. 3, the first person 301 is looking at the dog
350 at the first time instant ti such that it is determined that the plurality of people 301, . . 305 look one time at the dog 350 for an accumulated time tdog, ti. The second person 302, the third person 303, the fourth person 304 and the fifth person 305 are looking at the jogger 360 at the first time instant ti such that it is determined that the plurality of people 301, . . ., 305 look four times at the jogger 360 for an accumulated time tjogger, ti. Similarly, for the second time instant t2, the processing circuitry 120 may, e.g., determine how many times the people 301, . . ., 305 look at the dog and how long the people 301, . . ., 305 look at dog. As in indicated in Fig. 3, the first person 301, the second person 302 and the fifth person 305 are looking at the dog 350 at the second time instant t2 such that it is determined that the plurality of people 301, . . ., 305 look three times at the dog 350 for an accumulated time tdog,t2. The third person 303 is looking at the fourth person 304 at the second time instant t2 such that it is determined that the plurality of people 301, . . ., 305 look one time at the fourth person 304 for an accumulated time tperson4, t2. The fourth person 304 is looking at the jogger 360 at the second time instant t2 such that it is determined that the plurality of people 301, . . ., 305 look one time at the jogger 360 for an accumulated time tj0gger,t2. The same may be performed for the third time instant i by the processing circuitry 120.
The processing circuitry 120 may be further configured to determine the respective attention rating for the one or more region based on the at least one of the respective number of times and the respective accumulated time duration the plurality of people look at a respective one of the one or more region and/or object. For example, for the second time instant t2, the processing circuitry 120 may determine a respective attention rating for the dog 350, the fourth person 304 and the jogger 360 as the people look at these elements in the scene. The respective attention may be determined based on the number of times and the respective accumulated time duration the plurality of people look at the respective element in the scene. For example, the more often the people in the scene look at an element in the scene, the higher may be attention rating for the element. Similarly, the longer the accumulated time duration the plurality of people in the scene look at an element in the scene, the higher may be attention rating for the element. For the first time instant ti, processing circuitry 120 may, e.g., determine that the plurality of people 301, ..., 305 look one time at the dog 350 for an accumulated time tdog, ti and that the plurality of people 301, ..., 305 look four times at the jogger 360 for an accumulated time tjogger, ti. Accordingly, the processing circuitry 120 may further determine an attention rating for the dog 350 and the jogger 360. As most people look at the jogger 360, the attention rating for the jogger 360 may be high, whereas the at-
tention rating for the dog 350 may be lower. For the second time instant t2, the processing circuitry 120 may, e.g., determine an attention rating for the dog 350, the fourth person 304 and the jogger 360. As most people look at the dog 350, the attention rating for the dog 350 may be high, whereas the attention ratings for the fourth person 304 and the jogger 360 may be lower. The same may be performed for the third time instant t3 by the processing circuitry 120.
In accordance with the above explanations, the jogger 360 may be determined as viewing focus of the plurality of people 301, ..., 305 for the first time instant ti by the processing circuitry 120 and the dog 350 may be determined as viewing focus of the plurality of people 301, . . ., 305 for the second time instant t2. In other words, the processing circuitry 120 may create a heat map highlighting the one or more region and/or object in the scene that attract the attention of the plurality of people 301, . . ., 305.
The viewing directions of the plurality of people in the scene indicated by the second data 102 as well the analysis results may further be used for an enhanced interpretation of the scene. This is exemplary illustrated in Fig. 4 for the scene 300.
The first person 301 is looking at the dog 350 at all three instants ti, t2 and t3. Due to the first person 301’s constant attention on the dog 350, the processing circuitry 120 may determine that the first person 301 is probably the owner of the dog 350 and hence better understand the scene 300.
According to examples, the processing circuitry 120 may take into account further data for determining the respective attention rating for the determined one or more region and/or object in the scene. For example, the interface circuitry 110 may be further configured to determine the respective attention rating for the one or more region and/or object based on the third data 103, i.e., based on the contextual information of the scene. The contextual information may, e.g., indicate that the jogger 360 is running in the scene. Moving objects attract more attention than static objects. By considering such contextual information, the determination of the respective attention rating may be more accurate. For example, the attention rating for the jogger 360 may be increased for the first time instant ti as the contextual data indicate that the jogger is running (moving) in the scene 300.
As indicated above by the three time instants ti to t3, the processing circuitry 120 may be configured to continuously determine the respective attention rating for the one or more region and/or object. Continuously determining the respective attention rating allows to adapt the respective attention rating to the current state of the scene. According to examples, the processing circuitry 120 may be configured to determine the respective attention rating for the one or more region and/or object based on the attention rating for the one or more region and/or object determined for one or more previous time instant. Similarly, the processing circuitry 120 may be configured to determine the respective attention rating for the one or more region and/or object based on the current viewing directions of the plurality of people and viewing directions of the plurality of people for one or more previous time instant. More people look at the dog 350 at the time instant t2 than at the time instant ti. Accordingly, an attention shift towards the dog 350 takes place and the dog 350 becomes more interesting for the people 301, ..., 305. For example, the processing circuitry 120 may determine the attention ratio for the dog 350 for the second time instant t2 based on the attention ratio for the dog 350 for the first time instant ti. Accordingly, the development of the scene may be taken into account to increase the accuracy of the attention ratio determination.
At the third time instant t3, the second person 302 looks at the first person 301 after both looked at the dog 350 at the second time instant t2. Due to the history of the viewing directions of the first person 301 and the second person 302, the processing circuitry 120 may determine that the second person 302 is interested in a communication with the first person 301 about the dog 350 and hence better understand the scene 300.
At the third time instant t3, the third person 303 and the fourth person 304 look at each other after both looked at the jogger 360 at the first time instant ti. The third person 303 already looks at the fourth person 304 at the second time instant t2, while the fourth person 304 still looks at the jogger 360 at the second time instant t2. Due to the history of the viewing directions of the third person 303 and the fourth person 304, the processing circuitry 120 may determine that the third person 303 and the fourth person 304 are interested in a communication about the jogger 360 and hence better understand the scene 300.
The processing circuitry 120, may hence be configured to determine intentions of the people in the scene and determine, e.g., a respective probability for the determined intention. The
determined intentions as well as the probabilities for the determined intentions may be used for determining the respective attention ratio for the one or more object and/or region in the scene 300. According to examples, determined intentions as well as the probabilities for the determined intentions may further be used as contextual information of the scene for determining whether the user experiences negative emotions.
Adding contextual information to the heat map with the attention rating may allow to create a storyline (e.g., people’s attention drifts from the sunset to a swarm of birds because moving objects cause instantaneous attention and then back to the sunset) which can help in a CBT to explain why people focus attention on specific sequences of objects. It might also learn the input from the attention model to predict which scene will cause the most attention in the future in healthy user and in users with mental diseases.
In some examples, the processing circuitry 120 is configured to determine the at least one viewing focus of the plurality of people using a trained machine-learning model. The trained machine-learning model for determining the at least one viewing focus of the plurality of people in the scene may be integrated into the trained machine-learning model for determining whether the user experiences negative emotions or be a separate machine-learning model. The machine learning model may be trained analogously to what is described above. For example, for the machine-learning model to determine the at least one viewing focus, the machine-learning model may be trained using training data indicative of viewing directions of a plurality of people in a scene as input and one or more predefined viewing focus as target output. However, as described also other training methods such as semi-supervised learning, unsupervised learning or reinforcement learning may be used.
For example, the viewing directions of the plurality of people in the scene may be taken (extracted) from the second data 102 and input into the trained machine-learning model to determine the at least one viewing focus of the plurality of people. In some examples, the trained machine-learning model may be understood as an attention model. For example, the trained machine-learning model may be used analogously to what is described above to create a heatmap where objects or regions on the screen/scene get rated with different attention inputs. As described above, the rating may depend on how often and how long people look at the specific object or region. Furthermore, further features such as determined intentions of the people in the scene and probabilities for the determined intentions may be used. The
attention rating value may change over time such that a continuous update of the rating allows to determine at which time instant the specific object or region caused attraction.
The first data 101, the second data 102 and the third data 103 described above may be provided from various sources. For example, the first data 101 may at least in part be received from the user equipment of the user. The second data 102 may, e.g., at least in part be received from one or more user equipment of the plurality of people. Alternatively or additionally, in case the scene is a virtual reality scene, at least part of the first data 101 and/or at least part of the second data 102 may be received from a computing hardware (e.g., a server or a computing cloud) hosting the virtual world that includes the virtual reality scene. In case the scene is an augmented reality scene, at least part of the first data 101 and/or at least part of the second data 102 may be received from a computing hardware (e.g., a server or a computing cloud) providing the virtual objects augmented to the real world. Also the third data 103 may at least in part be received from the respective computing platform. Alternatively or additionally, as described above, at least part of the third data 103 may be generated by the processing circuitry 120 itself while performing viewing direction guidance according to the present disclosure.
In case the scene is a virtual reality scene, the apparatus 100 may, e.g., be part of or be coupled to a computing hardware (e.g., a server or a computing cloud) hosting the virtual world that includes the virtual reality scene. Similarly, case the scene is an augmented reality scene, the apparatus 100 may, e.g., be part of or be coupled to a computing hardware (e.g., a server or a computing cloud) providing the virtual objects augmented to the real world.
Fig. 5 schematically illustrates an exemplary user equipment 500 for the user.
The user equipment 500 comprises one or more sensor 510 configured to measure one or more physiological property of the user while the user is interacting in the same scene as a plurality of people. As described above, the one or more physiological property may be one or more of measured eye-tracking data, measured gaze-tracking data of the user, a heart (pulse) rate, a heart rate variability of the user, a respiration rate, a blood pressure and an electrodermal activity of the user. Accordingly, the one or more sensor 510 may be or be configured to perform the functionalities of one or more of an eye-tracking sensor, a gazetracking sensor, a Galvanic Skin Response (GSR) sensor, a PhotoPlethysmoGraphy (PPG)
sensor, a Laser Doppler Flowmetry (LDF) sensor, and an ElectroMyoGraphy (EMG) sensor.
Additionally, the user equipment 500 comprises interface circuitry 520 coupled to the one or more sensor 510. The interface circuitry 520 is configured to send data 501 indicative of (representing, encoded with) the measured one or more physiological property of the user to an external device 550 (i.e., a device external to / separate from the user equipment 500). For example, the external device 550 may be the apparatus 100 described above. If the scene is a virtual reality scene, the external device 550 may, e.g., be a computing hardware hosting the virtual world that includes the virtual reality scene. If the scene is an augmented reality scene, the external device 550 may, e.g., be a computing hardware providing the virtual objects augmented to the real world. The respective computing hardware may forward the data 501 to, e.g., the apparatus 100.
The interface circuitry 520 is further configured to receive, from the external device 550, control data 502 encoded with one or more commands for controlling the user equipment 500 to perform output for guiding the viewing direction of the user toward a determined viewing focus of the plurality of people in the scene. The control data 502 may, e.g., be determined by the apparatus 100 as described above.
The user equipment 500 further comprises a human-machine interface 530 configured to perform output for guiding the viewing direction of the user toward the viewing focus based on the control data 502.
For example, the human-machine interface 530 may comprise at least one of optical output means (structure, device) such as a display or a projector, an acoustical output means (structure, device) such as a loudspeaker or an earphone, and an haptic output means (structure, device) such as a vibrating device. For example, the graphical output may be one or more symbol or graphic element for drawing the user’s attention to the viewing focus (e.g., a text, an arrow or a frame highlighting the at least one viewing focus). The audio output may be one or more sounds or voice commands for drawing the user’s attention to the viewing focus. The haptic output may be one or more mechanical outputs by the human-machine interface 530 perceivable by the user (e.g., motion, vibration, pressure) for drawing the user’s attention to the viewing focus.
The user equipment 500 may allow to guide the user’s viewing direction toward the viewing focus based on the control data 502. In particular, the user equipment 500 may be used together with the apparatus 100 to make a user understand that the people in the scene are not focusing on him/her but on one or more other object or region in the scene. The user equipment 500 may, e.g., be a head-mounted equipment such as a virtual reality headset or smart glasses.
In an augmented reality scene not each person in the scene necessarily wears or uses a respective user equipment that allows to measure the viewing direction of the person. Therefore, the user equipment 500 may optionally further comprise one or more further sensor 540 configured to measure viewing directions of (at least part of) the plurality of people in the scene. For example, the one or more further sensor 540 may be or be configured to perform the functionalities of one or more of a camera, an eye-tracking sensor and a gazetracking sensor. Accordingly, the interface circuitry 520 may be further configured to send data 503 indicative of the measured viewing directions of the plurality of people to the external device 550. In this configuration, the user equipment 500 may allow to collect viewing directions of the people in the vicinity of the user such that the proposed viewing direction guidance may be performed even in case the other people in the scene do not wear or use a respective user equipment that allows to measure the viewing direction of the person.
In alternative examples, the interface circuitry 520 may be configured to receive from user equipments 560 of at least part of the other people in the scene data 504 indicative of measured viewing directions of the plurality of people in the scene. In the example of Fig. 5 one user equipment of another person in the scene is illustrated exemplarily. However, it is to be noted that the interface circuitry may receive data 504 indicative of measured viewing directions from plural other user equipments 560. Analogously to what is described above, the interface circuitry 520 may be further configured to send data 503 indicative of the measured viewing directions of the plurality of people to the external device 550. Further, the interface circuitry 520 may be configured to send data 505 indicative of the one or more physiological property of the user to the user equipments 560 of at least part of the people in the scene. The other user equipments 560 may, e.g., process the data 505 or forward the data 505 to the external device 550.
For further highlighting the viewing direction guidance described above, Fig. 6 illustrates a flowchart of a (e.g., computer-implemented) method 600 for guiding a viewing direction of a user. The method 600 comprises receiving 602 first data indicative of a physiological property of the user. In addition, the method 600 comprises receiving 604 second data indicative of viewing directions of a plurality of people. The plurality of people are interacting in the same scene as the user. The method 600 further comprises determining 606 at least one viewing focus of the plurality of people in the scene based on the second data. Additionally, the method 600 comprises determining 608 whether the user experiences negative emotions based on the first data. If it is determined that the user experiences negative emotions, the method 600 comprises causing 610 an output for guiding the viewing direction of the user toward one of the at least one viewing focus of the plurality of people in the scene by a user equipment of the user.
Analogously to what is described above, the method 600 may, e.g., enable the user to understand that the people are not focusing on him but on one or more other object or region in the scene. As the user is enabled to recognize that he/she is not the other people’s object or region of attention in the scene, the method 600 may allow to reduce the experience of negative emotions. In particular, the method 600 may be used for CBT as the method 600 may allow to make the user aware that his/her perception of the scene, which is based his/her inner beliefs, is erroneous.
More details and aspects of the method 600 are explained in connection with the proposed technique or one or more examples described above (e.g. Fig. 1 to Fig. 5). The method 600 may comprise one or more additional optional features corresponding to one or more aspects of the proposed technique or one or more examples described above. For example, if it is not determined that the user experiences negative emotions, the method 600 may comprise not causing 610 an output for guiding the viewing direction of the user as the user’s perception need not be changed. Accordingly, confusion or unnecessary disturbance of the user may be avoided.
The following examples pertain to further embodiments:
(1) An apparatus for guiding a viewing direction of a user, the apparatus comprising: interface circuitry configured to:
receive first data indicative of one or more physiological property of the user; and receive second data indicative of viewing directions of a plurality of people, the plurality of people interacting in the same scene as the user; and processing circuitry configured to: determine at least one viewing focus of the plurality of people in the scene based on the second data; determine whether the user experiences negative emotions based on the first data; and if it is determined that the user experiences negative emotions, cause an output for guiding the viewing direction of the user toward one of the at least one viewing focus of the plurality of people in the scene by a user equipment of the user.
(2) The apparatus of (1), wherein the first data comprises at least one of measured eyetracking data and measured gaze-tracking data of the user, and wherein, for determining whether the user experiences negative emotions, the processing circuitry is configured to: extract one or more feature from the at least one of the measured eye-tracking and the measured gaze-tracking data of the user; and determine whether the user experiences negative emotions based on the one or more extracted feature.
(3) The apparatus of (1) or (2), wherein the first data comprises at least measured sensor data indicative of one or more of a heart rate, a respiration rate, a blood pressure and an electrodermal activity of the user, and wherein, for determining determine whether the user experiences negative emotions, the processing circuitry is configured to: extract one or more feature from the measured sensor data; and determine whether the user experiences negative emotions based on the one or more extracted feature.
(4) The apparatus of any one of (1) to (3), wherein the interface circuitry is configured to receive third data indicative of contextual information about the scene, and wherein the processing circuitry is configured to determine whether the user experiences negative emotions further based on the third data.
(5) The apparatus of any one of (1) to (4), wherein the processing circuitry is configured to determine whether the user experiences negative emotions using a trained machinelearning model.
(6) The apparatus of (5), wherein, if it is determined that the user experiences negative emotions, the processing circuitry is further configured to: determine at least one of contextual information about a region and/or object in the scene causing the negative emotions of the user and information about one or more physiological property of the user while the user experiences the negative emotions; and further train the trained machine-learning model based on the at least one of the contextual information about the region and/or object in the scene causing the negative emotions of the user and the information about the one or more physiological property of the user while the user experiences the negative emotions.
(7) The apparatus of any one of (1) to (6), wherein the first data comprises at least one of measured eye-tracking data and measured gaze-tracking data of the plurality of people.
(8) The apparatus of any one of (1) to (7), wherein, for determining the at least one viewing focus of the plurality of people, the processing circuitry is, based on the second data, configured to: determine one or more region and/or object in the scene at which the people are looking; determine a respective attention rating for the one or more region and/or object; and determine one or more of the one or more region and/or object having the highest attention rating as the one or more viewing focus of the plurality of people.
(9) The apparatus of (8), wherein, for determining a respective attention rating for the one or more region and/or object, the processing circuitry is configured to: determine, based on the second data, at least one of a respective number of times and a respective accumulated time duration the plurality of people look at a respective one of the one or more region and/or object; and determine the respective attention rating for the one or more region based on the at least one of the respective number of times and the respective accumulated time duration the plurality of people look at a respective one of the one or more region and/or object.
(10) The apparatus of (8) or (9), wherein the interface circuitry is further configured to receive third data indicative of contextual information of the scene, and wherein the processing circuitry is further configured to determine the respective attention rating for the one or more region and/or object based on the contextual information of the scene.
(11) The apparatus of any one of (8) to (10), wherein the processing circuitry is configured to continuously determine the respective attention rating for the one or more region and/or object.
(12) The apparatus of any one of (1) to (11), wherein the processing circuitry is configured to determine the at least one viewing focus of the plurality of people using a trained machine-learning model.
(13) The apparatus of any one of (1) to (12), wherein the scene is a virtual reality scene.
(14) The apparatus of any one of (1) to (12), wherein the scene is an augmented reality scene.
(15) The apparatus of any one of (1) to (14), wherein, for causing the output by the user equipment, the processing circuitry is configured to generate control data for the user equipment, the control data being encoded with one or more commands for controlling the user equipment to perform the output, and wherein the interface circuitry is configured to send the control data to the user equipment.
(16) The apparatus of any one of (1) to (15), wherein the processing circuitry is configured to cause at least one of a graphic output, an audio output and a haptic output by the user equipment if it is determined that the user experiences negative emotions.
(17) A user equipment for a user, comprising: one or more sensor configured to measure one or more physiological property of the user while the user is interacting in the same scene as a plurality of people; interface circuitry configured to: send data indicative of the one or more physiological property of the user to an external device; and
receive, from the external device, control data encoded with one or more commands for controlling the user equipment to perform output for guiding the viewing direction of the user toward a determined viewing focus of the plurality of people in the scene; and a human-machine interface configured to perform output for guiding the viewing direction of the user toward the viewing focus based on the control data.
(18) The user equipment of (17), further comprising one or more further sensor configured to measure viewing directions of the plurality of people in the scene, wherein the interface circuitry is further configured to send data indicative of the measured viewing directions of the plurality of people to the external device.
(19) The user equipment of (17), wherein the interface circuitry is configured to receive from user equipments of at least part of the people data indicative of measured viewing directions of the plurality of people in the scene, and wherein the interface circuitry is configured to send data indicative of the one or more physiological property of the user to the user equipments of at least part of the people.
(20) A method for guiding a viewing direction of a user, the method comprising: receiving first data indicative of a physiological property of the user; receiving second data indicative of viewing directions of a plurality of people, the plurality of people interacting in the same scene as the user; determining at least one viewing focus of the plurality of people in the scene based on the second data; determining whether the user experiences negative emotions based on the first data; and if it is determined that the user experiences negative emotions, causing an output for guiding the viewing direction of the user toward one of the at least one viewing focus of the plurality of people in the scene by a user equipment of the user.
(21) A non-transitory machine-readable medium having stored thereon a program having a program code for performing the method according to (20), when the program is executed on a processor or a programmable hardware.
(22) A program having a program code for performing the method according to (20), when the program is executed on a processor or a programmable hardware.
The aspects and features described in relation to a particular one of the previous examples may also be combined with one or more of the further examples to replace an identical or similar feature of that further example or to additionally introduce the features into the further example.
Examples may further be or relate to a (computer) program including a program code to execute one or more of the above methods when the program is executed on a computer, processor or other programmable hardware component. Thus, steps, operations or processes of different ones of the methods described above may also be executed by programmed computers, processors or other programmable hardware components. Examples may also cover program storage devices, such as digital data storage media, which are machine-, processor- or computer-readable and encode and/or contain machine-executable, processorexecutable or computer-executable programs and instructions. Program storage devices may include or be digital storage devices, magnetic storage media such as magnetic disks and magnetic tapes, hard disk drives, or optically readable digital data storage media, for example. Other examples may also include computers, processors, control units, (field) programmable logic arrays ((F)PLAs), (field) programmable gate arrays ((F)PGAs), graphics processor units (GPU), ASICs, integrated circuits (ICs) or system-on-a-chip (SoCs) systems programmed to execute the steps of the methods described above.
It is further understood that the disclosure of several steps, processes, operations or functions disclosed in the description or claims shall not be construed to imply that these operations are necessarily dependent on the order described, unless explicitly stated in the individual case or necessary for technical reasons. Therefore, the previous description does not limit the execution of several steps or functions to a certain order. Furthermore, in further examples, a single step, function, process or operation may include and/or be broken up into several sub-steps, -functions, -processes or -operations.
If some aspects have been described in relation to a device or system, these aspects should also be understood as a description of the corresponding method. For example, a block, device or functional aspect of the device or system may correspond to a feature, such as a
method step, of the corresponding method. Accordingly, aspects described in relation to a method shall also be understood as a description of a corresponding block, a corresponding element, a property or a functional feature of a corresponding device or a corresponding system.
The following claims are hereby incorporated in the detailed description, wherein each claim may stand on its own as a separate example. It should also be noted that although in the claims a dependent claim refers to a particular combination with one or more other claims, other examples may also include a combination of the dependent claim with the sub- ject matter of any other dependent or independent claim. Such combinations are hereby explicitly proposed, unless it is stated in the individual case that a particular combination is not intended. Furthermore, features of a claim should also be included for any other independent claim, even if that claim is not directly defined as dependent on that other independent claim.
Claims
1. An apparatus for guiding a viewing direction of a user, the apparatus comprising: interface circuitry configured to: receive first data indicative of one or more physiological property of the user; and receive second data indicative of viewing directions of a plurality of people, the plurality of people interacting in the same scene as the user; and processing circuitry configured to: determine at least one viewing focus of the plurality of people in the scene based on the second data; determine whether the user experiences negative emotions based on the first data; and if it is determined that the user experiences negative emotions, cause an output for guiding the viewing direction of the user toward one of the at least one viewing focus of the plurality of people in the scene by a user equipment of the user.
2. The apparatus of claim 1, wherein the first data comprises at least one of measured eye-tracking data and measured gaze-tracking data of the user, and wherein, for determining whether the user experiences negative emotions, the processing circuitry is configured to: extract one or more feature from the at least one of the measured eye-tracking and the measured gaze-tracking data of the user; and determine whether the user experiences negative emotions based on the one or more extracted feature.
3. The apparatus of claim 1, wherein the first data comprises at least measured sensor data indicative of one or more of a heart rate, a respiration rate, a blood pressure and an electrodermal activity of the user, and wherein, for determining determine whether the user experiences negative emotions, the processing circuitry is configured to:
extract one or more feature from the measured sensor data; and determine whether the user experiences negative emotions based on the one or more extracted feature.
4. The apparatus of claim 1, wherein the interface circuitry is configured to receive third data indicative of contextual information about the scene, and wherein the processing circuitry is configured to determine whether the user experiences negative emotions further based on the third data.
5. The apparatus of claim 1, wherein the processing circuitry is configured to determine whether the user experiences negative emotions using a trained machine-learning model.
6. The apparatus of claim 5, wherein, if it is determined that the user experiences negative emotions, the processing circuitry is further configured to: determine at least one of contextual information about a region and/or object in the scene causing the negative emotions of the user and information about one or more physiological property of the user while the user experiences the negative emotions; and further train the trained machine-learning model based on the at least one of the contextual information about the region and/or object in the scene causing the negative emotions of the user and the information about the one or more physiological property of the user while the user experiences the negative emotions.
7. The apparatus of claim 1, wherein the first data comprises at least one of measured eye-tracking data and measured gaze-tracking data of the plurality of people.
8. The apparatus of claim 1, wherein, for determining the at least one viewing focus of the plurality of people, the processing circuitry is, based on the second data, configured to: determine one or more region and/or object in the scene at which the people are looking; determine a respective attention rating for the one or more region and/or object; and determine one or more of the one or more region and/or object having the highest attention rating as the one or more viewing focus of the plurality of people.
9. The apparatus of claim 8, wherein, for determining a respective attention rating for the one or more region and/or object, the processing circuitry is configured to: determine, based on the second data, at least one of a respective number of times and a respective accumulated time duration the plurality of people look at a respective one of the one or more region and/or object; and determine the respective attention rating for the one or more region based on the at least one of the respective number of times and the respective accumulated time duration the plurality of people look at a respective one of the one or more region and/or object.
10. The apparatus of claim 8, wherein the interface circuitry is further configured to receive third data indicative of contextual information of the scene, and wherein the processing circuitry is further configured to determine the respective attention rating for the one or more region and/or object based on the contextual information of the scene.
11. The apparatus of claim 8, wherein the processing circuitry is configured to continuously determine the respective attention rating for the one or more region and/or object.
12. The apparatus of claim 1, wherein the processing circuitry is configured to determine the at least one viewing focus of the plurality of people using a trained machine-learning model.
13. The apparatus of claim 1, wherein the scene is a virtual reality scene.
14. The apparatus of claim 1, wherein the scene is an augmented reality scene.
15. The apparatus of claim 1, wherein, for causing the output by the user equipment, the processing circuitry is configured to generate control data for the user equipment, the control data being encoded with one or more commands for controlling the user equipment to perform the output, and wherein the interface circuitry is configured to send the control data to the user equipment.
16. The apparatus of claim 1, wherein the processing circuitry is configured to cause at least one of a graphic output, an audio output and a haptic output by the user equipment if it is determined that the user experiences negative emotions.
17. A user equipment for a user, comprising:
one or more sensor configured to measure one or more physiological property of the user while the user is interacting in the same scene as a plurality of people; interface circuitry configured to: send data indicative of the one or more physiological property of the user to an external device; and receive, from the external device, control data encoded with one or more commands for controlling the user equipment to perform output for guiding the viewing direction of the user toward a determined viewing focus of the plurality of people in the scene; and a human-machine interface configured to perform output for guiding the viewing direction of the user toward the viewing focus based on the control data.
18. The user equipment of claim 17, further comprising one or more further sensor configured to measure viewing directions of the plurality of people in the scene, wherein the interface circuitry is further configured to send data indicative of the measured viewing directions of the plurality of people to the external device.
19. The user equipment of claim 17, wherein the interface circuitry is configured to receive from user equipments of at least part of the people data indicative of measured viewing directions of the plurality of people in the scene, and wherein the interface circuitry is configured to send data indicative of the one or more physiological property of the user to the user equipments of at least part of the people.
20. A method for guiding a viewing direction of a user, the method comprising: receiving first data indicative of a physiological property of the user; receiving second data indicative of viewing directions of a plurality of people, the plurality of people interacting in the same scene as the user; determining at least one viewing focus of the plurality of people in the scene based on the second data; determining whether the user experiences negative emotions based on the first data; and
if it is determined that the user experiences negative emotions, causing an output for guiding the viewing direction of the user toward one of the at least one viewing focus of the plurality of people in the scene by a user equipment of the user.
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| EP23155708 | 2023-02-09 | ||
| PCT/EP2024/052527 WO2024165421A1 (en) | 2023-02-09 | 2024-02-01 | Apparatus and method for guiding a viewing direction of a user, and user equipment |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP4662543A1 true EP4662543A1 (en) | 2025-12-17 |
Family
ID=85222002
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP24703165.1A Pending EP4662543A1 (en) | 2023-02-09 | 2024-02-01 | Apparatus and method for guiding a viewing direction of a user, and user equipment |
Country Status (2)
| Country | Link |
|---|---|
| EP (1) | EP4662543A1 (en) |
| WO (1) | WO2024165421A1 (en) |
Family Cites Families (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20140149177A1 (en) * | 2012-11-23 | 2014-05-29 | Ari M. Frank | Responding to uncertainty of a user regarding an experience by presenting a prior experience |
| US11022794B2 (en) * | 2018-12-27 | 2021-06-01 | Facebook Technologies, Llc | Visual indicators of user attention in AR/VR environment |
| US20240004464A1 (en) * | 2020-11-13 | 2024-01-04 | Magic Leap, Inc. | Transmodal input fusion for multi-user group intent processing in virtual environments |
-
2024
- 2024-02-01 EP EP24703165.1A patent/EP4662543A1/en active Pending
- 2024-02-01 WO PCT/EP2024/052527 patent/WO2024165421A1/en not_active Ceased
Also Published As
| Publication number | Publication date |
|---|---|
| WO2024165421A1 (en) | 2024-08-15 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| Fetsch et al. | Bridging the gap between theories of sensory cue integration and the physiology of multisensory neurons | |
| US11443645B2 (en) | Education reward system and method | |
| US20220254506A1 (en) | Extended reality systems and methods for special needs education and therapy | |
| Puviani et al. | A mathematical description of emotional processes and its potential applications to affective computing | |
| Kadyr et al. | Affective computing methods for simulation of action scenarios in video games | |
| CN118471438B (en) | Virtual digital image autism rehabilitation intervention system | |
| EP4662543A1 (en) | Apparatus and method for guiding a viewing direction of a user, and user equipment | |
| Taheri | Multimodal Multisensor attention modelling | |
| Reddy | The Cue-integrated sense of agency—as operationalized in experiments—is not a (Multisensory) perceptual effect, but is a judgment effect | |
| Akbari et al. | Distinct Holistic Processing Mechanisms for Faces and Line Patterns in DNNs | |
| JP2026033671A (en) | system | |
| JP2026033332A (en) | system | |
| Zhang | A Comprehensive Evaluation of User Experience in Eye-Controlled Interaction | |
| Kılıç | Analysis and Identification of Visual-Motor Integration in Human Motor Control | |
| JP2026018389A (en) | system | |
| JP2026038753A (en) | system | |
| JP2026024280A (en) | system | |
| JP2026024982A (en) | system | |
| JP2026018798A (en) | system | |
| JP2026025292A (en) | system | |
| JP2026033070A (en) | system | |
| JP2025051742A (en) | system | |
| Salous | User modeling for adaptation of cognitive systems | |
| JP2026061841A (en) | system | |
| JP2026030217A (en) | system |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: UNKNOWN |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE |
|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE |
|
| 17P | Request for examination filed |
Effective date: 20250909 |
|
| AK | Designated contracting states |
Kind code of ref document: A1 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR |