EP4453692A1 - Automatic and egocentric detection and recognition of tasks in gaze scenes - Google Patents
Automatic and egocentric detection and recognition of tasks in gaze scenesInfo
- Publication number
- EP4453692A1 EP4453692A1 EP22835037.7A EP22835037A EP4453692A1 EP 4453692 A1 EP4453692 A1 EP 4453692A1 EP 22835037 A EP22835037 A EP 22835037A EP 4453692 A1 EP4453692 A1 EP 4453692A1
- Authority
- EP
- European Patent Office
- Prior art keywords
- gaze
- participant
- characteristic
- scene
- objects
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F3/00—Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
- G06F3/01—Input arrangements or combined input and output arrangements for interaction between user and computer
- G06F3/011—Arrangements for interaction with the human body, e.g. for user immersion in virtual reality
- G06F3/013—Eye tracking input arrangements
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F3/00—Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
- G06F3/01—Input arrangements or combined input and output arrangements for interaction between user and computer
- G06F3/03—Arrangements for converting the position or the displacement of a member into a coded form
- G06F3/033—Pointing devices displaced or positioned by the user, e.g. mice, trackballs, pens or joysticks; Accessories therefor
- G06F3/038—Control and interface arrangements therefor, e.g. drivers or device-embedded control circuitry
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06Q—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES; SYSTEMS OR METHODS SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES, NOT OTHERWISE PROVIDED FOR
- G06Q10/00—Administration; Management
- G06Q10/06—Resources, workflows, human or project management; Enterprise or organisation planning; Enterprise or organisation modelling
- G06Q10/063—Operations research, analysis or management
- G06Q10/0631—Resource planning, allocation, distributing or scheduling for enterprises or organisations
Definitions
- the invention relates to the treatment of eye-tracking logs in the context of the analysis of the behavior of a participant in an open and uncontrolled environment.
- eye tracking In the field of industrial processes, the treatment of eye-tracking logs (or “eye tracking” in English) is becoming increasingly common, especially with the wider adoption of the ecosystem of connected objects and their connected users (often referred to as the "Internet of Things" or "IOT").
- IOT Internet of Things
- eye tracking presents advantages in respect of spatial and temporal resolution, the complexity of the operations and the revelation of movement (for example, movement of the head).
- the log is composed of two parts that are captured together: the position of the gaze (or “gaze input”) and the recording of what is being looked at (or “regarded scene”).
- known industrial processes work by attempting to control the latter, particularly by seeking to superpose the measured elements in real space with known elements in order to locate an object and determine precisely its spatial configuration.
- This treatment is applicable to controlled environments (for example, a screen or a scene with known references such as barcodes). In all these cases, the control of a situation in an uncontrolled environment remains unknown.
- This person performs one or more tasks with each task consisting of two types of information for action recognition: the verb representing the action, and the noun representing the object(s) related to this action (see “Learning Human-Object Interactions by Graph Parsing Neural Networks”, Qi, S.; Wang, W.; Jia, B.; Shen, J.; and Zhu, S.-C, Proceedings of the European Conference on Computer Vision (ECCV), 2018, pp. 401-417).
- the verb class concentrates on the classification of the actions that the participant performs (for example, “place” or “open”).
- the noun class serves to classify the object with which the participant is interacting.
- the predictions of both classes are generally merged without interactions for the action classification (for example, by means of CNN).
- physiological gaze measurements to be correlated with the physical environment, particularly in order to measure discrepancies in the use of connected objects during the participation of participants in open and uncontrolled environments (for example, tasks performed in a factory, experiments performed in laboratories, pathology monitoring with the search for crisis triggers, studies of customer behavior in supermarkets; and other tasks and activities performed in other uncontrolled environments).
- open and uncontrolled environments for example, tasks performed in a factory, experiments performed in laboratories, pathology monitoring with the search for crisis triggers, studies of customer behavior in supermarkets; and other tasks and activities performed in other uncontrolled environments.
- Even for a human it can be difficult to recognize an action by only looking at the objects while ignoring the participant’s intention, or by considering only the changes in movement without knowing the object in interaction. It would be useful to have an inference of the tasks and their associated activities in real time, in order to decide whether or not to interrupt an activity of the participant depending on its priority.
- the disclosed invention relates to a solution for dissociating the gaze input and the regarded scene in uncontrolled environments.
- the invention relates to a process that obtains the data input from the scene regarded by a participant, this input being captured by video during a navigation session performed by the participant.
- the problem addressed by the invention is that of automatic task recognition so as to recognize the associated activities in a predetermined context and the order of their performance.
- a first-person video recording system integrating the position of the gaze (for example, an eye-tracking system incorporating a wearable device associating the gaze with a physiological measurement and a forehead-mounted camera that captures the tasks from a first-person viewpoint), it becomes possible to identify useful measurement ranges, i.e., the start and the end of each activity, depending on the presence of characteristic objects associated with each task. Labelled periods of activity interruption are also considered to identify breaks and the resumption based upon completed tasks.
- the invention relates to a process for automatic and egocentric detection and recognition of tasks in gaze scenes of a participant in an uncontrolled environment, the process implemented by a computer system to recognize, by logging gaze scene data of the participant and objects associated with gaze scenes seen during a navigation session performed by the participant, tasks associated with a predetermined activity and the order of their performance, characterized in that the process includes the following steps: an eye-tracking acquisition step during which gaze scene data of the participant is collected using an eye-tracking system of the computer system that records the participant’s behavior captured in a gaze scene of a navigation session performed by the participant; a step of performance of a process of constructing a model for detecting characteristic objects on the basis of one or more influential parameters in the gaze scene of the participant; and a step of performance of a process of constructing on of a model for detecting constituent tasks of activities associated with the characteristic objects identified in the gaze scene; such that the computer system detects tasks performed based on the recognized characteristic objects output from the characteristic object detection model.
- the process of construction of the characteristic object detection model includes the following steps: a step of video acquisition of at least one navigation session performed by the participant, during which at least one video is acquired by the eye-tracking system, each video consisting of images of the gaze scene and the gaze input of the participant; a step of creation of a training database of characteristic objects including data corresponding to the gaze scene, during which the training database of characteristic objects is introduced into the characteristic object detection model; a step of classification of searched characteristic objects to identify searched characteristic objects that exist in images of the participant’s gaze scene; a step of annotation of images of searched characteristic objects to automatically detect boundaries between profiles of the characteristic objects and the gaze scene; a step of tracking characteristic objects classified during the classification step, allowing generation of a plurality of images from an annotation of images of searched characteristic objects performed during the annotation step; and a step of training the characteristic object detection model to recognize characteristic objects corresponding to the completion of the associated tasks, during which step a machinelearning process obtains data from the training database of characteristic objects; so that the computer system
- the process of construction of the task and activity detection model includes the following steps: a step of introduction of a task description database associated with the predetermined activity, wherein the task description database includes a description of the characteristic objects associated with each task according to the classification of the searched characteristic objects created during the classification step of the process of construction of the characteristic object detection model, according to an order of appearance of the objects, according to durations of appearance of the objects, and according to a co-occurrence of certain objects in the participant’s gaze scene; a step of chronological visualization of the tasks and the associated characteristic objects, during which a temporal synchronization is obtained between the participant’s gaze scene data and a navigation session performed by the participant; a step of creation of a training database of tasks including a reference database of tasks performed during a predetermined activity and a description of the characteristic objects associated with each task; a step of correction and validation of the presence of one or more characteristic objects associated with the participant’s gaze scene during the navigation session, this step including a step of comparison of the chronological visualization of the participant
- the temporal synchronization is obtained using temporal-synchronization points recorded by the eye-tracking system that correspond in the task detection model to time of a known task of the navigation session; so that the task and the temporal-synchronization points are used to synchronize the participant’s gaze scene data over time with the navigation session so that each gaze scene datum is mapped to a characteristic object at the appropriate time on a common clock.
- the task detection model uses the output of the object detection model in navigation sessions performed by the participant in order to place the coordinates of boxed regions, including bounding boxes surrounding the characteristic objects, in correlation with the coordinates of the gaze positions from the eyetracking system.
- the coordinates of the boxed regions are used to compute characteristics of a visual information stream in order to transmit the boxed regions and the characteristics of the visual information stream to the task detection model.
- the eye-tracking system includes at least one gaze input device coupled to the computer system that performs the process, the gaze input device including: a tracking component integrated into the gaze input device for tracking the participant’s gaze to generate the gaze scene data input from the participant; and at least one camera facing the physical environment around the participant’s gaze scene to capture still and video images incorporating the participant’s gaze scene.
- the process further includes a step of task-based measurement step to detect, based upon the recognized characteristic objects output from the object detection model, streams of interactions between the participant and the characteristic objects.
- the invention also relates to a computer system that performs the disclosed process, including: an eye tracking system including at least one gaze input device and a tracking component for tracking the participant’s gaze; one or more modules implemented by the computer system and executable by one or more processors, the one or more modules including: a receiving module configured to receive the gaze scene data input received from the gaze input device, the gaze scene data input representing a gaze input associated with at least one characteristic object associated with the participant’s gaze scene; an extraction module configured to extract a set of searched characteristic objects; and an analysis module configured to identify, based at least partially on the set of characteristic objects associated with the gaze scene, the particular characteristic object associated with a task in order to recognize one or more associated activities in a predetermined order.
- the gaze input device includes glasses having a camera that faces the physical environment to capture still and video images of the participant’s gaze scene.
- Figure 1 represents a flowchart of an embodiment of a process of the invention for the automatic and egocentric recognition and detection of tasks in a participant’s gaze scene in an uncontrolled environment.
- Figure 2 represents a schematic view of a computer system that performs the process of Figure 1.
- Figure 3 represents a perspective view of an embodiment of a gaze input device of the computer system of Figure 2.
- Figure 4 represents a navigation session performed in a factory for the production of tires.
- Figure 1 represents an embodiment of a process for automatic and egocentric recognition and detection of tasks in a participant’s gaze scenes in an uncontrolled environment (or “process”) 10 of the invention.
- the process 10 is implemented by a computer system (or “system”) 100 to recognize, on the basis of one or more gaze scenes of a participant and of characteristic objects associated with the gaze scenes seen during a navigation session performed by the participant, tasks associated with a predetermined activity and the order of their performance.
- a set of tasks form an established way of doing something (or a “procedure”), and one or more procedures define the state in which things occur or are done (or an “activity”) in a given context.
- a “characteristic object” associated with a predetermined task includes a characteristic object associated with a predetermined activity of which the task forms part. It will also be understood that a “characteristic object” associated with a predetermined task includes a characteristic object associated with one or more processes that form part of a predetermined activity of which the task forms part.
- the process 10 of the invention uses processes and tools based on artificial intelligence (or “Al”) to associate the visual objects and the performed tasks on the basis of scenes regarded by the participant.
- the algorithm allows continuous improvement in all of the gaze scenes (including the incorporation of the gaze input), ensuring that the system 100 improves from the experience that it acquires via correction or validation by the participant, especially in the choice of the order of tasks to be detected and/or or to be done. Therefore, the gaze treatment performed by the disclosed invention allows to differentiate the recognition of objects associated with each task, of tasks associated with a procedure and/or a particular activity, and the recognition of an activity in its entirety.
- the input of data in the participant’s gaze scene may facilitate an identification of objects in the participant’s gaze scene.
- objects identified in the gaze scenes may be associated with one or more predetermined tasks.
- the input of data in the gaze - o - scene captured during one or more navigation sessions of the participant 108 may include a position of the participant’s gaze derived from a visual input and relating to a given particular visual object.
- This captured input of data on the gaze scene may also include a position of the participant’s gaze associated with at least some of the visual objects with which the participant has the intention of interacting.
- the local interaction elements must be sufficiently dense.
- the process 10 therefore includes the extraction of characteristics of a gaze scene.
- these characteristics are at least partially based on the input of the participant’s gaze and on the characteristic objects associated with the gaze scenes seen during a navigation session.
- the disclosed process uses machine-learning protocols (for example, training of neural networks) to recognize characteristic objects associated with the tasks performed (or with the environment to be recognized), and to describe the tasks, processes and/or activities on the basis of the presence or absence of characteristic objects.
- acquisition of video gives a detailed comprehension of the tasks that determine which objects are involved in the procedures and in the activities, the nature of these objects and the results obtained.
- the succession of tasks which may vary partially both in the order of certain procedures and certain activities and in their duration, may be interrupted.
- the object of the interaction may exit from the field of view of the camera during the task (for example, if the participant is interacting with a person, he/she turns his/her head to interact with the person and to look elsewhere during the performance of the task).
- an activity may therefore be identified on the basis of a succession of detections of characteristic objects associated with a task and/or on the basis of the order of the tasks performed during a procedure and/or a predetermined activity (including during a sequence of procedures that form part of a predetermined activity).
- the process 10 includes an acquisition step 12 in which an eye-tracking system is used to record the participant’s behavior captured in a gaze scene of a navigation session performed by the participant.
- the eye-tracking system obtains data on a gaze scene potentially associated with an object characteristic of a particular activity in which the participant is participating.
- the obtained gaze scene data may be associated with an object characteristic of a procedure (or of a series of procedures) characteristic of a particular activity in which the participant is participating.
- any reference to an object characteristic of a particular procedure (or of a series of procedures) includes the association of such a characteristic object with the corresponding particular activity of which the particular procedure (or series of procedures) forms part.
- participant refers to a single participant or to a group of participants.
- a participant includes, without limitation, an individual participant in a particular activity, an individual member of a team or of a group that is participating in a particular activity, one or more machines and/or vehicles associated with an individual or a team that is participating in at least one task of a particular activity, a digital community associated with a particular activity, and combinations and equivalents thereof.
- the participant may be a spectator watching a task, physically or digitally (for example, for live control of a particular activity remotely in order to see the gaze scene in real time).
- “participant” (or “subject”) may also refer to one or more others that analyze data collected in another environment.
- “participant” (or “subject”) may also refer to any electronic apparatus or system configured to receive control input and configured to automatically send data to at least one other participant.
- process may include one or more steps performed by at least one computer system including one or more processors for executing instructions that perform the steps (including the identification of the activities). Any sequence of steps may be given by way of example, and the described processes are not limited to any particular sequence.
- Figure 2 represents a computer system (or “system”) 100 that performs the process 10.
- the system 100 includes a communication network 102 incorporating one or more communication servers (or “servers”) 102a that manage the data input into the system 100 from various sources.
- the communication network 102 may include wired or wireless links and may employ any data- transfer protocol known to those skilled in the art.
- wireless links may include, without limitation, radiofrequency (RF) links, satellite links, (analogue or digital) mobile telephone links, Bluetooth® links, Wi-Fi links, infrared links, ZigBee links, local-area- network (LAN) links, wireless-local-area-network (WLAN) links, wide-area-network (WAN) links, near-field-communication (NFC) links, links according to other wirelesscommunication standards and configurations, their equivalents, and a combination of these elements.
- RF radiofrequency
- satellite links satellite links
- Bluetooth® links Wi-Fi links
- infrared links ZigBee links
- LAN local-area- network
- WLAN wireless-local-area-network
- WAN wide-area-network
- NFC near-field-communication
- processor refers to one or more devices capable of processing and analyzing data and including one or more software packages for processing same (for example, one or more integrated circuits known to those skilled in the art as being included in a computer, one or more controllers, one or more microcontrollers, one or more microcomputers, one or more programmable logic controllers (or “PLCs”), one or more application-specific integrated circuits, one or more neural networks, and/or one or more other known equivalent programmable circuits).
- PLCs programmable logic controllers
- the processor includes software for processing the data captured by the elements associated with the system 100 (and the corresponding data obtained) as well as software for identifying and locating variance and identifying the sources thereof in order to correct same.
- the communication network 102 of the system 100 includes an eye-tracking system (or gazetracking system) including at least one gaze input device and a tracking component for tracking the participant’s gaze.
- the gaze input device is represented in Figure 3 by glasses 104 including a frame 104b with a lateral arm 104b’ for resting on each of the ears of the participant.
- One or both of the arms 104b’ include a control circuit 104c for controlling the glasses 104 that supplies various electronic elements (including, by way of example, loud-speakers, inertial sensors, a GPS transceiver and/or a temperature sensor).
- the glasses 104 incorporate inertial sensors, the latter serving to detect the position, orientation and sudden acceleration of the glasses 104 mounted on the participant’s head.
- the frame 104b includes a video camera (or “camera”) 106 that faces the physical environment and that is able to capture still and video images of the participant’s gaze scene.
- image in the singular or in the plural, refers to still images and to images captured in video sequences, respectively.
- the camera 106 may be chosen from commercially available cameras, including, without limitation, depth cameras and cameras employing visible light (RGB camera).
- RGB camera visible light
- the terms “video camera”, “camera”, “photographic apparatus”, “sensor” and “optical sensor” may be used interchangeably and may refer to one or more apparatuses configured to capture still and video images.
- one or more cameras may operate as a peripheral of the glasses 104 (for example, one or more cameras capturing videos originating from a moving participant; for example, cameras borne or worn by subjects sitting in a wheelchair that transmit the images to the servers 102a).
- the glasses are typically worn on the participant’s head so that the latter may see through an optical system 104a for each eye, therefore giving the participant a real direct view of the space in front of him/her.
- the term “real direct view” refers to the ability to see objects of a physical environment with the human eye directly and not representations of images created with the objects (for example, watching a video of an object on a screen is not a real direct view of the object).
- the gaze input device may include one or more complements and/or equivalents of the glasses 104, including one or more portable or wearable devices such as a mobile network device (including “augmented reality” and/or “virtual reality” devices, and/or any combinations and/or any equivalents) (see, for example, the head-mounted display device of publication WO2021/006978).
- a mobile network device including “augmented reality” and/or “virtual reality” devices, and/or any combinations and/or any equivalents
- the employed eye-tracking system records the participant’s behavior.
- the participant’s behavior includes the basic behavioral actions of the participant’s interaction with an object of the gaze scene (including an object that does not appear in the gaze scene but that is associated with the gaze scene to perform a predetermined task). It is not necessary to map objects in the physical environment represented in the gaze scene nor to identify them by positions relative to geographic locations or using the map of the gaze scene captured by the gaze input device (for example, the glasses 104).
- the gaze scenes, including the objects associated with a predetermined task are recorded in videos so as to provide the participant with feedback on the physical environment captured in the gaze scene.
- the gaze input device may communicate by wire or wirelessly (for example, via Wi-Fi, Bluetooth, infrared, RFID transmission, Universal Serial Bus (USB), cellular transmission, and/or other wireless communication means) with the communication network 102.
- Wi-Fi Wireless Fidelity
- Bluetooth Wireless Fidelity
- RFID Wireless Fidelity
- USB Universal Serial Bus
- the functionality of the processor may be integrated into a unit borne or worn on the participant’s body (for example, a unit worn on the wrist or borne in a pocket).
- the functionality of the processor may be integrated into a separate apparatus, including one or more portable apparatuses connected to the communication network 102 (represented in Figure 2).
- the functionality of the processor may be integrated into software and/or hardware of the glasses 104.
- the communication network 102 of the system 100 may include one or more other communication devices (or “devices”) that capture and transmit data collected from the participant’s gaze scene to the server 102a.
- the one or more communication devices may include one or more portable devices such as a mobile network device 116 of the identified participant (for example, a mobile telephone, a laptop computer, one or more portable or wearable devices connected to the network, including “augmented reality” and/or “virtual reality” devices, and/or any combinations and/or any equivalents).
- the one or more communication devices may also include one or more remote computers 118 able to transfer data via the communication network 102.
- a portable device 116 of the system 100 may transmit data collected from the participant’s gaze scene to a remote computer 118 of the system 100.
- the remote computer 118 may transmit, to the portable device 116, a detection of tasks associated with a predetermined activity in the course of being performed by the participant 108.
- the gaze represents a direction in which the eyes of a participant 108 are oriented during a gaze input.
- a tracking component integrated into the glasses 104 may track the gaze of the participant 108 in order to generate a gaze input.
- the tracking component may take into account a predetermined distance between the eyes, the head and/or the nose and a characteristic object in the participant’s gaze scene.
- the system 100 is able to identify objects associated with the physical, regarded scene with which the participant 108 intended to interact.
- the gaze input may include an input in respect of ocular gaze, an input in respect of head position/orientation, and/or an input in respect of nasal orientation.
- the ocular gaze may include a position/orientation of one or more eyes of a participant 108 in the gaze input.
- the head posture may include a head position/orientation in the gaze input.
- the input regarding nasal orientation may include a nasal position/orientation in the gaze input.
- the tracking component may be a commercially available device (for example, a device from Gaze Intelligence (see https://gazeintelligence.com), a Google Glass® or an augmented-reality device (for example, a HoloLens® from Microsoft)).
- the system 100 includes one or more modules implemented by the computer system and executable by one or more processors.
- the modules include a receiving module 120 configured to receive the gaze input received from the gaze input device (for example, the glasses 104).
- the input of data from the gaze scene is a gaze input associated with at least one characteristic object associated with the gaze scene.
- the modules also include an extraction module 122 configured to extract a set of searched characteristic objects. This set of objects may include characteristic objects that appear in the gaze scene and characteristic objects that do not appear in the gaze scene but that are associated with predetermined tasks.
- the modules further include an analysis module 124 configured to identify, based at least partially on the set of characteristic objects associated with the gaze scene, the particular characteristic object associated with a task in order to recognize one or more associated activities in a predetermined order.
- the process includes an eye-tracking-related acquisition step 12 during which data on the participant’s gaze scene are collected using the eye-tracking system of the system 100.
- this step is performed with the glasses 104.
- the eye-tracking- related acquisition step 12 incorporates recording of the participant’s behavior and of the participant’s gaze input, this being a gaze input associated with at least one characteristic object associated with the captured gaze scene.
- the process 10 includes a step of performance of a process 140 of construction of a model for detecting characteristic objects on the basis of one or more influential parameters in the participant’s gaze scene.
- the process 10 also includes a step of performance of a process 160 of construction of a model for detecting tasks associated with the detected objects.
- the process of the invention may be performed in any physical environment without prior knowledge of such an environment.
- the characteristic object detection model is constructed on the basis of one or more influential parameters in the participant’s gaze scene.
- the process 140 of construction of the characteristic object detection model includes a step 142 of video acquisition of at least one navigation session performed by the participant.
- Each navigation session includes at least one video containing images of the gaze scene and of one or more gaze inputs of the participant.
- Each video includes one or more characteristic objects associated with the participant’s gaze scene.
- the influential parameters are extracted from the acquired videos.
- the obtained influential parameters include data corresponding to general information on the gaze scene recorded in the video. These data are stored (for example, in one or more databases of the system 100), and they are updated on a continuous basis or on an intermittent basis.
- the step 142 of video acquisition includes a step of acquisition of characteristic objects associated with a predetermined task and searched during navigation sessions (or “searched characteristic objects”). During this step, images of searched characteristic objects are captured to create a training database (or “database”) of characteristic objects. This step may be performed in advance of other steps of the process 10 with a view to feeding one or more neural networks with actual expected configurations of the characteristic objects via analysis of the images of the captured videos. To decrease the required computational resources, regions of interest (or “ROIs”) may be employed to filter out the parts of the data that do not correspond to the ROIs.
- ROIs regions of interest
- one or more known devices incorporate one or more programming modes, including learning-based programming modes, for feeding, modifying and training a neural network.
- the processor may use databases of annotated data (or “ground truth”) to train and/or develop the neural network in order to automatically detect the participant’s gaze scene where the searched characteristic object (for example, the profile of the searched characteristic object) is expected to be found.
- the ground truth is represented in the training database.
- the software used to identify searched characteristic objects may convert a captured video into a set of images.
- the obtained variations in images which reveal one or more positions of the profile boundary of the searched characteristic object, train the neural network to identify all the positions of the profile boundary of the searched characteristic object.
- the aim of the algorithm is to automatically locate and indicate the external profile of the searched characteristic object and its position in the participant’s gaze scene.
- Figure 4 represents an embodiment of a navigation session performed by the participant 108.
- the gaze scene of the participant 108 reveals the interaction of the participant with one or more stations associated with tasks of an activity of production of tires that is performed in a factory 200.
- a participant may perform the process 100 of the invention in any physical environment (including, without limitation, in a workshop, in a town, in a hospital, on an unknown site, etc.).
- each station incorporates characteristic objects recorded for completing one or more predetermined tasks of the activity of tire production (being “the predetermined activity”).
- the characteristic objects contribute to the general information of the gaze scene of the participant 108 managing the tasks in the factory 200.
- characteristic objects may include containers or conveyors for transporting raw materials (for example, rubber, chemicals, oil, etc.) and/or reinforcing fillers (for example, carbon black or silica) intended for mixing processes.
- raw materials for example, rubber, chemicals, oil, etc.
- reinforcing fillers for example, carbon black or silica
- characteristic objects may also include containers of metal wires intended for activities associated with a station 204 where a bead wire is constructed.
- characteristic objects may also include containers of metal wires and/or fabrics intended for processes associated with a calender-rolling station 206.
- the raw materials and the reinforcing fillers may be associated with one or more mixers (for example, Banbury mixers) that represent characteristic objects of a mixing station 208.
- mixers for example, Banbury mixers
- characteristic objects may also include handling means (for example, a forklift truck, a lift, a conveyor and/or pallet truck, etc.) that transport the mixtures output from the mixers.
- each handling means may be associated with at least one other station that performs one or more processes of tire production (for example, an extruding station 210, a cutting station 212, a building station 214 and/or a vulcanizing station 216).
- Each station incorporates characteristic objects for completing the associated activity (for example, one or more extruders in the case of the extruding station 210, one or more blades in the case of the cutting station 212, one or more building drums in the case of the building station 214 and/or one or more vulcanizing presses in the case of the vulcanizing station 216).
- the process 140 of construction of the characteristic object detection model further includes a classification step 143 for identifying the searched characteristic objects (and other properties) that exist in the image of the gaze scene of the participant 108.
- This step includes a step of creation of classes of characteristic objects associated with the predetermined activities on the basis of a known classification algorithm.
- the process 140 of construction of the characteristic object detection model further includes a step 144 of annotation of the searched characteristic objects in order to automatically detect profile boundaries between the characteristic objects and the gaze scene.
- the dataset may be annotated on the basis of data input by the participant to create the ground-truth data.
- all of the data of the images in the videos is annotated, and the known variations are identified manually, on the basis of the knowledge of tire-manufacturing professionals.
- the ground-truth data may include data originating from a plurality of sources, including a plurality of professionals located in separate locations, to develop the neural network.
- a feedback loop allows the annotated videos and images to be updated with additional ground-truth data over time in order to improve the accuracy of the process 10.
- the process 140 of construction of the characteristic object detection model further includes a step 146 of tracking classified characteristic objects, allowing a plurality of images to be generated based on an annotation.
- each image may include a variation moved with respect to the preceding image, so that all the points of the variation are shown with respect to the gaze scene of the participant 108.
- the process 140 of construction of the characteristic object detection model further includes a step of creation of a training database 130 of characteristic objects including data corresponding to the gaze scene of the participant 108.
- the training database 130 of characteristic objects is introduced into the characteristic object detection model.
- the created training database 130 of characteristic objects may include a reference database of characteristic objects for performing associated tasks (for example, building drums associated with the building station 214 of the factory 200 of Figure 4).
- the training database 130 of characteristic objects may include an already created reference (for example, a table of characteristic objects and a table of associated tasks recorded in a reference database 131 of current objects).
- the training database 130 of characteristic objects may include parameters corresponding to the distances between characteristic objects forming part of the tasks performed in a predetermined order.
- the specific source of the characteristic objects is not essential to the process 200 described here, which will work equally well whether “current” data or “stored” data extracted from electronic storage or a database are used, provided that a sufficient number of data points are available.
- the system 100 implements the process 10 by retrieving “stored” data to prepare personalized reminders to be sent to the participant in advance to plan one or more subsequent tasks.
- the process 140 of construction of the characteristic object detection model further includes a step 149 of training the characteristic object detection model to recognize the characteristic objects corresponding to the completion of the associated tasks. These data are classified in order to automatically detect the profile boundaries of the characteristic objects and/or the participant’s gaze scene.
- This step includes a step of identification of the characteristic objects captured in the gaze scene.
- the processor trains the neural network on the basis of newly input data corresponding to the navigation sessions and to the videos of the searched characteristic objects.
- a machine-learning process receives as input the gaze scene of the participant 108 and the data of the training database of characteristic objects so that the processor may recognize known characteristic objects corresponding to each task of the predetermined activity.
- An analysis of the machine learning is performed using a machinelearning model such as an artificial neural network including a plurality of layers.
- the machine-learning process employed includes a supervised-learning process (for example, one using a deep-learning system built around one or more CNNs used to process the images).
- supervised-learning process for example, one using a deep-learning system built around one or more CNNs used to process the images.
- models employing linear regression, logistic regression, decision trees, support vector machines, naive Bayes classification, K-nearest- neighbour (kNN) classification, K-means clustering, random-forest classification, dimensionality reduction algorithms, gradient-descent algorithms, neural networks (for example, auto-encoding networks, CNNs, RNNs, perceptrons, long short-term memory (LSTM), Hopfield networks, Boltzmann machines, deep-learning networks, deconvolutional networks, and complements and equivalents thereof.
- neural networks for example, auto-encoding networks, CNNs, RNNs, perceptrons, long short-term memory (LSTM), Hopfield networks, Boltzmann machines, deep-learning networks, deconvolutional networks, and complements and equivalents thereof.
- the one or more neural networks may be trained with ground-truth data that are generated using data representative of the influential parameters (for example, the data incorporated into the training database of characteristic objects that was described above).
- the neural network may be trained to localize and segment the characteristic object and/or the participant’s gaze scene in which the characteristic object is found.
- the neural network may also be trained to allow identification of characteristic objects in navigation sessions.
- On the basis of the validation data and on the basis of the newly input data it is possible to use software that allows adaptation in real time or almost real time (i.e., in times that are industrially acceptable) to the predictions of the characteristic objects output from the detecting model.
- the process 10 of the invention includes a process 160 of construction of the task detection model.
- the process 160 of construction of the task detection model may include a step 162 of introduction of a database 132 of descriptors of tasks associated with the predetermined activity.
- the database 132 of task descriptors includes a description of the characteristic objects associated with each task (for example, a mixing process that forms part of an activity of tire manufacture in the factory 200 of Figure 4).
- the output of the model for detecting characteristic objects and the database of task descriptors are used to construct the task detection model.
- the process 160 of construction of the task detection model further includes a step 164 of chronological visualization of the tasks and of the associated characteristic objects.
- a temporal synchronization is obtained between data on the gaze scene of the participant 108 (representing the “egocentric” perspective) and a navigation session performed by the participant (which contributes to the history of navigation sessions).
- the temporal synchronization is obtained using temporalsynchronization points recorded by the gaze input device (for example, the glasses 104) that correspond in time to a known task of the navigation session.
- the task and the temporalsynchronization points are used to align the participant’s gaze data and the timestamps of the tasks of the navigation session on the same clock. In this way, the data from the participant’s gaze scene are synchronized in time with the navigation session so that each gaze scene datum is mapped with a characteristic object at the appropriate time on the clock.
- this step may include a step of classification of the status of a task associated with the navigation session (for example, the task has not started, the task is in progress, the task has been completed, etc.).
- time may be used as an explicit signal (for example, to identify the relationship between the objects in the field of view of the camera and the objects absent from the field of view of the camera).
- To perform a classification of a video causal events are identified, and these affect their labels. It will be understood that the classifications may be associated with navigation sessions in progress and the historic navigation sessions, to ensure the incorporation of the characteristic objects in a given slot.
- a trained neural network makes it possible to recognize characteristic objects in navigation sessions and to create framing boxes (or “boxed regions”) around recognized characteristic objects.
- the boxed regions allow characteristics of the gaze to be measured (for example, distances between a framing box and a point of fixation of the participant at the start of the input of the gaze scene, the frequency with which the participant looked at the framing box, the total time for which the participant looked at the framing box, etc.).
- the coordinates of the framing box of the characteristic object are correlated with the coordinates of the gaze positions generated by the gaze input device (for example, the glasses 104). These coordinates are used to compute characteristics of a stream of visual information (including, without limitation, object appearance time, the number of times the participant’s gaze passed over the characteristic object, entry time, the number of fixations on an object, fixation times, and fixation dwell times).
- the boxed regions are transmitted to a neural network (for example, one or more CNNs) to allow the representation of the task to be learned (see step 168 of training the task detection detecting tasks described hereinbelow).
- the boxed regions and the characteristics of the stream of visual information may be transmitted with the corresponding temporal synchronization.
- the data corresponding to the characteristics of the stream of visual information may be classified depending on the predetermined activity in the course of being performed during the navigation session (for example, “a building drum of the building station 214”, “a mixer of the mixing station 208”, “a computer installed at the vulcanizing station 216”, etc., associated with an activity of tire production) (see Figure 4).
- the process 160 of construction of the task detection model further includes a step of creation of a training database 134 of tasks.
- the training database 134 of tasks is introduced into the task detection model.
- the training database 134 of tasks may include a reference database of tasks performed during a predetermined activity (for example, a mixing process performed at the mixing station 208 of the factory 200 of Figure 4).
- the training database 134 of tasks may include a reference database of tasks performed in a predetermined order (for example, a table of predetermined processes incorporating associated tasks).
- the created training database 134 of tasks may include an already created reference (for example, the database of task descriptors described hereinabove).
- the process 160 of construction of the task detection model further includes a step 167 of correction and validation of the presence of one or more characteristic objects associated with the participant’s gaze scene during the navigation session.
- This step includes a step of comparison of the chronological visualization of the participant’s gaze scene (that was established during the chronological visualization step 164) with the training database 134 of tasks.
- a machine-learning process obtains the obtained influential parameters (being the data corresponding to the navigation sessions in progress and the data of the training database 134 of tasks).
- the process 160 of construction of the task detection model further includes a step 168 of training the model for detecting tasks on the basis of the characteristic objects detected in the navigation session.
- This step includes a step of association of the identified characteristic objects with one or more tasks of the predetermined activity.
- the machinelearning process obtains the known characteristic objects corresponding to each task of the predetermined activity and of the data of the training database 134 of tasks.
- the process may include a step 170 of task-based measurement for detecting, on the basis of the recognized characteristic objects output from the characteristic object detection model, streams of interactions between the participant and the characteristic objects.
- the given tasks may be identified on the basis of data corresponding to the parameters of the tasks subsequent to the tasks captured in the gaze scene. It is understood that the data of subsequent tasks may be hypothetical or real depending on how the tasks are managed by the participant in his physical environment and on his ability to ascertain the progression of the activities (and the processes that form part of the activities) with which a task will be associated.
- a discrepancy between the actual characteristic objects (those found in a navigation session in progress) and the expected characteristic objects (found in the historic navigation sessions) is denoted by a computed error detected by the participant (for example, during the correction and validation step 167).
- the detection errors may be corrected in the training database of tasks (described above) to improve the predictive ability of the task detection model.
- this simulation aims to simulate the streams of interactions using historical task data and associated characteristic objects. This simulation may be performed by way of a Markov chain or using other equivalent processes. In each embodiment of the process of the invention, one or more steps may be performed iteratively.
- the system 100 may include pre-programmed control information. For example, an adjustment of a predetermined activity may be associated with the parameters of the typical physical environments in which the system 100 operates.
- the system 100 (and/or an installation incorporating the system 100) may receive voice commands or other audio data representing, for example, a command to start or stop capturing videos.
- the command may include a request to know the current state of the predetermined activity.
- a generated response may be represented audibly, visually, in a tactile manner (for example, by way of a haptic interface) and/or in a virtual and/or augmented manner. This response, together with the corresponding data, may be recorded in the neural network.
- a monitoring system could be provided, at least part of which may be supplied in a portable device such as a mobile-network device (for example, a mobile telephone, a laptop computer, one or more portable or wearable devices connected to the network (including “augmented reality” and/or “virtual reality” devices, wearable clothing connected to the network and/or any combinations and/or any equivalents)). It is conceivable for the detection and comparison steps to be able to be performed iteratively.
- a mobile-network device for example, a mobile telephone, a laptop computer, one or more portable or wearable devices connected to the network (including “augmented reality” and/or “virtual reality” devices, wearable clothing connected to the network and/or any combinations and/or any equivalents). It is conceivable for the detection and comparison steps to be able to be performed iteratively.
Landscapes
- Engineering & Computer Science (AREA)
- General Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Human Computer Interaction (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- Image Analysis (AREA)
Abstract
The invention relates to a process (10) for automatic and egocentric recognition and detection of tasks in gaze scene of a participant (108) in an uncontrolled environment, the process (10) being implemented by a computer system (100) to recognize, by recording data from the participant's gaze scene and from objects associated with the gaze scenes seen during a navigation session performed by the participant, tasks associated with a predetermined activity and the order of their performance. The invention also relates to a computer system (100) that performs the disclosed process.
Description
Description
Title: Automatic and Egocentric Detection and Recognition of Tasks in Gaze Scenes
Technical Field
The invention relates to the treatment of eye-tracking logs in the context of the analysis of the behavior of a participant in an open and uncontrolled environment.
Context
In the field of industrial processes, the treatment of eye-tracking logs (or “eye tracking” in English) is becoming increasingly common, especially with the wider adoption of the ecosystem of connected objects and their connected users (often referred to as the "Internet of Things" or "IOT"). When measuring the way in which visual processing resources are allocated during the exploration of a scene, eye tracking presents advantages in respect of spatial and temporal resolution, the complexity of the operations and the revelation of movement (for example, movement of the head).
The log is composed of two parts that are captured together: the position of the gaze (or “gaze input”) and the recording of what is being looked at (or “regarded scene”). With respect to the environment, known industrial processes work by attempting to control the latter, particularly by seeking to superpose the measured elements in real space with known elements in order to locate an object and determine precisely its spatial configuration. This treatment is applicable to controlled environments (for example, a screen or a scene with known references such as barcodes). In all these cases, the control of a situation in an uncontrolled environment remains unknown.
With respect to the gaze, existing treatments allow virtual images to be combined with the physical environment. These treatments (including, for example, measurement of the frequency and duration of blinks of the eyelids, of dilation of the pupil, of the time of fixation of an object and the frequency of jumps) is often related to the gaze of a person (or “participant”) in a way that is said to be “egocentric”. This person performs one or more tasks with each task consisting of two types of information for action recognition: the verb representing the action, and the noun representing the object(s) related to this action (see “Learning Human-Object Interactions by Graph Parsing Neural Networks”, Qi, S.; Wang, W.;
Jia, B.; Shen, J.; and Zhu, S.-C, Proceedings of the European Conference on Computer Vision (ECCV), 2018, pp. 401-417).
In video analysis, with the emergence of deep convolutional neural networks (or “CNNs”) and of large-scale datasets, the performance of action recognition has considerably improved. The egocentric comprehension of daily human activities disclosed in the document EPIC- Kitchens (see “Scaling Egocentric Vision: The EPIC-KITCHENS Dataset”, Dima Damen et el., 31 July 2018 (https://arxiv.org/abs/1804.02748)) is mentioned by way of example. This dataset relates to the natural interactions encompassing objects and actions. With respect to the recognition of actions by a third person, the object with which the person interacts is distinguished from various distracting objects. In EPIC-Kitchens, because of the vastness of the action vocabulary, the verb and noun classes are generally formed separately. The verb class concentrates on the classification of the actions that the participant performs (for example, “place” or “open”). The noun class serves to classify the object with which the participant is interacting. The predictions of both classes are generally merged without interactions for the action classification (for example, by means of CNN).
In certain real-world scenarios, there is a need for physiological gaze measurements to be correlated with the physical environment, particularly in order to measure discrepancies in the use of connected objects during the participation of participants in open and uncontrolled environments (for example, tasks performed in a factory, experiments performed in laboratories, pathology monitoring with the search for crisis triggers, studies of customer behavior in supermarkets; and other tasks and activities performed in other uncontrolled environments). Even for a human, it can be difficult to recognize an action by only looking at the objects while ignoring the participant’s intention, or by considering only the changes in movement without knowing the object in interaction. It would be useful to have an inference of the tasks and their associated activities in real time, in order to decide whether or not to interrupt an activity of the participant depending on its priority. This requires an architecture with a connection of an eye-tracking system to a server to make the inference in real time. Thus, the disclosed invention relates to a solution for dissociating the gaze input and the regarded scene in uncontrolled environments. The invention relates to a process that obtains the data input from the scene regarded by a participant, this input being captured by video during a navigation session performed by the participant. The problem addressed by the invention is that of automatic task recognition so as to recognize the associated activities in a
predetermined context and the order of their performance. On the basis of the data generated by a first-person video recording system integrating the position of the gaze (for example, an eye-tracking system incorporating a wearable device associating the gaze with a physiological measurement and a forehead-mounted camera that captures the tasks from a first-person viewpoint), it becomes possible to identify useful measurement ranges, i.e., the start and the end of each activity, depending on the presence of characteristic objects associated with each task. Labelled periods of activity interruption are also considered to identify breaks and the resumption based upon completed tasks.
Summary of the invention
The invention relates to a process for automatic and egocentric detection and recognition of tasks in gaze scenes of a participant in an uncontrolled environment, the process implemented by a computer system to recognize, by logging gaze scene data of the participant and objects associated with gaze scenes seen during a navigation session performed by the participant, tasks associated with a predetermined activity and the order of their performance, characterized in that the process includes the following steps: an eye-tracking acquisition step during which gaze scene data of the participant is collected using an eye-tracking system of the computer system that records the participant’s behavior captured in a gaze scene of a navigation session performed by the participant; a step of performance of a process of constructing a model for detecting characteristic objects on the basis of one or more influential parameters in the gaze scene of the participant; and a step of performance of a process of constructing on of a model for detecting constituent tasks of activities associated with the characteristic objects identified in the gaze scene; such that the computer system detects tasks performed based on the recognized characteristic objects output from the characteristic object detection model.
In some embodiments of the process of the invention, the process of construction of the characteristic object detection model includes the following steps: a step of video acquisition of at least one navigation session performed by the participant, during which at least one video is acquired by the eye-tracking system, each video consisting of images of the gaze scene and the gaze input of the participant;
a step of creation of a training database of characteristic objects including data corresponding to the gaze scene, during which the training database of characteristic objects is introduced into the characteristic object detection model; a step of classification of searched characteristic objects to identify searched characteristic objects that exist in images of the participant’s gaze scene; a step of annotation of images of searched characteristic objects to automatically detect boundaries between profiles of the characteristic objects and the gaze scene; a step of tracking characteristic objects classified during the classification step, allowing generation of a plurality of images from an annotation of images of searched characteristic objects performed during the annotation step; and a step of training the characteristic object detection model to recognize characteristic objects corresponding to the completion of the associated tasks, during which step a machinelearning process obtains data from the training database of characteristic objects; so that the computer system recognizes characteristic objects corresponding to each task of the predetermined activity.
In some embodiments of the process of the invention, the process of construction of the task and activity detection model includes the following steps: a step of introduction of a task description database associated with the predetermined activity, wherein the task description database includes a description of the characteristic objects associated with each task according to the classification of the searched characteristic objects created during the classification step of the process of construction of the characteristic object detection model, according to an order of appearance of the objects, according to durations of appearance of the objects, and according to a co-occurrence of certain objects in the participant’s gaze scene; a step of chronological visualization of the tasks and the associated characteristic objects, during which a temporal synchronization is obtained between the participant’s gaze scene data and a navigation session performed by the participant; a step of creation of a training database of tasks including a reference database of tasks performed during a predetermined activity and a description of the characteristic objects associated with each task; a step of correction and validation of the presence of one or more characteristic objects associated with the participant’s gaze scene during the navigation session, this step including
a step of comparison of the chronological visualization of the participant’s gaze scene obtained during the chronological visualization step and of the training database of tasks; and a step of training the task detection model from the characteristic objects detected in the navigation session, to recognize the characteristic objects in the navigation session, during which step a machine-learning process obtains the recognized characteristic objects corresponding to each task of the predetermined activity and the data from the training database of tasks.
In some embodiments of the process of the invention, during the chronological visualization step of the process of construction of the task detection model, the temporal synchronization is obtained using temporal-synchronization points recorded by the eye-tracking system that correspond in the task detection model to time of a known task of the navigation session; so that the task and the temporal-synchronization points are used to synchronize the participant’s gaze scene data over time with the navigation session so that each gaze scene datum is mapped to a characteristic object at the appropriate time on a common clock.
In some embodiments of the process of the invention, during the chronological visualization step of the process of construction of the task detection model, the task detection model uses the output of the object detection model in navigation sessions performed by the participant in order to place the coordinates of boxed regions, including bounding boxes surrounding the characteristic objects, in correlation with the coordinates of the gaze positions from the eyetracking system.
In some embodiments of the process of the invention, the coordinates of the boxed regions are used to compute characteristics of a visual information stream in order to transmit the boxed regions and the characteristics of the visual information stream to the task detection model.
In some embodiments of the process of the invention, the eye-tracking system includes at least one gaze input device coupled to the computer system that performs the process, the gaze input device including: a tracking component integrated into the gaze input device for tracking the participant’s gaze to generate the gaze scene data input from the participant; and at least one camera facing the physical environment around the participant’s gaze scene to capture still and video images incorporating the participant’s gaze scene.
In some embodiments of the process of the invention, the process further includes a step of task-based measurement step to detect, based upon the recognized characteristic objects
output from the object detection model, streams of interactions between the participant and the characteristic objects.
The invention also relates to a computer system that performs the disclosed process, including: an eye tracking system including at least one gaze input device and a tracking component for tracking the participant’s gaze; one or more modules implemented by the computer system and executable by one or more processors, the one or more modules including: a receiving module configured to receive the gaze scene data input received from the gaze input device, the gaze scene data input representing a gaze input associated with at least one characteristic object associated with the participant’s gaze scene; an extraction module configured to extract a set of searched characteristic objects; and an analysis module configured to identify, based at least partially on the set of characteristic objects associated with the gaze scene, the particular characteristic object associated with a task in order to recognize one or more associated activities in a predetermined order.
In some embodiments of the computer system of the invention, the gaze input device includes glasses having a camera that faces the physical environment to capture still and video images of the participant’s gaze scene.
Further aspects of the invention will become obvious from the following detailed description.
Brief description of the drawings
The nature and various advantages of the invention will become more evident from reading the following detailed description, and from studying the attached drawings, in which the same reference numerals designate identical parts throughout, and in which:
[Fig 1] Figure 1 represents a flowchart of an embodiment of a process of the invention for the automatic and egocentric recognition and detection of tasks in a participant’s gaze scene in an uncontrolled environment.
[Fig 2] Figure 2 represents a schematic view of a computer system that performs the process of Figure 1.
[Fig 3] Figure 3 represents a perspective view of an embodiment of a gaze input device of the computer system of Figure 2.
[Fig 4] Figure 4 represents a navigation session performed in a factory for the production of tires.
Detailed description
With reference now to the figures, in which the same numbers identify identical elements, Figure 1 represents an embodiment of a process for automatic and egocentric recognition and detection of tasks in a participant’s gaze scenes in an uncontrolled environment (or “process”) 10 of the invention. The process 10 is implemented by a computer system (or “system”) 100 to recognize, on the basis of one or more gaze scenes of a participant and of characteristic objects associated with the gaze scenes seen during a navigation session performed by the participant, tasks associated with a predetermined activity and the order of their performance. As used herein, a set of tasks form an established way of doing something (or a “procedure”), and one or more procedures define the state in which things occur or are done (or an “activity”) in a given context. It will be understood that a “characteristic object” associated with a predetermined task includes a characteristic object associated with a predetermined activity of which the task forms part. It will also be understood that a “characteristic object” associated with a predetermined task includes a characteristic object associated with one or more processes that form part of a predetermined activity of which the task forms part.
The process 10 of the invention uses processes and tools based on artificial intelligence (or “Al”) to associate the visual objects and the performed tasks on the basis of scenes regarded by the participant. The algorithm allows continuous improvement in all of the gaze scenes (including the incorporation of the gaze input), ensuring that the system 100 improves from the experience that it acquires via correction or validation by the participant, especially in the choice of the order of tasks to be detected and/or or to be done. Therefore, the gaze treatment performed by the disclosed invention allows to differentiate the recognition of objects associated with each task, of tasks associated with a procedure and/or a particular activity, and the recognition of an activity in its entirety.
The input of data in the participant’s gaze scene, which data are acquired during a navigation session in an uncontrolled environment (and include at least one video incorporating one or more gaze scenes of the participant) may facilitate an identification of objects in the participant’s gaze scene. In embodiments of the process 10, objects identified in the gaze scenes may be associated with one or more predetermined tasks. The input of data in the gaze
- o - scene captured during one or more navigation sessions of the participant 108 may include a position of the participant’s gaze derived from a visual input and relating to a given particular visual object. This captured input of data on the gaze scene may also include a position of the participant’s gaze associated with at least some of the visual objects with which the participant has the intention of interacting. The local interaction elements must be sufficiently dense.
The process 10 therefore includes the extraction of characteristics of a gaze scene. In one embodiment, these characteristics are at least partially based on the input of the participant’s gaze and on the characteristic objects associated with the gaze scenes seen during a navigation session. The disclosed process uses machine-learning protocols (for example, training of neural networks) to recognize characteristic objects associated with the tasks performed (or with the environment to be recognized), and to describe the tasks, processes and/or activities on the basis of the presence or absence of characteristic objects.
In the process 10, acquisition of video gives a detailed comprehension of the tasks that determine which objects are involved in the procedures and in the activities, the nature of these objects and the results obtained. The succession of tasks, which may vary partially both in the order of certain procedures and certain activities and in their duration, may be interrupted. When a moving camera is used, the object of the interaction may exit from the field of view of the camera during the task (for example, if the participant is interacting with a person, he/she turns his/her head to interact with the person and to look elsewhere during the performance of the task). In the context of the process 10, an activity may therefore be identified on the basis of a succession of detections of characteristic objects associated with a task and/or on the basis of the order of the tasks performed during a procedure and/or a predetermined activity (including during a sequence of procedures that form part of a predetermined activity).
With reference once again to Figure 1, the process 10 includes an acquisition step 12 in which an eye-tracking system is used to record the participant’s behavior captured in a gaze scene of a navigation session performed by the participant. During this step, the eye-tracking system obtains data on a gaze scene potentially associated with an object characteristic of a particular activity in which the participant is participating. It will be understood that the obtained gaze scene data may be associated with an object characteristic of a procedure (or of a series of procedures) characteristic of a particular activity in which the participant is participating.
Thus, any reference to an object characteristic of a particular procedure (or of a series of procedures) includes the association of such a characteristic object with the corresponding particular activity of which the particular procedure (or series of procedures) forms part. As used here, “participant” (or “subject”) refers to a single participant or to a group of participants. A participant includes, without limitation, an individual participant in a particular activity, an individual member of a team or of a group that is participating in a particular activity, one or more machines and/or vehicles associated with an individual or a team that is participating in at least one task of a particular activity, a digital community associated with a particular activity, and combinations and equivalents thereof. The participant may be a spectator watching a task, physically or digitally (for example, for live control of a particular activity remotely in order to see the gaze scene in real time). As used here, “participant” (or “subject”) may also refer to one or more others that analyze data collected in another environment. As used herein, “participant” (or “subject”) may also refer to any electronic apparatus or system configured to receive control input and configured to automatically send data to at least one other participant.
As used herein, the term “process” or “procedure” may include one or more steps performed by at least one computer system including one or more processors for executing instructions that perform the steps (including the identification of the activities). Any sequence of steps may be given by way of example, and the described processes are not limited to any particular sequence.
With reference once again to Figure 1, and also to Figures 2 and 3, Figure 2 represents a computer system (or “system”) 100 that performs the process 10. The system 100 includes a communication network 102 incorporating one or more communication servers (or “servers”) 102a that manage the data input into the system 100 from various sources. The communication network 102 may include wired or wireless links and may employ any data- transfer protocol known to those skilled in the art. Examples of wireless links may include, without limitation, radiofrequency (RF) links, satellite links, (analogue or digital) mobile telephone links, Bluetooth® links, Wi-Fi links, infrared links, ZigBee links, local-area- network (LAN) links, wireless-local-area-network (WLAN) links, wide-area-network (WAN) links, near-field-communication (NFC) links, links according to other wirelesscommunication standards and configurations, their equivalents, and a combination of these elements.
The term “processor” (or, alternatively, the term “programmable logic circuit”) refers to one or more devices capable of processing and analyzing data and including one or more software packages for processing same (for example, one or more integrated circuits known to those skilled in the art as being included in a computer, one or more controllers, one or more microcontrollers, one or more microcomputers, one or more programmable logic controllers (or “PLCs”), one or more application-specific integrated circuits, one or more neural networks, and/or one or more other known equivalent programmable circuits). The processor includes software for processing the data captured by the elements associated with the system 100 (and the corresponding data obtained) as well as software for identifying and locating variance and identifying the sources thereof in order to correct same.
The communication network 102 of the system 100 includes an eye-tracking system (or gazetracking system) including at least one gaze input device and a tracking component for tracking the participant’s gaze. By way of example, the gaze input device is represented in Figure 3 by glasses 104 including a frame 104b with a lateral arm 104b’ for resting on each of the ears of the participant. One or both of the arms 104b’ include a control circuit 104c for controlling the glasses 104 that supplies various electronic elements (including, by way of example, loud-speakers, inertial sensors, a GPS transceiver and/or a temperature sensor). In some embodiments, the glasses 104 incorporate inertial sensors, the latter serving to detect the position, orientation and sudden acceleration of the glasses 104 mounted on the participant’s head.
The frame 104b includes a video camera (or “camera”) 106 that faces the physical environment and that is able to capture still and video images of the participant’s gaze scene. As used herein, “image”, in the singular or in the plural, refers to still images and to images captured in video sequences, respectively. The camera 106 may be chosen from commercially available cameras, including, without limitation, depth cameras and cameras employing visible light (RGB camera). In the description that follows, the terms “video camera”, “camera”, “photographic apparatus”, “sensor” and “optical sensor” may be used interchangeably and may refer to one or more apparatuses configured to capture still and video images. It will be understood that one or more cameras may operate as a peripheral of the glasses 104 (for example, one or more cameras capturing videos originating from a moving participant; for example, cameras borne or worn by subjects sitting in a wheelchair that transmit the images to the servers 102a).
In the example of the glasses 104, the glasses are typically worn on the participant’s head so that the latter may see through an optical system 104a for each eye, therefore giving the participant a real direct view of the space in front of him/her. As used herein, the term “real direct view” refers to the ability to see objects of a physical environment with the human eye directly and not representations of images created with the objects (for example, watching a video of an object on a screen is not a real direct view of the object).
It will be understood that the gaze input device may include one or more complements and/or equivalents of the glasses 104, including one or more portable or wearable devices such as a mobile network device (including “augmented reality” and/or “virtual reality” devices, and/or any combinations and/or any equivalents) (see, for example, the head-mounted display device of publication WO2021/006978).
The employed eye-tracking system records the participant’s behavior. The participant’s behavior includes the basic behavioral actions of the participant’s interaction with an object of the gaze scene (including an object that does not appear in the gaze scene but that is associated with the gaze scene to perform a predetermined task). It is not necessary to map objects in the physical environment represented in the gaze scene nor to identify them by positions relative to geographic locations or using the map of the gaze scene captured by the gaze input device (for example, the glasses 104). The gaze scenes, including the objects associated with a predetermined task, are recorded in videos so as to provide the participant with feedback on the physical environment captured in the gaze scene.
The gaze input device (for example, the glasses 104) may communicate by wire or wirelessly (for example, via Wi-Fi, Bluetooth, infrared, RFID transmission, Universal Serial Bus (USB), cellular transmission, and/or other wireless communication means) with the communication network 102.
In embodiments of the system 100, the functionality of the processor may be integrated into a unit borne or worn on the participant’s body (for example, a unit worn on the wrist or borne in a pocket). In other embodiments, the functionality of the processor may be integrated into a separate apparatus, including one or more portable apparatuses connected to the communication network 102 (represented in Figure 2). In other embodiments of the system 100, the functionality of the processor may be integrated into software and/or hardware of the glasses 104.
The communication network 102 of the system 100 may include one or more other
communication devices (or “devices”) that capture and transmit data collected from the participant’s gaze scene to the server 102a. The one or more communication devices may include one or more portable devices such as a mobile network device 116 of the identified participant (for example, a mobile telephone, a laptop computer, one or more portable or wearable devices connected to the network, including “augmented reality” and/or “virtual reality” devices, and/or any combinations and/or any equivalents). The one or more communication devices may also include one or more remote computers 118 able to transfer data via the communication network 102. By way of example, a portable device 116 of the system 100 may transmit data collected from the participant’s gaze scene to a remote computer 118 of the system 100. On the basis of the transmitted data, the remote computer 118 may transmit, to the portable device 116, a detection of tasks associated with a predetermined activity in the course of being performed by the participant 108.
With reference once again to Figures 1 to 3, and further to Figure 4, the gaze represents a direction in which the eyes of a participant 108 are oriented during a gaze input. A tracking component integrated into the glasses 104 may track the gaze of the participant 108 in order to generate a gaze input. The tracking component may take into account a predetermined distance between the eyes, the head and/or the nose and a characteristic object in the participant’s gaze scene. On the basis of the gaze input captured in images of the gaze scene, the system 100 is able to identify objects associated with the physical, regarded scene with which the participant 108 intended to interact. The gaze input may include an input in respect of ocular gaze, an input in respect of head position/orientation, and/or an input in respect of nasal orientation. The ocular gaze may include a position/orientation of one or more eyes of a participant 108 in the gaze input. The head posture may include a head position/orientation in the gaze input. The input regarding nasal orientation may include a nasal position/orientation in the gaze input. The tracking component may be a commercially available device (for example, a device from Gaze Intelligence (see https://gazeintelligence.com), a Google Glass® or an augmented-reality device (for example, a HoloLens® from Microsoft)).
With reference once again to Figures 1 to 4, the system 100 includes one or more modules implemented by the computer system and executable by one or more processors. The modules include a receiving module 120 configured to receive the gaze input received from the gaze input device (for example, the glasses 104). The input of data from the gaze scene is a gaze input associated with at least one characteristic object associated with the gaze scene. The
modules also include an extraction module 122 configured to extract a set of searched characteristic objects. This set of objects may include characteristic objects that appear in the gaze scene and characteristic objects that do not appear in the gaze scene but that are associated with predetermined tasks. The modules further include an analysis module 124 configured to identify, based at least partially on the set of characteristic objects associated with the gaze scene, the particular characteristic object associated with a task in order to recognize one or more associated activities in a predetermined order.
With reference once again to Figure 1, at the start of the process 10 of the invention, the process includes an eye-tracking-related acquisition step 12 during which data on the participant’s gaze scene are collected using the eye-tracking system of the system 100. In one embodiment of the process, this step is performed with the glasses 104. The eye-tracking- related acquisition step 12 incorporates recording of the participant’s behavior and of the participant’s gaze input, this being a gaze input associated with at least one characteristic object associated with the captured gaze scene.
The process 10 includes a step of performance of a process 140 of construction of a model for detecting characteristic objects on the basis of one or more influential parameters in the participant’s gaze scene. The process 10 also includes a step of performance of a process 160 of construction of a model for detecting tasks associated with the detected objects. Of course, the process of the invention may be performed in any physical environment without prior knowledge of such an environment.
The characteristic object detection model is constructed on the basis of one or more influential parameters in the participant’s gaze scene. Thus, the process 140 of construction of the characteristic object detection model includes a step 142 of video acquisition of at least one navigation session performed by the participant. Each navigation session includes at least one video containing images of the gaze scene and of one or more gaze inputs of the participant. Each video includes one or more characteristic objects associated with the participant’s gaze scene. During this step, the influential parameters are extracted from the acquired videos. The obtained influential parameters include data corresponding to general information on the gaze scene recorded in the video. These data are stored (for example, in one or more databases of the system 100), and they are updated on a continuous basis or on an intermittent basis. The step 142 of video acquisition includes a step of acquisition of characteristic objects associated with a predetermined task and searched during navigation sessions (or “searched
characteristic objects”). During this step, images of searched characteristic objects are captured to create a training database (or “database”) of characteristic objects. This step may be performed in advance of other steps of the process 10 with a view to feeding one or more neural networks with actual expected configurations of the characteristic objects via analysis of the images of the captured videos. To decrease the required computational resources, regions of interest (or “ROIs”) may be employed to filter out the parts of the data that do not correspond to the ROIs.
To capture the images of searched characteristic objects, one or more known devices (for example, one or more cameras or other portable devices) incorporate one or more programming modes, including learning-based programming modes, for feeding, modifying and training a neural network. The processor may use databases of annotated data (or “ground truth”) to train and/or develop the neural network in order to automatically detect the participant’s gaze scene where the searched characteristic object (for example, the profile of the searched characteristic object) is expected to be found. The ground truth is represented in the training database.
The software used to identify searched characteristic objects may convert a captured video into a set of images. The obtained variations in images, which reveal one or more positions of the profile boundary of the searched characteristic object, train the neural network to identify all the positions of the profile boundary of the searched characteristic object. The aim of the algorithm is to automatically locate and indicate the external profile of the searched characteristic object and its position in the participant’s gaze scene.
By way of example, Figure 4 represents an embodiment of a navigation session performed by the participant 108. In this example, the gaze scene of the participant 108 reveals the interaction of the participant with one or more stations associated with tasks of an activity of production of tires that is performed in a factory 200. Of course, a participant may perform the process 100 of the invention in any physical environment (including, without limitation, in a workshop, in a town, in a hospital, on an unknown site, etc.).
In the example represented in Figure 4, each station incorporates characteristic objects recorded for completing one or more predetermined tasks of the activity of tire production (being “the predetermined activity”). The characteristic objects contribute to the general information of the gaze scene of the participant 108 managing the tasks in the factory 200. For example, with a station 202 where raw materials are present, characteristic objects may
include containers or conveyors for transporting raw materials (for example, rubber, chemicals, oil, etc.) and/or reinforcing fillers (for example, carbon black or silica) intended for mixing processes. At the station 202 where raw materials are present, characteristic objects may also include containers of metal wires intended for activities associated with a station 204 where a bead wire is constructed. At the station 202 where raw materials are present, characteristic objects may also include containers of metal wires and/or fabrics intended for processes associated with a calender-rolling station 206.
In a video of the factory 200 capturing tasks of the process for tire production, the raw materials and the reinforcing fillers may be associated with one or more mixers (for example, Banbury mixers) that represent characteristic objects of a mixing station 208. At the mixing station 208, characteristic objects may also include handling means (for example, a forklift truck, a lift, a conveyor and/or pallet truck, etc.) that transport the mixtures output from the mixers.
In a video of the factory 200 capturing tasks of the activity of tire production, each handling means may be associated with at least one other station that performs one or more processes of tire production (for example, an extruding station 210, a cutting station 212, a building station 214 and/or a vulcanizing station 216). Each station incorporates characteristic objects for completing the associated activity (for example, one or more extruders in the case of the extruding station 210, one or more blades in the case of the cutting station 212, one or more building drums in the case of the building station 214 and/or one or more vulcanizing presses in the case of the vulcanizing station 216). It will be understood that the various characteristic objects are capable of being captured in the gaze inputs during the navigation sessions. With reference once again to Figure 1, the process 140 of construction of the characteristic object detection model further includes a classification step 143 for identifying the searched characteristic objects (and other properties) that exist in the image of the gaze scene of the participant 108. This step includes a step of creation of classes of characteristic objects associated with the predetermined activities on the basis of a known classification algorithm. The process 140 of construction of the characteristic object detection model further includes a step 144 of annotation of the searched characteristic objects in order to automatically detect profile boundaries between the characteristic objects and the gaze scene. Before being recorded, the dataset may be annotated on the basis of data input by the participant to create the ground-truth data. For example, in certain embodiments of the process 10, to assist the
neural network with detecting and identifying the profile boundaries of a characteristic object and/or the participant’s gaze scene (for example, the factory 200 of Figure 4), all of the data of the images in the videos is annotated, and the known variations are identified manually, on the basis of the knowledge of tire-manufacturing professionals. The ground-truth data may include data originating from a plurality of sources, including a plurality of professionals located in separate locations, to develop the neural network. A feedback loop allows the annotated videos and images to be updated with additional ground-truth data over time in order to improve the accuracy of the process 10.
With reference once again to Figure 1, the process 140 of construction of the characteristic object detection model further includes a step 146 of tracking classified characteristic objects, allowing a plurality of images to be generated based on an annotation. During this step, for images of the navigation session that have been labelled as incorporating one or more searched characteristic objects, each image may include a variation moved with respect to the preceding image, so that all the points of the variation are shown with respect to the gaze scene of the participant 108.
The process 140 of construction of the characteristic object detection model further includes a step of creation of a training database 130 of characteristic objects including data corresponding to the gaze scene of the participant 108. During this step, the training database 130 of characteristic objects is introduced into the characteristic object detection model. The created training database 130 of characteristic objects may include a reference database of characteristic objects for performing associated tasks (for example, building drums associated with the building station 214 of the factory 200 of Figure 4). The training database 130 of characteristic objects may include an already created reference (for example, a table of characteristic objects and a table of associated tasks recorded in a reference database 131 of current objects). The training database 130 of characteristic objects may include parameters corresponding to the distances between characteristic objects forming part of the tasks performed in a predetermined order. The specific source of the characteristic objects is not essential to the process 200 described here, which will work equally well whether “current” data or “stored” data extracted from electronic storage or a database are used, provided that a sufficient number of data points are available. The system 100 implements the process 10 by retrieving “stored” data to prepare personalized reminders to be sent to the participant in advance to plan one or more subsequent tasks.
The process 140 of construction of the characteristic object detection model further includes a step 149 of training the characteristic object detection model to recognize the characteristic objects corresponding to the completion of the associated tasks. These data are classified in order to automatically detect the profile boundaries of the characteristic objects and/or the participant’s gaze scene. This step includes a step of identification of the characteristic objects captured in the gaze scene. During this step, the processor trains the neural network on the basis of newly input data corresponding to the navigation sessions and to the videos of the searched characteristic objects.
During the training step 149, a machine-learning process receives as input the gaze scene of the participant 108 and the data of the training database of characteristic objects so that the processor may recognize known characteristic objects corresponding to each task of the predetermined activity. An analysis of the machine learning is performed using a machinelearning model such as an artificial neural network including a plurality of layers. In one embodiment, the machine-learning process employed includes a supervised-learning process (for example, one using a deep-learning system built around one or more CNNs used to process the images). Although the embodiments described herein concern the use of neural networks by way of supervised-learning models, other types of machine-learning models may be used. These include, without limitation, models employing linear regression, logistic regression, decision trees, support vector machines, naive Bayes classification, K-nearest- neighbour (kNN) classification, K-means clustering, random-forest classification, dimensionality reduction algorithms, gradient-descent algorithms, neural networks (for example, auto-encoding networks, CNNs, RNNs, perceptrons, long short-term memory (LSTM), Hopfield networks, Boltzmann machines, deep-learning networks, deconvolutional networks, and complements and equivalents thereof.
The one or more neural networks may be trained with ground-truth data that are generated using data representative of the influential parameters (for example, the data incorporated into the training database of characteristic objects that was described above). The neural network may be trained to localize and segment the characteristic object and/or the participant’s gaze scene in which the characteristic object is found. The neural network may also be trained to allow identification of characteristic objects in navigation sessions. On the basis of the validation data and on the basis of the newly input data, it is possible to use software that
allows adaptation in real time or almost real time (i.e., in times that are industrially acceptable) to the predictions of the characteristic objects output from the detecting model. With reference once again to Figure 1, the process 10 of the invention includes a process 160 of construction of the task detection model. The process 160 of construction of the task detection model may include a step 162 of introduction of a database 132 of descriptors of tasks associated with the predetermined activity. The database 132 of task descriptors includes a description of the characteristic objects associated with each task (for example, a mixing process that forms part of an activity of tire manufacture in the factory 200 of Figure 4). Thus, the output of the model for detecting characteristic objects and the database of task descriptors are used to construct the task detection model.
With reference once again to Figure 1, the process 160 of construction of the task detection model further includes a step 164 of chronological visualization of the tasks and of the associated characteristic objects. During this step, a temporal synchronization is obtained between data on the gaze scene of the participant 108 (representing the “egocentric” perspective) and a navigation session performed by the participant (which contributes to the history of navigation sessions). The temporal synchronization is obtained using temporalsynchronization points recorded by the gaze input device (for example, the glasses 104) that correspond in time to a known task of the navigation session. The task and the temporalsynchronization points are used to align the participant’s gaze data and the timestamps of the tasks of the navigation session on the same clock. In this way, the data from the participant’s gaze scene are synchronized in time with the navigation session so that each gaze scene datum is mapped with a characteristic object at the appropriate time on the clock.
In one embodiment of the process 10, this step may include a step of classification of the status of a task associated with the navigation session (for example, the task has not started, the task is in progress, the task has been completed, etc.). As a video concerns a temporal sequence, time may be used as an explicit signal (for example, to identify the relationship between the objects in the field of view of the camera and the objects absent from the field of view of the camera). To perform a classification of a video, causal events are identified, and these affect their labels. It will be understood that the classifications may be associated with navigation sessions in progress and the historic navigation sessions, to ensure the incorporation of the characteristic objects in a given slot.
During the chronological visualization step 164 of the process 10, a trained neural network
makes it possible to recognize characteristic objects in navigation sessions and to create framing boxes (or “boxed regions”) around recognized characteristic objects. The boxed regions allow characteristics of the gaze to be measured (for example, distances between a framing box and a point of fixation of the participant at the start of the input of the gaze scene, the frequency with which the participant looked at the framing box, the total time for which the participant looked at the framing box, etc.).
During this training, the coordinates of the framing box of the characteristic object are correlated with the coordinates of the gaze positions generated by the gaze input device (for example, the glasses 104). These coordinates are used to compute characteristics of a stream of visual information (including, without limitation, object appearance time, the number of times the participant’s gaze passed over the characteristic object, entry time, the number of fixations on an object, fixation times, and fixation dwell times). The boxed regions (and, optionally, the visual information characteristics) are transmitted to a neural network (for example, one or more CNNs) to allow the representation of the task to be learned (see step 168 of training the task detection detecting tasks described hereinbelow).
The boxed regions and the characteristics of the stream of visual information may be transmitted with the corresponding temporal synchronization. During this step, the data corresponding to the characteristics of the stream of visual information may be classified depending on the predetermined activity in the course of being performed during the navigation session (for example, “a building drum of the building station 214”, “a mixer of the mixing station 208”, “a computer installed at the vulcanizing station 216”, etc., associated with an activity of tire production) (see Figure 4).
The process 160 of construction of the task detection model further includes a step of creation of a training database 134 of tasks. During this step, the training database 134 of tasks is introduced into the task detection model. The training database 134 of tasks may include a reference database of tasks performed during a predetermined activity (for example, a mixing process performed at the mixing station 208 of the factory 200 of Figure 4). The training database 134 of tasks may include a reference database of tasks performed in a predetermined order (for example, a table of predetermined processes incorporating associated tasks). The created training database 134 of tasks may include an already created reference (for example, the database of task descriptors described hereinabove).
The process 160 of construction of the task detection model further includes a step 167 of
correction and validation of the presence of one or more characteristic objects associated with the participant’s gaze scene during the navigation session. This step includes a step of comparison of the chronological visualization of the participant’s gaze scene (that was established during the chronological visualization step 164) with the training database 134 of tasks. During this step, a machine-learning process obtains the obtained influential parameters (being the data corresponding to the navigation sessions in progress and the data of the training database 134 of tasks).
The process 160 of construction of the task detection model further includes a step 168 of training the model for detecting tasks on the basis of the characteristic objects detected in the navigation session. This step includes a step of association of the identified characteristic objects with one or more tasks of the predetermined activity. During this step, the machinelearning process obtains the known characteristic objects corresponding to each task of the predetermined activity and of the data of the training database 134 of tasks.
In every embodiment of the process 10 of the invention, the process may include a step 170 of task-based measurement for detecting, on the basis of the recognized characteristic objects output from the characteristic object detection model, streams of interactions between the participant and the characteristic objects. The given tasks may be identified on the basis of data corresponding to the parameters of the tasks subsequent to the tasks captured in the gaze scene. It is understood that the data of subsequent tasks may be hypothetical or real depending on how the tasks are managed by the participant in his physical environment and on his ability to ascertain the progression of the activities (and the processes that form part of the activities) with which a task will be associated. A discrepancy between the actual characteristic objects (those found in a navigation session in progress) and the expected characteristic objects (found in the historic navigation sessions) is denoted by a computed error detected by the participant (for example, during the correction and validation step 167). The detection errors may be corrected in the training database of tasks (described above) to improve the predictive ability of the task detection model. In these embodiments, this simulation aims to simulate the streams of interactions using historical task data and associated characteristic objects. This simulation may be performed by way of a Markov chain or using other equivalent processes. In each embodiment of the process of the invention, one or more steps may be performed iteratively.
The system 100 may include pre-programmed control information. For example, an
adjustment of a predetermined activity may be associated with the parameters of the typical physical environments in which the system 100 operates. In embodiments of the process 10, the system 100 (and/or an installation incorporating the system 100) may receive voice commands or other audio data representing, for example, a command to start or stop capturing videos. The command may include a request to know the current state of the predetermined activity. A generated response may be represented audibly, visually, in a tactile manner (for example, by way of a haptic interface) and/or in a virtual and/or augmented manner. This response, together with the corresponding data, may be recorded in the neural network.
A monitoring system could be provided, at least part of which may be supplied in a portable device such as a mobile-network device (for example, a mobile telephone, a laptop computer, one or more portable or wearable devices connected to the network (including “augmented reality” and/or “virtual reality” devices, wearable clothing connected to the network and/or any combinations and/or any equivalents)). It is conceivable for the detection and comparison steps to be able to be performed iteratively.
The terms “at least one” and “one or more” are used interchangeably. The ranges given as lying “between a and b” encompass the values “a” and “b”.
Although particular embodiments of the disclosed apparatus have been illustrated and described, it will be appreciated that various changes, additions and modifications can be made without departing from either the spirit or scope of the present description. Therefore, no limitation should be imposed on the scope of the invention described, apart from those set out in the appended claims.
Claims
1. A process (10) for automatic and egocentric recognition and detection of tasks in gaze scenes of a participant (108) in an uncontrolled environment, the process (10) being implemented by a computer system (100) to recognize, by recording data from the participant’s gaze scene and from objects associated with the gaze scenes seen during a navigation session performed by the participant, tasks associated with a predetermined activity and the order of their performance, characterized in that the process comprises the following steps: an eye-tracking-related acquisition step (12) during which data from the participant’s gaze scene are collected using an eye-tracking system of the computer system (100) that records the behavior of the participant captured in a gaze scene of a navigation session performed by the participant; a step of performance of a process (140) of construction of a model for detecting characteristic objects on the basis of one or more influential parameters in the participant’s gaze scene; and a step of performance of a process (160) of construction of a model for detecting constituent tasks of activities associated with characteristic objects identified in the gaze scene; such that the computer system (100) detects performed tasks on the basis of the recognized characteristic objects output from the characteristic object detection model.
2. The process (10) of claim 1, wherein the process (140) of construction of the characteristic object detection model comprises the following steps: a step (142) of video acquisition of at least one navigation session performed by the participant, during which step at least one video is acquired by the eye-tracking system, each video consisting of images of the gaze scene and of the gaze input of the participant; a step of creation of a training database (130) of characteristic objects comprising data corresponding to the gaze scene, during which step the training database (130) of characteristic objects is introduced into the characteristic object detection model; a step (143) of classification of searched characteristic objects to identify searched characteristic objects that exist in images of the participant’s gaze scene; a step (144) of annotation of images of searched characteristic objects in order to
automatically detect profile boundaries of the characteristic objects and the gaze scene; a step (146) of tracking characteristic objects classified during the classification step (143), allowing a plurality of images to be generated on the basis of an annotation of images of searched characteristic objects that is performed during the annotation step (144); and a step (149) of training the characteristic object detection model to recognize characteristic objects corresponding to completion of the associated tasks, during which step a machine-learning process obtains data from the training database (130) of characteristic objects; such that the computer system recognizes characteristic objects corresponding to each task of the predetermined activity.
3. The process (10) of claim 2, wherein the process (160) of construction of the task detection model comprises the following steps: a step (162) of introduction of a database (132) of descriptors of tasks associated with the predetermined activity, in which step the database (132) of task descriptors comprises a description of the characteristic objects associated with each task according to the classification of the searched characteristic objects created during the classification step (142) of the process (140) of construction of the characteristic object detection model, according to an order of appearance of the objects, according to the times of appearance of the objects, and according to a co-occurrence of certain objects in the participant’s gaze scene; a step (164) of chronological visualization of the tasks and of the associated characteristic objects, during which step a temporal synchronization is obtained between data from the gaze scene of the participant (108) and a navigation session performed by the participant; a step of creation of a training database (134) of tasks comprising a reference database of tasks performed during a predetermined activity and a description of the characteristic objects associated with each task; a step (167) of correction and validation of the presence of one or more characteristic objects associated with the participant’s gaze scene during the navigation session, this step comprising a step of comparison of the chronological visualization of the participant’s gaze scene that was obtained during the chronological visualization step (164) and of the training database (134) of tasks; and
a step (168) of training, on the basis of the characteristic objects detected in the navigation session, the task detection model to recognize the characteristic objects in the navigation session, during which step a machine-learning process obtains the recognized characteristic objects corresponding to each task of the predetermined activity and to the data of the training database (134) of tasks.
4. The process (10) of claim 3, wherein, during the chronological visualization step (164) of the process (160) of construction of the task detection model, the temporal synchronization is obtained using temporal-synchronization points recorded by the eye-tracking system that correspond in time to a known task of the navigation session; such that the task and the temporal-synchronization points are used to synchronize in time the data from the participant’s gaze scene with the navigation session so that each gaze scene datum is mapped with a characteristic object at the appropriate time on a common clock.
5. The process (10) of claim 3 or of claim 4, wherein, during the chronological visualization step (164) of the process (160) of construction of the task detection model, the task detection model uses the results output from the characteristic object detection model in navigation sessions performed by the participant (108) in order to correlate the coordinates of boxed regions, comprising boxes framing characteristic objects, with the coordinates of gaze positions generated by the eye-tracking system.
6. The process (10) of claim 5, wherein the coordinates of the boxed regions are used to compute characteristics of a visual information stream in order to transmit the boxed regions and the characteristics of the visual information stream to the task detection model.
7. The process (10) of any one of claims 1 to 6, wherein the eye-tracking system comprises at least one gaze input device coupled to the computer system (100) that performs the process, the gaze input device comprising: a tracking component integrated into the gaze input system for tracking the gaze of the participant so as to generate the input of data from the participant’s gaze scene; and at least one camera (106) facing the physical environment around the participant’s gaze scene so as to capture still and video images incorporating the participant’s gaze scene.
- 25 -
8. The process (10) of any one of claims 1 to 7, further comprising a step (170) of taskbased measurement for detecting, on the basis of the recognized characteristic objects output from the characteristic object detection model, streams of interactions between the participant and the characteristic objects.
9. A computer system (100) that performs the process of any one of claims 1 to 8, comprising: an eye-tracking system comprising at least one gaze input device and a tracking component for tracking the gaze of the participant; one or more modules implemented by the computer system (100) and executable by one or more processors, the one or more modules comprising: a receiving module (120) configured to receive the input of data from the gaze scene, the input of gaze scene data being received from the gaze input device and representing a gaze input associated with at least one characteristic object associated with the participant’s gaze scene; an extraction module (122) configured to extract a set of searched characteristic objects; and an analysis module (124) configured to identify, based at least partially on the set of characteristic objects associated with the gaze scene, the particular characteristic object associated with a task in order to recognize one or more associated activities in a predetermined order.
10. The computer system (100) of claim 9, wherein the gaze input device comprises glasses (104a) comprising a camera (106) that faces the physical environment to capture still and video images of the participant’s gaze scene.
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| FR2114058 | 2021-12-21 | ||
| PCT/EP2022/085780 WO2023117613A1 (en) | 2021-12-21 | 2022-12-14 | Automatic and egocentric detection and recognition of tasks in gaze scenes |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP4453692A1 true EP4453692A1 (en) | 2024-10-30 |
Family
ID=81346332
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP22835037.7A Pending EP4453692A1 (en) | 2021-12-21 | 2022-12-14 | Automatic and egocentric detection and recognition of tasks in gaze scenes |
Country Status (2)
| Country | Link |
|---|---|
| EP (1) | EP4453692A1 (en) |
| WO (1) | WO2023117613A1 (en) |
Families Citing this family (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN119203019B (en) * | 2024-07-29 | 2025-10-17 | 电子科技大学 | Zero-sample multi-mode first visual angle behavior recognition method based on visual language knowledge introduction |
Family Cites Families (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP6103765B2 (en) * | 2013-06-28 | 2017-03-29 | Kddi株式会社 | Action recognition device, method and program, and recognizer construction device |
| US11017231B2 (en) | 2019-07-10 | 2021-05-25 | Microsoft Technology Licensing, Llc | Semantically tagged virtual and physical objects |
-
2022
- 2022-12-14 WO PCT/EP2022/085780 patent/WO2023117613A1/en not_active Ceased
- 2022-12-14 EP EP22835037.7A patent/EP4453692A1/en active Pending
Also Published As
| Publication number | Publication date |
|---|---|
| WO2023117613A1 (en) | 2023-06-29 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US20210081754A1 (en) | Error correction in convolutional neural networks | |
| JP2022546644A (en) | Systems and methods for automatic anomaly detection in mixed human-robot manufacturing processes | |
| JP2019522300A (en) | Mobile and wearable video capture and feedback platform for the treatment of mental disorders | |
| US20100218094A1 (en) | Second-person avatars | |
| EP2933066A1 (en) | Activity monitoring of a robot | |
| US20210304001A1 (en) | Multi-head neural network model to simultaneously predict multiple physiological signals from facial RGB video | |
| CN113748389B (en) | Method and apparatus for monitoring industrial process steps | |
| US20230343040A1 (en) | Personal protective equipment training system with user-specific augmented reality content construction and rendering | |
| KR20240032779A (en) | Electric device, method for control thereof | |
| CN113569671A (en) | Abnormal behavior alarm method and device | |
| US20250091214A1 (en) | Systems and methods for robotic teleoperation intention estimation | |
| Akbari et al. | Facilitating human activity data annotation via context-aware change detection on smartwatches | |
| Banerjee et al. | Robot classification of human interruptibility and a study of its effects | |
| EP4453692A1 (en) | Automatic and egocentric detection and recognition of tasks in gaze scenes | |
| US11340703B1 (en) | Smart glasses based configuration of programming code | |
| US10910096B1 (en) | Augmented reality computing system for displaying patient data | |
| US20250209822A1 (en) | Method and apparatus for feedbacking human-machine collaboration state based on virtual-real integration, and electronic device | |
| US12456204B1 (en) | Computer vision-driven interactive full-body motion tracking | |
| CN120260080B (en) | Pet posture and expression tracking method and device, electronic equipment and storage medium | |
| Al-Amin | Sensor data based adaptive models for assembly worker training in cyber manufacturing | |
| US20260094106A1 (en) | System | |
| JP2026045448A (en) | system | |
| US20260051233A1 (en) | System | |
| JP2026073419A (en) | system | |
| JP2026019072A (en) | system |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: UNKNOWN |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE |
|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE |
|
| 17P | Request for examination filed |
Effective date: 20240722 |
|
| AK | Designated contracting states |
Kind code of ref document: A1 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR |
|
| DAV | Request for validation of the european patent (deleted) | ||
| DAX | Request for extension of the european patent (deleted) |