EP3506829A1 - Prédiction de l'attention d'un auditoire lors d'une présentation - Google Patents
Prédiction de l'attention d'un auditoire lors d'une présentationInfo
- Publication number
- EP3506829A1 EP3506829A1 EP17772084.4A EP17772084A EP3506829A1 EP 3506829 A1 EP3506829 A1 EP 3506829A1 EP 17772084 A EP17772084 A EP 17772084A EP 3506829 A1 EP3506829 A1 EP 3506829A1
- Authority
- EP
- European Patent Office
- Prior art keywords
- attention
- level
- presentation
- speaker
- evolution
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F3/00—Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
- G06F3/01—Input arrangements or combined input and output arrangements for interaction between user and computer
- G06F3/011—Arrangements for interaction with the human body, e.g. for user immersion in virtual reality
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L25/00—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
- G10L25/48—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use
- G10L25/51—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use for comparison or discrimination
- G10L25/63—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use for comparison or discrimination for estimating an emotional state
-
- A—HUMAN NECESSITIES
- A61—MEDICAL OR VETERINARY SCIENCE; HYGIENE
- A61B—DIAGNOSIS; SURGERY; IDENTIFICATION
- A61B5/00—Measuring for diagnostic purposes; Identification of persons
- A61B5/103—Measuring devices for testing the shape, pattern, colour, size or movement of the body or parts thereof, for diagnostic purposes
- A61B5/11—Measuring movement of the entire body or parts thereof, e.g. head or hand tremor or mobility of a limb
- A61B5/1126—Measuring movement of the entire body or parts thereof, e.g. head or hand tremor or mobility of a limb using a particular sensing technique
- A61B5/1128—Measuring movement of the entire body or parts thereof, e.g. head or hand tremor or mobility of a limb using a particular sensing technique using image analysis
-
- A—HUMAN NECESSITIES
- A61—MEDICAL OR VETERINARY SCIENCE; HYGIENE
- A61B—DIAGNOSIS; SURGERY; IDENTIFICATION
- A61B5/00—Measuring for diagnostic purposes; Identification of persons
- A61B5/16—Devices for psychotechnics; Testing reaction times ; Devices for evaluating the psychological state
- A61B5/168—Evaluating attention deficit, hyperactivity
-
- A—HUMAN NECESSITIES
- A61—MEDICAL OR VETERINARY SCIENCE; HYGIENE
- A61B—DIAGNOSIS; SURGERY; IDENTIFICATION
- A61B5/00—Measuring for diagnostic purposes; Identification of persons
- A61B5/48—Other medical applications
- A61B5/4803—Speech analysis specially adapted for diagnostic purposes
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F3/00—Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
- G06F3/01—Input arrangements or combined input and output arrangements for interaction between user and computer
- G06F3/017—Gesture based interaction, e.g. based on a set of recognized hand gestures
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N20/00—Machine learning
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V20/00—Scenes; Scene-specific elements
- G06V20/50—Context or environment of the image
- G06V20/52—Surveillance or monitoring of activities, e.g. for recognising suspicious objects
-
- G—PHYSICS
- G09—EDUCATION; CRYPTOGRAPHY; DISPLAY; ADVERTISING; SEALS
- G09B—EDUCATIONAL OR DEMONSTRATION APPLIANCES; APPLIANCES FOR TEACHING, OR COMMUNICATING WITH, THE BLIND, DEAF OR MUTE; MODELS; PLANETARIA; GLOBES; MAPS; DIAGRAMS
- G09B19/00—Teaching not covered by other main groups of this subclass
- G09B19/04—Speaking
-
- G—PHYSICS
- G09—EDUCATION; CRYPTOGRAPHY; DISPLAY; ADVERTISING; SEALS
- G09B—EDUCATIONAL OR DEMONSTRATION APPLIANCES; APPLIANCES FOR TEACHING, OR COMMUNICATING WITH, THE BLIND, DEAF OR MUTE; MODELS; PLANETARIA; GLOBES; MAPS; DIAGRAMS
- G09B5/00—Electrically-operated educational appliances
- G09B5/06—Electrically-operated educational appliances with both visual and audible presentation of the material to be studied
- G09B5/062—Combinations of audio and printed presentations, e.g. magnetically striped cards, talking books, magnetic tapes with printed texts thereon
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L25/00—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
- G10L25/48—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use
-
- A—HUMAN NECESSITIES
- A61—MEDICAL OR VETERINARY SCIENCE; HYGIENE
- A61B—DIAGNOSIS; SURGERY; IDENTIFICATION
- A61B2503/00—Evaluating a particular growth phase or type of persons or animals
- A61B2503/12—Healthy persons not otherwise provided for, e.g. subjects of a marketing survey
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F2203/00—Indexing scheme relating to G06F3/00 - G06F3/048
- G06F2203/01—Indexing scheme relating to G06F3/01
- G06F2203/011—Emotion or mood input determined on the basis of sensed human body parameters such as pulse, heart rate or beat, temperature of skin, facial expressions, iris, voice pitch, brain activity patterns
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V40/00—Recognition of biometric, human-related or animal-related patterns in image or video data
- G06V40/10—Human or animal bodies, e.g. vehicle occupants or pedestrians; Body parts, e.g. hands
- G06V40/16—Human faces, e.g. facial parts, sketches or expressions
- G06V40/174—Facial expression recognition
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16H—HEALTHCARE INFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR THE HANDLING OR PROCESSING OF MEDICAL OR HEALTHCARE DATA
- G16H50/00—ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics
- G16H50/30—ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics for calculating health indices; for individual health risk assessment
Definitions
- the present invention relates to the field of systems and methods for predicting the attention of an audience and more particularly to a presentation by at least one speaker after a learning phase on a set of presentations already made.
- the purpose of these measures is to detect a decrease in the person's attention in order to intervene either to stimulate the person or to change the visualized content or the context in which the person is.
- the present invention improves the situation.
- the method is such that it comprises the following steps:
- the speaker of the presentation has a prediction information on the attention that the audience will or will carry to the presentation he is performing. For this, it is not necessary with this method to make real-time measurements of the attention of the audience. Similarly, if a presentation is recorded for later distribution, the process allows to be informed on an estimate of the evolution of the level of attention that an audience can have so as to adapt the rest of the presentation as necessary. .
- the information relating to the evolution of the level of attention includes a probability on the evolution of the level of attention and this probability is presented to the at least one speaker.
- the probability thus presented allows the speaker to know to what extent he has to rely on the prediction of the evolution of the level of attention he has received. It can thus best adapt its future actions.
- the information relating to the evolution of the level of attention is corrected based on context information of the audience.
- the attention of the audience may differ depending on whether the audience includes one or more people, depending on the location in which the presentation is made or broadcast, according to the presentation schedule, the temperature of the location of an audience, depending on the type of audience present in the audience, whether or not the members of that audience made a substantial meal, etc., the attention of the audience may differ .
- the context information thus makes it possible to improve or modify the estimate of the level of attention that has been measured.
- the information relating to the evolution of the level of attention is corrected according to measurements of emotion associated with the characteristics measured.
- the attention of the audience can evolve significantly depending on the content of the presentation, especially key words or phrases, or as a variant of particular images that can generate emotion, which has to refocus the attention of the audience.
- Keywords and / or keyphrases, or images generating emotion are determined by analyzing the audio and video signals of the presentation, for example by means of voice recognition and / or image recognition.
- the database includes information to link these keywords with a measure of emotion.
- the attention of the audience may also differ depending on whether or not the characteristic elements are related to an additional emotion measure.
- the information relating to the content of the presentation coupled with an associated emotion measurement therefore makes it possible to improve or modify the estimate of the level of attention that has been measured.
- This measure of emotion can also be different depending on the type of audience present or the context of the audience. The two emotion and context information can then be taken into account to improve or modify the estimate of the level of attention.
- the method further comprises a step of determining recommendations for actions to be performed by the speaker to change the level of attention of the at least one audience according to the information relating to the evolution of the level of attention recovered and a step of presentation to the at least one speaker of the recommendations determined.
- the speaker knows how to adapt his presentation to increase the level of attention of his present or future audience. It can best optimize the current presentation.
- a learning phase is implemented.
- the invention thus relates to a method of learning information of evolution of the level of attention of at least one audience of presentations.
- the learning phase is such that it comprises the following steps:
- determining changes in attention levels by analyzing associations determined for a set of characteristics or groups of characteristics and according to at least one parameter of duration or occurrence of these characteristics;
- This learning method can be implemented on a plurality of presentations made by the same speaker or by different speakers so as to have a panel representative of the possible characteristics of presentations and speakers.
- the resulting database can be enriched as new measurements are made for new presentations, so it can evolve.
- This learning process therefore makes it possible to associate information on the evolution of the level of attention with characteristics related to the presentation in progress.
- the resulting database can be saved in the terminal implementing the prediction method, for example the presenter's terminal so that it has the attention level evolution information in a simple manner and without it. it is necessary to have measuring devices or even network access.
- the information relating to the level of attention comprises a probability of evolution calculated from the analysis of a repeatability rate of the evolutions determined on all the presentations.
- the learning method also takes into account the characteristics of the contexts of the audiences, which will make it possible, when using said database, to select audience contexts corresponding to those expected for a presentation on which the method will be used. applied.
- the same context characteristics may also be taken into account if we also apply emotion measures, these measures may also differ from one context to another.
- the invention provides a device for predicting the attention of at least one audience of a presentation made by at least one speaker.
- the device is such that it comprises:
- a measurement and detection module for detecting vocal or gestural characteristics of the at least one speaker of the current presentation and / or the content characteristics of the current presentation and for measuring at least one parameter of duration or of occurrence detected characteristics
- a database consultation module for determining information relating to the evolution of the level of attention, the database comprising a correspondence between voice or gestural characteristics of the speaker and / or presentation content characteristics. duration or occurrence parameters related to these characteristics and information relating to changes in the level of attention for these characteristics and parameters;
- a user interface for presenting to the at least one speaker of the presentation a prediction of level of attention from the information relating to the evolution of the level of attention retrieved.
- the invention relates to a terminal which comprises a prediction device as described.
- This terminal and this device have the same advantages as the method described above.
- the invention is aimed at a prediction system such that it comprises a prediction device described above and a learning device comprising:
- an indexing module for indexing, on the one hand, audience level of attention measurements, presentations of the set and, on the other hand, indexing by measures of voice or gestural characteristics of the speakers and / or measures of content characteristics of presentations, presentations of the set;
- a synchronization module for synchronizing the respective indexings to determine associations between characteristics and attention level measures for the presentations of the set;
- an analysis module for determining changes in attention levels by analyzing associations determined for a set of characteristics or groups of characteristics and according to a parameter of duration or occurrence of these characteristics;
- a database recording module for recording matches between speaker speech or gesture characteristics and / or presentation content characteristics, duration or time parameters related to these characteristics and information relating to the evolution of the level of attention for these characteristics and these parameters.
- This learning device can advantageously be inserted into a server of a communication network. It can also be inserted into a terminal.
- This device has the same advantages as the learning method described above, which it implements.
- the invention finally relates to a prediction system such that it comprises a learning device as described and a prediction device as described.
- the invention also relates to a computer program comprising code instructions for implementing the steps of the prediction method as described and / or the learning method as described above, when these instructions are executed by a processor.
- It also relates to a processor readable information medium on which is recorded such a computer program comprising instructions for performing the steps of the prediction method and / or the learning method as described.
- FIGS. 1a and 1b illustrate examples of a system for predicting the attention of an audience in presentation or conference contexts in real time or recorded, animated by a speaker, face-to-face in a room with an audience or in line through a network communication;
- FIG. 2a illustrates, in flowchart form, the main steps of a learning process of attention level evolution information, prior to the prediction phase, in one embodiment of the invention;
- FIG. 2b illustrates in flowchart form the main steps of a prediction method according to an embodiment of the invention
- FIG. 3 illustrates a hardware configuration of a learning device able to implement the learning method according to one embodiment of the invention.
- FIG. 4 illustrates a hardware configuration of a prediction device according to one embodiment of the invention.
- Figure la represents an example of a system and context in which the prediction method according to the invention can be implemented.
- An Ul speaker is presenting to an AU audience. It realizes its presentation using an El screen and a terminal, here a computer Tl.
- the computer Tl is for example connected to an R network of the internet type and can thus be connected to a server S on which a learning method has been implemented to constitute a database DB2.
- the learning method may also in another exemplary embodiment be implemented in the terminal T1. It will be described in more detail with reference to FIG. 2a.
- the terminal T1 or the server S implements a prediction method according to the invention. This will be described later with reference to Figure 2b.
- the terminal T1 To implement the prediction method and / or learning in the terminal T1, it is associated with at least one unrepresented microphone capable of capturing the oral presentation of the speaker. The captured sound will then be analyzed to determine speaker sound characteristics.
- the terminal T1 is also associated with a not shown camera that films and detects the movements of the speaker. These movements can also be analyzed to determine other characteristics of the speaker during the current presentation.
- Figure lb describes another context of a prediction system according to the invention.
- the speaker Ul performing a presentation or online training type MOOC for "Massive Open Online Course” in English
- MOOC for "Massive Open Online Course” in English
- a server S may in one embodiment implement the learning method and / or the prediction method according to the invention.
- the prediction method is implemented in the terminal T1 and the learning method in the server S or both the prediction method and the learning method are implemented in the terminal Tl.
- the terminal T1 is associated for example with a microphone and a camera to detect both the speaker's sound characteristics and the movement characteristics.
- the DB2 database is fed as a result of the learning phase and includes correspondences between speaker characteristic elements such as voice or gestural characteristics of the speaker and / or presentation as content characteristics of the presentation, duration or occurrence parameters related to these characteristics or elements and information relating to the evolution of the level of attention for these characteristics and these parameters.
- MOOC type presentation may be recorded by the presenter of that MOOC for later online broadcast or to be recorded on the network for viewing at any time.
- the audience consists of a single person who consults the presentation in isolation and when he wishes.
- This learning method constitutes a learning phase implemented prior to the steps of the prediction method
- a set of already recorded presentations is available for example on the network or in a database of either the network or the equipment implementing this learning phase.
- a server of the network or a terminal of a user for example, of the speaker of the presentation can implement this phase.
- an analysis is performed on each of the presentations P, of the set of presentations Pi to P N which will be called a reference set or set of reference presentations.
- An analysis is performed on the characteristics of the presenter, speaker of the presentation.
- an audio measurement sensor measures during the course of the presentation, the speaker's sound level, the prosody characteristics over time, that is to say the phenomena of accentuation and intonation (variation height, duration and intensity) of the speaker's voice.
- Another video-type sensor can measure the gestures made by the speaker during his recorded presentation and the breaks he can perform.
- Other analysis elements make it possible to measure, for example, the possible sound effects during the presentation.
- the analysis performed also determines the characteristics of the content itself, such as the way the presentation was filmed, the evolution of the framing, the presence of keywords, images or keyframe sequences. , for example using an image analysis algorithm. All these analysis elements are indexed and indexed, in step E21b, on the timeline of the flow of the reference presentation. From this same reference set, a step E20a is implemented to measure the attention of an audience.
- the learning device measuring this level of attention is for example equipped with a camera capable of detecting the movements of the face, the blinking of the eyes, the frequency of yawning, etc.
- the attention level measurements thus obtained are then indexed, in step E20b, on the time line of the progress of the reference presentation.
- a technique for measuring the level of attention is for example a technique based on the analysis of the faces of the audience. For example, when the audience consists of people consulting their computer to attend a conference or online training, the capture of the image of the face of the viewer can see when it turns away from its screen, if it moves away, moves or is replaced by another face. In all these cases, it means that the attention of the user has decreased.
- Another possible measure is based on measuring the frequency of blinking of audience members. When the number of blinks exceeds a certain threshold or when the eyelids of the user are too long closed, it means that the user is in the beginning phase of sleepiness, and therefore loss of attention.
- Yawning is a typical reaction induced by fatigue. This results in a prolonged and uncontrolled opening of the mouth very different from other deformations of the lips and which can be measured by image analysis techniques. The opening of the mouth during a yawn is more important than the opening of the mouth during speech.
- Such a technique is for example described in the article entitled "Yawning detection for fatigue driver monitoring of the authors Xiao Fan, Bao-Cai Yin, Yan- Feng Sun in" Proceedings of the Sixth International Conference on Machine Learning and Cybernetics "in Hong Kong , 19-22 August 2007.
- An orientation change detection of an auditor's head can also reveal a decrease in attention. Indeed a head fall forward is indicative of the fatigue of the person. If this detection is further correlated with other detections described above, then that person's loss of attention is revealed.
- the level of ambient chatter noise can also be detected and thus can reveal that the audience is not attentive to the presentation that is offered to them.
- an individual measure may be preferred over a global approach.
- the attention measure is performed for each of the people in the audience, the overall level of attention being then determined by the aggregate of the unit attention levels.
- context information of the audience is associated with the level of attention measurement. Indeed, depending on the context of the audience, the measure of attention may vary.
- the level of attention of a user may be different for the same presentation. It is indeed known that a state of drowsiness can be promoted early digestion within one hour after a meal while alertness reaches its maximum two to three hours after a meal. If one measures the level of attention of the same presentation at different times and for a similar audience, one can determine the correction to be made to the level of attention measured according to the time.
- a date, a duration of sunshine, the heat in a room, or the number of people attending the presentation may be background information to correct the level of attention being measured.
- the type of audience present in the audience can also cause the level of attention to be different, for example if the audience is old, young, of different culture, speaking a different language, etc.
- This type of measurement is for example performed by known techniques of facial analysis detecting for example a smile, a particular grimace, crying, etc.
- This indexing of emotion of the reference presentations is compared with indexings of the characteristics of the content of the presentations, for example the existence of key words, images or sequences of key images.
- step E22 a reconciliation of the different indexings performed in steps E20b and E21b is performed.
- a synchronization of the two types of indexing is implemented so that the audience attention measure, indexed at a time instant of the presentation, is associated with the characteristics of the speaker and / or presentation for this same temporal moment of the presentation.
- the synchronization will be limited to coincide the beginnings of said time lines.
- the resynchronizations can be periodic based on sequences detected as common (for example by analysis of the soundtrack and comparison) .
- step E23 the different synchronizations performed for each presentation between speaker characteristics, presentation characteristics and attention level measurements, are used by an analysis module to determine the evolution of the level of attention. .
- This module determines probabilities of correlation between a decrease or a rise observed on the attention measure and different groups of characteristic elements of the presentation and / or the speaker.
- a cause-and-effect duration parameter between groups of speaker and / or presentation characteristic elements is also determined, as are changes in attention measurements to distinguish, for example, the groups of elements that generate either immediately after a period of repetition of these elements, a rate of loss of attention or a rate of increase of attention.
- At this step is also determined the influence of an occurrence parameter of appearance of a group of characteristic elements in a presentation.
- step E23 makes it possible to determine an evolution of the level of attention according to a group of speaker characteristic elements or of the content of the presentation or both and according to at least one duration parameter. or occurrence of these characteristic elements.
- a monotone of a speaker of a presentation lasting several minutes progressively changes the level of attention downward while the pronunciation of key words or projection of key images (eg violence or beautiful landscape) can change the level of attention brutally upward.
- thresholds of increase and decrease of attention are defined so as to retain only significant characteristics of the speaker and / or the presentation.
- the threshold may for example be 1 or 2%.
- the step E23 also implements a verification of the repeatability of the evolutions determined for each of the reference representations.
- this correspondence is recorded in a database of DB2 data, also called learning base.
- a calculation of the probability of evolution of the level of attention can be carried out from this analysis of the rate of repeatability of the evolutions of level of attention determined on the set of reference. This probability can then be recorded in the database DB2, in association with the correspondence evolution / characteristics which corresponds to it.
- step E24 is recorded in a database DB2, a set of information relating to the evolution of the level of attention (evolution of the level up or down, rate of evolution, ie a progressive progressivity index, for example abrupt or progressive, a probability of evolution, for example the repeatability rate in the reference set, etc.) in correspondence with elements or groups of characteristic elements speaker and / or presentation and at least one parameter of duration or occurrence of these elements.
- a set of information relating to the evolution of the level of attention evolution of the level up or down, rate of evolution, ie a progressive progressivity index, for example abrupt or progressive, a probability of evolution, for example the repeatability rate in the reference set, etc.
- Said database DB2 can in its simplified version be limited to separate backup files, or a retention of information in a relational database table separate from other tables constituting DB1.
- An advantage of said distinction DB2 base is of course that it can subsequently be used distinctly from the DB1 database as part of the prediction process described with reference to Figure 2b. Rather than having to use the very large DB1 database with the different indexing of presentations, only the results of the analyzes contained in the DB2 database, namely the list of groups of characteristic elements and the associated duration or occurrence parameters causing a probability of evolution of attention and the associated evolution information (as described above) is necessary.
- the small potential size of the DB2 database thus allows autonomous uses in embedded mode, without the need for a network connection to a server dedicated to the DB1 database.
- several reference sets can be provided. The different sets are for example created according to the themes of the presentations or depending on the type of audience.
- An example of a record in the database DB2 may be, for a silence characteristic of the speaker with a duration parameter of a few seconds, a correspondence with information relating to the evolution of the level of attention which is an immediate increase attention.
- Another example is a correspondence between a sound level of the speaker's voice that remains invariant for several minutes and a gradual decrease in the level of attention.
- the change of speaker may for example be associated with an immediate increase in the level of attention, so the change of framing of the display of the presentation may be associated with an immediate increase in attention.
- a rate of rise or fall in the level of attention that is to say a progressivity index of the evolution can also be associated with the triggering characteristic elements.
- the database DB2 is enriched by a set of information relating to the evolution of the level of attention in correspondence with elements or groups of presentation and presentation characteristic elements. / or speakers and parameters of duration or occurrence of these elements.
- the attention level evolution information is characterized by a trend of evolution, decrease or increase, if necessary, a rate of evolution, that is to say an index linked to the progressiveness of the evolution. of attention, to distinguish the immediate effects and the effects smoothed over a longer period of time and the likelihood that this trend will apply. Information about the average time of occurrence of the evolution of attention can also be recorded.
- FIG. 2b illustrates the steps implemented during the prediction method according to the invention. This method is implemented for example in the terminal T1 of the presenter or in a server S of the communication network R. It applies to a presentation in progress, animated by at least one speaker. We will talk about current presentation Pc.
- a first step E25 performs an analysis of this presentation. This analysis concerns, for example, speaker characteristics, voice, sound level, gestures, pause or breathing time, etc.
- a voice analysis module is provided in the prediction device, on the sound picked up by a microphone associated with the equipment of the presentation.
- the analysis can also include the content of the presentation, what is presented on the screen, the frequency of page change, the framing of what is shown, the colors used, the detection of keywords, images keys, etc.
- This type of analysis can be performed for example by detecting an action of the presenter for the page change, by an image analyzer to detect colors or movements or keyframes, etc.
- a search in the DB2 database of these speaker and / or presentation features and associated parameters is performed in E27 to find information on the evolution of the corresponding probable attention level.
- This information therefore makes it possible to obtain a prediction of the level of attention that the presentation will have if the corresponding characteristic elements persist during the associated duration or are repeated according to the associated occurrence and if the speaker does not change his presentation or characteristics.
- This attention level prediction information is presented to the speaker of the presentation in E28 so that he can react in real time to his presentation.
- the prediction of the level of attention is associated with a determination of recommendations of actions to be performed on the presentation to change the level of attention in the desired direction, followed by the presentation of these recommendations to speaker.
- An example of a recommendation is to ask to increase the sound level of the speaker's voice if it has been detected that the level of the voice decreases over time and that the time for which the level of attention falls, is exceeded.
- an interface to select the desired direction of change in the level of attention to lower the attention (if for example the presenter must absolutely evoke such a subject, but he prefers that nobody remembers) or the increase.
- a default mode simplifying the interface from the point of view of the speaker would be to improve the attention with respect to a given relevant level, fixed for example with reference to the beginning of the presentation, phase where the attention is classically considered to be its maximum.
- the recommendation could consist in proposing to disseminate an image, for example of a beautiful landscape, whose impact on the probability of evolution of attention is known.
- This suggestion could also consist of proposing groups of key words to be pronounced.
- the presenter can modify his presentation according to the recommendations and thus improve the level of attention of his audience.
- the presenter is thus informed of potential changes in audience attention even if there is no ongoing measure of audience attention or even if there is no audience .
- the presentation may be simply being recorded without anyone in front for subsequent broadcast before an audience.
- the presenter may simply be rehearsing the presentation that he will make later, in order to be more efficient at the appropriate time.
- attention-measuring equipment it is not necessary to have attention-measuring equipment to be informed of the evolution of the level of attention in real time.
- this information on the evolution of the level of attention can be corrected according to context information related to the present or planned audience.
- This information can be for example the schedule of the presentation or that scheduled to be broadcast, the number of people in the audience, the temperature of the room in which the presentation is made, etc.
- the correction to be made is for example recorded in the DB1 and DB2 databases in association with the characteristics of the speaker and / or the presentation.
- a weighting of the level of attention can be provided and recorded in the database DB2. This weighting is then applied to the information relating to the evolution of the level of attention obtained during the prediction process when the triggering characteristic elements are associated with emotion measurements as previously described.
- the presentation of the level of attention prediction can be made in different forms. In an exemplary embodiment it can be presented to the speaker by a symbol of different color. For example, the color intensity may correspond to the rate of change of the level of attention. This presentation can be made on the personal screen of the presenter or, if it is light, on the microphone of it.
- the evolution of the level of attention can also be represented by an arrow pointing upwards in the event of a rise and downwards in the event of a fall, of a greater or lesser height depending on the rate of change associated.
- Another way to display the result of this prediction is, for example, to display at the beginning of the sequence an average value of the level of attention and then as the presentation progresses, to represent the level predictions.
- the prediction on the evolution of the level is then well readable by the presenter.
- a% representing the accuracy rate found on the basis of training data for the current forecast may be presented.
- the delay in which the prediction of the evolution of the attention is expected can also be presented in seconds for example.
- such a prediction method may allow the presenter of a training or presentation to improve it by taking into account the changes in level of attention that are presented to him. He can for example train before a real presentation in order to optimize his intervention and to avoid the declines of level of attention. It can also provide different presentations based on different audience context information. For example, depending on the presentation's broadcast schedule, it can make the more dynamic presentation with changes in speaker or tone, if the presentation is broadcast at a time of digestion and predict less dynamic otherwise.
- the prediction and suggestion process could lead to the diffusion of 3 different variants of the same MOOC video, one of which is both shorter and more dynamized because the session is planned at the beginning of afternoon on an area and period for which a high temperature is expected.
- FIG. 3 represents a simplified hardware architecture of an embodiment of a learning device implementing the learning method described with reference to FIG. 2a.
- module and/or entity
- This device is equipped with a measurement collection interface 320 able to collect the measurements captured by the C1 to CN sensors represented here at 310 1 , 310 2 , 310 3 and 310 N.
- These sensors are provided on the one hand to measure the vocal characteristics of the speaker or speakers, for example by means of one or more microphones, for measuring the speaker's movement characteristics, for example by means of a camera and on the other hand for measuring the level. of audience attention.
- a camera can be provided for this also a camera.
- the device comprises a processing unit 330 equipped with a processor and driven by a computer program Pg 345 stored in a memory 340 and implementing the learning phase according to the invention.
- the code instructions of the computer program Pg are for example loaded into a RAM not shown and executed by the processor of the processing unit 330.
- the processor of the processing unit 330 implements the steps of the learning method described above with reference to Figure 2a, according to the instructions of the computer program Pg.
- the device 300 thus comprises an input interface for receiving already recorded presentations of a DB database comprising one or more sets of reference presentations.
- the indexing module receives measurements collected by the interface 320 and made by the sensors C1 to CN to determine the loudness level of the presenter, its tone, its silences or the level of surrounding loudness. It also receives information about the changes in the content presented, for example a change in the framing, a change of presentation page, a zoom on the image, a keyword, a keyframe, coming from the interface 320 .
- the learning device also includes a module for indexing attention level measurements. These attention measurement levels are obtained by the interface 320 which collects the measurements carried out by the sensors C1 to CN and in particular the data measured by one or more cameras from which algorithms for detecting blinking or yawning or Still head positioning is implemented to obtain a level of attention measurement.
- This attention level measure is indexed to the current presentation of the set of reference presentations.
- a synchronization module 370 is also provided to synchronize the two types of indexing and to obtain an association between the speaker and / or presentation features of the module 350 and the attention level measurements of the indexing module 360.
- This combination of elements can be stored in a DB1 database integrated in the device or available via a communication network via a communication module 390.
- An analysis module 380 driven by the processor 330, analyzes the associations of measured and characteristic levels of attention for the reference presentations and determines an evolution of the level of attention according to at least one parameter of duration or a parameter of occurrence of characteristics.
- the analysis makes it possible to associate a characteristic or a succession of characteristics with a change in level of attention. It also makes it possible to determine a duration or number of occurrences for which the measured characteristic changes the level of attention.
- a correspondence is made between characteristic elements or groups of characteristic elements of the speaker and / or presentation, duration or occurrence parameters associated with these elements with information relating to the changing level of attention of the audience.
- the analysis module determines in a particular embodiment, the repeatability rate of the defined associations. Only matches with sufficient repeatability can be stored in the DB2 database.
- This database can be stored on a remote server accessible via a communication network via the communication module 390 of the device.
- the communication network is for example an IP network.
- this database DB2 is integrated in the learning device. It can also be sent or downloaded to a terminal, for example that of a presentation speaker.
- This learning device is either a network server communicating with the presenter's terminal or the presenter's terminal itself.
- FIG. 4 represents a simplified hardware architecture of an embodiment of a prediction device 400 implementing the prediction method described with reference to FIG. 2b.
- module and/or hardware components, capable of implementing the or the functions described for the module or entity concerned.
- This device is equipped with an input interface able to consult a DB2 database internal to the device or available on a communication network and including correspondences between speaker and / or presentation features, duration or time parameters. occurrence related to these elements and information relating to the evolution of the level of audience attention for these elements and parameters and as learned during a learning phase as described with reference to Figure 2a.
- the device comprises a processing unit 430 equipped with a processor and driven by a computer program Pg 445 stored in a memory 440 and implementing the prediction method according to the invention.
- the code instructions of the computer program Pg are for example loaded into a RAM not shown and executed by the processor of the processing unit 430.
- the processor of the processing unit 430 implements the steps of the prediction method described above, according to the instructions of the computer program Pg.
- the device 400 therefore comprises an input interface for receiving the data stream of the current presentation Pc.
- This interface can also receive context information from the audience of this presentation (Inf.Ctx).
- the processor 430 implements the module for determining information relating to the evolution of the level of attention by searching in the database DB2, via the interface 420 or via the memory 440, whether a correspondence with the detected element and the associated parameter is saved. If necessary, a prediction on the evolution of the level of attention resulting from the information relating to the evolution of the level of attention thus determined is sent to the user interface 470 so that a presentation of this prediction of evolution is made to the speaker of the current presentation. Recommendations for actions to be performed by the speaker can also be sent on this user interface so that it changes the level of attention of its presentation.
- This prediction device can be included in the speaker terminal of the presentation.
- the prediction is directly displayed on the screen of its terminal via the user interface or on an accessory connected to its terminal, such as a microphone.
- the device can also be integrated into a server of a communication network, for example an IP network; in this case, the prediction is presented to the speaker of the presentation via a communication module 490 which transmits the information to the presenter's terminal.
- a server of a communication network for example an IP network
- the context information of the audience can be used by the determination module 460 to correct the determined evolution if necessary.
- both the learning device and the prediction device are included in the same equipment, either the speaker's terminal or a server of the network. In another embodiment, these two devices are remote, the learning method and the prediction method being implemented in a system comprising the two devices communicating with each other via a network.
Landscapes
- Engineering & Computer Science (AREA)
- Health & Medical Sciences (AREA)
- Physics & Mathematics (AREA)
- Life Sciences & Earth Sciences (AREA)
- Theoretical Computer Science (AREA)
- General Health & Medical Sciences (AREA)
- General Physics & Mathematics (AREA)
- Medical Informatics (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Multimedia (AREA)
- General Engineering & Computer Science (AREA)
- Surgery (AREA)
- Molecular Biology (AREA)
- Business, Economics & Management (AREA)
- Educational Technology (AREA)
- Veterinary Medicine (AREA)
- Public Health (AREA)
- Biophysics (AREA)
- Pathology (AREA)
- Biomedical Technology (AREA)
- Heart & Thoracic Surgery (AREA)
- Animal Behavior & Ethology (AREA)
- Human Computer Interaction (AREA)
- Educational Administration (AREA)
- Hospice & Palliative Care (AREA)
- Child & Adolescent Psychology (AREA)
- Developmental Disabilities (AREA)
- Psychiatry (AREA)
- Signal Processing (AREA)
- Computational Linguistics (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Acoustics & Sound (AREA)
- Software Systems (AREA)
- Psychology (AREA)
- Entrepreneurship & Innovation (AREA)
- Radiology & Medical Imaging (AREA)
- Physiology (AREA)
- Dentistry (AREA)
- Oral & Maxillofacial Surgery (AREA)
- Nuclear Medicine, Radiotherapy & Molecular Imaging (AREA)
Abstract
Description
Claims
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| FR1658105A FR3055203A1 (fr) | 2016-09-01 | 2016-09-01 | Prediction de l'attention d'un auditoire lors d'une presentation |
| PCT/FR2017/052314 WO2018042133A1 (fr) | 2016-09-01 | 2017-08-31 | Prédiction de l'attention d'un auditoire lors d'une présentation |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP3506829A1 true EP3506829A1 (fr) | 2019-07-10 |
Family
ID=59381310
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP17772084.4A Pending EP3506829A1 (fr) | 2016-09-01 | 2017-08-31 | Prédiction de l'attention d'un auditoire lors d'une présentation |
Country Status (4)
| Country | Link |
|---|---|
| US (1) | US10942563B2 (fr) |
| EP (1) | EP3506829A1 (fr) |
| FR (1) | FR3055203A1 (fr) |
| WO (1) | WO2018042133A1 (fr) |
Families Citing this family (11)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN108882480B (zh) * | 2018-06-20 | 2020-06-05 | 新华网股份有限公司 | 舞台灯光和装置调节方法和系统 |
| US20200092339A1 (en) * | 2018-09-17 | 2020-03-19 | International Business Machines Corporation | Providing device control instructions for increasing conference participant interest based on contextual data analysis |
| US11514805B2 (en) * | 2019-03-12 | 2022-11-29 | International Business Machines Corporation | Education and training sessions |
| TWI780333B (zh) * | 2019-06-03 | 2022-10-11 | 緯創資通股份有限公司 | 動態處理並播放多媒體內容的方法及多媒體播放裝置 |
| JP7410557B2 (ja) * | 2020-02-04 | 2024-01-10 | 株式会社Agama-X | 情報処理装置及びプログラム |
| US11514924B2 (en) * | 2020-02-21 | 2022-11-29 | International Business Machines Corporation | Dynamic creation and insertion of content |
| EP3975181B1 (fr) * | 2020-09-29 | 2023-02-22 | Bull Sas | Évaluation de la qualité d'une session de communication dans un réseau de télécommunication |
| US12026948B2 (en) * | 2020-10-30 | 2024-07-02 | Microsoft Technology Licensing, Llc | Techniques for presentation analysis based on audience feedback, reactions, and gestures |
| JP7513534B2 (ja) * | 2021-01-12 | 2024-07-09 | 株式会社Nttドコモ | 情報処理装置及び情報処理システム |
| WO2023002496A1 (fr) * | 2021-07-18 | 2023-01-26 | Doshi Payal | Système d'apprentissage en ligne intelligent utilisant la distribution adaptative de cours vidéo en fonction de l'attention du spectateur |
| CN114358135B (zh) * | 2021-12-10 | 2024-02-09 | 西北大学 | 一种利用数据增强和特征加权实现的mooc辍学预测方法 |
Citations (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20140006326A1 (en) * | 2012-06-28 | 2014-01-02 | Nokia Corporation | Method and apparatus for providing rapport management |
Family Cites Families (9)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US8392503B2 (en) * | 2007-06-19 | 2013-03-05 | Cisco Technology, Inc. | Reporting participant attention level to presenter during a web-based rich-media conference |
| US20090138332A1 (en) * | 2007-11-23 | 2009-05-28 | Dimitri Kanevsky | System and method for dynamically adapting a user slide show presentation to audience behavior |
| US20110263946A1 (en) * | 2010-04-22 | 2011-10-27 | Mit Media Lab | Method and system for real-time and offline analysis, inference, tagging of and responding to person(s) experiences |
| US8670018B2 (en) * | 2010-05-27 | 2014-03-11 | Microsoft Corporation | Detecting reactions and providing feedback to an interaction |
| US8965822B2 (en) * | 2011-05-11 | 2015-02-24 | Ari M. Frank | Discovering and classifying situations that influence affective response |
| US9525952B2 (en) * | 2013-06-10 | 2016-12-20 | International Business Machines Corporation | Real-time audience attention measurement and dashboard display |
| US20150332166A1 (en) * | 2013-09-20 | 2015-11-19 | Intel Corporation | Machine learning-based user behavior characterization |
| US10446055B2 (en) * | 2014-08-13 | 2019-10-15 | Pitchvantage Llc | Public speaking trainer with 3-D simulation and real-time feedback |
| US10187694B2 (en) * | 2016-04-07 | 2019-01-22 | At&T Intellectual Property I, L.P. | Method and apparatus for enhancing audience engagement via a communication network |
-
2016
- 2016-09-01 FR FR1658105A patent/FR3055203A1/fr active Pending
-
2017
- 2017-08-31 EP EP17772084.4A patent/EP3506829A1/fr active Pending
- 2017-08-31 US US16/329,429 patent/US10942563B2/en active Active
- 2017-08-31 WO PCT/FR2017/052314 patent/WO2018042133A1/fr not_active Ceased
Patent Citations (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20140006326A1 (en) * | 2012-06-28 | 2014-01-02 | Nokia Corporation | Method and apparatus for providing rapport management |
Non-Patent Citations (3)
| Title |
|---|
| GAN TIAN GANTIAN@COMP NUS EDU SG ET AL: "Multi-sensor Self-Quantification of Presentations", MULTIMEDIA, ACM, 2 PENN PLAZA, SUITE 701 NEW YORK NY 10121-0701 USA, 13 October 2015 (2015-10-13), pages 601 - 610, XP058509717, ISBN: 978-1-4503-3459-4, DOI: 10.1145/2733373.2806252 * |
| JOHN R ZHANG ET AL: "Correlating Speaker Gestures in Political Debates with Audience Engagement Measured via EEG", MULTIMEDIA, ACM, 2 PENN PLAZA, SUITE 701 NEW YORK NY 10121-0701 USA, 3 November 2014 (2014-11-03), pages 387 - 396, XP058058674, ISBN: 978-1-4503-3063-3, DOI: 10.1145/2647868.2654909 * |
| See also references of WO2018042133A1 * |
Also Published As
| Publication number | Publication date |
|---|---|
| FR3055203A1 (fr) | 2018-03-02 |
| US20190212811A1 (en) | 2019-07-11 |
| US10942563B2 (en) | 2021-03-09 |
| WO2018042133A1 (fr) | 2018-03-08 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| EP3506829A1 (fr) | Prédiction de l'attention d'un auditoire lors d'une présentation | |
| US20260010560A1 (en) | Responding to remote media classification queries using classifier models and context parameters | |
| US9058384B2 (en) | System and method for identification of highly-variable vocalizations | |
| JP2011253374A (ja) | 情報処理装置、および情報処理方法、並びにプログラム | |
| EP3140786A1 (fr) | Liaison d'activités utilisateur courantes à des collections de contenus multimédias stockées associées | |
| CN102244788A (zh) | 信息处理方法、信息处理装置、场景元数据提取装置、丢失恢复信息生成装置和程序 | |
| CN107493501B (zh) | 一种音视频内容过滤系统及方法 | |
| Ramsay et al. | The intrinsic memorability of everyday sounds | |
| US11157549B2 (en) | Emotional experience metadata on recorded images | |
| CN115866339A (zh) | 电视节目推荐方法、装置、智能设备及可读存储介质 | |
| Wu et al. | Cold start problem for automated live video comments | |
| CN105284121A (zh) | 多媒体流和社交网络线程之间的同步 | |
| EP3556102B1 (fr) | Procede d'enregistrement d'un programme telediffuse a venir | |
| EP3107302B1 (fr) | Procédé et dispositif de substitution d'une partie d'une sequence video | |
| CA2592994A1 (fr) | Procede de recherche d'informations dans une base de donnees | |
| EP4198971A1 (fr) | Method for selecting voice contents recorded in a database, according to their veracity factor | |
| FR3130422A1 (fr) | Procédé de sélection de contenus vocaux en- registrés dans une base de données, en fonction de leur facteur de véracité. | |
| JP7572200B2 (ja) | キーワード抽出装置、キーワード抽出プログラム及び発話生成装置 | |
| CN120689477B (zh) | 一种基于数据驱动的数字人优质视频生成方法及系统 | |
| WO2024208593A1 (fr) | Procédé de génération d'une séquence temporelle d'évaluations d'une situation et dispositif associé | |
| WO2018077987A1 (fr) | Procédé de traitement de données audio issues d'un échange vocal, système et programme d'ordinateur correspondant | |
| JP6852191B2 (ja) | フィンガープリントを変換して不正なメディアコンテンツアイテムを検出するための方法、システムおよび媒体 | |
| JP2026074997A (ja) | システム | |
| EP4375899A1 (fr) | Procede et dispositif de recommandation d'activites a au moins un utilisateur | |
| WO2023007061A1 (fr) | Procédé de traitement d'informations, terminal de télécommunication et programme d'ordinateur |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: UNKNOWN |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE |
|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE |
|
| 17P | Request for examination filed |
Effective date: 20190122 |
|
| AK | Designated contracting states |
Kind code of ref document: A1 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC MK MT NL NO PL PT RO RS SE SI SK SM TR |
|
| AX | Request for extension of the european patent |
Extension state: BA ME |
|
| DAV | Request for validation of the european patent (deleted) | ||
| DAX | Request for extension of the european patent (deleted) | ||
| RAP1 | Party data changed (applicant data changed or rights of an application transferred) |
Owner name: ORANGE |
|
| RAP3 | Party data changed (applicant data changed or rights of an application transferred) |
Owner name: ORANGE |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: EXAMINATION IS IN PROGRESS |
|
| 17Q | First examination report despatched |
Effective date: 20221104 |