WO2018079106A1 - 感情推定装置、感情推定方法、記憶媒体、および感情カウントシステム - Google Patents
感情推定装置、感情推定方法、記憶媒体、および感情カウントシステム Download PDFInfo
- Publication number
- WO2018079106A1 WO2018079106A1 PCT/JP2017/032819 JP2017032819W WO2018079106A1 WO 2018079106 A1 WO2018079106 A1 WO 2018079106A1 JP 2017032819 W JP2017032819 W JP 2017032819W WO 2018079106 A1 WO2018079106 A1 WO 2018079106A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- information
- emotion
- utterance
- analysis unit
- pulse
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- A—HUMAN NECESSITIES
- A61—MEDICAL OR VETERINARY SCIENCE; HYGIENE
- A61B—DIAGNOSIS; SURGERY; IDENTIFICATION
- A61B5/00—Measuring for diagnostic purposes; Identification of persons
- A61B5/02—Detecting, measuring or recording for evaluating the cardiovascular system, e.g. pulse, heart rate, blood pressure or blood flow
-
- A—HUMAN NECESSITIES
- A61—MEDICAL OR VETERINARY SCIENCE; HYGIENE
- A61B—DIAGNOSIS; SURGERY; IDENTIFICATION
- A61B5/00—Measuring for diagnostic purposes; Identification of persons
- A61B5/02—Detecting, measuring or recording for evaluating the cardiovascular system, e.g. pulse, heart rate, blood pressure or blood flow
- A61B5/024—Measuring pulse rate or heart rate
- A61B5/0245—Measuring pulse rate or heart rate by using sensing means generating electric signals, i.e. ECG signals
-
- A—HUMAN NECESSITIES
- A61—MEDICAL OR VETERINARY SCIENCE; HYGIENE
- A61B—DIAGNOSIS; SURGERY; IDENTIFICATION
- A61B5/00—Measuring for diagnostic purposes; Identification of persons
- A61B5/103—Measuring devices for testing the shape, pattern, colour, size or movement of the body or parts thereof, for diagnostic purposes
- A61B5/107—Measuring physical dimensions, e.g. size of the entire body or parts thereof
-
- A—HUMAN NECESSITIES
- A61—MEDICAL OR VETERINARY SCIENCE; HYGIENE
- A61B—DIAGNOSIS; SURGERY; IDENTIFICATION
- A61B5/00—Measuring for diagnostic purposes; Identification of persons
- A61B5/16—Devices for psychotechnics; Testing reaction times ; Devices for evaluating the psychological state
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T7/00—Image analysis
- G06T7/20—Analysis of motion
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/08—Speech classification or search
- G10L15/10—Speech classification or search using distance or distortion measures between unknown speech and reference templates
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L25/00—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
- G10L25/48—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use
- G10L25/51—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use for comparison or discrimination
- G10L25/63—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use for comparison or discrimination for estimating an emotional state
Definitions
- Embodiments described herein relate generally to an emotion estimation device, an emotion estimation method, a storage medium, and an emotion count system.
- a technique for detecting an emotion of a target person based on a face image has been studied.
- a facial expression is determined from a face image, and an emotion is detected from the facial expression of the determination result.
- emotions can be detected only from appearance features such as smiles, and actual people such as people who are difficult to express emotions or people who intentionally hide emotions. It may be difficult to detect emotions.
- an emotion estimation device an emotion estimation method, a storage medium, and an estimation device capable of estimating an accurate emotion of an estimation target by comprehensively determining a plurality of pieces of information related to the estimation target. It is to provide an emotion counting system.
- the emotion estimation apparatus of the embodiment includes a pulse analysis unit, a facial expression analysis unit, an utterance content analysis unit, an utterance tone analysis unit, and an emotion estimation unit.
- the pulse analysis unit generates pulse information by analyzing the pulse of the estimated target.
- the facial expression analysis unit generates facial expression information by analyzing the facial expression of the estimated target based on image information of the face of the estimated target.
- the utterance content analysis unit analyzes the utterance content of the estimated target based on the speech information of the estimated target and generates utterance content information.
- the utterance tone analysis unit analyzes the utterance tone of the estimated target based on the voice information and generates utterance tone information.
- the emotion estimation unit estimates the emotion of the estimation target based on at least two of the pulse information, the facial expression information, the utterance content information, and the utterance tone information.
- the figure which shows an example of the hardware constitutions of the emotion estimation system using the emotion estimation apparatus of 2nd Embodiment The flowchart which shows an example of the process of the emotion estimation apparatus of 2nd Embodiment.
- the emotion estimation device of the embodiment inputs not only the facial expression of the estimation target face that is the target of emotion estimation but also multimodal output information such as pulse, utterance content, utterance tone, etc., and comprehensively determines them. Estimate emotions that are more sensitive and more accurate. In the following, an example in which “human” is targeted as an estimation target will be described.
- FIG. 1 is a diagram illustrating an example of a hardware configuration of an emotion estimation system A1 using the emotion estimation device 1 of the present embodiment.
- the emotion estimation system A1 includes, for example, an emotion estimation device 1, a network N, a target person terminal P1, and a user terminal P2.
- the emotion estimation device 1 estimates the emotion of the target person T based on the image information, audio information, and the like of the face of the target person T (estimated target).
- the network N connects the subject terminal P1 and the user terminal P2 so that they can communicate, and connects the user terminal P2 and the emotion estimation device 1 so that they can communicate.
- the network N is realized by an IP communication network, various telephone networks, other wireless communication networks, a wired communication network, and the like.
- the target person terminal P1 acquires the image information and voice information of the target person T, and transmits them to the user terminal P2 via the network N.
- the user terminal P2 transmits the image information and audio information received from the target person terminal P1 to the emotion estimation device 1 via the network N.
- the emotion estimation device 1 estimates the emotion of the target person T based on the image information and audio information received from the user terminal P2, and transmits the estimated emotion information to the user terminal P2.
- the target person terminal P1 and the user terminal P2 are, for example, a PC terminal or a mobile phone having a videophone function.
- the user terminal P ⁇ b> 2 may be installed with an application that realizes an interface that can communicate with the emotion estimation device 1 and that can display emotion information received from the emotion estimation device 1.
- FIG. 2 is a functional block diagram showing an example of the emotion estimation apparatus 1 of the present embodiment.
- the emotion estimation device 1 includes, for example, a communication unit 10, an analysis unit 12, a history storage unit D1, and a dictionary storage unit D3.
- the communication unit 10 receives, for example, image information and voice information of the face of the target person T from the user terminal P2 via the network N, and outputs them to the analysis unit 12. For example, the communication unit 10 transmits the estimated emotion information to the user terminal P2 via the network N.
- the analysis unit 12 performs emotion estimation of the subject T based on the image information and voice information of the subject T input from the communication unit 10.
- the analysis unit 12 includes, for example, an image analysis unit 20, a voice analysis unit 30, and an emotion estimation unit 40.
- the image analysis unit 20 analyzes the image information of the subject T input from the communication unit 10, and outputs the analysis result to the emotion estimation unit 40. As shown in FIG. 3, for example, the image analysis unit 20 analyzes the image information, generates pulse information related to the pulse of the face of the subject T, and outputs the pulse information to the emotion estimation unit 40. And a facial expression analysis unit 24 that generates facial expression information related to the facial expression of the subject T and outputs the facial expression information to the emotion estimation unit 40.
- the pulse analysis unit 22 estimates the pulse wave propagation time based on, for example, a signal absorbed by hemoglobin in the subcutaneous blood of the face of the subject T acquired from the image information.
- the pulse analysis unit 22 generates pulse information related to the pulse of the subject T using the estimated pulse wave propagation time having a positive correlation with blood pressure fluctuation.
- the pulse information may include numerical data (index value) of the pulse of the subject T.
- the pulse analysis unit 22 outputs the image information to the emotion estimation unit 40 without performing the above-described pulse analysis processing, and uses it as a part of information for estimating the total emotion in the emotion estimation unit 40. Also good.
- the pulse analysis unit 22 can also analyze the pulse from the image information of the body other than the face.
- the pulse analysis unit 22 can also analyze the pulse from a partial image such as image information of a part of the face.
- the pulse analysis unit 22 may analyze the pulse from a sphygmomanometer, an electrocardiogram result, or a wearable pulse wave sensor.
- the facial expression analysis unit 24 associates emotion types such as laughter, anger, sadness, and joy with the above image information, and outputs facial expression information related to the facial expression of the subject T.
- emotion types such as laughter, anger, sadness, and joy
- the facial expression analysis unit 24 associates the above image information with a more detailed type considering the degree of “Laughter”, “Smile”, etc. As a result, more detailed types such as “furious” and “small anger” may be linked.
- the facial expression information may include data (index value) in which the emotion of the subject T is digitized.
- the data in which this emotion is digitized is a numerical value representing the degree of each emotion such as laughter, anger, sadness, and joy.
- the facial expression analysis unit 24 outputs the image information to the emotion estimation unit 40 without performing the facial expression analysis process, and uses it as a part of information for estimating the total emotion in the emotion estimation unit 40. Also good.
- the voice analysis unit 30 analyzes the voice information of the target person T input from the communication unit 10 and outputs the analysis result to the emotion estimation unit 40.
- the speech analysis unit 30 for example, analyzes speech information, generates speech content information related to the speech content of the target person T, and outputs the speech content information to the emotion estimation unit 40.
- An utterance tone analysis unit 34 that analyzes information and generates utterance tone information related to the tone of the utterance of the target person T and outputs the utterance tone information to the emotion estimation unit 40.
- the utterance content analysis unit 32 converts the speech data included in the speech information into text data and refers to the dictionary data stored in the dictionary storage unit D3 to generate the utterance content information.
- the utterance content information may include data (index value) in which the emotion of the subject T is digitized. For example, when the utterance content of the target person T includes a negative expression such as “Funny” or “No”, the utterance content analysis unit 32 sets a high value indicating “anger”. On the other hand, the utterance content analysis unit 32 increases the numerical value indicating “joy” when the utterance content of the target person T includes positive expressions such as “thank you” and “saved”, for example. Set.
- the utterance content analysis unit 32 outputs the speech information to the emotion estimation unit 40 without performing the above utterance content analysis process, and uses it as a part of information for estimating the total emotion in the emotion estimation unit 40. It may be.
- the utterance tone analysis unit 34 generates utterance tone information based on at least one of the volume, speed, and frequency of the utterance, for example.
- the utterance tone information may include data (index value) in which the emotion of the subject T is digitized.
- the utterance tone analysis unit 34 acquires an average value of each of the volume, speed, and frequency for the target person T based on continuously acquired audio information (time-series audio information), and calculates the average value from the average value.
- Utterance tone information is generated based on the deviation.
- the utterance tone analysis unit 34 generates utterance tone information based on the rate of change of each of volume, speed, and frequency.
- the utterance tone analysis unit 34 determines that the emotion has increased, for example, when the volume has increased rapidly. The utterance tone analysis unit 34 determines that the emotion has increased when the speaking speed has increased. The utterance tone analysis unit 34 refers to the frequency that is a parameter representing the pitch of the voice, and determines that the emotion has increased when the frequency has increased. If the utterance tone analysis unit 34 determines that the emotion has increased, for example, the utterance tone analysis unit 34 sets a high numerical value indicating “anger” as a candidate for emotion. Alternatively, the utterance tone analysis unit 34 acquires volume, speed, and frequency from audio information at a certain moment, and at least one of the acquired volume, speed, and frequency uses a predetermined reference value for each volume, speed, and frequency.
- the utterance tone information is generated based on whether it has been exceeded.
- the utterance tone analysis unit 34 determines that the emotion has increased when the volume is larger than the reference value, for example.
- the utterance tone analysis unit 34 outputs voice information to the emotion estimation unit 40 without performing the utterance tone analysis process described above, and uses it as a part of information for estimating the total emotion in the emotion estimation unit 40. It may be.
- the emotion estimation unit 40 is based on at least two pieces of information from among the image information of the subject T, the speech information of the subject T, and the text information obtained by converting the speech data included in the speech information of the subject T into text data. Emotion estimation of the target person T is performed.
- the emotion estimation unit 40 receives the pulse information input from the pulse analysis unit 22, the facial expression information input from the facial expression analysis unit 24, the utterance content information input from the utterance content analysis unit 32, and the utterance tone analysis unit 34. Based on at least two of the input utterance tone information, emotion estimation of the target person T is performed.
- the emotion estimation unit 40 transmits the estimated emotion information of the target person T to the user terminal P ⁇ b> 2 via the communication unit 10.
- the user U can know the emotion of the target person T by confirming the emotion information displayed on the display of the user terminal P2.
- a multilayer neural network (Deep Neural Network: DNN), a convolutional neural network (Convolutional Neural Network: CNN), a recurrent neural network (RNN), or the like is used. Deep learning technology may be adopted.
- the emotion estimation unit 40 may perform in advance a supervised learning process in which data relating to pulse information, facial expression information, utterance content information, and utterance tone information and emotion estimation result data are input. In this case, the emotion estimation unit 40 assigns a weighting as to which information of the pulse information, facial expression information, utterance content information, and utterance tone information is important for each attribute (gender, age, nationality, etc.) of the target person T. You may go to
- the emotion estimation unit 40 may perform pulse analysis processing based on the image information.
- the facial expression analysis unit 24 outputs image information to the emotion estimation unit 40 without performing facial expression analysis processing
- the emotion estimation unit 40 may perform facial expression analysis processing based on the image information.
- the utterance content analysis unit 32 outputs voice information to the emotion estimation unit 40 without performing the utterance content analysis processing
- the emotion estimation unit 40 may perform the utterance content analysis processing based on the voice information.
- the utterance tone analysis unit 34 outputs the voice information to the emotion estimation unit 40 without performing the utterance tone analysis process
- the emotion estimation unit 40 may perform the utterance tone analysis process based on the voice information.
- the emotion estimation unit 40 may perform comprehensive emotion estimation using the information obtained from each analysis unit as it is, without performing each analysis process from the information obtained from each analysis unit.
- the history storage unit D1 stores various information related to the target person T.
- the history storage unit D1 includes, for example, image information of the subject T, voice information, pulse information analyzed by the pulse analysis unit 22, facial expression information analyzed by the facial expression analysis unit 24, and speech analyzed by the speech content analysis unit 32
- the content information, the utterance tone information analyzed by the utterance tone analysis unit 34, the emotion information estimated by the emotion estimation unit 40, the telephone number, the name, and the like are stored.
- the dictionary storage unit D3 stores, for example, dictionary data for interpreting the content of the voice information of the subject T.
- FIG. 5 is a diagram illustrating an example of dictionary data stored in the dictionary storage unit D3.
- the dictionary storage unit D3 stores dictionary data in which a character string is associated with a numerical value representing the degree of each emotion such as laughter, anger, sadness, and joy.
- a numerical value representing the degree of each emotion such as laughter, anger, sadness, and joy.
- the numerical value of the degree of each emotion for example, a numerical value from 1 to 5 is assigned, and the larger the numerical value, the higher the degree of each emotion.
- the dictionary storage unit D3 has an anger value of “3”, a laughter value of “1”, and a joy value of “0” as information associated with the character string “Funny”. Information with a sadness value of “1” is stored. That is, the character string “Funny” indicates a high tendency to express anger.
- the dictionary storage unit D3 may store a large number of dictionary data in addition to the data shown in FIG. For example, the dictionary storage unit D3 may further store information on other items for more accurately interpreting the content of the voice information.
- Each of the history storage unit D1 and the dictionary storage unit D3 is realized by a ROM (Read Only Memory), a RAM (Random Access Memory), an HDD (Hard Disk Drive), a flash memory, or the like.
- the history storage unit D1 and the dictionary storage unit D3 may be configured by a single piece of hardware.
- Each of the history storage unit D1 and the dictionary storage unit D3 may be provided outside the emotion estimation device 1.
- Some or all of the functional units of the emotion estimation apparatus 1 may be realized by a processor executing a program (software) stored in a program storage unit (not shown).
- the program may be installed in advance when the operation of the emotion estimation apparatus 1 is started, may be downloaded from another computer, or may be installed from a portable storage medium such as a compact disk.
- Some or all of the functional units of the emotion estimation device 1 may be realized by hardware such as LSI (Large Scale Integration), ASIC (Application Specific Integrated Circuit), or FPGA (Field-Programmable Gate Array). It may be realized by a combination of software and hardware.
- FIG. 6 is a flowchart illustrating an example of processing of the emotion estimation apparatus 1.
- the user terminal P2 While the user U and the target person T are communicating (for example, during a conversation by videophone), the user terminal P2 is connected to the target person T from the target terminal P1 via the network N. Face image information and audio information are continuously acquired.
- the user terminal P2 transmits the acquired image information and audio information to the emotion estimation apparatus 1.
- the communication part 10 of the emotion estimation apparatus 1 acquires the image information and audio
- the communication unit 10 outputs the acquired image information to the image analysis unit 20 and outputs the acquired audio information to the audio analysis unit 30.
- the pulse analysis unit 22 of the image analysis unit 20 analyzes the image information to generate pulse information related to the pulse of the subject T (step S103). For example, the pulse analysis unit 22 estimates the pulse wave propagation time based on a signal absorbed by hemoglobin in the subcutaneous blood of the face of the subject T acquired from the image information. The pulse analysis unit 22 generates pulse information by utilizing the fact that the estimated pulse wave propagation time is positively correlated with blood pressure fluctuation. The pulse analysis unit 22 outputs the generated pulse information to the emotion estimation unit 40. When the past pulse information of the subject T is stored in the history storage unit D1, the pulse analysis unit 22 generates more accurate pulse information by referring to the past pulse information together. Good.
- the facial expression analysis unit 24 of the image analysis unit 20 analyzes the image information to generate facial expression information related to the facial expression of the subject T (step S105).
- the facial expression analysis unit 24 generates facial expression information by associating image information with types of emotions such as laughter, anger, sadness, and joy.
- the facial expression analysis unit 24 may associate a more detailed type considering the degree of “Laughter”, “Smile”, etc., as the emotion of “laughter” with respect to the image information.
- the facial expression analysis unit 24 outputs the generated facial expression information to the emotion estimation unit 40.
- the facial expression analysis unit 24 When the facial expression information of the subject T is stored in the history storage unit D1, the facial expression analysis unit 24 generates more accurate facial expression information by referring to the past facial expression information together. Good.
- the speech content analysis unit 32 of the speech analysis unit 30 analyzes the speech information and generates speech content information related to the speech content of the target person T (step S107). For example, the utterance content analysis unit 32 converts the speech data included in the speech information into text data, and generates the utterance content information by referring to the dictionary data stored in the dictionary storage unit D3. The utterance content analysis unit 32 determines whether or not each character string included in the converted text data is stored in the dictionary storage unit D3. If the character string is stored, the utterance content analysis unit 32 calculates a numerical value indicating the degree of each emotion. Extract.
- the utterance content analysis unit 32 may use a value obtained by summing up numerical values for each emotion as the utterance content information.
- the utterance content analysis unit 32 outputs the generated utterance content information to the emotion estimation unit 40.
- the utterance content analysis unit 32 refers to the past utterance content information together, thereby more accurate utterance content information. May be generated.
- the utterance tone analysis unit 34 of the speech analysis unit 30 analyzes the speech information and generates utterance tone information related to the utterance tone of the target person T (step S109). For example, the utterance tone analysis unit 34 generates utterance tone information based on changes in the volume, speed, and frequency of the utterance. The utterance tone analysis unit 34 outputs the generated utterance tone information to the emotion estimation unit 40.
- the utterance tone analysis unit 34 refers to the past utterance tone information together, thereby more accurate utterance tone information. May be generated.
- the analysis processes of the pulse analysis unit 22, the expression analysis unit 24, the utterance content analysis unit 32, and the utterance tone analysis unit 34 may be executed in an arbitrary order or may be executed in parallel. .
- the emotion estimation unit 40 includes pulse information input from the pulse analysis unit 22, facial expression information input from the facial expression analysis unit 24, utterance content information input from the utterance content analysis unit 32, and utterance tone analysis unit 34. Based on the utterance tone information input from, the emotion estimation of the target person T is performed (step S111).
- the emotion estimation unit 40 transmits the estimated emotion information to the user terminal P2 via the communication unit 10 (step S113).
- the user terminal P2 displays the emotion information received from the communication unit 10, for example, on the display of the user terminal P2.
- FIG. 7 is a diagram illustrating an example of an emotion estimation result displayed on the display of the user terminal P2.
- the estimated emotion information of the target person T is displayed in the area S1 of the display screen of the display, the image of the target person T is displayed in the area S2, and the target object is displayed in the area S3.
- the profile information of the person T is displayed.
- the estimated emotion information of the subject T is “L1: small anger”.
- the user U can know the emotion of the target person T by confirming the emotion information displayed on the display.
- the profile information of the target person T may be input and edited by the user U using the user terminal P2.
- the profile information of the target person T estimated using a technique for estimating age sex and the like from face information may be displayed in the region S3. Thus, the process of this flowchart is completed.
- the process of the above flowchart is repeatedly executed, and emotion information of the target person T corresponding to the situation is obtained. It is displayed on the display of the user terminal P2.
- the estimated target stress level, excitement level (degree of calmness), etc. can be found.
- Direct emotions appearing on the appearance of the estimated target can be estimated from the facial expression information.
- the content of the utterance content information it can be inferred whether the utterance of the presumed target is derived from a negative emotion such as “anger” or a positive emotion such as “joy”.
- the utterance tone information it is possible to infer emotions such as whether the subject T is excited, calm, angry, or happy.
- an accurate emotion of the estimation target by comprehensively determining a plurality of pieces of multimodal information as described above. For example, even with the same “anger” emotion, it is possible to read an accurate emotion such as the nuance of the detailed emotion. A certain amount of guess can be made by just the appearance of the facial expression. However, for example, when it is necessary to read a customer's high-level emotions such as call center operations, it may be a serious situation if a mistake is made. According to the present embodiment, since accurate and detailed emotions of the estimation target can be estimated, it is possible to support customer service operations such as call center operations and customer service operations. By estimating the driver's feelings, the driver can be encouraged to drive safely. When the pulse analysis unit 22 analyzes the image information and generates the pulse information, it is possible to perform emotion estimation in consideration of the user's biometric information without using a device that acquires biometric information of the estimation target such as a biometric sensor.
- the second embodiment Compared to the first embodiment, the second embodiment is different in that the emotion estimation device 1 is mounted on the user terminal P2. For this reason, about the structure etc., the figure and related description which were demonstrated in 1st Embodiment are used, and detailed description is abbreviate
- FIG. 8 is a diagram illustrating an example of a hardware configuration of an emotion estimation system A2 using the emotion estimation device 1 of the present embodiment.
- the emotion estimation system A2 includes, for example, a subject terminal P1, a user terminal P2 on which the emotion estimation device 1 is mounted (hereinafter referred to as “emotion estimation device 1”), a subject terminal P1, the emotion estimation device 1, and the like. And a network N connecting the two.
- the emotion estimation device 1 receives the image information and audio information of the subject T from the subject terminal P1 via the network N.
- FIG. 9 is a flowchart illustrating an example of processing of the emotion estimation apparatus 1.
- the communication unit 10 of the emotion estimation apparatus 1 acquires, for example, image information and audio information of the face of the subject T from the subject terminal P1 via the network N (step S201).
- the communication unit 10 outputs the acquired image information to the image analysis unit 20 and outputs the acquired audio information to the audio analysis unit 30.
- the pulse analysis unit 22 of the image analysis unit 20 analyzes the image information and generates pulse information related to the pulse of the subject T (step S203).
- the pulse analysis unit 22 outputs the generated pulse information to the emotion estimation unit 40.
- the facial expression analysis unit 24 of the image analysis unit 20 analyzes the image information and generates facial expression information related to the facial expression of the subject T (step S205).
- the facial expression analysis unit 24 outputs the generated facial expression information to the emotion estimation unit 40.
- the speech content analysis unit 32 of the speech analysis unit 30 analyzes the speech information and generates speech content information related to the speech content of the target person T (step S207).
- the utterance content analysis unit 32 outputs the generated utterance content information to the emotion estimation unit 40.
- the utterance tone analysis unit 34 of the speech analysis unit 30 analyzes the speech information and generates utterance tone information related to the tone of the utterance of the target person T (step S209).
- the utterance tone analysis unit 34 outputs the generated utterance tone information to the emotion estimation unit 40.
- the analysis processes of the pulse analysis unit 22, the expression analysis unit 24, the utterance content analysis unit 32, and the utterance tone analysis unit 34 may be executed in an arbitrary order or may be executed in parallel. .
- the emotion estimation unit 40 includes pulse information input from the pulse analysis unit 22, facial expression information input from the facial expression analysis unit 24, utterance content information input from the utterance content analysis unit 32, and utterance tone analysis unit 34. Based on the utterance tone information input from, emotion estimation of the target person T is performed (step S211).
- the emotion estimation unit 40 displays the estimated emotion information of the subject T on, for example, a display unit (not shown). Thus, the process of this flowchart is completed.
- the emotion estimation unit 40 may perform emotion estimation and generate support information for supporting the user U's action determination and display it on the user terminal P2. For example, when a customer-facing operator at a call center uses the emotion estimation device 1, the emotion estimation unit 40 generates support information including answer contents to customers, proposal contents, and the like and displays them on the user terminal P 2. Also good.
- a stand-alone emotion estimation device 1 that does not require communication with other terminals via a network is used as a smile mentor as a system for educating smiles, utterances, etc., which are desirable for customer service. May be.
- the emotion estimation unit 40 may perform emotion estimation in consideration of information related to the surrounding situation of the subject T (weather, temperature, time zone, location, etc.).
- the dictionary storage unit D3 may store dictionary data customized according to the usage of the emotion estimation apparatus 1. For example, when the target person T is in an area where the temperature is very high, the target person T is estimated to be a little unhappy.
- the information regarding the surrounding situation of the target person T is acquired from, for example, the background image, position information, shooting time, and the like of the target person T.
- the estimation target may include various living bodies such as humans and other animals, or non-living bodies such as artificial intelligence.
- the third embodiment is different in that a function for counting the total number of facial expressions (smiling faces, angry faces, etc.) of the subject is implemented in the emotion estimation apparatus. For this reason, about the structure etc., the figure and related description which were demonstrated in 1st Embodiment are used, and detailed description is abbreviate
- FIG. 10 is a diagram illustrating an example of a hardware configuration of an emotion count system A3 using the emotion estimation apparatus 2 of the present embodiment.
- the emotion count system A3 includes, for example, the emotion estimation device 2, and at least one terminal (terminal R1, terminal R2, and terminal R3 in the example of FIG. 10) connected to the emotion estimation device 2 via the network N.
- a camera C1, a camera C2, and a camera C3 are connected to each of the terminal R1, the terminal R2, and the terminal R3.
- the camera C1, the camera C2, and the camera C3 are installed in various facilities such as an exhibition hall, for example, and are used in an environment where a large number of unspecified subjects T are photographed.
- the camera C1, the camera C2, and the camera C3 are installed in different places in various facilities.
- the camera C1 captures the face image of the subject T and outputs the captured face image to the terminal R1.
- the terminal R1 transmits the face image input from the camera C1 to the emotion estimation device 2.
- the camera C2 captures the face image of the subject T and outputs the captured face image to the terminal R2.
- the terminal R2 transmits the face image input from the camera C2 to the emotion estimation device 2.
- the camera C3 captures the face image of the subject T and outputs the captured face image to the terminal R3.
- the terminal R3 transmits the face image input from the camera C3 to the emotion estimation device 2.
- the target person T photographed by the camera C2 and the camera C3 may be different persons or the same person.
- FIG. 11 is a functional block diagram illustrating an example of the emotion estimation apparatus 2 of the present embodiment.
- the emotion estimation device 2 further includes a counting unit 50 that counts the total number of facial expressions.
- the facial expression analysis unit 24 is configured to output facial expression information to the counting unit 50.
- the counting unit 50 counts the total number for each facial expression analyzed by the analyzing unit 12 (facial expression analyzing unit 24). For example, when the facial expression analysis unit 24 determines that the facial expression of the face image of the subject T received from the terminal R1 is “smile”, the counting unit 50 includes, for example, a memory (not shown) provided therein. Is stored with the total number “1” of “smile” and the facial expression information of “smile” and the total number “1” are transmitted to the terminal R 1 via the communication unit 10. After that, the facial expression analysis unit 24 performs facial expression analysis processing on the facial image of the other target person T received from the terminal R1, and when it is determined that the facial expression of the other target person T is “smile”, the count is performed. The unit 50 stores “2” obtained by incrementing the total number of “smiles” by 1 in a memory provided therein, and also stores the facial expression information of “smiles” and the total number “2” via the communication unit 10 to the terminal R1. Send.
- the counting unit 50 may perform the counting process for each terminal, or may perform the counting process for a plurality of terminals and calculate the sum for each facial expression.
- FIG. 12 is a diagram illustrating an example of a “smile” count result displayed on the terminal R1 (display device) of the present embodiment.
- the count result of “smile” of images acquired by a plurality of terminals (terminal R1, terminal R2, terminal R3) is displayed on the display of terminal R1.
- the total number “128” in which the facial expression is determined to be “smile” with respect to the face image photographed by the camera C1 connected to the terminal R1 is displayed in the area W1.
- the total number of facial expressions photographed by the camera C1 with the facial expression determined as “smiling”, the total number of facial expressions photographed by the camera C2 with the facial expression determined as “smiling”, and the camera C3 The total “539” of the total number of facial expressions determined as “smile” for the face image photographed by the above is displayed in the area W2.
- An image of the subject T is displayed in the area W3, and other information is displayed in the area W4.
- the display result shown in FIG. 12 is an example, and for example, it may be displayed as a graph indicating the transition of the total count of “smile” in time series. Of the count results by the counting unit 50, the count number for each facial expression counted within a predetermined time may be displayed.
- the total number for each facial expression of the estimated target can be counted.
- Emotion estimation device 2 may be implemented in terminal R1, terminal R2, and terminal R3, respectively.
- the terminal R1, the terminal R2, and the terminal R3 in which the emotion estimation device 2 is mounted communicate with each other to transmit and receive information on the total number. Good.
- the facial expression analysis unit 24 can analyze the facial expression of each target person T when a plurality of target persons T are included in one image. For example, the facial expression analysis unit 24 outputs information indicating that there are three “smiles” to the counting unit 50 when three images of the smile target person T are included in one image.
- the counting unit 50 stores the total number “3” of “smiles” in an internal memory based on the information indicating that there are three “smiles” input from the facial expression analysis unit 24 (or In addition, the expression information of “smile” and the total number “3” (or a numerical value incremented by 3) are transmitted to the terminal R1 via the communication unit 10.
- the counting unit 50 counts the total number analyzed as “smile” by the facial expression analysis unit 24.
- the counting unit 50 may count facial expressions other than smiles.
- the analysis unit 12 may estimate the emotion of the target person T based on not only the analysis result by the facial expression analysis unit 24 but also other analysis results, and the counting unit 50 may count the total number for each estimated emotion.
- the analysis unit 12 may estimate the emotion of the target person T using facial expressions, utterance contents, and utterance tones, and the counting unit 50 may count a person who makes an angry utterance. That is, the emotion estimation device 2 of the present embodiment can estimate emotions by comprehensively judging a plurality of information related to the estimation target, and can count the total number of each estimated emotion.
- the counting unit 50 may count the total number for each index value included in each of pulse information, facial expression information, utterance content information, and utterance tone information. For example, by counting the total number for each index value (pulse value) included in the pulse information, it is possible to estimate the stress level, excitement level, etc. of the subject T.
- the emotion estimation apparatus 2 of the present embodiment may track the same person from the face image and count for each target person T photographed by different cameras, identify a person with a high degree of tension, and perform a disturbing movement. It may be used to detect suspicious persons.
- a pulse analysis unit that analyzes the pulse of the estimated target and generates pulse information, and analyzes the facial expression of the estimated target based on the image information of the face of the estimated target.
- a facial expression analysis unit that generates facial expression information
- an utterance content analysis unit that analyzes utterance content of the estimation target and generates utterance content information based on the speech information of the estimation target
Landscapes
- Health & Medical Sciences (AREA)
- Engineering & Computer Science (AREA)
- Life Sciences & Earth Sciences (AREA)
- Physics & Mathematics (AREA)
- General Health & Medical Sciences (AREA)
- Surgery (AREA)
- Veterinary Medicine (AREA)
- Pathology (AREA)
- Animal Behavior & Ethology (AREA)
- Molecular Biology (AREA)
- Medical Informatics (AREA)
- Heart & Thoracic Surgery (AREA)
- Biomedical Technology (AREA)
- Public Health (AREA)
- Biophysics (AREA)
- Multimedia (AREA)
- Cardiology (AREA)
- Acoustics & Sound (AREA)
- Signal Processing (AREA)
- Psychiatry (AREA)
- Hospice & Palliative Care (AREA)
- Physiology (AREA)
- Child & Adolescent Psychology (AREA)
- Human Computer Interaction (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Computational Linguistics (AREA)
- Developmental Disabilities (AREA)
- Dentistry (AREA)
- Oral & Maxillofacial Surgery (AREA)
- Social Psychology (AREA)
- Psychology (AREA)
- Educational Technology (AREA)
- Computer Vision & Pattern Recognition (AREA)
- General Physics & Mathematics (AREA)
- Theoretical Computer Science (AREA)
- Measuring Pulse, Heart Rate, Blood Pressure Or Blood Flow (AREA)
- Measurement Of The Respiration, Hearing Ability, Form, And Blood Characteristics Of Living Organisms (AREA)
- Image Analysis (AREA)
- User Interface Of Digital Computer (AREA)
Abstract
実施形態の感情推定装置は、脈拍解析部と、表情解析部と、発話内容解析部と、発話トーン解析部と、感情推定部とを持つ。脈拍解析部は、推定ターゲットの脈拍を解析して脈拍情報を生成する。表情解析部は、推定ターゲットの顔の画像情報に基づいて、推定ターゲットの表情を解析して表情情報を生成する。発話内容解析部は、推定ターゲットの音声情報に基づいて、推定ターゲットの発話内容を解析して発話内容情報を生成する。発話トーン解析部は、音声情報に基づいて、推定ターゲットの発話トーンを解析して発話トーン情報を生成する。感情推定部は、脈拍情報、表情情報、発話内容情報、および発話トーン情報のうち少なくとも2つに基づいて、推定ターゲットの感情を推定する。
Description
本発明の実施形態は、感情推定装置、感情推定方法、記憶媒体、および感情カウントシステムに関する。
従来、顔画像に基づいて対象となる人物の感情を検出する技術が研究されている。この従来技術は、例えば顔画像から表情を判定し、判定結果の表情から感情を検出する。しかしながら、上記の従来技術においては、笑顔などの外見上の特徴からしか感情を検出することができず、感情が表情に出にくい人物や、意図的に感情を隠そうとする人物などの実際の感情を検出することが難しい場合がある。
本発明が解決しようとする課題は、推定ターゲットに関連する複数の情報を総合的に判断することによって推定ターゲットの正確な感情を推定することができる感情推定装置、感情推定方法、記憶媒体、および感情カウントシステムを提供することである。
実施形態の感情推定装置は、脈拍解析部と、表情解析部と、発話内容解析部と、発話トーン解析部と、感情推定部とを持つ。前記脈拍解析部は、推定ターゲットの脈拍を解析して脈拍情報を生成する。前記表情解析部は、前記推定ターゲットの顔の画像情報に基づいて、前記推定ターゲットの表情を解析して表情情報を生成する。前記発話内容解析部は、前記推定ターゲットの音声情報に基づいて、前記推定ターゲットの発話内容を解析して発話内容情報を生成する。前記発話トーン解析部は、前記音声情報に基づいて、前記推定ターゲットの発話トーンを解析して発話トーン情報を生成する。前記感情推定部は、前記脈拍情報、前記表情情報、前記発話内容情報、および前記発話トーン情報のうち少なくとも2つに基づいて、前記推定ターゲットの感情を推定する。
以下、実施形態の感情推定装置、感情推定方法、記憶媒体、および感情カウントシステムを、図面を参照して説明する。
(第1の実施形態)
実施形態の感情推定装置は、感情推定の対象とする推定ターゲットの顔の表情だけでなく、脈拍や、発話内容、発話トーンなど、マルチモーダルな出力情報を入力として、それらを総合的に判断して、より機微であり、より正確な感情を推定する。以下においては、推定ターゲットとして「人間」を対象とする例を説明する。
図1は、本実施形態の感情推定装置1を用いた感情推定システムA1のハードウェア構成の一例を示す図である。感情推定システムA1は、例えば、感情推定装置1と、ネットワークNと、対象者端末P1と、利用者端末P2とを備える。感情推定装置1は、対象者T(推定ターゲット)の顔の画像情報、音声情報などに基づいて、対象者Tの感情を推定する。ネットワークNは、対象者端末P1と、利用者端末P2とを通信可能に接続し、利用者端末P2と、感情推定装置1とを通信可能に接続する。ネットワークNは、IP通信網、各種電話網、その他の無線通信網、有線通信網などにより実現される。
対象者端末P1は、対象者Tの画像情報および音声情報を取得し、ネットワークNを介して、利用者端末P2に送信する。利用者端末P2は、ネットワークNを介して、対象者端末P1から受信した画像情報および音声情報を感情推定装置1に送信する。感情推定装置1は、利用者端末P2から受信した画像情報および音声情報に基づいて、対象者Tの感情を推定し、推定した感情情報を利用者端末P2に送信する。対象者端末P1および利用者端末P2は、例えば、テレビ電話機能を備えたPC端末、携帯電話などである。利用者端末P2には、感情推定装置1と通信可能であり、感情推定装置1から受信した感情情報を表示することが可能なインターフェースを実現するアプリケーションがインストールされていてよい。
図2は、本実施形態の感情推定装置1の一例を示す機能ブロック図である。感情推定装置1は、例えば、通信部10と、解析部12と、履歴記憶部D1と、辞書記憶部D3とを備える。
通信部10は、例えば、ネットワークNを介して、利用者端末P2から対象者Tの顔の画像情報および音声情報を受信し、解析部12に出力する。通信部10は、例えば、ネットワークNを介して、推定した感情情報を利用者端末P2に送信する。
解析部12は、通信部10から入力された対象者Tの画像情報および音声情報に基づいて、対象者Tの感情推定を行う。解析部12は、例えば、画像解析部20と、音声解析部30と、感情推定部40とを備える。
画像解析部20は、通信部10から入力された対象者Tの画像情報を解析し、解析結果を感情推定部40に出力する。図3に示すように、画像解析部20は、例えば、画像情報を解析して対象者Tの顔の脈拍に関する脈拍情報を生成して感情推定部40に出力する脈拍解析部22と、画像情報を解析して対象者Tの表情に関する表情情報を生成して感情推定部40に出力する表情解析部24とを備える。
脈拍解析部22は、例えば、画像情報から取得される対象者Tの顔の皮下の血液中のヘモグロビンが吸収する信号に基づいて、脈波伝搬時間の推定を行う。脈拍解析部22は、この推定された脈波伝搬時間が血圧変動と正の相関をすることを利用して、対象者Tの脈拍に関する脈拍情報を生成する。脈拍情報には、対象者Tの脈拍の数値データ(指標値)が含まれてよい。脈拍解析部22は、上記の脈拍解析処理を行わずに、画像情報を感情推定部40に出力し、感情推定部40において総合的な感情を推定するための情報の一部として使うようにしてもよい。脈拍解析部22は、顔以外の身体の画像情報から脈拍を解析することも可能である。脈拍解析部22は、顔の一部の画像情報などの部分画像から脈拍を解析することも可能である。脈拍解析部22は、血圧計、心電図の結果、もしくはウェアラブルの脈波センサ等から脈拍を解析してもよい。
表情解析部24は、上記の画像情報に対して、笑い、怒り、悲しみ、喜びなどの感情の種類を紐付けし、対象者Tの表情に関する表情情報を出力する。例えば、表情解析部24は、上記の画像情報に対して、「笑い」の感情として、「大笑い」、「微笑」などの程度を考慮したより細かな種類を紐付けし、「怒り」の感情として、「激怒」、「小さな怒り」などのより細かな種類を紐付けしてもよい。表情情報には、対象者Tの感情が数値化されたデータ(指標値)が含まれてもよい。この感情が数値化されたデータとは、笑い、怒り、悲しみ、喜びなどの各感情の程度を数値で表したものである。表情解析部24は、上記の表情解析処理を行わずに、画像情報を感情推定部40に出力し、感情推定部40において総合的な感情を推定するための情報の一部として使うようにしてもよい。
音声解析部30は、通信部10から入力された対象者Tの音声情報を解析し、解析結果を感情推定部40に出力する。図4に示すように、音声解析部30は、例えば、音声情報を解析して対象者Tの発話内容に関する発話内容情報を生成して感情推定部40に出力する発話内容解析部32と、音声情報を解析して対象者Tの発話のトーン関する発話トーン情報を生成して感情推定部40に出力する発話トーン解析部34とを備える。
発話内容解析部32は、例えば、音声情報に含まれる音声データをテキストデータに変換して、辞書記憶部D3に記憶された辞書データを参照することにより、発話内容情報を生成する。発話内容情報には、対象者Tの感情が数値化されたデータ(指標値)が含まれてよい。発話内容解析部32は、例えば、対象者Tの発話内容に、「おかしい」、「違います」などの否定的な表現が含まれている場合には「怒り」を示す数値を高く設定する。
一方、発話内容解析部32は、例えば、対象者Tの発話内容に、「ありがとう」、「助かりました」などの肯定的な表現が含まれている場合には「喜び」を示す数値を高く設定する。発話内容解析部32は、上記の発話内容解析処理を行わずに、音声情報を感情推定部40に出力し、感情推定部40において総合的な感情を推定するための情報の一部として使うようにしてもよい。
一方、発話内容解析部32は、例えば、対象者Tの発話内容に、「ありがとう」、「助かりました」などの肯定的な表現が含まれている場合には「喜び」を示す数値を高く設定する。発話内容解析部32は、上記の発話内容解析処理を行わずに、音声情報を感情推定部40に出力し、感情推定部40において総合的な感情を推定するための情報の一部として使うようにしてもよい。
発話トーン解析部34は、例えば、発話の音量、速度、および周波数のうち少なくとも1つに基づいて、発話トーン情報を生成する。発話トーン情報には、対象者Tの感情が数値化されたデータ(指標値)が含まれてよい。例えば、発話トーン解析部34は、対象者Tに関して、連続的に取得した音声情報(時系列の音声情報)に基づいて音量、速度、周波数の各々の平均値を取得し、その平均値からのずれに基づいて発話トーン情報を生成する。或いは、発話トーン解析部34は、音量、速度、周波数の各々の変化率に基づいて発話トーン情報を生成する。発話トーン解析部34は、例えば、音量が急増した場合に、感情が高ぶってきたという判断する。発話トーン解析部34は、話す速度が上がった場合に、感情が高ぶってきたという判断する。発話トーン解析部34は、声の高低を表すパラメーターである周波数を参照し、周波数が上がった場合に、感情が高ぶってきたと判断する。発話トーン解析部34は、感情が高ぶってきたという判断をした場合、例えば、感情の候補として「怒り」を示す数値を高く設定する。或いは、発話トーン解析部34は、ある瞬間の音声情報から音量、速度、周波数を取得し、取得した音量、速度、周波数の少なくとも一つが、音量、速度、周波数ごとにあらかじめ定められた基準値を超えたかどうかに基づいて発話トーン情報を生成する。発話トーン解析部34は、例えば、音量が基準値より大きい場合に、感情が高ぶってきたという判断する。発話トーン解析部34は、上記の発話トーン解析処理を行わずに、音声情報を感情推定部40に出力し、感情推定部40において総合的な感情を推定するための情報の一部として使うようにしてもよい。
感情推定部40は、対象者Tの画像情報、対象者Tの音声情報、および対象者Tの音声情報に含まれる音声データをテキストデータに変換したテキスト情報のうち少なくとも2つの情報に基づいて、対象者Tの感情推定を行う。例えば、感情推定部40は、脈拍解析部22から入力された脈拍情報、表情解析部24から入力された表情情報、発話内容解析部32から入力された発話内容情報、および発話トーン解析部34から入力された発話トーン情報のうち少なくとも2つに基づいて、対象者Tの感情推定を行う。感情推定部40は、推定した対象者Tの感情情報を、通信部10を介して、利用者端末P2に送信する。
利用者Uは、利用者端末P2のディスプレイなどに表示された感情情報を確認することで、対象者Tの感情を知ることができる。
利用者Uは、利用者端末P2のディスプレイなどに表示された感情情報を確認することで、対象者Tの感情を知ることができる。
感情推定部40における感情推定処理においては、多層構造のニューラルネットワーク(Deep Neural Network:DNN)、畳み込みニューラルネットワーク(Convolutional Neural Network:CNN)、再帰型ニューラルネットワーク(Recurrent Neural Network:RNN)などを用いたディープラーニング技術を採用してもよい。感情推定部40は、脈拍情報、表情情報、発話内容情報、および発話トーン情報に関するデータと、感情推定結果のデータとを入力とする教師あり学習処理を予め行ってよい。この場合、感情推定部40は、脈拍情報、表情情報、発話内容情報、および発話トーン情報のいずれの情報を重視するのかについての重み付けを、対象者Tの属性(性別、年齢、国籍など)毎に行ってもよい。
脈拍解析部22が脈拍解析処理を行わずに画像情報を感情推定部40に出力している場合、感情推定部40は、画像情報に基づいて、脈拍解析処理を行ってよい。表情解析部24が表情解析処理を行わずに画像情報を感情推定部40に出力している場合、感情推定部40は、画像情報に基づいて、表情解析処理を行ってよい。発話内容解析部32が発話内容解析処理を行わずに音声情報を感情推定部40に出力している場合、感情推定部40は、音声情報に基づいて、発話内容解析処理を行ってよい。発話トーン解析部34が発話トーン解析処理を行わずに音声情報を感情推定部40に出力している場合、感情推定部40は、音声情報に基づいて、発話トーン解析処理を行ってよい。例えば、感情推定部40は、各解析部から得られた情報からそれぞれの解析処理を行わず、各解析部から得られた情報をそのまま使って、総合的な感情推定を行ってもよい。
履歴記憶部D1は、対象者Tに関する各種情報を記憶する。履歴記憶部D1は、例えば、対象者Tの画像情報、音声情報、脈拍解析部22によって解析された脈拍情報、表情解析部24によって解析された表情情報、発話内容解析部32によって解析された発話内容情報、および発話トーン解析部34によって解析された発話トーン情報、感情推定部40によって推定された感情情報、電話番号、氏名などを記憶する。
辞書記憶部D3は、例えば、対象者Tの音声情報の内容を解釈するための辞書データを記憶する。図5は、辞書記憶部D3に記憶されている辞書データの一例を示す図である。
辞書記憶部D3は、文字列と、笑い、怒り、悲しみ、喜びなどの各感情の程度を表す数値とを対応付けた辞書データを記憶する。各感情の程度を数値は、例えば、1から5の数値が割り振られており、数値が大きくなるほど、各感情の程度が高いことを示す。
辞書記憶部D3は、文字列と、笑い、怒り、悲しみ、喜びなどの各感情の程度を表す数値とを対応付けた辞書データを記憶する。各感情の程度を数値は、例えば、1から5の数値が割り振られており、数値が大きくなるほど、各感情の程度が高いことを示す。
例えば、辞書記憶部D3は、文字列「おかしい」と関連付けされた情報として、怒りの数値が「3」であり、笑いの数値が「1」であり、喜びの数値が「0」であり、悲しみ数値が「1」である情報を記憶している。すなわち、文字列「おかしい」は、怒りの感情を表している傾向が高いことを示している。辞書記憶部D3は、図5に示したデータ以外にも、多数の辞書データを記憶してよい。辞書記憶部D3は、例えば、音声情報の内容をより正確に解釈するための他の項目の情報をさらに記憶してもよい。
履歴記憶部D1および辞書記憶部D3の各々は、ROM(Read Only Memory)やRAM(Random Access Memory)、HDD(Hard Disk Drive)、フラッシュメモリなどで実現される。履歴記憶部D1および辞書記憶部D3は、一つのハードウェアで構成されていてもよい。履歴記憶部D1および辞書記憶部D3の各々は、感情推定装置1の外部に設けられてもよい。
上記の感情推定装置1の各機能部のうち一部または全部は、プロセッサがプログラム記憶部(図示しない)に記憶されたプログラム(ソフトウェア)を実行することにより実現されてよい。プログラムは、感情推定装置1の作動開始時に予めインストールされていてもよいし、他のコンピュータからダウンロードされてよいし、コンパクトディスクなどの可搬型記憶媒体からインストールされてもよい。感情推定装置1の各機能部のうち一部または全部は、LSI(Large Scale Integration)やASIC(Application Specific Integrated Circuit)、FPGA(Field-Programmable Gate Array)などのハードウェアによって実現されてもよいし、ソフトウェアとハードウェアの組み合わせによって実現されてもよい。
次に、本実施形態の感情推定装置1の動作について説明する。図6は、感情推定装置1の処理の一例を示すフローチャートである。
利用者Uと対象者Tとが通信を行っている間(例えば、テレビ電話で会話を行っている間)、利用者端末P2は、ネットワークNを介して、対象者端末P1から対象者Tの顔の画像情報および音声情報を継続的に取得する。利用者端末P2は、取得した画像情報および音声情報を感情推定装置1に送信する。これにより、感情推定装置1の通信部10は、対象者Tの画像情報および音声情報を取得する(ステップS101)。通信部10は、取得した画像情報を画像解析部20に出力し、取得した音声情報を音声解析部30に出力する。
次に、画像解析部20の脈拍解析部22は、画像情報を解析して対象者Tの脈拍に関する脈拍情報を生成する(ステップS103)。例えば、脈拍解析部22は、画像情報から取得される対象者Tの顔の皮下の血液中のヘモグロビンが吸収する信号に基づいて、脈波伝搬時間の推定を行う。脈拍解析部22は、この推定された脈波伝搬時間が血圧変動と正の相関をすることを利用して、脈拍情報を生成する。脈拍解析部22は、生成した脈拍情報を感情推定部40に出力する。履歴記憶部D1に対象者Tの過去の脈拍情報が記憶されている場合には、脈拍解析部22は、この過去の脈拍情報を合わせて参照することにより、より正確な脈拍情報を生成してよい。
次に、画像解析部20の表情解析部24は、画像情報を解析して対象者Tの表情に関する表情情報を生成する(ステップS105)。例えば、表情解析部24は、画像情報に対して、笑い、怒り、悲しみ、喜びなどの感情の種類を紐付けし、表情情報を生成する。表情解析部24は、画像情報に対する「笑い」の感情として、「大笑い」、「微笑」などの程度を考慮したより細かな種類を紐付けしてもよい。表情解析部24は、生成した表情情報を感情推定部40に出力する。履歴記憶部D1に対象者Tの過去の表情情報が記憶されている場合には、表情解析部24は、この過去の表情情報を合わせて参照することにより、より正確な表情情報を生成してよい。
次に、音声解析部30の発話内容解析部32は、音声情報を解析して対象者Tの発話内容に関する発話内容情報を生成する(ステップS107)。例えば、発話内容解析部32は、音声情報に含まれる音声データをテキストデータに変換して、辞書記憶部D3に記憶された辞書データを参照することにより、発話内容情報を生成する。発話内容解析部32は、変換したテキストデータに含まれる各文字列が、辞書記憶部D3に記憶されているか否かを判定し、記憶されている場合には各感情の程度を表した数値を抽出する。発話内容解析部32は、変換したテキストデータが辞書記憶部D3に記憶された文字列を複数含んでいる場合には、感情毎に数値を合計した値を発話内容情報としてよい。発話内容解析部32は、生成した発話内容情報を感情推定部40に出力する。履歴記憶部D1に対象者Tの過去の発話内容情報が記憶されている場合には、発話内容解析部32は、この過去の発話内容情報を合わせて参照することにより、より正確な発話内容情報を生成してよい。
次に、音声解析部30の発話トーン解析部34は、音声情報を解析して対象者Tの発話のトーン関する発話トーン情報を生成する(ステップS109)。例えば、発話トーン解析部34は、発話の音量、速度、周波数の各々の変化に基づいて、発話トーン情報を生成する。発話トーン解析部34は、生成した発話トーン情報を感情推定部40に出力する。履歴記憶部D1に対象者Tの過去の発話トーン情報が記憶されている場合には、発話トーン解析部34は、この過去の発話トーン情報を合わせて参照することにより、より正確な発話トーン情報を生成してよい。
上記の脈拍解析部22、表情解析部24、発話内容解析部32、および発話トーン解析部34の各々の解析処理は、任意の順序で実行してもよいし、並行して実行してもよい。
次に、感情推定部40は、脈拍解析部22から入力された脈拍情報、表情解析部24から入力された表情情報、発話内容解析部32から入力された発話内容情報、および発話トーン解析部34から入力された発話トーン情報に基づいて、対象者Tの感情推定を行う(ステップS111)。
次に、感情推定部40は、通信部10を介して、推定した感情情報を利用者端末P2に送信する(ステップS113)。利用者端末P2は、通信部10から受信した感情情報を、例えば、利用者端末P2のディスプレイなどに表示させる。
図7は、利用者端末P2のディスプレイに表示された感情推定結果の一例を示す図である。図7に示す例では、ディスプレイの表示画面の領域S1に、推定した対象者Tの感情情報が表示されており、領域S2に、対象者Tの画像が表示されており、領域S3に、対象者Tのプロフィール情報が表示されている。この例では、推定した対象者Tの感情情報が「L1:小さな怒り」である状態を示している。利用者Uは、ディスプレイに表示された感情情報を確認することで、対象者Tの感情を知ることができる。対象者Tのプロフィール情報は、利用者Uが利用者端末P2を用いて入力および編集してもよい。年齢性別などを顔情報から推定する技術を用いて推定された対象者Tのプロフィール情報を領域S3に表示してもよい。以上により、本フローチャートの処理を終了する。
利用者Uが利用者端末P2を利用している間(例えば、テレビ電話で会話を行っている間)、上記のフローチャートの処理が繰り返し実行され、その状況に応じた対象者Tの感情情報が利用者端末P2のディスプレイに表示される。
以上説明した第1の実施形態によれば、推定ターゲットに関連する複数の情報を総合的に判断することによって推定ターゲットの正確な感情を推定することができる。
例えば、脈拍情報から、推定ターゲットのストレス具合、興奮度(落ち着き度)などがわかる。表情情報から、推定ターゲットの外見上に表れた直接的な感情を推測できる。発話内容情報の内容から、推定ターゲットの発言が、「怒り」などの負の感情に由来するものであるのか、「喜び」などの正の感情に由来するものであるのかを推測できる。発話トーン情報から、対象者Tが興奮した状態なのか、それもと落ち着いた状態なのか、怒っているのか、喜んでいるのか、などの感情を推測できる。
本実施形態によれば、上記のようなマルチモーダルな複数の情報を総合的に判断することによって、推定ターゲットの正確な感情を推定することができる。例えば、同じ「怒り」の感情でも、その程度、その詳細な感情のニュアンスといった正確な感情を読み取ることができる。外見上に表れた表情だけでもある程度の推測はできるが、例えば、コールセンター業務などの顧客の高度な感情を読み取る必要がある場合、対応を間違えると大変な事態に陥ることがある。本実施形態によれば、推定ターゲットの正確かつ詳細な感情を推定することができるため、コールセンター業務や接客業などの顧客対応業務を支援することができる。運転手の感情を推定することにより、運転手に安全運転を促すこともできる。脈拍解析部22が画像情報を解析して脈拍情報を生成する場合、生体センサなど推定ターゲットの生体情報を取得する装置を用いることなくユーザの生体情報を考慮した感情推定を行うことができる。
(第2の実施形態)
以下、第2の実施形態について説明する。第1の実施形態と比較して、第2の実施形態は、感情推定装置1が利用者端末P2に実装されている点が異なる。このため、構成などについては第1の実施形態で説明した図および関連する記載を援用し、詳細な説明を省略する。
以下、第2の実施形態について説明する。第1の実施形態と比較して、第2の実施形態は、感情推定装置1が利用者端末P2に実装されている点が異なる。このため、構成などについては第1の実施形態で説明した図および関連する記載を援用し、詳細な説明を省略する。
図8は、本実施形態の感情推定装置1を用いた感情推定システムA2のハードウェア構成の一例を示す図である。感情推定システムA2は、例えば、対象者端末P1と、感情推定装置1が実装された利用者端末P2(以下、「感情推定装置1」と呼ぶ)と、対象者端末P1と感情推定装置1とをつなぐネットワークNとを備える。感情推定装置1は、ネットワークNを介して、対象者端末P1から対象者Tの画像情報および音声情報を受信する。
次に、本実施形態の感情推定装置1の動作について説明する。図9は、感情推定装置1の処理の一例を示すフローチャートである。
まず、感情推定装置1の通信部10は、例えば、ネットワークNを介して、対象者端末P1から対象者Tの顔の画像情報および音声情報を取得する(ステップS201)。通信部10は、取得した画像情報を画像解析部20に出力し、取得した音声情報を音声解析部30に出力する。
次に、画像解析部20の脈拍解析部22は、画像情報を解析して対象者Tの脈拍に関する脈拍情報を生成する(ステップS203)。脈拍解析部22は、生成した脈拍情報を感情推定部40に出力する。
次に、画像解析部20の表情解析部24は、画像情報を解析して対象者Tの表情に関する表情情報を生成する(ステップS205)。表情解析部24は、生成した表情情報を感情推定部40に出力する。
次に、音声解析部30の発話内容解析部32は、音声情報を解析して対象者Tの発話内容に関する発話内容情報を生成する(ステップS207)。発話内容解析部32は、生成した発話内容情報を感情推定部40に出力する。
次に、音声解析部30の発話トーン解析部34は、音声情報を解析して対象者Tの発話のトーン関する発話トーン情報を生成する(ステップS209)。発話トーン解析部34は、生成した発話トーン情報を感情推定部40に出力する。
上記の脈拍解析部22、表情解析部24、発話内容解析部32、および発話トーン解析部34の各々の解析処理は、任意の順序で実行してもよいし、並行して実行してもよい。
次に、感情推定部40は、脈拍解析部22から入力された脈拍情報、表情解析部24から入力された表情情報、発話内容解析部32から入力された発話内容情報、および発話トーン解析部34から入力された発話トーン情報に基づいて、対象者Tの感情推定を行う(ステップS211)。感情推定部40は、推定した対象者Tの感情情報を、例えば、表示部(図示しない)に表示させる。以上により、本フローチャートの処理を終了する。
以上説明した第2の実施形態によれば、推定ターゲットに関連する複数の情報を総合的に判断することによって推定ターゲットの正確な感情を推定することができる。
感情推定部40は、感情推定を行うとともに、利用者Uの動作決定を支援するための支援情報を生成して利用者端末P2に表示させてもよい。例えば、コールセンターの顧客対応オペレータが感情推定装置1を利用する場合には、感情推定部40は、顧客への回答内容、提案内容などを含む支援情報を生成して利用者端末P2に表示させてもよい。別の実施形態として、ネットワークを介した他の端末との通信を必要としないスタンドアローン型の感情推定装置1を、笑顔メンターとして、接客員にとって望ましい、笑顔や発話などを教育するシステムとして利用してもよい。
感情推定部40は、対象者Tの周辺状況に関する情報(天候、温度、時間帯、場所など)を合わせて考慮して、感情推定を行ってもよい。辞書記憶部D3は、感情推定装置1の利用用途に応じてカスタマイズした辞書データを記憶するようにしてもよい。
例えば、対象者Tが非常に温度の高い地域にいる場合、対象者Tは少し不機嫌であると推定する。対象者Tの周辺状況に関する情報は、例えば、対象者Tの背景画像や位置情報や撮影時間などから取得する。推定ターゲットには、人間、その他動物などの多様な生体、或いは、人工知能などの非生体が含まれてもよい。
例えば、対象者Tが非常に温度の高い地域にいる場合、対象者Tは少し不機嫌であると推定する。対象者Tの周辺状況に関する情報は、例えば、対象者Tの背景画像や位置情報や撮影時間などから取得する。推定ターゲットには、人間、その他動物などの多様な生体、或いは、人工知能などの非生体が含まれてもよい。
(第3の実施形態)
以下、第3の実施形態について説明する。第1の実施形態と比較して、第3の実施形態は、対象者の表情毎(笑顔、怒り顔など)の総数をカウントする機能が感情推定装置に実装されている点が異なる。このため、構成などについては第1の実施形態で説明した図および関連する記載を援用し、詳細な説明を省略する。
以下、第3の実施形態について説明する。第1の実施形態と比較して、第3の実施形態は、対象者の表情毎(笑顔、怒り顔など)の総数をカウントする機能が感情推定装置に実装されている点が異なる。このため、構成などについては第1の実施形態で説明した図および関連する記載を援用し、詳細な説明を省略する。
図10は、本実施形態の感情推定装置2を用いた感情カウントシステムA3のハードウェア構成の一例を示す図である。感情カウントシステムA3は、例えば、感情推定装置2と、感情推定装置2とネットワークNを介して接続された少なくとも1つの端末(図10の例では、端末R1、端末R2、端末R3)とを備える。端末R1、端末R2、および端末R3の各々には、カメラC1、カメラC2、およびカメラC3が接続されている。カメラC1、カメラC2、およびカメラC3は、例えば、展示場などの各種施設などに設置され、不特定多数の対象者Tを撮影するような環境で使用される。カメラC1、カメラC2、およびカメラC3は、例えば、各種施設内の互いに異なる場所に設置される。
カメラC1は、対象者Tの顔画像を撮影し、撮影した顔画像を端末R1に出力する。端末R1は、カメラC1から入力された顔画像を感情推定装置2に送信する。カメラC2は、対象者Tの顔画像を撮影し、撮影した顔画像を端末R2に出力する。端末R2は、カメラC2から入力された顔画像を感情推定装置2に送信する。カメラC3は、対象者Tの顔画像を撮影し、撮影した顔画像を端末R3に出力する。端末R3は、カメラC3から入力された顔画像を感情推定装置2に送信する。カメラC1は、カメラC2、およびカメラC3によって撮影される対象者Tは、互いに異なる人物であってもよいし、同一人物であってもよい。
感情推定装置2は、端末R1、端末R2、および端末R3の各々から受信した顔画像に対して表情解析処理を行い、表情毎の総数をカウントする。図11は、本実施形態の感情推定装置2の一例を示す機能ブロック図である。第1の実施形態と比較して、感情推定装置2は、表情毎の総数をカウントする計数部50をさらに備えている。表情解析部24は、表情情報を計数部50に出力するように構成されている。
計数部50は、解析部12(表情解析部24)によって解析された表情毎の総数をカウントする。例えば、表情解析部24によって、端末R1から受信した対象者Tの顔画像の表情が「笑顔」であると判定された場合、計数部50は、例えば、内部に設けられたメモリ(図示しない)に「笑顔」の総数「1」を記憶させるとともに、通信部10を介して、「笑顔」の表情情報および総数「1」を端末R1に送信する。この後、表情解析部24によって、端末R1から受信した他の対象者Tの顔画像に対する表情解析処理が行われ、他の対象者Tの表情が「笑顔」であると判定された場合、計数部50は、内部に設けられたメモリに「笑顔」の総数を1つインクリメントした「2」を記憶させるとともに、通信部10を介して「笑顔」の表情情報および総数「2」を端末R1に送信する。
計数部50は、端末毎に上記のカウント処理を行ってもよいし、複数の端末に対して上記のカウント処理を行い、その表情毎の合計を算出してもよい。図12は、本実施形態の端末R1(表示装置)に表示された「笑顔」のカウント結果の一例を示す図である。
図12に示す例では、複数の端末(端末R1、端末R2、端末R3)によって取得された画像の「笑顔」のカウント結果が端末R1のディスプレイなどに表示されている。ここでは、端末R1に接続されたカメラC1によって撮影された顔画像に対してその表情が「笑顔」と判定された総数「128」が領域W1に表示されている。カメラC1によって撮影された顔画像に対してその表情が「笑顔」と判定された総数と、カメラC2によって撮影された顔画像に対してその表情が「笑顔」と判定された総数と、カメラC3によって撮影された顔画像に対してその表情が「笑顔」と判定された総数と、の合計「539」が領域W2に表示されている。領域W3には、対象者Tの画像が表示され、領域W4には、その他の情報が表示されている。図12に示す表示結果は一例であり、例えば「笑顔」のカウント総数の時系列での推移を示すグラフで表示してもよい。計数部50によるカウント結果のうち、所定の時間内にカウントした表情毎のカウント数を表示してもよい。
図12に示す例では、複数の端末(端末R1、端末R2、端末R3)によって取得された画像の「笑顔」のカウント結果が端末R1のディスプレイなどに表示されている。ここでは、端末R1に接続されたカメラC1によって撮影された顔画像に対してその表情が「笑顔」と判定された総数「128」が領域W1に表示されている。カメラC1によって撮影された顔画像に対してその表情が「笑顔」と判定された総数と、カメラC2によって撮影された顔画像に対してその表情が「笑顔」と判定された総数と、カメラC3によって撮影された顔画像に対してその表情が「笑顔」と判定された総数と、の合計「539」が領域W2に表示されている。領域W3には、対象者Tの画像が表示され、領域W4には、その他の情報が表示されている。図12に示す表示結果は一例であり、例えば「笑顔」のカウント総数の時系列での推移を示すグラフで表示してもよい。計数部50によるカウント結果のうち、所定の時間内にカウントした表情毎のカウント数を表示してもよい。
以上説明した第3の実施形態によれば、推定ターゲットの表情毎の総数をカウントすることができる。感情推定装置2は、端末R1、端末R2、および端末R3にそれぞれ実装されていてもよい。ここで、複数の端末における表情毎の合計を算出する場合には、感情推定装置2が実装された端末R1、端末R2、および端末R3が互いに通信して総数に関する情報を送受信するようにしてもよい。
表情解析部24は、1枚の画像に複数人の対象者Tが含まれている場合、その各対象者Tの表情を解析することができる。例えば、表情解析部24は、1枚の画像に3人の笑顔の対象者Tが含まれている場合、「笑顔」が3つである旨を示す情報を計数部50に出力する。計数部50は、表情解析部24から入力された「笑顔」が3つである旨を示す情報に基づいて、内部に設けられたメモリに「笑顔」の総数「3」を記憶させる(或いは、3つインクリメントさせる)とともに、通信部10を介して、「笑顔」の表情情報および総数「3」(或いは、3つインクリメントさせた数値)を端末R1に送信する。
上記の本実施形態では、計数部50が、表情解析部24により「笑顔」と解析された総数をカウントする例を示したが、計数部50は笑顔以外の表情をカウントしてもよい。解析部12は表情解析部24による解析結果だけでなく、その他の解析結果に基づいて対象者Tの感情を推定し、計数部50が、推定された感情毎に総数をカウントしてもよい。例えば、表情、発話内容、発話トーンを用いて解析部12は対象者Tの感情を推定し、怒り気味の発話をした人を計数部50がカウントしてもよい。すなわち、本実施形態の感情推定装置2は、推定ターゲットに関連する複数の情報を総合的に判断して感情を推定し、推定した感情毎の総数をカウントできる。これにより、例えば、展示場などの会場内の参加者の盛り上がり度合いを推定するなども可能である。計数部50は、脈拍情報、表情情報、発話内容情報、および発話トーン情報の各々に含まれる指標値毎の総数をカウントしてもよい。例えば、脈拍情報に含まれる指標値(脈拍値)毎の総数をカウントすることで、対象者Tのストレス具合や、興奮具合などを推定することが可能である。
本実施形態の感情推定装置2は、顔画像から同一人物をトラックし、異なるカメラによって撮影された対象者T毎のカウントをしてもよいし、緊張度の高い人物を特定し不穏な動きをする不審者の検出に使用してもよい。
以上説明した少なくとも一つの実施形態によれば、推定ターゲットの脈拍を解析して脈拍情報を生成する脈拍解析部と、前記推定ターゲットの顔の画像情報に基づいて、前記推定ターゲットの表情を解析して表情情報を生成する表情解析部と、前記推定ターゲットの音声情報に基づいて、前記推定ターゲットの発話内容を解析して発話内容情報を生成する発話内容解析部と、前記音声情報に基づいて、前記推定ターゲットの発話トーンを解析して発話トーン情報を生成する発話トーン解析部と、前記脈拍情報、前記表情情報、前記発話内容情報、および前記発話トーン情報のうち少なくとも2つに基づいて、前記推定ターゲットの感情を推定する感情推定部とを備えることで、推定ターゲットに関連する複数の情報を総合的に判断して推定ターゲットの正確な感情を推定することができる。
本発明のいくつかの実施形態を説明したが、これらの実施形態は、例として提示したものであり、発明の範囲を限定することは意図していない。これら実施形態は、その他の様々な形態で実施されることが可能であり、発明の要旨を逸脱しない範囲で、種々の省略、置き換え、変更を行うことができる。これら実施形態やその変形は、発明の範囲や要旨に含まれると同様に、特許請求の範囲に記載された発明とその均等の範囲に含まれるものである。
Claims (11)
- 推定ターゲットの脈拍を解析して脈拍情報を生成する脈拍解析部と、
前記推定ターゲットの顔の画像情報に基づいて、前記推定ターゲットの表情を解析して表情情報を生成する表情解析部と、
前記推定ターゲットの音声情報に基づいて、前記推定ターゲットの発話内容を解析して発話内容情報を生成する発話内容解析部と、
前記音声情報に基づいて、前記推定ターゲットの発話トーンを解析して発話トーン情報を生成する発話トーン解析部と、
前記脈拍情報、前記表情情報、前記発話内容情報、および前記発話トーン情報のうち少なくとも2つに基づいて、前記推定ターゲットの感情を推定する感情推定部と
を備える感情推定装置。 - 前記脈拍解析部は、前記推定ターゲットの画像情報に基づいて、前記推定ターゲットの脈拍を解析して前記脈拍情報を生成する、
請求項1に記載の感情推定装置。 - 文字列と、感情の程度を表す数値とを対応付けた辞書データを記憶する辞書記憶部をさらに備え、
前記発話内容解析部は、前記音声情報に含まれる音声データをテキストデータに変換したテキスト情報に含まれる文字列が前記辞書記憶部に記憶されているか否かを判定し、前記テキスト情報に含まれる前記文字列が前記辞書記憶部に記憶されている場合には、前記数値に基づいて前記発話内容情報を生成する、
請求項1または請求項2に記載の感情推定装置。 - 前記発話トーン解析部は、前記音声情報に含まれる音声データの音量、速度、および周波数のうち少なくとも1つに基づいて、前記発話トーン情報を生成する、
請求項1または請求項2に記載の感情推定装置。 - 前記推定ターゲットの前記脈拍情報、前記表情情報、前記発話内容情報、および前記発話トーン情報を記憶した履歴記憶部をさらに備え、
前記脈拍解析部、前記表情解析部、前記発話内容解析部、および前記発話トーン解析部の各々は、前記履歴記憶部を参照して解析処理を行う、
請求項1または請求項2に記載の感情推定装置。 - 前記感情推定部によって推定された前記推定ターゲットの感情に基づいて、前記推定ターゲットの感情毎の総数をカウントする計数部をさらに備える、
請求項1または請求項2に記載の感情推定装置。 - 前記脈拍情報、前記表情情報、前記発話内容情報、および前記発話トーン情報の各々に含まれる指標値毎の総数をカウントする計数部をさらに備える、
請求項1または請求項2に記載の感情推定装置。 - 前記計数部は、前記表情解析部によって生成された前記表情情報に基づいて、前記推定ターゲットの表情毎の総数をカウントする、請求項7に記載の感情推定装置。
- 推定ターゲットの脈拍を解析して脈拍情報を生成し、
前記推定ターゲットの顔の画像情報に基づいて、前記推定ターゲットの表情を解析して表情情報を生成し、
前記推定ターゲットの音声情報に基づいて、前記推定ターゲットの発話内容を解析して発話内容情報を生成し、
前記音声情報に基づいて、前記推定ターゲットの発話トーンを解析して発話トーン情報を生成し、
前記脈拍情報、前記表情情報、前記発話内容情報、および前記発話トーン情報のうち少なくとも2つに基づいて、前記推定ターゲットの感情を推定する
感情推定方法。 - コンピュータに、
推定ターゲットの脈拍を解析して脈拍情報を生成させ、
前記推定ターゲットの顔の画像情報に基づいて、前記推定ターゲットの表情を解析して表情情報を生成させ、
前記推定ターゲットの音声情報に基づいて、前記推定ターゲットの発話内容を解析して発話内容情報を生成させ、
前記音声情報に基づいて、前記推定ターゲットの発話トーンを解析して発話トーン情報を生成させ、
前記脈拍情報、前記表情情報、前記発話内容情報、および前記発話トーン情報のうち少なくとも2つに基づいて、前記推定ターゲットの感情を推定させる
感情推定プログラムを記憶した記憶媒体。 - 請求項6に記載の感情推定装置と、
前記計数部によってカウントされた総数を表示する表示装置と
を備える、感情カウントシステム。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP2016-211871 | 2016-10-28 | ||
| JP2016211871A JP6524049B2 (ja) | 2016-10-28 | 2016-10-28 | 感情推定装置、感情推定方法、感情推定プログラム、および感情カウントシステム |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2018079106A1 true WO2018079106A1 (ja) | 2018-05-03 |
Family
ID=62024821
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/JP2017/032819 Ceased WO2018079106A1 (ja) | 2016-10-28 | 2017-09-12 | 感情推定装置、感情推定方法、記憶媒体、および感情カウントシステム |
Country Status (2)
| Country | Link |
|---|---|
| JP (1) | JP6524049B2 (ja) |
| WO (1) | WO2018079106A1 (ja) |
Cited By (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN109171773A (zh) * | 2018-09-30 | 2019-01-11 | 合肥工业大学 | 基于多通道数据的情感分析方法和系统 |
| CN113330477A (zh) * | 2019-01-31 | 2021-08-31 | 日立系统股份有限公司 | 有害行为检测系统及方法 |
| CN113950683A (zh) * | 2019-06-24 | 2022-01-18 | 松下知识产权经营株式会社 | 空间设计计划的提出方法和空间设计计划的提出系统 |
| WO2023053267A1 (ja) * | 2021-09-29 | 2023-04-06 | 日本電気株式会社 | 情報処理システム、情報処理装置、情報処理方法、プログラムが格納された非一時的なコンピュータ可読媒体 |
| CN116110435A (zh) * | 2021-11-11 | 2023-05-12 | 株式会社日立制作所 | 情绪识别系统及情绪识别方法 |
Families Citing this family (47)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP7125050B2 (ja) * | 2018-06-08 | 2022-08-24 | 株式会社ニコン | 推定装置、推定システム、推定方法および推定プログラム |
| JP2020074947A (ja) * | 2018-11-08 | 2020-05-21 | 株式会社Nttドコモ | 情報処理装置、低次メンタル状態推定システム、及び低次メンタル状態推定方法 |
| US10835823B2 (en) * | 2018-12-27 | 2020-11-17 | Electronic Arts Inc. | Sensory-based dynamic game-state configuration |
| JP6580281B1 (ja) * | 2019-02-20 | 2019-09-25 | ソフトバンク株式会社 | 翻訳装置、翻訳方法、および翻訳プログラム |
| JP7154164B2 (ja) * | 2019-03-20 | 2022-10-17 | ヤフー株式会社 | 生成装置、生成方法及び生成プログラム |
| KR102183280B1 (ko) * | 2019-03-22 | 2020-11-26 | 한국과학기술원 | 멀티모달 데이터를 이용한 주의집중의 순환 신경망 기반 전자 장치 및 그의 동작 방법 |
| EP3942552A1 (en) * | 2019-04-05 | 2022-01-26 | Huawei Technologies Co., Ltd. | Methods and systems that provide emotion modifications during video chats |
| JP6782329B1 (ja) * | 2019-05-15 | 2020-11-11 | 株式会社Nttドコモ | 感情推定装置、感情推定システム、及び感情推定方法 |
| JP7092104B2 (ja) * | 2019-11-07 | 2022-06-28 | 日本電気株式会社 | 認証装置、認証システム、認証方法及びコンピュータプログラム |
| JP6815667B1 (ja) * | 2019-11-15 | 2021-01-20 | 株式会社Patic Trust | 情報処理装置、情報処理方法、プログラム及びカメラシステム |
| CN114786567A (zh) * | 2019-12-10 | 2022-07-22 | 皇家飞利浦有限公司 | 用于检测基于心率模式的潮热的系统和方法 |
| GB2590473B (en) * | 2019-12-19 | 2022-07-27 | Samsung Electronics Co Ltd | Method and apparatus for dynamic human-computer interaction |
| WO2021214841A1 (ja) * | 2020-04-20 | 2021-10-28 | 三菱電機株式会社 | 感情認識装置、事象認識装置、及び感情認識方法 |
| WO2022024194A1 (ja) * | 2020-07-27 | 2022-02-03 | 株式会社I’mbesideyou | 感情解析システム |
| DE102020210748A1 (de) * | 2020-08-25 | 2022-03-03 | Siemens Aktiengesellschaft | System und Verfahren zur emotionalen Erkennung |
| JP7724550B2 (ja) * | 2020-09-24 | 2025-08-18 | 株式会社I’mbesideyou | ビデオミーティング評価システム及びビデオミーティング評価サーバ |
| WO2022064619A1 (ja) * | 2020-09-24 | 2022-03-31 | 株式会社I’mbesideyou | ビデオミーティング評価システム及びビデオミーティング評価サーバ |
| WO2022064622A1 (ja) * | 2020-09-24 | 2022-03-31 | 株式会社I’mbesideyou | 感情解析システム |
| WO2022064618A1 (ja) * | 2020-09-24 | 2022-03-31 | 株式会社I’mbesideyou | ビデオミーティング評価システム及びビデオミーティング評価サーバ |
| JPWO2022064621A1 (ja) * | 2020-09-24 | 2022-03-31 | ||
| WO2022064620A1 (ja) * | 2020-09-24 | 2022-03-31 | 株式会社I’mbesideyou | ビデオミーティング評価システム及びビデオミーティング評価サーバ |
| JP7629226B2 (ja) * | 2020-10-08 | 2025-02-13 | 株式会社I’mbesideyou | ビデオミーティング評価端末、ビデオミーティング評価システム及びビデオミーティング評価プログラム |
| JP7130290B2 (ja) * | 2020-10-27 | 2022-09-05 | 株式会社I’mbesideyou | 情報抽出装置 |
| KR102428916B1 (ko) * | 2020-11-20 | 2022-08-03 | 오로라월드 주식회사 | 멀티-모달 기반의 감정 분류장치 및 방법 |
| JP7445331B2 (ja) * | 2020-11-26 | 2024-03-07 | 株式会社I’mbesideyou | ビデオミーティング評価端末及びビデオミーティング評価方法 |
| JP7477909B2 (ja) * | 2020-12-25 | 2024-05-02 | 株式会社I’mbesideyou | ビデオミーティング評価端末、ビデオミーティング評価システム及びビデオミーティング評価プログラム |
| JPWO2022145043A1 (ja) * | 2020-12-31 | 2022-07-07 | ||
| WO2022145044A1 (ja) * | 2020-12-31 | 2022-07-07 | 株式会社I’mbesideyou | 反応通知システム |
| JPWO2022145038A1 (ja) * | 2020-12-31 | 2022-07-07 | ||
| WO2022145042A1 (ja) * | 2020-12-31 | 2022-07-07 | 株式会社I’mbesideyou | ビデオミーティング評価端末、ビデオミーティング評価システム及びビデオミーティング評価プログラム |
| KR102670492B1 (ko) * | 2021-01-21 | 2024-05-30 | 주식회사 에스알유니버스 | 인공지능을 활용한 심리 상담 방법 및 장치 |
| JP6985767B1 (ja) * | 2021-01-27 | 2021-12-22 | 株式会社comipro | 診断装置、診断方法、診断システム、及びプログラム |
| WO2022168182A1 (ja) * | 2021-02-02 | 2022-08-11 | 株式会社I’mbesideyou | ビデオセッション評価端末、ビデオセッション評価システム及びビデオセッション評価プログラム |
| JP7694961B2 (ja) * | 2021-02-02 | 2025-06-18 | 株式会社I’mbesideyou | ビデオセッション評価端末、ビデオセッション評価システム及びビデオセッション評価プログラム |
| WO2022168179A1 (ja) * | 2021-02-02 | 2022-08-11 | 株式会社I’mbesideyou | ビデオセッション評価端末、ビデオセッション評価システム及びビデオセッション評価プログラム |
| JP7694960B2 (ja) * | 2021-02-02 | 2025-06-18 | 株式会社I’mbesideyou | ビデオセッション評価端末、ビデオセッション評価システム及びビデオセッション評価プログラム |
| WO2022180855A1 (ja) * | 2021-02-26 | 2022-09-01 | 株式会社I’mbesideyou | ビデオセッション評価端末、ビデオセッション評価システム及びビデオセッション評価プログラム |
| WO2022180856A1 (ja) * | 2021-02-26 | 2022-09-01 | 株式会社I’mbesideyou | ビデオセッション評価端末、ビデオセッション評価システム及びビデオセッション評価プログラム |
| WO2022180857A1 (ja) * | 2021-02-26 | 2022-09-01 | 株式会社I’mbesideyou | ビデオセッション評価端末、ビデオセッション評価システム及びビデオセッション評価プログラム |
| WO2022180853A1 (ja) * | 2021-02-26 | 2022-09-01 | 株式会社I’mbesideyou | ビデオセッション評価端末、ビデオセッション評価システム及びビデオセッション評価プログラム |
| JP7152825B1 (ja) * | 2021-02-26 | 2022-10-13 | 株式会社I’mbesideyou | ビデオセッション評価端末、ビデオセッション評価システム及びビデオセッション評価プログラム |
| WO2022180854A1 (ja) * | 2021-02-26 | 2022-09-01 | 株式会社I’mbesideyou | ビデオセッション評価端末、ビデオセッション評価システム及びビデオセッション評価プログラム |
| JP7405436B2 (ja) * | 2021-03-30 | 2023-12-26 | Necフィールディング株式会社 | 健康管理装置、健康管理方法及びプログラム |
| EP4428655A4 (en) | 2021-11-01 | 2025-02-26 | Sony Group Corporation | Information processing device, communication assistance device, and communication assistance system |
| JP7169030B1 (ja) * | 2022-05-16 | 2022-11-10 | 株式会社RevComm | プログラム、情報処理装置、情報処理システム、情報処理方法、情報処理端末 |
| JP2024119141A (ja) * | 2023-02-22 | 2024-09-03 | 株式会社日立製作所 | 生体計測データ処理装置、及び生体計測データ処理方法 |
| KR102929139B1 (ko) * | 2025-09-04 | 2026-02-24 | 주식회사 스포클립에이아이 | 운동 자세에 따른 부상 위험을 예측하고 안내하기 위한 장치 및 방법 |
Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2005122549A (ja) * | 2003-10-17 | 2005-05-12 | Aruze Corp | 会話制御装置及び会話制御方法 |
| JP2006127057A (ja) * | 2004-10-27 | 2006-05-18 | Canon Inc | 推定装置、及びその制御方法 |
| JP2016103786A (ja) * | 2014-11-28 | 2016-06-02 | 日立マクセル株式会社 | 撮像システム |
-
2016
- 2016-10-28 JP JP2016211871A patent/JP6524049B2/ja active Active
-
2017
- 2017-09-12 WO PCT/JP2017/032819 patent/WO2018079106A1/ja not_active Ceased
Patent Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2005122549A (ja) * | 2003-10-17 | 2005-05-12 | Aruze Corp | 会話制御装置及び会話制御方法 |
| JP2006127057A (ja) * | 2004-10-27 | 2006-05-18 | Canon Inc | 推定装置、及びその制御方法 |
| JP2016103786A (ja) * | 2014-11-28 | 2016-06-02 | 日立マクセル株式会社 | 撮像システム |
Cited By (8)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN109171773A (zh) * | 2018-09-30 | 2019-01-11 | 合肥工业大学 | 基于多通道数据的情感分析方法和系统 |
| CN109171773B (zh) * | 2018-09-30 | 2021-05-18 | 合肥工业大学 | 基于多通道数据的情感分析方法和系统 |
| CN113330477A (zh) * | 2019-01-31 | 2021-08-31 | 日立系统股份有限公司 | 有害行为检测系统及方法 |
| CN113950683A (zh) * | 2019-06-24 | 2022-01-18 | 松下知识产权经营株式会社 | 空间设计计划的提出方法和空间设计计划的提出系统 |
| WO2023053267A1 (ja) * | 2021-09-29 | 2023-04-06 | 日本電気株式会社 | 情報処理システム、情報処理装置、情報処理方法、プログラムが格納された非一時的なコンピュータ可読媒体 |
| JPWO2023053267A1 (ja) * | 2021-09-29 | 2023-04-06 | ||
| JP7677440B2 (ja) | 2021-09-29 | 2025-05-15 | 日本電気株式会社 | 情報処理システム、情報処理装置、情報処理方法、プログラム |
| CN116110435A (zh) * | 2021-11-11 | 2023-05-12 | 株式会社日立制作所 | 情绪识别系统及情绪识别方法 |
Also Published As
| Publication number | Publication date |
|---|---|
| JP2018068618A (ja) | 2018-05-10 |
| JP6524049B2 (ja) | 2019-06-05 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| JP6524049B2 (ja) | 感情推定装置、感情推定方法、感情推定プログラム、および感情カウントシステム | |
| CN110875032B (zh) | 语音交互系统和方法、程序、学习模型生成装置和方法 | |
| JP6617053B2 (ja) | 感情分類によって文脈意味の理解精度を高める発話意味分析プログラム、装置及び方法 | |
| US10448887B2 (en) | Biometric customer service agent analysis systems and methods | |
| US8715179B2 (en) | Call center quality management tool | |
| JP6715410B2 (ja) | 評価方法、評価装置、評価プログラム、および、評価システム | |
| US8715178B2 (en) | Wearable badge with sensor | |
| US9138186B2 (en) | Systems for inducing change in a performance characteristic | |
| KR102199928B1 (ko) | 사용자 페르소나를 고려한 대화형 에이전트 장치 및 방법 | |
| JP2016149063A (ja) | 感情推定装置及び感情推定方法 | |
| CN109793526B (zh) | 测谎方法、装置、计算机设备和存储介质 | |
| KR20120044910A (ko) | 지능형 감성단어 확장장치 및 그 확장방법 | |
| JP2019058625A (ja) | 感情読み取り装置及び感情解析方法 | |
| KR20220106029A (ko) | 인공지능을 활용한 심리 상담 방법 및 장치 | |
| US20190371344A1 (en) | Apparatus and method for predicting/recognizing occurrence of personal concerned context | |
| JP7021488B2 (ja) | 情報処理装置、及びプログラム | |
| JP6213476B2 (ja) | 不満会話判定装置及び不満会話判定方法 | |
| CN109599127A (zh) | 信息处理方法、信息处理装置以及信息处理程序 | |
| JP7278972B2 (ja) | 表情解析技術を用いた商品に対するモニタの反応を評価するための情報処理装置、情報処理システム、情報処理方法、及び、プログラム | |
| JP7040593B2 (ja) | 接客支援装置、接客支援方法、及び、接客支援プログラム | |
| CN115666376B (zh) | 用于高血压监测的系统和方法 | |
| JP7138456B2 (ja) | 印象導出システム、印象導出方法および印象導出プログラム | |
| JP7014761B2 (ja) | 認知機能推定方法、コンピュータプログラム及び認知機能推定装置 | |
| JP2021163329A (ja) | 翻訳システム | |
| JP2021135363A (ja) | 制御システム、制御装置、制御方法及びコンピュータプログラム |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 17864298 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 17864298 Country of ref document: EP Kind code of ref document: A1 |