WO2022239234A1 - 感情状態推定学習装置、感情状態推定装置、それらの方法、およびプログラム - Google Patents

感情状態推定学習装置、感情状態推定装置、それらの方法、およびプログラム Download PDF

Info

Publication number
WO2022239234A1
WO2022239234A1 PCT/JP2021/018419 JP2021018419W WO2022239234A1 WO 2022239234 A1 WO2022239234 A1 WO 2022239234A1 JP 2021018419 W JP2021018419 W JP 2021018419W WO 2022239234 A1 WO2022239234 A1 WO 2022239234A1
Authority
WO
WIPO (PCT)
Prior art keywords
stimulus
emotional state
learning
estimation
filter coefficient
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/JP2021/018419
Other languages
English (en)
French (fr)
Inventor
藍李 太田
信哉 志水
奏 山本
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
NTT Inc
Original Assignee
Nippon Telegraph and Telephone Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Nippon Telegraph and Telephone Corp filed Critical Nippon Telegraph and Telephone Corp
Priority to PCT/JP2021/018419 priority Critical patent/WO2022239234A1/ja
Priority to JP2023520721A priority patent/JP7540588B2/ja
Publication of WO2022239234A1 publication Critical patent/WO2022239234A1/ja
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N20/00Machine learning

Definitions

  • the present invention relates to technology for estimating how an emotional state changes over time in response to a given stimulus.
  • Non-Patent Document 1 is known as a technique for estimating changes in the emotional state over time in response to a given stimulus by emotional dimension evaluation.
  • Non-Patent Document 1 changes in emotions are caused by stimuli, but in fact they are also affected by the emotional state when receiving stimuli.
  • the emotional state when receiving a stimulus since the emotional state when receiving a stimulus is not considered, there are cases where natural changes in the emotional state cannot be estimated.
  • An object of the present invention is to provide an emotional state estimation device that takes into consideration the emotional state when receiving a stimulus, an emotional state estimation learning device that learns parameters used in the emotional state estimation device, a method thereof, and a program. .
  • an emotional state estimation learning device uses a parameter w to convert an emotional state EL (t) for learning at time t into a filter coefficient estimator for estimating filter coefficients for correcting one or more feature values extracted from the learning stimulus S L (t) to one or more feature values that are actually perceived and recognized; a stimulus modification unit that filters the feature quantity extracted from the stimulus S L (t) using the stimulus S L (t) and modifies it to the stimulus S' L (t) that is the feature quantity that is actually perceived and recognized;
  • S' L (t) and the emotional state E L (t+1) for learning at time t+1 the emotional state E L (t+1) and the modified stimulus S'(t) are a learning unit that learns the parameter w so that the distance from the estimated emotional state value ⁇ EL (t+1) at time t+1 estimated using the parameter w becomes smaller.
  • an emotional state estimating device extracts a stimulus S(t) from an emotional state E(t) at time t based on a learned parameter w. ), a filter coefficient estimating unit for estimating a filter coefficient for correcting one or more feature values extracted from ) to one or more feature values that are actually perceived and recognized, and using the filter coefficients, the stimulus S(t) Using the stimulus modifier that filters the feature extracted from the stimulus S'(t), which is the feature that is actually perceived and perceived, and the modified stimulus S'(t), the time point an emotion estimating unit that estimates the emotional state E(t+1) at t+1, and the learned parameter w is the distance between the emotional state E L (t+1) for learning and the estimated value is It has been learned to be small.
  • the present invention it is possible to consider the emotional state when receiving a stimulus, and to estimate a more natural change in the emotional state.
  • the figure which shows the structural example of an emotional state estimation system Functional block diagram of the emotional state estimation learning device.
  • the figure which shows the example of the processing flow of an emotional state estimation apparatus. The figure which shows the structural example of the computer which applies this method.
  • Emotions are mental and physical states altered by the perception of received emotion-related stimuli.
  • the emotion-related stimulus is a stimulus that affects an existing emotional state or evokes a new emotion. It has been reported that the perception of emotion-related stimuli is affected by the current emotional state. Below are three examples.
  • Negativity bias (negativity effect) Information with negative emotions tends to have a higher information value than information with positive emotions, and has a greater impact on humans. Humans are more sensitive to negative emotions than positive emotions, and negative emotions tend to last longer than positive emotions.
  • Example 2 Mood Matching Judgment Effect A more positive evaluation/judgment is made when the person has a positive emotion (mood), and a negative evaluation/judgment is more likely to be made when the person has a negative emotion (mood).
  • Example 3 Kuleshov Effect When an image or photograph is edited in a cinematic fashion, it has the property of affecting the meaning of other images positioned before and after it. Even in a series of images and photographs that have no context, humans unconsciously associate the connections between the scenes and interpret the meaning.
  • the emotional state at the time of receiving the stimulus is considered, and a more natural change in the emotional state is estimated.
  • the stimulus S(t) received during the emotional state E(t) is perceived under the influence of the emotional state E(t).
  • a filter coefficient estimator estimates a filter coefficient for correcting the input stimulus S(t) to a stimulus that is actually perceived/recognised, using the emotional state E(t).
  • the stimulus modifier the stimulus S(t) is filtered, and the emotional state E(t+1) induced by the filtered stimulus is estimated using a conventional emotion estimation technique.
  • FIG. 1 shows a configuration example of an emotional state estimation system.
  • the emotional state estimation system includes an emotional state estimation learning device 100 and an emotional state estimation device 200 .
  • Emotional state estimation learning device 100 receives learning stimulus S L (t) and learning emotional states E L (t) and E L (t+1) as inputs, and learns parameters used in emotional state estimation device 200. and outputs the learned parameters.
  • t is an index indicating the time point.
  • the emotional state E L (t+1) is the emotional state at time t+1 obtained from the stimulus S L (t) and the emotional state E L (t) at time t, and corresponds to correct data.
  • Emotional state estimation learning apparatus 100 receives, for example, multiple sets of learning stimulus SL (t) and learning emotional states EL (t), EL (t+1).
  • the stimulus S L (t) is a stimulus for arousing a specific emotion, and is given a label and a level score for the arousing emotion.
  • a plurality of pairs of emotional states E L (t) and E L (t+1) for learning when receiving the “stimulus S L (t)_emotion label_emotional level score_stimulus type” are prepared in advance. , are input to the emotional state estimation learning device 100 .
  • Emotion labels, level scores, and stimulus types are described below. However, hereinafter, “stimulus S L (t)_emotion label_emotion level score_stimulus type” is simply referred to as stimulus S L (t).
  • Emotional state estimation device 200 receives learned parameters prior to emotion estimation processing.
  • the learned parameters are parameters output from the emotional state estimation learning device 100, for example.
  • Emotional state estimating device 200 receives stimulus S(t) and emotional state E(t) as inputs, and derives emotional state E L (t ) is estimated and output.
  • Emotional state estimation learning device and emotional state estimation device for example, a special program is loaded into a publicly known or dedicated computer having a central processing unit (CPU: Central Processing Unit), a main memory (RAM: Random Access Memory), etc. It is a special device constructed
  • the emotional state estimation learning device and the emotional state estimation device for example, execute each process under the control of the central processing unit.
  • the data input to the emotional state estimation learning device and the emotional state estimation device and the data obtained in each process are stored, for example, in a main memory device, and the data stored in the main memory device are subjected to central processing as necessary. It is read out to the device and used for other processing.
  • each processing unit of the emotional state estimation learning device and the emotional state estimation device may be configured by hardware such as an integrated circuit.
  • Each storage unit included in the emotional state estimation learning device and the emotional state estimation device can be configured by, for example, a main storage device such as RAM (Random Access Memory), or middleware such as a relational database or key-value store.
  • each storage unit does not necessarily have to be provided inside the emotional state estimation learning device and the emotional state estimation device, and is an auxiliary storage device composed of a semiconductor memory device such as a hard disk, an optical disk, or a flash memory. and provided outside the emotional state estimation learning device and the emotional state estimation device.
  • FIG. 2 is a functional block diagram of the emotional state estimation learning device 100 according to the first embodiment, and FIG. 3 shows its processing flow.
  • the emotional state estimation learning device 100 includes a filter coefficient estimation unit 120 , a stimulus modification unit 130 and a learning unit 140 .
  • the emotional state estimation learning device 100 receives the learning stimulus S L (t) and the learning emotional states E L (t) and E L (t+1) as inputs, and the emotional state estimation device 200 Learn the parameters to use and output the learned parameters.
  • a learning stimulus S L (t), which is an input to emotional state estimation learning apparatus 100, is a stimulus for arousing a specific emotion.
  • the stimulus S L (t) is a stimulus that can be received by human sense organs such as visual and auditory organs, and is, for example, video data, image data, audio data, and the like.
  • the type of this data corresponds to the aforementioned stimulus type.
  • the stimulus S L (t) is given a label of an evoked emotion and a level score.
  • the stimulus S L (t) is given a label indicating an emotion such as emotion, or a numerical value indicating the level of emotion such as emotion (for example, each emotion level is indicated by a numerical value from 1 to 5). ing.
  • Emotional state E L (t) for learning which is input to emotional state estimation learning apparatus 100, indicates the emotional state at time t when stimulus S L (t) at time t is input. or level score.
  • Emotional states can be obtained using conventional techniques. For example, (1-1) the emotional state estimation learning device 100 presents video data that evokes a specific emotion, which is a learning stimulus S L (t), to the subject, and (1-2) the subject However, by inputting the emotional state at that time through the input device while viewing the video data, the emotional state estimation learning device 100 can acquire the emotional state. By having the subject perform this work in a neutral emotional state, the subject's emotional state can be ignored. At this time, the learning stimulus S L (t) may be the same as the stimulus used for filter coefficient estimation.
  • the emotional state estimation learning device 100 may acquire the emotional state EL (t) for learning using an emotional state estimation model trained using conventional technology.
  • This emotional state estimation model is a model that receives a stimulus S(t) as an input and outputs an emotional state E(t) in response to the stimulus S(t).
  • Emotional state E L (t+1) is an emotional state aroused by perceiving and perceiving input stimulus S L (t) at time t, and is the time after change from emotional state E L (t) Emotional state at t+1.
  • the emotional state estimation device 200 determines that if the emotional state E(t) is neutral, the emotional state E(t+1) after perception of the stimulus S(t) is not affected by the emotional state E(t). , the emotional state E(t+1) can be estimated based on the emotional label and level score assigned to the stimulus S(t). However, since various emotional states can actually occur at time t, emotional state estimation device 200 estimates the emotional state at time t+1 in consideration of the influence of emotional state E L (t).
  • the learned parameters output from the emotional state estimation learning device 100 are the parameters used by the emotional state estimation device 200.
  • the filter coefficient estimation unit 120 receives the emotional state E L (t) for learning and the parameters w1, w2, . . . output from the learning unit 140, and uses the parameters w1, w2, . t), one or more feature values extracted from the learning stimulus S L (t) are converted to one or more feature values that are actually perceived or perceived (in other words, one or more feature values that correspond to the actually perceived or perceived stimulus
  • the filter coefficients v1, v2, . . . are estimated (S120) and output.
  • the stimulus modification unit 130 receives the learning stimulus S L (t ) and the filter coefficients v1, v2, . . . , and uses the filter coefficients v1, v2, . is filtered to modify it to a stimulus that is a feature quantity that is actually perceived and recognized (S130) and output.
  • the corrected stimulus is hereinafter denoted as S' L (t).
  • the stimulus modifier 130 extracts the feature quantity Q L (t) from the stimulus S L (t).
  • a conventional method may be used as a method for extracting feature amounts.
  • the method of Non-Patent Document 1 may be used to extract feature amounts related to facial expression, pupil diameter, heartbeat, etc. from video data that is the stimulus S L (t).
  • feature amounts related to volume, tone, voice, etc. may be extracted from the speech data that is the stimulus S L (t), and the emotional labels and level scores assigned to the stimulus S L (t) may be extracted. may be extracted as a feature amount.
  • the modified feature quantity Q' L (t) corresponds to the modified stimulus S' L (t) described above.
  • the learning unit 140 receives the corrected stimulus S' L (t) and the emotional state E L (t+1) for learning as inputs, and uses these values to obtain the emotional state E L (t+1) and , the above-mentioned filter coefficient estimating unit 120 so that the distance from the estimated emotional state ⁇ EL (t+1) at time t+1 estimated using the modified stimulus S′(t) becomes small. learn and output model parameters w1, w2, . . . used when estimating filter coefficients v1, v2, .
  • the learning unit 140 uses the modified stimulus S′ L (t) to obtain the estimated emotional state ⁇ EL (t+1) at time t+1 ( S140-1 ).
  • an estimated emotional state ⁇ EL (t+1) induced by the stimulus S'L (t) modified by the stimulus modification unit 130 is obtained using a trained estimation model of a conventional emotion estimation method.
  • a trained estimation model is a model learned by machine learning such as a neural network, and the feature value obtained from the stimulus S(t) at time t is input, and the emotional state at time t+1 is estimated. It is a model that Note that an emotion label or level score may be used as the estimation result.
  • the stimulus modification unit 130 when extracting feature amounts related to facial expression, pupil diameter, heart rate, etc. from video data that is the stimulus S L (t) by the method of Non-Patent Document 1, Using the trained estimation model, the estimated emotional state ⁇ E L (t+1) is obtained from the corrected stimulus S' L (t).
  • the learning unit 140 calculates the distance D(t) between the learning emotional state E L (t+1) at time t+1 and its estimated value ⁇ E L (t+1) (S140- 2). For example, Euclidean distance and Mahalanobis distance can be used to calculate the distance D(t).
  • the learning unit 140 determines whether or not the distance D(t) is equal to or less than a predetermined threshold (S140-3), and if it is greater than the predetermined threshold (no in S140-3), for example, gradient descent or the like is used. is used to update the parameters w1, w2, . On the other hand, if the distance D(t) is equal to or less than the predetermined threshold (yes in S140-3), emotional state estimation and learning apparatus 100 processes a set (learning data) that has not yet been processed.
  • a predetermined threshold S140-3
  • the predetermined threshold for example, gradient descent or the like
  • the learning unit 140 receives all the sets ( learning data) has been processed (S140-5), and if there is a group (learning data) that has not been processed yet (no in S140-5), S120 to S140-4 are controlled to be repeated. , if there is no unprocessed set (learning data) (yes in S140-5), the parameters w1, w2, . . . at that time are output as learned parameters.
  • FIG. 4 is a functional block diagram of the emotional state estimation device 200 according to the first embodiment, and FIG. 5 shows its processing flow.
  • the emotional state estimation device 200 includes a filter coefficient estimation section 220 , a stimulus modification section 230 and an emotion estimation section 240 .
  • the emotional state estimation device 200 receives learned parameters prior to emotion estimation processing.
  • Emotional state estimating device 200 receives stimulus S(t) and emotional state E(t) as inputs, and estimates emotional state E at time t+1 from stimulus S(t) and emotional state E(t) at time t. Get (t+1) and print it.
  • the emotional state estimate E ⁇ (t+1) is, for example, an emotional label or level score.
  • Filter coefficient estimation section 220 receives learned parameters prior to emotion estimation processing.
  • the filter coefficient estimation unit 220 receives the emotional state E(t) as an input, and actually perceives one or more feature amounts extracted from the stimulus S(t) from the emotional state E(t) based on the learned parameters.
  • the stimulus modifier 230 receives the stimulus S(t) and the filter coefficients v1, v2, . . . , and uses the filter coefficients v1, v2, .
  • the stimulus is corrected to a feature quantity that is actually perceived and recognized (S230) and output. Specifically, the same processing as S130 of the stimulus modification unit 130 is performed.
  • the corrected stimulus is hereinafter denoted as S'(t).
  • the emotion estimation unit 240 receives the modified stimulus S′(t) as an input, uses the modified stimulus S′(t), and estimates the time point t+ using the modified stimulus S′(t). An estimated value E(t+1) of the emotional state of 1 is obtained (S240) and output. Specifically, the same processing as S140-1 of the learning unit 140 is performed.
  • Emotional state estimation is one of the important technologies in the field of HCI (human-computer interaction), such as interaction between humans and robots.
  • HCI human-computer interaction
  • the emotional state is indicated by an emotional label or level score, but a parameter or the like indicating an emotion that is likely to continue may be added to the value indicating the emotional state. For example, add parameters that indicate emotional changes that are responsive immediately and emotional changes that are slow.
  • Fig. 1 shows a configuration example of an emotional state estimation system.
  • the emotional state estimation system includes an emotional state estimation learning device 300 and an emotional state estimation device 400 .
  • FIG. 2 is a functional block diagram of the emotional state estimation learning device 300 according to the first embodiment, and FIG. 3 shows its processing flow.
  • the emotional state estimation learning device 300 includes a feature extraction unit 310 , a filter coefficient estimation unit 320 , a stimulus correction unit 130 and a learning unit 140 .
  • the processing contents of the stimulus modification unit 130 and the learning unit 140 are the same as those in the first embodiment, so the description is omitted.
  • the feature amount extraction unit 310 receives the learning stimulus S L (t), extracts the feature amount P L (t) from the stimulus S L (t) (S310), and outputs the result.
  • feature amounts used in conventional emotion estimation techniques may be extracted, or other feature amounts may be extracted.
  • the feature amount may be the same as the feature amount extracted by the stimulus modification unit 130 described above, or may be a different feature amount.
  • the feature amount P L (t) is a feature amount related to facial expression, pupil diameter, heartbeat, etc. extracted from video data that is the stimulus S L (t), or from audio data that is the stimulus S L (t).
  • the emotional state estimation model (stimulus S(t) is input, and the emotional state E(t) ) is used, the feature quantity used when estimating the emotional state from the learning stimulus S L (t) is extracted as the aforementioned feature quantity P L .
  • the technique of references 1 and 2 can be used to extract the feature quantity P L .
  • the filter coefficient estimation unit 320 receives the emotional state E L (t) for learning, the feature amount P L (t), and the parameters w1, w2, . Using the emotional state E L (t) and the feature amount P L (t), one or more feature amounts extracted from the learning stimulus S L (t) are actually perceived and recognized. Filter coefficients v1, v2, . As the initial values of the parameters w1, w2, . . . , default values between 0 and 1 may be given.
  • FIG. 4 is a functional block diagram of the emotional state estimation device 400 according to the first embodiment, and FIG. 5 shows its processing flow.
  • the emotional state estimation device 400 includes a feature extraction unit 410 , a filter coefficient estimation unit 420 , a stimulus correction unit 230 and an emotion estimation unit 240 .
  • the processing contents of the stimulus modification unit 230 and the emotion estimation unit 240 are the same as those in the first embodiment, so the description is omitted.
  • the feature quantity extraction unit 410 receives the stimulus S(t), extracts the feature quantity P(t) from the stimulus S(t) (S410), and outputs it. Specifically, the same processing as S310 of the feature amount extraction unit 310 is performed.
  • Filter coefficient estimation section 420 receives learned parameters prior to emotion estimation processing.
  • the filter coefficient estimation unit 420 receives the emotional state E(t) and the feature amount P(t) as input, and based on the learned parameters w1, w2, . . . , one or more feature values extracted from the stimulus S(t) to one or more feature values that are actually perceived or perceived (in other words, one or more feature values corresponding to the actually perceived or perceived stimuli)
  • Filter coefficients v1, v2, . . . for correction are estimated (S420) and output. Specifically, the same processing as S320 of the filter coefficient estimation unit 320 is performed.
  • the same effects as those of the first embodiment can be obtained. Furthermore, when estimating the filter coefficients, the stimulus can be modified more appropriately by considering the feature amount of the stimulus.
  • the present invention is not limited to the above embodiments and modifications.
  • the various types of processing described above may not only be executed in chronological order according to the description, but may also be executed in parallel or individually according to the processing capacity of the device that executes the processing or as necessary.
  • appropriate modifications are possible without departing from the gist of the present invention.
  • a program that describes this process can be recorded on a computer-readable recording medium.
  • Any computer-readable recording medium may be used, for example, a magnetic recording device, an optical disk, a magneto-optical recording medium, a semiconductor memory, or the like.
  • this program is carried out, for example, by selling, assigning, lending, etc. portable recording media such as DVDs and CD-ROMs on which the program is recorded.
  • the program may be distributed by storing the program in the storage device of the server computer and transferring the program from the server computer to other computers via the network.
  • a computer that executes such a program for example, first stores the program recorded on a portable recording medium or the program transferred from the server computer once in its own storage device. Then, when executing the process, this computer reads the program stored in its own recording medium and executes the process according to the read program. Also, as another execution form of this program, the computer may read the program directly from a portable recording medium and execute processing according to the program, and the program is transferred from the server computer to this computer. Each time, the processing according to the received program may be executed sequentially. In addition, the above-mentioned processing is executed by a so-called ASP (Application Service Provider) type service, which does not transfer the program from the server computer to this computer, and realizes the processing function only by its execution instruction and result acquisition. may be It should be noted that the program in this embodiment includes information that is used for processing by a computer and that conforms to the program (data that is not a direct instruction to the computer but has the property of prescribing the processing of the computer, etc.).
  • ASP
  • the device is configured by executing a predetermined program on a computer, but at least part of these processing contents may be implemented by hardware.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Software Systems (AREA)
  • Data Mining & Analysis (AREA)
  • Evolutionary Computation (AREA)
  • Medical Informatics (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Physics & Mathematics (AREA)
  • Computing Systems (AREA)
  • General Engineering & Computer Science (AREA)
  • General Physics & Mathematics (AREA)
  • Mathematical Physics (AREA)
  • Artificial Intelligence (AREA)
  • Measurement Of The Respiration, Hearing Ability, Form, And Blood Characteristics Of Living Organisms (AREA)

Abstract

刺激を受けた際の感情状態を考慮した感情状態推定装置で用いるパラメータを学習する感情状態推定学習装置等を提供する。感情状態推定学習装置は、パラメータwを用いて、ある時点tの学習用の感情状態EL(t)から、時点tの学習用の刺激SL(t)から抽出される1以上の特徴量を実際に知覚、認知される1以上の特徴量に修正するためのフィルタ係数を推定するフィルタ係数推定部と、フィルタ係数を用いて、刺激SL(t)から抽出される特徴量をフィルタリングして、実際に知覚、認知される特徴量である刺激S'L(t)に修正する刺激修正部と、修正後の刺激S'L(t)と時点t+1の学習用の感情状態EL(t+1)とを用いて、感情状態EL(t+1)と、修正後の刺激S'(t)を用いて推定される時点t+1の感情状態の推定値^EL(t+1)との距離が小さくなるように、パラメータwを学習する学習部と、を含む。

Description

感情状態推定学習装置、感情状態推定装置、それらの方法、およびプログラム
 本発明は、与えられた刺激に対し、感情状態が経時的にどのように変化するかを推定する技術に関する。
 与えられた刺激に対し、感情状態の経時的な変化を感情次元評価で推定する技術として非特許文献1が知られている。
Kenta Masui, Takumi Nagasawa, Hirokazu Doi, Norimichi Tsumura, "Continuous estimation of emotional change using multimodal affective responses", CVPR2020.
 感情の変化は非特許文献1に記載されているように刺激に起因するが、実際には刺激を受けた際の感情状態にも影響を受ける。非特許文献1では、刺激を受けた際の感情状態が考慮されていないため、自然な感情状態の変化を推定できない場合がある。
 本発明は、刺激を受けた際の感情状態を考慮した感情状態推定装置、感情状態推定装置で用いるパラメータを学習する感情状態推定学習装置、それらの方法、およびプログラムを提供することを目的とする。
 上記の課題を解決するために、本発明の一態様によれば、感情状態推定学習装置は、パラメータwを用いて、ある時点tの学習用の感情状態EL(t)から、時点tの学習用の刺激SL(t)から抽出される1以上の特徴量を実際に知覚、認知される1以上の特徴量に修正するためのフィルタ係数を推定するフィルタ係数推定部と、フィルタ係数を用いて、刺激SL(t)から抽出される特徴量をフィルタリングして、実際に知覚、認知される特徴量である刺激S'L(t)に修正する刺激修正部と、修正後の刺激S'L(t)と時点t+1の学習用の感情状態EL(t+1)とを用いて、感情状態EL(t+1)と、修正後の刺激S'(t)を用いて推定される時点t+1の感情状態の推定値^EL(t+1)との距離が小さくなるように、パラメータwを学習する学習部と、を含む。
 上記の課題を解決するために、本発明の他の態様によれば、感情状態推定装置は、学習済みのパラメータwに基づきある時点tの感情状態E(t)から時点tの刺激S(t)から抽出される1以上の特徴量を実際に知覚、認知される1以上の特徴量に修正するためのフィルタ係数を推定するフィルタ係数推定部と、フィルタ係数を用いて、刺激S(t)から抽出される特徴量をフィルタリングして、実際に知覚、認知される特徴量である刺激S'(t)に修正する刺激修正部と、修正後の刺激S'(t)を用いて、時点t+1の感情状態E(t+1)を推定する感情推定部と、を含み、学習済みのパラメータwは、学習用の感情状態EL(t+1)とその推定値との距離が小さくなるように学習されたものである。
 本発明によれば、刺激を受けた際の感情状態を考慮し、より自然な感情状態の変化を推定することができるという効果を奏する。
感情状態推定システムの構成例を示す図。 感情状態推定学習装置の機能ブロック図。 感情状態推定学習装置の処理フローの例を示す図。 感情状態推定装置の機能ブロック図。 感情状態推定装置の処理フローの例を示す図。 本手法を適用するコンピュータの構成例を示す図。
 以下、本発明の実施形態について、説明する。なお、以下の説明に用いる図面では、同じ機能を持つ構成部や同じ処理を行うステップには同一の符号を記し、重複説明を省略する。以下の説明において、テキスト中で使用する記号「^」等は、本来直後の文字の真上に記載されるべきものであるが、テキスト記法の制限により、当該文字の直前に記載する。式中においてはこれらの記号は本来の位置に記述している。また、ベクトルや行列の各要素単位で行われる処理は、特に断りが無い限り、そのベクトルやその行列の全ての要素に対して適用されるものとする。
<第一実施形態のポイント>
 感情とは、受けた感情関連刺激の知覚によって変化した心身の状態のことである。なお、感情関連刺激とは、既存の感情状態に影響を与える、あるいは新たな感情を喚起させる刺激である。感情関連刺激の知覚には、そのときの感情状態が影響するという知見が報告されている。以下に3つの例を示す。
 (例1)ネガティビティ・バイアス(ネガティビティ効果)
 ポジティブな感情を持つ情報より、ネガティブな感情を持つ情報のほうが情報価が高く、ヒトに対して大きな影響を及ぼす傾向がある。ヒトはポジティブな感情よりネガティブな感情に対して敏感で、ネガティブ感情はポジティブ感情より長く続く傾向がある。
 (例2)気分一致判断効果
 ポジティブな感情(気分)時に、より肯定的な評価・判断がなされ、ネガティブな感情(気分)時には否定的な評価・判断がなされやすい。
 (例3)クレショフ効果
 一つの画像や写真が映画的に編集されることによって、その画像や写真がその前後に位置する他の映像の意味に対して及ぼすという性質がある。前後の脈絡がない映像や写真の羅列でも、人間は前後のつながりを無意識に関連づけて意味を解釈してしまう。
 人はこのような感情特性を持つことから、刺激入力時のベースとなる感情状態を考慮して感情状態の推定を行うことで、感情状態の推定精度が向上すると考えられる。
 しかしながら、現在、感情関連刺激を受ける前の感情状態と、それに伴う受けた感情関連刺激の知覚認知への変化を考慮した感情状態変化の推定技術は存在しない。そのため、ある状況である事象が起こったためある感情になったというような時間の流れを考慮した感情の理解や、自然な感情状態変化の説明が困難であるといった課題が存在する。
 本実施形態では、入力された感情関連刺激をフィルタリング処理することで、刺激を受けた際の感情状態を考慮し、より自然な感情状態の変化を推定する。
 より詳しく説明すると、感情状態E(t)のときに受けた刺激S(t)は、感情状態E(t)の影響を受けて知覚される。後述するフィルタ係数推定部において、入力刺激S(t)を実際に知覚・認知される刺激に修正するためのフィルタ係数を感情状態E(t)を用いて推定する。刺激修正部において刺激S(t)のフィルタリングを行い、従来の感情推定手法を用いてフィルタで修正された刺激が誘発させる感情状態E(t+1)を推定する。
<第一実施形態>
 図1は感情状態推定システムの構成例を示す。感情状態推定システムは、感情状態推定学習装置100と感情状態推定装置200を含む。
 感情状態推定学習装置100は、学習用の刺激SL(t)と学習用の感情状態EL(t)、EL(t+1)を入力とし、感情状態推定装置200で用いるパラメータを学習し、学習済みのパラメータを出力する。ただし、tは時点を示すインデックスである。なお、感情状態EL(t+1)は、時点tにおける刺激SL(t)と感情状態EL(t)で得られる時点t+1における感情状態であり、正解データに相当する。感情状態推定学習装置100は、例えば、学習用の刺激SL(t)と学習用の感情状態EL(t)、EL(t+1)の組を複数入力とする。具体的には、刺激SL(t)は特定の感情を喚起させるための刺激であり、喚起させる感情のラベルやレベルスコアが付与されているものとする。付与されている感情のラベルやレベルスコア毎、かつ、刺激の種別毎に識別可能な状態で用意された学習用の「刺激SL(t)_感情ラベル_感情レベルスコア_刺激種別」と、当該「刺激SL(t)_感情ラベル_感情レベルスコア_刺激種別」を受ける際の学習用の感情状態EL(t)、EL(t+1)の組を予め複数用意しておき、感情状態推定学習装置100の入力とする。感情のラベル、レベルスコア、刺激の種別については後述する。ただし、以下では、「刺激SL(t)_感情ラベル_感情レベルスコア_刺激種別」を単に刺激SL(t)と記載する。
 感情状態推定装置200は、感情推定処理に先立ち、学習済みのパラメータを受け取る。学習済みのパラメータは、例えば感情状態推定学習装置100から出力されるパラメータである。感情状態推定装置200は、刺激S(t)と感情状態E(t)を入力とし、時点tの刺激S(t)と感情状態E(t)から時点t+1の感情状態EL(t)を推定し、出力する。
 感情状態推定学習装置および感情状態推定装置は、例えば、中央演算処理装置(CPU: Central Processing Unit)、主記憶装置(RAM: Random Access Memory)などを有する公知又は専用のコンピュータに特別なプログラムが読み込まれて構成された特別な装置である。感情状態推定学習装置および感情状態推定装置は、例えば、中央演算処理装置の制御のもとで各処理を実行する。感情状態推定学習装置および感情状態推定装置に入力されたデータや各処理で得られたデータは、例えば、主記憶装置に格納され、主記憶装置に格納されたデータは必要に応じて中央演算処理装置へ読み出されて他の処理に利用される。感情状態推定学習装置および感情状態推定装置の各処理部は、少なくとも一部が集積回路等のハードウェアによって構成されていてもよい。感情状態推定学習装置および感情状態推定装置が備える各記憶部は、例えば、RAM(Random Access Memory)などの主記憶装置、またはリレーショナルデータベースやキーバリューストアなどのミドルウェアにより構成することができる。ただし、各記憶部は、必ずしも感情状態推定学習装置および感情状態推定装置がその内部に備える必要はなく、ハードディスクや光ディスクもしくはフラッシュメモリ(Flash Memory)のような半導体メモリ素子により構成される補助記憶装置により構成し、感情状態推定学習装置および感情状態推定装置の外部に備える構成としてもよい。
 まず、感情状態推定学習装置について説明する。
<感情状態推定学習装置100>
 図2は第一実施形態に係る感情状態推定学習装置100の機能ブロック図を、図3はその処理フローを示す。
 感情状態推定学習装置100は、フィルタ係数推定部120と刺激修正部130と学習部140とを含む。
 上述の通り、感情状態推定学習装置100は、学習用の刺激SL(t)と学習用の感情状態EL(t)、EL(t+1)を入力とし、感情状態推定装置200で用いるパラメータを学習し、学習済みのパラメータを出力する。
 感情状態推定学習装置100の入力である学習用の刺激SL(t)は、特定の感情を喚起させるための刺激である。刺激SL(t)は、ヒトの視覚器、聴覚器等の感覚器で受け取ることができる刺激であり、例えば、映像データ、画像データ、音声データ等である。このデータの種別が、前述の刺激種別に相当する。なお、前述の通り、刺激SL(t)には、喚起させる感情のラベルやレベルスコアが付与されている。例えば、喜怒哀楽等の感情を示すラベルや、喜怒哀楽等の感情のレベルを示す数値(例えば、各感情のレベルを1~5の数値で示す)が刺激SL(t)に付与されている。
 感情状態推定学習装置100の入力である学習用の感情状態EL(t)は、時点tの刺激SL(t)が入力されるときの時点tの感情状態を示し、例えば、感情のラベルやレベルスコアで表現される。従来技術を用いて感情状態を取得することができる。例えば、(1-1)感情状態推定学習装置100が、対象者に対して学習用の刺激SL(t)である特定の感情を喚起させる映像データを提示し、(1-2)対象者が、入力装置を介して、映像データを視聴しながらそのときの感情状態を入力することで、感情状態推定学習装置100は感情状態を取得することができる。なお、対象者がこの作業をニュートラルな感情状態で行うことで、対象者の感情状態を無視できるようにする。このとき学習用の刺激SL(t)は、フィルタ係数推定に用いる刺激と同一のものでもよい。また、例えば、(2)感情状態推定学習装置100は従来技術を用いて学習した感情状態推定モデルを用いて学習用の感情状態EL(t)を取得してもよい。この感情状態推定モデルは、刺激S(t)を入力とし、刺激S(t)に対する感情状態E(t)を出力するモデルである。
 時点tが経過することで、時点t+1の感情状態EL(t+1)も同様に取得できる。感情状態EL(t+1)は、時点tにおいて入力刺激SL(t)を知覚・認知したことによって喚起された感情状態であって、感情状態EL(t)から変化した後の時点t+1の感情状態である。後述する感情状態推定装置200は、感情状態E(t)がニュートラルであれば、刺激S(t)の知覚後の感情状態E(t+1)は感情状態E(t)の影響を受けないと想定し、刺激S(t)に付与されている感情のラベルやレベルスコアに基づき感情状態E(t+1)を推定できる。しかし実際は時点tにおいてさまざまな感情状態をとりうるため、感情状態推定装置200は、感情状態EL(t)の影響を考慮して時点t+1の感情状態を推定する。
 感情状態推定学習装置100の出力である学習済みのパラメータは、感情状態推定装置200で用いるパラメータである。
 以下、各部について説明する。
<フィルタ係数推定部120>
 フィルタ係数推定部120は、学習用の感情状態EL(t)と学習部140から出力されたパラメータw1,w2,…を入力とし、パラメータw1,w2,…を用いて、感情状態EL(t)から、学習用の刺激SL(t)から抽出される1以上の特徴量を実際に知覚、認知される1以上の特徴量(言い換えると、実際に知覚、認知される刺激に対応する1以上の特徴量)に修正するためのフィルタ係数v1,v2,…を推定し(S120)、出力する。なお、パラメータw1,w2,…の初期値として、例えば、0~1の間の既定値を与えればよい。例えば、フィルタ係数推定部120は、感情状態EL(t)を入力として、次式のように感情状態EL(t)にパラメータw1,w2,…をかけてフィルタ係数v1,v2,…を生成する。
EL(t)・w1=v1
<刺激修正部130>
 刺激修正部130は、学習用の刺激SL(t)とフィルタ係数v1,v2,…を入力とし、フィルタ係数v1,v2,…を用いて、刺激SL(t)から抽出される特徴量をフィルタリングして、実際に知覚、認知される特徴量である刺激に修正し(S130)、出力する。以下、修正後の刺激をS'L(t)と表記する。
 例えば、刺激修正部130は、刺激SL(t)から特徴量QL(t)を抽出する。特徴量の抽出方法としては従来手法を用いればよい。例えば、非特許文献1の方法により、刺激SL(t)である映像データから表情、瞳孔径、心拍等に係る特徴量を抽出してもよい。また、例えば、刺激SL(t)である音声データから音量、口調、声音等に係る特徴量を抽出してもよいし、刺激SL(t)に付与されている感情のラベルやレベルスコアを特徴量として抽出してもよい。
 次に、刺激修正部130は、刺激SL(t)から抽出した特徴量QL(t)にフィルタ係数v1,v2,…を乗じることで、次式のように特徴量を修正する。
v1・QL(t)=Q'L(t)
修正された特徴量Q'L(t)が上述の修正後の刺激S'L(t)に相当する。
<学習部140>
 学習部140は、修正後の刺激S'L(t)と学習用の感情状態EL(t+1)とを入力とし、これらの値を用いて、感情状態EL(t+1)と、修正後の刺激S'(t)を用いて推定される時点t+1の感情状態の推定値^EL(t+1)との距離が小さくなるように、上述のフィルタ係数推定部120でフィルタ係数v1,v2,…を推定する際に用いるモデルのパラメータw1,w2,…を学習し、出力する。
 まず、学習部140は、修正後の刺激S'L(t)を用いて、時点t+1の感情状態の推定値^EL(t+1)を得る(S140-1)。例えば、従来の感情推定手法の学習済みの推定モデルを用いて刺激修正部130で修正された刺激S'L(t)が誘発させる感情状態の推定値^EL(t+1)を得る。例えば、学習済みの推定モデルは、ニューラルネットワーク等の機械学習により学習されたモデルであって、時点tの刺激S(t)から得られる特徴量を入力とし、時点t+1の感情状態を推定するモデルである。なお、感情のラベルやレベルスコアを推定結果としてすればよい。例えば、刺激修正部130において、非特許文献1の方法により、刺激SL(t)である映像データから表情、瞳孔径、心拍等に係る特徴量を抽出する場合には、非特許文献1の学習済みの推定モデルを用いて、修正後の刺激S'L(t)から感情状態の推定値^EL(t+1)を得る。
 次に、学習部140は、時点t+1の学習用の感情状態EL(t+1)とその推定値^EL(t+1)との距離D(t)を計算する(S140-2)。例えば、距離D(t)の計算には、ユークリッド距離、マハラノビス距離が利用可能である。
 さらに、学習部140は、距離D(t)が所定の閾値以下か否かを判定し(S140-3)、所定の閾値より大きい場合(S140-3のno)、例えば、勾配降下法等を用いて、距離D(t)が小さくなるようにパラメータw1,w2,…を更新し(S140-4)、フィルタ係数推定部120に出力し、S120~S140-3を繰り返すように制御する。一方、距離D(t)が所定の閾値以下の場合(S140-3のyes)、感情状態推定学習装置100は、まだ処理していない組(学習データ)について処理を行う。
 学習部140は、感情状態推定学習装置100に入力された、学習用の刺激SL(t)と学習用の感情状態EL(t)とEL(t+1)のすべての組(学習データ)について処理を行ったか否かを判定し(S140-5)、まだ処理していない組(学習データ)がある場合(S140-5のno)、S120~S140-4を繰り返すように制御し、処理していない組(学習データ)がない場合(S140-5のyes)、そのときのパラメータw1,w2,…を学習済みのパラメータとして出力する。
 次に、感情状態推定装置について説明する。
<感情状態推定装置200>
 図4は第一実施形態に係る感情状態推定装置200の機能ブロック図を、図5はその処理フローを示す。
 感情状態推定装置200は、フィルタ係数推定部220と刺激修正部230と感情推定部240とを含む。
 上述の通り、感情状態推定装置200は、感情推定処理に先立ち、学習済みのパラメータを受け取る。感情状態推定装置200は、刺激S(t)と感情状態E(t)を入力とし、時点tの刺激S(t)と感情状態E(t)から時点t+1の感情状態の推定値E(t+1)を得、出力する。感情状態の推定値E^(t+1)は、例えば、感情のラベルやレベルスコアである。
 以下、各部の処理について説明する。
<フィルタ係数推定部220>
 フィルタ係数推定部220は、感情推定処理に先立ち、学習済みのパラメータを受け取る。フィルタ係数推定部220は、感情状態E(t)を入力とし、学習済みのパラメータに基づき感情状態E(t)から、刺激S(t)から抽出される1以上の特徴量を実際に知覚、認知される1以上の特徴量(言い換えると、実際に知覚、認知される刺激に対応する1以上の特徴量)に修正するためのフィルタ係数v1,v2,…を推定し(S220)、出力する。具体的には、フィルタ係数推定部120のS120と同様の処理を行う。
<刺激修正部230>
 刺激修正部230は、刺激S(t)とフィルタ係数v1,v2,…を入力とし、フィルタ係数v1,v2,…を用いて、刺激S(t)から抽出される特徴量をフィルタリングして、実際に知覚、認知される特徴量である刺激に修正し(S230)、出力する。具体的には、刺激修正部130のS130と同様の処理を行う。以下、修正後の刺激をS'(t)と表記する。
<感情推定部240>
 感情推定部240は、修正後の刺激S'(t)を入力とし、修正後の刺激S'(t)を用いて、修正後の刺激S'(t)を用いて推定される時点t+1の感情状態の推定値^E(t+1)を得(S240)、出力する。具体的には、学習部140のS140-1と同様の処理を行う。
<効果>
 以上の構成により、刺激を受けた際の感情状態を考慮し、より自然な感情状態の変化を推定することができる。
 感情状態推定は人間とロボットの交流など、HCI(human-computer interaction)分野における重要な技術の一つである。デジタル化が進む現在、人間とロボットが共存する社会に向けた世界的な取り組みは活発に行われており、人間と自然な相互作用を起こすロボットの開発が目指されている。感情状態の変化を正確に推定する技術により、ロボットが相手の人間に気を配り、感情状態を考慮した行動をとることができれば、さらに自然で豊かなインタラクションが行えると考えられる。
<変形例>
 本実施形態では、感情状態を感情のラベルやレベルスコアで示しているが、感情状態を示す値に継続しやすい感情を示すパラメータ等を追加してもよい。例えば、すぐに反応がある感情変化や、ゆっくりとした感情変化を示すパラメータを追加する。
<第二実施形態>
 第一実施形態と異なる部分を中心に説明する。
 本実施形態では、フィルタ係数推定部においてフィルタ係数を推定する際に、時点tの感情状態だけでなく、時点tの刺激から得られる特徴量も利用する。
 図1は感情状態推定システムの構成例を示す。感情状態推定システムは、感情状態推定学習装置300と感情状態推定装置400を含む。
<感情状態推定学習装置300>
 図2は第一実施形態に係る感情状態推定学習装置300の機能ブロック図を、図3はその処理フローを示す。
 感情状態推定学習装置300は、特徴量抽出部310とフィルタ係数推定部320と刺激修正部130と学習部140とを含む。刺激修正部130と学習部140の処理内容は第一実施形態と同様なので説明を省略する。
<特徴量抽出部310>
 特徴量抽出部310は、学習用の刺激SL(t)を入力とし、刺激SL(t)から特徴量PL(t)を抽出し(S310)、出力する。例えば、従来技術の感情推定技術で用いる特徴量を抽出してもよいし、他の特徴量であってもよい。また、上述の刺激修正部130で抽出する特徴量と同じ特徴量でもよいし、異なる特徴量であってもよい。例えば、特徴量PL(t)は、刺激SL(t)である映像データから抽出される表情、瞳孔径もしくは心拍等に係る特徴量、または、刺激SL(t)である音声データから抽出される音量、口調もしくは声音等に係る特徴量、または、刺激SL(t)に付与されている感情のラベルやレベルスコア、のいずれか1つ以上からなる。また、例えば、上述の<感情状態推定学習装置100>で説明した従来技術を用いて学習した感情状態推定モデル(刺激S(t)を入力とし、刺激S(t)に対する感情状態E(t)を出力するモデル)を用いる場合、学習用の刺激SL(t)から感情状態を推定する際に用いる特徴量を前述の特徴量PLとして抽出する。例えば、参考文献1,2の技術を用いて特徴量PLを抽出することができる。
(参考文献1)Ali Mollahosseini, Behzad Hasani, Mohammad H. Mahoor,"AffectNet: A Database for Facial Expression, Valence, and Arousal Computing in the Wild", IEEE Transactions on Affective Computing, 2017.
(参考文献2)Wafa Mellouk, Wahida Handouzi, "Facial emotion recognition using deep learning: review and insights", Procedia Computer Science, Volume 175, Pages 689-694, 2020.
<フィルタ係数推定部320>
 フィルタ係数推定部320は、学習用の感情状態EL(t)と特徴量PL(t)と学習部140から出力されたパラメータw1,w2,…を入力とし、パラメータw1,w2,…を用いて、感情状態EL(t)と特徴量PL(t)から、学習用の刺激SL(t)から抽出される1以上の特徴量を実際に知覚、認知される1以上の特徴量(言い換えると、実際に知覚、認知される刺激に対応する1以上の特徴量)に修正するためのフィルタ係数v1,v2,…を推定し(S320)、出力する。なお、パラメータw1,w2,…の初期値として、例えば、0~1の間の既定値を与えればよい。例えば、フィルタ係数推定部320は、感情状態EL(t)と特徴量PL(t)とを入力として、次式のように感情状態EL(t)と特徴量PL(t)とパラメータw1,w2,…をかけてフィルタ係数v1,v2,…を生成する。
EL(t)・PL(t)・w1=v1
<感情状態推定装置400>
 図4は第一実施形態に係る感情状態推定装置400の機能ブロック図を、図5はその処理フローを示す。
 感情状態推定装置400は、特徴量抽出部410とフィルタ係数推定部420と刺激修正部230と感情推定部240とを含む。刺激修正部230と感情推定部240の処理内容は第一実施形態と同様なので説明を省略する。
<特徴量抽出部410>
 特徴量抽出部410は、刺激S(t)を入力とし、刺激S(t)から特徴量P(t)を抽出し(S410)、出力する。具体的には、特徴量抽出部310のS310と同様の処理を行う。
<フィルタ係数推定部420>
 フィルタ係数推定部420は、感情推定処理に先立ち、学習済みのパラメータを受け取る。フィルタ係数推定部420は、感情状態E(t)と特徴量P(t)を入力とし、学習済みのパラメータw1,w2,…に基づき、感情状態E(t)と特徴量P(t)から、刺激S(t)から抽出される1以上の特徴量を実際に知覚、認知される1以上の特徴量(言い換えると、実際に知覚、認知される刺激に対応する1以上の特徴量)に修正するためのフィルタ係数v1,v2,…を推定し(S420)、出力する。具体的には、フィルタ係数推定部320のS320と同様の処理を行う。
<効果>
 このような構成とすることで、第一実施形態と同様の効果を得ることができる。さらに、フィルタ係数を推定する際に、刺激の特徴量を考慮することで、より適切に刺激を修正することができる。
<その他の変形例>
 本発明は上記の実施形態及び変形例に限定されるものではない。例えば、上述の各種の処理は、記載に従って時系列に実行されるのみならず、処理を実行する装置の処理能力あるいは必要に応じて並列的にあるいは個別に実行されてもよい。その他、本発明の趣旨を逸脱しない範囲で適宜変更が可能である。
<プログラム及び記録媒体>
 上述の各種の処理は、図6に示すコンピュータの記憶部2020に、上記方法の各ステップを実行させるプログラムを読み込ませ、制御部2010、入力部2030、出力部2040などに動作させることで実施できる。
 この処理内容を記述したプログラムは、コンピュータで読み取り可能な記録媒体に記録しておくことができる。コンピュータで読み取り可能な記録媒体としては、例えば、磁気記録装置、光ディスク、光磁気記録媒体、半導体メモリ等どのようなものでもよい。
 また、このプログラムの流通は、例えば、そのプログラムを記録したDVD、CD-ROM等の可搬型記録媒体を販売、譲渡、貸与等することによって行う。さらに、このプログラムをサーバコンピュータの記憶装置に格納しておき、ネットワークを介して、サーバコンピュータから他のコンピュータにそのプログラムを転送することにより、このプログラムを流通させる構成としてもよい。
 このようなプログラムを実行するコンピュータは、例えば、まず、可搬型記録媒体に記録されたプログラムもしくはサーバコンピュータから転送されたプログラムを、一旦、自己の記憶装置に格納する。そして、処理の実行時、このコンピュータは、自己の記録媒体に格納されたプログラムを読み取り、読み取ったプログラムに従った処理を実行する。また、このプログラムの別の実行形態として、コンピュータが可搬型記録媒体から直接プログラムを読み取り、そのプログラムに従った処理を実行することとしてもよく、さらに、このコンピュータにサーバコンピュータからプログラムが転送されるたびに、逐次、受け取ったプログラムに従った処理を実行することとしてもよい。また、サーバコンピュータから、このコンピュータへのプログラムの転送は行わず、その実行指示と結果取得のみによって処理機能を実現する、いわゆるASP(Application Service Provider)型のサービスによって、上述の処理を実行する構成としてもよい。なお、本形態におけるプログラムには、電子計算機による処理の用に供する情報であってプログラムに準ずるもの(コンピュータに対する直接の指令ではないがコンピュータの処理を規定する性質を有するデータ等)を含むものとする。
 また、この形態では、コンピュータ上で所定のプログラムを実行させることにより、本装置を構成することとしたが、これらの処理内容の少なくとも一部をハードウェア的に実現することとしてもよい。

Claims (8)

  1.  パラメータwを用いて、ある時点tの学習用の感情状態EL(t)から、前記時点tの学習用の刺激SL(t)から抽出される1以上の特徴量を実際に知覚、認知される1以上の特徴量に修正するためのフィルタ係数を推定するフィルタ係数推定部と、
     前記フィルタ係数を用いて、前記刺激SL(t)から抽出される特徴量をフィルタリングして、実際に知覚、認知される特徴量である刺激S'L(t)に修正する刺激修正部と、
     修正後の刺激S'L(t)と時点t+1の学習用の感情状態EL(t+1)とを用いて、前記感情状態EL(t+1)と、修正後の刺激S'(t)を用いて推定される時点t+1の感情状態の推定値^EL(t+1)との距離が小さくなるように、前記パラメータwを学習する学習部と、を含む、
     感情状態推定学習装置。
  2.  請求項1の感情状態推定学習装置であって、
     前記刺激SL(t)から特徴量PL(t)を抽出する特徴量抽出部を含み、
     前記フィルタ係数推定部は、前記パラメータwを用いて、前記感情状態EL(t)と前記特徴量PL(t)から、前記フィルタ係数を推定し、
     前記特徴量PL(t)は、映像データから抽出される表情、瞳孔径もしくは心拍に係る特徴量、または、音声データから抽出される音量、口調もしくは声音に係る特徴量、または、感情のラベルもしくは感情のレベルスコア、のいずれか1つ以上からなる、
     感情状態推定学習装置。
  3.  請求項1または請求項2の感情状態推定学習装置であって、
     当該感情状態推定学習装置は、学習用の前記刺激SL(t)と学習用の前記感情状態EL(t)、EL(t+1)の組を複数入力とし、
     学習用の前記刺激SL(t)は、特定の感情を喚起させるための刺激であり、
     学習用の前記刺激SL(t)は、学習用の前記刺激SL(t)に付与されている感情のラベル、感情レベルスコア、刺激の種別、それぞれ毎に識別可能な状態で用意されたデータである、
     感情状態推定学習装置。
  4.  学習済みのパラメータwに基づきある時点tの感情状態E(t)から前記時点tの刺激S(t)から抽出される1以上の特徴量を実際に知覚、認知される1以上の特徴量に修正するためのフィルタ係数を推定するフィルタ係数推定部と、
     前記フィルタ係数を用いて、前記刺激S(t)から抽出される特徴量をフィルタリングして、実際に知覚、認知される特徴量である刺激S'(t)に修正する刺激修正部と、
     修正後の刺激S'(t)を用いて、時点t+1の感情状態E(t+1)を推定する感情推定部と、を含み、
     前記学習済みのパラメータwは、学習用の感情状態EL(t+1)とその推定値との距離が小さくなるように学習されたものである、
     感情状態推定装置。
  5.  請求項4の感情状態推定装置であって、
     前記刺激S(t)から特徴量P(t)を抽出する特徴量抽出部を含み、
     前記フィルタ係数推定部は、前記学習済みのパラメータwを用いて、前記感情状態E(t)と前記特徴量P(t)から、前記フィルタ係数を推定し、
     前記特徴量P(t)は、映像データから抽出される表情、瞳孔径もしくは心拍に係る特徴量、または、音声データから抽出される音量、口調もしくは声音に係る特徴量、または、感情のラベルもしくは感情のレベルスコア、のいずれか1つ以上からなる、
     感情状態推定装置。
  6.  感情状態推定学習装置が、パラメータwを用いて、ある時点tの学習用の感情状態EL(t)から、前記時点tの学習用の刺激SL(t)から抽出される1以上の特徴量を実際に知覚、認知される1以上の特徴量に修正するためのフィルタ係数を推定するフィルタ係数推定ステップと、
     感情状態推定学習装置が、前記フィルタ係数を用いて、前記刺激SL(t)から抽出される特徴量をフィルタリングして、実際に知覚、認知される特徴量である刺激S'L(t)に修正する刺激修正ステップと、
     感情状態推定学習装置が、修正後の刺激S'L(t)と時点t+1の学習用の感情状態EL(t+1)とを用いて、前記感情状態EL(t+1)と、修正後の刺激S'(t)を用いて推定される時点t+1の感情状態の推定値^EL(t+1)との距離が小さくなるように、前記パラメータwを学習する学習ステップと、を含む、
     感情状態推定学習方法。
  7.  感情状態推定装置が、学習済みのパラメータwに基づきある時点tの感情状態E(t)から前記時点tの刺激S(t)から抽出される1以上の特徴量を実際に知覚、認知される1以上の特徴量に修正するためのフィルタ係数を推定するフィルタ係数推定ステップと、
     感情状態推定装置が、前記フィルタ係数を用いて、前記刺激S(t)から抽出される特徴量をフィルタリングして、実際に知覚、認知される特徴量である刺激S'(t)に修正する刺激修正ステップと、
     感情状態推定装置が、修正後の刺激S'(t)を用いて、時点t+1の感情状態E(t+1)を推定する感情推定ステップと、を含み、
     前記学習済みのパラメータwは、学習用の感情状態EL(t+1)とその推定値との距離が小さくなるように学習されたものである、
     感情状態推定方法。
  8.  請求項1から請求項3の何れかの感情状態推定学習装置、または、請求項4もしくは請求項5の感情状態推定装置として、コンピュータを機能させるためのプログラム。
PCT/JP2021/018419 2021-05-14 2021-05-14 感情状態推定学習装置、感情状態推定装置、それらの方法、およびプログラム Ceased WO2022239234A1 (ja)

Priority Applications (2)

Application Number Priority Date Filing Date Title
PCT/JP2021/018419 WO2022239234A1 (ja) 2021-05-14 2021-05-14 感情状態推定学習装置、感情状態推定装置、それらの方法、およびプログラム
JP2023520721A JP7540588B2 (ja) 2021-05-14 2021-05-14 感情状態推定学習装置、感情状態推定装置、それらの方法、およびプログラム

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
PCT/JP2021/018419 WO2022239234A1 (ja) 2021-05-14 2021-05-14 感情状態推定学習装置、感情状態推定装置、それらの方法、およびプログラム

Publications (1)

Publication Number Publication Date
WO2022239234A1 true WO2022239234A1 (ja) 2022-11-17

Family

ID=84028957

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/JP2021/018419 Ceased WO2022239234A1 (ja) 2021-05-14 2021-05-14 感情状態推定学習装置、感情状態推定装置、それらの方法、およびプログラム

Country Status (2)

Country Link
JP (1) JP7540588B2 (ja)
WO (1) WO2022239234A1 (ja)

Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2014144052A (ja) * 2013-01-28 2014-08-14 Nippon Telegr & Teleph Corp <Ntt> 感情推定方法、装置及びプログラム
US20180285641A1 (en) * 2014-11-06 2018-10-04 Samsung Electronics Co., Ltd. Electronic device and operation method thereof

Family Cites Families (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
TWI221574B (en) * 2000-09-13 2004-10-01 Agi Inc Sentiment sensing method, perception generation method and device thereof and software
EP3409205A1 (en) * 2017-05-30 2018-12-05 Koninklijke Philips N.V. Device, system and method for determining an emotional state of a user
WO2019017124A1 (ja) * 2017-07-19 2019-01-24 パナソニックIpマネジメント株式会社 眠気推定装置及び覚醒誘導装置

Patent Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2014144052A (ja) * 2013-01-28 2014-08-14 Nippon Telegr & Teleph Corp <Ntt> 感情推定方法、装置及びプログラム
US20180285641A1 (en) * 2014-11-06 2018-10-04 Samsung Electronics Co., Ltd. Electronic device and operation method thereof

Also Published As

Publication number Publication date
JP7540588B2 (ja) 2024-08-27
JPWO2022239234A1 (ja) 2022-11-17

Similar Documents

Publication Publication Date Title
CN115861995A (zh) 一种视觉问答方法、装置及电子设备和存储介质
US9724824B1 (en) Sensor use and analysis for dynamic update of interaction in a social robot
JP6450138B2 (ja) 情報処理装置及び発話内容出力方法
CN113705792B (zh) 基于深度学习模型的个性化推荐方法、装置、设备及介质
CN108074203A (zh) 一种教学调整方法和装置
CN109155110A (zh) 信息处理装置以及其控制方法、计算机程序
CN117252739B (zh) 一种评卷方法、系统、电子设备及存储介质
CN111783473B (zh) 医疗问答中最佳答案的识别方法、装置和计算机设备
CN119166854A (zh) 面向复杂不完备数据场景的短视频谣言检测方法及系统
Tripathi et al. Facial emotion-based song recommender system using CNN
CN119168842A (zh) 图像生成方法、装置、电子设备和计算机可读存储介质
WO2022239234A1 (ja) 感情状態推定学習装置、感情状態推定装置、それらの方法、およびプログラム
TWI874785B (zh) 資訊處理方法、資訊處理裝置以及計算機系統
CN117334296A (zh) 一种语义记忆辅助训练方法及设备
CN113536809B (zh) 一种基于语义的无监督常识问答方法及系统
JP2024019892A (ja) 脳応答空間生成装置、評価装置、及び脳応答空間生成方法
JP6856965B1 (ja) 画像出力装置及び画像出力方法
CN108985456B (zh) 层数增减深度学习神经网络训练方法、系统、介质和设备
JP2022148878A (ja) プログラム、情報処理装置、及び方法
CN119694325B (zh) 基于细粒度对比学习的副语言信息识别方法及系统
US20260087019A1 (en) Method and apparatus for personalization of generative language model
CN119004391B (zh) 基于情感感知表征解耦与偏移的多模态情感分析方法及产品
US20250259561A1 (en) Interactive digital learning system
JP7634243B2 (ja) 情報処理装置及び情報処理方法
JP7843978B1 (ja) 情報処理装置、情報処理方法及びプログラム

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 21941964

Country of ref document: EP

Kind code of ref document: A1

WWE Wipo information: entry into national phase

Ref document number: 2023520721

Country of ref document: JP

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 21941964

Country of ref document: EP

Kind code of ref document: A1