WO2020085070A1 - パラ言語情報推定装置、パラ言語情報推定方法、およびプログラム - Google Patents

パラ言語情報推定装置、パラ言語情報推定方法、およびプログラム Download PDF

Info

Publication number
WO2020085070A1
WO2020085070A1 PCT/JP2019/039572 JP2019039572W WO2020085070A1 WO 2020085070 A1 WO2020085070 A1 WO 2020085070A1 JP 2019039572 W JP2019039572 W JP 2019039572W WO 2020085070 A1 WO2020085070 A1 WO 2020085070A1
Authority
WO
WIPO (PCT)
Prior art keywords
feature amount
feature
paralinguistic information
model
information estimation
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/JP2019/039572
Other languages
English (en)
French (fr)
Inventor
厚志 安藤
歩相名 神山
哲 小橋川
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
NTT Inc
Original Assignee
Nippon Telegraph and Telephone Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Nippon Telegraph and Telephone Corp filed Critical Nippon Telegraph and Telephone Corp
Priority to US17/287,102 priority Critical patent/US11798578B2/en
Publication of WO2020085070A1 publication Critical patent/WO2020085070A1/ja
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F40/00Handling natural language data
    • G06F40/30Semantic analysis
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N20/00Machine learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/044Recurrent networks, e.g. Hopfield networks
    • G06N3/0442Recurrent networks, e.g. Hopfield networks characterised by memory or gating, e.g. long short-term memory [LSTM] or gated recurrent units [GRU]
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/08Learning methods
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/08Learning methods
    • G06N3/09Supervised learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/70Arrangements for image or video recognition or understanding using pattern recognition or machine learning
    • G06V10/82Arrangements for image or video recognition or understanding using pattern recognition or machine learning using neural networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V40/00Recognition of biometric, human-related or animal-related patterns in image or video data
    • G06V40/10Human or animal bodies, e.g. vehicle occupants or pedestrians; Body parts, e.g. hands
    • G06V40/16Human faces, e.g. facial parts, sketches or expressions
    • G06V40/168Feature extraction; Face representation
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V40/00Recognition of biometric, human-related or animal-related patterns in image or video data
    • G06V40/20Movements or behaviour, e.g. gesture recognition
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/02Feature extraction for speech recognition; Selection of recognition unit
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/08Speech classification or search
    • G10L15/16Speech classification or search using artificial neural networks
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L25/00Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
    • G10L25/48Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use
    • G10L25/51Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use for comparison or discrimination
    • G10L25/63Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use for comparison or discrimination for estimating an emotional state
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F40/00Handling natural language data
    • G06F40/20Natural language analysis
    • G06F40/279Recognition of textual entities
    • G06F40/284Lexical analysis, e.g. tokenisation or collocates
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F40/00Handling natural language data
    • G06F40/20Natural language analysis
    • G06F40/279Recognition of textual entities
    • G06F40/289Phrasal analysis, e.g. finite state techniques or chunking

Definitions

  • the present invention relates to a technique for estimating paralinguistic information from speech.
  • Paralinguistic information for example, whether the utterance intention is doubtful or humorous, or whether the emotion is joy, sadness, anger, or calmness
  • Paralinguistic information is, for example, advanced sophistication of speech translation (for example, for a Japanese utterance "Tomorrow", understand the question intention "Tomorrow?" And translate it into English “Is it tomorrow?" Understands the intent "Tomorrow.” And translates it into English, such as "It is tomorrow.”
  • the present invention can be applied to dialogue control that takes into account the feeling of the other party (for example, changing the topic if the other party is angry).
  • Non-Patent Document 1 and the like show a paralinguistic information estimation technique using a plurality of independent feature amounts.
  • a speaker's emotional dimension value (Valence: pleasantness / unpleasantness, Arousal: arousal) based on a voice feature (voice waveform) and a video feature (image sequence of multiple frames).
  • -Sleep There is also known a technique for estimating the paralinguistic information of a speaker based on time-series information of prosodic features such as pitch of voice for each short time of speech and time-series information of language features such as spoken words. Has been.
  • the technique of combining a plurality of feature amounts can recognize paralinguistic information with higher accuracy than the technique of using the feature amount alone.
  • FIG. 1 illustrates a conventional technique of a paralinguistic information estimation model using a plurality of independent feature quantities.
  • the paralinguistic information estimation model 900 includes a feature amount submodel 101 that estimates paralinguistic information from each feature amount, and a result integration submodel 104 that integrates the outputs of the feature amount submodel 101 and outputs a final paralinguistic information estimation result. Composed of.
  • this configuration includes whether the prosodic feature includes a question or a characteristic of normalization (for example, whether the ending is raised) or whether the linguistic feature shows a question or a characteristic of normality (for example, a questionable question). It is equivalent to the process of estimating whether the utterance intention is doubtful or normal by integrating the results after estimating.
  • paralinguistic information estimation models based on deep learning in which each sub-model is composed of a model based on deep learning and the entire paralinguistic information estimation model is integrally learned, have become mainstream.
  • Paragraph information does not necessarily show the characteristic in all feature quantities, but the characteristic of paralinguistic information may appear in only some feature quantities.
  • the utterance intention there is an utterance in which the way of speaking is a final word, but the sentence is a plain sentence (that is, the characteristic of the question utterance appears only in the prosody characteristics), and such an utterance is regarded as the question utterance.
  • the sentence is a plain sentence (that is, the characteristic of the question utterance appears only in the prosody characteristics)
  • such an utterance is regarded as the question utterance.
  • emotion there is an utterance that appears calm from the facial expression, but angry is strongly expressed as a way of speaking or a word, and such an utterance is regarded as an angry emotional utterance.
  • model learning is performed as if all feature quantities exhibit the same characteristics of paralinguistic information. For example, in the case of learning a question utterance, the learning is performed as if the characteristics of the question utterance appear in both the prosody features and the language features. Therefore, even if a utterance whose characteristic is a question utterance is expressed only in the prosodic feature, model learning is performed assuming that the characteristic of the question utterance is also expressed in the language feature. It becomes a noise in learning correctly.
  • the present invention includes, in the paralinguistic information estimation using a plurality of independent feature amounts, utterances in which the characteristics of the paralinguistic information appear in only some of the feature amounts in the learning data.
  • the objective is to correctly learn the paralinguistic information estimation model and correctly estimate the paralinguistic information.
  • a paralinguistic information estimation device is a paralinguistic information estimation device that estimates paralinguistic information from an input utterance, and uses a plurality of independent feature quantities as input.
  • the paralinguistic information estimation model storage unit that stores the paralinguistic information estimation model that outputs the linguistic information estimation result, the feature amount extraction unit that extracts multiple independent feature amounts from the input utterance, and the paralinguistic information estimation model
  • a paralinguistic information estimation unit that estimates paralinguistic information of the input utterance from a plurality of independent feature quantities extracted from the input utterance, and the paralinguistic information estimation model includes only that feature quantity for each of the plurality of independent feature quantities.
  • the feature amount sub-model Based on the output result of the feature amount sub-model for each of a plurality of independent feature amounts, the feature amount sub-model that outputs information used for estimating paralinguistic information based on A feature amount weight calculation unit that calculates a feature amount weight that indicates whether or not to use it for estimating the language information, and outputs the output result of the feature amount submodel for each of a plurality of independent feature amounts by weighting with the feature amount weight.
  • a feature amount gate and a result integrated sub-model that estimates paralinguistic information based on the output results of all feature amount gates are included.
  • a paralinguistic information estimation model is correctly learned even for utterances in which the characteristics of paralinguistic information appear only in some feature quantities. , Will be able to estimate paralinguistic information correctly. As a result, the accuracy of paralinguistic information estimation is improved.
  • FIG. 1 is a diagram illustrating a conventional paralinguistic information estimation model.
  • FIG. 2 is a diagram illustrating a paralinguistic information estimation model of the present invention.
  • FIG. 3 is a diagram illustrating a functional configuration of the paralinguistic information estimation model learning device.
  • FIG. 4 is a diagram illustrating a processing procedure of the paralinguistic information estimation model learning method.
  • FIG. 5 is a figure which illustrates the paralinguistic information estimation model of 1st embodiment.
  • FIG. 6 is a diagram illustrating a functional configuration of the paralinguistic information estimation model learning unit.
  • FIG. 7 is a diagram illustrating a functional configuration of the paralinguistic information estimation device.
  • FIG. 8 is a diagram illustrating a processing procedure of the paralinguistic information estimation method.
  • FIG. 9 is a diagram illustrating a paralinguistic information estimation model according to the second embodiment.
  • the point of the present invention is to consider the possibility that the characteristic of the paralinguistic information appears only in a part of the feature amount, and introduce a feature amount gate that determines whether or not to use the information of each feature amount for paralinguistic information estimation.
  • a feature amount gate that determines whether or not to use the information of each feature amount for paralinguistic information estimation.
  • this selection mechanism is realized in the form of a feature amount gate.
  • FIG. 2 shows an example of the paralinguistic information estimation model of the present invention.
  • the paralinguistic information estimation model 100 includes a feature amount sub-model 101 similar to the conventional one, a feature amount gate 103 for determining whether or not to use the output of the feature amount sub model 101 for para-linguistic information estimation, and a feature amount gate. And a result integration sub-model 104 that outputs a final paralinguistic information estimation result based on the output of 103.
  • the feature amount gate 103 has a role of determining whether to output the output of each feature amount sub-model 101 to the result integrated sub-model 104.
  • the feature amount gate 103 determines the output based on the equation (1).
  • y k is a feature gate output vector
  • x k is a feature gate input vector (feature submodel output result)
  • w k is a feature gate Weight vector
  • the feature amount gate weight vector w k When the feature amount gate weight vector w k is a unit vector, the feature amount submodel output result x k is output to the result integrated submodel 104 as it is. When the feature amount gate weight vector w k is a zero vector, the feature amount sub model output result x k is converted to zero and output to the result integrated sub model 104. In this way, by controlling the feature amount gate weight vector w k corresponding to each feature amount, one feature amount is used but another feature amount is not used. It becomes possible to estimate information. In the case of a paralinguistic information estimation model based on deep learning, the feature amount gate weight vector w k can be regarded as one of the model parameters. Therefore, the entire model including the feature amount gate weight vector w k must be integrally learned. Is possible.
  • the paralinguistic information is estimated by the following procedure.
  • a paralinguistic information estimation model that consists of a sub-model for each feature amount, a feature amount gate for each feature amount, and a result integrated sub-model is prepared by inputting multiple independent feature amounts.
  • the entire model including the weight vector of the feature amount gate is integrally learned by the error backpropagation method.
  • the weight vector of the feature amount gate is determined by a rule manually. For example, when the output result of the sub-model for each feature is the distance from the discrimination plane, if the absolute value of the distance from the discrimination plane is 0.5 or less, the weight vector of the feature gate is zero vector, and the absolute value of the distance from the discrimination plane. If is greater than 0.5, the weight vector of the feature gate is a unit vector. In this case, two-stage learning is performed, in which the sub-model for each feature quantity is first learned and then the resultant integrated sub-model is learned.
  • a plurality of independent feature quantities are input to the learned paralinguistic information estimation model, and the paralinguistic information estimation result is obtained for each utterance.
  • the input utterance refers to both the voice waveform information of the utterance and the image information of the facial expression (face) of the speaker of the utterance.
  • the feature amount used for paralinguistic information estimation in the present invention may be two or more independent feature amounts that can be extracted from human speech, but in the present embodiment, prosodic features, language features, and video features are independent of each other. Three types of feature quantities shall be used. However, of these three types of feature amounts, only two types of feature amounts may be used. Further, as long as it is independent of other feature amounts, a feature amount using information such as biological signal information (pulse, skin potential, etc.) may be additionally used.
  • the paralinguistic information probability for each feature amount can be received as the output result of the sub-model for each feature amount, but the intermediate information necessary for estimating the paralinguistic information probability for each feature amount (for example, It is also possible to receive the output value of the middle layer in the deep neural network).
  • the weight vector is not a fixed value for all inputs, and the weight vector can be dynamically changed every time the input changes. Specifically, the weight vector is dynamically changed by calculating the weight vector from the input using the equation (2) or the equation (3).
  • x k is a feature gate input vector (feature sub model output result)
  • w k is a feature gate weight vector
  • w x is a feature gate A weight vector calculation matrix
  • b x is a feature amount gate weight vector calculation bias
  • is an activation function (for example, the sigmoid function of Expression (4)).
  • w x and b x are determined in advance by learning.
  • the degree of use of the output result of the sub-model for each feature amount is changed according to the speaker of the input utterance and the utterance environment (for example, prosodic features are more prominent for people who easily show paralinguistic information in intonation).
  • Paralinguistic information can be estimated with emphasis). Therefore, compared to a general estimation method based on a weighted sum of paralinguistic information probabilities for each feature amount, it is possible to estimate paralinguistic information with high accuracy even for more diverse inputs. That is, the accuracy of paralinguistic information estimation for various speech environments is improved.
  • the paralinguistic information estimation model learning device of the first embodiment learns the paralinguistic information estimation model from utterances to which a teacher label is added.
  • the paralinguistic information estimation model learning device includes a speech storage unit 10-1, a teacher label storage unit 10-2, a prosody feature extraction unit 11-1, a language feature extraction unit 11-2, and a video feature.
  • An extraction unit 11-3, a paralinguistic information estimation model learning unit 12, and a paralinguistic information estimation model storage unit 20 are provided.
  • the prosody feature extraction unit 11-1, the language feature extraction unit 11-2, and the video feature extraction unit 11-3 may be collectively referred to as a feature amount extraction unit 11.
  • the feature quantity extraction unit 11 changes the configuration such as the number and the processing content according to the type of feature quantity used for paralinguistic information estimation.
  • This paralinguistic information estimation model learning device implements the paralinguistic information estimation model learning method of the first embodiment by performing the processing of each step illustrated in FIG.
  • the paralinguistic information estimation model learning device is configured by loading a special program into a known or dedicated computer having, for example, a central processing unit (CPU: Central Processing Unit) and a main storage device (RAM: Random Access Memory). It is a special device.
  • the paralinguistic information estimation model learning device executes each process under the control of the central processing unit, for example.
  • the data input to the paralinguistic information estimation model learning device and the data obtained by each process are stored in, for example, the main storage device, and the data stored in the main storage device is read to the central processing unit as necessary. Issued and used for other processing.
  • At least a part of each processing unit of the paralinguistic information estimation model learning device may be configured by hardware such as an integrated circuit.
  • Each storage unit included in the paralinguistic information estimation model learning device is, for example, a main storage device such as a RAM (Random Access Memory) or an auxiliary storage configured by a semiconductor memory device such as a hard disk, an optical disk, or a flash memory (Flash Memory). It can be configured by a device or middleware such as a relational database or a key-value store.
  • a main storage device such as a RAM (Random Access Memory) or an auxiliary storage configured by a semiconductor memory device such as a hard disk, an optical disk, or a flash memory (Flash Memory). It can be configured by a device or middleware such as a relational database or a key-value store.
  • the utterance storage unit 10-1 stores utterances (hereinafter, also referred to as “learning utterances”) used for learning the paralinguistic information estimation model.
  • the utterance is assumed to be composed of voice waveform information in which a human uttered voice is recorded and video information in which a facial expression of a speaker of the utterance is recorded. What kind of information the utterance is composed of is determined according to what kind of feature amount is used for estimating the paralinguistic information.
  • the teacher label storage unit 10-2 stores a teacher label representing the correct answer value of the paralinguistic information given to each utterance stored in the utterance storage unit 10-1.
  • the teacher label may be added to the utterance manually or by using a well-known label classification technique. Specifically, what kind of teacher label is given depends on what kind of feature amount is used to estimate the paralinguistic information.
  • the prosody feature extraction unit 11-1 extracts prosody features from the speech waveform information of each utterance stored in the utterance storage unit 10-1.
  • the prosodic features are one or more of the fundamental frequency, short time power, MFCC (Mel-frequency Cepstral Coefficients), zero crossing rate, Harmonics-to-Noise-Ratio (HNR), and mel filter bank output. It is a vector that contains. Further, it may be a series vector for each time (for each frame), or a vector for statistics (average, variance, maximum value, minimum value, gradient, etc.) of the entire utterance.
  • the prosody feature extraction unit 11-1 outputs the extracted prosody features to the paralinguistic information estimation model learning unit 12.
  • the language feature extraction unit 11-2 extracts language features from the speech waveform information of each utterance stored in the utterance storage unit 10-1.
  • a word string acquired by a speech recognition technique or a phoneme string acquired by a phoneme recognition technique is used.
  • the language feature may be expressed as a sequence vector of these word strings or phoneme strings, or may be a vector representing the number of appearances of a specific word in the entire utterance.
  • the language feature extraction unit 11-2 outputs the extracted language feature to the paralinguistic information estimation model learning unit 12.
  • the video feature extraction unit 11-3 extracts video features from the video information of each utterance stored in the utterance storage unit 10-1.
  • the video features are one or more of the position coordinates of the facial feature points in each frame, the velocity component for each small area calculated from the optical flow, and the local image gradient histogram (Histograms of Oriented Gradients: HOG). It is a vector that contains. Further, it may be a sequence vector for each time (frame) at these fixed intervals, or a vector for statistics (average, variance, maximum value, minimum value, gradient, etc.) of the entire utterance. Good.
  • the video feature extraction unit 11-3 outputs the extracted video features to the paralinguistic information estimation model learning unit 12.
  • step S12 the paralinguistic information estimation model learning unit 12 uses the prosody features, the language features, and the video features that have been input, and the teacher labels stored in the teacher label storage unit 10-2 to make a plurality of independent labels.
  • a paralinguistic information estimation model that outputs the paralinguistic information estimation result by inputting the feature quantity is learned.
  • the paralinguistic information estimation model learning unit 12 stores the learned paralinguistic information estimation model in the paralinguistic information estimation model storage unit 20.
  • FIG. 5 shows a configuration example of the paralinguistic information estimation model used in this embodiment.
  • This paralinguistic information estimation model includes a prosody feature sub-model 101-1, a language feature sub-model 101-2, a video feature sub-model 101-3, a prosody feature weight calculator 102-1, a language feature weight calculator 102-2, The image feature weight calculator 102-3, the prosody feature gate 103-1, the language feature gate 103-2, the image feature gate 103-3, and the result integrated sub-model 104 are provided.
  • the prosody feature sub-model 101-1, the language feature sub-model 101-2, and the video feature sub-model 101-3 are the feature amount sub-model 101, the prosody feature weight calculator 102-1 and the language feature weight calculator 102-.
  • the feature amount sub-model 101 estimates the paralinguistic information based only on the input feature amount, and outputs the paralinguistic estimation result or an intermediate value generated during the paralinguistic estimation (hereinafter, referred to as “para-linguistic information estimation. (Also referred to as "information to be used").
  • the feature amount weight calculation unit 102 based on the output result of the feature amount sub-model 101, represents a feature amount gate weight vector (hereinafter, also referred to as “feature amount weight”) that indicates whether the feature amount is used for estimating paralinguistic information. ) Is calculated.
  • feature amount gate 103 weights the output result of the feature amount sub-model 101 with the feature amount gate weight vector output from the feature amount weight calculation unit 102 and outputs the weighted result.
  • the result integrated sub-model 104 estimates the paralinguistic information based on the output results of all the feature amount gates 103.
  • the paralinguistic information estimation model may be, for example, Deep Neural Network (DNN) based on deep learning or Support Vector Machine (SVM).
  • DNN Deep Neural Network
  • SVM Support Vector Machine
  • an estimation model such as Long Short-Term Memory Recurrent Neural Network (LSTM-RNN) that can consider sequences may be used.
  • LSTM-RNN Long Short-Term Memory Recurrent Neural Network
  • the paralinguistic information estimation model is constructed by a deep learning-based method that includes DNN and LSTM-RNN, the entire model including the weight vector of the feature gate should be regarded as a single network (classification model). Therefore, it is possible to integrally learn the entire paralinguistic information estimation model by the error back propagation method.
  • the paralinguistic information estimation model learning unit 12 includes a prosody feature submodel learning unit 121-1, a language feature submodel learning unit 121-2, a video feature submodel learning unit 121-3, and a prosody feature weight calculating unit 122-. 1, language feature weight calculator 122-2, video feature weight calculator 122-3, prosody feature gate processor 123-1, language feature gate processor 123-2, video feature gate processor 123-3, and result integration
  • the sub model learning unit 124 is provided.
  • the prosody feature submodel learning unit 121-1 learns a prosody feature submodel that estimates paralinguistic information based on only the prosody features from the set of the prosody features and the teacher label.
  • the prosody feature sub-model uses, for example, SVM, but other machine learning methods capable of classifying may be used.
  • the output result of the prosody feature sub-model refers to the distance from the discrimination plane if the prosody feature sub-model is SVM, for example.
  • the language feature submodel learning unit 121-2 and the image feature submodel learning unit 121-3 learn the language feature submodel and the image feature submodel in the same manner as the prosody feature submodel learning unit 121-1.
  • the prosody feature weight calculation unit 122-1 calculates a prosody feature gate weight vector from the output result of the prosody feature sub model using the feature amount gate rule.
  • the feature amount gate rule refers to a set of a rule for determining the feature amount gate and a weight vector of the feature amount gate. If the prosody feature sub-model is an example of SVM, “If the absolute value of the distance from the discrimination plane is 0.5 or less in the output result of the prosody feature sub-model, the prosody feature gate weight vector is zero vector, If the value is greater than 0.5, the prosodic feature gate weight vector is a unit vector.
  • the language feature weight calculation unit 122-2 and the image feature weight calculation unit 122-3 calculate the language feature weight vector and the image feature weight vector in the same manner as the prosody feature weight calculation unit 122-1.
  • the prosody feature gate processing unit 123-1 calculates the above formula (1) using the output result of the prosody feature submodel and the prosody feature gate weight vector to obtain the prosody feature gate output vector.
  • the language feature gate processing unit 123-2 and the video feature gate processing unit 123-3 calculate the language feature gate output vector and the video feature gate output vector in the same manner as the prosody feature gate processing unit 123-1.
  • the result integrated submodel learning unit 124 learns the result integrated submodel from a set of prosody feature gate output vector, language feature gate output vector, video feature gate output vector, and teacher label.
  • the result integration sub-model uses, for example, SVM, but other machine learning methods capable of classifying may be used.
  • the paralinguistic information estimation device of the first embodiment estimates paralinguistic information from an input utterance using a learned paralinguistic information estimation model.
  • the paralinguistic information estimating device includes a prosody feature extracting unit 11-1, a language feature extracting unit 11-2, a video feature extracting unit 11-3, a paralinguistic information estimating model storage unit 20, and a paralinguistic information estimating model storage unit 20.
  • the language information estimation unit 21 is provided.
  • the paralinguistic information estimating method of the first embodiment is realized by the paralinguistic information estimating device performing the processing of each step illustrated in FIG. 8.
  • the paralinguistic information estimation device is configured by loading a special program into a publicly known or dedicated computer having, for example, a central processing unit (CPU: Central Processing Unit) and a main storage device (RAM: Random Access Memory). It is a special device.
  • the paralinguistic information estimation device executes each process under the control of the central processing unit, for example.
  • the data input to the paralinguistic information estimation device and the data obtained by each process are stored in, for example, the main storage device, and the data stored in the main storage device is read to the central processing unit as necessary. Used for other processing.
  • At least a part of each processing unit of the paralinguistic information estimation device may be configured by hardware such as an integrated circuit.
  • Each storage unit included in the paralinguistic information estimation device is, for example, a main storage device such as a RAM (Random Access Memory), an auxiliary storage device configured by a semiconductor memory element such as a hard disk, an optical disc, or a flash memory (Flash Memory), Alternatively, it can be configured by middleware such as a relational database or a key value store.
  • a main storage device such as a RAM (Random Access Memory)
  • auxiliary storage device configured by a semiconductor memory element such as a hard disk, an optical disc, or a flash memory (Flash Memory)
  • middleware such as a relational database or a key value store.
  • step S11-1 the prosody feature extraction unit 11-1 extracts prosody features from the speech waveform information of the input utterance.
  • the extraction of prosody features may be performed in the same manner as in the paralinguistic information estimation model learning device.
  • the prosody feature extraction unit 11-1 outputs the extracted prosody features to the paralinguistic information estimation unit 21.
  • step S11-2 the language feature extraction unit 11-2 extracts language features from the speech waveform information of the input utterance. Extraction of language features may be performed in the same manner as in the paralinguistic information estimation model learning device. The language feature extraction unit 11-2 outputs the extracted language feature to the paralinguistic information estimation unit 21.
  • step S11-3 the video feature extraction unit 11-3 extracts video features from the video information of the input utterance. Extraction of video features may be performed in the same manner as in the paralinguistic information estimation model learning device. The video feature extraction unit 11-3 outputs the extracted video features to the paralinguistic information estimation unit 21.
  • step S21 the paralinguistic information estimation unit 21 estimates the paralinguistic information of the utterance based on the prosodic feature, the linguistic feature, and the video feature extracted from the input utterance.
  • the learned paralinguistic information estimation model stored in the paralinguistic information estimation model storage unit 20 is used.
  • the paralinguistic information estimation model is a model based on deep learning, the paralinguistic information estimation result is obtained by forward propagating each feature amount.
  • each feature amount is input to the feature amount sub-model, and the feature amount gate rule is applied to the output result of each feature amount sub-model to obtain the feature amount gate weight vector, ),
  • the result of taking the element product of the feature amount gate weight vector and the output result of the feature amount sub-model is input to the result integrated sub-model to obtain the paralinguistic information estimation result.
  • the feature amount gate weight vector of a certain feature amount is determined from the output result of the feature amount sub-model of that feature amount. This is a configuration in which, for example, when it is determined that the characteristics of specific paralinguistic information strongly appear in the prosodic features, the prosody features are used for paralinguistic information estimation.
  • the feature amount gate weight vector of a certain feature amount is determined from the output results of the feature amount submodels of all feature amounts.
  • each feature amount sub-model 101 for example, the prosody feature sub-model 101-1
  • all feature amount weight calculation units 102 that is, The prosody feature weight calculator 102-1, the language feature weight calculator 102-2, and the video feature weight calculator 102-3) are input.
  • Each feature amount weight calculation unit 102 for example, the prosody feature weight calculation unit 102-1
  • the outputs of the sub-models 101-3 are compared to determine the feature amount gate weight vector of the feature amount (that is, the prosody feature gate weight vector).
  • the paralinguistic information estimation model learning device and the paralinguistic information estimation device of the second embodiment use the paralinguistic information estimation model shown in FIG. It is possible to learn and estimate paralinguistic information.
  • the program describing this processing content can be recorded in a computer-readable recording medium.
  • the computer-readable recording medium may be any recording medium such as a magnetic recording device, an optical disc, a magneto-optical recording medium, or a semiconductor memory.
  • distribution of this program is performed by selling, transferring, or lending a portable recording medium such as a DVD or a CD-ROM in which the program is recorded.
  • the program may be stored in a storage device of a server computer and transferred from the server computer to another computer via a network to distribute the program.
  • a computer that executes such a program first stores, for example, the program recorded in a portable recording medium or the program transferred from the server computer in its own storage device. Then, when executing the process, this computer reads the program stored in its own storage device and executes the process according to the read program.
  • a computer may directly read the program from a portable recording medium and execute processing according to the program, and the program is transferred from the server computer to this computer. Each time, the processing according to the received program may be sequentially executed.
  • the above-mentioned processing is executed by the so-called ASP (Application Service Provider) type service that realizes the processing function only by the execution instruction and result acquisition without transferring the program from the server computer to this computer. May be It should be noted that the program in this embodiment includes information that is used for processing by an electronic computer and that conforms to the program (data that is not a direct command to a computer but has the property of defining computer processing).
  • the device is configured by executing a predetermined program on a computer, but at least a part of the processing contents may be realized by hardware.
  • Paralinguistic information estimation model 101 Features Quantity sub model 102 feature quantity weight vector 103 feature quantity gate 104 result integrated sub model 121 feature quantity sub model learning unit 122 feature quantity weight calculation unit 123 feature quantity gate processing unit 124 result integrated sub model learning unit

Landscapes

  • Engineering & Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Health & Medical Sciences (AREA)
  • Theoretical Computer Science (AREA)
  • General Health & Medical Sciences (AREA)
  • Computational Linguistics (AREA)
  • General Physics & Mathematics (AREA)
  • Evolutionary Computation (AREA)
  • Artificial Intelligence (AREA)
  • Multimedia (AREA)
  • Human Computer Interaction (AREA)
  • Software Systems (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Computing Systems (AREA)
  • General Engineering & Computer Science (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Acoustics & Sound (AREA)
  • Mathematical Physics (AREA)
  • Data Mining & Analysis (AREA)
  • Biophysics (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Molecular Biology (AREA)
  • Biomedical Technology (AREA)
  • Psychiatry (AREA)
  • Child & Adolescent Psychology (AREA)
  • Hospice & Palliative Care (AREA)
  • Signal Processing (AREA)
  • Oral & Maxillofacial Surgery (AREA)
  • Medical Informatics (AREA)
  • Social Psychology (AREA)
  • Databases & Information Systems (AREA)
  • Machine Translation (AREA)
  • Image Analysis (AREA)

Abstract

パラ言語情報推定の精度を向上する。パラ言語情報推定モデル記憶部20は、複数の独立した特徴量を入力としてパラ言語情報推定結果を出力するパラ言語情報推定モデルを記憶する。特徴量抽出部11は、入力発話から特徴量を抽出する。パラ言語情報推定部20は、パラ言語情報推定モデルを用いて入力発話から抽出した特徴量から入力発話のパラ言語情報を推定する。パラ言語情報推定モデルは、特徴量ごとにその特徴量のみに基づいてパラ言語情報の推定に用いる情報を出力する特徴量サブモデルと、特徴量ごとに特徴量サブモデルの出力結果に基づいて特徴量重みを算出する特徴量重み算出部と、特徴量ごとに特徴量サブモデルの出力結果を特徴量重みで重み付けして出力する特徴量ゲートと、すべての特徴量ゲートの出力結果に基づいてパラ言語情報を推定する結果統合サブモデルと、を含む。

Description

パラ言語情報推定装置、パラ言語情報推定方法、およびプログラム
 この発明は、音声からパラ言語情報を推定する技術に関する。
 音声からパラ言語情報(例えば、発話意図が疑問か平叙か、感情が喜び・悲しみ・怒り・平静のいずれか)を推定する技術が求められている。パラ言語情報は、例えば、音声翻訳の高度化(例えば、「明日」という日本語の発話に対して、疑問意図「明日?」と理解して「Is it tomorrow?」と英語に翻訳したり、平叙意図「明日。」と理解して「It is tomorrow.」と英語に翻訳したりと、フランクな発話に対しても発話者の意図を正しく理解した日英翻訳ができる)や、音声対話における話し相手の感情を考慮した対話制御(例えば、相手が怒っていれば話題を変える)などに応用可能である。
 従来技術として、複数の独立した特徴量を用いたパラ言語情報推定技術が非特許文献1などに示されている。非特許文献1では、音声特徴(音声波形)と映像特徴(複数フレームの画像系列)に基づいて、話者の感情次元値(Valence(感情価):快-不快、Arousal(覚醒度):覚醒-睡眠、の二種)を推定する。また、音声の短時間ごとの声の高さなどの韻律特徴の時系列情報と、話した単語などの言語特徴の時系列情報とに基づいて、話者のパラ言語情報を推定する技術も知られている。これらの複数の特徴量を組み合わせる技術は、特徴量単体を利用する技術に比べて高い精度でパラ言語情報を認識できる。
 図1に、複数の独立した特徴量を用いたパラ言語情報推定モデルの従来技術を例示する。このパラ言語情報推定モデル900は、各特徴量からパラ言語情報を推定する特徴量サブモデル101と、それらの出力を統合して最終的なパラ言語情報推定結果を出力する結果統合サブモデル104とで構成される。この構成は、例えば発話意図推定においては、韻律特徴に疑問や平叙の特性が含まれるか(例えば、語尾が上がっているか否か)、言語特徴に疑問や平叙の特性が表れるか(例えば、疑問詞が含まれるか否か)を推定した後、それらの結果を統合して発話意図が疑問か平叙かを推定する処理に相当する。近年では、各サブモデルを深層学習に基づくモデルで構成し、パラ言語情報推定モデル全体を一体的に学習する、深層学習に基づくパラ言語情報推定モデルが主流となっている。
Panagiotis Tzirakis, George Trigeorgis, Mihalis A. Nicolaou, Bjorn W. Schuller, Stefanos Zafeiriou, "End-to-End Multimodal Emotion Recognition Using Deep Neural Networks," IEEE Journal of Selected Topics in Signal Processing, vol. 11, No. 8, pp. 1301-1309, 2017.
 パラ言語情報はすべての特徴量にその特性が表れるとは限らず、一部の特徴量だけにパラ言語情報の特性が表れることがある。例えば発話意図では、話し方は語尾上がりだが文章が平叙文である(すなわち、韻律特徴にのみ疑問発話の特性が表れる)発話が存在し、このような発話は疑問発話とみなされる。また、例えば感情では、表情からは平静にみえるが話し方や単語として怒りが強く表れている発話が存在し、このような発話は怒り感情発話とみなされる。
 しかしながら、従来技術では、一部の特徴量だけにパラ言語情報の特性が表れる発話を正しく学習することは困難である。これは、従来技術のパラ言語情報推定モデルでは、すべての特徴量が同じパラ言語情報の特性を示すかのようにモデル学習を行うためである。例えば、疑問発話の学習を行う場合、韻律特徴でも言語特徴でも疑問発話の特性が表れているかのように学習を行ってしまう。このため、韻律特徴にのみ疑問発話の特性が表れている発話でも、言語特徴にも疑問発話の特性が表れているとみなしてモデル学習をしてしまい、この発話は言語特徴における疑問発話の特性を正しく学習する上でのノイズとなる。その結果、従来技術において、一部の特徴量だけにパラ言語情報の特性が表れる発話が学習データに含まれると、パラ言語情報推定モデルを正しく学習することができず、パラ言語情報推定精度が低下する。
 この発明は、上記のような技術的課題を鑑みて、複数の独立した特徴量を用いたパラ言語情報推定において、一部の特徴量だけにパラ言語情報の特性が表れる発話が学習データに含まれる場合でも、正しくパラ言語情報推定モデルを学習し、正しくパラ言語情報を推定することを目的とする。
 上記の課題を解決するために、この発明の一態様のパラ言語情報推定装置は、入力発話からパラ言語情報を推定するパラ言語情報推定装置であって、複数の独立した特徴量を入力としてパラ言語情報推定結果を出力するパラ言語情報推定モデルを記憶するパラ言語情報推定モデル記憶部と、入力発話から複数の独立した特徴量を抽出する特徴量抽出部と、パラ言語情報推定モデルを用いて入力発話から抽出した複数の独立した特徴量から入力発話のパラ言語情報を推定するパラ言語情報推定部と、を含み、パラ言語情報推定モデルは、複数の独立した特徴量ごとにその特徴量のみに基づいてパラ言語情報の推定に用いる情報を出力する特徴量サブモデルと、複数の独立した特徴量ごとに特徴量サブモデルの出力結果に基づいてその特徴量をパラ言語情報の推定に用いるか否かを表す特徴量重みを算出する特徴量重み算出部と、複数の独立した特徴量ごとに特徴量サブモデルの出力結果を特徴量重みで重み付けして出力する特徴量ゲートと、すべての特徴量ゲートの出力結果に基づいてパラ言語情報を推定する結果統合サブモデルと、を含む。
 この発明によれば、複数の独立した特徴量を用いたパラ言語情報推定において、一部の特徴量だけにパラ言語情報の特性が表れる発話に対しても、正しくパラ言語情報推定モデルを学習し、正しくパラ言語情報を推定することができるようになる。その結果、パラ言語情報推定の精度が向上する。
図1は、従来のパラ言語情報推定モデルを例示する図である。 図2は、本発明のパラ言語情報推定モデルを例示する図である。 図3は、パラ言語情報推定モデル学習装置の機能構成を例示する図である。 図4は、パラ言語情報推定モデル学習方法の処理手順を例示する図である。 図5は、第一実施形態のパラ言語情報推定モデルを例示する図である。 図6は、パラ言語情報推定モデル学習部の機能構成を例示する図である。 図7は、パラ言語情報推定装置の機能構成を例示する図である。 図8は、パラ言語情報推定方法の処理手順を例示する図である。 図9は、第二実施形態のパラ言語情報推定モデルを例示する図である。
 以下、この発明の実施の形態について詳細に説明する。なお、図面中において同じ機能を有する構成部には同じ番号を付し、重複説明を省略する。
 本発明のポイントは、一部の特徴量だけにパラ言語情報の特性が表れる可能性を考慮し、各特徴量の情報をパラ言語情報推定に利用するかどうかを決定する特徴量ゲートを導入する点にある。一部の特徴量だけにパラ言語情報の特性が表れる発話に対してモデル学習を行うためには、特徴量ごとにパラ言語情報推定に利用するかどうかを選択できる機構を導入すればよいと考えられる。例えば、ある特徴量で特定のパラ言語情報の特性が強く表れている場合はその特徴量を利用してパラ言語情報推定を行うが、別の特徴量で特定のパラ言語情報の特性が表れていない場合は、その特徴量を利用せずにパラ言語情報推定を行う、といった機構を導入する。この選択機構を本発明では特徴量ゲートという形で実現する。
 図2に、本発明のパラ言語情報推定モデルの例を示す。このパラ言語情報推定モデル100は、従来と同様の特徴量サブモデル101と、特徴量サブモデル101の出力をパラ言語情報推定に利用するか否かを決定する特徴量ゲート103と、特徴量ゲート103の出力に基づいて最終的なパラ言語情報推定結果を出力する結果統合サブモデル104とで構成される。
 特徴量ゲート103は、各特徴量サブモデル101の出力を結果統合サブモデル104に入力するかどうかを決定する役割を持つ。特徴量ゲート103は、式(1)に基づいて出力を決定する。
Figure JPOXMLDOC01-appb-M000003
ここで、kは特徴量番号(k=1, 2, …)、ykは特徴量ゲート出力ベクトル、xkは特徴量ゲート入力ベクトル(特徴量サブモデル出力結果)、wkは特徴量ゲート重みベクトル、
Figure JPOXMLDOC01-appb-M000004
は、要素積を表す。特徴量ゲート重みベクトルwkが単位ベクトルのとき、特徴量サブモデル出力結果xkがそのまま結果統合サブモデル104へ出力される。特徴量ゲート重みベクトルwkがゼロベクトルのとき、特徴量サブモデル出力結果xkがゼロに変換されて結果統合サブモデル104へ出力される。このように、各特徴量に対応する特徴量ゲート重みベクトルwkを制御することで、ある特徴量を利用するが別の特徴量は利用しないというようにパラ言語情報推定モデルの学習やパラ言語情報の推定を行うことが可能となる。なお、深層学習に基づくパラ言語情報推定モデルの場合、特徴量ゲート重みベクトルwkもモデルパラメータの一つであるとみなせるため、特徴量ゲート重みベクトルwkも含めてモデル全体を一体学習することが可能である。
 具体的には、以下の手順によりパラ言語情報の推定を行う。
 1.複数の独立した特徴量を入力とし、特徴量ごとのサブモデル、特徴量ごとの特徴量ゲート、結果統合サブモデルから構成されるパラ言語情報推定モデルを用意する。
 2.パラ言語情報推定モデルの学習を行う。深層学習に基づくパラ言語情報推定モデルの場合、特徴量ゲートの重みベクトルを含めたモデル全体を誤差逆伝搬法により一体学習する。それ以外の場合では特徴量ゲートは学習できないため、特徴量ゲートの重みベクトルは人手によるルールで決定する。例えば、特徴量ごとのサブモデルの出力結果が識別平面からの距離の場合、識別平面からの距離の絶対値が0.5以下なら特徴量ゲートの重みベクトルはゼロベクトル、識別平面からの距離の絶対値が0.5より大きいなら特徴量ゲートの重みベクトルは単位ベクトルとする、というルールを定める。この場合、特徴量ごとのサブモデルを先に学習し、その後結果統合サブモデルを学習するという二段階の学習を行う。
 3.学習済みのパラ言語情報推定モデルに複数の独立した特徴量を入力し、発話ごとにパラ言語情報推定結果を得る。
 [第一実施形態]
 本実施形態において、入力発話とは、当該発話の音声波形情報および当該発話の話者の表情(顔)の映像情報の両方を指すものとする。本発明でパラ言語情報推定に用いる特徴量は、人間の発話から抽出できる独立した二以上の特徴量であればよいが、本実施形態では、韻律特徴、言語特徴、および映像特徴の互いに独立な三種類の特徴量を用いるものとする。ただし、これら三種類の特徴量のうち、いずれか二種類の特徴量のみを用いてもよい。また、他特徴量と互いに独立であれば、例えば生体信号情報(脈拍、皮膚電位など)などの情報を用いた特徴量を追加で利用してもよい。
 本実施形態では、特徴量ごとのサブモデルの出力結果として特徴量ごとのパラ言語情報確率を受け取ることもできるが、特徴量ごとのパラ言語情報確率の推定のために必要な中間情報(例えば、ディープニューラルネットワークにおける中間層の出力値)を受け取ることもできる。また、特徴量ゲートの重みベクトルも含めて学習を行う場合、重みベクトルはすべての入力に対して固定値ではなく、入力が変わるたびに動的に重みベクトルを変えることもできる。具体的には、式(2)または式(3)を用いて、入力から重みベクトルを算出することで、重みベクトルを動的に変化させる。
Figure JPOXMLDOC01-appb-M000005
ここで、kは特徴量番号(k=1, 2, …)、xkは特徴量ゲート入力ベクトル(特徴量サブモデル出力結果)、wkは特徴量ゲート重みベクトル、wxは特徴量ゲート重みベクトル算出用行列、bxは特徴量ゲート重みベクトル算出用バイアス、σは活性化関数(例えば、式(4)のシグモイド関数)を表す。wx, bxは予め学習により決定しておく。なお、式(4)においてxがベクトルの場合、ベクトルの各要素に対して式(4)を適用する。
Figure JPOXMLDOC01-appb-M000006
 上記のように構成することにより、入力発話の話者や発話環境に応じて特徴量ごとのサブモデルの出力結果の利用度合いを変える(例えば、抑揚にパラ言語情報が表れやすい人では韻律特徴を重視してパラ言語情報推定を行う、など)ことができる。そのため、一般的な特徴量ごとのパラ言語情報確率の重み付け和に基づく推定手法に比べて、より多様な入力に対しても高精度にパラ言語情報を推定することが可能となる。すなわち、多様な発話環境に対するパラ言語情報推定精度が向上する。
 <パラ言語情報推定モデル学習装置>
 第一実施形態のパラ言語情報推定モデル学習装置は、教師ラベルが付与された発話からパラ言語情報推定モデルを学習する。パラ言語情報推定モデル学習装置は、図3に例示するように、発話記憶部10-1、教師ラベル記憶部10-2、韻律特徴抽出部11-1、言語特徴抽出部11-2、映像特徴抽出部11-3、パラ言語情報推定モデル学習部12、およびパラ言語情報推定モデル記憶部20を備える。以下、韻律特徴抽出部11-1、言語特徴抽出部11-2、および映像特徴抽出部11-3を特徴量抽出部11と総称することもある。特徴量抽出部11はパラ言語情報推定に用いる特徴量の種類に応じて数や処理内容等の構成を変更する。このパラ言語情報推定モデル学習装置が、図4に例示する各ステップの処理を行うことにより第一実施形態のパラ言語情報推定モデル学習方法が実現される。
 パラ言語情報推定モデル学習装置は、例えば、中央演算処理装置(CPU: Central Processing Unit)、主記憶装置(RAM: Random Access Memory)などを有する公知又は専用のコンピュータに特別なプログラムが読み込まれて構成された特別な装置である。パラ言語情報推定モデル学習装置は、例えば、中央演算処理装置の制御のもとで各処理を実行する。パラ言語情報推定モデル学習装置に入力されたデータや各処理で得られたデータは、例えば、主記憶装置に格納され、主記憶装置に格納されたデータは必要に応じて中央演算処理装置へ読み出されて他の処理に利用される。パラ言語情報推定モデル学習装置の各処理部は、少なくとも一部が集積回路等のハードウェアによって構成されていてもよい。パラ言語情報推定モデル学習装置が備える各記憶部は、例えば、RAM(Random Access Memory)などの主記憶装置、ハードディスクや光ディスクもしくはフラッシュメモリ(Flash Memory)のような半導体メモリ素子により構成される補助記憶装置、またはリレーショナルデータベースやキーバリューストアなどのミドルウェアにより構成することができる。
 発話記憶部10-1には、パラ言語情報推定モデルの学習に用いる発話(以下、「学習発話」ともいう)が記憶されている。本実施形態では、発話は人間の発話音声を収録した音声波形情報と、その発話の話者の表情を収録した映像情報とからなるものとする。発話が具体的にどのような情報から構成されるかはパラ言語情報の推定にどのような特徴量を用いるかに応じて決定される。
 教師ラベル記憶部10-2には、発話記憶部10-1に記憶された各発話に付与されるパラ言語情報の正解値を表す教師ラベルが記憶されている。発話に対する教師ラベルの付与は、人手で行ってもよいし、周知のラベル分類技術を用いて行ってもよい。具体的にどのような教師ラベルを付与するかはパラ言語情報の推定にどのような特徴量を用いるかに応じて決定する。
 ステップS11-1において、韻律特徴抽出部11-1は、発話記憶部10-1に記憶された各発話の音声波形情報から韻律特徴を抽出する。韻律特徴は、基本周波数、短時間パワー、MFCC(Mel-frequency Cepstral Coefficients)、ゼロ交差率、Harmonics-to-Noise-Ratio(HNR)、メルフィルタバンク出力、のいずれか一つ以上の特徴量を含むベクトルである。また、これらの時間ごと(フレームごと)の系列ベクトルであってもよいし、これらの発話全体の統計量(平均、分散、最大値、最小値、勾配など)のベクトルであってもよい。韻律特徴抽出部11-1は、抽出した韻律特徴をパラ言語情報推定モデル学習部12へ出力する。
 ステップS11-2において、言語特徴抽出部11-2は、発話記憶部10-1に記憶された各発話の音声波形情報から言語特徴を抽出する。言語特徴の抽出には、音声認識技術により取得した単語列または音素認識技術により取得した音素列を利用する。言語特徴はこれらの単語列または音素列を系列ベクトルとして表現したものであってもよいし、発話全体での特定単語の出現数などを表すベクトルとしてもよい。言語特徴抽出部11-2は、抽出した言語特徴をパラ言語情報推定モデル学習部12へ出力する。
 ステップS11-3において、映像特徴抽出部11-3は、発話記憶部10-1に記憶された各発話の映像情報から映像特徴を抽出する。映像特徴は、各フレームでの顔の特徴点の位置座標、オプティカルフローから算出した小領域ごとの速度成分、局所的な画像勾配のヒストグラム(Histograms of Oriented Gradients: HOG)のいずれか一つ以上を含むベクトルである。また、これらの一定間隔の時間ごと(フレームごと)の系列ベクトルであってもよいし、これらの発話全体の統計量(平均、分散、最大値、最小値、勾配など)のベクトルであってもよい。映像特徴抽出部11-3は、抽出した映像特徴をパラ言語情報推定モデル学習部12へ出力する。
 ステップS12において、パラ言語情報推定モデル学習部12は、入力された韻律特徴、言語特徴、および映像特徴と、教師ラベル記憶部10-2に記憶された教師ラベルとを用いて、複数の独立した特徴量を入力としてパラ言語情報推定結果を出力するパラ言語情報推定モデルを学習する。パラ言語情報推定モデル学習部12は、学習済みのパラ言語情報推定モデルをパラ言語情報推定モデル記憶部20へ記憶する。
 図5に、本実施形態で利用するパラ言語情報推定モデルの構成例を示す。このパラ言語情報推定モデルは、韻律特徴サブモデル101-1、言語特徴サブモデル101-2、映像特徴サブモデル101-3、韻律特徴重み算出部102-1、言語特徴重み算出部102-2、映像特徴重み算出部102-3、韻律特徴ゲート103-1、言語特徴ゲート103-2、映像特徴ゲート103-3、および結果統合サブモデル104を備える。以下、韻律特徴サブモデル101-1、言語特徴サブモデル101-2、および映像特徴サブモデル101-3を特徴量サブモデル101と、韻律特徴重み算出部102-1、言語特徴重み算出部102-2、および映像特徴重み算出部102-3を特徴量重み算出部102と、韻律特徴ゲート103-1、言語特徴ゲート103-2、および映像特徴ゲート103-3を特徴量ゲート103と総称することもある。特徴量サブモデル101は、入力された特徴量のみに基づいてパラ言語情報の推定を行い、パラ言語推定結果もしくはパラ言語推定の際に生成される中間値(以下、「パラ言語情報の推定に用いる情報」ともいう)を出力する。特徴量重み算出部102は、特徴量サブモデル101の出力結果に基づいてその特徴量をパラ言語情報の推定に用いるか否かを表す特徴量ゲート重みベクトル(以下、「特徴量重み」ともいう)を算出する。特徴量ゲート103は、特徴量サブモデル101の出力結果を特徴量重み算出部102が出力する特徴量ゲート重みベクトルで重み付けして出力する。結果統合サブモデル104は、すべての特徴量ゲート103の出力結果に基づいてパラ言語情報を推定する。
 パラ言語情報推定モデルは、例えば深層学習に基づくDeep Neural Network(DNN)であってもよいし、Support Vector Machine(SVM)であってもよい。また、時間ごとの系列ベクトルを特徴量に用いる場合、Long Short-Term Memory Recurrent Neural Network(LSTM-RNN)などの系列を考慮できる推定モデルを用いてもよい。なお、パラ言語情報推定モデルがすべてDNNやLSTM-RNNを含む深層学習に基づく手法によって構成される場合、特徴量ゲートの重みベクトルも含めてモデル全体を単一のネットワーク(分類モデル)と見なすことができるため、パラ言語情報推定モデル全体を誤差逆伝搬法により一体学習することが可能である。
 パラ言語情報推定モデルが深層学習に基づく手法以外を含む場合(例えば各特徴量のサブモデルがSVMによって構成される場合)、特徴量ゲートの重みベクトルの数値や重みベクトルの決定規則は人手により与える必要がある。またこの場合、特徴量ごとのサブモデルや結果統合サブモデルは別々に学習する必要がある。このような場合でのパラ言語情報推定モデル学習部12の構成を図6に示す。この場合のパラ言語情報推定モデル学習部12は、韻律特徴サブモデル学習部121-1、言語特徴サブモデル学習部121-2、映像特徴サブモデル学習部121-3、韻律特徴重み算出部122-1、言語特徴重み算出部122-2、映像特徴重み算出部122-3、韻律特徴ゲート処理部123-1、言語特徴ゲート処理部123-2、映像特徴ゲート処理部123-3、および結果統合サブモデル学習部124を備える。
 韻律特徴サブモデル学習部121-1は、韻律特徴と教師ラベルとの組から、韻律特徴のみに基づいてパラ言語情報を推定する韻律特徴サブモデルを学習する。韻律特徴サブモデルは例えばSVMを用いるが、クラス分類が可能な他の機械学習手法を用いてもよい。また、韻律特徴サブモデルの出力結果とは、例えば韻律特徴サブモデルがSVMであれば識別平面からの距離を指す。
 言語特徴サブモデル学習部121-2および映像特徴サブモデル学習部121-3は、韻律特徴サブモデル学習部121-1と同様にして、言語特徴サブモデルおよび映像特徴サブモデルを学習する。
 韻律特徴重み算出部122-1は、特徴量ゲートルールを用いて、韻律特徴サブモデルの出力結果から韻律特徴ゲート重みベクトルを算出する。特徴量ゲートルールとは、特徴量ゲートを決定する規則と、特徴量ゲートの重みベクトルとの組を指す。韻律特徴サブモデルがSVMの例であれば、「韻律特徴サブモデルの出力結果において、識別平面からの距離の絶対値が0.5以下なら韻律特徴ゲート重みベクトルはゼロベクトル、識別平面からの距離の絶対値が0.5より大きいなら韻律特徴ゲート重みベクトルは単位ベクトル」といった、人手により与えたルールを指す。これは、SVMの識別平面からの距離が推定結果の尤もらしさであるとみなし、推定結果が尤もらしい(ある特徴量で特定のパラ言語情報の特性が強く表れている可能性が高い)場合は特徴量ゲート重みベクトルを単位ベクトルに、そうでない場合はゼロベクトルに設定する処理に等しい。この人手により与えたルールを韻律特徴サブモデルの出力結果に適用し、出力結果に対する韻律特徴ゲート重みベクトルを算出する。なお、韻律特徴ゲート重みベクトルの次元数は韻律特徴サブモデル出力結果と同じとする(SVMの例であれば1次元のベクトルとする)。
 言語特徴重み算出部122-2および映像特徴重み算出部122-3は、韻律特徴重み算出部122-1と同様にして、言語特徴重みベクトルおよび映像特徴重みベクトルを算出する。
 韻律特徴ゲート処理部123-1は、韻律特徴サブモデルの出力結果と、韻律特徴ゲート重みベクトルとを用いて、上記式(1)を計算し、韻律特徴ゲート出力ベクトルを求める。
 言語特徴ゲート処理部123-2および映像特徴ゲート処理部123-3は、韻律特徴ゲート処理部123-1と同様にして、言語特徴ゲート出力ベクトルおよび映像特徴ゲート出力ベクトルを算出する。
 結果統合サブモデル学習部124は、韻律特徴ゲート出力ベクトル、言語特徴ゲート出力ベクトル、映像特徴ゲート出力ベクトル、および教師ラベルの組から、結果統合サブモデルを学習する。結果統合サブモデルは例えばSVMを用いるが、クラス分類が可能な他の機械学習手法を用いてもよい。
 <パラ言語情報推定装置>
 第一実施形態のパラ言語情報推定装置は、学習済みのパラ言語情報推定モデルを用いて入力発話からパラ言語情報を推定する。パラ言語情報推定装置は、図7に例示するように、韻律特徴抽出部11-1、言語特徴抽出部11-2、映像特徴抽出部11-3、パラ言語情報推定モデル記憶部20、およびパラ言語情報推定部21を備える。このパラ言語情報推定装置が、図8に例示する各ステップの処理を行うことにより第一実施形態のパラ言語情報推定方法が実現される。
 パラ言語情報推定装置は、例えば、中央演算処理装置(CPU: Central Processing Unit)、主記憶装置(RAM: Random Access Memory)などを有する公知又は専用のコンピュータに特別なプログラムが読み込まれて構成された特別な装置である。パラ言語情報推定装置は、例えば、中央演算処理装置の制御のもとで各処理を実行する。パラ言語情報推定装置に入力されたデータや各処理で得られたデータは、例えば、主記憶装置に格納され、主記憶装置に格納されたデータは必要に応じて中央演算処理装置へ読み出されて他の処理に利用される。パラ言語情報推定装置の各処理部は、少なくとも一部が集積回路等のハードウェアによって構成されていてもよい。パラ言語情報推定装置が備える各記憶部は、例えば、RAM(Random Access Memory)などの主記憶装置、ハードディスクや光ディスクもしくはフラッシュメモリ(Flash Memory)のような半導体メモリ素子により構成される補助記憶装置、またはリレーショナルデータベースやキーバリューストアなどのミドルウェアにより構成することができる。
 ステップS11-1において、韻律特徴抽出部11-1は、入力発話の音声波形情報から韻律特徴を抽出する。韻律特徴の抽出は、パラ言語情報推定モデル学習装置と同様に行えばよい。韻律特徴抽出部11-1は、抽出した韻律特徴をパラ言語情報推定部21へ出力する。
 ステップS11-2において、言語特徴抽出部11-2は、入力発話の音声波形情報から言語特徴を抽出する。言語特徴の抽出は、パラ言語情報推定モデル学習装置と同様に行えばよい。言語特徴抽出部11-2は、抽出した言語特徴をパラ言語情報推定部21へ出力する。
 ステップS11-3において、映像特徴抽出部11-3は、入力発話の映像情報から映像特徴を抽出する。映像特徴の抽出は、パラ言語情報推定モデル学習装置と同様に行えばよい。映像特徴抽出部11-3は、抽出した映像特徴をパラ言語情報推定部21へ出力する。
 ステップS21において、パラ言語情報推定部21は、入力発話から抽出した韻律特徴、言語特徴、および映像特徴に基づいて、当該発話のパラ言語情報を推定する。推定にはパラ言語情報推定モデル記憶部20に記憶された学習済みのパラ言語情報推定モデルを用いる。パラ言語情報推定モデルが深層学習に基づくモデルである場合、各特徴量を順伝播することでパラ言語情報推定結果が得られる。深層学習に基づくモデルでない場合、各特徴量をそれぞれ特徴量サブモデルに入力し、各特徴量サブモデルの出力結果に特徴量ゲートルールを適用して特徴量ゲート重みベクトルを求め、上記式(1)に従って特徴量ゲート重みベクトルと特徴量サブモデルの出力結果との要素積を取った結果を結果統合サブモデルに入力することでパラ言語情報推定結果が得られる。
 [第二実施形態]
 第一実施形態では、ある特徴量の特徴量ゲート重みベクトルは、その特徴量の特徴量サブモデルの出力結果から決定している。これは、例えば韻律特徴において特定のパラ言語情報の特性が強く表れていると判断されたとき、韻律特徴をパラ言語情報推定に利用するという構成である。
 第二実施形態では、ある特徴量の特徴量ゲート重みベクトルは、すべての特徴量の特徴量サブモデルの出力結果から決定する。すべての特徴量の特徴量サブモデルの出力結果を考慮して特徴量ゲート重みベクトルを決定することで、どの特徴量の情報をパラ言語情報推定に利用すべきかを区別しやすくなり、各特徴量にわずかにパラ言語情報の特性が表れる発話に対してもパラ言語情報推定精度が向上する。例えば、韻律特徴でも言語特徴でも特定のパラ言語情報の特性がわずかに表れるような場合、韻律特徴と言語特徴の特性の現れ方を比較し、特性がより強く表れている方の特徴量をパラ言語情報推定に利用できるようになるためである。
 第二実施形態のパラ言語情報推定モデルは、図9に示すように、各特徴量サブモデル101(例えば、韻律特徴サブモデル101-1)の出力をすべての特徴量重み算出部102(すなわち、韻律特徴重み算出部102-1、言語特徴重み算出部102-2、および映像特徴重み算出部102-3)に入力するように構成する。各特徴量重み算出部102(例えば、韻律特徴重み算出部102-1)は、すべての特徴量サブモデル101(すなわち、韻律特徴サブモデル101-1、言語特徴サブモデル101-2、および映像特徴サブモデル101-3)の出力を比較して、その特徴量の特徴量ゲート重みベクトル(すなわち、韻律特徴ゲート重みベクトル)を決定する。
 第二実施形態のパラ言語情報推定モデル学習装置およびパラ言語情報推定装置は、図9に示すパラ言語情報推定モデルを用いることで、第一実施形態と同様の手順により、パラ言語情報推定モデルの学習やパラ言語情報の推定が可能である。
 以上、この発明の実施の形態について説明したが、具体的な構成は、これらの実施の形態に限られるものではなく、この発明の趣旨を逸脱しない範囲で適宜設計の変更等があっても、この発明に含まれることはいうまでもない。実施の形態において説明した各種の処理は、記載の順に従って時系列に実行されるのみならず、処理を実行する装置の処理能力あるいは必要に応じて並列的にあるいは個別に実行されてもよい。
 [プログラム、記録媒体]
 上記実施形態で説明した各装置における各種の処理機能をコンピュータによって実現する場合、各装置が有すべき機能の処理内容はプログラムによって記述される。そして、このプログラムをコンピュータで実行することにより、上記各装置における各種の処理機能がコンピュータ上で実現される。
 この処理内容を記述したプログラムは、コンピュータで読み取り可能な記録媒体に記録しておくことができる。コンピュータで読み取り可能な記録媒体としては、例えば、磁気記録装置、光ディスク、光磁気記録媒体、半導体メモリ等どのようなものでもよい。
 また、このプログラムの流通は、例えば、そのプログラムを記録したDVD、CD-ROM等の可搬型記録媒体を販売、譲渡、貸与等することによって行う。さらに、このプログラムをサーバコンピュータの記憶装置に格納しておき、ネットワークを介して、サーバコンピュータから他のコンピュータにそのプログラムを転送することにより、このプログラムを流通させる構成としてもよい。
 このようなプログラムを実行するコンピュータは、例えば、まず、可搬型記録媒体に記録されたプログラムもしくはサーバコンピュータから転送されたプログラムを、一旦、自己の記憶装置に格納する。そして、処理の実行時、このコンピュータは、自己の記憶装置に格納されたプログラムを読み取り、読み取ったプログラムに従った処理を実行する。また、このプログラムの別の実行形態として、コンピュータが可搬型記録媒体から直接プログラムを読み取り、そのプログラムに従った処理を実行することとしてもよく、さらに、このコンピュータにサーバコンピュータからプログラムが転送されるたびに、逐次、受け取ったプログラムに従った処理を実行することとしてもよい。また、サーバコンピュータから、このコンピュータへのプログラムの転送は行わず、その実行指示と結果取得のみによって処理機能を実現する、いわゆるASP(Application Service Provider)型のサービスによって、上述の処理を実行する構成としてもよい。なお、本形態におけるプログラムには、電子計算機による処理の用に供する情報であってプログラムに準ずるもの(コンピュータに対する直接の指令ではないがコンピュータの処理を規定する性質を有するデータ等)を含むものとする。
 また、この形態では、コンピュータ上で所定のプログラムを実行させることにより、本装置を構成することとしたが、これらの処理内容の少なくとも一部をハードウェア的に実現することとしてもよい。
10-1 発話記憶部
10-2 教師ラベル記憶部
11 特徴量抽出部
12 パラ言語情報推定モデル学習部
20 パラ言語情報推定モデル記憶部
21 パラ言語情報推定部
100,900 パラ言語情報推定モデル
101 特徴量サブモデル
102 特徴量重みベクトル
103 特徴量ゲート
104 結果統合サブモデル
121 特徴量サブモデル学習部
122 特徴量重み算出部
123 特徴量ゲート処理部
124 結果統合サブモデル学習部

Claims (8)

  1.  入力発話からパラ言語情報を推定するパラ言語情報推定装置であって、
     複数の独立した特徴量を入力としてパラ言語情報推定結果を出力するパラ言語情報推定モデルを記憶するパラ言語情報推定モデル記憶部と、
     入力発話から上記複数の独立した特徴量を抽出する特徴量抽出部と、
     上記パラ言語情報推定モデルを用いて上記入力発話から抽出した上記複数の独立した特徴量から上記入力発話のパラ言語情報を推定するパラ言語情報推定部と、
     を含み、
     上記パラ言語情報推定モデルは、
     上記複数の独立した特徴量ごとにその特徴量のみに基づいてパラ言語情報の推定に用いる情報を出力する特徴量サブモデルと、
     上記複数の独立した特徴量ごとに上記特徴量サブモデルの出力結果に基づいてその特徴量をパラ言語情報の推定に用いるか否かを表す特徴量重みを算出する特徴量重み算出部と、
     上記複数の独立した特徴量ごとに上記特徴量サブモデルの出力結果を上記特徴量重みで重み付けして出力する特徴量ゲートと、
     すべての上記特徴量ゲートの出力結果に基づいて上記パラ言語情報を推定する結果統合サブモデルと、
     を含むパラ言語情報推定装置。
  2.  請求項1に記載のパラ言語情報推定装置であって、
     上記特徴量重み算出部は、すべての上記特徴量の上記特徴量サブモデルの出力結果に基づいて上記特徴量重みを算出するものである、
     パラ言語情報推定装置。
  3.  請求項1または2に記載のパラ言語情報推定装置であって、
     上記特徴量重み算出部は、kを特徴量番号とし、xkを上記特徴量サブモデルの出力結果とし、wkを上記特徴量重みとし、wxをあらかじめ学習した行列とし、bxをあらかじめ学習したバイアスとし、σを活性化関数とし、
    Figure JPOXMLDOC01-appb-M000001

    または
    Figure JPOXMLDOC01-appb-M000002

    により上記特徴量重みを算出するものである、
     パラ言語情報推定装置。
  4.  請求項1から3のいずれかに記載のパラ言語情報推定装置であって、
     上記パラ言語情報推定モデルは、ニューラルネットワークに基づくモデルであり、
     上記特徴量重みは、固定値または入力に応じた関数であり、
     上記特徴量サブモデルと上記特徴量重みと上記結果統合サブモデルとは、複数の学習発話から抽出した上記複数の独立した特徴量と上記学習発話に付与された教師ラベルとを用いて一体で学習したものである、
     パラ言語情報推定装置。
  5.  請求項1から3のいずれかに記載のパラ言語情報推定装置であって、
     上記特徴量サブモデルは、複数の学習発話から抽出した上記複数の独立した特徴量と上記学習発話に付与された教師ラベルとから学習したものであり、
     上記特徴量重みは、上記特徴量ごとにあらかじめ定められたルールに従って算出されるものであり、
     上記結果統合サブモデルは、すべての上記特徴量ゲートの出力結果と上記教師ラベルとから学習したものである、
     パラ言語情報推定装置。
  6.  入力発話からパラ言語情報を推定するパラ言語情報推定方法であって、
     パラ言語情報推定モデル記憶部に、複数の独立した特徴量を入力としてパラ言語情報推定結果を出力するパラ言語情報推定モデルが記憶されており、
     特徴量抽出部が、入力発話から上記複数の独立した特徴量を抽出し、
     パラ言語情報推定部が、上記パラ言語情報推定モデルを用いて上記入力発話から抽出した上記複数の独立した特徴量から上記入力発話のパラ言語情報を推定し、
     上記パラ言語情報推定モデルは、
     上記複数の独立した特徴量ごとにその特徴量のみに基づいてパラ言語情報の推定に用いる情報を出力する特徴量サブモデルと、
     上記複数の独立した特徴量ごとに上記特徴量サブモデルの出力結果に基づいてその特徴量をパラ言語情報の推定に用いるか否かを表す特徴量重みを算出する特徴量重み算出部と、
     上記複数の独立した特徴量ごとに上記特徴量サブモデルの出力結果を上記特徴量重みで重み付けして出力する特徴量ゲートと、
     すべての上記特徴量ゲートの出力結果に基づいて上記パラ言語情報を推定する結果統合サブモデルと、
     を含むパラ言語情報推定方法。
  7.  請求項6に記載のパラ言語情報推定方法であって、
     上記パラ言語情報推定モデルは、ニューラルネットワークに基づくモデルであり、
     上記特徴量重みは、固定値または入力に応じた関数であり、
     上記特徴量サブモデルと上記特徴量重みと上記結果統合サブモデルとは、複数の学習発話から抽出した上記複数の独立した特徴量と上記学習発話に付与された教師ラベルとを用いて一体で学習したものである、
     パラ言語情報推定方法。
  8.  請求項1から5のいずれかに記載のパラ言語情報推定装置としてコンピュータを機能させるためのプログラム。
PCT/JP2019/039572 2018-10-22 2019-10-08 パラ言語情報推定装置、パラ言語情報推定方法、およびプログラム Ceased WO2020085070A1 (ja)

Priority Applications (1)

Application Number Priority Date Filing Date Title
US17/287,102 US11798578B2 (en) 2018-10-22 2019-10-08 Paralinguistic information estimation apparatus, paralinguistic information estimation method, and program

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
JP2018198427A JP6992725B2 (ja) 2018-10-22 2018-10-22 パラ言語情報推定装置、パラ言語情報推定方法、およびプログラム
JP2018-198427 2018-10-22

Publications (1)

Publication Number Publication Date
WO2020085070A1 true WO2020085070A1 (ja) 2020-04-30

Family

ID=70331153

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/JP2019/039572 Ceased WO2020085070A1 (ja) 2018-10-22 2019-10-08 パラ言語情報推定装置、パラ言語情報推定方法、およびプログラム

Country Status (3)

Country Link
US (1) US11798578B2 (ja)
JP (1) JP6992725B2 (ja)
WO (1) WO2020085070A1 (ja)

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPWO2022168297A1 (ja) * 2021-02-08 2022-08-11
WO2025099939A1 (ja) * 2023-11-10 2025-05-15 日本電信電話株式会社 学習装置、生成装置、学習方法、生成方法、及びプログラム

Families Citing this family (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2022176124A1 (ja) * 2021-02-18 2022-08-25 日本電信電話株式会社 学習装置、推定装置、それらの方法、およびプログラム
CN113380238A (zh) * 2021-06-09 2021-09-10 阿波罗智联(北京)科技有限公司 处理音频信号的方法、模型训练方法、装置、设备和介质

Citations (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2018146898A (ja) * 2017-03-08 2018-09-20 パナソニックIpマネジメント株式会社 装置、ロボット、方法、及びプログラム

Family Cites Families (10)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US10515629B2 (en) * 2016-04-11 2019-12-24 Sonde Health, Inc. System and method for activation of voice interactive services based on user state
US10135989B1 (en) * 2016-10-27 2018-11-20 Intuit Inc. Personalized support routing based on paralinguistic information
US10049664B1 (en) * 2016-10-27 2018-08-14 Intuit Inc. Determining application experience based on paralinguistic information
US10475530B2 (en) * 2016-11-10 2019-11-12 Sonde Health, Inc. System and method for activation and deactivation of cued health assessment
US20180032612A1 (en) * 2017-09-12 2018-02-01 Secrom LLC Audio-aided data collection and retrieval
JP7052866B2 (ja) * 2018-04-18 2022-04-12 日本電信電話株式会社 自己訓練データ選別装置、推定モデル学習装置、自己訓練データ選別方法、推定モデル学習方法、およびプログラム
US10872602B2 (en) * 2018-05-24 2020-12-22 Dolby Laboratories Licensing Corporation Training of acoustic models for far-field vocalization processing systems
JP7111017B2 (ja) * 2019-02-08 2022-08-02 日本電信電話株式会社 パラ言語情報推定モデル学習装置、パラ言語情報推定装置、およびプログラム
JP7332024B2 (ja) * 2020-02-21 2023-08-23 日本電信電話株式会社 認識装置、学習装置、それらの方法、およびプログラム
US12596956B2 (en) * 2020-04-08 2026-04-07 Sony Group Corporation Information processing device and information processing method for presenting reaction-adaptive explanation of automatic operations

Patent Citations (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2018146898A (ja) * 2017-03-08 2018-09-20 パナソニックIpマネジメント株式会社 装置、ロボット、方法、及びプログラム

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPWO2022168297A1 (ja) * 2021-02-08 2022-08-11
WO2025099939A1 (ja) * 2023-11-10 2025-05-15 日本電信電話株式会社 学習装置、生成装置、学習方法、生成方法、及びプログラム

Also Published As

Publication number Publication date
US11798578B2 (en) 2023-10-24
US20210398552A1 (en) 2021-12-23
JP6992725B2 (ja) 2022-01-13
JP2020067500A (ja) 2020-04-30

Similar Documents

Publication Publication Date Title
Sun End-to-end speech emotion recognition with gender information
Lozano-Diez et al. An analysis of the influence of deep neural network (DNN) topology in bottleneck feature based language recognition
Dahl et al. Context-dependent pre-trained deep neural networks for large-vocabulary speech recognition
US9058811B2 (en) Speech synthesis with fuzzy heteronym prediction using decision trees
WO2021104099A1 (zh) 一种基于情景感知的多模态抑郁症检测方法和系统
WO2019102884A1 (ja) ラベル生成装置、モデル学習装置、感情認識装置、それらの方法、プログラム、および記録媒体
US12211496B2 (en) Method and apparatus with utterance time estimation
JP7268711B2 (ja) 信号処理システム、信号処理装置、信号処理方法、およびプログラム
WO2020085070A1 (ja) パラ言語情報推定装置、パラ言語情報推定方法、およびプログラム
Tsuchiya et al. Speaker invariant feature extraction for zero-resource languages with adversarial learning
US20230095088A1 (en) Emotion recognition apparatus, emotion recognition model learning apparatus, methods and programs for the same
Song et al. MPSA-DenseNet: A novel deep learning model for English accent classification
Ahmed et al. CNN-based speech segments endpoints detection framework using short-time signal energy features
Ault et al. On speech recognition algorithms
Elbarougy Speech emotion recognition based on voiced emotion unit
Punithavathi et al. [Retracted] Empirical Investigation for Predicting Depression from Different Machine Learning Based Voice Recognition Techniques
Ivanko et al. An experimental analysis of different approaches to audio–visual speech recognition and lip-reading
JP2022147397A (ja) 感情分類器の訓練装置及び訓練方法
Vetráb et al. Aggregation strategies of Wav2vec 2.0 embeddings for computational paralinguistic tasks
JP6220733B2 (ja) 音声分類装置、音声分類方法、プログラム
Sefara et al. Gender identification in sepedi speech corpus
Higuchi et al. Speaker Adversarial Training of DPGMM-Based Feature Extractor for Zero-Resource Languages.
JP7111017B2 (ja) パラ言語情報推定モデル学習装置、パラ言語情報推定装置、およびプログラム
Maddali et al. Classification of disordered patient’s voice by using pervasive computational algorithms
Oruh et al. Deep learning with optimization techniques for the classification of spoken English digit

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 19875813

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 19875813

Country of ref document: EP

Kind code of ref document: A1