WO2020209072A1 - 対話行為推定装置、対話行為推定方法、対話行為推定モデル学習装置及びプログラム - Google Patents
対話行為推定装置、対話行為推定方法、対話行為推定モデル学習装置及びプログラム Download PDFInfo
- Publication number
- WO2020209072A1 WO2020209072A1 PCT/JP2020/013445 JP2020013445W WO2020209072A1 WO 2020209072 A1 WO2020209072 A1 WO 2020209072A1 JP 2020013445 W JP2020013445 W JP 2020013445W WO 2020209072 A1 WO2020209072 A1 WO 2020209072A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- utterance
- utterance sentence
- sentence
- feature amount
- dialogue action
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/30—Semantic analysis
- G06F40/35—Discourse or dialogue representation
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/90—Details of database functions independent of the retrieved data types
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/20—Natural language analysis
- G06F40/279—Recognition of textual entities
- G06F40/284—Lexical analysis, e.g. tokenisation or collocates
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/20—Natural language analysis
- G06F40/205—Parsing
- G06F40/216—Parsing using statistical methods
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/20—Natural language analysis
- G06F40/279—Recognition of textual entities
Definitions
- the present invention relates to a dialogue action estimation device, a dialogue action estimation method, a dialogue action estimation model learning device, and a program.
- Dialogue estimation is the estimation of the type of dialogue that indicates the intent of the utterance in the dialogue. For example, by correctly estimating the type of dialogue action "apology" for the utterance “I'm sorry", the dialogue action "acceptance” should be responded to the utterance "I'm sorry” of the user. It becomes possible to control.
- dialogue action types dialogue action system
- ISO24617-2 a dialogue action system called ISO24617-2
- a model for estimating the dialogue action learned in advance based on the supervised learning is used, and the user's utterance sentence is used as a feature amount at that time.
- the dialogue act immediately before the utterance sentence, the number of characters, the word n-gram, etc. are used (for example, Non-Patent Document 1).
- methods used for learning for example, support vector machine (SVM), conditional random field (CRF), logistic regression and the like have been reported.
- the response utterance generation in the dialogue system is generally generated by applying the response utterance generation logic for each estimated dialogue action type. From this point of view, it is desirable to be able to estimate the dialogue action system at the particle size corresponding to the utterance sentence generation logic to be responded.
- the present invention has been made in view of the above points, and an object of the present invention is to provide a dialogue action estimation device, a dialogue action estimation method, and a program capable of accurately estimating a dialogue action type in consideration of an utterance target. To do. Another object of the present invention is to provide a dialogue action estimation model learning device for accurately estimating a dialogue action type in consideration of an utterance target.
- the dialogue action estimation device receives an input of a first utterance sentence and a second utterance sentence which is a utterance sentence before the first utterance sentence including at least the utterance sentence immediately before the first utterance sentence. For each of the first utterance sentence and the second utterance sentence, a feature amount including the utterance target feature amount, which is a feature amount related to the utterance target of the utterance sentence, is extracted, and the extracted first utterance sentence and the first utterance sentence and the first utterance sentence are extracted.
- a feature amount extraction unit that aggregates the feature amounts for each of the spoken sentences into an aggregated feature amount, the aggregated feature amount, and the types of dialogue actions that have been learned in advance in consideration of the utterance target of the spoken sentence. It is configured to include a dialogue action estimation unit that estimates the dialogue action type of the first speech sentence by using a dialogue action estimation model for estimating the dialogue action type.
- the input unit is a second utterance sentence which is a utterance sentence before the first utterance sentence including the first utterance sentence and the utterance sentence at least immediately before the first utterance sentence.
- the feature amount extraction unit extracts and extracts the feature amount including the utterance target feature amount, which is the feature amount related to the utterance target of the utterance sentence, for each of the first utterance sentence and the second utterance sentence.
- the feature amounts for each of the first utterance sentence and the second utterance sentence are aggregated into an aggregated feature amount, and the dialogue action estimation unit uses the aggregated feature amount and a pre-learned utterance target of the utterance sentence.
- the dialogue action type of the first utterance sentence is estimated by using the dialogue action estimation model for estimating the dialogue action type indicating the type of dialogue action in consideration of.
- the input unit inputs the first utterance sentence and the second utterance sentence which is a utterance sentence before the first utterance sentence including at least the utterance sentence immediately before the first utterance sentence.
- the feature amount extraction unit extracts, for each of the first utterance sentence and the second utterance sentence, a feature amount including the utterance target feature amount, which is a feature amount related to the utterance target of the utterance sentence, and extracts the first
- the feature quantities for each of the one spoken sentence and the second spoken sentence were aggregated into an aggregated feature quantity, and the dialogue action estimation unit considered the aggregated feature quantity and the utterance target of the spoken sentence learned in advance.
- the input unit accepts the input of the first utterance sentence and the second utterance sentence which is the utterance sentence immediately before the first utterance sentence.
- the feature amount extraction unit extracts the utterance target feature amount, which is the feature amount related to the utterance target of the utterance sentence, for each of the first utterance sentence and the second utterance sentence, and extracts the first utterance sentence and the second utterance sentence.
- the utterance target feature amount for each is aggregated into an aggregated feature amount.
- the dialogue action estimation unit uses the aggregated feature quantity and the dialogue action estimation model for estimating the dialogue action type indicating the type of the dialogue action in consideration of the utterance target of the utterance sentence. Estimate the dialogue type of one utterance sentence.
- the characteristics regarding the utterance target of the utterance sentence As described above, for each of the first utterance sentence and the second utterance sentence which is the utterance sentence before the first utterance sentence including at least the utterance sentence immediately before the first utterance sentence, the characteristics regarding the utterance target of the utterance sentence.
- the feature quantity extraction unit of the dialogue action estimation device identifies the main utterance phrase, which is the phrase that most expresses the content of the utterance sentence, for each of the first utterance sentence and the second utterance sentence.
- the first utterance is based on the functional feature amount extraction unit for extracting the target feature amount and the utterance main verses for each of the first utterance sentence and the second utterance sentence specified by the utterance main utterance identification unit.
- the dialogue action estimation model learning device includes a first utterance sentence, a second utterance sentence which is a utterance sentence before the first utterance sentence including at least the utterance sentence immediately before the first utterance sentence, and a second utterance sentence.
- An input unit that accepts input of learning data including a dialogue action type indicating the type of dialogue action considering the utterance target of the first utterance sentence, and utterance sentences for each of the first utterance sentence and the second utterance sentence.
- the feature amount including the utterance target feature amount, which is the feature amount related to the utterance target, is extracted, and the feature amounts for each of the extracted first utterance sentence and the second utterance sentence are aggregated into an aggregated feature amount.
- the dialogue action type of the first utterance sentence estimated based on the dialogue action estimation model for estimation matches the dialogue action type of the first utterance sentence included in the learning data. It is configured to include a model learning unit that learns the parameters of the dialogue action estimation model.
- the second utterance sentence is a utterance sentence before the first utterance sentence including the first utterance sentence and the utterance sentence at least immediately before the first utterance sentence.
- the feature amount including the utterance target feature amount which is the feature amount related to the utterance target of the utterance sentence, is extracted, and the feature amounts for each of the extracted first utterance sentence and the second utterance sentence are aggregated.
- the parameters of the dialogue action estimation model are set so that the dialogue action type of the first speech sentence estimated based on the feature quantity and the dialogue action estimation model matches the dialogue action type of the first speech sentence included in the training data.
- the dialogue action estimation device the dialogue action estimation method, and the program of the present invention, it is possible to accurately estimate the dialogue action type in consideration of the utterance target. Further, according to the dialogue action estimation model learning device of the present invention, it is possible to learn a dialogue action estimation model for accurately estimating the dialogue action type in consideration of the utterance target.
- FIG. 1 is a block diagram showing a schematic configuration of a computer functioning as the dialogue action estimation model learning device 100 according to the embodiment of the present invention.
- FIG. 2 is a block diagram showing the configuration of the dialogue action estimation model learning device 100 according to the embodiment of the present invention.
- the dialogue action estimation model learning device 100 includes a CPU 11, a memory 12 such as a RAM, a communication interface (IF) unit 13, and an input unit 14 such as a keyboard.
- a computer including a display unit 15 such as a display, and a storage unit 16 such as a ROM that stores a program 17 for executing a dialogue action estimation model learning processing routine described later.
- the CPU 11, the memory 12, the communication IF unit 13, the input unit 14, the display unit 15, and the storage unit 16 are connected via the bus 10.
- the communication IF unit 13 can be connected to an external terminal by a communication line such as a LAN cable.
- the dialogue action estimation model learning device 100 has a dialogue action between the input unit 110, the text analysis unit 120, the feature amount extraction unit 130, and the model learning unit 140. It is configured to include an estimation model storage unit 150.
- the input unit 110 sets the second utterance sentence, which is a utterance sentence before the first utterance sentence, including the first utterance sentence and the utterance sentence immediately immediately before the first utterance sentence, and the utterance target of the first utterance sentence.
- Accepts input of learning data including a dialogue type indicating the type of dialogue that was considered.
- the learning data includes a history of utterance sentences and a dialogue action type of each utterance sentence, and the input unit 110 accepts input of a plurality of learning data.
- the history of the utterance sentence includes at least a pair consisting of the first utterance sentence which is the last utterance sentence and the second utterance sentence which is the previous utterance sentence, and the utterance sentence from the start of the dialogue act to the present time. And. However, if the first utterance sentence is the first utterance sentence at the start of the utterance, the second utterance sentence which is the previous utterance sentence becomes empty. As long as the pair is included, as a set of utterance sentences, a predetermined period or a predetermined number, for example, N utterance sentences from the latest utterance sentence may be used as the history of the utterance sentences. Further, the first utterance sentence and the second utterance sentence are utterance sentences in the dialogue system, the second utterance sentence is the utterance sentence of the system, and the first utterance sentence is the utterance sentence by the user's utterance.
- the first utterance sentence and the second utterance sentence need to have the dialogue action system itself as a system considering the utterance target.
- the system considering the utterance target is a system in which the conventional dialogue act is refined for each utterance target. For example, in a system that considers the utterance target, regarding the question of dialogue, Question: I is a question to the first person, Question: II is a question to the second person, and Question: III is a question to the third person. It is a system that is detailed as follows.
- the utterance target of the utterance sentence is classified into the first person I who is the speaker (user), the second person II who is the other party (system), and the third person III who is another person or thing. ..
- Question: I to III is a dialogue action type indicating the type of dialogue action in consideration of the utterance target of the utterance sentence.
- a system in which the utterance target is taken into consideration will be described as an example of the Question of the dialogue act.
- Example 2 As a concrete example of training data (Example 1) the second utterance sentence: “Hello, have you want to hear something?", The first utterance sentence: “. I'd like to hear about the services that are under contract now, but", and of the first utterance sentence Dialogue type: "Question: III”, (Example 2) the second utterance sentence: “Hello, have you want to hear something?", The first utterance sentence: "Nani is your name?”, Dialogue act type of the first utterance sentence: "Question: II " Can be mentioned.
- Example 1 since the utterance target of the first utterance sentence is a question about the third party "service", the dialogue action type "Question: III" indicating a question to the third party is learned as the correct answer. Given in the data. Further, in (Example 2), since the utterance target of the first utterance sentence is a question about "you" who is the second person, the dialogue action type "Question: II” indicating a question to the second person is the correct answer. Is given to the training data as.
- the input unit 110 transfers the first utterance sentence and the second utterance sentence included in the received learning data to the text analysis unit 120, and the dialogue action type of the first utterance sentence included in the learning data to the model learning unit 140. Pass each one.
- the text analysis unit 120 requests the morpheme information and the dependency information of the utterance sentence for each of the first utterance sentence and the second utterance sentence.
- the text analysis unit 120 obtains morphological information and dependency information for each of the first utterance sentence and the second utterance sentence by morphological analysis and dependency analysis, which are known techniques.
- the morpheme information is information related to morphemes such as part of speech and terminal form, and the morpheme information includes information of "phrase ID, destination clause ID / relationship type, head morpheme number / function word morpheme number".
- the following table shows an analysis example of the first utterance sentence "I would like to ask about the service you are currently subscribed to" in the above (Example 1).
- the text analysis unit 120 passes the morpheme information and the dependency information obtained for each of the first utterance sentence and the second utterance sentence to the feature amount extraction unit 130.
- the feature amount extraction unit 130 extracts the utterance target feature amount, which is the feature amount related to the utterance target of the utterance sentence, for each of the first utterance sentence and the second utterance sentence, and extracts the first utterance sentence and the second utterance sentence.
- the utterance target feature amount for each is aggregated into an aggregated feature amount.
- the feature amount extraction unit 130 includes the word n-gram extraction unit 131, the utterance main phrase identification unit 132, the functional feature amount extraction unit 133, and the utterance target feature amount extraction. It is configured to include a unit 134 and a feature amount aggregation unit 135.
- the word n-gram extraction unit 131 extracts n-grams for each of the first utterance sentence and the second utterance sentence.
- the word n-gram extraction unit 131 extracts the morpheme notation n-gram from the morpheme information and the dependency information for each of the first utterance sentence and the second utterance sentence obtained by the text analysis unit 120. Extract.
- the 5-gram of the first utterance sentence "I would like to ask about the service you are currently subscribed to" in the above (Example 1) is as follows.
- "BOS” and "EOS” are added to the beginning and end of the sentence, respectively.
- ⁇ 5-gram >> BOS-Now BOS-Now-Contract BOS-Now-Contract-BOS-Now-Contract-Now-Contract-Contract-I ... (Omitted) ...
- the word n-gram extraction unit 131 passes the extracted n-gram to the feature amount aggregation unit 135.
- the word n-gram extraction unit 131 may extract the n-gram by using a standard notation or a terminal form instead of the morpheme notation.
- the utterance main phrase identification unit 132 identifies the utterance main phrase, which is the phrase that most expresses the content of the utterance sentence, for each of the first utterance sentence and the second utterance sentence.
- the utterance main phrase identification unit 132 for each of the first utterance sentence and the second utterance sentence, the final phrase including the predicate of the main clause is the utterance main phrase.
- the utterance main clause specifying unit 132 sets the clause including the last independent word of the utterance sentence as the utterance main clause. For example, the utterance main clause specifying unit 132, utterance sentence "very much Hello" is to identify the "Hello" as the utterance main clause.
- the utterance main phrase specifying unit 132 passes the utterance main phrase for each of the specified first utterance sentence and the second utterance sentence to the functional feature amount extraction unit 133 and the utterance target feature amount extraction unit 134.
- the functional feature extraction unit 133 is a function that is a functional feature of the utterance sentence included in the utterance main phrase for each of the first utterance sentence and the second utterance sentence specified by the utterance main phrase identification unit 132. Extract the feature quantity.
- the functional feature amount extraction unit 133 describes the feature amount related to the function, such as the part of speech, tense, and modality of the words included in the main utterance clause of each utterance sentence for each of the first utterance sentence and the second utterance sentence. Is extracted. More specifically, the functional feature amount extraction unit 133 applies the following rules (1) to (3) to the main utterance clause, and collects the feature amounts extracted to obtain a functional feature amount. (1) When the part of speech of the head of the main utterance phrase is "adjective stem”, "verb stem”, “noun: action”, “noun: adjective", the corresponding part of speech is combined with "MPOS_" to form a feature quantity. ..
- the functional feature extraction unit 133 changes from “listening", which is the head of the main utterance phrase, to "listening".
- "MOD_WNT” is extracted as a feature from "MPOS_verb stem” and "tai”, and these features are collectively used as a functional feature.
- the functional feature amount extraction unit 133 also extracts the functional feature amount for the second utterance sentence. Then, the functional feature amount extraction unit 133 passes the functional feature amount for each of the extracted first utterance sentence and the second utterance sentence to the feature amount aggregation unit 135.
- the utterance target feature amount extraction unit 134 is based on each of the first utterance sentence and the second utterance sentence for each of the first utterance sentence and the second utterance sentence specified by the utterance main phrase identification unit 132, respectively. Extract the feature amount to be spoken.
- the utterance target feature amount extraction unit 134 includes case particles such as “ga”, “ha”, “mo”, “o”, “about”, and “to” related to the main utterance phrase, and continuous particles. Items with (hereinafter collectively referred to as case notation) are extracted, and features are generated by the following procedure.
- case notation Items with (hereinafter collectively referred to as case notation) are extracted, and features are generated by the following procedure.
- the term here refers to a content word related to a main utterance phrase accompanied by a case particle and a continuous particle.
- the utterance target feature amount extraction unit 134 passes the utterance target feature amount for each of the extracted first utterance sentences and the second utterance sentences to the feature amount aggregation unit 135.
- the feature amount aggregation unit 135 includes n-grams for each of the first utterance sentence and the second utterance sentence extracted by the word n-gram extraction unit 131, and the first feature amount extraction unit 133 extracted by the functional feature amount extraction unit 133.
- the functional feature amount for each of the utterance sentence and the second utterance sentence and the utterance target feature amount for each of the first utterance sentence and the second utterance sentence extracted by the utterance target feature amount extraction unit 134 are aggregated. It is an aggregate feature quantity.
- the feature amount aggregation unit 135 aggregates the word n-gram feature amount, the functional feature amount, and the utterance target feature amount into one feature amount. At that time, the feature amount aggregation unit 135 distinguishes each feature amount for the first utterance sentence and each feature amount for the second utterance sentence by giving labels such as "TARGET" and "PRE". If there are two or more previous utterance sentences in the utterance sentence history, they are distinguished by adding different labels such as "PRE2" and "PRE3". This is because the first utterance sentence and the second utterance sentence, which is the utterance sentence including the utterance sentence at least immediately before (one before) the first utterance sentence, are important in the embodiment of the present invention. A separate label is given to make it possible.
- the feature amount aggregating unit 135, "TARGET_BOS- now TARGET_BOS- now - contract ... PRE_BOS- Hello ... PRE_TARGET_verb stem ... TARGET_MPOS_verb stem TARGET_MOD_WNT TARGET_III_ About PRE_MOD_Q PRE_III_ is the aggregate feature quantity.
- the feature amount aggregation unit 135 is "TARGET_BOS-you TARGET_BOS-you-...
- PRE_masu-?-EOS TARGET_MOD_Q TARGET_II_ "PRE_MOD_Q PRE_III_ha” is used as the aggregate feature amount. Then, the feature amount aggregation unit 135 passes the aggregated feature amount to the model learning unit 140.
- the model learning unit 140 is the first utterance estimated based on the aggregated feature amount of the first utterance sentence and the second utterance sentence included in the learning data extracted by the feature amount extraction unit 130 and the dialogue action estimation model.
- the parameters of the dialogue action estimation model are learned so that the dialogue action type of the sentence matches the dialogue action type of the first utterance sentence included in the training data.
- the model learning unit 140 learns the dialogue action estimation model using the existing machine learning model.
- the case of learning using logistic regression will be described as an example, but a support vector machine (SVM), a conditional random field (CRF), or the like may be used.
- SVM support vector machine
- CRF conditional random field
- the model learning unit 140 correctly estimates the dialogue action in consideration of the speech target, that is, the dialogue action type estimated when the aggregated feature amount extracted by the feature amount extraction unit 130 is input to the dialogue action estimation model. And the parameters of the dialogue action estimation model are learned so that the dialogue action type of the first spoken sentence included in the learning data matches.
- the model learning unit 140 repeats the learning process until a predetermined end condition, for example, a condition such as a case where the learning process is repeated for a predetermined number of learning data, is satisfied. Then, the model learning unit 140 stores the parameters of the learned dialogue action estimation model in the dialogue action estimation model storage unit 150.
- a predetermined end condition for example, a condition such as a case where the learning process is repeated for a predetermined number of learning data.
- the dialogue action estimation model storage unit 150 stores the dialogue action estimation model and the parameters of the dialogue action estimation model learned by the model learning unit 140.
- FIG. 4 is a flowchart showing a dialogue action estimation model learning routine according to the embodiment of the present invention.
- the dialogue action estimation model learning processing routine shown in FIG. 4 is executed in the dialogue action estimation model learning device 100.
- step S100 the input unit 110 considers the first utterance sentence, the second utterance sentence which is the utterance sentence immediately before the first utterance sentence, and the type of dialogue action in consideration of the utterance target of the first utterance sentence. Accepts input of learning data including dialogue type indicating.
- step S110 the text analysis unit 120 requests the morpheme information and the dependency information of the utterance sentence for each of the first utterance sentence and the second utterance sentence.
- step S120 the word n-gram extraction unit 131 extracts n-grams for each of the first utterance sentence and the second utterance sentence input in step S110.
- step S130 the utterance main phrase specifying unit 132 specifies the utterance main phrase, which is the phrase that most expresses the content of the utterance sentence, for each of the first utterance sentence and the second utterance sentence input in step S110.
- the functional feature amount extraction unit 133 is a functional feature amount of the utterance sentence included in the utterance main bunsetsu for each of the first utterance sentence and the second utterance sentence specified in step S130. Extract functional features.
- step S150 the utterance target feature amount extraction unit 134 of the first utterance sentence and the second utterance sentence based on the utterance main bunsetsu for each of the first utterance sentence and the second utterance sentence specified in step S130. Each utterance target feature amount is extracted.
- the feature amount aggregation unit 135 includes n-grams for each of the first utterance sentence and the second utterance sentence extracted in step S120, and the first utterance sentence and the second utterance sentence extracted in step S140.
- the functional feature amount for each of the utterance sentences and the utterance target feature amount for each of the first utterance sentence and the second utterance sentence extracted in step S150 are aggregated into an aggregated feature amount.
- step S170 the model learning unit 140 is estimated based on the aggregated feature quantities of the first utterance sentence and the second utterance sentence included in the learning data extracted in step S160 and the dialogue action estimation model.
- the parameters of the dialogue action estimation model are learned so that the dialogue action type of the one-speech sentence matches the dialogue action type of the first speech sentence included in the learning data input in step S110.
- step S180 the model learning unit 140 determines whether or not the end condition is satisfied. If the end condition is not satisfied (NO in step S180), the process returns to step S100 and the processes of steps S100 to S180 are repeated. On the other hand, when the end condition is satisfied (YES in step S180), in step S190, the model learning unit 140 stores the parameters of the learned dialogue action estimation model in the dialogue action estimation model storage unit 150.
- the first utterance sentence and the utterance sentence immediately before the first utterance sentence are included before the first utterance sentence.
- the feature amount including the utterance target feature amount which is the feature amount related to the utterance target of the utterance sentence, is extracted, and for each of the extracted first utterance sentence and the second utterance sentence.
- Dialogue so that the dialogue action type of the first speech sentence estimated based on the aggregated feature quantity that aggregates the feature quantities and the dialogue action estimation model matches the dialogue action type of the first speech sentence included in the learning data.
- the dialogue action estimation device 200 includes a CPU 11, a memory 12 such as a RAM, a communication interface (IF) unit 13, an input unit 14 such as a keyboard, and a display. Etc. 15 and a storage unit 16 such as a ROM that stores a program 27 for executing a dialogue action estimation processing routine described later.
- the CPU 11, the memory 12, the communication IF unit 13, the input unit 14, the display unit 15, and the storage unit 16 are connected via the bus 10.
- the communication IF unit 13 can be connected to an external terminal by a communication line such as a LAN cable.
- the dialogue action estimation device 200 has a dialogue between the input unit 210, the text analysis unit 120, the feature amount extraction unit 130, and the dialogue action estimation model storage unit 150. It is configured to include an action estimation unit 260 and an output unit 270.
- the dialogue action estimation model storage unit 150 stores the dialogue action estimation model and the parameters of the dialogue action estimation model learned in advance by the dialogue action estimation model learning device 100.
- the input unit 210 accepts the input of the first utterance sentence and the second utterance sentence which is the utterance sentence before the first utterance sentence including the utterance sentence at least immediately before the first utterance sentence. Then, the input unit 210 passes the received first utterance sentence and the second utterance sentence to the text analysis unit 120.
- the dialogue action estimation unit 260 uses the aggregated feature quantity and the dialogue action estimation model for estimating the dialogue action type indicating the type of dialogue action in consideration of the utterance target of the utterance sentence, and the first Estimate the type of dialogue in the utterance.
- the dialogue action estimation unit 260 first acquires the dialogue action estimation model and the parameters of the dialogue action estimation model from the dialogue action estimation model storage unit 150. Next, the dialogue action estimation unit 260 estimates the dialogue action type of the first utterance sentence based on the aggregated feature amount extracted by the feature amount extraction unit 130 and the acquired dialogue action estimation model. Then, the dialogue action estimation unit 260 passes the estimated dialogue action type to the output unit 270.
- the output unit 270 outputs the dialogue action type estimated by the dialogue action estimation unit 260.
- FIG. 6 is a flowchart showing a dialogue action estimation processing routine according to the embodiment of the present invention. Note that the same processing as the dialogue action estimation model learning processing routine according to the embodiment of the present invention is designated by the same reference numerals and detailed description thereof will be omitted.
- step S200 the input unit 210 accepts the input of the first utterance sentence and the second utterance sentence which is the utterance sentence before the first utterance sentence including at least the utterance sentence immediately before the first utterance sentence.
- step S270 the dialogue action estimation unit 260 estimates the dialogue action type from the dialogue action estimation model storage unit 150, which indicates the type of dialogue action in consideration of the utterance target of the utterance sentence. Get the model and the parameters of the dialogue behavior estimation model.
- step S280 the dialogue action estimation unit 260 estimates the dialogue action type of the first utterance sentence by using the aggregated feature amount and the dialogue action estimation model acquired in step S270.
- step S290 the dialogue action type of the first utterance sentence estimated by step S280 is output.
- the dialogue action estimation device it is a utterance sentence before the first utterance sentence including the first utterance sentence and the utterance sentence at least immediately before the first utterance sentence.
- the feature amounts including the utterance target feature amount which is the feature amount related to the utterance target of the utterance sentence, are extracted, and the feature amounts for each of the extracted first utterance sentence and the second utterance sentence are aggregated.
- the dialogue action type of the first speech sentence is used by using the aggregated feature amount and the dialogue action estimation model for estimating the dialogue action type indicating the type of dialogue action considering the speech target of the speech sentence. By estimating, it is possible to accurately estimate the dialogue action type considering the speech target. Then, the dialogue system can appropriately select the response generation logic based on the dialogue action type estimated in this way, so that the dialogue accuracy of the entire dialogue system can be improved.
- the conventional dialogue action type since the aggregated feature amount includes n-gram, the conventional dialogue action type has a self-evident utterance target such as "greeting” or "Feedback". As for, the conventional system can be used as it is.
- the present invention is not limited to the above-described embodiment, and various modifications and applications are possible within a range that does not deviate from the gist of the present invention.
- program has been described as an embodiment in which the program is pre-installed in the specification of the present application, it is also possible to store the program in a computer-readable recording medium and provide the program.
- Dialogue action estimation model learning device 10 bus 11 CPU 12 Memory 13 Communication IF unit 14 Input unit 15 Display unit 16 Storage unit 17 Program 27 Program 100 Dialogue action estimation model learning device 110 Input unit 120 Text analysis unit 130 Feature quantity extraction unit 131 Word n-gram extraction unit 132 Speaking main phrase identification Unit 133 Functional feature amount extraction unit 134 Speaking target feature amount extraction unit 135 Feature amount aggregation unit 140 Model learning unit 150 Dialogue action estimation model storage unit 200 Dialogue action estimation device 210 Input unit 260 Dialogue action estimation unit 270 Output unit
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Physics & Mathematics (AREA)
- General Engineering & Computer Science (AREA)
- General Physics & Mathematics (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Computational Linguistics (AREA)
- General Health & Medical Sciences (AREA)
- Health & Medical Sciences (AREA)
- Artificial Intelligence (AREA)
- Databases & Information Systems (AREA)
- Data Mining & Analysis (AREA)
- Machine Translation (AREA)
- Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
Abstract
発話対象を考慮した対話行為タイプを精度よく推定することができるようにする。 特徴量抽出部130が、第1発話文と当該第1発話文の少なくとも直前の発話文を含む当該第1発話文より前の発話文である第2発話文との各々について、発話文の発話対象に関する特徴量である発話対象特徴量を含む特徴量を抽出し、対話行為推定部260が、抽出した第1発話文及び第2発話文の各々についての特徴量を集約した集約特徴量と、予め学習された、発話文の発話対象を考慮した対話行為の種類を示す対話行為タイプを推定するための対話行為推定モデルとを用いて、第1発話文の対話行為タイプを推定する。
Description
本発明は、対話行為推定装置、対話行為推定方法、対話行為推定モデル学習装置及びプログラムに関する。
従来から、対話システムがユーザの意図を理解して応答を生成するために重要な技術の一つである、対話行為推定が研究されている。対話行為推定とは、対話におけるその発話文の意図を示す対話行為のタイプを推定することである。例えば、「ごめんなさい」という発話文に対して「謝罪」という対話行為のタイプを正しく推定することで、ユーザの「ごめんなさい」という発話文に対して「謝罪受理」という対話行為の応答をすべき、という制御が可能となる。対話行為タイプのセット(対話行為体系)は、各々の研究で研究者が独自に開発したものが用いられることが多いが、最近ではISO24617-2という対話行為体系が提案されている。
また、従来の対話行為推定技術では、教師有り学習に基づいてあらかじめ学習した対話行為を推定するためのモデル(対話行為推定モデル)を使用しており、その際の特徴量として、ユーザの発話文を形態素解析し、発話文に含まれる形態素や発話文の直前の対話行為、文字数、単語n-gram等を用いている(例えば非特許文献1)。学習に用いる手法は、例えばサポートベクトルマシン(SVM)、条件付き確率場(CRF)、ロジスティック回帰等が報告されている。
福岡知隆,白井清昭,対話行為に固有の特徴を考慮した自由対話システムにおける対話行為推定,自然言語処理 Vol.24, No.4,2017.
対話システムにおける応答発話文の生成は、推定された対話行為タイプごとに応答発話文生成ロジックを適用する方法が一般的である。この観点から、応答すべき発話文生成ロジックに対応した粒度での対話行為体系が推定できることが望ましい。
しかしながら、従来の対話行為推定ではその粒度が対応していないという課題がある。例えば、ISO24617-2では「Question」という対話行為タイプが存在するが、当該対話行為タイプには「あなたの名前は?」のようにシステム(第2者)に関する発話文と、「首相の名前は?」のように第3者に関する発話文との両方が含まれる。前者は予め用意したシステムのパーソナルデータベースを検索して回答を生成し、後者は一般のインターネットにある情報を検索して回答を生成するという異なる生成ロジックが想定されるため、これら二つを区別することが必要であるが、従来の対話行為推定は「何について・誰について(以下、発話対象)」は考慮されていない、という問題があった。
本発明は上記の点に鑑みてなされたものであり、発話対象を考慮した対話行為タイプを精度よく推定することができる対話行為推定装置、対話行為推定方法、及びプログラムを提供することを目的とする。また、本発明は、発話対象を考慮した対話行為タイプを精度よく推定するための対話行為推定モデル学習装置を提供することを目的とする。
本発明に係る対話行為推定装置は、第1発話文と前記第1発話文の少なくとも直前の発話文を含む前記第1発話文より前の発話文である第2発話文との入力を受け付ける入力部と、前記第1発話文及び前記第2発話文の各々について、発話文の発話対象に関する特徴量である発話対象特徴量を含む特徴量を抽出し、抽出した前記第1発話文及び前記第2発話文の各々についての前記特徴量を集約して集約特徴量とする特徴量抽出部と、前記集約特徴量と、予め学習された、発話文の発話対象を考慮した対話行為の種類を示す対話行為タイプを推定するための対話行為推定モデルとを用いて、前記第1発話文の前記対話行為タイプを推定する対話行為推定部と、を備えて構成される。
また、本発明に係る対話行為推定方法は、入力部が、第1発話文と前記第1発話文の少なくとも直前の発話文を含む前記第1発話文より前の発話文である第2発話文との入力を受け付け、特徴量抽出部が、前記第1発話文及び前記第2発話文の各々について、発話文の発話対象に関する特徴量である発話対象特徴量を含む特徴量を抽出し、抽出した前記第1発話文及び前記第2発話文の各々についての前記特徴量を集約して集約特徴量とし、対話行為推定部が、前記集約特徴量と、予め学習された、発話文の発話対象を考慮した対話行為の種類を示す対話行為タイプを推定するための対話行為推定モデルとを用いて、前記第1発話文の前記対話行為タイプを推定する。
また、本発明に係るプログラムは、入力部が、第1発話文と前記第1発話文の少なくとも直前の発話文を含む前記第1発話文より前の発話文である第2発話文との入力を受け付け、特徴量抽出部が、前記第1発話文及び前記第2発話文の各々について、発話文の発話対象に関する特徴量である発話対象特徴量を含む特徴量を抽出し、抽出した前記第1発話文及び前記第2発話文の各々についての前記特徴量を集約して集約特徴量とし、対話行為推定部が、前記集約特徴量と、予め学習された、発話文の発話対象を考慮した対話行為の種類を示す対話行為タイプを推定するための対話行為推定モデルとを用いて、前記第1発話文の前記対話行為タイプを推定することを含む処理をコンピュータに実行させるためのプログラムである。
本発明に係る対話行為推定装置、対話行為推定方法及びプログラムによれば、入力部が、第1発話文と当該第1発話文の直前の発話文である第2発話文との入力を受け付け、特徴量抽出部が、第1発話文及び前記第2発話文の各々について、発話文の発話対象に関する特徴量である発話対象特徴量を抽出し、抽出した第1発話文及び第2発話文の各々についての発話対象特徴量を集約して集約特徴量とする。
そして、対話行為推定部が、集約特徴量と、予め学習された、発話文の発話対象を考慮した対話行為の種類を示す対話行為タイプを推定するための対話行為推定モデルとを用いて、第1発話文の対話行為タイプを推定する。
このように、第1発話文と当該第1発話文の少なくとも直前の発話文を含む当該第1発話文より前の発話文である第2発話文との各々について、発話文の発話対象に関する特徴量である発話対象特徴量を含む特徴量を抽出し、抽出した第1発話文及び第2発話文の各々についての特徴量を集約した集約特徴量と、予め学習された、発話文の発話対象を考慮した対話行為の種類を示す対話行為タイプを推定するための対話行為推定モデルとを用いて、第1発話文の対話行為タイプを推定することにより、発話対象を考慮した対話行為タイプを精度よく推定することができる。
また、本発明に係る対話行為推定装置の前記特徴量抽出部は、前記第1発話文と前記第2発話文との各々について、発話文の内容を最も表す文節である発話主要文節を特定する発話主要文節特定部と、前記発話主要文節特定部により特定された前記第1発話文及び前記第2発話文の各々についての発話主要文節に含まれる、発話文の機能的な特徴量である機能的特徴量を抽出する機能的特徴量抽出部と、前記発話主要文節特定部により特定された前記第1発話文及び前記第2発話文の各々についての発話主要文節に基づいて、前記第1発話文及び前記第2発話文の各々の前記発話対象特徴量を抽出する発話対象特徴量抽出部と、前記機能的特徴量抽出部により抽出された前記第1発話文及び前記第2発話文の各々についての前記機能的特徴量と、前記発話対象特徴量抽出部により抽出された前記第1発話文及び前記第2発話文の各々についての前記発話対象特徴量とを集約して前記集約特徴量とする特徴量集約部を含むことができる。
また、本発明に係る対話行為推定モデル学習装置は、第1発話文と前記第1発話文の少なくとも直前の発話文を含む前記第1発話文より前の発話文である第2発話文と、前記第1発話文の発話対象を考慮した対話行為の種類を示す対話行為タイプとを含む学習データの入力を受け付ける入力部と、前記第1発話文及び前記第2発話文の各々について、発話文の発話対象に関する特徴量である発話対象特徴量を含む特徴量を抽出し、抽出した前記第1発話文及び前記第2発話文の各々についての前記特徴量を集約して集約特徴量とする特徴量抽出部と、前記特徴量抽出部により抽出された前記第1発話文及び前記第2発話文についての集約特徴量と、発話文の発話対象を考慮した対話行為の種類を示す対話行為タイプを推定するための対話行為推定モデルとに基づいて推定される前記第1発話文の前記対話行為タイプが、前記学習データに含まれる前記第1発話文の前記対話行為タイプと一致するように、前記対話行為推定モデルのパラメータを学習するモデル学習部と、を備えて構成される。
このように、本発明に係る対話行為推定モデル学習装置によれば、第1発話文と当該第1発話文の少なくとも直前の発話文を含む当該第1発話文より前の発話文である第2発話文との各々について、発話文の発話対象に関する特徴量である発話対象特徴量を含む特徴量を抽出し、抽出した第1発話文及び第2発話文の各々についての特徴量を集約した集約特徴量と、対話行為推定モデルとに基づいて推定される第1発話文の対話行為タイプが、学習データに含まれる第1発話文の対話行為タイプと一致するように対話行為推定モデルのパラメータを学習することにより、発話対象を考慮した対話行為タイプを精度よく推定するための対話行為推定モデルを学習することができる。
本発明の対話行為推定装置、対話行為推定方法、及びプログラムによれば、発話対象を考慮した対話行為タイプを精度よく推定することができる。また、本発明の対話行為推定モデル学習装置によれば、発話対象を考慮した対話行為タイプを精度よく推定するための対話行為推定モデルを学習することができる。
<本発明の実施の形態に係る対話行為推定モデル学習装置の構成>
図1及び図2を参照して、本発明の実施の形態に係る対話行為推定モデル学習装置100の構成について説明する。図1は、本発明の実施の形態に係る対話行為推定モデル学習装置100として機能するコンピュータの概略構成を示すブロック図である。図2は、本発明の実施の形態に係る対話行為推定モデル学習装置100の構成を示すブロック図である。
図1及び図2を参照して、本発明の実施の形態に係る対話行為推定モデル学習装置100の構成について説明する。図1は、本発明の実施の形態に係る対話行為推定モデル学習装置100として機能するコンピュータの概略構成を示すブロック図である。図2は、本発明の実施の形態に係る対話行為推定モデル学習装置100の構成を示すブロック図である。
図1に示すように、本発明の実施の形態に係る対話行為推定モデル学習装置100は、CPU11と、RAM等のメモリ12と、通信インターフェース(IF)部13と、キーボード等の入力部14と、ディスプレイ等の表示部15と、後述する対話行為推定モデル学習処理ルーチンを実行するためのプログラム17を記憶したROM等の記憶部16とを備えたコンピュータで構成されている。また、CPU11、メモリ12、通信IF部13、入力部14、表示部15、及び記憶部16は、バス10を介して接続されている。また、通信IF部13は、LANケーブル等の通信回線により外部端末と接続することができる。
図2に示すように、本発明の実施の形態に係る対話行為推定モデル学習装置100は、入力部110と、テキスト解析部120と、特徴量抽出部130と、モデル学習部140と、対話行為推定モデル記憶部150とを備えて構成される。
入力部110は、第1発話文と当該第1発話文の少なくとも直前の発話文を含む当該第1発話文より前の発話文である第2発話文と、当該第1発話文の発話対象を考慮した対話行為の種類を示す対話行為タイプとを含む学習データの入力を受け付ける。具体的には、学習データには、発話文の履歴と、各発話文の対話行為タイプとが含まれており、入力部110は複数の学習データの入力を受け付ける。発話文の履歴には、最後の発話文である第1発話文と、その一つ前の発話文である第2発話文とからなる対を少なくとも含み、対話行為の開始から現時点までの発話文とする。ただし、第1発話文が発話開始の1発話目であった場合、その1つ前の発話文である第2発話文は空となる。当該対を含むものであれば、発話文の集合として、所定期間または所定数、例えば直近の発話文からN個の発話文を発話文の履歴として用いるようにしてもよい。また、第1発話文と第2発話文とは、対話システムにおける発話文であり、第2発話文がシステムの発話、第1発話文がユーザの発話による発話文である。
発話対象を考慮した対話行為推定を実現するためには、第1発話文と第2発話文とは、その対話行為の体系自体が、発話対象を考慮した体系となっている必要がある。発話対象を考慮した体系とは、従来の対話行為が、発話対象毎に詳細化されている体系である。例えば、発話対象を考慮した体系は、対話行為のQuestionについて、Question:Iは第1者への質問、Question:IIは第2者への質問、Question:IIIは第3者への質問、というように詳細化されている体系である。すなわち、発話文の発話対象を、話者(ユーザ)である第1者I、話相手(システム)である第2者II、それ以外の人や物である第3者IIIに分類すると定義する。ここで、Question:I~IIIは、発話文の発話対象を考慮した対話行為の種類を示す対話行為タイプとする。以下、本実施の形態では、上記対話行為のQuestionについて発話対象を考慮した体系を例に説明する。
学習データの具体例として、
(例1)第2発話文:「こんにちは、何か聞きたいことはありますか?」、第1発話文:「今契約しているサービスについて聞きたいのですが。」、及び第1発話文の対話行為タイプ:「Question:III」、
(例2)第2発話文:「こんにちは、何か聞きたいことはありますか?」、第1発話文:「あなたの名前はなあに?」、第1発話文の対話行為タイプ:「Question:II」
が挙げられる。
(例1)第2発話文:「こんにちは、何か聞きたいことはありますか?」、第1発話文:「今契約しているサービスについて聞きたいのですが。」、及び第1発話文の対話行為タイプ:「Question:III」、
(例2)第2発話文:「こんにちは、何か聞きたいことはありますか?」、第1発話文:「あなたの名前はなあに?」、第1発話文の対話行為タイプ:「Question:II」
が挙げられる。
(例1)では、第1発話文の発話対象は、第3者である「サービス」についてのQuestionであるから、第3者への質問を示す対話行為タイプ「Question:III」が正解として学習データに与えられている。また、(例2)では、第1発話文の発話対象は、第2者である「あなた」についてのQuestionであるから、第2者への質問を示す対話行為タイプ「Question:II」が正解として学習データに与えられている。
そして、入力部110は、受け付けた学習データに含まれる第1発話文及び第2発話文をテキスト解析部120に、当該学習データに含まれる第1発話文の対話行為タイプをモデル学習部140にそれぞれ渡す。
テキスト解析部120は、第1発話文及び第2発話文の各々について、発話文の形態素情報及び係り受け情報を求める。
具体的には、テキスト解析部120は、第1発話文及び第2発話文の各々について、既知の技術である形態素解析、係り受け解析により、形態素情報及び係り受け情報を求める。形態素情報は、品詞、終止形等の形態素に関する情報であり、文節情報は「文節ID、係り先文節ID/係りタイプ、主辞形態素番号/機能語形態素番号」の情報を含む。上記(例1)の第1発話文「今契約しているサービスについて聞きたいのですが」の解析例を下記表に示す。
そして、テキスト解析部120は、第1発話文及び第2発話文の各々について求めた形態素情報及び係り受け情報を、特徴量抽出部130に渡す。
特徴量抽出部130は、第1発話文及び第2発話文の各々について、発話文の発話対象に関する特徴量である発話対象特徴量を抽出し、抽出した第1発話文及び第2発話文の各々についての発話対象特徴量を集約して集約特徴量とする。
具体的には、図3に示すように、特徴量抽出部130は、単語n-gram抽出部131と、発話主要文節特定部132と、機能的特徴量抽出部133と、発話対象特徴量抽出部134と、特徴量集約部135とを備えて構成される。
単語n-gram抽出部131は、第1発話文と第2発話文との各々についてのn-gramを抽出する。
具体的には、単語n-gram抽出部131は、テキスト解析部120により求められた第1発話文及び第2発話文の各々についての形態素情報及び係り受け情報から、形態素表記のn-gramを抽出する。例えば上記(例1)の第1発話文「今契約しているサービスについて聞きたいのですが」の5-gramは、以下のようになる。なお、文頭と文末にはそれぞれ「BOS」、「EOS」を付与する。
<<5-gram>>
BOS-今
BOS-今-契約
BOS-今-契約-し
BOS-今-契約-し-て
今-契約-し-て-い
…(中略)…
た-い-の-です-が
い-の-です-が-EOS
の-です-が-EOS
です-が-EOS
<<5-gram>>
BOS-今
BOS-今-契約
BOS-今-契約-し
BOS-今-契約-し-て
今-契約-し-て-い
…(中略)…
た-い-の-です-が
い-の-です-が-EOS
の-です-が-EOS
です-が-EOS
そして、単語n-gram抽出部131は、抽出したn-gramを特徴量集約部135に渡す。なお、単語n-gram抽出部131は、形態素表記の代わりに標準表記や終止形を使用してn-gramを抽出してもよい。
発話主要文節特定部132は、第1発話文と第2発話文との各々について、発話文の内容を最も表す文節である発話主要文節を特定する。
具体的には、発話主要文節特定部132は、第1発話文及び第2発話文の各々について、主節の述語が含まれる最終文節が発話主要文節とする。発話主要文節特定部132は、主節の述語が存在しない場合(例えば独立詞等)、発話文の最後の独立詞等が含まれる文節を発話主要文節とする。例えば、発話主要文節特定部132は、「どうもこんにちは」という発話文については、「こんにちは」を発話主要文節として特定する。
そして、発話主要文節特定部132は、特定した第1発話文及び第2発話文の各々についての発話主要文節を、機能的特徴量抽出部133及び発話対象特徴量抽出部134に渡す。
機能的特徴量抽出部133は、発話主要文節特定部132により特定された第1発話文及び第2発話文の各々についての発話主要文節に含まれる、発話文の機能的な特徴量である機能的特徴量を抽出する。
具体的には、機能的特徴量抽出部133は、第1発話文及び第2発話文の各々について、各発話文の発話主要文節に含まれる語の品詞、テンス、モダリティ等、機能に関する特徴量を抽出する。より具体的には、機能的特徴量抽出部133は、下記(1)から(3)の規則を発話主要文節に適用して抽出された特徴量をまとめて、機能的特徴量とする。
(1)発話主要文節の主辞の品詞が「形容詞語幹」、「動詞語幹」、「名詞:動作」、「名詞:形容」の場合、該当する品詞を「MPOS_」と結合して特徴量とする。
(2)発話文がただ一つの文節しかもたない場合、「ONLY」を特徴量とする。
(3)発話主要文節の主辞より後に出現する機能語を抽出し、下記(3-A)、(3-B)に該当する情報があればテンス情報(過去)、モダリティ情報(願望・意志・命令・禁止・疑問等)の特徴量として抽出する。
(3-A)テンス情報の抽出
述語の後ろに品詞に「接尾辞:終止」を含む形態素表記「た」が存在する場合、「PAST_T」を出力する。
(3-B)モダリティ情報の抽出
・『願望』:述語の後ろに、終止形が「たい」となる形態素が存在すれば「MOD_WNT」を出力する。
・『命令』:動詞が「しろ」、「帰れ」のような命令形であれば「MOD_IMP」を出力する。
・『禁止』:述語が動詞の基本形で、その直後に「な」が存在すれば「MOD_FBD」を出力する。
・『疑問』:文節の末尾形態素が「?」もしくは疑問を表す終助詞「か」、疑問詞「何」「どこ」「誰」等の場合、「MOD_Q」を出力する。
・『依頼』:述語が動詞で、直後の形態素表記が「て」の場合、下記リストに含まれるいずれかの表記が後続するか、又は後続する表記が何も存在しない場合は「MOD_REQ」を出力する。
[リスト]:「くれ」、「ください」、「いただく」、「ちょうだい」、「もらう」、「ほしい」、「もらいたい」
(1)発話主要文節の主辞の品詞が「形容詞語幹」、「動詞語幹」、「名詞:動作」、「名詞:形容」の場合、該当する品詞を「MPOS_」と結合して特徴量とする。
(2)発話文がただ一つの文節しかもたない場合、「ONLY」を特徴量とする。
(3)発話主要文節の主辞より後に出現する機能語を抽出し、下記(3-A)、(3-B)に該当する情報があればテンス情報(過去)、モダリティ情報(願望・意志・命令・禁止・疑問等)の特徴量として抽出する。
(3-A)テンス情報の抽出
述語の後ろに品詞に「接尾辞:終止」を含む形態素表記「た」が存在する場合、「PAST_T」を出力する。
(3-B)モダリティ情報の抽出
・『願望』:述語の後ろに、終止形が「たい」となる形態素が存在すれば「MOD_WNT」を出力する。
・『命令』:動詞が「しろ」、「帰れ」のような命令形であれば「MOD_IMP」を出力する。
・『禁止』:述語が動詞の基本形で、その直後に「な」が存在すれば「MOD_FBD」を出力する。
・『疑問』:文節の末尾形態素が「?」もしくは疑問を表す終助詞「か」、疑問詞「何」「どこ」「誰」等の場合、「MOD_Q」を出力する。
・『依頼』:述語が動詞で、直後の形態素表記が「て」の場合、下記リストに含まれるいずれかの表記が後続するか、又は後続する表記が何も存在しない場合は「MOD_REQ」を出力する。
[リスト]:「くれ」、「ください」、「いただく」、「ちょうだい」、「もらう」、「ほしい」、「もらいたい」
例えば、上記(例1)の第1発話文「今契約しているサービスについて聞きたいのですが」の場合、機能的特徴量抽出部133は、発話主要文節の主辞である「聞く」から「MPOS_動詞語幹」、「たい」から「MOD_WNT」を特徴量として抽出し、これらの特徴量をまとめて機能的特徴量とする。機能的特徴量抽出部133は、第2発話文についても同様に機能的特徴量を抽出する。そして、機能的特徴量抽出部133は、抽出した第1発話文及び第2発話文の各々についての機能的特徴量を、特徴量集約部135に渡す。
発話対象特徴量抽出部134は、発話主要文節特定部132により特定された第1発話文及び第2発話文の各々についての発話主要文節に基づいて、第1発話文及び第2発話文の各々の発話対象特徴量を抽出する。
具体的には、発話対象特徴量抽出部134は、発話主要文節に係る「が」、「は」、「も」、「を」、「について」、「という」等の格助詞や、連用助詞(以下、まとめて格表記という)を伴う項を抽出し、以下の手順で特徴量を生成する。なお、ここでの項は、格助詞や連用助詞を伴って発話主要文節に係る内容語を指す。
<<手順>>
格表記の前に出現する名詞相当(品詞が名詞、もしくは未知語)の連続を項の表記として抽出し、以下の(A)~(E)の処理を実施する。
(A)項の表記が「あなた」「お前」「てめえ」「あんた」等の第2者を表す場合、「II_格表記」を発話対象特徴量とする。なお、「格表記」は、該当する表記に置き換えられる。
(B)項の表記が「わたし」「私」「俺」「オレ」等の第1者を表す場合、「I_格表記」を発話対象特徴量とする。
(C)項の表記が上記以外の場合、対象の項に「の」を伴って係る項がある場合、その項について上記(A)(B)を適用する。適用されない場合は「III_格表記」を発話対象特徴量として抽出する。例えば、例1:「サービスについて」→「III_について」、例2:「あなたの名前」→「II_の」とする。
(D)項の表記が存在せず、かつ、発話が対話の先頭(直前に発話が存在しない)の場合、「II_ELM」を発話対象特徴量として抽出する。
(E)項の表記が存在せず、かつ、上記(D)以外の場合、「SBJ_UNK」を発話対象特徴量とする。
格表記の前に出現する名詞相当(品詞が名詞、もしくは未知語)の連続を項の表記として抽出し、以下の(A)~(E)の処理を実施する。
(A)項の表記が「あなた」「お前」「てめえ」「あんた」等の第2者を表す場合、「II_格表記」を発話対象特徴量とする。なお、「格表記」は、該当する表記に置き換えられる。
(B)項の表記が「わたし」「私」「俺」「オレ」等の第1者を表す場合、「I_格表記」を発話対象特徴量とする。
(C)項の表記が上記以外の場合、対象の項に「の」を伴って係る項がある場合、その項について上記(A)(B)を適用する。適用されない場合は「III_格表記」を発話対象特徴量として抽出する。例えば、例1:「サービスについて」→「III_について」、例2:「あなたの名前」→「II_の」とする。
(D)項の表記が存在せず、かつ、発話が対話の先頭(直前に発話が存在しない)の場合、「II_ELM」を発話対象特徴量として抽出する。
(E)項の表記が存在せず、かつ、上記(D)以外の場合、「SBJ_UNK」を発話対象特徴量とする。
そして、発話対象特徴量抽出部134は、抽出した第1発話文及び第2発話文の各々についての発話対象特徴量を、特徴量集約部135に渡す。
特徴量集約部135は、単語n-gram抽出部131により抽出された第1発話文と第2発話文との各々についてのn-gramと、機能的特徴量抽出部133により抽出された第1発話文及び第2発話文の各々についての機能的特徴量と、発話対象特徴量抽出部134により抽出された第1発話文及び第2発話文の各々についての発話対象特徴量とを集約して集約特徴量とする。
具体的には、特徴量集約部135は、単語n-gram特徴量、機能的特徴量、発話対象特徴量を集約して一つの特徴量とする。その際、特徴量集約部135は、第1発話文についての各特徴量と第2発話文についての各特徴量とは、「TARGET」、「PRE」等のラベルを付与することで区別する。なお、発話文の履歴に、二つ以上前の発話文がある場合には、「PRE2」、「PRE3」等の別ラベルを付与することで区別する。これは、第1発話文と当該第1発話文の少なくとも直前(1つ前)の発話文を含む発話文である第2発話文が本発明の実施の形態において重要であるため、それらを区別可能にするために別ラベルを付与するものである。
例えば、上記(例1)の第1発話文「今契約しているサービスについて聞きたいのですが」の場合、特徴量集約部135は、「TARGET_BOS-今 TARGET_BOS-今-契約…PRE_BOS-こんにちは…PRE_TARGET_動詞語幹…TARGET_MPOS_動詞語幹 TARGET_MOD_WNT TARGET_III_について PRE_MOD_Q PRE_III_は」を集約特徴量とする。同様に、上記(例2)の第1発話文「あなたの名前はなあに?」の場合、特徴量集約部135は「TARGET_BOS-あなた TARGET_BOS-あなた-の…PRE_ます-か-?-EOS TARGET_MOD_Q TARGET_II_の PRE_MOD_Q PRE_III_は」を集約特徴量とする。そして、特徴量集約部135は、集約特徴量をモデル学習部140に渡す。
モデル学習部140は、特徴量抽出部130により抽出された学習データに含まれる第1発話文及び第2発話文についての集約特徴量と、対話行為推定モデルとに基づいて推定される第1発話文の対話行為タイプが、学習データに含まれる第1発話文の対話行為タイプと一致するように対話行為推定モデルのパラメータを学習する。
具体的には、モデル学習部140は、既存の機械学習モデルを用いて対話行為推定モデルを学習する。本実施の形態では、ロジスティック回帰を用いて学習する場合を例に説明するが、サポートベクトルマシン(SVM)、条件付き確率場(CRF)等を用いてもよい。モデル学習部140は、発話対象を考慮した対話行為を正しく推定するように、すなわち、特徴量抽出部130により抽出された集約特徴量を対話行為推定モデルに入力した場合に推定される対話行為タイプと、学習データに含まれる第1発話文の対話行為タイプとが一致するように、対話行為推定モデルのパラメータを学習する。モデル学習部140は、所定の終了条件、例えば所定数の学習データについて学習処理を繰り返した場合等の条件を満たすまで、学習処理を繰り返す。そして、モデル学習部140は、学習した対話行為推定モデルのパラメータを、対話行為推定モデル記憶部150に格納する。
対話行為推定モデル記憶部150には、対話行為推定モデルとモデル学習部140により学習された対話行為推定モデルのパラメータとが格納されている。
<本発明の実施の形態に係る対話行為推定モデル学習装置の作用>
図4は、本発明の実施の形態に係る対話行為推定モデル学習ルーチンを示すフローチャートである。入力部110に学習データが入力されると、対話行為推定モデル学習装置100おいて、図4に示す対話行為推定モデル学習処理ルーチンが実行される。
図4は、本発明の実施の形態に係る対話行為推定モデル学習ルーチンを示すフローチャートである。入力部110に学習データが入力されると、対話行為推定モデル学習装置100おいて、図4に示す対話行為推定モデル学習処理ルーチンが実行される。
まず、ステップS100において、入力部110は、第1発話文と、当該第1発話文の直前の発話文である第2発話文と、当該第1発話文の発話対象を考慮した対話行為の種類を示す対話行為タイプとを含む学習データの入力を受け付ける。
ステップS110において、テキスト解析部120は、第1発話文及び第2発話文の各々について、発話文の形態素情報及び係り受け情報を求める。
ステップS120において、単語n-gram抽出部131は、上記ステップS110により入力された第1発話文と第2発話文との各々についてのn-gramを抽出する。
ステップS130において、発話主要文節特定部132は、上記ステップS110により入力された第1発話文と第2発話文との各々について、発話文の内容を最も表す文節である発話主要文節を特定する。
ステップS140において、機能的特徴量抽出部133は、上記ステップS130により特定された第1発話文及び第2発話文の各々についての発話主要文節に含まれる、発話文の機能的な特徴量である機能的特徴量を抽出する。
ステップS150において、発話対象特徴量抽出部134は、上記ステップS130により特定された第1発話文及び第2発話文の各々についての発話主要文節に基づいて、第1発話文及び第2発話文の各々の発話対象特徴量を抽出する。
ステップS160において、特徴量集約部135は、上記ステップS120により抽出された第1発話文及び第2発話文の各々についてのn-gramと、上記ステップS140により抽出された第1発話文及び第2発話文の各々についての機能的特徴量と、上記ステップS150により抽出された第1発話文及び第2発話文の各々についての発話対象特徴量とを集約して集約特徴量とする。
ステップS170において、モデル学習部140は、上記ステップS160により抽出された学習データに含まれる第1発話文及び第2発話文についての集約特徴量と、対話行為推定モデルとに基づいて推定される第1発話文の対話行為タイプが、上記ステップS110により入力された学習データに含まれる第1発話文の対話行為タイプと一致するように対話行為推定モデルのパラメータを学習する。
ステップS180において、モデル学習部140は、終了条件を満たすか否かを判定する。終了条件を満たしていない場合(上記ステップS180のNO)、上記ステップS100に戻り、ステップS100~S180の処理を繰り返す。一方、終了条件を満たしている場合(上記ステップS180のYES)、ステップS190において、モデル学習部140は、学習した対話行為推定モデルのパラメータを、対話行為推定モデル記憶部150に格納する。
以上説明したように、本発明の実施の形態に係る対話行為推定モデル学習装置によれば、第1発話文と当該第1発話文の少なくとも直前の発話文を含む当該第1発話文より前の発話文である第2発話文との各々について、発話文の発話対象に関する特徴量である発話対象特徴量を含む特徴量を抽出し、抽出した第1発話文及び第2発話文の各々についての特徴量を集約した集約特徴量と、対話行為推定モデルとに基づいて推定される第1発話文の対話行為タイプが、学習データに含まれる第1発話文の対話行為タイプと一致するように対話行為推定モデルのパラメータを学習することにより、発話対象を考慮した対話行為タイプを精度よく推定するための対話行為推定モデルを学習することができる。
<本発明の実施の形態に係る対話行為推定装置の構成>
次に、図1及び図5を参照して、本発明の実施の形態に係る対話行為推定装置200の構成について説明する。なお、本発明の実施の形態に係る対話行為推定モデル学習装置100と同様の構成については、同一の符号を付して詳細な説明は省略する。
次に、図1及び図5を参照して、本発明の実施の形態に係る対話行為推定装置200の構成について説明する。なお、本発明の実施の形態に係る対話行為推定モデル学習装置100と同様の構成については、同一の符号を付して詳細な説明は省略する。
図1に示すように、本発明の実施の形態に係る対話行為推定装置200は、CPU11と、RAM等のメモリ12と、通信インターフェース(IF)部13と、キーボード等の入力部14と、ディスプレイ等の表示部15と、後述する対話行為推定処理ルーチンを実行するためのプログラム27を記憶したROM等の記憶部16とを備えたコンピュータで構成されている。また、CPU11、メモリ12、通信IF部13、入力部14、表示部15、及び記憶部16は、バス10を介して接続されている。また、通信IF部13は、LANケーブル等の通信回線により外部端末と接続することができる。
図5に示すように、本発明の実施の形態に係る対話行為推定装置200は、入力部210と、テキスト解析部120と、特徴量抽出部130と、対話行為推定モデル記憶部150と、対話行為推定部260と、出力部270とを備えて構成される。
対話行為推定モデル記憶部150には、対話行為推定モデルと対話行為推定モデル学習装置100により予め学習された対話行為推定モデルのパラメータとが格納されている。
入力部210は、第1発話文と当該第1発話文の少なくとも直前の発話文を含む当該第1発話文より前の発話文である第2発話文との入力を受け付ける。そして、入力部210は、受け付けた第1発話文及び第2発話文を、テキスト解析部120に渡す。
対話行為推定部260は、集約特徴量と、予め学習された、発話文の発話対象を考慮した対話行為の種類を示す対話行為タイプを推定するための対話行為推定モデルとを用いて、第1発話文の対話行為タイプを推定する。
具体的には、対話行為推定部260は、まず、対話行為推定モデル記憶部150から、対話行為推定モデルと対話行為推定モデルのパラメータとを取得する。次に、対話行為推定部260は、特徴量抽出部130により抽出された集約特徴量と、取得した対話行為推定モデルに基づいて、第1発話文の対話行為タイプを推定する。そして、対話行為推定部260は、推定した対話行為タイプを出力部270に渡す。
出力部270は、対話行為推定部260により推定された対話行為タイプを出力する。
<本発明の実施の形態に係る対話行為推定装置の作用>
図6は、本発明の実施の形態に係る対話行為推定処理ルーチンを示すフローチャートである。なお、本発明の実施の形態に係る対話行為推定モデル学習処理ルーチンと同様の処理については、同一の符号を付して詳細な説明は省略する。
図6は、本発明の実施の形態に係る対話行為推定処理ルーチンを示すフローチャートである。なお、本発明の実施の形態に係る対話行為推定モデル学習処理ルーチンと同様の処理については、同一の符号を付して詳細な説明は省略する。
ステップS200において、入力部210は、第1発話文と当該第1発話文の少なくとも直前の発話文を含む当該第1発話文より前の発話文である第2発話文との入力を受け付ける。
ステップS270において、対話行為推定部260は、対話行為推定モデル記憶部150から、予め学習された、発話文の発話対象を考慮した対話行為の種類を示す対話行為タイプを推定するための対話行為推定モデルと対話行為推定モデルのパラメータとを取得する。
ステップS280において、対話行為推定部260は、集約特徴量と、上記ステップS270により取得した対話行為推定モデルとを用いて、第1発話文の対話行為タイプを推定する。
ステップS290において、上記ステップS280により推定された第1発話文の対話行為タイプを出力する。
以上説明したように、本実施の形態に係る対話行為推定装置によれば、第1発話文と当該第1発話文の少なくとも直前の発話文を含む当該第1発話文より前の発話文である第2発話文との各々について、発話文の発話対象に関する特徴量である発話対象特徴量を含む特徴量を抽出し、抽出した第1発話文及び第2発話文の各々についての特徴量を集約した集約特徴量と、予め学習された、発話文の発話対象を考慮した対話行為の種類を示す対話行為タイプを推定するための対話行為推定モデルとを用いて、第1発話文の対話行為タイプを推定することにより、発話対象を考慮した対話行為タイプを精度よく推定することができる。そして、このように推定した対話行為タイプに基づいて対話システムが応答生成ロジックを適切に選択できるようになることにより、対話システム全体の対話精度を向上できる。
また、本実施の形態に係る対話行為推定装置では、集約特徴量にn-gramも含まれるため、従来の対話行為タイプには「挨拶」や「Feedback」のように、発話対象が自明のものについては、従来の体系をそのまま用いることができる。
なお、本発明は、上述した実施の形態に限定されるものではなく、この発明の要旨を逸脱しない範囲内で様々な変形や応用が可能である。
また、本願明細書中において、プログラムが予めインストールされている実施形態として説明したが、当該プログラムを、コンピュータ読み取り可能な記録媒体に格納して提供することも可能である。
10 バス
11 CPU
12 メモリ
13 通信IF部
14 入力部
15 表示部
16 記憶部
17 プログラム
27 プログラム
100 対話行為推定モデル学習装置
110 入力部
120 テキスト解析部
130 特徴量抽出部
131 単語n-gram抽出部
132 発話主要文節特定部
133 機能的特徴量抽出部
134 発話対象特徴量抽出部
135 特徴量集約部
140 モデル学習部
150 対話行為推定モデル記憶部
200 対話行為推定装置
210 入力部
260 対話行為推定部
270 出力部
11 CPU
12 メモリ
13 通信IF部
14 入力部
15 表示部
16 記憶部
17 プログラム
27 プログラム
100 対話行為推定モデル学習装置
110 入力部
120 テキスト解析部
130 特徴量抽出部
131 単語n-gram抽出部
132 発話主要文節特定部
133 機能的特徴量抽出部
134 発話対象特徴量抽出部
135 特徴量集約部
140 モデル学習部
150 対話行為推定モデル記憶部
200 対話行為推定装置
210 入力部
260 対話行為推定部
270 出力部
Claims (5)
- 第1発話文と前記第1発話文の少なくとも直前の発話文を含む前記第1発話文より前の発話文である第2発話文との入力を受け付ける入力部と、
前記第1発話文及び前記第2発話文の各々について、発話文の発話対象に関する特徴量である発話対象特徴量を含む特徴量を抽出し、抽出した前記第1発話文及び前記第2発話文の各々についての前記特徴量を集約して集約特徴量とする特徴量抽出部と、
前記集約特徴量と、予め学習された、発話文の発話対象を考慮した対話行為の種類を示す対話行為タイプを推定するための対話行為推定モデルとを用いて、前記第1発話文の前記対話行為タイプを推定する対話行為推定部と、
を含む対話行為推定装置。 - 前記特徴量抽出部は、
前記第1発話文と前記第2発話文との各々について、発話文の内容を最も表す文節である発話主要文節を特定する発話主要文節特定部と、
前記発話主要文節特定部により特定された前記第1発話文及び前記第2発話文の各々についての発話主要文節に含まれる、発話文の機能的な特徴量である機能的特徴量を抽出する機能的特徴量抽出部と、
前記発話主要文節特定部により特定された前記第1発話文及び前記第2発話文の各々についての発話主要文節に基づいて、前記第1発話文及び前記第2発話文の各々の前記発話対象特徴量を抽出する発話対象特徴量抽出部と、
前記機能的特徴量抽出部により抽出された前記第1発話文及び前記第2発話文の各々についての前記機能的特徴量と、前記発話対象特徴量抽出部により抽出された前記第1発話文及び前記第2発話文の各々についての前記発話対象特徴量とを集約して前記集約特徴量とする特徴量集約部
を含む請求項1記載の対話行為推定装置。 - 第1発話文と前記第1発話文の少なくとも直前の発話文を含む前記第1発話文より前の発話文である第2発話文と、前記第1発話文の発話対象を考慮した対話行為の種類を示す対話行為タイプとを含む学習データの入力を受け付ける入力部と、
前記第1発話文及び前記第2発話文の各々について、発話文の発話対象に関する特徴量である発話対象特徴量を含む特徴量を抽出し、抽出した前記第1発話文及び前記第2発話文の各々についての前記特徴量を集約して集約特徴量とする特徴量抽出部と、
前記特徴量抽出部により抽出された前記第1発話文及び前記第2発話文についての集約特徴量と、発話文の発話対象を考慮した対話行為の種類を示す対話行為タイプを推定するための対話行為推定モデルとに基づいて推定される前記第1発話文の前記対話行為タイプが、前記学習データに含まれる前記第1発話文の前記対話行為タイプと一致するように、前記対話行為推定モデルのパラメータを学習するモデル学習部と、
を含む対話行為推定モデル学習装置。 - 入力部が、第1発話文と前記第1発話文の少なくとも直前の発話文を含む前記第1発話文より前の発話文である第2発話文との入力を受け付け、
特徴量抽出部が、前記第1発話文及び前記第2発話文の各々について、発話文の発話対象に関する特徴量である発話対象特徴量を含む特徴量を抽出し、抽出した前記第1発話文及び前記第2発話文の各々についての前記特徴量を集約して集約特徴量とし、
対話行為推定部が、前記集約特徴量と、予め学習された、発話文の発話対象を考慮した対話行為の種類を示す対話行為タイプを推定するための対話行為推定モデルとを用いて、前記第1発話文の前記対話行為タイプを推定する
対話行為推定方法。 - 入力部が、第1発話文と前記第1発話文の少なくとも直前の発話文を含む前記第1発話文より前の発話文である第2発話文との入力を受け付け、
特徴量抽出部が、前記第1発話文及び前記第2発話文の各々について、発話文の発話対象に関する特徴量である発話対象特徴量を含む特徴量を抽出し、抽出した前記第1発話文及び前記第2発話文の各々についての前記特徴量を集約して集約特徴量とし、
対話行為推定部が、前記集約特徴量と、予め学習された、発話文の発話対象を考慮した対話行為の種類を示す対話行為タイプを推定するための対話行為推定モデルとを用いて、前記第1発話文の前記対話行為タイプを推定する
ことを含む処理をコンピュータに実行させるためのプログラム。
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US17/602,281 US20220164545A1 (en) | 2019-04-10 | 2020-03-25 | Dialog action estimation device, dialog action estimation method, dialog action estimation model learning device, and program |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP2019075055A JP7180513B2 (ja) | 2019-04-10 | 2019-04-10 | 対話行為推定装置、対話行為推定方法、対話行為推定モデル学習装置及びプログラム |
| JP2019-075055 | 2019-04-10 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2020209072A1 true WO2020209072A1 (ja) | 2020-10-15 |
Family
ID=72751072
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/JP2020/013445 Ceased WO2020209072A1 (ja) | 2019-04-10 | 2020-03-25 | 対話行為推定装置、対話行為推定方法、対話行為推定モデル学習装置及びプログラム |
Country Status (3)
| Country | Link |
|---|---|
| US (1) | US20220164545A1 (ja) |
| JP (1) | JP7180513B2 (ja) |
| WO (1) | WO2020209072A1 (ja) |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP7460377B2 (ja) | 2020-01-28 | 2024-04-02 | 浜松ホトニクス株式会社 | レーザ加工装置及びレーザ加工方法 |
Families Citing this family (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP7692562B2 (ja) * | 2021-11-10 | 2025-06-16 | 日本電信電話株式会社 | 対話映像要約装置、対話映像要約方法、および、対話映像要約プログラム |
| US20240371369A1 (en) * | 2023-05-03 | 2024-11-07 | Origin8Cares, LLC | Transcript pairing |
Citations (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2016001242A (ja) * | 2014-06-11 | 2016-01-07 | 日本電信電話株式会社 | 質問文生成方法、装置、及びプログラム |
| JP2017228160A (ja) * | 2016-06-23 | 2017-12-28 | パナソニックIpマネジメント株式会社 | 対話行為推定方法、対話行為推定装置及びプログラム |
Family Cites Families (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US9262395B1 (en) * | 2009-02-11 | 2016-02-16 | Guangsheng Zhang | System, methods, and data structure for quantitative assessment of symbolic associations |
| US8375033B2 (en) * | 2009-10-19 | 2013-02-12 | Avraham Shpigel | Information retrieval through identification of prominent notions |
-
2019
- 2019-04-10 JP JP2019075055A patent/JP7180513B2/ja active Active
-
2020
- 2020-03-25 WO PCT/JP2020/013445 patent/WO2020209072A1/ja not_active Ceased
- 2020-03-25 US US17/602,281 patent/US20220164545A1/en not_active Abandoned
Patent Citations (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2016001242A (ja) * | 2014-06-11 | 2016-01-07 | 日本電信電話株式会社 | 質問文生成方法、装置、及びプログラム |
| JP2017228160A (ja) * | 2016-06-23 | 2017-12-28 | パナソニックIpマネジメント株式会社 | 対話行為推定方法、対話行為推定装置及びプログラム |
Non-Patent Citations (1)
| Title |
|---|
| KIMURA, SHINICHI ET AL.: "An estimation method of speech intention using the engagement relation", PROCEEDINGS OF THE 63RD (SECOND HALF OF 2001) NATIONAL CONVENTION (2, 26 September 2001 (2001-09-26), pages 2 - 197 , 2-198 * |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP7460377B2 (ja) | 2020-01-28 | 2024-04-02 | 浜松ホトニクス株式会社 | レーザ加工装置及びレーザ加工方法 |
Also Published As
| Publication number | Publication date |
|---|---|
| US20220164545A1 (en) | 2022-05-26 |
| JP7180513B2 (ja) | 2022-11-30 |
| JP2020173608A (ja) | 2020-10-22 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| JP7849489B2 (ja) | 検索エンジン結果を使用した機械学習言語モデルの拡張 | |
| CN110543643B (zh) | 文本翻译模型的训练方法及装置 | |
| US11636272B2 (en) | Hybrid natural language understanding | |
| US9582757B1 (en) | Scalable curation system | |
| CN112560510B (zh) | 翻译模型训练方法、装置、设备及存储介质 | |
| CN106601237B (zh) | 交互式语音应答系统及其语音识别方法 | |
| CN106649825B (zh) | 语音交互系统及其创建方法和装置 | |
| WO2023108994A1 (zh) | 一种语句生成方法及电子设备、存储介质 | |
| WO2022142121A1 (zh) | 摘要语句提取方法、装置、服务器及计算机可读存储介质 | |
| JP6370962B1 (ja) | 生成装置、生成方法および生成プログラム | |
| CN110019742B (zh) | 用于处理信息的方法和装置 | |
| CN101454826A (zh) | 语音识别词典/语言模型制作系统、方法、程序,以及语音识别系统 | |
| KR101677859B1 (ko) | 지식 베이스를 이용하는 시스템 응답 생성 방법 및 이를 수행하는 장치 | |
| JPWO2016151700A1 (ja) | 意図理解装置、方法およびプログラム | |
| WO2020209072A1 (ja) | 対話行為推定装置、対話行為推定方法、対話行為推定モデル学習装置及びプログラム | |
| JP2017125921A (ja) | 発話選択装置、方法、及びプログラム | |
| WO2022022049A1 (zh) | 文本长难句的压缩方法、装置、计算机设备及存储介质 | |
| CN110705212A (zh) | 文本序列的处理方法、处理装置、电子终端和介质 | |
| Lorenc et al. | Benchmark of public intent recognition services | |
| JP2020004054A (ja) | 出力装置、出力方法および出力プログラム | |
| JP2015225416A (ja) | モデル学習装置、ランキング装置、方法、及びプログラム | |
| CN109002498B (zh) | 人机对话方法、装置、设备及存储介质 | |
| CN113590747B (zh) | 用于意图识别的方法以及相应的系统、计算机设备和介质 | |
| JP5860439B2 (ja) | 言語モデル作成装置とその方法、そのプログラムと記録媒体 | |
| Vologina et al. | RAG and few-shot prompting in emotional text generation |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 20787486 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 20787486 Country of ref document: EP Kind code of ref document: A1 |
