WO2021181679A1 - 対話支援装置、対話支援方法及びプログラム - Google Patents
対話支援装置、対話支援方法及びプログラム Download PDFInfo
- Publication number
- WO2021181679A1 WO2021181679A1 PCT/JP2020/011193 JP2020011193W WO2021181679A1 WO 2021181679 A1 WO2021181679 A1 WO 2021181679A1 JP 2020011193 W JP2020011193 W JP 2020011193W WO 2021181679 A1 WO2021181679 A1 WO 2021181679A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- speaker
- dialogue
- knowledge level
- question sentence
- support device
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/22—Procedures used during a speech recognition process, e.g. man-machine dialogue
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L13/00—Speech synthesis; Text to speech systems
- G10L13/02—Methods for producing synthetic speech; Speech synthesisers
- G10L13/027—Concept to speech synthesisers; Generation of natural phrases from machine-based concepts
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/08—Speech classification or search
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/08—Speech classification or search
- G10L2015/088—Word spotting
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/22—Procedures used during a speech recognition process, e.g. man-machine dialogue
- G10L2015/223—Execution procedure of a spoken command
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/22—Procedures used during a speech recognition process, e.g. man-machine dialogue
- G10L2015/225—Feedback of the input speech
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/22—Procedures used during a speech recognition process, e.g. man-machine dialogue
- G10L2015/226—Procedures used during a speech recognition process, e.g. man-machine dialogue using non-speech characteristics
- G10L2015/227—Procedures used during a speech recognition process, e.g. man-machine dialogue using non-speech characteristics of the speaker; Human-factor methodology
Definitions
- the present invention relates to a dialogue support device, a dialogue support method, and a program.
- the present invention has been made in view of the above points, and an object of the present invention is to support facilitation of dialogue.
- the dialogue support device estimates the knowledge level of the second speaker who has a dialogue with the first speaker in the field related to the utterance content of the first speaker. From the estimation unit of the above and the storage unit that stores the question sentence in association with the keyword and the knowledge level, the question sentence corresponding to the keyword included in the utterance content and the knowledge level of the second speaker It has an acquisition unit for acquiring the corresponding question text and an output unit for outputting the acquired question text to the first speaker.
- FIG. 10 It is a figure which shows the hardware configuration example of the dialogue support apparatus 10 in embodiment of this invention. It is a figure which shows the functional structure example of the dialogue support apparatus 10 in embodiment of this invention. It is a flowchart for demonstrating an example of the processing procedure of the keyword extraction processing from the utterance content of speaker A. It is a flowchart for demonstrating an example of a process procedure of a dialogue support process. It is a figure which shows the structural example of the knowledge level DB 122. It is a figure which shows the structural example of the question sentence DB 123. In a specific example, the processing content executed by the dialogue support device 10 is shown. It is a figure which shows the situation which there are a plurality of persons corresponding to speaker B.
- a speaker A having a high literacy (high knowledge level) and a speaker having a relatively low literacy (low knowledge level) in a certain field for example, ICT (Information and Communication Technology)
- a situation is assumed in which B has a dialogue.
- the speaker A may be a person in charge at a counter of a certain store, and the speaker B may be a person who consults with the speaker A at the counter.
- the situation setting is for facilitating the understanding of the present embodiment, and does not mean that the situation in which the present embodiment is effective is limited to the above situation.
- a dialogue support device 10 for supporting the dialogue is installed at a place where the speaker A and the speaker B have a dialogue.
- the dialogue support device 10 may have the shape of a robot. However, a PC (Personal Computer), a smartphone, or the like may be used as the dialogue support device 10.
- FIG. 1 is a diagram showing a hardware configuration example of the dialogue support device 10 according to the embodiment of the present invention.
- the dialogue support device 10 of FIG. 1 includes a drive device 100, an auxiliary storage device 102, a memory device 103, a CPU 104, a microphone 105, a display device 106, a camera 107, and the like, which are connected to each other by a bus B, respectively.
- the program that realizes the processing in the dialogue support device 10 is provided by a recording medium 101 such as a CD-ROM.
- a recording medium 101 such as a CD-ROM.
- the program is installed in the auxiliary storage device 102 from the recording medium 101 via the drive device 100.
- the program does not necessarily have to be installed from the recording medium 101, and may be downloaded from another computer via the network.
- the auxiliary storage device 102 stores the installed program and also stores necessary files, data, and the like.
- the memory device 103 reads and stores the program from the auxiliary storage device 102 when the program is instructed to start.
- the CPU 104 realizes the function related to the dialogue support device 10 according to the program stored in the memory device 103.
- the microphone 105 is used for inputting the voice of the dialogue (particularly, the utterance content of the speaker A).
- the display device 106 is, for example, a liquid crystal display or the like, and when the speaker B cannot understand the utterance content of the speaker A, the display device 106 outputs (displays) the voice of the question sentence to the speaker A as described later. Used for The display device 106 may have a shape like a window installed between the speaker A and the speaker B, for example.
- the camera 107 is, for example, a digital camera, and is used for inputting an image of the face of speaker B (hereinafter, referred to as “face image”).
- face image an image of the face of speaker B
- the microphone 105, the display device 106, the camera 107, and the like may not be built in the dialogue support device 10, and may be connected to the dialogue support device 10 by, for example, wirelessly or by wire.
- FIG. 2 is a diagram showing a functional configuration example of the dialogue support device 10 according to the embodiment of the present invention.
- the dialogue support device 10 includes a keyword extraction unit 11, a comprehension level estimation unit 12, a knowledge level estimation unit 13, a question sentence acquisition unit 14, a question sentence output unit 15, and the like. Each of these parts is realized by a process of causing the CPU 104 to execute one or more programs installed in the dialogue support device 10.
- the dialogue support device 10 also uses storage units such as a keyword storage unit 121, a knowledge level DB 122 (Data Base), and a question sentence DB 123. These storage units can be realized by using, for example, a storage device that can be connected to the memory device 103, the auxiliary storage device 102, or the dialogue support device 10 via a network.
- FIG. 3 is a flowchart for explaining an example of the processing procedure of the keyword extraction processing from the utterance content of the speaker A.
- the processing procedure of FIG. 3 is started, for example, in response to the start of a dialogue between speaker A and speaker B.
- the keyword extraction unit 11 When the speaker A starts utterance, the keyword extraction unit 11 inputs the utterance voice of the speaker A via the microphone 105 (S101). For example, at the timing when the utterance ends, the keyword extraction unit 11 applies voice recognition to the utterance voice input for the utterance, and extracts one or more keywords from the text data obtained as a result of the voice recognition. (S102). For example, if the spoken voice is "Do you use tethering?", "Tethering" may be extracted as a keyword.
- Keywords may be extracted using the method described in 2003-SLP-048)), 21-28. ”.
- the keyword registered in the knowledge level DB 122 which will be described later, may be the extraction target.
- the keyword extraction unit 11 records the extracted keyword in the keyword storage unit 121 (S103), and waits for the next utterance by the speaker A (S101). Each keyword is recorded in the keyword storage unit 121 so that the extraction order (utterance order) of the keywords can be identified.
- FIG. 4 is a flowchart for explaining an example of the processing procedure of the dialogue support processing.
- the processing procedure of FIG. 4 is started, for example, in response to the start of a dialogue between speaker A and speaker B, and is executed in parallel (parallel) with the processing procedure of FIG.
- the comprehension estimation unit 12 inputs the face image of the speaker B continuously photographed by the camera 107 (S201), and the comprehension level of the speaker B with respect to the utterance content of the speaker A based on the face image. Is estimated (calculated) (S202). That is, when it is difficult to understand the utterance content of the speaker A, there is a high possibility that the facial expression of the speaker B will change. Therefore, the comprehension estimation unit 12 estimates the comprehension based on the facial expression of the speaker B. For the estimation of such comprehension, see, for example, "Atsushi Mimura, Masafumi Hagiwara, Comprehension estimation system obtained from facial expressions, IEEJ Transactions. C, 120 (2), 2000, 273-278.” It may be done using the techniques described.
- the degree of comprehension is estimated in five stages (0 to 4) from a state in which the person does not understand at all to a state in which the person fully understands.
- the degree of comprehension may be estimated by using the existing speech recognition technique or text analysis technique by inputting the utterance content of the speaker A or the speaker B.
- the comprehension estimation unit 12 estimates whether or not the comprehension of the speaker B is less than the threshold value (S203).
- the smaller the value of comprehension the lower the degree of comprehension. Therefore, in step S203, it is determined whether or not the degree of understanding of the speaker B is low.
- step S201 Return to.
- the knowledge level estimation unit 13 is based on one or more keywords stored in the keyword storage unit 121 and the knowledge level DB 122.
- the knowledge level of speaker B in a field related to the utterance content of A is estimated (S204). That is, it is estimated how much the speaker B has the knowledge about the field.
- FIG. 5 is a diagram showing a configuration example of the knowledge level DB 122.
- the knowledge level DB 122 stores the knowledge level in association with the keyword.
- an example in which the knowledge level is expressed by a numerical value (score) is shown, but the knowledge level is expressed by a label or the like (for example, “high”, “medium”, “low”, etc.). May be good.
- the keyword group used in step S104 (hereinafter referred to as "target keyword group”) may be limited to a predetermined number or less, for example, the latest N elements. In the dialogue with speaker A, the topic may change over time.
- the keyword storage unit 121 may be a FIFO (First-In First-Out) type storage area that can store only N keywords.
- the target keyword group does not have to be extracted only from the utterance content of the speaker A. For example, it may be extracted from the utterance contents of speaker A and speaker B in the latest M dialogues. In this case, in step S101 of FIG. 3, not only the uttered voice of the speaker A but also the uttered voice of the speaker B may be input.
- the comprehension estimation unit 12 acquires the knowledge level for each target keyword from the knowledge level DB 122, and sets the lowest value among the acquired knowledge levels as the speaker. It may be estimated as the knowledge level of B. Alternatively, the comprehension estimation unit 12 sets the highest value among the knowledge levels corresponding to any of the target keywords recorded in the keyword storage unit 121 before the comprehension is estimated to be less than the threshold value. It may be estimated as the knowledge level of B. This is because it is highly possible that speaker B understood the keywords before the comprehension level was estimated to be less than the threshold value.
- the technique disclosed in Japanese Patent Application Laid-Open No. 2013-167765 may be used.
- the history of the dialogue between the speaker A and the speaker B is recorded, and the knowledge level estimation unit 13 may estimate the knowledge level (knowledge amount) of the speaker B based on the history.
- the knowledge level of speaker B may be estimated using the technique disclosed in Japanese Patent Application Laid-Open No. 2019-28604.
- the question sentence acquisition unit 14 acquires the question sentence to be output to the speaker A from the question sentence DB 123 based on the target keyword group and the knowledge level group estimated for the speaker B (S205).
- FIG. 6 is a diagram showing a configuration example of the question sentence DB 123.
- the “question sentence” and the “number of outputs” are stored in association with the “keyword” and the “necessary knowledge level”.
- the "question text” of each record indicates a question text that should be output when a person who does not understand the "keyword” of the record has a knowledge level equal to or higher than the "necessary knowledge level” of the record.
- the "number of outputs" of each record indicates the number of times the "question text" of the record has been output in the past.
- the question sentence acquisition unit 14 is a record in which any of the keywords included in the target keyword group is included in the "keyword", and the "required knowledge level” is equal to or lower than the knowledge level of the speaker B. Get the "question text” of a record. When there are a plurality of corresponding “question sentences", for example, they may be sorted in descending order of "number of outputs".
- the question sentence output unit 15 outputs (displays) the question sentence acquired by the question sentence acquisition unit 14 to the display device 106 (S206).
- the display device 106 is arranged so that the speaker A and the speaker B can see.
- speaker A utters the answer to the question. It can be expected that the speaker B will be able to understand the utterance content of the speaker A who could not understand based on the question sentence and the answer.
- a specific example of the dialogue between the speaker A and the speaker B and the question text output by the dialogue support device 10 is shown below.
- a (1) "Do you use wireless LAN at home?" B (1) "Yes.”
- a (2) "Do you use tethering while you are out?”
- Dialogue support device 10 "Can I use the Internet on a laptop computer by using tethering?" A (3) "Yes.”
- B (3) "I don't use a laptop outside, so I don't think I use tethering.”
- Step S202 and subsequent steps are executed, and in step S206 executed as a result, the dialogue support device 10 "tethers.” If you use it, will you be able to use the Internet on your laptop computer? “Is output to speaker A on behalf of speaker B. According to Speaker A's answer (“Yes.”), Speaker B can answer to utterance A (2) (speech B (3)) without fully understanding the meaning of “tethering”. It is possible, and the dialogue between the two is smooth. That is, the dialogue between the two is engaged, and it is avoided that the dialogue collapses.
- FIG. 7 shows the processing content executed by the dialogue support device 10 in a specific example.
- the keyword extraction unit 11 extracts "wireless LAN” as a keyword.
- the comprehension estimation unit 12 estimates the comprehension of the speaker B as “high”. Therefore, the knowledge level estimation unit 13 and the question sentence acquisition unit 14 do not execute the process.
- the keyword extraction unit 11 extracts "tethering" as a keyword.
- the keyword storage unit 121 shows an example in which up to four keywords are stored. Therefore, at this point in time, the keyword storage unit 121 stores two keywords, "wireless LAN” and "tethering".
- the comprehension estimation unit 12 estimates the comprehension of the speaker B as “low”.
- the knowledge level estimation unit 13 estimates the knowledge level of speaker B to be "50”
- the question sentence acquisition unit 14 selects (2) from the record group shown in (1) of FIG. Get the indicated record from the question DB.
- a record group including any one of the target keyword groups (“wireless LAN”, “tethering”) in the “keyword” is shown.
- records having a "required knowledge level" of 50 or less are shown in the record group of (1).
- the question text output unit 15 outputs the question text "Can the Internet be used on a laptop computer by using tethering?" On behalf of the speaker B. ..
- the question sentence output unit 15 may output the question sentence by voice.
- the dialogue support device 10 may have a speaker.
- the person corresponding to speaker B (for example, the counselor for speaker A) is a group of a plurality of people (speaker B1 to speaker BN in FIG. 8).
- NS a threshold value for the number of speakers corresponding to speaker B is set in the dialogue support device 10, and the question sentence output unit 15 asks a question sentence when there are more people than the threshold value. May not be output.
- each speaker B can avoid being inferred by another speaker B that his / her knowledge level is low due to the output of the question sentence by the dialogue support device 10.
- the comprehension estimation unit 12 estimates the comprehension level of each speaker B (in parallel), and the knowledge level estimation unit 13 estimates the knowledge level of each speaker B (in parallel). good. Even if the question sentence acquisition unit 14 acquires the question sentence to be output to the speaker A from the question sentence DB 123 based on the lowest knowledge level among the plurality of estimated knowledge levels (in parallel). good. By doing so, the question sentence may be output according to the speaker B having the lowest knowledge level.
- the dialogue support device 10 replaces the speaker B.
- a question sentence corresponding to the knowledge level of speaker B is output (notified) to speaker A.
- the speaker B can respond to the utterance content based on the answer without completely understanding the utterance content that he / she could not understand. .. Therefore, it is possible to support the facilitation of dialogue.
- the knowledge level estimation unit 13 is an example of the first estimation unit in the mobile phone of the present implementation.
- the question sentence acquisition unit 14 is an example of the acquisition unit.
- the question sentence output unit 15 is an example of an output unit.
- the comprehension estimation unit 12 is an example of the second estimation unit.
- Speaker A is an example of the first speaker.
- Speaker B is an example of a second speaker.
- Dialogue support device 11 Keyword extraction unit 12 Understanding level estimation unit 13 Knowledge level estimation unit 14
- Question text acquisition unit 15
- Question text output unit 100
- Drive device 101
- Recording medium 102
- Auxiliary storage device 103
- Memory device 104
- Microphone 106
- Display device 107 and camera 121 Keyword storage unit 122
- Knowledge level DB 123
- Question text DB B bus
Landscapes
- Engineering & Computer Science (AREA)
- Computational Linguistics (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Human Computer Interaction (AREA)
- Physics & Mathematics (AREA)
- Acoustics & Sound (AREA)
- Multimedia (AREA)
- Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
- Machine Translation (AREA)
Abstract
対話支援装置は、第1の話者の発話内容に関連する分野に関して、前記第1の話者と対話を行う第2の話者の知識レベルを推定する第1の推定部と、キーワード及び知識レベルに対応付けて質問文を記憶する記憶部から、前記発話内容に含まれるキーワードに対応する質問文であって、かつ、前記第2の話者の知識レベルに対応する質問文を取得する取得部と、取得された質問文を前記第1の話者に対して出力する出力部と、を有することで、対話の円滑化を支援する。
Description
本発明は、対話支援装置、対話支援方法及びプログラムに関する。
2以上の話者が対話を行う場合に、話者同士が相手の知識レベルに応じて発話することは困難である。
例えば、ICT(Information and Communication Technology)に関して高リテラシな(ICT用語に対する理解度が高い)話者Aと、ICTに関して低リテラシな(ICT用語に対する理解度が低い)話者Bとの間でICTに関する対話が行われる場合、話者Aの発話内容を話者Bが理解することができず、当該対話が破綻する可能性がある。
従来、ユーザとロボットとの対話の破綻を防ぐ技術は考案されている。
しかしながら、従来技術では、対話を行う話者の知識レベルを考慮していないため、一方の話者の対話内容について他方の話者の理解を助けることはできなかった。その結果、対話の円滑化を支援するのが困難であった。
本発明は、上記の点に鑑みてなされたものであって、対話の円滑化を支援することを目的とする。
そこで上記課題を解決するため、対話支援装置は、第1の話者の発話内容に関連する分野に関して、前記第1の話者と対話を行う第2の話者の知識レベルを推定する第1の推定部と、キーワード及び知識レベルに対応付けて質問文を記憶する記憶部から、前記発話内容に含まれるキーワードに対応する質問文であって、かつ、前記第2の話者の知識レベルに対応する質問文を取得する取得部と、取得された質問文を前記第1の話者に対して出力する出力部と、を有する。
対話の円滑化を支援することができる。
以下、図面に基づいて本発明の実施の形態を説明する。本実施の形態では、或る分野(例えば、ICT(Information and Communication Technology)等)に関してリテラシの高い(知識レベルの高い)話者Aと、相対的にリテラシの低い(知識レベルの低い)話者Bとが対話を行う状況が想定される。例えば、話者Aは、或る店舗の窓口の担当者であり、話者Bが当該窓口において話者Aに対して相談を行う者であってもよい。なお、このような状況設定は、本実施の形態の理解を容易にするためのものであり、本実施の形態が有効な状況が、上記の状況に限定されることを意味するものではない。
話者Aと話者Bとが対話する場所には、当該対話を支援するための対話支援装置10が設置される。対話支援装置10は、ロボットの形状を有してもよい。但し、PC(Personal Computer)又はスマートフォン等が対話支援装置10として利用されてもよい。
図1は、本発明の実施の形態における対話支援装置10のハードウェア構成例を示す図である。図1の対話支援装置10は、それぞれバスBで相互に接続されているドライブ装置100、補助記憶装置102、メモリ装置103、CPU104、マイク105、表示装置106、及びカメラ107等を有する。
対話支援装置10での処理を実現するプログラムは、CD-ROM等の記録媒体101によって提供される。プログラムを記憶した記録媒体101がドライブ装置100にセットされると、プログラムが記録媒体101からドライブ装置100を介して補助記憶装置102にインストールされる。但し、プログラムのインストールは必ずしも記録媒体101より行う必要はなく、ネットワークを介して他のコンピュータよりダウンロードするようにしてもよい。補助記憶装置102は、インストールされたプログラムを格納すると共に、必要なファイルやデータ等を格納する。
メモリ装置103は、プログラムの起動指示があった場合に、補助記憶装置102からプログラムを読み出して格納する。CPU104は、メモリ装置103に格納されたプログラムに従って対話支援装置10に係る機能を実現する。マイク105は、対話の音声(特に、話者Aの発話内容)の入力に用いられる。表示装置106は、例えば、液晶ディスプレイ等であり、話者Bが話者Aの発話内容を理解できない場合等に、後述されるように、話者Aに対する質問文の音声を出力(表示)するために用いられる。なお、表示装置106は、例えば、話者Aと話者Bとの間に設置される窓のような形状を有していてもよい。カメラ107は、例えば、デジタルカメラであり、話者Bの顔の画像(以下、「顔画像」という)の入力に用いられる。なお、マイク105、表示装置106及びカメラ107等は、対話支援装置10に内蔵されていなくてもよく、例えば、無線又は有線によって、対話支援装置10と接続されてもよい。
図2は、本発明の実施の形態における対話支援装置10の機能構成例を示す図である。図2において、対話支援装置10は、キーワード抽出部11、理解度推定部12、知識レベル推定部13、質問文取得部14及び質問文出力部15等を有する。これら各部は、対話支援装置10にインストールされた1以上のプログラムが、CPU104に実行させる処理により実現される。対話支援装置10は、また、キーワード記憶部121、知識レベルDB122(Data Base)及び質問文DB123等の記憶部を利用する。これら記憶部は、例えば、メモリ装置103、補助記憶装置102、又は対話支援装置10にネットワークを介して接続可能な記憶装置等を用いて実現可能である。
以下、対話支援装置10が実行する処理について説明する。図3は、話者Aの発話内容からのキーワードの抽出処理の処理手順の一例を説明するためのフローチャートである。図3の処理手順は、例えば、話者Aと話者Bとの対話の開始に応じて開始される。
話者Aが発話を開始すると、キーワード抽出部11は、マイク105を介して話者Aの発話音声を入力する(S101)。例えば、当該発話が終了したタイミングで、キーワード抽出部11は、当該発話に関して入力された発話音声に対して音声認識を適用し、音声認識の結果として得られるテキストデータから1以上のキーワードを抽出する(S102)。例えば、発話音声が「テザリングは使っていますか。」であれば、「テザリング」がキーワードとして抽出されてもよい。
このようなキーワードの抽出は、公知技術を用いて行うことができる。例えば、「松下雅彦, 西崎博光, 宇津呂武仁, & 中川聖一、音声入力による Web 検索のための キーワード認識・抽出法の検討、情報処理学会研究報告音声言語情報処理 (SLP),2003(104 (2003-SLP-048)), 21-28.」に記載された方法を用いてキーワードの抽出が行われてもよい。又は、後述される知識レベルDB122に登録されているキーワードが抽出対象とされてもよい。
続いて、キーワード抽出部11は、抽出したキーワードをキーワード記憶部121に記録し(S103)、話者Aによる次の発話を待機する(S101)。なお、キーワード記憶部121には、キーワードの抽出順(発話順)が識別可能なように各キーワードが記録される。
図4は、対話の支援処理の処理手順の一例を説明するためのフローチャートである。図4の処理手順は、例えば、話者Aと話者Bとの対話の開始に応じて開始され、図3の処理手順と並行して(並列的に)実行される。
理解度推定部12は、カメラ107によって継続的に撮影されている話者Bの顔画像を入力して(S201)、当該顔画像に基づいて話者Aの発話内容に対する話者Bの理解度を推定(計算)する(S202)。すなわち、話者Aの発話内容の理解が困難な場合には、話者Bの表情が変化する可能性が高い。そこで、理解度推定部12は、話者Bの表情に基づいて、当該理解度を推定する。なお、斯かる理解度の推定は、例えば、「三村淳, 萩原将文、表情から得られる理解度の推定システム、電気学会論文誌. C, 120(2), 2000、273-278.」に記載されている技術を用いて行われてもよい。この場合、理解度は、全く理解していない状態から完全に理解している状態までの5段階(0~4)によって推定される。なお、本実施の形態では顔画像を入力として理解度の推定を行う場合を例に説明を行ったが、それに限定するものではない。話者Aまたは話者Bの発話内容を入力として、既存の音声認識技術やテキスト解析技術を用いて理解度の推定を行うようにしてもよい。
続いて、理解度推定部12は、話者Bの理解度が閾値未満であるか否かを推定する(S203)。なお、本実施の形態では、理解度の値が小さいほど、理解の程度が低いこととする。したがって、ステップS203では、話者Bの理解の程度が低いか否かが判定される。
話者Bの理解度が閾値以上である場合(S203でNo)、話者Bは話者Aの発話内容を理解できていることが推定され、話者Bを支援する必要はないためステップS201へ戻る。話者Bの理解度が閾値未満である場合(S203でYes)、知識レベル推定部13は、キーワード記憶部121に記憶されている1以上のキーワードと、知識レベルDB122とに基づいて、話者Aの発話内容に関連する分野(例えば、ICT等)に関する話者Bの知識レベルを推定する(S204)。すなわち、当該分野に関する知識を話者Bがどの程度有しているのかが推定される。
図5は、知識レベルDB122の構成例を示す図である。図5に示されるように、知識レベルDB122には、キーワードに対応付けて知識レベルが記憶されている。図5では、知識レベルが、数値(スコア)によって表現される例が示されているが、知識レベルは、ラベル等(例えば、「高」、「中」、「低」等)によって表現されてもよい。なお、ステップS104において用いられるキーワード群(以下、「対象キーワード群」という。)は、例えば、直近のN個等、所定数以下に限定されてもよい。話者Aとの対話において、時間の経過に応じて話題が変化している可能性がある。対象キーワード群を直近のN個に限定することで、現在の話題とは関係性が低い話題に関連するキーワードが対象キーワード群から除外され、現在の話題に関連する分野の知識レベルの推定精度の向上を期待することができる。なお、キーワード記憶部121は、N個のキーワードしか記憶できないFIFO(First-In First-Out)型の記憶領域であってもよい。なお、対象キーワード群は、話者Aの発話内容のみから抽出しなくてもよい。例えば、直近のM個の対話における、話者Aと話者Bの発話内容からも抽出してもよい。この場合、図3のステップS101では、話者Aの発話音声のみならず、話者Bの発話音声が入力されるようにしてもよい。
対象キーワード群に含まれるキーワードが複数である場合、例えば、理解度推定部12は、対象キーワードごとに知識レベルを知識レベルDB122から取得し、取得された知識レベルの中の最低値を、話者Bの知識レベルとして推定してもよい。又は、理解度推定部12は、理解度が閾値未満であると推定されるより前にキーワード記憶部121に記録されたいずれかの対象キーワードに対応する知識レベルの中で、最高値を話者Bの知識レベルとして推定してもよい。理解度が閾値未満であると推定されるより前のキーワードについては、話者Bが理解していた可能性が高いからである。
なお、上記以外に、特開2013-167765に開示された技術が用いられてもよい。この場合、話者Aと話者Bとの対話の履歴が記録され、知識レベル推定部13は、当該履歴に基づいて話者Bの知識レベル(知識量)を推定してもよい。又は、特開2019-28604に開示された技術が用いられて話者Bの知識レベルが推定されてもよい。
続いて、質問文取得部14は、対象キーワード群及び話者Bについて推定された知識レベル群に基づいて、話者Aに対して出力すべき質問文を質問文DB123から取得する(S205)。
図6は、質問文DB123の構成例を示す図である。図6に示されるように、質問文DB123の各レコードには、「キーワード」及び「必要知識レベル」に対応付けて「質問文」及び「出力数」が記憶されている。各レコードの「質問文」は、当該レコードの「キーワード」を理解していない者が当該レコードの「必要知識レベル」以上の知識レベルを有している場合に出力されるべき質問文を示す。また、各レコードの「出力数」は、当該レコードの「質問文」が過去に出力された回数を示す。
したがって、ステップS205において、質問文取得部14は、対象キーワード群に含まれるいずれかのキーワードを「キーワード」に含むレコードであって、かつ、「必要知識レベル」が話者Bの知識レベル以下であるレコードの「質問文」を取得する。該当する「質問文」が複数有る場合、例えば、「出力数」の降順にソートされてもよい。
続いて、質問文出力部15は、質問文取得部14によって取得された質問文を表示装置106に出力(表示)する(S206)。表示装置106は、話者A及び話者Bが視認可能なように配置される。
その後、話者Aは、質問文に対する回答を発話する。当該質問文と当該回答とに基づいて、話者Bが、理解できなかった話者Aの発話内容を理解できるようになることが期待できる。
話者A及び話者Bの対話、並びに対話支援装置10が出力する質問文の具体例を以下に示す。
A(1)「家で無線LANは使っていますか?」
B(1)「はい。」
A(2)「外出中にはテザリングは使いますか?」
B(2)「ええっと。」
対話支援装置10「テザリングを利用するとノートパソコンでもインターネットは使えるようになりますか。」
A(3)「はい。」
B(3)「外ではノートパソコンを使わないので、テザリングは使っていないと思います。」
上記において、A(m)(m=1~3)は、話者Aによる発話を示す。B(m)(m=1~3)は、話者Bによる発話を示す。上記では、話者Bが、「ええっと。」と発話したときの表情に基づいて、ステップS202以降が実行され、その結果として実行されたステップS206において、対話支援装置10が、「テザリングを利用するとノートパソコンでもインターネットは使えるようになりますか。」という質問文を、話者Bに代わって話者Aに対して出力している。それに対する話者Aの回答(「はい。」)によって、話者Bは、「テザリング」の意味を完全には理解せずとも発話A(2)に対する回答(発話B(3)を行うことができ、二人の対話が円滑に行われている。すなわち、二人の対話がかみあったものとなり、当該対話が破綻するのが回避されている。
A(1)「家で無線LANは使っていますか?」
B(1)「はい。」
A(2)「外出中にはテザリングは使いますか?」
B(2)「ええっと。」
対話支援装置10「テザリングを利用するとノートパソコンでもインターネットは使えるようになりますか。」
A(3)「はい。」
B(3)「外ではノートパソコンを使わないので、テザリングは使っていないと思います。」
上記において、A(m)(m=1~3)は、話者Aによる発話を示す。B(m)(m=1~3)は、話者Bによる発話を示す。上記では、話者Bが、「ええっと。」と発話したときの表情に基づいて、ステップS202以降が実行され、その結果として実行されたステップS206において、対話支援装置10が、「テザリングを利用するとノートパソコンでもインターネットは使えるようになりますか。」という質問文を、話者Bに代わって話者Aに対して出力している。それに対する話者Aの回答(「はい。」)によって、話者Bは、「テザリング」の意味を完全には理解せずとも発話A(2)に対する回答(発話B(3)を行うことができ、二人の対話が円滑に行われている。すなわち、二人の対話がかみあったものとなり、当該対話が破綻するのが回避されている。
なお、図7は、具体例において対話支援装置10が実行する処理内容を示すである。図7では、発話A(1)において、キーワード抽出部11は、「無線LAN」をキーワードとして抽出している。続く発話B(1)のタイミングで、理解度推定部12は、話者Bの理解度を「高」として推定している。したがって、知識レベル推定部13及び質問文取得部14は処理を実行しない。続く発話A(2)において、キーワード抽出部11は、「テザリング」をキーワードとして抽出している。なお、この例では、キーワード記憶部121には4個までのキーワードが記憶される例が示されている。したがって、この時点において、キーワード記憶部121には、「無線LAN」及び「テザリング」の2つのキーワードが記憶されている。続く発話B(2)のタイミングで、理解度推定部12は、話者Bの理解度を「低」として推定している。それに応じ、知識レベル推定部13は、話者Bの知識レベルを「50」と推定し、質問文取得部14が、図7の(1)に示されたレコード群の中から(2)に示されたレコードを質問DBから取得する。(1)には、対象キーワード群(「無線LAN」、「テザリング」)のいずれかを「キーワード」に含むレコード群が示されている。(2)には、(1)のレコード群の中で、「必要知識レベル」が50以下であるレコードが示されている。
したがって、この場合、上記の具体例のように、質問文出力部15は、「テザリングを利用するとノートパソコンでもインターネットは使えるようになりますか。」という質問文を話者Bに代わって出力する。なお、本実施の形態では、質問文の出力形態が表示である例を示したが、例えば、質問文出力部15は、質問文を音声によって出力してもよい。この場合、対話支援装置10は、スピーカを有していればよい。
なお、図8に示されるように、話者Bに相当する者(例えば、話者Aに対する相談者)が複数人のグループ(図8における話者B1~話者BN)であるケースも想定される。このようなケースを想定して、話者Bに相当する話者の人数に対する閾値が対話支援装置10に設定され、質問文出力部15は、当該閾値を超える人数が居る場合には、質問文を出力しないようにしてもよい。そうすることで、例えば、各話者Bは、対話支援装置10による質問文の出力により、自分の知識レベルが低いことが他の話者Bによって推察されるのを回避することができる。
但し、複数人の話者Bが居る場合であっても、上記の閾値に基づく質問文の出力の制限が行われなくてもよい。この場合、理解度推定部12は、(並列的に)各話者Bの理解度を推定し、知識レベル推定部13は、(並列的に)各話者Bの知識レベルを推定してもよい。質問文取得部14は、(並列的に)推定された複数の知識レベルのうち、最も低い知識レベルに基づいて、話者Aに対して出力すべき質問文を質問文DB123から取得してもよい。そうすることで、知識レベルが最も低い話者Bに合わせて質問文の出力が行われるようにしてもよい。
上述したように、本実施の形態によれば、話者Aによる発話内容(話者Aとの対話内容)を話者Bが理解できない場合に、対話支援装置10が話者Bに代わって、話者Bの知識レベルに応じた質問文を話者Aに対して出力(通知)する。話者Aが当該質問文に対して回答することにより、話者Bは、理解できなかった発話内容を完全に理解しなくても当該回答に基づいて、当該発話内容に対する応答を行うことができる。したがって、対話の円滑化を支援することができる。
なお、本実施の携帯において、知識レベル推定部13は、第1の推定部の一例である。質問文取得部14は、取得部の一例である。質問文出力部15は、出力部の一例である。理解度推定部12は、第2の推定部の一例である。話者Aは、第1の話者の一例である。話者Bは、第2の話者の一例である。
以上、本発明の実施の形態について詳述したが、本発明は斯かる特定の実施形態に限定されるものではなく、請求の範囲に記載された本発明の要旨の範囲内において、種々の変形・変更が可能である。
10 対話支援装置
11 キーワード抽出部
12 理解度推定部
13 知識レベル推定部
14 質問文取得部
15 質問文出力部
100 ドライブ装置
101 記録媒体
102 補助記憶装置
103 メモリ装置
104 CPU
105 マイク
106 表示装置
107 及びカメラ
121 キーワード記憶部
122 知識レベルDB
123 質問文DB
B バス
11 キーワード抽出部
12 理解度推定部
13 知識レベル推定部
14 質問文取得部
15 質問文出力部
100 ドライブ装置
101 記録媒体
102 補助記憶装置
103 メモリ装置
104 CPU
105 マイク
106 表示装置
107 及びカメラ
121 キーワード記憶部
122 知識レベルDB
123 質問文DB
B バス
Claims (4)
- 第1の話者の発話内容に関連する分野に関して、前記第1の話者と対話を行う第2の話者の知識レベルを推定する第1の推定部と、
キーワード及び知識レベルに対応付けて質問文を記憶する記憶部から、前記発話内容に含まれるキーワードに対応する質問文であって、かつ、前記第2の話者の知識レベルに対応する質問文を取得する取得部と、
取得された質問文を前記第1の話者に対して出力する出力部と、
を有することを特徴とする対話支援装置。 - 前記第1の話者の発話内容に対する前記第2の話者の理解度を推定する第2の推定部を有し、
前記出力部は、前記理解度に応じて、前記質問文を前記第1の話者に対して出力する、
ことを特徴とする請求項1記載の対話支援装置。 - 第1の話者の発話内容に関連する分野に関して、前記第1の話者と対話を行う第2の話者の知識レベルを推定する第1の推定手順と、
キーワード及び知識レベルに対応付けて質問文を記憶する記憶部から、前記発話内容に含まれるキーワードに対応する質問文であって、かつ、前記第2の話者の知識レベルに対応する質問文を取得する取得手順と、
取得された質問文を前記第1の話者に対して出力する出力手順と、
コンピュータが実行することを特徴とする対話支援方法。 - 請求項1又は2記載の対話支援装置としてコンピュータを機能させることを特徴とするプログラム。
Priority Applications (3)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP2022505705A JP7435733B2 (ja) | 2020-03-13 | 2020-03-13 | 対話支援装置、対話支援方法及びプログラム |
| PCT/JP2020/011193 WO2021181679A1 (ja) | 2020-03-13 | 2020-03-13 | 対話支援装置、対話支援方法及びプログラム |
| US17/910,802 US20230121148A1 (en) | 2020-03-13 | 2020-03-13 | Dialog support apparatus, dialog support method and program |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/JP2020/011193 WO2021181679A1 (ja) | 2020-03-13 | 2020-03-13 | 対話支援装置、対話支援方法及びプログラム |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2021181679A1 true WO2021181679A1 (ja) | 2021-09-16 |
Family
ID=77672210
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/JP2020/011193 Ceased WO2021181679A1 (ja) | 2020-03-13 | 2020-03-13 | 対話支援装置、対話支援方法及びプログラム |
Country Status (3)
| Country | Link |
|---|---|
| US (1) | US20230121148A1 (ja) |
| JP (1) | JP7435733B2 (ja) |
| WO (1) | WO2021181679A1 (ja) |
Families Citing this family (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2023144896A1 (ja) * | 2022-01-25 | 2023-08-03 | Nttテクノクロス株式会社 | 情報処理装置、情報処理方法及びプログラム |
Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPH06162090A (ja) * | 1992-11-20 | 1994-06-10 | Nippon Telegr & Teleph Corp <Ntt> | 質問文生成装置 |
| JP2003208439A (ja) * | 2002-01-11 | 2003-07-25 | Seiko Epson Corp | 会話応答支援方法及びシステム、コンピュータプログラム |
| JP2014002470A (ja) * | 2012-06-15 | 2014-01-09 | Ricoh Co Ltd | 処理装置、処理システム、出力方法およびプログラム |
Family Cites Families (19)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20040162724A1 (en) * | 2003-02-11 | 2004-08-19 | Jeffrey Hill | Management of conversations |
| US20080140421A1 (en) * | 2006-12-07 | 2008-06-12 | Motorola, Inc. | Speaker Tracking-Based Automated Action Method and Apparatus |
| US8837706B2 (en) * | 2011-07-14 | 2014-09-16 | Intellisist, Inc. | Computer-implemented system and method for providing coaching to agents in an automated call center environment based on user traits |
| US9779260B1 (en) * | 2012-06-11 | 2017-10-03 | Dell Software Inc. | Aggregation and classification of secure data |
| US9495350B2 (en) * | 2012-09-14 | 2016-11-15 | Avaya Inc. | System and method for determining expertise through speech analytics |
| US10395552B2 (en) * | 2014-12-19 | 2019-08-27 | International Business Machines Corporation | Coaching a participant in a conversation |
| US10530929B2 (en) * | 2015-06-01 | 2020-01-07 | AffectLayer, Inc. | Modeling voice calls to improve an outcome of a call between a representative and a customer |
| US10586539B2 (en) * | 2015-06-01 | 2020-03-10 | AffectLayer, Inc. | In-call virtual assistant |
| US9697198B2 (en) * | 2015-10-05 | 2017-07-04 | International Business Machines Corporation | Guiding a conversation based on cognitive analytics |
| US10366160B2 (en) * | 2016-09-30 | 2019-07-30 | International Business Machines Corporation | Automatic generation and display of context, missing attributes and suggestions for context dependent questions in response to a mouse hover on a displayed term |
| US10102846B2 (en) * | 2016-10-31 | 2018-10-16 | International Business Machines Corporation | System, method and computer program product for assessing the capabilities of a conversation agent via black box testing |
| WO2018230345A1 (ja) | 2017-06-15 | 2018-12-20 | 株式会社Caiメディア | 対話ロボットおよび対話システム、並びに対話プログラム |
| US10769732B2 (en) * | 2017-09-19 | 2020-09-08 | International Business Machines Corporation | Expertise determination based on shared social media content |
| US11615144B2 (en) * | 2018-05-31 | 2023-03-28 | Microsoft Technology Licensing, Llc | Machine learning query session enhancement |
| US11017001B2 (en) * | 2018-12-31 | 2021-05-25 | Dish Network L.L.C. | Apparatus, systems and methods for providing conversational assistance |
| US11427216B2 (en) * | 2019-06-06 | 2022-08-30 | GM Global Technology Operations LLC | User activity-based customization of vehicle prompts |
| US11210677B2 (en) * | 2019-06-26 | 2021-12-28 | International Business Machines Corporation | Measuring the effectiveness of individual customer representative responses in historical chat transcripts |
| WO2021033886A1 (en) * | 2019-08-22 | 2021-02-25 | Samsung Electronics Co., Ltd. | A system and method for providing assistance in a live conversation |
| US11245793B2 (en) * | 2019-11-22 | 2022-02-08 | Genesys Telecommunications Laboratories, Inc. | System and method for managing a dialog between a contact center system and a user thereof |
-
2020
- 2020-03-13 JP JP2022505705A patent/JP7435733B2/ja active Active
- 2020-03-13 US US17/910,802 patent/US20230121148A1/en not_active Abandoned
- 2020-03-13 WO PCT/JP2020/011193 patent/WO2021181679A1/ja not_active Ceased
Patent Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPH06162090A (ja) * | 1992-11-20 | 1994-06-10 | Nippon Telegr & Teleph Corp <Ntt> | 質問文生成装置 |
| JP2003208439A (ja) * | 2002-01-11 | 2003-07-25 | Seiko Epson Corp | 会話応答支援方法及びシステム、コンピュータプログラム |
| JP2014002470A (ja) * | 2012-06-15 | 2014-01-09 | Ricoh Co Ltd | 処理装置、処理システム、出力方法およびプログラム |
Non-Patent Citations (1)
| Title |
|---|
| MIYAWAKI, KENZABURO ET AL.: "Adaptive Embodied Entrainment Control and Interaction Design of the First Meeting Introducer Robot", 2009 IEEE INTERNATIONAL CONFERENCE ON SYSTEMS, MAN, AND CYBERNETICS, 4 December 2009 (2009-12-04), pages 430 - 435, XP031574887 * |
Also Published As
| Publication number | Publication date |
|---|---|
| US20230121148A1 (en) | 2023-04-20 |
| JPWO2021181679A1 (ja) | 2021-09-16 |
| JP7435733B2 (ja) | 2024-02-21 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US11669683B2 (en) | Speech recognition and summarization | |
| JP4838351B2 (ja) | キーワード抽出装置 | |
| US20190080687A1 (en) | Learning-type interactive device | |
| JP7255032B2 (ja) | 音声認識 | |
| US12200322B2 (en) | Systems and methods for generating a video summary of a virtual event | |
| KR102414159B1 (ko) | 보류 상태를 관리하기 위한 방법 및 장치 | |
| US20140081643A1 (en) | System and method for determining expertise through speech analytics | |
| US20060020471A1 (en) | Method and apparatus for robustly locating user barge-ins in voice-activated command systems | |
| US20110264444A1 (en) | Keyword display system, keyword display method, and program | |
| US12499882B2 (en) | Low-latency conversational large language models | |
| US20240257811A1 (en) | System and Method for Providing Real-time Speech Recommendations During Verbal Communication | |
| JP2018197924A (ja) | 情報処理装置、対話処理方法、及び対話処理プログラム | |
| JP5084297B2 (ja) | 会話解析装置および会話解析プログラム | |
| ES2751375T3 (es) | Análisis lingüístico basado en una selección de palabras y dispositivo de análisis lingüístico | |
| JP7435733B2 (ja) | 対話支援装置、対話支援方法及びプログラム | |
| CN116034427B (zh) | 使用计算机的会话支持方法 | |
| JP2016062333A (ja) | 検索サーバ、及び検索方法 | |
| EP4139784B1 (en) | Hierarchical context specific actions from ambient speech | |
| JP6383748B2 (ja) | 音声翻訳装置、音声翻訳方法、及び音声翻訳プログラム | |
| JP2020149073A (ja) | 通信システム、通信方法、サーバ装置、及びプログラム | |
| JP6962849B2 (ja) | 会議支援装置、会議支援制御方法およびプログラム | |
| JP7689787B1 (ja) | 情報処理装置、情報処理方法及びプログラム | |
| JP4760452B2 (ja) | 発話訓練装置、発話訓練システム、発話訓練支援方法およびプログラム | |
| US20250348521A1 (en) | Memory Assistant System | |
| US20260073921A1 (en) | Semiconductor manufacturing apparatus and support method |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 20924404 Country of ref document: EP Kind code of ref document: A1 |
|
| ENP | Entry into the national phase |
Ref document number: 2022505705 Country of ref document: JP Kind code of ref document: A |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 20924404 Country of ref document: EP Kind code of ref document: A1 |