WO2018015989A1 - 音声認識システム、音声認識方法及びプログラム - Google Patents

音声認識システム、音声認識方法及びプログラム Download PDF

Info

Publication number
WO2018015989A1
WO2018015989A1 PCT/JP2016/071110 JP2016071110W WO2018015989A1 WO 2018015989 A1 WO2018015989 A1 WO 2018015989A1 JP 2016071110 W JP2016071110 W JP 2016071110W WO 2018015989 A1 WO2018015989 A1 WO 2018015989A1
Authority
WO
WIPO (PCT)
Prior art keywords
text
voice
speech recognition
speech
selection
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/JP2016/071110
Other languages
English (en)
French (fr)
Inventor
俊二 菅谷
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Optim Corp
Original Assignee
Optim Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Optim Corp filed Critical Optim Corp
Priority to PCT/JP2016/071110 priority Critical patent/WO2018015989A1/ja
Publication of WO2018015989A1 publication Critical patent/WO2018015989A1/ja
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/22Procedures used during a speech recognition process, e.g. man-machine dialogue

Definitions

  • the present invention relates to a speech recognition system, a speech recognition method, and a program for converting speech into text.
  • a user inputs voice to a sound collecting device such as a microphone, and converts the inputted voice into text.
  • a sound collecting device such as a microphone
  • voice-to-text conversion is used, for example, for document input by a user, creation of minutes for meetings, and the like.
  • Patent Document 1 it is difficult to check whether or not the input voice matches the converted text after the voice is converted to text.
  • the present invention provides the following solutions.
  • the invention according to the first feature comprises sound collecting means for collecting sound; Voice recognition means for recognizing the collected voice; Text conversion means for converting the recognized speech into text; Display means for displaying the converted text; Partial selection receiving means for receiving selection of a desired portion from the displayed text; Voice output means for outputting voice corresponding to a portion of the text that has received the selection;
  • a speech recognition system characterized by comprising:
  • voice is collected, the collected voice is recognized, the recognized voice is converted into text, the converted text is displayed, and the desired text is displayed from the displayed text.
  • the selection is accepted for the part to be selected, and the voice corresponding to the part of the text for which the selection is accepted is output.
  • the invention according to the first feature is a category of the speech recognition system, but the same operation and effect corresponding to the category are exhibited in other categories such as a method or a program.
  • the invention according to the second feature is a text change accepting means for accepting a change of a part of the text that accepts the selection based on the result of the output voice;
  • a speech recognition system that is an invention according to a first feature is provided.
  • the speech recognition system that is the invention relating to the first feature accepts a change of a part of the text that has received the selection based on the result of the outputted speech.
  • the invention according to a third aspect is characterized in that the storage unit stores the text that has received the change in association with the voice that has been voice-recognized earlier.
  • the speech recognition system which is invention which concerns on the 2nd characteristic characterized by providing is provided.
  • the speech recognition system that is the invention relating to the second feature stores the text that has received the change in association with the speech that has been speech-recognized earlier.
  • the invention according to a fourth feature is the accuracy display means for displaying the accuracy rate of the converted text;
  • a speech recognition system that is an invention according to a first feature is provided.
  • the speech recognition system displays the accuracy rate of the converted text.
  • the invention includes a step of collecting sound; Recognizing the collected sound; Converting the recognized speech into text; Displaying the converted text; Accepting a selection for a desired portion from the displayed text; Outputting speech corresponding to a portion of the text that has received the selection;
  • a speech recognition method characterized by comprising:
  • the invention according to the sixth feature provides a speech recognition system comprising: Collecting sound, Recognizing the collected sound; Converting the recognized speech into text; Displaying the converted text; Accepting selection of a desired portion from the displayed text; Outputting speech corresponding to a portion of the text that has received the selection; Provide a program to execute.
  • the voice recognition enables the user to check by listening to the voice whether or not a part of the desired text has been wrongly converted with respect to the result of the voice converted into the text. It is an object to provide a system, a speech recognition method, and a program.
  • FIG. 1 is a diagram showing an outline of the speech recognition system 1.
  • FIG. 2 is an overall configuration diagram of the voice recognition system 1.
  • FIG. 3 is a functional block diagram of the user terminal 100 and the display terminal 200.
  • FIG. 4 is a diagram illustrating a voice recognition process executed by the user terminal 100 and the display terminal 200.
  • FIG. 5 is a diagram showing a voice recognition process executed by the user terminal 100 and the display terminal 200.
  • FIG. 6 is a diagram showing a voice recognition process executed by the user terminal 100 and the display terminal 200.
  • FIG. 7 is a diagram illustrating an example of a text display screen displayed on the display terminal 200.
  • FIG. 8 is a diagram illustrating an example of a text display screen displayed on the display terminal 200.
  • FIG. 9 is a diagram illustrating an example of a text display screen displayed on the display terminal 200.
  • FIG. 10 is a diagram illustrating an example of a text display screen displayed on the display terminal 200.
  • FIG. 1 is a diagram for explaining an outline of a speech recognition system 1 which is a preferred embodiment of the present invention.
  • the voice recognition system 1 includes a user terminal 100 and a display terminal 200.
  • the number of user terminals 100 or display terminals 200 is not limited to one and may be plural. Further, the user terminal 100 or the display terminal 200 is not limited to a real device, and may be a virtual device. Moreover, each process mentioned later may be implement
  • the user terminal 100 is a terminal device capable of data communication with the display terminal 200.
  • the user terminal 100 is, for example, a cellular phone, a portable information terminal, a tablet terminal, a personal computer, an electric appliance such as a netbook terminal, a slate terminal, an electronic book terminal, a portable music player, a smart glass, a head mounted display, or the like Wearable terminals and other items.
  • the display terminal 200 is a terminal device capable of data communication with the user terminal 100.
  • the display terminal 200 is a terminal device similar to the user terminal 100.
  • the user terminal 100 collects voice input from the user by a sound collecting device such as a microphone (step S01).
  • the user terminal 100 recognizes the collected voice (step S02).
  • the user terminal 100 recognizes the content of the voice uttered by the user by voice recognition.
  • the user terminal 100 converts the recognized voice into text (step S03).
  • the user terminal 100 transmits the recognized voice data and text data to the display terminal 200 (step S04).
  • the display terminal 200 receives the voice data and the text data, and displays the text on its display unit based on the text data (step S05).
  • voice data with a text may be sufficient.
  • the display terminal 200 receives a selection from the user for a part of the displayed text, and outputs a voice corresponding to the part of the text for which the selection has been received (step S06).
  • the structure corresponding to all the texts to display may be output not only for the part of text which received selection.
  • the display terminal 200 may be configured to accept a change from the user with respect to a part of the text for which selection has been accepted based on the output voice result. In this case, what is necessary is just the structure which displays the text which the user changed by receiving the change from a user. Further, the display terminal 200 may be configured to store the text that has received this change in association with the voice that was previously voice-recognized by the user terminal 100.
  • FIG. 2 is a diagram showing a system configuration of the speech recognition system 1 which is a preferred embodiment of the present invention.
  • the speech recognition system 1 includes a user terminal 100, a display terminal 200, and a public line network (Internet network, third and fourth generation communication network, etc.) 5.
  • the user terminal 100 or the display terminal 200 is not limited to one, and may be a plurality.
  • the user terminal 100 or the display terminal 200 is not limited to a real device, and may be a virtual device.
  • each process mentioned later may be implement
  • the user terminal 100 is the above-described terminal device having the functions described later.
  • the display terminal 200 is the above-described terminal device having the functions described below.
  • FIG. 3 is a functional block diagram of the user terminal 100 and the display terminal 200.
  • the user terminal 100 includes a CPU (Central Processing Unit), a RAM (Random Access Memory), a ROM (Read Only Memory), etc. as the control unit 110, and a device for enabling communication with other devices as the communication unit 120.
  • a WiFi (Wireless Fidelity) compatible device compliant with IEEE 802.11 is provided.
  • the user terminal 100 includes a data storage unit such as a hard disk, a semiconductor memory, a recording medium, or a memory card as the storage unit 130.
  • the user terminal 100 serves as an input / output unit 140 that outputs and displays data and images controlled by the control unit 110, an input unit such as a touch panel, a keyboard, and a mouse that receives input from the user, and collects audio.
  • a sound collection device such as a sound microphone, a conversion device that converts voice into text, a probability determination device that determines the accuracy rate of text, and the like are provided.
  • the control unit 110 when the control unit 110 reads a predetermined program, the data transmission / reception module 150 is realized in cooperation with the communication unit 120. Further, in the user terminal 100, the storage module 160 is realized in cooperation with the storage unit 130 by the control unit 110 reading a predetermined program. In the user terminal 100, the control unit 110 reads a predetermined program, thereby realizing the sound collection module 170, the voice recognition module 171, and the conversion module 172 in cooperation with the input / output unit 140.
  • the display terminal 200 includes a CPU, a RAM, a ROM, and the like as the control unit 210, and a WiFi compatible device for enabling communication with other devices as the communication unit 220, and a storage unit 230 includes a data storage unit, and the input / output unit 240 includes a display unit, an input unit, an accuracy determination device, and the like.
  • the control unit 210 when the control unit 210 reads a predetermined program, the data transmission / reception module 250 is realized in cooperation with the communication unit 220. In the display terminal 200, the control unit 210 reads a predetermined program, thereby realizing the storage module 260 in cooperation with the storage unit 230. In the display terminal 200, the control unit 210 reads a predetermined program, thereby realizing the input receiving module 270, the display module 271, the audio output module 272, and the conversion module 273 in cooperation with the input / output unit 240.
  • FIGS. 4, 5, and 6 are flowcharts of voice recognition processing executed by the user terminal 100 and the display terminal 200. The processing executed by the modules of each device described above will be described together with this processing.
  • the sound collection module 170 collects the voice uttered by the user and accepts the input of the voice from the user (step S10). For example, the sound collection module 170 collects sound such as sentences and phrases uttered by the user.
  • the voice recognition module 171 recognizes the voice that has received the input (step S11). In step S11, the voice recognition module 171 recognizes the content of the voice that has been accepted.
  • the speech recognition module 171 recognizes speech using, for example, an existing speech recognition technology or pattern recognition technology.
  • the speech recognition module 171 is not limited to the configuration described above, and may perform speech recognition using other configurations.
  • the conversion module 172 converts the speech-recognized content into text (step S12).
  • the conversion module 172 is, for example, the contents of the speech recognition, the case is a "Hello, my name is hot", as text, to convert the "Hello. It's hot.”.
  • the conversion module 172 converts the speech-recognized content into text based on, for example, the relationship between before and after, the segment break position, and the speech inflection.
  • Step S12 when there are a plurality of converted text candidates, the conversion module 172 converts each of the plurality of texts.
  • conversion module 172 is not limited to the configuration described above, and may be converted into text by another configuration.
  • the conversion module 172 calculates the accuracy of the converted text (step S13).
  • step S13 when there are a plurality of candidates in the converted text, the conversion module 172 determines each of the plurality of texts based on the relationship between before and after, phrase break position, speech inflection, vocabulary appearance frequency, and the like. Calculate the accuracy.
  • the accuracy of the text is the accuracy rate between the speech-recognized content and the converted text. For example, a vocabulary with a high frequency of appearance has high accuracy, and a vocabulary with a low frequency of appearance has a low accuracy, and the accuracy is numerically set for each vocabulary and phrase, and the accuracy is based on this value. It is the structure which calculates.
  • the conversion module 172 calculates the accuracy for each sentence, phrase, or word.
  • the conversion module 172 is not limited to the configuration described above, and the accuracy may be calculated by another configuration. Further, the conversion module 172 may calculate the accuracy even when a plurality of candidates do not exist.
  • the storage module 160 stores conversion information that correlates the input voice, the converted text, and the accuracy of the text (step S14). When there are a plurality of converted texts, the storage module 160 stores each of them. That is, conversion information in which one voice, a plurality of converted texts, and the accuracy of each text are associated with each other is stored.
  • the storage module 160 stores conversion information in which one voice, one converted text, and the accuracy of the text are associated with each other when a plurality of converted texts do not exist.
  • step S14 can be omitted. That is, the user terminal 100 may be configured to execute processing to be described later without storing the conversion information.
  • the data transmission / reception module 150 transmits the conversion information to the display terminal 200 (step S15).
  • the data transmission / reception module 250 receives the conversion information.
  • the input receiving module 270 determines whether or not an input of the accuracy of changing the text has been received (step S16). In step S ⁇ b> 16, the input receiving module 270 determines whether or not an input of accuracy serving as a reference for receiving a text change input to be described later is received. For example, the input reception module 270 determines whether an input of a numerical value such as 50% or 60% has been received as the accuracy input.
  • the numerical value of the accuracy is not limited to% notation and can be changed as appropriate, such as a ratio and a numerical value.
  • step S16 when the input reception module 270 determines that the input is not received (step S16 NO), the display module 271 displays the text and the accuracy based on the received conversion information (step S17).
  • FIG. 7 is a diagram showing an example of a text display screen displayed by the display module 271 in step S17.
  • the display module 271 displays a text display area 300 and a selection icon 310.
  • the text display area 300 is an area for displaying text based on the conversion information.
  • the text display area 300 displays the text divided into phrases or sentences.
  • the display module 271 the text display area 300, and "Hello," "Oh, I am with” to display the ⁇ people separated by.
  • the selection icon 310 accepts input from the user and accepts selection input to the corresponding text.
  • the voice output module 272 outputs the text that has received the input to the selection icon 310 based on the conversion information.
  • the display module 271 displays a selection icon 310 on each text regardless of the text accuracy.
  • step S16 determines in step S16 that the input has been received (step S16 YES)
  • the display module 271 displays the text and the accuracy based on the received conversion information (step S18).
  • FIG. 8 is a diagram showing an example of a text display screen displayed by the display module 271 in step S18.
  • FIG. 8 shows an example when an input with a predetermined accuracy of 50% is received.
  • the display module 271 displays a text display area 300 and a selection icon 310 as in FIG.
  • the difference between FIG. 7 and FIG. 8 is that the display module 271 displays the selection icon 310 for the text of all the accuracy in FIG. 7, whereas in FIG.
  • the difference is that the selection icon 310 is displayed only for the text of accuracy. That is, in FIG.
  • the display module 271 the accuracy is higher than a predetermined value "95%” for "Hello” is a non-display selection icon 310, accuracy is lower than a predetermined value " For “hot” which is “40%”, a selection icon 310 is displayed.
  • the input reception module 270 determines whether or not a text selection input has been received (step S19). In step S19, the input reception module 270 determines, for example, whether an input operation to the selection icon 310 described above has been received, whether a tap operation has been received on the text to be displayed, or the like.
  • step S19 when the input reception module 270 determines that the input is not received (step S19: NO), the storage module 260 stores the conversion information (step S20). In step S20, the storage module 260 stores the voice, the displayed text, and the accuracy in association with each other. The display terminal 200 ends this process after performing the process of step S20.
  • the voice output module 272 outputs the selected text as a voice based on the conversion information (step S21).
  • the input receiving module 270 determines whether or not a text change input has been received (step S22). In step S ⁇ b> 22, the input reception module 270 determines whether an input for changing to a text different from the text displayed by the display module 271 is received based on direct input, voice input, gesture input, or the like.
  • step S22 when the input receiving module 270 determines that no text change input is received (NO in step S22), the input receiving module 270 executes the process of step S20 described above.
  • step S22 when the input receiving module 270 determines in step S22 that the text change input has been received (YES in step S22), the conversion module 273 calculates the accuracy of the text after the change (step S23). In step S23, the conversion module 273 calculates the accuracy of the text after the change by the same process as the process of step S13 described above.
  • the conversion module 273 determines whether or not the accuracy is equal to or higher than a predetermined value (step S24). In step S24, the conversion module 273 determines whether or not it is greater than or equal to a preset value, whether or not it is greater than or equal to the value received in step S16 described above. In step S24, when the conversion module 273 determines that the value is equal to or greater than the predetermined value (YES in step S24), the display module 271 displays the changed text and its accuracy (step S25). In step S25, the display module 271 may be configured to display only the changed text.
  • FIG. 9 is a diagram showing an example of the text display screen after the change displayed by the display module 271.
  • the display module 271 displays a text display area 300.
  • the display module 271 displays in the text display area 300 the text that has not been changed, the changed text, and the accuracy of each. That is, in FIG. 9, the text of "Hello” is a text that has not been changed, "Oh, I am with (already changed)" is a modified text.
  • the display module 271 displays the changed text with a message indicating that the text has been changed.
  • the display module 271 may be configured to display only the changed text in the text display area 300. Further, the display module 271 may be configured to display the changed text without giving a message indicating that the change has been made.
  • the storage module 260 stores the post-change conversion information in which the voice included in the conversion information, the post-change text, and the accuracy are associated (step S26).
  • step S ⁇ b> 26 the storage module 260 stores not only the post-change conversion information but also the voice of the text that has not been changed, this text, and the accuracy. That is, in step S26, the storage module 260 stores all the text displayed by the display module 271 in step S25, the voice of the text, and the accuracy in association with each other.
  • the display terminal 200 ends this process after performing the process of step S26.
  • step S24 determines in step S24 that the value is not equal to or greater than the predetermined value (NO in step S24)
  • the display module 271 displays the changed text and a notification of a different conversion candidate (step S27). ).
  • FIG. 10 is a diagram showing an example of the text display screen after the change displayed by the display module 271.
  • FIG. 10 shows an example when the accuracy is not 50% or more.
  • the display module 271 displays a text display area 300, a selection icon 310, and a conversion candidate notification display area 400.
  • the display module 271 displays the text that has not been changed, the changed text, and the respective accuracy in the text display area 300, as in FIG.
  • the display module 271 displays a notification indicating that the accuracy is not equal to or higher than a predetermined value in the conversion candidate notification display area 400.
  • the display module 271 displays the conversion candidate text with the highest accuracy in the conversion candidate notification display area 400 and the accuracy of the conversion candidate text. That is, in FIG.
  • the display mode of this notification can be changed as appropriate, and only the conversion candidate text may be displayed, or a plurality of conversion candidate texts with a predetermined accuracy or higher may be displayed. Further, the accuracy may not be displayed.
  • the input reception module 270 determines whether or not a text change input has been received (step S28).
  • the process of step S28 is the same as the process of step S22 described above.
  • step S28 when the input reception module 270 determines that the input is not received (step S28: NO), the process of step S26 described above is executed.
  • step S28 determines in step S28 that the input has been received (step S28: YES)
  • step S23 the process of step S23 described above is executed.
  • the user terminal 100 recognizes voice
  • the display terminal 200 may recognize voice and convert text.
  • the configuration may be such that the audio received by the user terminal 100 is transmitted as audio data to the display terminal 200 and the display terminal 200 executes processing such as audio recognition.
  • the server and other configurations may be configured to execute voice recognition, text conversion, accuracy calculation, and the like.
  • the accuracy may be calculated based on the changed text and voice stored in the user terminal 100 or the display terminal 200, the text and voice before the change, and the like.
  • a configuration may be used in which conversion candidates for text to be changed this time are determined based on data of text that has been changed before or text that has not been changed.
  • the means and functions described above are realized by a computer (including a CPU, an information processing apparatus, and various terminals) reading and executing a predetermined program.
  • the program is provided in a form recorded on a computer-readable recording medium such as a flexible disk, CD (CD-ROM, etc.), DVD (DVD-ROM, DVD-RAM, etc.).
  • the computer reads the program from the recording medium, transfers it to the internal storage device or the external storage device, stores it, and executes it.
  • the program may be recorded in advance in a storage device (recording medium) such as a magnetic disk, an optical disk, or a magneto-optical disk, and provided from the storage device to a computer via a communication line.

Landscapes

  • Engineering & Computer Science (AREA)
  • Computational Linguistics (AREA)
  • Health & Medical Sciences (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Human Computer Interaction (AREA)
  • Physics & Mathematics (AREA)
  • Acoustics & Sound (AREA)
  • Multimedia (AREA)
  • Telephonic Communication Services (AREA)
  • User Interface Of Digital Computer (AREA)

Abstract

【課題】音声がテキストに変換された結果に対して、所望するテキストの一部分が間違った変換がなされていないか否かを、ユーザが音声を聞いてチェックすることが可能な音声認識システム、音声認識方法及びプログラムを提供することを目的とする。 【解決手段】音声を集音し、前記集音した音声を認識し、前記認識した音声をテキストに変換し、前記変換したテキストを表示し、前記表示したテキストから所望する一部分について、選択を受け付け、前記選択を受け付けた一部分のテキストに対応する音声を出力する。

Description

音声認識システム、音声認識方法及びプログラム
 本発明は、音声をテキストに変換する音声認識システム、音声認識方法及びプログラムに関する。
 従来、ユーザがマイク等の集音装置へ音声を入力し、この入力した音声をテキストに変換することが行われている。このような音声のテキスト変換は、例えば、ユーザによる文書入力、会議等の議事録作成等に用いられている。
 このような議事録の作成として、会議の出席者の音声をマイクにより集音し、この音声の音声データをテキストに変換し、議事録として記録する構成が開示されている(特許文献1参照)。
特開2013-222347号公報
 しかしながら、特許文献1の構成では、音声をテキストに変換した後に、入力された音声と変換したテキストとが一致しているか否かを、チェックすることが困難であった。
 本発明の目的は、音声がテキストに変換された結果に対して、所望するテキストの一部分が間違った変換がなされていないか否かを、ユーザが音声を聞いてチェックすることが可能な音声認識システム、音声認識方法及びプログラムを提供することを目的とする。
 本発明では、以下のような解決手段を提供する。
 第1の特徴に係る発明は、音声を集音する集音手段と、
 前記集音した音声を認識する音声認識手段と、
 前記認識した音声をテキストに変換するテキスト変換手段と、
 前記変換したテキストを表示する表示手段と、
 前記表示したテキストから所望する一部分について、選択を受け付ける部分選択受付手段と、
 前記選択を受け付けた一部分のテキストに対応する音声を出力する音声出力手段と、
 を備えることを特徴とする音声認識システムを提供する。
 第1の特徴に係る発明によれば、音声を集音し、前記集音した音声を認識し、前記認識した音声をテキストに変換し、前記変換したテキストを表示し、前記表示したテキストから所望する一部分について、選択を受け付け、前記選択を受け付けた一部分のテキストに対応する音声を出力する。
 ここで、第1の特徴に係る発明は、音声認識システムのカテゴリであるが、方法又はプログラム等の他のカテゴリにおいても、そのカテゴリに応じた同様の作用・効果を発揮する。
 第2の特徴に係る発明は、前記選択を受け付けた一部分のテキストを、前記出力した音声の結果に基づいて、変更を受け付けるテキスト変更受付手段と、
 を備えることを特徴とする第1の特徴に係る発明である音声認識システムを提供する。
 第2の特徴に係る発明によれば、第1の特徴に係る発明である音声認識システムは、前記選択を受け付けた一部分のテキストを、前記出力した音声の結果に基づいて、変更を受け付ける。
 第3の特徴に係る発明は、前記変更を受け付けたテキストを、先に音声認識した音声と対応付けて記憶する記憶手段と、
 を備えることを特徴とする第2の特徴に係る発明である音声認識システムを提供する。
 第3の特徴に係る発明によれば、第2の特徴に係る発明である音声認識システムは、前記変更を受け付けたテキストを、先に音声認識した音声と対応付けて記憶する。
 第4の特徴に係る発明は、前記変換したテキストの正解率を表示する確度表示手段と、
 を備えることを特徴とする第1の特徴に係る発明である音声認識システムを提供する。
 第4の特徴に係る発明によれば、第1の特徴に係る発明である音声認識システムは、前記変換したテキストの正解率を表示する。
 第5の特徴に係る発明は、音声を集音するステップと、
 前記集音した音声を認識するステップと、
 前記認識した音声をテキストに変換するステップと、
 前記変換したテキストを表示するステップと、
 前記表示したテキストから所望する一部分について、選択を受け付けるステップと、
 前記選択を受け付けた一部分のテキストに対応する音声を出力するステップと、
 を備えることを特徴とする音声認識方法を提供する。
 第6の特徴に係る発明は、音声認識システムに、
 音声を集音するステップ、
 前記集音した音声を認識するステップ、
 前記認識した音声をテキストに変換するステップ、
 前記変換したテキストを表示するステップ、
 前記表示したテキストから所望する一部分について、選択を受け付けるステップ、
 前記選択を受け付けた一部分のテキストに対応する音声を出力するステップ、
 を実行させるためのプログラムを提供する。
 本発明によれば、音声がテキストに変換された結果に対して、所望するテキストの一部分が間違った変換がなされていないか否かを、ユーザが音声を聞いてチェックすることが可能な音声認識システム、音声認識方法及びプログラムを提供することを目的とする。
図1は、音声認識システム1の概要を示す図である。 図2は、音声認識システム1の全体構成図である。 図3は、ユーザ端末100、表示端末200の機能ブロック図である。 図4は、ユーザ端末100、表示端末200が実行する音声認識処理を示す図である。 図5は、ユーザ端末100、表示端末200が実行する音声認識処理を示す図である。 図6は、ユーザ端末100、表示端末200が実行する音声認識処理を示す図である。 図7は、表示端末200が表示するテキスト表示画面の一例を示す図である。 図8は、表示端末200が表示するテキスト表示画面の一例を示す図である。 図9は、表示端末200が表示するテキスト表示画面の一例を示す図である。 図10は、表示端末200が表示するテキスト表示画面の一例を示す図である。
 以下、本発明を実施するための最良の形態について、図を参照しながら説明する。なお、これはあくまでも一例であって、本発明の技術的範囲はこれに限られるものではない。
 [音声認識システム1の概要]
 本発明の好適な実施形態の概要について、図1に基づいて説明する。図1は、本発明の好適な実施形態である音声認識システム1の概要を説明するための図である。音声認識システム1は、ユーザ端末100、表示端末200から構成される。
 なお、図1において、ユーザ端末100又は表示端末200は、1つに限らず複数であってもよい。また、ユーザ端末100又は表示端末200は、実在する装置に限らず仮想的な装置であってもよい。また、後述する各処理は、ユーザ端末100又は表示端末200のいずれか又は双方により実現されてもよい。
 ユーザ端末100は、表示端末200とデータ通信可能な端末装置である。ユーザ端末100は、例えば、携帯電話、携帯情報端末、タブレット端末、パーソナルコンピュータに加え、ネットブック端末、スレート端末、電子書籍端末、携帯型音楽プレーヤ等の電化製品や、スマートグラス、ヘッドマウントディスプレイ等のウェアラブル端末や、その他の物品である。
 表示端末200は、ユーザ端末100とデータ通信可能な端末装置である。表示端末200は、ユーザ端末100と同様の端末装置である。
 ユーザ端末100は、ユーザからの音声入力を、マイク等の集音装置により集音する(ステップS01)。
 ユーザ端末100は、集音した音声を音声認識する(ステップS02)。ユーザ端末100は、音声認識することにより、ユーザが発した音声の内容を認識する。
 ユーザ端末100は、認識した音声を、テキストに変換する(ステップS03)。
 ユーザ端末100は、認識した音声の音声データと、テキストのテキストデータとを、表示端末200に送信する(ステップS04)。
 表示端末200は、音声データとテキストデータとを受信し、このテキストデータに基づいて、テキストを自身の表示部に表示する(ステップS05)。なお、音声データを変換したテキストの正解率をテキストとともに表示する構成であってもよい。
 表示端末200は、表示したテキストの一部分に対して、ユーザからの選択を受け付け、選択を受け付けた一部分のテキストに対応する音声を出力する(ステップS06)。なお、選択を受け付けた一部分のテキストに限らず、表示する全てのテキストに対応する音声を出力する構成であってもよい。
 また、表示端末200は、出力した音声の結果に基づいて、選択を受け付けた一部分のテキストに対して、ユーザからの変更を受け付ける構成であってもよい。この場合、ユーザからの変更を受け付けることにより、ユーザが変更したテキストを、表示する構成であればよい。また、表示端末200は、この変更を受け付けたテキストを、先にユーザ端末100により音声認識した音声と対応付けて記憶する構成であってもよい。
 以上が、音声認識システム1の概要である。
 [音声認識システム1のシステム構成]
 図2に基づいて、本発明の好適な実施形態である音声認識システム1のシステム構成について説明する。図2は、本発明の好適な実施形態である音声認識システム1のシステム構成を示す図である。音声認識システム1は、ユーザ端末100、表示端末200、公衆回線網(インターネット網や、第3、第4世代通信網等)5から構成される。なお、ユーザ端末100又は表示端末200は、1つに限らず、複数であってもよい。また、ユーザ端末100又は表示端末200は、実在する装置に限らず、仮想的な装置であってもよい。また、後述する各処理は、ユーザ端末100又は表示端末200のいずれか又は双方により実現されてもよい。
 ユーザ端末100は、後述の機能を備えた上述した端末装置である。
 表示端末200は、後述の機能を備えた上述した端末装置である。
 [各機能の説明]
 図3に基づいて、本発明の好適な実施形態である音声認識システム1の機能について説明する。図3は、ユーザ端末100、表示端末200の機能ブロック図を示す図である。
 ユーザ端末100は、制御部110として、CPU(Central Processing Unit)、RAM(Random Access Memory)、ROM(Read Only Memory)等を備え、通信部120として、他の機器と通信可能にするためのデバイス、例えば、IEEE802.11に準拠したWiFi(Wireless Fidelity)対応デバイスを備える。また、ユーザ端末100は、記憶部130として、ハードディスクや半導体メモリ、記録媒体、メモリカード等によるデータのストレージ部を備える。また、ユーザ端末100は、入出力部140として、制御部110で制御したデータや画像を出力表示する表示部や、ユーザからの入力を受け付けるタッチパネルやキーボード、マウス等の入力部や、音声を集音するマイク等の集音デバイスや、音声をテキスト変換する変換デバイスや、テキストの正解率を判断する確度判断デバイス等を備える。
 ユーザ端末100において、制御部110が所定のプログラムを読み込むことにより、通信部120と協働して、データ送受信モジュール150を実現する。また、ユーザ端末100において、制御部110が所定のプログラムを読み込むことにより、記憶部130と協働して、記憶モジュール160を実現する。また、ユーザ端末100において、制御部110が所定のプログラムを読み込むことにより、入出力部140と協働して、集音モジュール170、音声認識モジュール171、変換モジュール172を実現する。
 表示端末200は、ユーザ端末100と同様に、制御部210として、CPU、RAM、ROM等を備え、通信部220として、他の機器と通信可能にするためのWiFi対応デバイス等を備え、記憶部230として、データのストレージ部を備え、入出力部240として、表示部や入力部や確度判断デバイス等を備える。
 表示端末200において、制御部210が所定のプログラムを読み込むことにより、通信部220と協働して、データ送受信モジュール250を実現する。また、表示端末200において、制御部210が所定のプログラムを読み込むことにより、記憶部230と協働して、記憶モジュール260を実現する。また、表示端末200において、制御部210が所定のプログラムを読み込むことにより、入出力部240と協働して、入力受付モジュール270、表示モジュール271、音声出力モジュール272、変換モジュール273を実現する。
 [音声認識処理]
 図4、図5及び図6に基づいて、音声認識システム1が実行する音声認識処理について説明する。図4、図5及び図6は、ユーザ端末100、表示端末200が実行する音声認識処理のフローチャートを示す図である。上述した各装置のモジュールが実行する処理について、本処理に併せて説明する。
 集音モジュール170は、ユーザが発した音声を集音し、ユーザからの音声の入力を受け付ける(ステップS10)。集音モジュール170は、例えば、ユーザが発した文章や文節等の音声を集音する。
 音声認識モジュール171は、入力を受け付けた音声を、音声認識する(ステップS11)。ステップS11において、音声認識モジュール171は、入力を受け付けた音声の内容を音声認識する。音声認識モジュール171は、例えば、既存の音声認識技術やパターン認識技術等により、音声を音声認識する。
 なお、音声認識モジュール171は、上述した構成に限らず、他の構成により音声認識を行ってもよい。
 変換モジュール172は、音声認識した内容を、テキストに変換する(ステップS12)。ステップS12において、変換モジュール172は、例えば、音声認識した内容が、「こんにちはあついです」である場合、テキストとして、「こんにちは。暑いですね。」と変換する。変換モジュール172は、テキストの変換として、例えば、前後の関係、文節の区切り位置、音声の抑揚等に基づいて、音声認識した内容をテキストに変換する。また、ステップS12において、変換モジュール172は、変換したテキストの候補が複数ある場合、複数のテキストを其々変換する。
 なお、変換モジュール172は、上述した構成に限らず、他の構成によりテキストに変換してもよい。
 変換モジュール172は、変換したテキストの確度を算出する(ステップS13)。ステップS13において、変換モジュール172は、変換したテキストに複数の候補が存在する場合、前後の関係、文節の区切り位置、音声の抑揚、語彙の登場頻度等に基づいて、複数のテキストにおける其々の確度を算出する。テキストの確度とは、音声認識した内容と、変換したテキストとの正解率である。例えば、登場頻度の多い語彙は、確度が高く、登場頻度の少ない語彙は、確度が低くなるといった構成や、各語彙や文節毎に確度を数値化して設定しておき、この数値に基づいて確度を算出する構成である。変換モジュール172は、文章、文節又は単語の其々に対して確度を算出する。
 なお、変換モジュール172は、上述した構成に限らず、他の構成により確度を算出してもよい。また、変換モジュール172は、複数の候補が存在しない場合であっても、確度を算出してもよい。
 記憶モジュール160は、入力を受け付けた音声と、この音声を変換したテキストと、このテキストの確度とを対応付けた変換情報を記憶する(ステップS14)。記憶モジュール160は、変換したテキストが複数存在する場合、其々記憶する。すなわち、一の音声と、複数の変換テキストと、各テキストの確度とを対応付けた変換情報を記憶する。
 なお、記憶モジュール160は、変換したテキストが複数存在しない場合、一の音声と、一の変換したテキストと、このテキストの確度とを対応付けた変換情報を記憶する。
 また、ステップS14の処理は、省略可能である。すなわち、ユーザ端末100は、変換情報を記憶せずに、後述する処理を実行する構成であってもよい。
 データ送受信モジュール150は、変換情報を表示端末200に送信する(ステップS15)。
 データ送受信モジュール250は、変換情報を受信する。入力受付モジュール270は、テキストを変更する確度の入力を受け付けたか否かを判断する(ステップS16)。ステップS16において、入力受付モジュール270は、後述するテキストの変更入力を受け付けるか否かの基準となる確度の入力を受け付けたか否かを判断する。例えば、入力受付モジュール270は、確度の入力として、50%、60%等の数値の入力を受け付けたか否かを判断する。確度の数値としては、%表記に限らず、割合、数値等適宜変更可能である。
 ステップS16において、入力受付モジュール270は、入力を受け付けていないと判断した場合(ステップS16 NO)、表示モジュール271は、受信した変換情報に基づいて、テキストと確度とを表示する(ステップS17)。
 図7は、ステップS17において、表示モジュール271が表示するテキスト表示画面の一例を示す図である。図7において、表示モジュール271は、テキスト表示領域300、選択アイコン310を表示する。テキスト表示領域300は、変換情報に基づいたテキストを表示する領域である。テキスト表示領域300は、テキストを文節毎又は文章毎に区切って表示する。図7において、表示モジュール271は、テキスト表示領域300に、「こんにちは」と「あ、ついですね」を其々区切って表示する。選択アイコン310は、ユーザからの入力を受け付け、該当するテキストへの選択入力を受け付ける。音声出力モジュール272は、選択アイコン310への入力を受け付けたテキストを、変換情報に基づいて音声出力する。図7において、表示モジュール271は、テキストの確度に関わらず、選択アイコン310を各テキストに表示させる。
 一方、ステップS16において、入力受付モジュール270は、入力を受け付けたと判断した場合(ステップS16 YES)、表示モジュール271は、受信した変換情報に基づいて、テキストと確度とを表示する(ステップS18)。
 図8は、ステップS18において、表示モジュール271が表示するテキスト表示画面の一例を示す図である。図8において、所定の確度を50%の入力を受け付けた場合での一例を示す。図8において、表示モジュール271は、図7と同様に、テキスト表示領域300、選択アイコン310を表示する。図7と図8との相違点は、表示モジュール271は、図7では、全ての確度のテキストに対して選択アイコン310を表示するのに対して、図8では、入力を受け付けた確度以下の確度のテキストに対してのみ選択アイコン310を表示する点が相違している。すなわち、図8において、表示モジュール271は、確度が所定の値よりも高い「95%」である「こんにちは」に対しては、選択アイコン310を非表示とし、確度が所定の値よりも低い「40%」である「あついですね」に対しては、選択アイコン310を表示する。
 入力受付モジュール270は、テキストの選択入力を受け付けたか否かを判断する(ステップS19)。ステップS19において、入力受付モジュール270は、例えば、上述した選択アイコン310への入力操作を受け付けたか否か、表示するテキストへのタップ操作を受け付けたか否か等により判断する。
 ステップS19において、入力受付モジュール270は、受け付けていないと判断した場合(ステップS19 NO)、記憶モジュール260は、変換情報を記憶する(ステップS20)。ステップS20において、記憶モジュール260は、音声と、表示したテキストと、確度とを対応付けて記憶する。表示端末200は、ステップS20の処理を実行後、本処理を終了する。
 一方、ステップS19において、入力受付モジュール270は、受け付けたと判断した場合(ステップS19 YES)、音声出力モジュール272は、選択されたテキストを変換情報に基づいて、音声出力する(ステップS21)。
 入力受付モジュール270は、テキストの変更入力を受け付けたか否かを判断する(ステップS22)。ステップS22において、入力受付モジュール270は、直接入力、音声入力、ジェスチャー入力等に基づいて、表示モジュール271が表示するテキストとは異なるテキストへ変更する入力を受け付けたか否かを判断する。
 ステップS22において、入力受付モジュール270は、テキストの変更入力を受け付けていないと判断した場合(ステップS22 NO)、上述したステップS20の処理を実行する。
 一方、ステップS22において、入力受付モジュール270は、テキストの変更入力を受け付けたと判断した場合(ステップS22 YES)、変換モジュール273は、変更後のテキストの確度を算出する(ステップS23)。ステップS23において、変換モジュール273は、上述したステップS13の処理と同様の処理により、変更後のテキストの確度を算出する。
 変換モジュール273は、確度が所定の値以上であるか否かを判断する(ステップS24)。ステップS24において、変換モジュール273は、予め設定された値以上であるか否か、上述したステップS16において受け付けた値以上であるか否か等を判断する。ステップS24において、変換モジュール273は、所定の値以上であると判断した場合(ステップS24 YES)、表示モジュール271は、変更後のテキストとその確度とを表示する(ステップS25)。なお、ステップS25において、表示モジュール271は、変更後のテキストのみを表示する構成であってもよい。
 図9は、表示モジュール271が表示する変更後のテキスト表示画面の一例を示す図である。図9において、表示モジュール271は、テキスト表示領域300を表示する。表示モジュール271は、テキスト表示領域300に、変更されなかったテキストと、変更されたテキストと、其々の確度とを表示する。すなわち、図9において、「こんにちは」のテキストは、変更されなかったテキストであり、「あ、ついですね(変更済)」は、変更されたテキストである。このように、表示モジュール271は、変更されたテキストには、変更済である旨のメッセージを付与して表示する。
 なお、表示モジュール271は、変更されたテキストのみをテキスト表示領域300に表示する構成であってもよい。また、表示モジュール271は、変更済である旨のメッセージを付与せずに、変更されたテキストを表示する構成であってもよい。
 記憶モジュール260は、変換情報に含まれる音声と、変更後のテキストと、確度とを対応付けた変更後変換情報を記憶する(ステップS26)。ステップS26において、記憶モジュール260は、変更後変換情報だけでなく、変更していないテキストの音声と、このテキストと、確度とを対応付けて記憶する。すなわち、ステップS26において、記憶モジュール260は、表示モジュール271がステップS25において表示した全てのテキストと、このテキストの音声と、確度とを対応付けて記憶する。表示端末200は、ステップS26の処理を実行した後、本処理を終了する。
 一方、ステップS24において、変換モジュール273は、所定の値以上ではないと判断した場合(ステップS24 NO)、表示モジュール271は、変更後のテキストと、異なる変換候補の通知とを表示する(ステップS27)。
 図10は、表示モジュール271が表示する変更後のテキスト表示画面の一例を示す図である。図10において、確度が50%以上でない場合における一例を示す。図10において、表示モジュール271は、テキスト表示領域300、選択アイコン310、変換候補通知表示領域400を表示する。表示モジュール271は、テキスト表示領域300に、図9と同様に、変更されなかったテキストと、変更されたテキストと、其々の確度とを表示する。ここで、表示モジュール271は、確度が所定の値以上でないことを示す旨の通知を、変換候補通知表示領域400に表示する。表示モジュール271は、この変換候補通知表示領域400に、最も確度が高い変換候補のテキストを表示するとともに、この変換候補のテキストの確度を表示する。すなわち、図10において、「暑いですね(変更済)」の確度が「40%」であることから、所定の確度を下回っているため、今回テキストの変更を受け付けたものの、他により確度が高い変換候補が存在する旨の通知として「変更していますが、今回は、「あ、ついですね。」の可能性が高いです(確度70%)。」を、対応するテキストの一部に重畳させて表示する。
 なお、この通知の表示態様は、適宜変更可能であり、変換候補のテキストのみが表示されてもよいし、所定の確度以上の複数の変換候補のテキストが表示されてもよい。また、確度は、表示されていなくてもよい。
 入力受付モジュール270は、テキストの変更入力を受け付けたか否かを判断する(ステップS28)。ステップS28の処理は、上述したステップS22の処理と同様である。
 ステップS28において、入力受付モジュール270は、受け付けていないと判断した場合(ステップS28 NO)、上述したステップS26の処理を実行する。
 一方、ステップS28において、入力受付モジュール270は、受け付けたと判断した場合(ステップS28 YES)、上述したステップS23の処理を実行する。
 以上が、音声認識処理である。
 なお、上述した音声認識処理において、ユーザ端末100が音声認識しているが、表示端末200が音声認識し、テキスト変換する構成であってもよい。例えば、ユーザ端末100が受け付けた音声を、音声データとして、表示端末200に送信し、表示端末200が音声認識等の処理を実行する構成であってもよい。また、サーバやその他の構成が、音声認識、テキスト変換、確度算出等を実行する構成であってもよい。
 また、ユーザ端末100又は表示端末200が記憶した変更後のテキスト及び音声、変更前のテキスト及び音声等に基づいて、確度を算出する構成であってもよい。例えば、以前に変更したテキストや変更しなかったテキストのデータに基づいて、今回変更するテキストの変換候補を判断する構成であってもよい。
 上述した手段、機能は、コンピュータ(CPU、情報処理装置、各種端末を含む)が、所定のプログラムを読み込んで、実行することによって実現される。プログラムは、例えば、フレキシブルディスク、CD(CD-ROMなど)、DVD(DVD-ROM、DVD-RAMなど)等のコンピュータ読取可能な記録媒体に記録された形態で提供される。この場合、コンピュータはその記録媒体からプログラムを読み取って内部記憶装置又は外部記憶装置に転送し記憶して実行する。また、そのプログラムを、例えば、磁気ディスク、光ディスク、光磁気ディスク等の記憶装置(記録媒体)に予め記録しておき、その記憶装置から通信回線を介してコンピュータに提供するようにしてもよい。
 以上、本発明の実施形態について説明したが、本発明は上述したこれらの実施形態に限るものではない。また、本発明の実施形態に記載された効果は、本発明から生じる最も好適な効果を列挙したに過ぎず、本発明による効果は、本発明の実施形態に記載されたものに限定されるものではない。
 1 音声認識システム、100 ユーザ端末、200 表示端末

Claims (6)

  1.  音声を集音する集音手段と、
     前記集音した音声を認識する音声認識手段と、
     前記認識した音声をテキストに変換するテキスト変換手段と、
     前記変換したテキストを表示する表示手段と、
     前記表示したテキストから所望する一部分について、選択を受け付ける部分選択受付手段と、
     前記選択を受け付けた一部分のテキストに対応する音声を出力する音声出力手段と、
     を備えることを特徴とする音声認識システム。
  2.  前記選択を受け付けた一部分のテキストを、前記出力した音声の結果に基づいて、変更を受け付けるテキスト変更受付手段と、
     を備えることを特徴とする請求項1に記載の音声認識システム。
  3.  前記変更を受け付けたテキストを、先に音声認識した音声と対応付けて記憶する記憶手段と、
     を備えることを特徴とする請求項2に記載の音声認識システム。
  4.  前記変換したテキストの正解率を表示する確度表示手段と、
     を備えることを特徴とする請求項1に記載の音声認識システム。
  5.  音声を集音するステップと、
     前記集音した音声を認識するステップと、
     前記認識した音声をテキストに変換するステップと、
     前記変換したテキストを表示するステップと、
     前記表示したテキストから所望する一部分について、選択を受け付けるステップと、
     前記選択を受け付けた一部分のテキストに対応する音声を出力するステップと、
     を備えることを特徴とする音声認識方法。
  6.  音声認識システムに、
     音声を集音するステップ、
     前記集音した音声を認識するステップ、
     前記認識した音声をテキストに変換するステップ、
     前記変換したテキストを表示するステップ、
     前記表示したテキストから所望する一部分について、選択を受け付けるステップ、
     前記選択を受け付けた一部分のテキストに対応する音声を出力するステップ、
     を実行させるためのプログラム。
PCT/JP2016/071110 2016-07-19 2016-07-19 音声認識システム、音声認識方法及びプログラム Ceased WO2018015989A1 (ja)

Priority Applications (1)

Application Number Priority Date Filing Date Title
PCT/JP2016/071110 WO2018015989A1 (ja) 2016-07-19 2016-07-19 音声認識システム、音声認識方法及びプログラム

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
PCT/JP2016/071110 WO2018015989A1 (ja) 2016-07-19 2016-07-19 音声認識システム、音声認識方法及びプログラム

Publications (1)

Publication Number Publication Date
WO2018015989A1 true WO2018015989A1 (ja) 2018-01-25

Family

ID=60992320

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/JP2016/071110 Ceased WO2018015989A1 (ja) 2016-07-19 2016-07-19 音声認識システム、音声認識方法及びプログラム

Country Status (1)

Country Link
WO (1) WO2018015989A1 (ja)

Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2005228178A (ja) * 2004-02-16 2005-08-25 Nec Corp 書き起こしテキスト作成支援システムおよびプログラム
JP2011232521A (ja) * 2010-04-27 2011-11-17 On Semiconductor Trading Ltd 音声認識装置

Patent Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2005228178A (ja) * 2004-02-16 2005-08-25 Nec Corp 書き起こしテキスト作成支援システムおよびプログラム
JP2011232521A (ja) * 2010-04-27 2011-11-17 On Semiconductor Trading Ltd 音声認識装置

Similar Documents

Publication Publication Date Title
US11727914B2 (en) Intent recognition and emotional text-to-speech learning
US10089974B2 (en) Speech recognition and text-to-speech learning system
JP7121461B2 (ja) コンピュータシステム、音声認識方法及びプログラム
JP6588637B2 (ja) 個別化されたエンティティ発音の学習
CN105592343B (zh) 针对问题和回答的显示装置和方法
TWI711967B (zh) 播報語音的確定方法、裝置和設備
US11282523B2 (en) Voice assistant management
CN109754783A (zh) 用于确定音频语句的边界的方法和装置
CN107680581A (zh) 用于名称发音的系统和方法
CN101099147A (zh) 对话支持装置
CN107077638A (zh) 基于先进的递归神经网络的“字母到声音”
CN113409761B (zh) 语音合成方法、装置、电子设备以及计算机可读存储介质
CN113012683A (zh) 语音识别方法及装置、设备、计算机可读存储介质
JP2020187340A (ja) 音声認識方法及び装置
KR102935957B1 (ko) 인공지능 가상 비서 서비스에서의 텍스트 출력 방법 및 이를 지원하는 전자 장치
JP2015106203A (ja) 情報処理装置、情報処理方法、及びプログラム
CN107767862B (zh) 语音数据处理方法、系统及存储介质
KR20130116128A (ko) 티티에스를 이용한 음성인식 질의응답 시스템 및 그것의 운영방법
CN112259076A (zh) 语音交互方法、装置、电子设备及计算机可读存储介质
WO2018015989A1 (ja) 音声認識システム、音声認識方法及びプログラム
CN102043771A (zh) 电子化翻译方法、便携式电子装置及翻译系统
JP2016009091A (ja) 複数の異なる対話制御部を同時に用いて応答文を再生する端末、プログラム及びシステム
KR101611224B1 (ko) 오디오 인터페이스
KR102238973B1 (ko) 대화 데이터베이스를 이용한 대화문장 추천 방법 및 그것이 적용된 음성대화장치
JP5049310B2 (ja) 音声学習・合成システム及び音声学習・合成方法

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 16909464

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

32PN Ep: public notification in the ep bulletin as address of the adressee cannot be established

Free format text: NOTING OF LOSS OF RIGHTS PURSUANT TO RULE 112(1) EPC (EPO FORM 1205A DATED 23/04/2019)

NENP Non-entry into the national phase

Ref country code: JP

122 Ep: pct application non-entry in european phase

Ref document number: 16909464

Country of ref document: EP

Kind code of ref document: A1