WO2007111162A1 - テキスト表示装置、テキスト表示方法およびプログラム - Google Patents

テキスト表示装置、テキスト表示方法およびプログラム Download PDF

Info

Publication number
WO2007111162A1
WO2007111162A1 PCT/JP2007/055374 JP2007055374W WO2007111162A1 WO 2007111162 A1 WO2007111162 A1 WO 2007111162A1 JP 2007055374 W JP2007055374 W JP 2007055374W WO 2007111162 A1 WO2007111162 A1 WO 2007111162A1
Authority
WO
WIPO (PCT)
Prior art keywords
recognition
word
importance
text
recognition result
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/JP2007/055374
Other languages
English (en)
French (fr)
Inventor
Ken Hanazawa
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
NEC Corp
Original Assignee
NEC Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by NEC Corp filed Critical NEC Corp
Priority to JP2008507433A priority Critical patent/JPWO2007111162A1/ja
Priority to US12/294,318 priority patent/US20090287488A1/en
Publication of WO2007111162A1 publication Critical patent/WO2007111162A1/ja
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/22Procedures used during a speech recognition process, e.g. man-machine dialogue
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/26Speech to text systems
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L21/00Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
    • G10L21/06Transformation of speech into a non-audible representation, e.g. speech visualisation or speech processing for tactile aids
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/43Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
    • H04N21/439Processing of audio elementary streams
    • H04N21/4394Processing of audio elementary streams involving operations for analysing the audio stream, e.g. detecting features or characteristics in audio streams
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/43Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
    • H04N21/44Processing of video elementary streams, e.g. splicing a video clip retrieved from local storage with an incoming video stream or rendering scenes according to encoded video stream scene graphs
    • H04N21/4402Processing of video elementary streams, e.g. splicing a video clip retrieved from local storage with an incoming video stream or rendering scenes according to encoded video stream scene graphs involving reformatting operations of video signals for household redistribution, storage or real-time display
    • H04N21/440236Processing of video elementary streams, e.g. splicing a video clip retrieved from local storage with an incoming video stream or rendering scenes according to encoded video stream scene graphs involving reformatting operations of video signals for household redistribution, storage or real-time display by media transcoding, e.g. video is transformed into a slideshow of still pictures, audio is converted into text
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/47End-user applications
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/47End-user applications
    • H04N21/488Data services, e.g. news ticker
    • H04N21/4884Data services, e.g. news ticker for displaying subtitles

Definitions

  • Text display device text display method and program
  • the present invention relates to a text display device that displays text in synchronization with voice input, a text display method, and a program for causing a computer to execute the method.
  • FIG. 5 is a block diagram showing a configuration example of a conventional caption display device.
  • a conventional caption display device includes a voice input unit 201 for inputting voice, a storage unit 230 in which a recognition dictionary 203 for voice recognition is stored, and a voice recognition unit 202 for recognizing the input voice.
  • the control unit 220 includes an output unit 204 for displaying text.
  • the voice input unit 201 is represented by a microphone.
  • the conventional caption display device shown in FIG. 5 generally performs speech recognition processing when a speaker's utterance is received, and displays a recognition result word or word string with a slight delay from the speech. If the next utterance has already started after the recognition result is displayed, the next recognition result may be displayed in the same manner after a certain period of time.
  • FIG. 6 is a diagram showing a specific example of input voice and its recognition result.
  • the mobile phone has the configuration shown in FIG.
  • the cellular phone performs recognition processing for each section where the voice is detected, and displays the subtitles 502a to 502d in order for a certain period of time as the recognition result.
  • the mobile phone displays each of the subtitles 502a to 502d for a certain period of time TO from time tl to t5. In this way, the recognized word or word string is displayed for a certain period of time or until the next recognition result is obtained! /.
  • Patent Document 1 Japanese Patent Application Laid-Open No. 2002-342311
  • the present invention has been made to solve the problems of the conventional techniques as described above, and is a text display device capable of efficiently transmitting information by voice to the user as text. It is an object to provide a display method and a program for causing a computer to execute the method.
  • a text display device of the present invention includes a voice input unit for inputting voice, a storage unit storing a recognition dictionary for converting voice information into text,
  • a recognition dictionary for converting voice information into text
  • an output unit for displaying text and a word or a word string corresponding to the voice are recognized with reference to the recognition dictionary, and a recognition result including the word or the word string and its A control unit that obtains an importance level, calculates a display time of the recognition result corresponding to the importance level, and causes the output unit to display the recognition result for a time longer than the calculated display time.
  • control unit is any one of the reliability of recognition of input speech, the word importance described in the recognition dictionary, and the word importance specified by the user 1 One or a combination thereof may determine the importance of the recognition result.
  • control unit may recognize another recognition result from the calculation result of the display time.
  • the recognition result that is determined to be displayed for a longer time than the recognition result is either underlined, highlighted, or changed in font, size, or color
  • One or a combination of these may be displayed on the output unit.
  • the recognition result is displayed on the output unit for a time longer than the time calculated according to the importance. Therefore, if a recognition result with a high degree of importance is displayed for a long time, important information can be easily transmitted to the user.
  • a text display method of the present invention for achieving the above object is a text display method by an information processing apparatus for converting speech into text, and stores a recognition dictionary for converting speech information into text.
  • a speech is input, a word or a word string corresponding to the speech is recognized with reference to the recognition dictionary, and a recognition result including the word or the word string and its importance are obtained.
  • the display time of the recognition result is calculated corresponding to the importance, and the recognition result is displayed on the output unit for the calculated display time or longer.
  • the text display method may be one of the reliability of the recognition of the input speech, the word importance described in the recognition dictionary, and the word importance specified by the user, or one of these.
  • the importance of the recognition result may be determined by a combination of.
  • the text display method may underline, highlight, or display a recognition result determined to be displayed for a longer time than the other recognition results from the display time calculation result. Any one of changing the size or color, or a combination of these may be highlighted and displayed on the output unit.
  • a program of the present invention for achieving the above object is a program for causing a computer to execute processing for converting speech into text and displaying it, and for recognizing speech information for conversion into text.
  • a step of storing a dictionary in a storage unit; a step of recognizing a word or a word string corresponding to the voice by referring to the recognition dictionary when a voice is input; and a recognition result including the word or the word string And a step of obtaining the importance level, a step of calculating a display time of the recognition result corresponding to the importance level, and a calculation And displaying the recognition result on the output unit for the display time or longer.
  • the program may be based on one of the reliability of recognition of input speech, the word importance described in the recognition dictionary, and the word importance specified by the user, or a combination thereof.
  • the method may include a step of determining the importance of the recognition result.
  • the program draws an underline, highlights, or displays the recognition result determined to be displayed for a longer time than the other recognition results from the calculation result of the display time.
  • any one of changing the size or color, or a step of emphasizing by a combination of these may be displayed on the output unit.
  • the recognition result with high importance is displayed for a long time by giving priority to the recognition result, so that the important recognition result remains in the output unit even when the display screen is switched. Therefore, even if the place and time for displaying the recognition result are not sufficient, information can be efficiently transmitted to the user.
  • FIG. 1 is a block diagram showing a configuration example of a text display device according to the present embodiment.
  • FIG. 2 is a flowchart showing an operation procedure of the text display device of the present embodiment.
  • FIG. 3 is a diagram illustrating a description example of a recognition dictionary according to the present embodiment.
  • FIG. 4 is a diagram showing an example of input speech and recognition results of the present embodiment.
  • FIG. 5 is a block diagram showing a configuration example of a conventional text display device.
  • FIG. 6 is a diagram showing a specific example of input speech and recognition results in the conventional case.
  • the text display device of the present invention obtains the recognition result recognized from the input speech and its importance, calculates the display time corresponding to the importance, and displays the recognition result for the calculated display time or longer. It is characterized by that.
  • FIG. 1 is a block diagram showing an example of the configuration of the text display device of the present embodiment.
  • the text display device recognizes the speech input unit 101 for inputting speech, the storage unit 130 storing the recognition dictionary 103, and the input speech using the recognition dictionary 103, and the recognition result.
  • a speech recognition unit 202 that outputs a word or a word string and its importance
  • a control unit 120 that includes a display time calculation unit 204 that calculates a display time from the importance
  • an output unit 105 that displays a recognition result
  • the control unit 120 causes the output unit 105 to display the display time and the recognition result calculated by the display time calculation unit 204 according to the importance.
  • the control unit 120 has a CPU (Central Processing Unit) that executes predetermined processing according to a program and a memory for storing the program.
  • the voice recognition unit 102 and the display time calculation unit 104 are virtually configured in the control unit 120 when the CPU executes a program.
  • FIG. 2 is a flowchart showing the operation procedure of the text display device.
  • step 301 when voice is input via the voice input means 101 (step 301) and the voice recognition means 102 receives voice data from the voice input means 101, it is stored in the storage unit 130.
  • the speech is recognized with reference to the recognized recognition dictionary 103 (step 302).
  • a recognition result including a word or a word string is output and its importance is obtained.
  • the recognition result and its importance are output to the display time calculation means 104.
  • the display time calculation unit 104 receives the recognition result and the importance level information from the voice recognition means 102
  • the display time calculation unit 104 calculates the display time of the recognition result corresponding to the importance level (step 303).
  • the control unit 120 causes the output unit 105 to display the display time and the recognition result calculated according to the importance (step 304).
  • FIG. 3 is a diagram showing a description example of the recognition dictionary. As shown in FIG. 3, in the recognition dictionary 103, the importance of the word “RSS” is “3.0”, the importance of the word “site” is “1.5”, and the word “ It is described that the importance of “John” is “0.9”.
  • the speech recognition means 102 When the speech recognition means 102 identifies a word or word string with reference to the recognition dictionary 103, the speech recognition means 102 reads the importance from the recognition dictionary 103, and the recognition result including the identified word or word string and information on the importance To the display time calculation means 104.
  • Cw is a value indicating the word importance of the word w.
  • p is a coefficient.
  • An example of p is a display area dependent constant of the system.
  • the display area-dependent constant is a value that is determined by restrictions on the screen display size. The smaller the screen display size, the smaller the value because there is no room in the place and time for displaying the recognition result.
  • the control unit 120 highlights the recognition result with a high importance level, so that the recognition result is underlined with a high importance level. Displayed on the output unit 105. In this embodiment, it is assumed that the output unit 105 is highlighted when the display time of the recognition result is equal to or greater than the first threshold.
  • the first threshold value is the time for the criterion for determining the force or power to be highlighted.
  • the control unit 120 does not display the recognition result on the output unit 105 as a low importance if the display time of the recognition result does not reach the second threshold value.
  • the second threshold is a time that is a criterion for determining whether or not to display the force.
  • the first threshold and the second threshold are stored in advance in the storage unit 1 Stored in 30.
  • FIG. 4 is a diagram showing an example of input speech and recognition results in this embodiment.
  • the input voice information is the same as in the conventional case shown in FIG.
  • the coefficient p in the above equation (1) is set to 3.0.
  • the first threshold is set to 3.5 seconds
  • the second threshold is set to 2.0 seconds.
  • the standard subtitle switching period is set to 3.5 seconds.
  • the speech recognition means 102 recognizes words or word strings by speech in order.
  • the importance “3.0” is read from the recognition dictionary 103, and information about the word “RSS” and the importance “3.0” is passed to the display time calculation means 104.
  • the display time calculation means 104 calculates the display time T1 of the word “13 ⁇ 43” from the above equation (1).
  • the control unit 120 recognizes that the word “RSS highlighting target”. RSS is the information until “... is in circulation,” and after confirming what is highlighted, the subtitle 402a is displayed on the output unit 105.
  • the voice recognition means 102 recognizes the voices up to "each" is continuing.
  • the importance is obtained for each recognized word or word string, and is passed to the display time calculation means 104.
  • the display time calculation means 104 calculates the display time for each word or word string.
  • the display time of the word sequence “continued” is 1.5 seconds. This time is less than the second threshold.
  • the word “RSS” displayed in the subtitle 402a is an object to be highlighted.
  • the control unit 120 displays the subtitle 402a for 3.5 seconds and then instructs the output unit 105 to switch to the next subtitle, the display time of the word "RSS" in the subtitle 402a is set to 9 seconds. Because it is less than that, the word “RSS” is displayed in an underlined state. Further, the control unit 120 causes the output unit 105 to display the next subtitle 402b except for the word string “following”. In this way, subtitles 402b as shown in FIG.
  • the voice recognition means 102 recognizes the next voice as "Weblog ". This is the same as described above.
  • the display time calculation means 104 calculates the display time for each word or word string, based on the recognition result that also receives the voice recognition means 102 power, the information on its importance, and the above equation (1).
  • the control unit 120 outputs the word “RSS” because the display time of the two subtitles 402a and 402b is 7 seconds, which is less than 9 seconds.
  • the subtitle 402c is displayed, the word “RSS” remains highlighted. In this way, the caption 402c shown in FIG.
  • the control unit 120 applies the word “ ⁇ ⁇ 1 0 ⁇ ” and the word string “site overview format” to the highlight object as in the case of the word “13 ⁇ 43”. Make a decision.
  • the voice “newly proposed” is recognized by the voice recognition means 102, and the display time is displayed by the display time calculation means 104 for each word or word string. Calculated. Thereafter, the control unit 120 displays the next subtitle to the output unit 105 because the display time of the word “RSS” is 10.5 seconds, which is the sum of the three subtitle display times of subtitles 402a to 402c. When displaying, delete the word "RSS" from the display. On the other hand, since the word “We blogj” and the word string “site summary format” are newly highlighted, the control unit 120 sends the word “Weblog” and the word string “site summary format” to the output unit 105. Highlight. As a result, the subtitle 402d is displayed on the output unit 105 as shown in FIG.
  • the display target and the display time of the recognition result are obtained in consideration of the level of importance and the display constraint, and the display screen where the text display location and time are not sufficient. Select the recognition result to be displayed on the screen. Even if the recognition results cannot be displayed in real time, the recognition results to be highlighted are displayed for a long time, and the recognition results with low importance are not displayed. Is possible.
  • the recognition result “RSS” to be highlighted is displayed for three subtitle display times (total display time 10.5 seconds), but the recognition result “RSS” is displayed. When the time reaches 9 seconds, the recognition result “RSS” may be deleted from the display screen.
  • the importance of the word or word string is described in advance in the recognition dictionary 103 as shown in FIG. However, it may be changed according to the user's profile. For example, even if a word is highly important when it is first registered in the recognition dictionary 103, if it is highlighted many times, the user can understand the meaning of that word. make low. You can specify a word with high importance by yourself or write a numerical value indicating the importance.
  • the recognition result display time may be obtained using the recognition reliability instead of the word importance.
  • the recognition reliability indicates the suitability between the speech data and the word or word string when the speech recognition means 102 refers to the recognition dictionary 103 and identifies the word or word string for the input speech. If the input speech is not clear ⁇ or if there are multiple registered words that are similar in reading, there is a high probability that the speech recognition means 102 will identify a word or word string that is different from the input speech. , Reliability is low. This is because the recognition result with low reliability may be misrecognized, and if such a recognition result is highlighted, the user may be confused.
  • the importance of a word or a word string may be obtained by a combination of a numerical value described in advance in the recognition dictionary 103 as shown in FIG. 3 and the recognition reliability.
  • the numerical value previously described in the recognition dictionary 103 is high, if the reliability of the recognition result is low, the possibility of misrecognition increases and the information is not displayed.
  • misrecognition results with low reliability errors in information transmission can be reduced. As a result, the accuracy of information transmission to the user is improved.
  • the importance of the word or the word string is obtained by any one of the importance registered in the recognition dictionary 103 in advance, the designation by the user, and the reliability of recognition, or a combination thereof.
  • the recognition dictionary 103 in advance, the designation by the user, and the reliability of recognition, or a combination thereof.
  • the recognition results “RSS” and “Weblog” are highlighted by highlighting words that are determined to have high importance and long display time.
  • the font, size, or color of the target text may be changed! The target text may be highlighted.
  • a combination of these methods may be used. As a result, the user can easily distinguish words having high importance and long display time from other words.
  • the text display device of the present invention recognizes the input speech information, the text display device displays the recognition result on the output unit for a time calculated in accordance with the importance.
  • the text display device of the present invention can be applied to uses such as caption display in TV broadcasting, videophone, and web conferences. Further, the present invention may be applied to a program for causing a computer to execute the text display method of the present invention.

Landscapes

  • Engineering & Computer Science (AREA)
  • Multimedia (AREA)
  • Signal Processing (AREA)
  • Physics & Mathematics (AREA)
  • Computational Linguistics (AREA)
  • Health & Medical Sciences (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Human Computer Interaction (AREA)
  • Acoustics & Sound (AREA)
  • Quality & Reliability (AREA)
  • Data Mining & Analysis (AREA)
  • User Interface Of Digital Computer (AREA)
  • Controls And Circuits For Display Device (AREA)

Abstract

 音声による情報をテキストでユーザに効率よく伝達することが可能なテキスト表示装置を提供する。音声を入力するための音声入力部101と、音声による情報をテキストに変換するための認識辞書が格納された記憶部130と、テキストを表示するための出力部105と、音声が入力されると、音声に対応する単語または単語列を上記認識辞書を参照して認識し、単語または単語列を含む認識結果およびその重要度を求め、重要度に対応して認識結果の表示時間を計算し、算出した表示時間以上、認識結果を出力部に表示させる制御部120とを有する構成である。

Description

テキスト表示装置、テキスト表示方法およびプログラム
技術分野
[0001] 本発明は、音声の入力に同期して、そのテキストを表示するテキスト表示装置、テキ スト表示方法、およびその方法をコンピュータに実行させるためのプログラムに関する 背景技術
[0002] TV放送、 TV電話、 Web会議などにぉ 、て、音声認識によるリアルタイム自動字幕 表示装置が考えられている (特許文献 1参照)。従来の字幕表示装置を簡単に説明 する。
[0003] 図 5は従来の字幕表示装置の一構成例を示すブロック図である。従来の字幕表示 装置は、音声を入力するための音声入力部 201と、音声認識のための認識辞書 203 が格納された記憶部 230と、入力された音声を認識する音声認識手段 202を含む制 御部 220と、テキストを表示するための出力部 204とを有する構成である。音声入力 部 201はマイクに代表される。図 5に示す従来の字幕表示装置は、話者の発声を受 け付けると音声認識処理を行い、認識結果の単語あるいは単語列を、当該音声から 少し遅れて表示するものが一般的である。当該認識結果が表示された後に次の発声 が既に始まっている場合には、一定時間の表示の後、次の認識結果を同様に表示 すること〖こなる。
[0004] 図 6は入力される音声とその認識結果の具体例を示す図である。ここでは、携帯電 話機を用いた Web会議の場合とし、携帯電話機は図 5に示した構成を有しているもの とする。携帯電話機は、図 6に示すような入力音声 501があると、音声検出した区間 毎に認識処理を行い、その認識結果として、字幕 502aから 502dまでを順にそれぞ れ一定時間表示する。図 6の表示時刻に示すように、携帯電話機は、時刻 tlから t5 の間に字幕 502aから 502dのそれぞれを一定時間 TOだけ表示している。このように して、認識した単語あるいは単語列をある一定時間、または次の認識結果が得られる までの時間だけ表示して!/、た。 [0005] 特許文献 1:特開 2002— 342311号公報
発明の開示
発明が解決しょうとする課題
[0006] 上述した従来の方法では、認識対象となる音声が連続的に行われ、認識結果を表 示する場所および時間が充分でないと、ユーザが音声を聞き逃したり、字幕の認識 結果を見逃したりすることがあった。この場合、音声および字幕中に重要な単語が含 まれていても、その重要度と無関係に字幕が次々と切り替わってしまうため、ユーザ は重要な単語を知覚できないという問題があった。特に、携帯電話機はラップトップ 型やデスクトップ型のパーソナルコンピュータなどの情報処理装置の中でも、表示画 面の大きさが一般的に最も小さいため、多くの字幕およびその履歴を表示しておくこ とは困難であり、上述の問題が起こりやすい。
[0007] 本発明は上述したような従来の技術が有する問題点を解決するためになされたも のであり、音声による情報をテキストでユーザに効率よく伝達することが可能なテキス ト表示装置、テキスト表示方法、およびその方法をコンピュータに実行させるためのプ ログラムを提供することを目的とする。
課題を解決するための手段
[0008] 上記目的を達成するための本発明のテキスト表示装置は、音声を入力するための 音声入力部と、音声による情報をテキストに変換するための認識辞書が格納された 記憶部と、前記テキストを表示するための出力部と、音声が入力されると、該音声に 対応する単語または単語列を前記認識辞書を参照して認識し、該単語または該単 語列を含む認識結果およびその重要度を求め、該重要度に対応して該認識結果の 表示時間を計算し、算出した表示時間以上、該認識結果を前記出力部に表示させ る制御部とを有する構成である。
[0009] また、テキスト表示装置は、前記制御部が、入力される音声に対する認識の信頼度 、前記認識辞書に記述された単語重要度、およびユーザによって指定された単語重 要度のいずれか 1つ、またはこれらの組み合わせにより、前記認識結果の重要度を 決定するものであってもよ 、。
[0010] また、テキスト表示装置は、前記制御部が、前記表示時間の計算結果から他の認 識結果と比較して長い時間表示すると決定した認識結果に対して、下線を引くこと、 反転表示すること、またはフォントもしくは大きさもしくは色を変えることのうちいずれか
1つ、またはこれらの組み合わせにより強調して前記出力部に表示させるものであつ てもよい。
[0011] 本発明では、入力された音声の情報が認識されると、認識結果がその重要度に対 応して算出された時間以上、出力部に表示される。そのため、重要度の高い認識結 果が長く表示されるようにすれば、重要な情報がユーザにより伝達しやすくなる。
[0012] 上記目的を達成するための本発明のテキスト表示方法は、音声をテキストに変換す る情報処理装置によるテキスト表示方法であって、音声による情報をテキストに変換 するための認識辞書を記憶部に格納し、音声が入力されると、該音声に対応する単 語または単語列を前記認識辞書を参照して認識し、前記単語または前記単語列を 含む認識結果およびその重要度を求め、前記重要度に対応して前記認識結果の表 示時間を計算し、算出した表示時間以上、前記認識結果を出力部に表示させるもの である。
[0013] また、テキスト表示方法は、入力される音声に対する認識の信頼度、前記認識辞書 に記述された単語重要度、およびユーザによって指定された単語重要度の!ヽずれか 1つ、またはこれらの組み合わせにより、前記認識結果の重要度を決定するものであ つてもよい。
[0014] また、テキスト表示方法は、前記表示時間の計算結果から他の認識結果と比較して 長い時間表示すると決定した認識結果に対して、下線を引くこと、反転表示すること、 またはフォントもしくは大きさもしくは色を変えることのうちいずれか 1つ、またはこれら の組み合わせにより強調して前記出力部に表示させるものであってもよい。
[0015] 上記目的を達成するための本発明のプログラムは、音声をテキストに変換して表示 する処理をコンピュータに実行させるためのプログラムであって、音声による情報をテ キストに変換するための認識辞書を記憶部に格納するステップと、音声が入力される と、該音声に対応する単語または単語列を前記認識辞書を参照して認識するステツ プと、前記単語または前記単語列を含む認識結果およびその重要度を求めるステツ プと、前記重要度に対応して前記認識結果の表示時間を計算するステップと、算出 した表示時間以上、前記認識結果を出力部に表示させるステップと、を有する処理 を前記コンピュータに実行させるものである。
[0016] また、プログラムは、入力される音声に対する認識の信頼度、前記認識辞書に記述 された単語重要度、およびユーザによって指定された単語重要度のいずれか 1つ、 またはこれらの組み合わせにより、前記認識結果の重要度を決定するステップを有 するものであってもよい。
[0017] また、プログラムは、前記表示時間の計算結果から他の認識結果と比較して長!、時 間表示すると決定した認識結果に対して、下線を引くこと、反転表示すること、または フォントもしくは大きさもしくは色を変えることのうちいずれか 1つ、またはこれらの組み 合わせにより強調して前記出力部に表示させるステップを有するものであってもよい 発明の効果
[0018] 本発明によれば、重要度の高!ヽ認識結果を優先して長!ヽ時間表示することで、表 示画面が切り替わっても重要な認識結果が履歴として出力部に残る。したがって、認 識結果を表示する場所および時間が充分でなくても、ユーザに効率よく情報を伝達 できる。
図面の簡単な説明
[0019] [図 1]本実施形態のテキスト表示装置の一構成例を示すブロック図である。
[図 2]本実施形態のテキスト表示装置の動作手順を示すフローチャートである。
[図 3]本実施例の認識辞書の記述例を示す図である。
[図 4]本実施例の入力音声と認識結果の一例を示す図である。
[図 5]従来のテキスト表示装置の一構成例を示すブロック図である。
[図 6]従来の場合の入力音声と認識結果の具体例を示す図である。
符号の説明
[0020] 101 音声入力部
102 音声認識手段
103 認識辞書
104 表示時間計算手段 105 出力部
120 制御部
130 記憶部
発明を実施するための最良の形態
[0021] 本発明のテキスト表示装置は、入力音声から認識した認識結果およびその重要度 を求め、その重要度に対応して表示時間を算出し、算出した表示時間以上、認識結 果を表示することを特徴とする。
[0022] 次に、本実施形態のテキスト表示装置について図面を参照して詳細に説明する。
[0023] 図 1は本実施形態のテキスト表示装置の一構成例を示すブロック図である。本実施 形態のテキスト表示装置は、音声を入力するための音声入力部 101と、認識辞書 10 3が格納された記憶部 130と、入力された音声を認識辞書 103を用いて認識し、認識 結果の単語または単語列およびその重要度を出力する音声認識手段 202および重 要度から表示時間を計算する表示時間計算手段 204を含む制御部 120と、認識結 果を表示するための出力部 105とを有する。制御部 120は、重要度に応じて表示時 間計算手段 204により計算された表示時間、認識結果を出力部 105に表示させる。
[0024] 制御部 120は、プログラムにしたがって所定の処理を実行する CPU (Central Proce ssing Unit)とプログラムを格納するためのメモリとを有する。音声認識手段 102および 表示時間計算手段 104は CPUがプログラムを実行することで制御部 120内に仮想 的に構成される。
[0025] 次に、本実施形態のテキスト表示装置の動作を説明する。図 2はテキスト表示装置 の動作手順を示すフローチャートである。
[0026] 図 2に示すように、音声入力手段 101を介して音声が入力され (ステップ 301)、音 声認識手段 102が、音声入力手段 101から音声のデータを受け取ると、記憶部 130 に格納された認識辞書 103を参照して音声を認識する (ステップ 302)。続いて、単 語または単語列を含む認識結果を出力するとともに、その重要度を求める。そして、 認識結果およびその重要度を表示時間計算手段 104に出力する。表示時間計算手 段 104は、認識結果およびその重要度の情報を音声認識手段 102から受け取ると、 その重要度に対応して認識結果の表示時間を計算する (ステップ 303)。その後、制 御部 120は、その重要度に応じて算出された表示時間、認識結果を出力部 105に 表示させる(ステップ 304)。
[0027] [実施例 1]
本実施例のテキスト表示装置の構成を説明する。本実施例のテキスト表示装置の 認識辞書 103には、登録された各単語に対応して重要度の情報が記述されている。 図 3は認識辞書の記述例を示す図である。図 3に示すように、認識辞書 103には、単 語「RSS」の重要度が「3. 0」であり、単語「サイト」の重要度が「1. 5」であり、単語「バ 一ジョン」の重要度が「0. 9」であることが記述されて 、る。
[0028] 音声認識手段 102は、認識辞書 103を参照して単語または単語列を特定すると、 その重要度を認識辞書 103から読み出し、特定した単語または単語列を含む認識 結果とその重要度の情報を表示時間計算手段 104に渡す。
[0029] 表示時間計算手段 104が計算する、単語 wの表示時間 Tを求める計算式の一例と して、以下のものがある。
T = Cw X p · · ·式 (1)
Cwは単語 wの単語重要度を示す値である。 pは係数である。 pの一例として、システム の表示領域依存の定数がある。表示領域依存の定数とは、画面表示サイズの制約 で決まる値であり、画面表示サイズが小さいほど認識結果を表示する場所と時間に 余裕がなくなるため、その値が小さくなる。表示時間計算手段 104は、音声認識手段 102から単語 wを含む認識結果と重要度の情報を受け取ると、上記(1)式を計算し、 認識結果の表示時間 Tを算出する。
[0030] 表示時間計算手段 104が認識結果の表示時間を算出すると、制御部 120は、重要 度の高 、認識結果を強調表示するために、重要度の高 、認識結果に対して下線が 引かれた状態で出力部 105に表示させる。本実施例では、その認識結果の表示時 間が第 1の閾値以上であると、出力部 105に強調表示させるものとする。第 1の閾値 は、強調表示する力否力の判定基準の時間となる。
[0031] 反対に、制御部 120は、認識結果の表示時間が第 2の閾値に達していないと、重 要度の低いものとして出力部 105に表示させないものとする。第 2の閾値は、表示す る力否かの判定基準となる時間となる。第 1の閾値および第 2の閾値は予め記憶部 1 30に格納されている。
[0032] 次に、本実施例において、音声入力からテキスト表示までの動作を説明する。図 4 は本実施例における、入力音声と認識結果の一例を示す図である。ここでは、入力さ れる音声の情報が図 6で示した従来の場合と同様とする。また、上記式(1)の係数 p を 3. 0とする。また、第 1の閾値を 3. 5秒とし、第 2の閾値を 2. 0秒とする。また、字幕 の標準切換周期を 3. 5秒とする。
[0033] 音声認識手段 102は、「RSSは' · '流通しており、」までの音声が入力されると、音声 による単語または単語列を順に認識する。単語「RSS」を認識すると、認識辞書 103か らその重要度「3. 0」を読み出し、単語「RSS」と重要度「3. 0」の情報を表示時間計算 手段 104に渡す。表示時間計算手段 104は、上記(1)式から単語「1¾3」の表示時間 T1を算出する。表示時間 T1は、 3. 0 X 3. 0 = 9. 0秒となる。そして、制御部 120は、 単語「RSS」の表示時間が第 1の閾値よりも大きいことから、単語「RSS 強調表示の 対象になることを認識する。このようにして、制御部 120は、「RSSは · · '流通しており、 」までの情報で、強調表示するものを確定した後、字幕 402aを出力部 105に表示さ せる。
[0034] 続いて、音声認識手段 102は、「それぞれを' · '続いています。」までの音声を認識 する。上述したようにして、認識した単語または単語列毎に重要度を求めて、表示時 間計算手段 104に渡す。そして、表示時間計算手段 104が、各認識結果とその重要 度の情報を受け取ると、単語または単語列毎に表示時間を計算する。ここで、「続い ています」の重要度を 0. 5とすると、単語列「続いています」の表示時間は 1. 5秒とな る。この時間は第 2の閾値よりも小さい。また、上述したように、字幕 402aで表示した 単語「RSS」は強調表示の対象になつている。
[0035] 制御部 120は、字幕 402aを 3. 5秒表示させた後、出力部 105に次の字幕への切 換指示をする際、字幕 402aにおける単語「RSS」の表示時間が 9秒に満たないことか ら、下線が引かれた状態で単語「RSS」を表示させる。また、制御部 120は、出力部 10 5に対して、単語列「続いています」を除いて次の字幕 402bを表示させる。このように して、図 4に示すような字幕 402bが出力部 105に表示される。
[0036] さらに、音声認識手段 102が、次の「Weblogに…として」の音声を認識する場合も 、上述したのと同様である。続いて、表示時間計算手段 104が、音声認識手段 102 力も受け取る認識結果とその重要度の情報と上記(1)式により、単語または単語列 毎に表示時間を計算する。表示時間が算出された後、制御部 120は、単語「RSS」の 表示時間が字幕 402a、 402bの 2つの字幕表示時間を合わせても 7秒であり 9秒に 満たないことから、出力部 105に対して字幕 402cを表示させる際、単語「RSS」を強 調表示させたままにしておく。このようにして、図 4に示す字幕 402cが出力部 105に 表示される。なお、詳細な説明を省略するが、制御部 120は、単語「\^ 1」と単語 列「サイト概要フォーマット」についても、単語「1¾3」の場合と同様に、強調表示の対 象にする判断を行う。
[0037] 続いて、上述と同様にして、「新たに' · ·提案されています」の音声が音声認識手段 102により認識され、単語または単語列毎に表示時間が表示時間計算手段 104によ り算出される。その後、制御部 120は、単語「RSS」の表示時間が字幕 402a〜402c の 3つの字幕表示時間を合わせた 10. 5秒あり 9秒よりも長いことから、次の字幕を出 力部 105に表示させる際、単語「RSS」を表示から消去させる。一方、新たに単語「We blogjと単語列「サイト概要フォーマット」が強調表示の対象になったことから、制御部 120は、出力部 105に対して単語「Weblog」および単語列「サイト概要フォーマット」を 強調表示させる。その結果、図 4に示すように、字幕 402dが出力部 105に表示され る。
[0038] 上述のようにして、本実施例では、重要度の高低と表示制約とを考慮して認識結果 の表示対象と表示時間を求め、テキストを表示する場所と時間が充分ではない表示 画面に表示する認識結果を取捨選択する。そして、認識結果を全てリアルタイムに表 示しきれない場合でも、強調表示の対象となる認識結果を長く表示し、重要度の低い 認識結果を表示しないため、ユーザに対して効率よく情報伝達を行うことが可能とな る。
[0039] なお、本実施例では、強調表示の対象となる認識結果「RSS」を 3つの字幕表示時 間 (合計表示時間 10. 5秒)だけ表示させたが、認識結果「RSS」の表示時間が 9秒に なったときに、認識結果「RSS」を表示画面から消去するようにしてもよい。
[0040] また、単語または単語列の重要度が図 3に示すように予め認識辞書 103に記述さ れていてもよいが、ユーザのプロファイルにより変更可能であってもよい。例えば、認 識辞書 103をはじめに登録した段階で重要度の高い単語であっても、何度も強調表 示すると、ユーザもその単語の意味を理解できるようになるため、その単語の重要度 を低くする。ユーザ自身が重要度の高い単語を指定したり、重要度を示す数値を記 述したりしてちょい。
[0041] また、単語の重要度の代わりに認識の信頼度を用いて、認識結果の表示時間を求 めるようにしてもよい。認識の信頼度とは、音声認識手段 102が認識辞書 103を参照 して入力音声に対する単語または単語列を特定したときの、音声データと単語または 単語列との適合性を示すものである。入力される音声が明瞭でな ヽ場合や読み方の 似ている単語が複数登録されている場合などは、音声認識手段 102が入力音声と異 なる単語または単語列を特定してしまう確率が高くなり、信頼度が低くなる。信頼度の 低 ヽ認識結果は誤認識されて ヽる可能性があり、このような認識結果が強調表示さ れてしまうと、かえってユーザは混乱してしまうおそれがあるからである。
[0042] また、図 3に示したような認識辞書 103に予め記述されている数値と、認識の信頼 度との組み合わせによって、単語または単語列の重要度を求めるようにしてもよい。こ の場合には、予め認識辞書 103に記述されている数値が高くても、認識結果の信頼 度が低ければ、誤認識の可能性が高くなり、表示しないことになる。信頼度の低い誤 認識結果を表示しないことで、情報伝達の誤りを低減できる。その結果、ユーザへの 情報伝達の精度が向上する。
[0043] さらに、予め認識辞書 103に登録された重要度、ユーザによる指定、および認識の 信頼度のうちいずれか、または、これらの組み合わせにより、単語または単語列の重 要度を求めるようにしてもょ 、。
[0044] 強調表示の方法については、図 4に示した例では、認識結果「RSS」および「Weblog 」など、重要度が高ぐ表示時間が長いと判定された単語に下線を引いて強調表示を したが、この方法に限られない。強調表示の方法は、その他にも、対象となるテキスト のフォントまたは大きさまたは色を変えてもよ!、し、対象となるテキストを反転表示して もよい。また、これらの方法の組み合わせでもよい。これにより、ユーザは、重要度が 高く表示時間が長い単語を、それ以外の単語と容易に区別することが可能となる。 [0045] 本発明のテキスト表示装置は、入力された音声の情報を認識すると、認識結果をそ の重要度に対応して算出された時間以上、出力部に表示させる。重要度の高い認識 結果を優先して長 、時間表示することで、表示画面が切り替わっても重要な認識結 果が履歴として出力部に残る。したがって、認識結果を表示する場所および時間が 充分でなくても、ユーザに効率よく情報を伝達できる。
[0046] 本発明のテキスト表示装置を、 TV放送、 TV電話および WEB会議などでの字幕表 示といった用途に適用できる。また、本発明のテキスト表示方法をコンピュータに実 行させるためのプログラムに適用してもよい。

Claims

請求の範囲
[1] 音声を入力するための音声入力部と、
音声による情報をテキストに変換するための認識辞書が格納された記憶部と、 前記テキストを表示するための出力部と、
音声が入力されると、該音声に対応する単語または単語列を前記認識辞書を参照 して認識し、該単語または該単語列を含む認識結果およびその重要度を求め、該重 要度に対応して該認識結果の表示時間を計算し、算出した表示時間以上、該認識 結果を前記出力部に表示させる制御部と
を有することを特徴とするテキスト表示装置。
[2] 前記制御部は、
入力される音声に対する認識の信頼度、前記認識辞書に記述された単語重要度、 およびユーザによって指定された単語重要度のいずれか 1つ、またはこれらの組み 合わせにより、前記認識結果の重要度を決定することを特徴とする請求項 1に記載の テキスト表示装置。
[3] 前記制御部は、
前記表示時間の計算結果から他の認識結果と比較して長い時間表示すると決定し た認識結果に対して、下線を引くこと、反転表示すること、またはフォントもしくは大き さもしくは色を変えることのうちいずれか 1つ、またはこれらの組み合わせにより強調し て前記出力部に表示させることを特徴とする請求項 1または 2に記載のテキスト表示 装置。
[4] 音声をテキストに変換する情報処理装置によるテキスト表示方法であって、
音声による情報をテキストに変換するための認識辞書を記憶部に格納し、 音声が入力されると、該音声に対応する単語または単語列を前記認識辞書を参照 して認識し、
前記単語または前記単語列を含む認識結果およびその重要度を求め、 前記重要度に対応して前記認識結果の表示時間を計算し、
算出した表示時間以上、前記認識結果を出力部に表示させることを特徴とするテキ スト表示方法。
[5] 入力される音声に対する認識の信頼度、前記認識辞書に記述された単語重要度、 およびユーザによって指定された単語重要度のいずれか 1つ、またはこれらの組み 合わせにより、前記認識結果の重要度を決定することを特徴とする請求項 4に記載の テキスト表示方法。
[6] 前記表示時間の計算結果から他の認識結果と比較して長!、時間表示すると決定し た認識結果に対して、下線を引くこと、反転表示すること、またはフォントもしくは大き さもしくは色を変えることのうちいずれか 1つ、またはこれらの組み合わせにより強調し て前記出力部に表示させることを特徴とする請求項 4または 5に記載のテキスト表示 方法。
[7] 音声をテキストに変換して表示する処理をコンピュータに実行させるためのプロダラ ムであって、
音声による情報をテキストに変換するための認識辞書を記憶部に格納するステップ と、
音声が入力されると、該音声に対応する単語または単語列を前記認識辞書を参照 して認識するステップと、
前記単語または前記単語列を含む認識結果およびその重要度を求めるステップと 前記重要度に対応して前記認識結果の表示時間を計算するステップと、 算出した表示時間以上、前記認識結果を出力部に表示させるステップと、 を有する処理を前記コンピュータに実行させるためのプログラム。
[8] 入力される音声に対する認識の信頼度、前記認識辞書に記述された単語重要度、 およびユーザによって指定された単語重要度のいずれか 1つ、またはこれらの組み 合わせにより、前記認識結果の重要度を決定するステップを有する請求項 7に記載 のプログラム。
[9] 前記表示時間の計算結果から他の認識結果と比較して長!、時間表示すると決定し た認識結果に対して、下線を引くこと、反転表示すること、またはフォントもしくは大き さもしくは色を変えることのうちいずれか 1つ、またはこれらの組み合わせにより強調し て前記出力部に表示させるステップを有する請求項 7または 8に記載のプログラム。
PCT/JP2007/055374 2006-03-24 2007-03-16 テキスト表示装置、テキスト表示方法およびプログラム Ceased WO2007111162A1 (ja)

Priority Applications (2)

Application Number Priority Date Filing Date Title
JP2008507433A JPWO2007111162A1 (ja) 2006-03-24 2007-03-16 テキスト表示装置、テキスト表示方法およびプログラム
US12/294,318 US20090287488A1 (en) 2006-03-24 2007-03-16 Text display, text display method, and program

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
JP2006082658 2006-03-24
JP2006-082658 2006-03-24

Publications (1)

Publication Number Publication Date
WO2007111162A1 true WO2007111162A1 (ja) 2007-10-04

Family

ID=38541082

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/JP2007/055374 Ceased WO2007111162A1 (ja) 2006-03-24 2007-03-16 テキスト表示装置、テキスト表示方法およびプログラム

Country Status (4)

Country Link
US (1) US20090287488A1 (ja)
JP (1) JPWO2007111162A1 (ja)
CN (1) CN101410790A (ja)
WO (1) WO2007111162A1 (ja)

Cited By (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2011099086A1 (ja) * 2010-02-15 2011-08-18 株式会社 東芝 会議支援装置
JP2012181358A (ja) * 2011-03-01 2012-09-20 Nec Corp テキスト表示時間決定装置、テキスト表示システム、方法およびプログラム
JP2013174718A (ja) * 2012-02-24 2013-09-05 Toshiba Corp 音声記録選択装置、音声記録選択方法及び音声記録選択プログラム
JP2019062332A (ja) * 2017-09-26 2019-04-18 株式会社Jvcケンウッド 表示態様決定装置、表示装置、表示態様決定方法及びプログラム
JP2022049984A (ja) * 2020-09-17 2022-03-30 Necソリューションイノベータ株式会社 出力方法

Families Citing this family (8)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP4957821B2 (ja) * 2010-03-18 2012-06-20 コニカミノルタビジネステクノロジーズ株式会社 会議システム、情報処理装置、表示方法および表示プログラム
CN102385861B (zh) * 2010-08-31 2013-07-31 国际商业机器公司 一种用于从语音内容生成文本内容提要的系统和方法
CN102566863B (zh) * 2010-12-25 2016-07-27 上海量明科技发展有限公司 在即时通信工具中设置辅助区的方法及系统
DE112012002190B4 (de) * 2011-05-20 2016-05-04 Mitsubishi Electric Corporation Informationsgerät
CN102693094A (zh) * 2012-06-12 2012-09-26 上海量明科技发展有限公司 即时通信中调整字符的方法、客户端及系统
JP5921722B2 (ja) * 2013-01-09 2016-05-24 三菱電機株式会社 音声認識装置および表示方法
CN112599130B (zh) * 2020-12-03 2022-08-19 安徽宝信信息科技有限公司 一种基于智慧屏的智能会议系统
CN114360530B (zh) * 2021-11-30 2024-08-27 北京罗克维尔斯科技有限公司 语音测试方法、装置、计算机设备和存储介质

Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH10123450A (ja) * 1996-10-15 1998-05-15 Sony Corp 音声認識機能付ヘッドアップディスプレイ装置
JPH10301927A (ja) * 1997-04-23 1998-11-13 Nec Software Ltd 電子会議発言整理装置
JP2006005861A (ja) * 2004-06-21 2006-01-05 Matsushita Electric Ind Co Ltd 文字スーパー表示装置および文字スーパー表示方法

Family Cites Families (8)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20030093790A1 (en) * 2000-03-28 2003-05-15 Logan James D. Audio and video program recording, editing and playback systems using metadata
GB2323693B (en) * 1997-03-27 2001-09-26 Forum Technology Ltd Speech to text conversion
US6839669B1 (en) * 1998-11-05 2005-01-04 Scansoft, Inc. Performing actions identified in recognized speech
US7164753B2 (en) * 1999-04-08 2007-01-16 Ultratec, Incl Real-time transcription correction system
US7953219B2 (en) * 2001-07-19 2011-05-31 Nice Systems, Ltd. Method apparatus and system for capturing and analyzing interaction based content
JP2004304601A (ja) * 2003-03-31 2004-10-28 Toshiba Corp Tv電話装置、tv電話装置のデータ送受信方法
JP3945778B2 (ja) * 2004-03-12 2007-07-18 インターナショナル・ビジネス・マシーンズ・コーポレーション 設定装置、プログラム、記録媒体、及び設定方法
US7729478B1 (en) * 2005-04-12 2010-06-01 Avaya Inc. Change speed of voicemail playback depending on context

Patent Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH10123450A (ja) * 1996-10-15 1998-05-15 Sony Corp 音声認識機能付ヘッドアップディスプレイ装置
JPH10301927A (ja) * 1997-04-23 1998-11-13 Nec Software Ltd 電子会議発言整理装置
JP2006005861A (ja) * 2004-06-21 2006-01-05 Matsushita Electric Ind Co Ltd 文字スーパー表示装置および文字スーパー表示方法

Cited By (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2011099086A1 (ja) * 2010-02-15 2011-08-18 株式会社 東芝 会議支援装置
JP2012181358A (ja) * 2011-03-01 2012-09-20 Nec Corp テキスト表示時間決定装置、テキスト表示システム、方法およびプログラム
JP2013174718A (ja) * 2012-02-24 2013-09-05 Toshiba Corp 音声記録選択装置、音声記録選択方法及び音声記録選択プログラム
JP2019062332A (ja) * 2017-09-26 2019-04-18 株式会社Jvcケンウッド 表示態様決定装置、表示装置、表示態様決定方法及びプログラム
JP2022049984A (ja) * 2020-09-17 2022-03-30 Necソリューションイノベータ株式会社 出力方法

Also Published As

Publication number Publication date
US20090287488A1 (en) 2009-11-19
JPWO2007111162A1 (ja) 2009-08-13
CN101410790A (zh) 2009-04-15

Similar Documents

Publication Publication Date Title
WO2007111162A1 (ja) テキスト表示装置、テキスト表示方法およびプログラム
JP5064404B2 (ja) モバイルデバイスにおける音声および代替入力手法の組み合わせ
CN111312231B (zh) 音频检测方法、装置、电子设备及可读存储介质
JP6751658B2 (ja) 音声認識装置、音声認識システム
CN106971723B (zh) 语音处理方法和装置、用于语音处理的装置
US10553206B2 (en) Voice keyword detection apparatus and voice keyword detection method
US8868419B2 (en) Generalizing text content summary from speech content
JPWO2015098109A1 (ja) 音声認識処理装置、音声認識処理方法、および表示装置
CN109036406A (zh) 一种语音信息的处理方法、装置、设备和存储介质
KR20130135410A (ko) 음성 인식 기능을 제공하는 방법 및 그 전자 장치
CN106250474A (zh) 一种语音控制的处理方法及系统
WO2016110068A1 (zh) 语音识别设备语音切换方法及装置
CN113763921B (zh) 用于纠正文本的方法和装置
CN112863496B (zh) 一种语音端点检测方法以及装置
US20250378286A1 (en) Application Programming Interfaces For On-Device Speech Services
US11217266B2 (en) Information processing device and information processing method
JP2008033198A (ja) 音声対話システム、音声対話方法、音声入力装置、プログラム
JP6260138B2 (ja) コミュニケーション処理装置、コミュニケーション処理方法、及び、コミュニケーション処理プログラム
CN110971505B (zh) 一种通讯信息处理方法、装置、终端及计算机可读介质
CN115394297A (zh) 一种语音识别方法、装置、电子设备及存储介质
US20080256071A1 (en) Method And System For Selection Of Text For Editing
CN108491183B (zh) 一种信息处理方法和电子设备
CN114564265B (zh) 有屏智能设备的交互方法、装置以及电子设备
US20200219482A1 (en) Electronic device for processing user speech and control method for electronic device
JP2020024310A (ja) 音声処理システム及び音声処理方法

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 07738819

Country of ref document: EP

Kind code of ref document: A1

DPE1 Request for preliminary examination filed after expiration of 19th month from priority date (pct application filed from 20040101)
WWE Wipo information: entry into national phase

Ref document number: 2008507433

Country of ref document: JP

WWE Wipo information: entry into national phase

Ref document number: 12294318

Country of ref document: US

Ref document number: 200780010487.X

Country of ref document: CN

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 07738819

Country of ref document: EP

Kind code of ref document: A1

DPE1 Request for preliminary examination filed after expiration of 19th month from priority date (pct application filed from 20040101)