WO2008062529A1 - Sentence reading-out device, method for controlling sentence reading-out device and program for controlling sentence reading-out device - Google Patents
Sentence reading-out device, method for controlling sentence reading-out device and program for controlling sentence reading-out device Download PDFInfo
- Publication number
- WO2008062529A1 WO2008062529A1 PCT/JP2006/323427 JP2006323427W WO2008062529A1 WO 2008062529 A1 WO2008062529 A1 WO 2008062529A1 JP 2006323427 W JP2006323427 W JP 2006323427W WO 2008062529 A1 WO2008062529 A1 WO 2008062529A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- information
- word
- speech
- display
- text
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L13/00—Speech synthesis; Text to speech systems
- G10L13/02—Methods for producing synthetic speech; Speech synthesisers
- G10L13/04—Details of speech synthesis systems, e.g. synthesiser structure or memory management
Definitions
- Text reading device control method for controlling text reading device, and control program for controlling text reading device
- the present invention relates to a technique for supplementing an unnatural portion of a reading voice in a text reading device that reads a text described in a text file or the like.
- the sound information is information obtained by encoding the sound of a word pronounced by a person.
- the phoneme of phoneme information here is the smallest unit of sound that is abstracted to form concrete speech.
- the phoneme information is information obtained by encoding a phoneme sound from which the sound power of a word pronounced by a person is also extracted.
- the synthesized speech information synthesized from the above phoneme information is used.
- This synthesized speech information is obtained by synthesizing phoneme information and adjusting accents and intonations for a more natural speech.
- synthesized speech using this synthesized speech information was something that people could hear while feeling unnatural.
- Patent Document 1 JP-A-8-87698
- Patent Document 2 JP 2005-265477
- a text-to-speech device that has storage means that stores speech information in units of words! Since the speech information is stored, the text has the function of supplementing words uttered with unnatural synthesized speech.
- a reading device is provided.
- a text-to-speech reading apparatus having a storage means for storing speech information in units of words
- an unstored word is stored in the storage means.
- a determination unit that determines whether or not the document is to be read out; and a display unit that highlights and displays the notation information of the unstored word based on the determination result of the determination unit.
- the display information is symbol information of an unstored word and an unstored word.
- the display means terminates the display of the notation information based on a request for external force.
- a control method for controlling a text-to-speech device having a storage means storing speech information in units of words is described.
- the display information is symbol information of an unstored word and an unstored word.
- the display step ends the display of the notation information based on a request from an external force.
- a seventh means for solving the problem to be solved by the above invention in a control program for controlling a text-to-speech device having a storage means storing voice information in units of words, the storage Is stored in the means! Judgment step for judging whether or not an unstored word exists in the document to be read out, and based on the judgment result of the judgment step, the notation information of the unstored word is emphasized.
- the display information is unstored words and symbol information of unstored words.
- the display step terminates the display of the notation information based on a request for an external force.
- the word read out by the synthesized voice and the symbol information of the word are displayed, so that the person who hears the synthesized voice can be read out by the synthesized voice. Even when a word cannot be understood only by displaying the word, the meaning of the word of the synthesized speech can be completely understood based on the symbol information.
- the voice information is stored! Therefore, by terminating the display of the word notation information read out by the synthesized voice, the synthesized voice is heard. This has the effect of adjusting the time required for a person to understand the meaning of a word read out in synthesized speech.
- FIG. 1 is a hardware configuration diagram of a document reading apparatus.
- FIG. 2 is a block diagram of a word DB.
- FIG. 3 is a block diagram of a phoneme DB.
- FIG. 4 Configuration diagram of symbol DB.
- FIG. 5 is a functional block diagram of a text-to-speech process.
- FIG. 6 is a flowchart (part 1) of a text reading process in the first embodiment.
- FIG. 7 is a flowchart (part 2) of the text reading process in the first embodiment.
- FIG. 8 is a flowchart of a text-to-speech process in the second embodiment.
- FIG. 9 A display example of a synthetic voice supplement screen.
- the word here means the smallest unit of a language that has a unified meaning and function in the grammar.
- the present invention that provides a function for supplementing words uttered by unnatural synthesized speech is required.
- FIG. 1 is a block diagram illustrating an example of a hardware configuration of the text-to-speech reading apparatus 1.
- the text-to-speech device 1 includes a CPU (Central Processing Unit) 3, a storage unit 5, an input unit 7, an output unit 9, and a bus 11.
- the CPU 3 controls each part and performs various calculations.
- the storage unit 5 stores a text reading program 51, a word DB53, a phoneme DB55, and a symbol DB57.
- RAM Random for executing programs and storing data
- ROM Read Only Memory
- an external storage device capable of storing a large amount of programs and data.
- a text-to-speech program 51 When a text-to-speech program 51 is given a text to be read and a reading request from the input unit 7, it reads out using the word DB53, phoneme DB55, and symbol DB57.
- This reading process includes a function for storing speech information and supplementing synthesized speech of words.
- the word DB53 stores voice information in units of words used for reading.
- the phoneme DB55 stores phoneme information used for reading.
- the symbol DB 57 stores symbol information for supplementing the above-described synthesized speech.
- the input unit 7 is used to give the text-to-speech device 1 a request for external power for the text-to-speech document and text-to-speech processing.
- the output unit 9 sends out reading voice and notation information related to the reading voice to the outside. Specifically, it operates as a speaker or monitor.
- the bus 11 is for exchanging data between the CPU 3 and the storage unit 5, the input unit 7, and the output unit 9.
- the text here refers to a group of letters and thoughts and feelings.
- Input section 7 receives a reading target document and a reading request for it.
- the CPU 3 expands the text-to-speech program 51 to the RAM, and executes the text-to-speech program 51. Then, the text-to-speech reading program 51 uses the reading target document given in (1) and the word DB 53, the phoneme DB 55, and the symbol DB 57 to generate the reading voice information of the reading target document and the notation information corresponding to the reading voice information.
- the output unit 9 sends out the reading voice information generated in (2) and the notation information corresponding to the reading voice information.
- FIG. 2 shows a word DB 53 that stores voice information of words.
- the word DB 53 is used by the text reading device 1 to extract voice information of words used in the target text.
- Word DB53 information elements are word name 531 and voice information 533, read aloud Time is 535.
- the word name 531 is information used when the text-to-speech reading device 1 searches for speech information of a word used in the target reading document.
- the voice information 533 is used when the voice reading device 1 sends out the sound of the word from the output unit 9 to the outside. This voice information is information obtained by encoding the voice of a word pronounced by a person, and it may be further compressed in some cases.
- the reading time 535 is the time taken to read the voice information 533. This reading time 535 is information used by the text reading device 1 to calculate a trigger for displaying notation information of words that are not stored in the word DB 53.
- Fig. 3 shows the phoneme DB55 that stores phoneme information.
- the phoneme DB55 is used by the text-to-speech reading device 1 for synthesizing the speech stored in the word DB53.
- the phoneme DB55 information elements are phoneme name 551, phoneme information 553, and reading time 555.
- the phoneme name 551 is used by the text-to-speech reading device 1 to extract phoneme information to be synthesized.
- the phoneme information 553 is used when the speech reading apparatus 1 synthesizes speech information of a single word that is not stored in the word DB 53.
- the phoneme information 553 is information obtained by encoding the phoneme sound extracted from the speech power of a word pronounced by a person, and may be further compressed in some cases.
- the reading time 555 is the time taken to read the phoneme information 553. This reading time 555 is information used by the text reading device 1 to calculate the trigger for displaying V, N, and word notation information stored in the word DB 53.
- FIG. 4 shows a symbol DB57 that stores symbols of words that are not stored in the word DB53.
- the symbol DB57 is used to display a symbol related to the meaning of the word used in the target reading document although the text reading device 1 is not stored in the word DB53.
- the symbol here means a sign other than a character.
- the information elements of symbol DB57 are word name 571 and symbol information 573.
- the character here means a sign representing a word.
- the word name 571 is information used when the text reading device 1 searches for symbol information of a word used in the target reading document.
- the symbol information 573 is used when the speech reading apparatus 1 sends out a symbol related to the meaning of the word from the output unit 9 to the outside.
- a company logo is stored as an example.
- FIG. 5 is a functional block diagram showing an example of the text-to-speech function.
- the text-to-speech function of the text-to-speech reading device 1 functions when the text-to-speech program 51 is executed.
- the text-to-speech function is composed of input means 2, determination means 4, storage means 6, speech means 8, and display means 10. Each means of the text reading function will be described below.
- the input means 2 gives a text-to-speech device 1 with a text to be read and a reading request for it. Also, a display information display end request to be described later is given to the display means 10.
- the judging means 4 performs the following operation.
- the entire speech information corresponding to the read text is generated. Also, when synthetic voice information is included in the whole voice information, an opportunity to read out the synthesized voice information to be monitored during utterance is set.
- the synthesized speech information here is generated by using the above phoneme information to generate speech information of unstored words for which speech information does not exist in the storage means. Then, the whole speech information is given to the speech means 8.
- the storage means 6 stores speech information and phoneme information in units of words and symbol information in units of words.
- the speech information in units of words corresponds to the word DB53.
- Phoneme information corresponds to phoneme DB55.
- the symbol information corresponds to the symbol DB57.
- the utterance means 8 sends out the whole voice information given from the judgment means 4 to the outside as a sound.
- the display means 10 sends the notation information given from the judgment means 4 to the outside as characters or symbols. Start out. Further, the process of sending out characters and symbols to the outside in response to the display end display request of the notation information given from the input means 2 is ended.
- the determination unit 4 analyzes the reading target sentence that is the reading information given from the input unit 2.
- the analysis refers to determining whether or not the speech DB53 stores the speech information of the words used in the reading target document.
- the determination means 4 extracts unstored words in which the speech information 533 is not stored in the speech DB 53 in which the central force of all words used in the reading target sentence is also divided in the determination of S501.
- the determination means 4 determines whether or not there is an unstored word in the speech DB 53 in which speech information is not stored. As a result of the determination, if there is an unstored word in which no voice information is stored, the process of S507 is performed. As a result of the determination, if the voice information is not stored and there is no unstored word, the process of S513 is performed.
- the determination means 4 extracts phoneme information corresponding to the unstored word extracted in S503 from the phoneme DB 55.
- a specific extraction method is as follows. Unregistered words are converted into Roman characters, which are information indicating how to read them, based on the rule information that the text-to-speech reading device 1 has. Then, phoneme information 553 corresponding to the phoneme name included in the Romaji is extracted from the phoneme DB 55.
- the determination unit 4 synthesizes the phoneme information 553 extracted in S507 to generate synthesized speech information of an unregistered word. Then, the synthesized speech is edited so that it falls within the amplitude threshold of the text-to-speech reading device 1. This editing is performed to adjust the prosody (rhythm) of the synthesized speech so that it can be heard naturally.
- the determination means 4 sets an opportunity to read the synthesized speech of the unstored word during the reading of the target document.
- a specific setting method is as follows.
- the individual reading time 535 of words existing before the unseen word is added to the word existing at the beginning of the reading target sentence, and the time required to speak the speech information is calculated.
- the calculated time is used as an opportunity to start displaying the unstored word.
- the time required for speaking the synthesized speech is calculated by adding the reading time 555 of the phoneme information used when the synthesized speech of the unstored word is generated.
- the calculated time and the time obtained by adding the display start trigger are stored in the storage unit 5 as a display end trigger for the unstored word. If there are multiple unstored words in the text to be read, repeat the above process.
- the determination means 4 generates a whole voice corresponding to the whole reading target sentence.
- the whole voice information may be generated by connecting only the voice information 533 of the word DB53, or may be generated by connecting the voice information 533 of the word DB53 and the synthesized voice information generated by S509. Then, the overall loudness and pitch of the sound is adjusted based on the rule information that the text-to-speech device 1 has. This adjustment is performed so that the sound of the entire audio information can be heard naturally.
- the determination unit 4 determines whether or not the entire voice information generated in S513 includes the synthesized voice information generated in S509. As a result of the determination, if the entire voice information generated in S513 includes the synthesized voice generated in S509, the process of S519 is performed. As a result of the determination, if the entire speech information generated in S513 does not include the synthesized speech information generated in S509, the utterance means 8 utters the entire speech information in the process of S517.
- the utterance means 8 starts uttering the entire voice synthesized in S513. This whole voice information is generated by connecting the voice information 533 of the word DB53 and the synthesized voice information synthesized in S509.
- the determination unit 4 monitors whether or not the elapsed time from the utterance of the entire voice information has reached the display start opportunity calculated in S511 in S519. This monitoring is performed until the elapsed time of the whole voice that starts speaking in S519 reaches the display start trigger calculated in S511. If, as a result of this monitoring, the elapsed time of the entire voice that started speaking in S519 has reached the display start timing calculated in S511, the processing in S523 is performed.
- the determination means 4 determines whether or not symbol information of an unstored word corresponding to the display start trigger exists in the symbol DB57. As a result of the determination, if the symbol information of the unstored word does not exist in the symbol DB 57, the display means 10 displays and outputs the character information of the unstored word extracted in S503 on the output unit 9 in S525. As a result of this determination, When the symbol information of the stored word exists in the symbol DB57, in S527, the display means 10 displays and outputs the character information of the unstored word extracted in S503 and the symbol information in the symbol DB57 on the output unit 9.
- S525 and S527 will be described with reference to FIG. Figure 9 assumes that the text-to-speech device 1 has been commercialized as a car navigation system with a navigation function.
- Reference numeral 901 denotes a car navigation system.
- Reference numeral 903 denotes a speaker that outputs a reading voice.
- Reference numeral 905 denotes a screen for displaying a map used for navigation.
- Reference numeral 907 denotes a map used for navigation.
- 909 indicates the character of the unstored word displayed in S525.
- personal names are shown as unremembered words.
- 911 indicates symbol information displayed in S527.
- the symbol information corresponding to 909 the company logo related to the name of 909 is shown.
- 9 13 shows a mail reading button.
- This e-mail reading button is used when the car navigation system 1 performs a process of reading the received e-mail.
- Reference numeral 915 denotes a setting button. This setting button is used to make various settings for the car navigation system.
- 919 is an indication of the position of the vehicle equipped with the car navigation system on the 907 map.
- S921 indicates a controller. This controller is used to specify the destination on the 907 map.
- the character information displayed in S525 corresponds to 909. Character information displayed by S527 is equivalent to 909, and symbol information is equivalent to 911.
- the determination means 4 monitors whether the elapsed time from the display start trigger detected in S521 has reached the display end trigger calculated in S511. This monitoring is performed until the elapsed time of the display start trigger detected in S521 reaches the display end trigger calculated in S511. As a result of this monitoring, when the elapsed time from the display start trigger detected in S521 has reached the display end trigger calculated in S511, the display of the information displayed on the display means 10 is ended in S530.
- the determination means 4 monitors whether the elapsed time from the display start trigger detected in S521 has reached the display end trigger calculated in S511. This monitoring is performed until the elapsed time of the display start trigger detected in S521 reaches the display end trigger calculated in S511. As a result of this monitoring, when the elapsed time from the display start trigger detected in S521 has reached the display end trigger calculated in S511, the process of S541 is performed.
- the determination means 4 determines whether or not it is the force received from the input means 2 from the outside to end display of an unstored word and a symbol corresponding to the unstored word. If the end request is received as a result of this determination, the display of the information displayed on the display means 10 is ended in S530. If the result of this determination is that an end request has not been received, processing in S543 is performed.
- the determination means 4 determines whether the elapsed time from the display end trigger detected in S531 has reached the extension time that the text reading device 1 has in the storage unit 5. This determination is performed until the elapsed time of the display end trigger detected at S531 reaches the extension time. As a result of this determination, when the elapsed time from the display end trigger detected in S531 has reached the extended time, the display of the information displayed on the display means 10 is ended in S530.
- the present invention has been described based on the embodiments.
- the present invention is not limited to the above-described embodiments, and may be implemented in any way as long as the configuration described in the claims is not changed. be able to.
- the present invention is a technique for supplementing a part in which a reading voice is unnatural in a text reading apparatus that reads a text described in a text file or the like, and a navigation system is a product such as a portable terminal. Applicable to.
Landscapes
- Engineering & Computer Science (AREA)
- Computational Linguistics (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Human Computer Interaction (AREA)
- Physics & Mathematics (AREA)
- Acoustics & Sound (AREA)
- Multimedia (AREA)
- Document Processing Apparatus (AREA)
Description
明 細 書
文章読上げ装置、文章読上げ装置を制御する制御方法及び文章読上げ 装置を制御する制御プログラム
技術分野
[0001] 本発明は、テキストファイルなどに記載された文章を読上げる文章読上げ装置にお いて、読上げ音声が不自然だった部分を補足する技術に関する。
背景技術
[0002] テキストファイルを表示しながら読上げるソフトウェアは既に販売されている。この読 上げは、単語と音声情報を記憶した単語 DB (DataBase)と、音素情報を記憶した音 素 DBを使う。ここでいう音声情報とは、人が発音した単語の音を符号化した情報であ る。また、ここでいう音素情報の音素とは、具体的な音声を形作るものとして抽象化し た音の最小単位である。音素情報とは、人が発音した単語の音力も抽出した音素の 音を符号ィ匕した情報である。読上げ対象文章中の単語が単語 DBに記憶されていた 場合、前述の音声情報を使うため、その音声は人が自然に聞き取れるものだった。 読上げ文章中の単語が単語 DBに記憶されていな力つた場合、前述の音素情報を 合成した合成音声情報を使う。この合成音声情報は音素情報を合成し、更に自然な 音声にするためにアクセントやイントネーションを調整したものである。しかし、この合 成音声情報を使った合成音声は、やはり人が不自然さを感じながら聞き取れるもの たった。
[0003] 先行技術文献として下記のものがある。
特許文献 1:特開平 8 -87698
特許文献 2:特開 2005 - 265477
発明の開示
[0004] (発明が解決しょうとする課題)
単語単位の音声情報を記憶した記憶手段を有する文章読上げ装置にお!、て、音 声情報が記憶されて ヽな ヽため、不自然な合成音声で発話された単語を補足する 機能を有する文章読上げ装置を提供する。
(課題を解決するための手段)
上記の発明が解決しょうとする課題を解決するための第一の手段として、単語単位 の音声情報を記憶した記憶手段を有する文章読み上げ装置において、記憶手段に 記憶されて 、な 、未記憶単語が読み上げ対象文書に存在するかどうかを判断する 判断手段と、判断手段の判断結果に基づ!、て未記憶単語の表記情報を強調して表 示する表示手段を有する。
[0005] 上記の発明が解決しょうとする課題を解決するための第二の手段として、上記の文 章読上げ装置において、表示情報は、未記憶単語と未記憶単語の記号情報である
[0006] 上記の発明が解決しょうとする課題を解決するための第三の手段として、上記の文 章読上げ装置において、表示手段は、外部力 の要求に基づいて表記情報の表示 を終了する。
[0007] 上記の発明が解決しょうとする課題を解決するための第四の手段として、単語単位 の音声情報を記憶した記憶手段を有する文章読み上げ装置を制御する制御方法に ぉ 、て、記憶手段に記憶されて!ヽな 、未記憶単語が読み上げ対象文書に存在する 力どうかを判断する判断ステップと、判断ステップの判断結果に基づ ヽて未記憶単語 の表記情報を強調して表示する表示ステップを有する。
[0008] 上記の発明が解決しょうとする課題を解決するための第五の手段として、上記制御 方法において、表示情報は、未記憶単語と未記憶単語の記号情報である。
[0009] 上記の発明が解決しょうとする課題を解決するための第六の手段として、上記制御 方法において、表示ステップは、外部力もの要求に基づいて表記情報の表示を終了 する。
[0010] 上記の発明が解決しょうとする課題を解決するための第七の手段として、単語単位 の音声情報を記憶した記憶手段を有する文章読み上げ装置を制御する制御プログ ラムにお 、て、記憶手段に記憶されて!、な 、未記憶単語が読み上げ対象文書に存 在するかどうかを判断する判断ステップと、判断ステップの判断結果に基づ 、て未記 憶単語の表記情報を強調して表示する表示ステップを有する。
[0011] 上記の発明が解決しょうとする課題を解決するための第八の手段として、上記制御
プログラムにおいて、表示情報は、未記憶単語と未記憶単語の記号情報である。
[0012] 上記の発明が解決しょうとする課題を解決するための第九の手段として、上記制御 プログラムにおいて、表示ステップは、外部力 の要求に基づいて表記情報の表示 を終了する。
(発明の効果)
音声情報が記憶されて 、な 、ため合成音声で読上げられた単語の意味を完全に 理解できる効果がある。
[0013] また、音声情報が記憶されて 、な!、ため合成音声で読上げられた単語とその単語 の記号情報を表示することで、その合成音声を聞!、た人が合成音声で読上げられた 単語の表示だけでは理解できなカゝつた場合も記号情報に基づいて合成音声の単語 の意味を完全に理解することができる効果がある。
[0014] また、外部からの要求に基づ 、て音声情報が記憶されて!、な 、ため合成音声で読 上げられた単語の表記情報の表示を終了することで、その合成音声を聞 、た人が合 成音声で読上げられた単語の意味を理解するために必要とする時間を調整できる効 果がある。
図面の簡単な説明
[0015] [図 1]文書読上げ装置のハードウェア構成図である。
[図 2]単語 DBの構成図である。
[図 3]音素 DBの構成図である。
[図 4]記号 DBの構成図である。
[図 5]文章読上げ処理の機能ブロック図である。
[図 6]実施例 1における文章読上げ処理のフローチャート(その 1)である。
[図 7]実施例 1における文章読上げ処理のフローチャート(その 2)である。
[図 8]実施例 2における文章読上げ処理のフローチャートである。
[図 9]合成音声の補足画面の表示例である。
符号の説明
[0016] 1 文章読上げ装置
5 §己' I思 B'|5
7 入力部
9 出力部
11 バス
51 文章読上げプログラム
53 単語 DB
55 音素 DB
57 記号 DB
発明を実施するための最良の形態
実施例を説明する前に本発明が必要とされる場面について説明する。上述の不自 然な合成音声で発話された単語を聞!ヽた人は、その単語の意味を直ぐに理解できな いときがある。特に以下の場面ではその単語の意味を直ぐに理解するのは難しいと 考えられる。ここでいう単語とは、文法上で、まとまった意味や機能をもつ言語の最小 単位を意味する。
(1)機械操作や移動しているときで、その単語の意味を確認する時間がない場面
(2)その単語が未知のもので、自然な音声で発話されても意味を理解できな!ヽ場面
(3)その単語を表示するハードウェアが小さぐその単語の文字を確認することが難 し ヽ ή
このため、不自然な合成音声で発話された単語を補足する機能を提供する本発明 が必要となる。
以下に図面を用いて本発明の実施例 1と実施例 2について説明する。
(実施例 1)
[1.ハードウェア構成のブロック図]
図 1は、文章読上げ装置 1のハードウェア構成の一例を示すブロック図である。文章 読上げ装置 1は、 CPU (CentralProcessing Unit) 3と記憶部 5、入力部 7、出力部 9、バス 11で構成されている。 CPU3は、各部の制御や各種の演算を行うものである 。記憶部 5は、文章読上げプログラム 51や単語 DB53、音素 DB55、記号 DB57を格 納するものである。そして、プログラムの実行やデータの記憶を行う RAM (Random
Access Memory)、プログラムやデータの記憶を行う ROM (Read Only Memory )、プログラムやデータを大量に記憶できる外部記憶装置として動作するものである。 文章読上げプログラム 51は、入力部 7から読上げ対象文書と読上げ要求を与えられ ると、単語 DB53や音素 DB55、記号 DB57を使って読上げ処理を行うものである。こ の読上げ処理は、音声情報が記憶されて 、な 、単語の合成音声を補足する機能を 含むものである。単語 DB53は、読上げに使う単語単位の音声情報を記憶したもの である。音素 DB55は、読上げに使う音素情報を記憶したものである。記号 DB57は 、上述の合成音声を補足するための記号情報を記憶したものである。入力部 7は、読 上げ対象文書や文章読上げ処理に対する外部力 の要求を文章読上げ装置 1に与 えるためのものである。具体的には、読上げ対象文書としての電子メールを入力する 通信インターフェースや対象文書の読上げや後述する表記情報の表示終了などの 要求のボタンとして動作可能なものである。出力部 9は、読上げ音声や読上げ音声に 関わる表記情報を外部に送り出すものである。具体的には、スピーカーやモニターと して動作するものである。バス 11は、 CPU3と記憶部 5、入力部 7、出力部 9の間でデ ータを交換するためのものである。また、ここでいう文章とは、文字を連ねて、思想や 感情をひとまとまりにしたものを意味する。
[0018] 以下に文章入力装置 1の動作を簡単に説明する。
(1)入力部 7から読上げ対象文書とそれに対する読上げ要求を与えられる。
(2) CPU3が文章読上げプログラム 51を RAMに展開し、文章読上げプログラム 51 を実行する。そして、文章読上げプログラム 51は、(1)で与えられた読上げ対象文書 と単語 DB53、音素 DB55、記号 DB57を使い、読上げ対象文書の読上げ音声情報 や読上げ音声情報に対応する表記情報を生成する。
(3)出力部 9が(2)で生成した読上げ音声情報や読上げ音声情報に対応する表記 情報を外部に送り出す。
[0019] [1. 1.単語 DBの構成図]
図 2は、単語の音声情報を記憶した単語 DB53を示している。単語 DB53は、文章 読上げ装置 1が対象読み上げ文章で使われている単語の音声情報を抽出するため に使うものである。単語 DB53の情報要素は、単語名 531と音声情報 533、読上げ
時間 535である。単語名 531は、文章読上げ装置 1が対象読み上げ文書で使用され ている単語の音声情報を探すときに使う情報である。音声情報 533は、音声読上げ 装置 1が単語の音を出力部 9から外部に送り出すときに使うものである。この音声情 報は、人が発音した単語の音声を符号化した情報であり、場合によってはそれを更 に圧縮処理したものである。読上げ時間 535は、音声情報 533の読上げに掛かる時 間である。この読上げ時間 535は、文章読上げ装置 1が単語 DB53に記憶されてい ない単語の表記情報を表示する契機を計算するために使用する情報である。
[0020] [1. 2.音素 DBの構成図]
図 3は、音素情報を記憶した音素 DB55を示している。音素 DB55は、文章読上げ 装置 1が単語 DB53に記憶されて 、な 、音声を合成するために使うものである。音素 DB55の情報要素は、音素名 551と、音素情報 553、読上げ時間 555である。音素 名 551は、文章読上げ装置 1が合成の対象となる音素情報を抽出するために使うも のである。音素情報 553は、音声読上げ装置 1が単語 DB53に記憶されていない単 語の音声情報を合成するときに使うものである。この音素情報 553は、人が発音した 単語の音声力 抽出した音素の音を符号ィ匕した情報であり、場合によってはそれを 更に圧縮処理したものである。読上げ時間 555は、音素情報 553の読上げに掛かる 時間である。この読上げ時間 555は、文章読上げ装置 1が単語 DB53に記憶されて V、な 、単語の表記情報を表示する契機を計算するために使用する情報である。
[0021] [1. 3.記号 DBの構成図]
図 4は、単語 DB53に記憶されていない単語の記号を記憶した記号 DB57を示して いる。記号 DB57は、文章読上げ装置 1が単語 DB53に記憶されていないが、対象 読み上げ文書で使用されている単語の意味に関連する記号を表示するために使うも のである。ここでいう記号とは、文字以外のしるしを意味する。記号 DB57の情報要素 は、単語名 571と記号情報 573である。また、ここでいう文字とは、言葉を表すしるし を意味する。単語名 571は、文章読上げ装置 1が対象読み上げ文書で使用されてい る単語の記号情報を探すときに使う情報である。記号情報 573は、音声読上げ装置 1が単語の意味に関連する記号を出力部 9から外部に送り出すときに使うものである 。ここでは、例として会社のロゴマークを格納している。
[0022] [2.機能ブロック図]
図 5は、文章読上げ機能の一例を示す機能ブロック図である。文章読上げ装置 1が 有する文章読上げ機能は、文章読上げプログラム 51が実行されることにより機能す る。その文章読上げ機能は、入力手段 2と判断手段 4、記憶手段 6、発話手段 8、表 示手段 10で構成される。以下に文章読上げ機能の各手段について説明する。
[0023] [入力手段]
入力手段 2は、読上げ対象文書とそれに対する読上げ要求を文章読上げ装置 1に 与える。また、後述する表記情報の表示終了要求を表示手段 10に与える。
[0024] [判断手段]
判断手段 4は、以下の動作を行う。
(1)入力手段 2から与えられた読上げ対象文書と記憶手段 6に記憶されている単語 単位の音声情報や音素情報を使って読上げ文章に対応する全体音声情報を生成 する。また、全体音声情報に合成音声情報が含まれるとき、発話中に監視する合成 音声情報を読上げる契機を設定する。ここでいう合成音声情報とは、記憶手段中に 音声情報が存在しない未記憶単語の音声情報を上述の音素情報を使って生成した ものである。そして全体音声情報を発話手段 8に与える。
(2)未記憶単語の合成音声情報を読み上げる契機を監視する。そして、その契機を 検知したとき、未記憶単語の文字や記号に相当する表記情報を表示手段 10に与え る。
[0025] [記憶手段]
記憶手段 6は、単語単位の音声情報や音素情報、単語単位の記号情報を記憶す る。単語単位の音声情報は、単語 DB53に対応するものである。音素情報は、音素 DB55に対応するものである。記号情報は、記号 DB57に対応するものである。
[0026] [発話手段]
発話手段 8は、判断手段 4から与えられた全体音声情報を音として外部に送り出す
[0027] [表示手段]
表示手段 10は、判断手段 4から与えられた表記情報を文字や記号として外部に送
り出す。また、入力手段 2から与えられた表記情報の表示終了要求により、文字や記 号を外部に送り出す処理を終了する。
[0028] [3.文章読上げ処理]
以下に図 6、 7を使って、実施例 1における文章読上げ処理を説明する。
[0029] S501において、判断手段 4は、入力手段 2から与えられた読上げ情報である読上 げ対象文章を解析する。ここでいう解析とは、読上げ対象文書で使われている単語 の音声情報が音声 DB53に記憶されているかどうかを判定することである。
[0030] S503において、判断手段 4は、読上げ対象文章で使われている全ての単語の中 力も S501の判定で分力つた音声 DB53に音声情報 533が記憶されていない未記憶 単語を抽出する。
[0031] S505において、判断手段 4は、音声 DB53に音声情報が記憶されていない未記 憶単語が存在するかどうかを判定する。判定の結果、音声情報が記憶されていない 未記憶単語が存在するときは S507の処理を行う。判定の結果、音声情報が記憶さ れて 、な 、未記憶単語が存在しな 、ときは S513の処理を行う。
[0032] S507において、判断手段 4は、 S503で抽出した未記憶単語に対応する音素情報 を音素 DB55から抽出する。具体的な抽出方法は、以下の通りである。未登録単語 を文章読上げ装置 1が有している規則情報に基づいて読み方を表す情報であるロー マ字に変換する。そして、そのローマ字に含まれる音素名に対応する音素情報 553 を音素 DB55から抽出する。
[0033] S509において、判断手段 4は、 S507で抽出した音素情報 553を合成して未登録 単語の合成音声情報を生成する。そして、この合成音声が文章読上げ装置 1が有す る振幅しきい値に収まるように編集する。この編集は、合成音声の韻律 (リズム)が自 然に聞こえるように調整するために行うものである。
[0034] S511において、判断手段 4は、対象文書の読上げの中で未記憶単語の合成音声 を読上げる契機を設定する。具体的な設定方法は、以下の通りである。
読上げ対象文章の初めに存在する単語から未記憶単語の前までに存在する単語の 個々の読上げ時間 535を加算し、それらの音声情報を発話するために必要な時間を 計算する。そして、その計算した時間を未記憶単語の表示開始契機として記憶部 5
に記憶する。そして、未記憶単語の合成音声を生成するときに使った音素情報の読 上げ時間 555を加算して合成音声を発話するために必要な時間を計算する。そして 、その計算した時間と上記表示開始契機を加算した時間を未記憶単語の表示終了 契機として記憶部 5に記憶する。読上げ対象文章中に未記憶単語が複数存在すると きは、上述の処理を繰り返す。
[0035] S513において、判断手段 4は、読上げ対象文章全体に対応する全体音声を生成 する。全体音声情報は、単語 DB53の音声情報 533のみをつなぎ合わせて生成する 場合と、単語 DB53の音声情報 533と S509で生成した合成音声情報をつなぎ合わ せて生成する場合がある。そして、この全体音声情報全体としての音の大きさや高さ を文章読上げ装置 1が有する規則情報に基づいて調整する。この調整は、全体音声 情報の音が自然に聞こえるようにするために行うものである。
[0036] S515において、判断手段 4は、 S513で生成した全体音声情報が S509で生成し た合成音声情報を含むものかどうかを判定する。判定の結果、 S513で生成した全体 音声情報が S509で生成した合成音声を含むものときは S519の処理を行う。判定の 結果、 S513で生成した全体音声情報が S509で生成した合成音声情報を含まない もののときは S517の処理において、発話手段 8が全体音声情報を発話する。
[0037] S519において、発話手段 8は、 S513で合成した全体音声の発話を開始する。こ の全体音声情報は、単語 DB53の音声情報 533と S509で合成した合成音声情報 つなぎ合わせて生成したものである。
[0038] S521において、判断手段 4は、 S519で全体音声情報の発話からの経過時間が S 511で計算した表示開始契機に達したかどうかを監視する。この監視は、 S519で発 話を開始した全体音声の経過時間が S511で計算した表示開始契機に達するまで 行う。この監視の結果、 S519で発話を開始した全体音声の経過時間が S511で計 算した表示開始契機に達していたときは、 S523の処理を行う。
[0039] S523において、判断手段 4は、表示開始契機に対応する未記憶単語の記号情報 が記号 DB57中に存在するかどうかを判定する。この判定の結果、未記憶単語の記 号情報が記号 DB57中に存在しないときは、 S525において、表示手段 10は、 S503 で抽出した未記憶単語の文字情報を出力部 9に表示出力する。この判定の結果、未
記憶単語の記号情報が記号 DB57中に存在すときは、 S527において、表示手段 1 0は、 S503で抽出した未記憶単語の文字情報と記号 DB57中の記号情報を出力部 9に表示出力する。
[0040] 以下に図 9を使って、 S525と S527の具体例を説明する。図 9は文章読上げ装置 1 がナビゲーシヨン機能を有するカーナビゲーシヨンシステムとして製品化された場合 を想定したものである。 901は、カーナビゲーシヨンシステムを示している。 903は、読 上げ音声を出力するスピーカーを示している。 905はナビゲーシヨンに使う地図など を表示する画面を示している。 907は、ナビゲーシヨンに使う地図を示している。 909 は、 S525で表示した未記憶単語の文字を示している。ここでは、未記憶単語として 人名を示している。 911は、 S527で表示した記号情報を示している。ここでは、 909 に対応する記号情報として 909の人名に関連する会社のロゴマークを示している。 9 13は、メール読上げボタンを示している。このメール読上げボタンは、カーナビゲー シヨンシステム 1に受信した電子メールを読上げる処理を行わせるときに使うものであ る。 915は、設定ボタンを示している。この設定ボタンは、カーナビゲーシヨンシステム の各種設定を行うときに使うものである。 919は、 907の地図上でのカーナビゲーショ ンシステムを搭載した乗り物の位置を示すしるしである。 S921は、コントローラーを示 している。このコントローラ一は、 907の地図上で目的地を指定するために使うもので ある。 S525で表示する文字情報は、 909に相当するものである。 S527で表示する 文字情報は 909、記号情報は 911に相当するものである。
[0041] S529において、判断手段 4は、 S521で検知した表示開始契機からの経過時間が S511で計算した表示終了契機に達した力どうかを監視する。この監視は、 S521で 検知した表示開始契機力もの経過時間が S511で計算した表示終了契機に達する まで行う。この監視の結果、 S521で検知した表示開始契機からの経過時間が S511 で計算した表示終了契機に達していたときは、 S530において、表示手段 10に表示 して 、る情報の表示を終了する。
(実施例 2)
実施例 2では、実施例 1とは未記憶単語やその未記憶単語に対応する記号の表示 を終了する契機が異なる文章読上げ処理について説明する。
[0042] 未記憶単語表示又は未記憶単語と記号情報表示以前の処理については、実施例 1を同一であるため、その説明を省略する。
[0043] 以下に図 8を使って、実施例 2における文章読上げ処理を説明する。
[0044] S531において、判断手段 4は、 S521で検知した表示開始契機からの経過時間が S511で計算した表示終了契機に達した力どうかを監視する。この監視は、 S521で 検知した表示開始契機力もの経過時間が S511で計算した表示終了契機に達する まで行う。この監視の結果、 S521で検知した表示開始契機からの経過時間が S511 で計算した表示終了契機に達していたときは、 S541の処理を行う。
[0045] S541において、判断手段 4は、外部から未記憶単語やその未記憶単語に対応す る記号の表示を終了させるための終了要求を入力手段 2から受信した力どうかを判 定する。この判定の結果、終了要求を受信したときは、 S530において、表示手段 10 に表示している情報の表示を終了する。この判定の結果、終了要求を受信していな いときは、 S543の処理を行う。
[0046] S543において、判断手段 4は、 S531で検出した表示終了契機からの経過時間が 文章読上げ装置 1が記憶部 5中に有する延長時間に達した力どうかを判定する。この 判定は、 S531で検出した表示終了契機力ゝらの経過時間が延長時間に達するまで行 う。この判定の結果、 S531で検出した表示終了契機からの経過時間が延長時間に 達していたときは、 S530において、表示手段 10に表示している情報の表示を終了 する。
[0047] 以上、本発明を実施例に基づいて説明したが、本発明は前記の実施例に限定され るものではなぐ特許請求の範囲に記載した構成を変更しない限りどのようにでも実 施することができる。
産業上の利用可能性
[0048] 本発明は、テキストファイルなどに記載された文章を読上げる文章読上げ装置にお いて、読上げ音声が不自然だった部分を補足する技術であり、ナビゲーシヨンシステ ムゃ携帯端末などの製品に適用できる。
Claims
[1] 単語単位の音声情報を記憶した記憶手段を有する文章読み上げ装置において、 該記憶手段に記憶されて 、な 、未記憶単語が読み上げ対象文書に存在するかど うかを判断する判断手段と、
該判断手段の判断結果に基づいて未記憶単語の表記情報を強調して表示する表 示手段と、
を有することを特徴とする文章読み上げ装置。
[2] 該表示情報は、該未記憶単語と該未記憶単語の記号情報であることを特徴とする 請求項 1記載の文章読み上げ装置。
[3] 該表示手段は、外部からの要求に基づいて該表記情報の表示を終了することを特 徴とする請求項 1記載の文章読上げ装置。
[4] 単語単位の音声情報を記憶した記憶手段を有する文章読み上げ装置を制御する 制御方法において、
該記憶手段に記憶されて 、な 、未記憶単語が読み上げ対象文書に存在するかど うかを判断する判断ステップと、
該判断ステップの判断結果に基づいて未記憶単語の表記情報を強調して表示す る表示ステップと、
を有することを特徴とする制御方法。
[5] 該表示情報は、該未記憶単語と該未記憶単語の記号情報であることを特徴とする 請求項 4記載の制御方法。
[6] 該表示ステップは、外部からの要求に基づいて該表記情報の表示を終了すること を特徴とする請求項 4記載の制御方法。
[7] 単語単位の音声情報を記憶した記憶手段を有する文章読み上げ装置を制御する 制御プログラムにおいて、
該記憶手段に記憶されて 、な 、未記憶単語が読み上げ対象文書に存在するかど うかを判断する判断ステップと、
該判断ステップの判断結果に基づいて未記憶単語の表記情報を強調して表示す る表示ステップと、
を有することを特徴とする制御プログラム。
[8] 該表示情報は、該未記憶単語と該未記憶単語の記号情報であることを特徴とする 請求項 7記載の制御プログラム。
[9] 該表示ステップは、外部からの要求に基づいて該表記情報の表示を終了すること を特徴とする請求項 7記載の制御プログラム。
Priority Applications (3)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/JP2006/323427 WO2008062529A1 (en) | 2006-11-24 | 2006-11-24 | Sentence reading-out device, method for controlling sentence reading-out device and program for controlling sentence reading-out device |
| JP2008545287A JP4973664B2 (ja) | 2006-11-24 | 2006-11-24 | 文書読上げ装置、文書読上げ装置を制御する制御方法及び文書読上げ装置を制御する制御プログラム |
| US12/463,532 US8315873B2 (en) | 2006-11-24 | 2009-05-11 | Sentence reading aloud apparatus, control method for controlling the same, and control program for controlling the same |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/JP2006/323427 WO2008062529A1 (en) | 2006-11-24 | 2006-11-24 | Sentence reading-out device, method for controlling sentence reading-out device and program for controlling sentence reading-out device |
Related Child Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| US12/463,532 Continuation US8315873B2 (en) | 2006-11-24 | 2009-05-11 | Sentence reading aloud apparatus, control method for controlling the same, and control program for controlling the same |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2008062529A1 true WO2008062529A1 (en) | 2008-05-29 |
Family
ID=39429471
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/JP2006/323427 Ceased WO2008062529A1 (en) | 2006-11-24 | 2006-11-24 | Sentence reading-out device, method for controlling sentence reading-out device and program for controlling sentence reading-out device |
Country Status (3)
| Country | Link |
|---|---|
| US (1) | US8315873B2 (ja) |
| JP (1) | JP4973664B2 (ja) |
| WO (1) | WO2008062529A1 (ja) |
Families Citing this family (11)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP6045175B2 (ja) * | 2012-04-05 | 2016-12-14 | 任天堂株式会社 | 情報処理プログラム、情報処理装置、情報処理方法及び情報処理システム |
| US9942396B2 (en) * | 2013-11-01 | 2018-04-10 | Adobe Systems Incorporated | Document distribution and interaction |
| US9544149B2 (en) | 2013-12-16 | 2017-01-10 | Adobe Systems Incorporated | Automatic E-signatures in response to conditions and/or events |
| US9703982B2 (en) | 2014-11-06 | 2017-07-11 | Adobe Systems Incorporated | Document distribution and interaction |
| US9531545B2 (en) | 2014-11-24 | 2016-12-27 | Adobe Systems Incorporated | Tracking and notification of fulfillment events |
| US9432368B1 (en) | 2015-02-19 | 2016-08-30 | Adobe Systems Incorporated | Document distribution and interaction |
| US9935777B2 (en) | 2015-08-31 | 2018-04-03 | Adobe Systems Incorporated | Electronic signature framework with enhanced security |
| US9626653B2 (en) | 2015-09-21 | 2017-04-18 | Adobe Systems Incorporated | Document distribution and interaction with delegation of signature authority |
| US10347215B2 (en) | 2016-05-27 | 2019-07-09 | Adobe Inc. | Multi-device electronic signature framework |
| US10503919B2 (en) | 2017-04-10 | 2019-12-10 | Adobe Inc. | Electronic signature framework with keystroke biometric authentication |
| KR20210102617A (ko) | 2020-02-12 | 2021-08-20 | 삼성전자주식회사 | 전자 장치 및 그 제어 방법 |
Citations (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPH0635913A (ja) * | 1992-07-21 | 1994-02-10 | Canon Inc | 文章読み上げ装置 |
| JPH10171485A (ja) * | 1996-12-12 | 1998-06-26 | Matsushita Electric Ind Co Ltd | 音声合成装置 |
| JPH10340095A (ja) * | 1997-06-09 | 1998-12-22 | Brother Ind Ltd | 文章読み上げ装置 |
| JP2003308085A (ja) * | 2002-04-15 | 2003-10-31 | Canon Inc | 音声処理装置およびその制御方法、ならびにプログラム |
| JP2004171174A (ja) * | 2002-11-19 | 2004-06-17 | Brother Ind Ltd | 文章読み上げ装置、読み上げのためのプログラム及び記録媒体 |
| JP2006313176A (ja) * | 2005-05-06 | 2006-11-16 | Hitachi Ltd | 音声合成装置 |
Family Cites Families (9)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPH07140996A (ja) * | 1993-11-16 | 1995-06-02 | Fujitsu Ltd | 音声規則合成装置 |
| JPH0887698A (ja) | 1994-09-16 | 1996-04-02 | Alpine Electron Inc | 車載用ナビゲーション装置 |
| JPH10228471A (ja) * | 1996-12-10 | 1998-08-25 | Fujitsu Ltd | 音声合成システム,音声用テキスト生成システム及び記録媒体 |
| US6446041B1 (en) * | 1999-10-27 | 2002-09-03 | Microsoft Corporation | Method and system for providing audio playback of a multi-source document |
| GB2357943B (en) * | 1999-12-30 | 2004-12-08 | Nokia Mobile Phones Ltd | User interface for text to speech conversion |
| US7451087B2 (en) * | 2000-10-19 | 2008-11-11 | Qwest Communications International Inc. | System and method for converting text-to-voice |
| US7913176B1 (en) * | 2003-03-03 | 2011-03-22 | Aol Inc. | Applying access controls to communications with avatars |
| JP4287785B2 (ja) * | 2003-06-05 | 2009-07-01 | 株式会社ケンウッド | 音声合成装置、音声合成方法及びプログラム |
| JP2005265477A (ja) | 2004-03-16 | 2005-09-29 | Matsushita Electric Ind Co Ltd | 車載ナビゲーションシステム |
-
2006
- 2006-11-24 WO PCT/JP2006/323427 patent/WO2008062529A1/ja not_active Ceased
- 2006-11-24 JP JP2008545287A patent/JP4973664B2/ja active Active
-
2009
- 2009-05-11 US US12/463,532 patent/US8315873B2/en not_active Expired - Fee Related
Patent Citations (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPH0635913A (ja) * | 1992-07-21 | 1994-02-10 | Canon Inc | 文章読み上げ装置 |
| JPH10171485A (ja) * | 1996-12-12 | 1998-06-26 | Matsushita Electric Ind Co Ltd | 音声合成装置 |
| JPH10340095A (ja) * | 1997-06-09 | 1998-12-22 | Brother Ind Ltd | 文章読み上げ装置 |
| JP2003308085A (ja) * | 2002-04-15 | 2003-10-31 | Canon Inc | 音声処理装置およびその制御方法、ならびにプログラム |
| JP2004171174A (ja) * | 2002-11-19 | 2004-06-17 | Brother Ind Ltd | 文章読み上げ装置、読み上げのためのプログラム及び記録媒体 |
| JP2006313176A (ja) * | 2005-05-06 | 2006-11-16 | Hitachi Ltd | 音声合成装置 |
Also Published As
| Publication number | Publication date |
|---|---|
| US20090222269A1 (en) | 2009-09-03 |
| JPWO2008062529A1 (ja) | 2010-03-04 |
| US8315873B2 (en) | 2012-11-20 |
| JP4973664B2 (ja) | 2012-07-11 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US8315873B2 (en) | Sentence reading aloud apparatus, control method for controlling the same, and control program for controlling the same | |
| US7490042B2 (en) | Methods and apparatus for adapting output speech in accordance with context of communication | |
| US7062439B2 (en) | Speech synthesis apparatus and method | |
| US7062440B2 (en) | Monitoring text to speech output to effect control of barge-in | |
| US20080319754A1 (en) | Text-to-speech apparatus | |
| CN107871503A (zh) | 语音对话系统以及发声意图理解方法 | |
| JP5431282B2 (ja) | 音声対話装置、方法、プログラム | |
| WO2017068826A1 (ja) | 情報処理装置、情報処理方法、およびプログラム | |
| US20070088547A1 (en) | Phonetic speech-to-text-to-speech system and method | |
| US7792673B2 (en) | Method of generating a prosodic model for adjusting speech style and apparatus and method of synthesizing conversational speech using the same | |
| JP2012163692A (ja) | 音声信号処理システム、音声信号処理方法および音声信号処理方法プログラム | |
| CN106471569B (zh) | 语音合成设备、语音合成方法及其存储介质 | |
| KR20050015585A (ko) | 향상된 음성인식 장치 및 방법 | |
| JP4798039B2 (ja) | 音声対話装置および方法 | |
| JP2007140200A (ja) | 語学学習装置およびプログラム | |
| JP2015087649A (ja) | 発話制御装置、方法、発話システム、プログラム、及び発話装置 | |
| JP4953767B2 (ja) | 音声生成装置 | |
| EP1116217B1 (en) | Voice command navigation of electronic mail reader | |
| JP6825485B2 (ja) | 説明支援プログラム、説明支援方法及び情報処理端末 | |
| JP2006259641A (ja) | 音声認識装置及び音声認識用プログラム | |
| JP3575919B2 (ja) | テキスト音声変換装置 | |
| JP3685648B2 (ja) | 音声合成方法及び音声合成装置、並びに音声合成装置を備えた電話機 | |
| JP4056647B2 (ja) | 波形接続型音声合成装置および方法 | |
| JP7542826B2 (ja) | 音声認識プログラム及び音声認識装置 | |
| JP2009053522A (ja) | 音声出力装置 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 06833231 Country of ref document: EP Kind code of ref document: A1 |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 2008545287 Country of ref document: JP |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 06833231 Country of ref document: EP Kind code of ref document: A1 |