EP1960996A1 - Synthese vocale par concatenation d'unites acoustiques - Google Patents
Synthese vocale par concatenation d'unites acoustiquesInfo
- Publication number
- EP1960996A1 EP1960996A1 EP06841948A EP06841948A EP1960996A1 EP 1960996 A1 EP1960996 A1 EP 1960996A1 EP 06841948 A EP06841948 A EP 06841948A EP 06841948 A EP06841948 A EP 06841948A EP 1960996 A1 EP1960996 A1 EP 1960996A1
- Authority
- EP
- European Patent Office
- Prior art keywords
- text
- elementary
- processing
- operator
- synthesized
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Granted
Links
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L13/00—Speech synthesis; Text to speech systems
- G10L13/02—Methods for producing synthetic speech; Speech synthesisers
- G10L13/033—Voice editing, e.g. manipulating the voice of the synthesiser
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L13/00—Speech synthesis; Text to speech systems
- G10L13/06—Elementary speech units used in speech synthesisers; Concatenation rules
Definitions
- the present invention relates to a system and method for voice synthesis by concatenation of acoustic units and a computer program for implementing the method.
- a speech synthesis system based on a text conventionally comprises input means of the text to be synthesized and linguistic processing means of this text to transform it into a series of phonemes accompanied by prosodic indications.
- These linguistic treatments include syntactic treatments, grapheme-phoneme translations as well as prosodic treatments. They rely on dictionaries as well as rulesets.
- It also includes concatenation synthesis means of prerecorded elements for generating an acoustic signal according to the sequence of phonemes provided by the linguistic processing.
- the Lexitool tool that is part of the catalog of the company Elan Speech, allows to manage an exceptional lexicon.
- the operator enriches the data of the system by adding in the lexicon the words that the system does not pronounce correctly and associating with them the expected pronunciation.
- the object of the invention is therefore to overcome this drawback by proposing an interactive speech synthesis system and method that is easy to use for an operator.
- the object of the invention is a voice synthesis system by concatenation of acoustic units comprising:
- synthesis means by concatenating pre-recorded elements to restore an acoustic signal, as a function of the series of phonemes,
- the linguistic processing means comprise at least one elementary processing unit generating intermediate results of linguistic processing of said text, said elementary processing unit being associated with an editor of the means of inputting and editing, allowing an operator to modify the results of the elementary processing unit, and said voice synthesis system further comprises means for setting the text to be synthesized according to the results modified by the operator, and said linguistic processing means adapting the linguistic processing of the text according to said parameterization.
- the text setting includes tags inserted into the text to be synthesized
- the or each unit of elementary treatment is adapted to perform one of the elementary treatments of all the elementary treatments of: a) validation of the text to be synthesized, b) cutting of the text into sentences, c) cutting of the text in groups of breath, d) - division of text into words, e) - modification of a lexicon of exceptions, f) - phonetization of words, g) - grammatical analysis, h) - prosody.
- the linguistic processing means comprise elementary processing means for performing all of the elementary treatments of said set of elementary processes.
- Another object is a method of concatenating acoustic voice synthesis comprising the steps of:
- the modification of the parameters consists of creating / modifying tags in the text to be synthesized
- the step of generating intermediate results comprises one of the elementary treatment sub-stages:
- said method further comprises a step of selecting the elementary treatment substep to be performed from among the set of elementary treatment substeps; it is executed successively 8 times and each time, a different elementary treatment sub-step is selected in the following order:
- Another object is a computer program comprising program code instructions for performing the steps of the method when said program is executed on a computer.
- the linguistic processing is decomposed for the operator into a series of elementary processes allowing him to control all the parameters having an impact on the quality of the sound flow produced.
- the operator Being able to select the elementary step on which he wishes to intervene, the operator advantageously controls the speech synthesis tool in what appears to him to be the detail of its operation.
- sequence of elementary treatments proposes a logic order of treatment well adapted to the mode of operation of the operator while it does not correspond to the internal operation of the synthesis system.
- FIG. 1 is a block diagram of a speech synthesis system according to one embodiment of the invention.
- FIG. 2 is a flow chart of a speech synthesis method according to one embodiment of the invention.
- FIG. 3 is a variant of the method according to FIG. 2;
- FIG. 4 is a flow chart of a speech synthesis method using the method of FIG. 3 according to an order of presentation of elementary processes.
- a voice synthesis system 1 comprises means 2 for inputting a text to be synthesized. This text is stored in a buffer memory 3 in the form of a record comprising the actual coded text, for example, according to the ISO / IEC 10646 standard, as well as linguistic processing aid parameters, for example in the form of tags. SSML.
- the buffer memory 3 is connected to linguistic processing means 4 of this text. These linguistic processing means 4 are connected to a second buffer 5 in which they store the result of the linguistic processing in the form of a series of phonemes accompanied by prosodic indications.
- This second memory 5 is connected to synthesis means 6 by concatenation of prerecorded elements to restore an acoustic signal as a function of the sequence of phonemes.
- the acoustic signal is transformed into sounds by speakers 7.
- the voice synthesis system 1 comprises means 8 for inputting and editing.
- These input and edit means 8 comprise keyboard-type input means 9 and a pointing tool 10 such as a mouse. They also comprise a display screen 11 and means 12 for controlling these devices 9, 10, 11.
- these input and edit means 8 present to an operator of the voice synthesis system 1 a user-friendly graphical interface.
- the linguistic processing means 4 comprise a unit processing unit chain 4A, 4B, 4C, each of which processes a particular element of the linguistic processing chain such as the division of the text into sentences, the division of the sentences into words, the phonetization of words, grammatical analysis, prosody ...
- Each unit 4A, 4B, 4C of elementary treatment is connected to a specialized editor 8A, 8B, 8C 8 means of input and editing allowing the operator to intervene on the elementary results of the corresponding unit 4A, 4B, 4C to modify them.
- Each pair consisting of a unit 4A, 4B, 4C of elementary processing and its editor 8A, 8B, 8C, constitutes a module 13A, 13B, 13C of processing and editing for a determined stage of linguistic processing.
- the voice synthesis system 1 comprises parameterization means 14 connected to the first buffer memory 3 and to the elementary processing modules 13A, 13B, 13C.
- These setting means 14 add, modify or delete the linguistic processing aid parameters contained in the recording stored in the buffer memory according to the modifications made by the operator on the elementary results of the unit 4A, 4B 1 4C. of elementary processing so that during a subsequent processing of the recording by the same elementary processing units, the elementary result obtained at the output of each unit is the result modified by the operator.
- the means 14 are not suitable for acting on the actual parameter setting of the elementary processing units, nor on the synthesis means 6.
- the speech synthesis system 1 comprises 8 modules corresponding to 8 stages of the linguistic processing of the text.
- the first module deals with the text itself. It allows the operator to validate that the text to be synthesized suits him. Optionally, this module enriches the text with change of voice tags.
- this first module is described in the state of the art, for example in the standardization of the W3C SSML language.
- the second module deals with the division of text into phases.
- the editor shows the operator which phase boundaries can be deleted, moved or inserted.
- the third module deals with splitting into breath groups.
- the publisher highlights breath groups and break times between groups.
- the operator can change the placement of breaks and their durations.
- the fourth module deals with the division into words.
- the publisher highlights the groupings of words that have a link.
- the operator can separate words or group others to form phrases.
- the fifth module deals with the lexicon.
- the operator intervenes on the data by adding, modifying or deleting entries of the exception lexicon.
- the sixth module deals with the phonetization of words.
- the editor presents to the operator the phonetic form or forms of each word on which the system is based to vocalize the text.
- the operator intervenes on the choice of the variants of pronunciation, the connections, the e dumb, ... It should be noted that this module differs from the preceding module on the lexicon in that it does not modify the data but the result of the phonetization process.
- the seventh module deals with grammatical analysis.
- the editor presents the operator with the result of the grammar analysis and the rules that resulted in this result.
- the operator can modify the choice of grammar rules and markers associated with each word or group of words.
- the eighth module is about prosody.
- the editor presents the operator with prosodic information in the form of curves or tables of values that the operator can modify.
- This synthesis 21 comprises successively a linguistic processing step 22 and a concatenation synthesis step 23 as explained above.
- one of the units 4A 1 4B, 4C of elementary processing generates in 24 intermediate results.
- the grammatical analysis means generate a grammatical analysis result accompanied by the rules used.
- the sound result and the intermediate results obtained are presented to the operator at 25.
- the sound result and / or intermediate results are not in accordance with the expectations of the operator, it modifies the intermediate results in 28 using the corresponding interface module.
- the improvement process loops until the operator is satisfied with the result obtained.
- the speech synthesis method further includes a step 30 of selecting the elementary processing module whose intermediate results will be analyzed and possibly modified by the operation.
- the operator can advantageously choose the type of elementary treatment which he wishes to analyze and modify the results.
- FIG. 4 the modifications are made in the order of presentation of the following elementary treatment units.
- the operator starts at 40 by editing the text via the first module associated with the basic processing units of the text itself.
- the operator launches at 41 the second module for cutting the text into sentences.
- This embodiment is remarkable in that it follows a logical order for the operator but does not correspond to the organization of the processing within a linguistic analyzer of a conventional speech synthesis system.
- the operator can also go back to modify the intermediate results of one of the modules already treated, for example because he noticed a mistake late.
Landscapes
- Engineering & Computer Science (AREA)
- Computational Linguistics (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Human Computer Interaction (AREA)
- Physics & Mathematics (AREA)
- Acoustics & Sound (AREA)
- Multimedia (AREA)
- Machine Translation (AREA)
- Document Processing Apparatus (AREA)
Abstract
Description
Claims
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| FR0512854A FR2895133A1 (fr) | 2005-12-16 | 2005-12-16 | Systeme et procede de synthese vocale par concatenation d'unites acoustiques et programme d'ordinateur pour la mise en oeuvre du procede. |
| PCT/FR2006/002745 WO2007071834A1 (fr) | 2005-12-16 | 2006-12-15 | Synthese vocale par concatenation d'unites acoustiques |
Publications (2)
| Publication Number | Publication Date |
|---|---|
| EP1960996A1 true EP1960996A1 (fr) | 2008-08-27 |
| EP1960996B1 EP1960996B1 (fr) | 2010-02-24 |
Family
ID=36716805
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP06841948A Active EP1960996B1 (fr) | 2005-12-16 | 2006-12-15 | Synthese vocale par concatenation d'untes acoustiques |
Country Status (4)
| Country | Link |
|---|---|
| EP (1) | EP1960996B1 (fr) |
| DE (1) | DE602006012540D1 (fr) |
| FR (1) | FR2895133A1 (fr) |
| WO (1) | WO2007071834A1 (fr) |
Families Citing this family (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US11532312B2 (en) * | 2020-12-15 | 2022-12-20 | Microsoft Technology Licensing, Llc | User-perceived latency while maintaining accuracy |
Family Cites Families (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US5860064A (en) * | 1993-05-13 | 1999-01-12 | Apple Computer, Inc. | Method and apparatus for automatic generation of vocal emotion in a synthetic text-to-speech system |
| US6006187A (en) * | 1996-10-01 | 1999-12-21 | Lucent Technologies Inc. | Computer prosody user interface |
-
2005
- 2005-12-16 FR FR0512854A patent/FR2895133A1/fr active Pending
-
2006
- 2006-12-15 EP EP06841948A patent/EP1960996B1/fr active Active
- 2006-12-15 WO PCT/FR2006/002745 patent/WO2007071834A1/fr not_active Ceased
- 2006-12-15 DE DE602006012540T patent/DE602006012540D1/de active Active
Non-Patent Citations (1)
| Title |
|---|
| See references of WO2007071834A1 * |
Also Published As
| Publication number | Publication date |
|---|---|
| WO2007071834A1 (fr) | 2007-06-28 |
| EP1960996B1 (fr) | 2010-02-24 |
| DE602006012540D1 (de) | 2010-04-08 |
| FR2895133A1 (fr) | 2007-06-22 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US8825486B2 (en) | Method and apparatus for generating synthetic speech with contrastive stress | |
| US10347238B2 (en) | Text-based insertion and replacement in audio narration | |
| US9424833B2 (en) | Method and apparatus for providing speech output for speech-enabled applications | |
| US7280968B2 (en) | Synthetically generated speech responses including prosodic characteristics of speech inputs | |
| US20100312565A1 (en) | Interactive tts optimization tool | |
| KR100811568B1 (ko) | 대화형 음성 응답 시스템들에 의해 스피치 이해를 방지하기 위한 방법 및 장치 | |
| US20170300182A9 (en) | Systems and methods for multiple voice document narration | |
| US20130144625A1 (en) | Systems and methods document narration | |
| US8914291B2 (en) | Method and apparatus for generating synthetic speech with contrastive stress | |
| US20030154080A1 (en) | Method and apparatus for modification of audio input to a data processing system | |
| JP2007249212A (ja) | テキスト音声合成のための方法、コンピュータプログラム及びプロセッサ | |
| JP2003295882A (ja) | 音声合成用テキスト構造、音声合成方法、音声合成装置及びそのコンピュータ・プログラム | |
| EP4710326A1 (fr) | Clonage vocal prosodique interlingual dans une pluralite de langues | |
| US7895037B2 (en) | Method and system for trimming audio files | |
| GB2444539A (en) | Altering text attributes in a text-to-speech converter to change the output speech characteristics | |
| JP2003186489A (ja) | 音声情報データベース作成システム,録音原稿作成装置および方法,録音管理装置および方法,ならびにラベリング装置および方法 | |
| EP1960996B1 (fr) | Synthese vocale par concatenation d'untes acoustiques | |
| JP4409279B2 (ja) | 音声合成装置及び音声合成プログラム | |
| JP2009020264A (ja) | 音声合成装置及び音声合成方法並びにプログラム | |
| EP1846918B1 (fr) | Procede d'estimation d'une fonction de conversion de voix | |
| JP2012163721A (ja) | 読み記号列編集装置および読み記号列編集方法 | |
| Mac Lochlainn | Sintéiseoir 1.0: a multidialectical TTS application for Irish | |
| JP6159436B2 (ja) | 読み記号列編集装置および読み記号列編集方法 | |
| WO2007028871A1 (fr) | Systeme de synthese vocale ayant des parametres prosodiques modifiables par un operateur | |
| JPS63208098A (ja) | 音声合成装置および方法 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| 17P | Request for examination filed |
Effective date: 20080613 |
|
| AK | Designated contracting states |
Kind code of ref document: A1 Designated state(s): DE FR GB |
|
| 17Q | First examination report despatched |
Effective date: 20081121 |
|
| DAX | Request for extension of the european patent (deleted) | ||
| RBV | Designated contracting states (corrected) |
Designated state(s): DE FR GB |
|
| GRAP | Despatch of communication of intention to grant a patent |
Free format text: ORIGINAL CODE: EPIDOSNIGR1 |
|
| GRAS | Grant fee paid |
Free format text: ORIGINAL CODE: EPIDOSNIGR3 |
|
| GRAA | (expected) grant |
Free format text: ORIGINAL CODE: 0009210 |
|
| AK | Designated contracting states |
Kind code of ref document: B1 Designated state(s): DE FR GB |
|
| REG | Reference to a national code |
Ref country code: GB Ref legal event code: FG4D Free format text: NOT ENGLISH |
|
| REF | Corresponds to: |
Ref document number: 602006012540 Country of ref document: DE Date of ref document: 20100408 Kind code of ref document: P |
|
| PLBE | No opposition filed within time limit |
Free format text: ORIGINAL CODE: 0009261 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: NO OPPOSITION FILED WITHIN TIME LIMIT |
|
| 26N | No opposition filed |
Effective date: 20101125 |
|
| REG | Reference to a national code |
Ref country code: FR Ref legal event code: PLFP Year of fee payment: 10 |
|
| REG | Reference to a national code |
Ref country code: FR Ref legal event code: PLFP Year of fee payment: 11 |
|
| REG | Reference to a national code |
Ref country code: FR Ref legal event code: PLFP Year of fee payment: 12 |
|
| PGFP | Annual fee paid to national office [announced via postgrant information from national office to epo] |
Ref country code: DE Payment date: 20251126 Year of fee payment: 20 |
|
| PGFP | Annual fee paid to national office [announced via postgrant information from national office to epo] |
Ref country code: GB Payment date: 20251120 Year of fee payment: 20 |
|
| PGFP | Annual fee paid to national office [announced via postgrant information from national office to epo] |
Ref country code: FR Payment date: 20251120 Year of fee payment: 20 |