WO2020196979A1 - 특징 제어 가능 음성 모사를 위한 전자 장치 및 그의 동작 방법 - Google Patents
특징 제어 가능 음성 모사를 위한 전자 장치 및 그의 동작 방법 Download PDFInfo
- Publication number
- WO2020196979A1 WO2020196979A1 PCT/KR2019/004270 KR2019004270W WO2020196979A1 WO 2020196979 A1 WO2020196979 A1 WO 2020196979A1 KR 2019004270 W KR2019004270 W KR 2019004270W WO 2020196979 A1 WO2020196979 A1 WO 2020196979A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- feature
- speaker
- embedding information
- information
- electronic device
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L13/00—Speech synthesis; Text to speech systems
- G10L13/02—Methods for producing synthetic speech; Speech synthesisers
- G10L13/033—Voice editing, e.g. manipulating the voice of the synthesiser
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L13/00—Speech synthesis; Text to speech systems
- G10L13/08—Text analysis or generation of parameters for speech synthesis out of text, e.g. grapheme to phoneme translation, prosody generation or stress or intonation determination
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L17/00—Speaker identification or verification techniques
- G10L17/02—Preprocessing operations, e.g. segment selection; Pattern representation or modelling, e.g. based on linear discriminant analysis [LDA] or principal components; Feature selection or extraction
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L25/00—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
- G10L25/27—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the analysis technique
- G10L25/30—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the analysis technique using neural networks
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L25/00—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
- G10L25/48—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use
- G10L25/51—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use for comparison or discrimination
Definitions
- Various embodiments relate to an electronic device for feature controllable voice simulation and an operating method thereof.
- voice is a meaningful sound used as a means of human communication. Attempts to implement communication through voice between humans and machines have been steadily since the past, and recently, technologies that effectively process voice are applied in real life.
- speech processing techniques include speech recognition, speech synthesis, speaker identification and verification, and the like.
- Speech recognition is a technology that recognizes speech uttered from a speaker and converts it into text
- speech synthesis is a technology that converts text into speech
- speaker authentication is a technology that estimates or authenticates a speaker based on the speech.
- An operation method of an electronic device is for simulating a feature controllable voice, an operation of encoding text, an operation of inferring embedding information related to a speaker's voice signal and feature information, and the inference of the encoded text. It may include an operation of generating a speech signal by decoding together with the embedded information.
- An electronic device is for feature controllable speech simulation, a text encoder for encoding text, a combiner for inferring embedding information related to a speaker's voice signal and feature information, and the inferred embedding of the encoded text It may include a decoder for decoding together with information, and a vocoder for generating a speech signal corresponding to the decoded text.
- the electronic device may convert text into various voices. That is, the electronic device can simulate the voices of multiple speakers by selectively using the speaker's voice signal. In addition, the electronic device may variably simulate the speaker's voice by selectively controlling the characteristics of the speaker's voice signal. For example, the electronic device may express voices with various emotions.
- FIG. 1 is a diagram illustrating an electronic device according to various embodiments.
- FIG. 2 is a diagram illustrating the processor of FIG. 1.
- FIG. 3 is a diagram illustrating a method of operating an electronic device according to various embodiments.
- FIG. 1 is a diagram illustrating an electronic device 100 according to various embodiments.
- an electronic device 100 may include at least one of an input module 110, an output module 120, a memory 130, and a processor 140.
- the input module 110 may receive commands or data to be used for components of the electronic device 100 from outside the electronic device 100.
- the input module 110 includes at least one of an input device configured to directly input a command or data to the electronic device 100 or a communication device configured to receive commands or data by communicating with an external electronic device by wire or wirelessly. It can contain either.
- the input device may include at least one of a microphone, a mouse, a keyboard, and a camera.
- the communication device may include at least one of a wired communication device or a wireless communication device, and the wireless communication device may include at least one of a short-range communication device and a long-distance communication device.
- the output module 120 may provide information to the outside of the electronic device 100.
- the output module 120 includes at least one of an audio output device configured to audibly output information, a display device configured to visually output information, or a communication device configured to transmit information by wired or wireless communication with an external electronic device. It can contain either.
- the communication device may include at least one of a wired communication device or a wireless communication device
- the wireless communication device may include at least one of a short-range communication device and a long-distance communication device.
- the memory 130 may store data used by components of the electronic device 100.
- the data may include input data or output data for a program or a command related thereto.
- the memory 130 may include at least one of a volatile memory or a nonvolatile memory.
- the processor 140 may execute a program in the memory 130 to control components of the electronic device 100 and perform data processing or operation.
- the processor 140 may convert text into an audio signal.
- the processor 140 may convert text into a speech signal based on a neural network for deep learning.
- the processor 140 may simulate the voice of a specific speaker in converting the text into a voice signal.
- the processor 140 may convert text into a voice signal having a speaker's tone. Through this, the processor 140 may generate a voice signal as if text is spoken by a specific speaker.
- the processor 140 may control at least one characteristic in the voice of a specific speaker in converting the text into a voice signal.
- the processor 140 converts text into a speech signal, and may use a speaker speech signal stored in advance and feature information stored in advance.
- the processor 140 may remove the feature element corresponding to the feature information from the speaker's voice signal and apply the feature information instead of the removed feature element. Through this, a correlation between a speaker's voice signal and feature information may be removed.
- the feature information may include at least one of emotion, sex, and age.
- FIG. 2 is a diagram illustrating the processor 140 of FIG. 1.
- the processor 140 includes at least one of a text encoder 210, a speaker encoder 220, a feature encoder 230, a combiner 240, a decoder 250, or a vocoder 260. can do.
- the text encoder 210 may encode text.
- the text may be input through the input module 110.
- text may be directly input from a user through an input device.
- the text may be received from an external electronic device through a communication device.
- the text is stored in the memory 130 and may be fetched from the memory 130 by the processor 140.
- the speaker encoder 220 may encode a speaker voice signal.
- the speaker voice signal is previously stored in the memory 130 and may have a variable length.
- the processor 140 may remove the speech content from the speaker speech signal.
- the speaker encoder 220 may generate speaker embedding information from the speaker voice signal.
- the speaker embedding information may have a fixed length.
- the speaker encoder 220 uses at least one of a recurrent neural network (RNN), a gated recurrent unit (GRU), a long short term memory network (LSTM), or a convolutional neural network (CNN), and speaker embedding information Can be inferred.
- RNN recurrent neural network
- GRU gated recurrent unit
- LSTM long short term memory network
- CNN convolutional neural network
- the speaker encoder 220 may perform additional learning for speaker identification of the speaker's speech signal.
- the speaker encoder 220 may apply a linear discriminate analysis (LDA) to distinguish the speaker of the speaker's speech signal.
- LDA linear discriminate analysis
- the same speaker embedding information may be generated from a plurality of speaker voice signals by the same speaker.
- the feature encoder 230 may encode feature information.
- the characteristic information is previously stored in the memory 130 and may have a variable length.
- the feature information may have discrete or continuous feature variables.
- the feature encoder 230 may generate feature embedding information from the feature information. That is, the feature encoder 230 may infer feature embedding information based on discrete or continuous feature variables related to feature information.
- the feature embedding information may have a fixed length.
- the feature encoder 230 may infer feature embedding information using at least one of RNN, GRU, LSTM, and CNN. For example, the feature encoder 230 may assign an intensity level to the feature information by using the product of one control value related to the feature information, and a plurality of control values related to the feature information are combined. You can also get results.
- the combiner 240 may infer embedding information related to the speaker's voice signal and feature information. In this case, the combiner 240 may generate embedding information by combining speaker embedding information and feature embedding information.
- the combiner 240 may combine speaker embedding information and feature embedding information using at least one of a weight sum, multiplication, or a neural network.
- the neural network may include at least one of RNN, GRU, LSTM, and CNN.
- the decoder 250 may decode the encoded text together with embedding information. That is, the decoder 250 may synthesize the encoded test and embedding information.
- the vocoder 260 may generate a voice signal corresponding to the decoded text.
- the electronic device 100 is for feature controllable speech simulation, and includes a text encoder 210 for encoding text, a combiner 240 for inferring embedding information related to a speaker's voice signal and feature information, It may include a decoder 250 for decoding the encoded text together with the inferred embedding information, and a vocoder 260 for generating a speech signal corresponding to the decoded text.
- the electronic device 100 includes a speaker encoder 220 for inferring speaker embedding information by encoding a speaker voice signal, and a feature encoder 230 for inferring feature embedding information by encoding feature information. ) May be further included.
- the combiner 240 may generate embedding information by combining speaker embedding information and feature embedding information.
- the combiner 240 may combine speaker embedding information and feature embedding information by using at least one of a weight sum, a multiplication, or a neural network.
- the electronic device 100 further includes a processor 140 for removing a feature element corresponding to the feature information from the speaker voice signal so that the correlation between the speaker embedding information and the feature embedding information is removed. I can.
- the feature information may include at least one of emotion, gender, and age.
- the feature encoder 230 may infer feature embedding information based on a discrete or continuous feature variable related to feature information.
- the speaker voice signal may have a variable length
- the speaker embedding information may have a fixed length
- the feature information may have a variable length
- the feature embedding information may have a fixed length
- FIG. 3 is a diagram illustrating a method of operating an electronic device 100 according to various embodiments.
- the electronic device 100 may encode text in operation 310.
- the processor 140 may encode text through the text encoder 210.
- the text may be input through the input module 110.
- text may be directly input from a user through an input device.
- the text may be received from an external electronic device through a communication device.
- the text is stored in the memory 130, and the processor 140 may read the text from the memory 130.
- the electronic device 100 may infer speaker embedding information and feature embedding information, respectively, in operation 320.
- the memory 130 at least one speaker voice signal and at least one control value related to characteristic information may be stored.
- a control value related to a speaker's voice signal and feature information may have a variable length.
- the speaker speech signal of the memory 130 has the same condition that the distribution of speaker embedding information generated from the speaker speech signal must be minimized, and the speaker embedding information generated through shuffle by dividing the speaker speech signal into blocks. The condition of having a value and the condition of generating the same speaker embedding information from a plurality of speaker voice signals by the same speaker may be satisfied.
- the processor 140 may select at least one of a speaker's voice signal and a control value related to feature information.
- the processor 140 may infer speaker embedding information by encoding the speaker voice signal selected through the speaker encoder 220.
- the speaker embedding information may have a fixed length.
- the processor 140 may remove the feature element corresponding to the feature information from the speaker voice signal, and then encode the speaker voice signal. Through this, a correlation between a speaker's voice signal and feature information may be removed.
- the speaker encoder 220 may infer speaker embedding information using at least one of RNN, GRU, LSTM, and CNN.
- the processor 140 may infer feature embedding information by encoding feature information of a control value selected through the feature encoder 230.
- the feature embedding information may have a fixed length.
- the feature encoder 230 may infer feature embedding information using at least one of RNN, GRU, LSTM, and CNN.
- the electronic device 100 may infer embedding information related to the speaker's voice signal and feature information.
- feature information may be applied in place of the feature element removed from the speaker's voice signal. That is, the electronic device 100 may generate embedding information from speaker embedding information and feature embedding information.
- the processor 140 may generate embedding information by combining speaker embedding information and feature embedding information through the combiner 240.
- the combiner 240 may combine speaker embedding information and feature embedding information by using at least one of weight sum, multiplication, and neural network.
- the neural network may include at least one of RNN, GRU, LSTM, and CNN.
- the electronic device 100 may decode the encoded text together with the embedding information in operation 340.
- the processor 140 may decode the text encoded through the decoder 250 together with the embedding information.
- the electronic device 100 may generate a voice signal corresponding to the decoded text in operation 350.
- the processor 140 may generate a voice signal through the vocoder 260.
- the operation method of the electronic device 100 is for feature controllable speech simulation, and includes an operation of encoding text, an operation of inferring embedding information related to a speaker's speech signal and characteristic information, and an encoded text. It may include an operation of generating a speech signal by decoding together with the inferred embedding information.
- the embedding information deduction operation includes an operation of encoding a speaker voice signal to infer speaker embedding information, an operation of encoding feature information to infer feature embedding information, and speaker embedding information and feature embedding information. By combining them, it may include an operation of generating embedding information.
- the operation of generating embedding information may include combining speaker embedding information and feature embedding information by using at least one of weight sum, multiplication, and neural network.
- the operation of inferring speaker embedding information may include an operation of removing a feature element corresponding to the feature information from the speaker voice signal so that a correlation between the speaker embedding information and the feature embedding information is removed.
- the feature information may include at least one of emotion, gender, and age.
- the feature embedding information deduction operation may include an operation of inferring feature embedding information based on a discrete or continuous feature variable related to the feature information.
- the speaker voice signal may have a variable length
- the speaker embedding information may have a fixed length
- the feature information may have a variable length
- the feature embedding information may have a fixed length
- the electronic device 100 may convert text into various voices. That is, the electronic device 100 may simulate the voices of a plurality of speakers by selectively using the speaker's voice signal. In addition, the electronic device 100 may variably simulate the speaker's voice by selectively controlling the characteristics of the speaker's voice signal. For example, the electronic device 100 may express a speaker's voice with various emotions.
- the components are not limited.
- a certain (eg, first) component is “(functionally or communicatively) connected” or “connected” to another (eg, second) component
- the certain component is It may be directly connected to the component, or may be connected through another component (eg, a third component).
- module used in this document includes a unit composed of hardware, software, or firmware, and may be used interchangeably with terms such as, for example, logic, logic blocks, parts, or circuits.
- a module may be an integrally configured component or a minimum unit or a part of one or more functions.
- the module may be configured as an application-specific integrated circuit (ASIC).
- ASIC application-specific integrated circuit
- Various embodiments of the present document are implemented as software including one or more instructions stored in a storage medium (eg, memory 130) readable by a machine (eg, electronic device 100).
- a storage medium eg, memory 130
- the processor of the device may call at least one instruction from among one or more instructions stored from a storage medium and execute it. This enables the device to be operated to perform at least one function according to the at least one command invoked.
- the one or more instructions may include code generated by a compiler or code that can be executed by an interpreter.
- a storage medium that can be read by a device may be provided in the form of a non-transitory storage medium.
- non-transient only means that the storage medium is a tangible device and does not contain a signal (e.g., electromagnetic wave), and this term refers to the case where data is semi-permanently stored in the storage medium. It does not distinguish between temporary storage cases.
- a signal e.g., electromagnetic wave
- each component eg, a module or program of the described components may include a singular number or a plurality of entities.
- one or more components or operations among the above-described corresponding components may be omitted, or one or more other components or operations may be added.
- a plurality of components eg, a module or a program
- the integrated component may perform one or more functions of each component of the plurality of components in the same or similar to that performed by the corresponding component among the plurality of components prior to integration.
- operations performed by a module, program, or other component may be sequentially, parallel, repeatedly, or heuristically executed, or one or more of the operations may be executed in a different order, or omitted. , Or one or more other actions may be added.
Landscapes
- Engineering & Computer Science (AREA)
- Acoustics & Sound (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Human Computer Interaction (AREA)
- Physics & Mathematics (AREA)
- Multimedia (AREA)
- Computational Linguistics (AREA)
- Signal Processing (AREA)
- Artificial Intelligence (AREA)
- Evolutionary Computation (AREA)
- Computer Vision & Pattern Recognition (AREA)
- User Interface Of Digital Computer (AREA)
Abstract
다양한 실시예들에 따른 전자 장치 및 그의 동작 방법은 특징 제어 가능 음성 모사를 위한 것으로, 텍스트를 인코딩하고, 화자 음성 신호 및 특징 정보와 관련된 임베딩 정보를 추론하고, 인코딩된 텍스트를 추론된 임베딩 정보와 함께 디코딩하여, 음성 신호를 발생시키도록 구성될 수 있다.
Description
다양한 실시예들은 특징 제어 가능 음성 모사를 위한 전자 장치 및 그의 동작 방법에 관한 것이다.
일반적으로, 음성은 인간의 의사 소통 수단으로서 이용되는 의미 있는 소리이다. 인간과 기계 사이의 음성을 통한 통신 구현에 대한 시도는 과거부터 꾸준히 있어 왔으며, 최근 음성을 효과적으로 처리하는 기술이 실생활에 적용되고 있다. 예를 들면, 음성을 처리하는 기술은 음성 인식(speech recognition), 음성 합성(speech synthesis), 화자 인증(speaker identification and verification) 등을 포함한다. 음성 인식은 화자로부터 발화된 음성을 인식하여 텍스트로 변환하는 기술이고, 음성 합성은 텍스트를 음성으로 변환하는 기술이며, 화자 인증은 음성에 기반하여 화자를 추정하거나 인증하는 기술이다.
그런데, 상기와 같은 음성 합성에 따르면, 텍스트가 미리 정해진 음성으로만 변환될 뿐이다. 즉 텍스트가 하나의 화자의 음색을 갖는 음성으로만 변환될 뿐이다. 따라서, 텍스트를 다양한 음성으로 변환할 수 있는 방안이 요구된다.
다양한 실시예들에 따른 전자 장치의 동작 방법은 특징 제어 가능 음성 모사를 위한 것으로, 텍스트를 인코딩하는 동작, 화자 음성 신호 및 특징 정보와 관련된 임베딩 정보를 추론하는 동작, 및 상기 인코딩된 텍스트를 상기 추론된 임베딩 정보와 함께 디코딩하여, 음성 신호를 발생시키는 동작을 포함할 수 있다.
다양한 실시예들에 따른 전자 장치는 특징 제어 가능 음성 모사를 위한 것으로, 텍스트를 인코딩하는 텍스트 인코더, 화자 음성 신호 및 특징 정보와 관련된 임베딩 정보를 추론하는 결합부, 상기 인코딩된 텍스트를 상기 추론된 임베딩 정보와 함께 디코딩하는 디코더, 및 상기 디코딩된 텍스트에 대응하는 음성 신호를 발생시키는 보코더를 포함할 수 있다.
다양한 실시예들에 따르면, 전자 장치는 텍스트를 다양한 음성으로 변환할 수 있다. 즉 전자 장치는 화자 음성 신호를 선택적으로 이용하여, 다수의 화자들의 음성을 모사할 수 있다. 아울러, 전자 장치는 화자 음성 신호의 특징을 선택적으로 제어하여, 화자들의 음성을 가변적으로 모사할 수 있다. 예를 들면, 전자 장치는 음성을 다양한 감정으로 표현할 수 있다.
도 1은 다양한 실시예들에 따른 전자 장치를 도시하는 도면이다.
도 2는 도 1의 프로세서를 도시하는 도면이다.
도 3은 다양한 실시예들에 따른 전자 장치의 동작 방법을 도시하는 도면이다.
이하, 본 문서의 다양한 실시예들이 첨부된 도면을 참조하여 설명된다.
도 1은 다양한 실시예들에 따른 전자 장치(100)를 도시하는 도면이다.
도 1을 참조하면, 다양한 실시예들에 따른 전자 장치(100)는, 입력 모듈(110), 출력 모듈(120), 메모리(130) 또는 프로세서(140) 중 적어도 어느 하나를 포함할 수 있다.
입력 모듈(110)은 전자 장치(100)의 구성 요소에 사용될 명령 또는 데이터를 전자 장치(100)의 외부로부터 수신할 수 있다. 입력 모듈(110)은, 사용자가 전자 장치(100)에 직접적으로 명령 또는 데이터를 입력하도록 구성되는 입력 장치 또는 외부 전자 장치와 유선 또는 무선으로 통신하여 명령 또는 데이터를 수신하도록 구성되는 통신 장치 중 적어도 어느 하나를 포함할 수 있다. 예를 들면, 입력 장치는 마이크로폰(microphone), 마우스(mouse), 키보드(keyboard) 또는 카메라(camera) 중 적어도 어느 하나를 포함할 수 있다. 예를 들면, 통신 장치는 유선 통신 장치 또는 무선 통신 장치 중 적어도 어느 하나를 포함하며, 무선 통신 장치는 근거리 통신 장치 또는 원거리 통신 장치 중 적어도 어느 하나를 포함할 수 있다.
출력 모듈(120)은 전자 장치(100)의 외부로 정보를 제공할 수 있다. 출력 모듈(120)은 정보를 청각적으로 출력하도록 구성되는 오디오 출력 장치, 정보를 시각적으로 출력하도록 구성되는 표시 장치 또는 외부 전자 장치와 유선 또는 무선으로 통신하여 정보를 전송하도록 구성되는 통신 장치 중 적어도 어느 하나를 포함할 수 있다. 예를 들면, 통신 장치는 유선 통신 장치 또는 무선 통신 장치 중 적어도 어느 하나를 포함하며, 무선 통신 장치는 근거리 통신 장치 또는 원거리 통신 장치 중 적어도 어느 하나를 포함할 수 있다.
메모리(130)는 전자 장치(100)의 구성 요소에 의해 사용되는 데이터를 저장할 수 있다. 데이터는 프로그램 또는 이와 관련된 명령에 대한 입력 데이터 또는 출력 데이터를 포함할 수 있다. 예를 들면, 메모리(130)는 휘발성 메모리 또는 비휘발성 메모리 중 적어도 어느 하나를 포함할 수 있다.
프로세서(140)는 메모리(130)의 프로그램을 실행하여, 전자 장치(100)의 구성 요소를 제어할 수 있고, 데이터 처리 또는 연산을 수행할 수 있다. 프로세서(140)는 텍스트를 음성 신호로 변환할 수 있다. 여기서, 프로세서(140)는 딥러닝을 위한 신경 회로망을 기반으로, 텍스트를 음성 신호로 변환할 수 있다. 이 때 프로세서(140)는 텍스트를 음성 신호로 변환하는 데 있어서, 특정 화자의 음성을 모사할 수 있다. 예를 들면, 프로세서(140)는 텍스트를 화자의 음색을 갖는 음성 신호로 변환할 수 있다. 이를 통해, 프로세서(140)는, 특정 화자에 의해 텍스트가 발화되는 것과 같이, 음성 신호를 생성할 수 있다. 아울러, 프로세서(140)는 텍스트를 음성 신호로 변환하는 데 있어서, 특정 화자의 음성에서 적어도 하나의 특징을 제어할 수 있다. 이를 위해, 프로세서(140)는 텍스트를 음성 신호로 변환하는 데, 미리 저장된 화자 음성 신호와 미리 저장된 특징 정보를 이용할 수 있다. 이 때 프로세서(140)는 화자 음성 신호에서 특징 정보에 대응하는 특징 요소를 제거하고, 제거된 특징 요소를 대신하여 특징 정보를 적용할 수 있다. 이를 통해, 화자 음성 신호와 특징 정보 간 상관 관계가 제거될 수 있다. 예를 들면, 특징 정보는 감정, 성별 또는 연령 중 적어도 어느 하나를 포함할 수 있다.
도 2는 도 1의 프로세서(140)를 도시하는 도면이다.
도 2를 참조하면, 프로세서(140)는 텍스트 인코더(210), 화자 인코더(220), 특징 인코더(230), 결합부(240), 디코더(250) 또는 보코더(260) 중 적어도 어느 하나를 포함할 수 있다.
텍스트 인코더(210)는 텍스트를 인코딩할 수 있다. 이 때 텍스트는 입력 모듈(110)을 통해 입력될 수 있다. 일 예로, 텍스트는 입력 장치를 통해 사용자로부터 직접적으로 입력될 수 있다. 다른 예로, 텍스트는 통신 장치를 통해 외부 전자 장치로부터 수신될 수 있다. 또는 텍스트는 메모리(130)에 저장되어 있으며, 프로세서(140)에 의해 메모리(130)로부터 인출될 수 있다.
화자 인코더(220)는 화자 음성 신호를 인코딩할 수 있다. 여기서, 화자 음성 신호는 메모리(130)에 미리 저장되어 있으며, 가변적 길이를 가질 수 있다. 예를 들면, 화자 음성 신호는 발화 내용을 포함하지 않으며, 화자 음성 신호가 발화 내용을 포함하는 경우, 프로세서(140)가 화자 음성 신호로부터 발화 내용을 제거할 수 있다. 이를 통해, 화자 인코더(220)는 화자 음성 신호로부터 화자 임베딩 정보를 생성할 수 있다. 여기서, 화자 임베딩 정보는 고정된 길이를 가질 수 있다. 이 때 화자 인코더(220)는 RNN(recurrent neural network), GRU(gated recurrent unit), LSTM(long short term memory network) 또는 CNN(convolutional neural network) 중 적어도 어느 하나의 기법을 사용하여, 화자 임베딩 정보를 추론할 수 있다. 어떤 실시예에서는, 화자 인코더(220)는 화자 음성 신호의 화자 인식(speaker identification)을 위한 추가 학습을 수행할 수 있다. 화자 인코더(220)는 LDA(linear discriminate analysis)를 적용하여, 화자 음성 신호의 화자를 구분할 수 있다. 예를 들면, 동일한 화자에 의한 복수 개의 화자 음성 신호들로부터 동일한 화자 임베딩 정보가 생성될 수 있다.
특징 인코더(230)는 특징 정보를 인코딩할 수 있다. 여기서, 특징 정보는 메모리(130)에 미리 저장되어 있으며, 가변적 길이를 가질 수 있다. 예를 들면, 특징 정보는 이산적인 또는 연속적 특징 변수를 가질 수 있다. 이를 통해, 특징 인코더(230)는 특징 정보로부터 특징 임베딩 정보를 생성할 수 있다. 즉 특징 인코더(230)는 특징 정보와 관련된 이산적 또는 연속적 특징 변수를 기반으로, 특징 임베딩 정보를 추론할 수 있다. 여기서, 특징 임베딩 정보는 고정된 길이를 가질 수 있다. 이 때 특징 인코더(230)는 RNN, GRU, LSTM 또는 CNN 중 적어도 어느 하나의 기법을 사용하여, 특징 임베딩 정보를 추론할 수 있다. 예를 들면, 특징 인코더(230)는 특징 정보와 관련되는 하나의 제어 값의 곱을 사용하여, 특징 정보에 대한 세기 정도를 부여할 수 있으며, 특징 정보와 관련되는 복수 개의 제어 값들을에 대해 복합적인 결과를 획득할 수도 있다.
결합부(240)는 화자 음성 신호 및 특징 정보와 관련된 임베딩 정보를 추론할 수 있다. 이 때 결합부(240)는 화자 임베딩 정보와 특징 임베딩 정보를 결합하여, 임베딩 정보를 생성할 수 있다. 결합부(240)는 가중치 합, 곱셈 또는 신경망(neural network) 중 적어도 어느 하나를 사용하여, 화자 임베딩 정보와 특징 임베딩 정보를 결합할 수 있다. 예를 들면, 신경망은 RNN, GRU, LSTM 또는 CNN 중 적어도 어느 하나를 포함할 수 있다.
디코더(250)는 인코딩된 텍스트를 임베딩 정보와 함께 디코딩할 수 있다. 즉 디코더(250)가 인코딩된 테스트와 임베딩 정보를 합성할 수 있다.
보코더(260)는 디코딩된 텍스트에 대응하는 음성 신호를 발생시킬 수 있다.
다양한 실시예들에 따른 전자 장치(100)는 특징 제어 가능 음성 모사를 위한 것으로, 텍스트를 인코딩하는 텍스트 인코더(210), 화자 음성 신호 및 특징 정보와 관련된 임베딩 정보를 추론하는 결합부(240), 인코딩된 텍스트를 추론된 임베딩 정보와 함께 디코딩하는 디코더(250), 및 디코딩된 텍스트에 대응하는 음성 신호를 발생시키는 보코더(260)를 포함할 수 있다.
다양한 실시예들에 따르면, 전자 장치(100)는, 화자 음성 신호를 인코딩하여, 화자 임베딩 정보를 추론하는 화자 인코더(220), 및 특징 정보를 인코딩하여, 특징 임베딩 정보를 추론하는 특징 인코더(230)를 더 포함할 수 있다.
다양한 실시예들에 따르면, 결합부(240)는, 화자 임베딩 정보와 특징 임베딩 정보를 결합하여, 임베딩 정보를 생성할 수 있다.
다양한 실시예들에 따르면, 결합부(240)는, 가중치 합, 곱셈 또는 신경망 중 적어도 어느 하나를 사용하여, 화자 임베딩 정보와 특징 임베딩 정보를 결합할 수 있다.
다양한 실시예들에 따르면, 전자 장치(100)는 화자 임베딩 정보와 특징 임베딩 정보 간 상과 관계가 제거되도록, 화자 음성 신호로부터 특징 정보에 대응하는 특징 요소를 제거하는 프로세서(140)를 더 포함할 수 있다.
다양한 실시예들에 따르면, 특징 정보는 감정, 성별 또는 연령 중 적어도 어느 하나를 포함할 수 있다.
다양한 실시예들에 따르면, 특징 인코더(230)는 특징 정보와 관련된 이산적 또는 연속적 특징 변수를 기반으로, 특징 임베딩 정보를 추론할 수 있다.
다양한 실시예들에 따르면, 화자 음성 신호는 가변적 길이를 갖고, 화자 임베딩 정보는 고정된 길이를 가질 수 있다.
다양한 실시예들에 따르면, 특징 정보는 가변적 길이를 갖고, 특징 임베딩 정보는 고정된 길이를 가질 수 있다.
도 3은 다양한 실시예들에 따른 전자 장치(100)의 동작 방법을 도시하는 도면이다.
도 3을 참조하면, 전자 장치(100)는 310 동작에서 텍스트를 인코딩할 수 있다. 이 때 프로세서(140)가 텍스트 인코더(210)를 통하여 텍스트를 인코딩할 수 있다. 이 때 텍스트는 입력 모듈(110)을 통해 입력될 수 있다. 일 예로, 텍스트는 입력 장치를 통해 사용자로부터 직접적으로 입력될 수 있다. 다른 예로, 텍스트는 통신 장치를 통해 외부 전자 장치로부터 수신될 수 있다. 또는 텍스트는 메모리(130)에 저장되어 있으며, 프로세서(140)가 메모리(130)로부터 텍스트를 읽어 올 수 있다.
전자 장치(100)는 320 동작에서 화자 임베딩 정보와 특징 임베딩 정보를 각각 추론할 수 있다. 메모리(130)에, 적어도 하나의 화자 음성 신호와 특징 정보와 관련된 적어도 하나의 제어 값이 저장되어 있을 수 있다. 여기서, 화자 음성 신호와 특징 정보와 관련된 제어 값은 가변적인 길이를 가질 수 있다. 메모리(130)의 화자 음성 신호는, 화자 음성 신호로부터 생성되는 화자 임베딩 정보의 분산이 최소가 되어야 하는 조건, 화자 음성 신호가 블록 단위로 분할되어 셔플(shuffle)을 통해 생성되는 화자 임베딩 정보가 동일한 값을 가져야 하는 조건 및 동일한 화자에 의한 복수 개의 화자 음성 신호들로부터 동일한 화자 임베딩 정보가 생성되어야 하는 조건을 충족할 수 있다. 프로세서(140)는 화자 음성 신호 중 어느 하나와 특징 정보와 관련된 제어 값 중 적어도 어느 하나를 선택할 수 있다. 프로세서(140)는 화자 인코더(220)를 통해 선택된 화자 음성 신호를 인코딩하여, 화자 임베딩 정보를 추론할 수 있다. 여기서, 화자 임베딩 정보는 고정된 길이를 가질 수 있다. 이 때 프로세서(140)는 화자 음성 신호에서 특징 정보에 대응하는 특징 요소를 제거한 다음, 화자 음성 신호를 인코딩할 수 있다. 이를 통해, 화자 음성 신호와 특징 정보 간 상관 관계가 제거될 수 있다. 화자 인코더(220)는 RNN, GRU, LSTM 또는 CNN 중 적어도 어느 하나의 기법을 사용하여, 화자 임베딩 정보를 추론할 수 있다. 프로세서(140)는 특징 인코더(230)를 통해 선택된 제어 값의 특징 정보를 인코딩하여, 특징 임베딩 정보를 추론할 수 있다. 여기서, 특징 임베딩 정보는 고정된 길이를 가질 수 있다. 특징 인코더(230)는 RNN, GRU, LSTM 또는 CNN 중 적어도 어느 하나의 기법을 사용하여, 특징 임베딩 정보를 추론할 수 있다.
전자 장치(100)는 330 동작에서 화자 음성 신호 및 특징 정보와 관련된 임베딩 정보를 추론할 수 있다. 이 때 화자 음성 신호에서 제거된 특징 요소를 대신하여 특징 정보가 적용될 수 있다. 즉 전자 장치(100)는 화자 임베딩 정보와 특징 임베딩 정보로부터 임베딩 정보를 생성할 수 있다. 이를 위해, 프로세서(140)가 결합부(240)를 통해 화자 임베딩 정보와 특징 임베딩 정보를 결합하여, 임베딩 정보를 생성할 수 있다. 결합부(240)는 가중치 합, 곱셈 또는 신경망 중 적어도 어느 하나를 사용하여, 화자 임베딩 정보와 특징 임베딩 정보를 결합할 수 있다. 예를 들면, 신경망은 RNN, GRU, LSTM 또는 CNN 중 적어도 어느 하나를 포함할 수 있다.
전자 장치(100)는 340 동작에서 인코딩된 텍스트를 임베딩 정보와 함께 디코딩할 수 있다. 프로세서(140)는 디코더(250)를 통해 인코딩된 텍스트를 임베딩 정보와 함께 디코딩할 수 있다.
전자 장치(100)는 350 동작에서 디코딩된 텍스트에 대응하는 음성 신호를 발생시킬 수 있다. 프로세서(140)는 보코더(260)를 통해 음성 신호를 발생시킬 수 있다.
다양한 실시예들에 따른 전자 장치(100)의 동작 방법은 특징 제어 가능 음성 모사를 위한 것으로, 텍스트를 인코딩하는 동작, 화자 음성 신호 및 특징 정보와 관련된 임베딩 정보를 추론하는 동작, 및 인코딩된 텍스트를 추론된 임베딩 정보와 함께 디코딩하여, 음성 신호를 발생시키는 동작을 포함할 수 있다.
다양한 실시예들에 따르면, 임베딩 정보 추론 동작은, 화자 음성 신호를 인코딩하여, 화자 임베딩 정보를 추론하는 동작, 특징 정보를 인코딩하여, 특징 임베딩 정보를 추론하는 동작, 및 화자 임베딩 정보와 특징 임베딩 정보를 결합하여, 임베딩 정보를 생성하는 동작을 포함할 수 있다.
다양한 실시예들에 따르면, 임베딩 정보 생성 동작은, 가중치 합, 곱셈 또는 신경망 중 적어도 어느 하나를 사용하여, 화자 임베딩 정보와 특징 임베딩 정보를 결합하는 동작을 포함할 수 있다.
다양한 실시예들에 따르면, 화자 임베딩 정보 추론 동작은, 화자 임베딩 정보와 특징 임베딩 정보 간 상과 관계가 제거되도록, 화자 음성 신호로부터 특징 정보에 대응하는 특징 요소를 제거하는 동작을 포함할 수 있다.
다양한 실시예들에 따르면, 특징 정보는 감정, 성별 또는 연령 중 적어도 어느 하나를 포함할 수 있다.
다양한 실시예들에 따르면, 특징 임베딩 정보 추론 동작은, 특징 정보와 관련된 이산적 또는 연속적 특징 변수를 기반으로, 특징 임베딩 정보를 추론하는 동작을 포함할 수 있다.
다양한 실시예들에 따르면, 화자 음성 신호는 가변적 길이를 갖고, 화자 임베딩 정보는 고정된 길이를 가질 수 있다.
다양한 실시예들에 따르면, 특징 정보는 가변적 길이를 갖고, 상기 특징 임베딩 정보는 고정된 길이를 가질 수 있다.
다양한 실시예들에 따르면, 전자 장치(100)는 텍스트를 다양한 음성으로 변환할 수 있다. 즉 전자 장치(100)는 화자 음성 신호를 선택적으로 이용하여, 다수의 화자들의 음성을 모사할 수 있다. 아울러, 전자 장치(100)는 화자 음성 신호의 특징을 선택적으로 제어하여, 화자들의 음성을 가변적으로 모사할 수 있다. 예를 들면, 전자 장치(100)는 한 화자의 음성을 다양한 감정으로 표현할 수 있다.
본 문서의 다양한 실시예들 및 이에 사용된 용어들은 본 문서에 기재된 기술을 특정한 실시 형태에 대해 한정하려는 것이 아니며, 해당 실시 예의 다양한 변경, 균등물, 및/또는 대체물을 포함하는 것으로 이해되어야 한다. 도면의 설명과 관련하여, 유사한 구성요소에 대해서는 유사한 참조 부호가 사용될 수 있다. 단수의 표현은 문맥상 명백하게 다르게 뜻하지 않는 한, 복수의 표현을 포함할 수 있다. 본 문서에서, "A 또는 B", "A 및/또는 B 중 적어도 하나", "A, B 또는 C" 또는 "A, B 및/또는 C 중 적어도 하나" 등의 표현은 함께 나열된 항목들의 모든 가능한 조합을 포함할 수 있다. "제 1", "제 2", "첫째" 또는 "둘째" 등의 표현들은 해당 구성요소들을, 순서 또는 중요도에 상관없이 수식할 수 있고, 한 구성요소를 다른 구성요소와 구분하기 위해 사용될 뿐 해당 구성요소들을 한정하지 않는다. 어떤(예: 제 1) 구성요소가 다른(예: 제 2) 구성요소에 "(기능적으로 또는 통신적으로) 연결되어" 있다거나 "접속되어" 있다고 언급된 때에는, 상기 어떤 구성요소가 상기 다른 구성요소에 직접적으로 연결되거나, 다른 구성요소(예: 제 3 구성요소)를 통하여 연결될 수 있다.
본 문서에서 사용된 용어 "모듈"은 하드웨어, 소프트웨어 또는 펌웨어로 구성된 유닛을 포함하며, 예를 들면, 로직, 논리 블록, 부품, 또는 회로 등의 용어와 상호 호환적으로 사용될 수 있다. 모듈은, 일체로 구성된 부품 또는 하나 또는 그 이상의 기능을 수행하는 최소 단위 또는 그 일부가 될 수 있다. 예를 들면, 모듈은 ASIC(application-specific integrated circuit)으로 구성될 수 있다.
본 문서의 다양한 실시예들은 기기(machine)(예: 전자 장치(100))에 의해 읽을 수 있는 저장 매체(storage medium)(예: 메모리(130))에 저장된 하나 이상의 명령어들을 포함하는 소프트웨어로서 구현될 수 있다. 예를 들면, 기기의 프로세서(예: 프로세서(140))는, 저장 매체로부터 저장된 하나 이상의 명령어들 중 적어도 하나의 명령을 호출하고, 그것을 실행할 수 있다. 이것은 기기가 호출된 적어도 하나의 명령어에 따라 적어도 하나의 기능을 수행하도록 운영되는 것을 가능하게 한다. 하나 이상의 명령어들은 컴파일러에 의해 생성된 코드 또는 인터프리터에 의해 실행될 수 있는 코드를 포함할 수 있다. 기기로 읽을 수 있는 저장매체 는, 비일시적(non-transitory) 저장매체의 형태로 제공될 수 있다. 여기서, ‘비일시적’은 저장매체가 실재(tangible)하는 장치이고, 신호(signal)(예: 전자기파)를 포함하지 않는다는 것을 의미할 뿐이며, 이 용어는 데이터가 저장매체에 반영구적으로 저장되는 경우와 임시적으로 저장되는 경우를 구분하지 않는다.
다양한 실시예들에 따르면, 기술한 구성요소들의 각각의 구성요소(예: 모듈 또는 프로그램)는 단수 또는 복수의 개체를 포함할 수 있다. 다양한 실시예들에 따르면, 전술한 해당 구성요소들 중 하나 이상의 구성요소들 또는 동작들이 생략되거나, 또는 하나 이상의 다른 구성요소들 또는 동작들이 추가될 수 있다. 대체적으로 또는 추가적으로, 복수의 구성요소들(예: 모듈 또는 프로그램)은 하나의 구성요소로 통합될 수 있다. 이런 경우, 통합된 구성요소는 복수의 구성요소들 각각의 구성요소의 하나 이상의 기능들을 통합 이전에 복수의 구성요소들 중 해당 구성요소에 의해 수행되는 것과 동일 또는 유사하게 수행할 수 있다. 다양한 실시예들에 따르면, 모듈, 프로그램 또는 다른 구성요소에 의해 수행되는 동작들은 순차적으로, 병렬적으로, 반복적으로, 또는 휴리스틱하게 실행되거나, 동작들 중 하나 이상이 다른 순서로 실행되거나, 생략되거나, 또는 하나 이상의 다른 동작들이 추가될 수 있다.
Claims (15)
- 특징 제어 가능 음성 모사를 위한 전자 장치의 동작 방법에 있어서,텍스트를 인코딩하는 동작;화자 음성 신호 및 특징 정보와 관련된 임베딩 정보를 추론하는 동작; 및상기 인코딩된 텍스트를 상기 추론된 임베딩 정보와 함께 디코딩하여, 음성 신호를 발생시키는 동작을 포함하는 방법.
- 제 1 항에 있어서, 상기 임베딩 정보 추론 동작은,상기 화자 음성 신호를 인코딩하여, 화자 임베딩 정보를 추론하는 동작;상기 특징 정보를 인코딩하여, 특징 임베딩 정보를 추론하는 동작; 및상기 화자 임베딩 정보와 상기 특징 임베딩 정보를 결합하여, 상기 임베딩 정보를 생성하는 동작을 포함하는 방법.
- 제 2 항에 있어서, 상기 임베딩 정보 생성 동작은,가중치 합, 곱셈 또는 신경망 중 적어도 어느 하나를 사용하여, 상기 화자 임베딩 정보와 상기 특징 임베딩 정보를 결합하는 동작을 포함하는 방법.
- 제 2 항에 있어서, 상기 화자 임베딩 정보 추론 동작은,상기 화자 임베딩 정보와 상기 특징 임베딩 정보 간 상과 관계가 제거되도록, 상기 화자 음성 신호로부터 상기 특징 정보에 대응하는 특징 요소를 제거하는 동작을 포함하는 방법.
- 제 1 항에 있어서,상기 특징 정보는 감정, 성별 또는 연령 중 적어도 어느 하나를 포함하는 방법.
- 제 2 항에 있어서, 상기 특징 임베딩 정보 추론 동작은,상기 특징 정보와 관련된 이산적 또는 연속적 특징 변수를 기반으로, 상기 특징 임베딩 정보를 추론하는 동작을 포함하는 방법.
- 제 2 항에 있어서,상기 화자 음성 신호와 상기 특징 정보는 각각 가변적 길이를 갖고,상기 화자 임베딩 정보와 상기 특징 임베딩 정보는 각각 고정된 길이를 갖는 방법.
- 특징 제어 가능 음성 모사를 위한 전자 장치에 있어서,텍스트를 인코딩하는 텍스트 인코더;화자 음성 신호 및 특징 정보와 관련된 임베딩 정보를 추론하는 결합부;상기 인코딩된 텍스트를 상기 추론된 임베딩 정보와 함께 디코딩하는 디코더; 및상기 디코딩된 텍스트에 대응하는 음성 신호를 발생시키는 보코더를 포함하는 전자 장치.
- 제 8 항에 있어서,상기 화자 음성 신호를 인코딩하여, 화자 임베딩 정보를 추론하는 화자 인코더; 및상기 특징 정보를 인코딩하여, 특징 임베딩 정보를 추론하는 특징 인코더를 더 포함하는 전자 장치.
- 제 9 항에 있어서, 상기 결합부는,상기 화자 임베딩 정보와 상기 특징 임베딩 정보를 결합하여, 상기 임베딩 정보를 생성하는 전자 장치.
- 제 10 항에 있어서, 상기 결합부는,가중치 합, 곱셈 또는 신경망 중 적어도 어느 하나를 사용하여, 상기 화자 임베딩 정보와 상기 특징 임베딩 정보를 결합하는 전자 장치.
- 제 10 항에 있어서,상기 화자 임베딩 정보와 상기 특징 임베딩 정보 간 상관 관계가 제거되도록, 상기 화자 음성 신호로부터 상기 특징 정보에 대응하는 특징 요소를 제거하는 프로세서를 더 포함하는 전자 장치.
- 제 8 항에 있어서,상기 특징 정보는 감정, 성별 또는 연령 중 적어도 어느 하나를 포함하는 전자 장치.
- 제 9 항에 있어서, 상기 특징 인코더는,상기 특징 정보와 관련된 이산적 또는 연속적 특징 변수를 기반으로, 상기 특징 임베딩 정보를 추론하는 전자 장치.
- 제 9 항에 있어서,상기 화자 음성 신호와 상기 특징 정보는 각각 가변적 길이를 갖고,상기 화자 임베딩 정보와 상기 특징 임베딩 정보는 각각 고정된 길이를 갖는 전자 장치.
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| KR10-2019-0033403 | 2019-03-25 | ||
| KR1020190033403A KR102221260B1 (ko) | 2019-03-25 | 2019-03-25 | 특징 제어 가능 음성 모사를 위한 전자 장치 및 그의 동작 방법 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2020196979A1 true WO2020196979A1 (ko) | 2020-10-01 |
Family
ID=72610636
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/KR2019/004270 Ceased WO2020196979A1 (ko) | 2019-03-25 | 2019-04-10 | 특징 제어 가능 음성 모사를 위한 전자 장치 및 그의 동작 방법 |
Country Status (2)
| Country | Link |
|---|---|
| KR (1) | KR102221260B1 (ko) |
| WO (1) | WO2020196979A1 (ko) |
Families Citing this family (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| KR102477444B1 (ko) * | 2020-12-07 | 2022-12-15 | 서울대학교산학협력단 | 화자 외 정보가 제거된 화자 임베딩 장치 및 방법 |
| KR20240059350A (ko) * | 2022-10-27 | 2024-05-07 | 삼성전자주식회사 | 음성 신호 비식별화 처리 방법 및 그 전자 장치 |
Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2003233388A (ja) * | 2002-02-07 | 2003-08-22 | Sharp Corp | 音声合成装置および音声合成方法、並びに、プログラム記録媒体 |
| US20090319275A1 (en) * | 2007-03-20 | 2009-12-24 | Fujitsu Limited | Speech synthesizing device, speech synthesizing system, language processing device, speech synthesizing method and recording medium |
| KR101097186B1 (ko) * | 2010-03-03 | 2011-12-22 | 미디어젠(주) | 대화체 앞뒤 문장정보를 이용한 다국어 음성합성 시스템 및 방법 |
| KR20150017662A (ko) * | 2013-08-07 | 2015-02-17 | 삼성전자주식회사 | 텍스트-음성 변환 방법, 장치 및 저장 매체 |
| KR20190016889A (ko) * | 2017-08-09 | 2019-02-19 | 한국과학기술원 | 텍스트-음성 변환 방법 및 시스템 |
-
2019
- 2019-03-25 KR KR1020190033403A patent/KR102221260B1/ko active Active
- 2019-04-10 WO PCT/KR2019/004270 patent/WO2020196979A1/ko not_active Ceased
Patent Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2003233388A (ja) * | 2002-02-07 | 2003-08-22 | Sharp Corp | 音声合成装置および音声合成方法、並びに、プログラム記録媒体 |
| US20090319275A1 (en) * | 2007-03-20 | 2009-12-24 | Fujitsu Limited | Speech synthesizing device, speech synthesizing system, language processing device, speech synthesizing method and recording medium |
| KR101097186B1 (ko) * | 2010-03-03 | 2011-12-22 | 미디어젠(주) | 대화체 앞뒤 문장정보를 이용한 다국어 음성합성 시스템 및 방법 |
| KR20150017662A (ko) * | 2013-08-07 | 2015-02-17 | 삼성전자주식회사 | 텍스트-음성 변환 방법, 장치 및 저장 매체 |
| KR20190016889A (ko) * | 2017-08-09 | 2019-02-19 | 한국과학기술원 | 텍스트-음성 변환 방법 및 시스템 |
Also Published As
| Publication number | Publication date |
|---|---|
| KR20200113364A (ko) | 2020-10-07 |
| KR102221260B1 (ko) | 2021-03-02 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| CN112735373A (zh) | 语音合成方法、装置、设备及存储介质 | |
| WO2020196978A1 (ko) | 멀티스케일 음성 감정 인식을 위한 전자 장치 및 그의 동작 방법 | |
| CN114338623B (zh) | 音频的处理方法、装置、设备及介质 | |
| WO2020256475A1 (ko) | 텍스트를 이용한 발화 동영상 생성 방법 및 장치 | |
| WO2021132797A1 (ko) | 반지도 학습 기반 단어 단위 감정 임베딩과 장단기 기억 모델을 이용한 대화 내에서 발화의 감정 분류 방법 | |
| CN116386605B (zh) | 模型训练方法和装置、语音合成方法、设备及存储介质 | |
| JP2017204067A (ja) | 手話会話支援システム | |
| CN113555032A (zh) | 多说话人场景识别及网络训练方法、装置 | |
| KR102137523B1 (ko) | 텍스트-음성 변환 방법 및 시스템 | |
| Mian Qaisar | Isolated speech recognition and its transformation in visual signs | |
| WO2020196976A1 (ko) | 멀티모달 데이터를 이용한 주의집중의 순환 신경망 기반 전자 장치 및 그의 동작 방법 | |
| KR102221260B1 (ko) | 특징 제어 가능 음성 모사를 위한 전자 장치 및 그의 동작 방법 | |
| WO2023229117A1 (ko) | 대화형 가상 아바타의 구현 방법 | |
| JP2021177228A (ja) | 多言語多話者個性表現音声合成のための電子装置およびこの処理方法 | |
| CN119580695B (zh) | 一种自适应情感驱动的音色克隆文字转语音方法及装置 | |
| WO2023095988A1 (ko) | 대화 상대방의 성격정보를 고려하여 신뢰도 증강을 위한 맞춤형 대화 생성 시스템 및 그 방법 | |
| KR20190140803A (ko) | 감정 임베딩과 순환형 신경망을 이용한 대화 시스템 및 방법 | |
| WO2022145611A1 (ko) | 전자 장치 및 그 제어 방법 | |
| CN112151073B (zh) | 一种语音处理方法、系统、设备及介质 | |
| US20260113307A1 (en) | Soundless speech recognition method, system and device | |
| CN116959496A (zh) | 语音情感变化识别方法、装置、电子设备及介质 | |
| CN114945105A (zh) | 一种结合声音补偿下的无线耳机音频滞后性抵消方法 | |
| WO2021045434A1 (ko) | 전자 장치 및 이의 제어 방법 | |
| WO2021060591A1 (ko) | 캐릭터 발화 맥락에 따른 음성합성 모델 변경장치 | |
| CN114708871A (zh) | 声纹识别模型的训练方法、系统以及声纹识别方法和系统 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 19922047 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 19922047 Country of ref document: EP Kind code of ref document: A1 |