WO2024193227A1 - 语音编辑方法、装置、存储介质及电子装置 - Google Patents
语音编辑方法、装置、存储介质及电子装置 Download PDFInfo
- Publication number
- WO2024193227A1 WO2024193227A1 PCT/CN2024/074070 CN2024074070W WO2024193227A1 WO 2024193227 A1 WO2024193227 A1 WO 2024193227A1 CN 2024074070 W CN2024074070 W CN 2024074070W WO 2024193227 A1 WO2024193227 A1 WO 2024193227A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- audio
- target
- voice
- text
- editing
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L13/00—Speech synthesis; Text to speech systems
- G10L13/02—Methods for producing synthetic speech; Speech synthesisers
- G10L13/027—Concept to speech synthesisers; Generation of natural phrases from machine-based concepts
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L13/00—Speech synthesis; Text to speech systems
- G10L13/08—Text analysis or generation of parameters for speech synthesis out of text, e.g. grapheme to phoneme translation, prosody generation or stress or intonation determination
- G10L13/10—Prosody rules derived from text; Stress or intonation
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L19/00—Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
- G10L19/04—Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis using predictive techniques
- G10L19/16—Vocoder architecture
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L25/00—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
- G10L25/27—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the analysis technique
- G10L25/30—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the analysis technique using neural networks
Definitions
- the present disclosure relates to computer technology and artificial intelligence technology, and in particular to a voice editing method, device, storage medium and electronic device.
- At least some embodiments of the present disclosure provide a speech editing method, apparatus, storage medium and electronic device to at least solve the technical problem that the speech editing method provided in the related art has a mismatch between training and testing, resulting in low fluency and poor realism in the speech editing results.
- a speech editing method comprising: obtaining original audio and target text to be processed, wherein the target text is used to determine text content to be edited into the original audio; performing speech masking on a portion of the audio to be edited in the original audio to obtain a first masked audio; and performing speech editing on the target text and the first masked audio to obtain the target audio.
- a voice editing method comprises: in response to a trigger operation performed on the voice editing control, a voice editing interface pops up; in response to an input operation performed on the voice editing interface, original audio and target text are imported, wherein the target text is used to determine the text content to be edited into the original audio; in response to An editing operation is performed on the voice editing interface to select a portion of the audio to be edited from the original audio; in response to a play operation performed on the voice editing interface, a target audio is played in a game scene, wherein the target audio is obtained by voice editing the target text and the masked audio, and the masked audio is obtained by voice masking the portion of the audio to be edited.
- a model training method including: obtaining training audio and training text to be processed, wherein the training text is used to determine the text content to be edited into the training audio; performing speech masking on the portion of the audio to be edited in the training audio to obtain the masked training audio; using the masked training audio and training text to train an initial speech editing model to obtain a target speech editing model, wherein the target speech editing model is used to perform speech editing on the target text and the masked original audio to obtain the target audio, and the masked original audio is obtained by performing speech masking on the portion of the audio to be edited in the original audio.
- a speech editing device including: an acquisition module, used to acquire the original audio and target text to be processed, wherein the target text is used to determine the text content to be edited into the original audio; a masking module, used to perform speech masking on the part of the audio to be edited in the original audio to obtain a first masked audio; and an editing module, used to perform speech editing on the target text and the first masked audio to obtain the target audio.
- a computer-readable storage medium in which a computer program is stored, wherein the computer program is configured to execute the above-mentioned speech editing method or model training method when running.
- an electronic device including: a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the above-mentioned speech editing method or model training method.
- the target text is used to determine the text content to be edited into the original audio; a first masked audio is obtained by voice masking the part of the audio to be edited in the original audio; and the target text and the first masked audio are further voice edited to obtain the target audio, thereby achieving the purpose of obtaining the target audio by first voice masking the original audio to be voice edited and then voice editing, thereby achieving the technical effect of improving the fluency and realism of the voice editing results, and further solving the technical problem of low fluency and poor realism of the voice editing results caused by the mismatch between training and testing of the voice editing method provided in the related art.
- FIG1 is a hardware structure block diagram of a mobile terminal of a voice editing method according to one embodiment of the present disclosure
- FIG2 is a flow chart of a voice editing method according to one embodiment of the present disclosure.
- FIG3 is a schematic diagram of an optional voice editing process according to one embodiment of the present disclosure.
- FIG4 is a schematic diagram of an optional acoustic feature extraction process according to one embodiment of the present disclosure.
- FIG5 is a flow chart of an optional voice editing method according to one embodiment of the present disclosure.
- FIG6 is a schematic diagram of an optional voice editing on a cloud server according to one embodiment of the present disclosure.
- FIG7 is a flow chart of another voice editing method according to one embodiment of the present disclosure.
- FIG8 is a flow chart of a model training method according to one embodiment of the present disclosure.
- FIG9 is a structural block diagram of a voice editing device according to one embodiment of the present disclosure.
- FIG10 is a structural block diagram of an optional voice editing device according to one embodiment of the present disclosure.
- FIG. 11 is a schematic diagram of an electronic device according to one embodiment of the present disclosure.
- Text-to-Speech is a technology that uses a text-to-speech algorithm to convert text data into human voice audio.
- TTS can be used to generate a virtual human voice with specific vocal cavity, emotions, and characteristics based on an existing voice library.
- Voice editing refers to the processing and modification of voice to achieve the required format and quality.
- Voice editing may include: processing voice files (such as replacement, deletion, insertion, cutting, connection, calibration, etc.); changing sound characteristics (such as enhancing low frequencies or amplitude, reducing background noise, etc.); setting expression marks (including position, annotations, unit names). The above voice editing can help understand the specific voices or instruments used in the work.
- the game scenario applied to the embodiment of the present disclosure can be any application scenario involving voice synthesis or voice editing in the field of computer technology or artificial intelligence technology, and the game types targeted can be action, adventure, simulation, role-playing, leisure, etc.
- the disclosed embodiment proposes a speech editing method, which adopts the technical concept of performing speech masking on the original audio before speech editing to unify the speech editing training objectives and test objectives, thereby achieving the technical effect of improving the fluency and realism of the speech editing results, and further solving the technical problem that the speech editing methods provided in the related arts have mismatches in training and testing, resulting in low fluency and poor realism of the speech editing results.
- a terminal device e.g., a mobile terminal, a computer terminal, or a similar computing device.
- the mobile terminal can be a smart phone, a tablet computer, a PDA, a mobile Internet device, a game console, and other terminal devices.
- FIG1 is a hardware structure block diagram of a mobile terminal of a voice editing method according to one embodiment of the present disclosure.
- the mobile terminal may include one or more (only one is shown in FIG1 ) processors 102, memory 104, transmission device 106, input/output device 108, and display device 110.
- the processor 102 calls and runs a computer program stored in the memory 104 to execute the voice editing method, and the generated target audio is transmitted to the input/output device 108 and/or the display device 110 through the transmission device 106, and then the target audio is provided to the player.
- the processor 102 may include but is not limited to: a central processing unit (CPU); Processing devices such as CPU, Graphics Processing Unit (GPU), Digital Signal Processing (DSP) chip, Microcontroller Unit (MCU), Field Programmable Gate Array (FPGA), Neural-Network Processing Unit (NPU), Tensor Processing Unit (TPU), Artificial Intelligence (AI) type processor, etc.
- CPU central processing unit
- GPU Graphics Processing Unit
- DSP Digital Signal Processing
- MCU Microcontroller Unit
- FPGA Field Programmable Gate Array
- NPU Neural-Network Processing Unit
- TPU Tensor Processing Unit
- AI Artificial Intelligence
- FIG1 is merely illustrative and does not limit the structure of the mobile terminal.
- the mobile terminal may include more or fewer components than those shown in FIG1 , or may have a different configuration than that shown in FIG1 .
- the above-mentioned terminal device can also provide a human-computer interaction interface with a touch-sensitive surface, which can sense finger contact and/or gestures to perform human-computer interaction with a graphical user interface (Graphical User Interface, GUI).
- the human-computer interaction function can include the following interactions: creating web pages, drawing, word processing, making electronic documents, games, video conferencing, instant messaging, sending and receiving emails, call interface, playing digital videos, playing digital music and/or web browsing, etc.
- the executable instructions for executing the above-mentioned human-computer interaction functions are configured/stored in a computer program product executable by one or more processors or a readable storage medium.
- the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (Content Delivery Network, CDN), and basic cloud computing services such as big data and artificial intelligence platforms.
- the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (Content Delivery Network, CDN), and basic cloud computing services such as big data and artificial intelligence platforms.
- CDN Content Delivery Network
- the electronic game server can obtain the target audio in the electronic game scene based on the voice editing method, and provide the target audio to the player (for example, it can be rendered and displayed on the display screen of the player's terminal, or provided to the player through holographic projection, etc.).
- an embodiment of a voice editing method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
- FIG. 2 is a flow chart of a voice editing method according to one embodiment of the present disclosure. As shown in FIG. 2 , the method includes the following steps:
- Step S21 obtaining the original audio and target text to be processed, wherein the target text is used to determine the original audio to be edited.
- the voice editing method provided by the present disclosure can be applied to, but is not limited to, the following application scenarios: electronic games (for example, editing the original audio corresponding to game characters or game scenes), voice navigation systems (for example, used in automated equipment such as cars and robots to convert text into voice prompts to users of their current location and operation steps), digital telephone services (for example, in voice calls, online chats, intelligent customer service, etc., quickly converting text into voice as part of digital customer service), intelligent virtual assistants, technical education and lectures (for example, quickly generating corresponding audio from text information in slides for lectures), and audio books/news/advertisements (for example, quickly generating audio content from text content).
- electronic games for example, editing the original audio corresponding to game characters or game scenes
- voice navigation systems for example, used in automated equipment such as cars and robots to convert text into voice prompts to users of their current location and operation steps
- digital telephone services for example, in voice calls, online chats, intelligent customer service, etc., quickly converting text into voice as part of digital customer service
- the above-mentioned original audio is the audio data to be processed in the above-mentioned application scenario.
- the above-mentioned target text is the text content to be edited to the original audio in the above-mentioned application scenario.
- the above-mentioned voice editing method of the embodiment of the present disclosure can be run on the client.
- the above-mentioned designation can be specified by the user through the client, or it can be specified by a computer application running on the client to generate a control instruction according to the scenario requirements.
- the above-mentioned voice editing method of the embodiment of the present disclosure can also be run on the server, which can be an independent server, a server cluster or a cloud server.
- SaaS Software as a Service
- Step S22 performing voice masking on the part of the audio to be edited in the original audio to obtain a first masked audio
- the speech editing model used in the above-mentioned speech editing process of the present disclosure introduces a mask training strategy, that is, the speech synthesis model that solves the speech synthesis or speech editing problem is converted into a speech editing model.
- the above-mentioned speech editing model including the mask training strategy has the same goals in the model training stage and the model testing stage (or model application stage), thereby being able to improve the speech editing effect of the speech editing model.
- the function of the voice mask is to locally modify the part of the audio to be edited in the original audio
- the voice editing method of the local modification may include at least one of the following: replacement method, deletion method and insertion method. That is, the part of the audio to be edited may be any audio segment selected by the user from the original audio, and the original audio content corresponding to the part of the audio to be edited is erased from the original audio by using the voice masking technical means, and the remaining context audio content in the original audio except the part of the audio to be edited is retained.
- Step S23 performing voice editing on the target text and the first masked audio to obtain the target audio.
- the above-mentioned voice editing of the target text and the first masked audio may be voice editing of the first masked audio according to the text content to be edited into the original audio determined by the target text.
- the target audio obtained by adopting the above scheme performs better than the voice editing results of the related art in terms of voice fluency and realism.
- the technical solution of the above method of the embodiment of the present disclosure is further introduced.
- the above method is run on the client corresponding to the intelligent virtual assistant.
- the client records the original voice to be voice edited by the user (that is, obtains the original audio), obtains the target text input by the user, and determines the text content to be edited to the original audio based on the target text.
- the content of the original voice to be voice edited recorded by the original audio is "voice editing based on deep learning”
- the text content determined by the target text is "machine learning”.
- the voice editing requirement in the above scenario is: replace “deep learning" in the original audio with "machine learning”.
- the part of the audio to be edited in the original audio is the audio corresponding to "deep learning”.
- the target text is used to determine the text content to be edited into the original audio; a first masked audio is obtained by voice masking the part of the audio to be edited in the original audio; and the target text and the first masked audio are further voice edited to obtain the target audio, thereby achieving the purpose of obtaining the target audio by first voice masking the original audio to be voice edited and then voice editing, thereby achieving the technical effect of improving the fluency and realism of the voice editing results, and further solving the technical problem that the voice editing method provided in the related art has a mismatch between training and testing, resulting in low fluency and poor realism of the voice editing results.
- the original audio may be the original game voice to be used in a game application (APP), and the first masked audio may be the masked game voice.
- voice masking may be performed on the part of the original game voice to be used in the game application to be edited to obtain the masked game voice
- voice editing may be performed on the game content text to be used in the game application and the masked game voice to obtain the target game voice.
- the voice "Warrior is coming" to be edited in the original game voice "Welcome our great warriors here! to be used in the game application is voice masked to obtain the masked game voice "Welcome our great [XXXX] here!
- the game content text to be used in the game application (the name of the warrior player character is coming) and the masked game voice "Welcome our great [XXXX] here! are voice edited to obtain the target game voice "Welcome our great [Warrior player character name is coming] here!.
- personalized audio can be played for different game players.
- the original audio may be original multimedia dubbing to be used in a dubbing application (e.g., dubbing of a film or TV series, dubbing of an animation, etc.), and the first masked audio may be masked multimedia dubbing.
- voice masking may be performed on the part of the dubbing to be edited in the original multimedia dubbing to be used in the dubbing application to obtain the masked multimedia dubbing
- voice editing may be performed on the dubbing content text and the masked multimedia dubbing to be used in the dubbing application to obtain the target multimedia dubbing.
- the dubbing actor In the existing dubbing scene, after the dubbing is completed according to the script of the film and television drama, the dubbing actor sometimes finds that some words are omitted in the original multimedia dubbing without dubbing. At this point, if the entire dubbing is re-recorded, it will not only consume manpower and time, but also be difficult to ensure that the dubbing re-recorded will not be omitted again. For this reason, the above-mentioned technical scheme proposed by the present disclosure can be applied, and in the above-mentioned original film and television drama dubbing, the omitted words in the script are inserted.
- the original film and television drama dubbing "Well, let's go together! to be used in the dubbing application is voice masked in the part of the dubbing "Let's go together” to be edited, and the multimedia dubbing "Well, [XXXX] go! after the masking is obtained, and the dubbing content text (let's go together tomorrow) to be used in the dubbing application and the multimedia dubbing "Well, [XXXX] go! after the masking are voice edited to obtain the target multimedia dubbing "Well, let's go together tomorrow!.
- the intelligent virtual assistant performs voice editing on the original voice of the user as an example
- the technical solution of the above method of the embodiment of the present disclosure is further introduced.
- the above method is run on the client corresponding to the intelligent virtual assistant.
- step S22 voice masking is performed on the part of the audio to be edited in the original audio to obtain the first masked audio, which may include the following execution steps:
- Step S221 obtaining the position information of the audio portion to be edited in the original audio
- Step S222 perform voice masking on the part of the audio to be edited in the original audio based on the position information to obtain a first masked audio.
- the client obtains the above-mentioned position information by comparing the original audio and the audio portion to be edited, and the position information may be determined by the position IDs of multiple phonemes of the audio portion to be edited in the original speech corresponding to the original audio.
- the mask part in the speech editing model is used to perform speech masking on the audio portion to be edited in the original audio based on the position information to obtain the first masked audio.
- the purpose of voice masking the part of the original audio to be edited is to extract the part of the audio to be edited from the original audio. For example, if the audio corresponding to "deep learning” is extracted from the original audio corresponding to "voice editing based on deep learning", the content corresponding to the first masked audio is "voice editing based on [XXXX]".
- FIG. 3 is a schematic diagram of an optional voice editing process according to one embodiment of the present disclosure. As shown in FIG. 3 , After obtaining the original audio, the mask part is used to perform mask processing on the original audio to obtain the first masked audio.
- the position information of "deep learning” in “voice editing based on deep learning” is determined, and the position information can be the grapheme ID of the content corresponding to the part of the audio to be edited. Then, based on the scenario requirement of replacing "deep learning” in the original audio with “machine learning” and the above position information, the audio "deep learning” in the audio "voice editing based on deep learning” is voice masked to obtain the audio "voice editing based on [XXXX]".
- step S23 voice editing is performed on the target text and the first masked audio to obtain the target audio, which may include the following execution steps:
- Step S231 performing speech conversion on the target text to obtain intermediate audio
- Step S232 perform voice splicing on the intermediate audio and the first masked audio to obtain the target audio.
- the target text can be voice-converted to obtain the intermediate audio.
- the text content determined by the target text is "machine learning”
- the intermediate audio obtained by voice-converting the target text is the audio segment corresponding to "machine learning”.
- the first masked audio "voice editing based on XXXX” and the intermediate audio, namely "machine learning” are voice-concatenated to obtain the target audio "voice editing based on machine learning”.
- step S231 performing speech conversion on the target text to obtain the target audio may include the following execution steps:
- Step S2311 performing voice editing on the target text and the first masked audio to obtain a target acoustic feature, wherein the target acoustic feature is used to determine the audio segment corresponding to the target text;
- Step S2312 Perform voice coding conversion on the target acoustic features to obtain intermediate audio.
- voice editing is performed on “machine learning” and the audio "voice editing based on [XXXX]” to obtain the target acoustic features corresponding to "machine learning”; the target acoustic features corresponding to “machine learning” are converted into voice code to obtain the audio "machine learning” (i.e., intermediate audio); the audio "machine learning” and the audio “voice editing based on [XXXX]” are spliced to obtain the target audio, and the target audio is "voice editing based on machine learning”.
- the purpose of the above audio splicing is: based on the corresponding position of "deep learning” in the original audio, the content of the target text “machine learning” is spliced with the content corresponding to the audio after the first mask "voice editing based on [XXXX]” to obtain the content corresponding to the target audio "voice editing based on machine learning”.
- speech editing is performed based on the target text and the first masked audio to obtain the target acoustic features, and then the target acoustic features are vocoded using a neural vocoder to obtain an intermediate audio, and then the intermediate audio is compared with the first masked audio.
- the encoded audio is spliced to obtain the target audio.
- the target acoustic features corresponding to the target text are generated using the speech editing method, thereby obtaining the intermediate audio. Since the above first masked audio is obtained based on the position information of the part of the audio to be edited in the original audio, the target audio obtained by splicing the intermediate audio with the first masked audio can have better performance in speech fluency and realism.
- step S2311 voice editing is performed on the target text and the first masked audio to obtain the target acoustic feature, which may include the following execution steps:
- Step S23111 converting the target text into graphemes to obtain a phoneme sequence
- Step S23112 use the target speech editing model to perform speech editing on the phoneme sequence and the first masked audio to obtain target acoustic features, wherein the target speech editing model is obtained through deep learning training using multiple sets of data, and the multiple sets of data include: training audio and training text, and the training text is the text corresponding to the audio part to be edited in the training audio.
- the target text is converted from graphemes to phonemes using a conversion part to obtain a phoneme sequence corresponding to the target text.
- the above-mentioned conversion part is a graph to tree (G2T) conversion part, which is used to represent the target text (which can be a single sentence or multiple segmented sentences) as a grammar tree through natural language processing technology.
- the above-mentioned phoneme sequence can correspond to multiple nodes of the grammar tree.
- the above-mentioned target speech editing model is a sequence-to-sequence speech editing part, which uses the sequence-to-sequence speech editing part to perform speech editing on the phoneme sequence and the first masked audio to obtain the target acoustic features.
- the above-mentioned multiple sets of data used to train the target voice editing model can be historical voice editing results, that is, data obtained by voice editing the specified training audio using the voice editing model in the application scenario.
- the training audio included in each set of data in the multiple sets of data corresponds to the training text.
- the target text "machine learning” is converted from grapheme to phoneme to obtain the phoneme sequence corresponding to "machine learning”; the target voice editing model is used to perform voice editing on the phoneme sequence corresponding to "machine learning” and the audio "voice editing based on [XXXX]” to obtain the target acoustic features.
- the above voice editing method may further include the following execution steps:
- Step S241 performing voice masking on the part of the audio to be edited in the training audio to obtain a second masked audio
- Step S242 using the initial speech editing model to perform speech editing on the second masked audio and the training text to obtain predicted acoustic features
- Step S243 determining the target loss by predicting the acoustic features and the real acoustic features corresponding to the training text
- Step S244 using the target loss to update the parameters of the initial speech editing model to obtain a target speech editing model.
- the target loss of the model training is calculated based on the training audio and the training text, and the model parameters of the initial speech editing model are optimized and updated using the target loss to obtain the target speech editing model.
- the loss function corresponding to the above target loss can be any commonly used loss function, and the embodiments of the present disclosure do not limit the calculation method of the target loss.
- the target speech editing model includes: an encoder, a feature regulator and a decoder.
- the target speech editing model is used to perform speech editing on the phoneme sequence and the first masked audio to obtain the target acoustic feature, which may include the following execution steps:
- Step S23113 using an encoder to perform text feature space encoding on the phoneme sequence to obtain text features
- Step S23114 using a feature adjuster to perform feature adjustment on the text feature and the first masked audio to obtain a first auditory perception feature, wherein the first auditory perception feature is an auditory perception feature corresponding to the target text;
- Step S23115 Use a decoder to acoustically decode the first auditory perception feature to obtain a target acoustic feature.
- Figure 4 is a schematic diagram of an optional acoustic feature extraction process according to one embodiment of the present disclosure.
- the sequence-to-sequence speech editing part in the target speech editing model includes: an encoder, a feature adjuster and a decoder, wherein the decoder is an acoustic decoder.
- the above-mentioned encoder is used to perform text feature space encoding on the factor sequence corresponding to the target text to obtain text features, and the text features are simultaneously passed to the feature adjuster and the acoustic decoder.
- the above-mentioned feature adjuster is used to perform feature adjustment on the text features and the first masked audio to obtain the first auditory perceptual feature, wherein the first auditory perceptual feature is the auditory perceptual feature corresponding to the target text.
- the above-mentioned acoustic decoder is used to acoustically decode the text features and the first auditory perceptual feature to obtain the above-mentioned target acoustic features.
- the encoder maps the phoneme sequence to a high-dimensional text feature space for encoding through a nonlinear transformation to obtain the text feature.
- the feature adjuster performs feature prediction and feature adjustment based on the text feature and the first masked audio through a nonlinear transformation prediction method to obtain the first auditory perception feature.
- the acoustic decoder performs acoustic feature prediction on the text feature and the first auditory perception feature through a nonlinear transformation prediction method to obtain the target acoustic feature.
- the encoder is used to encode the phoneme sequence corresponding to "machine learning” in the text feature space to obtain the text features corresponding to "machine learning”;
- the feature adjuster is used to perform feature adjustment on the text features corresponding to "machine learning” and the audio "voice editing based on [XXXX]” to obtain the auditory perception features corresponding to "machine learning” in the target text;
- the decoder is used to perform feature adjustment on the text features corresponding to "machine learning” and the audio “voice editing based on [XXXX]” to obtain the auditory perception features corresponding to "machine learning” in the target text;
- the auditory perception features corresponding to "Xi” are acoustically decoded to obtain the target acoustic features.
- step S23114 using a feature adjuster to perform feature adjustment on the text feature and the first masked audio to obtain a first auditory perception feature may include the following execution steps:
- Step S23116 extracting a second auditory perception feature from the first masked audio, wherein the second auditory perception feature is an auditory perception feature corresponding to the context audio associated with the portion of audio to be edited in the original audio;
- Step S23117 Use a feature adjuster to perform feature adjustment on the text feature and the second auditory perception feature to obtain the first auditory perception feature.
- the above-mentioned feature adjuster can also obtain the first auditory perception feature through the following steps: extract features from the first masked audio to obtain the second auditory perception feature. Use the feature adjuster to perform feature adjustment on the text feature and the second auditory perception feature to obtain the first auditory perception feature.
- the above-mentioned second auditory perception feature is the auditory perception feature corresponding to "voice editing based on [XXXX]", and the auditory perception feature at least includes: pitch feature, energy feature and duration feature.
- the first auditory perception feature includes at least one of the following: a pitch corresponding to the target text; an energy corresponding to the target text; a duration corresponding to the target text.
- the first auditory perception feature includes at least one of the following: the pitch feature, energy feature, and duration feature corresponding to the text content "machine learning" in the target text.
- the voice editing methods corresponding to the scenario requirements of the voice editing scenario may include, in addition to the above-mentioned replacement method, insertion method and deletion method.
- the voice editing method corresponding to the scene requirement is to insert "technology" between “learning” and “of”
- the part of the audio to be edited in the original audio is the context content of the insertion position, that is, "learning”
- the text content corresponding to the target text is "learning technology”.
- the target text and the first masked audio are voice edited to obtain the target audio, and the content corresponding to this target audio is "voice editing based on deep learning technology”.
- the voice editing method corresponding to the scenario requirement is to delete "voice” after "deep learning”
- the part of the audio to be edited in the above original audio is the context content of the deletion position, that is, "voice editing”
- the target text is "Editor”.
- a processing method similar to the above replacement method is used to perform voice editing on the target text and the first masked audio to obtain the target audio, and the content corresponding to this target audio is "Editor based on deep learning technology".
- the speech editing method provided by the embodiment of the present disclosure introduces audio mask processing, which can unify the training goal and the test goal in the speech editing process, so that the target audio (especially at the audio splicing position) obtained after editing the audio to be edited in the original audio has higher fluency and realism.
- the trained target speech editing model can predict the partial audio corresponding to the target text (i.e., the intermediate audio, i.e., the partial audio to be spliced to the first masked audio) based on the first masked audio. Therefore, the training phase of the above-mentioned target speech editing model is consistent with the objectives of the testing phase (or scenario application phase), avoiding the problem of poor speech editing effect caused by the mismatch between the training objectives and the testing objectives of the model in the related art. In addition, the above-mentioned training process can also ensure smoother audio splicing.
- FIG5 is a flow chart of an optional voice editing method according to one embodiment of the present disclosure. As shown in FIG5 , the voice editing method includes:
- Step S51 receiving the original audio and target text to be processed from the client, wherein the target text is used to determine the text content to be edited into the original audio;
- Step S52 performing voice masking on the part of the audio to be edited in the original audio to obtain a first masked audio, and performing voice editing on the target text and the first masked audio to obtain a target audio;
- Step S53 Feedback the target audio to the client.
- FIG6 is a schematic diagram of an optional voice editing on a cloud server according to one embodiment of the present disclosure.
- the client uploads the original audio and the target text to the cloud server, wherein the target text is used to determine the text content to be edited into the original audio; the cloud server performs voice masking on the part of the audio to be edited in the original audio to obtain the first masked audio, and performs voice editing on the target text and the first masked audio to obtain the target audio.
- the cloud server will feedback the target audio to the above-mentioned client, and the final target audio will be provided to the user through the graphical user interface of the client.
- the above-mentioned voice editing method provided in the embodiment of the present disclosure can be applied to, but not limited to, actual application scenarios such as voice navigation systems, digital telephone services, intelligent virtual assistants, technical education/lectures, and audio books/news/advertisements, through the interaction between the SaaS server and the client, using the client to provide the server with the original audio and The server performs voice masking on the part of the audio to be edited in the original audio to obtain the first masked audio, and performs voice editing on the target text and the first masked audio to obtain the target audio. The server returns the target audio to the client and provides it to the user.
- actual application scenarios such as voice navigation systems, digital telephone services, intelligent virtual assistants, technical education/lectures, and audio books/news/advertisements
- One embodiment of the present disclosure further provides another voice editing method, which provides a graphical user interface through a terminal device, and the content displayed by the graphical user interface includes a voice editing control.
- FIG. 7 is a flow chart of another voice editing method according to one embodiment of the present disclosure. As shown in FIG. 7, the voice editing method includes:
- Step S71 in response to a trigger operation performed on a voice editing control, a voice editing interface pops up;
- Step S72 in response to an input operation performed on the voice editing interface, importing the original audio and the target text, wherein the target text is used to determine the text content to be edited into the original audio;
- Step S73 in response to the editing operation performed on the voice editing interface, selecting a portion of the audio to be edited from the original audio;
- Step S74 in response to the play operation performed on the voice editing interface, the target audio is played in the game scene, wherein the target audio is obtained by voice editing the target text and the masked audio, and the masked audio is obtained by voice masking the audio to be edited.
- At least a voice editing control is displayed in the above-mentioned graphical user interface.
- the user triggers the voice editing control to pop up a voice editing interface in the above-mentioned graphical user interface.
- the user imports the original audio and target text to be voice-edited through the voice editing interface, wherein the target text is used to determine the text content to be edited to the original audio; and selects the part of the audio to be edited from the original audio, and then voice-edits the target text and the masked audio, and plays the target audio in the game scene.
- the trigger operation, input operation, editing operation and play operation can all be touch operations, and the touch operation can include single-point touch and multi-point touch, wherein the touch operation of each touch point can include click, long press, heavy press, swipe, etc.
- the trigger operation, input operation, editing operation and play operation can also be an operation implemented by an input device such as a mouse and a keyboard.
- the above-mentioned input operation corresponds to the first control or the first touch area in the graphical user interface (such as a typing box, a handwriting input area, an input button (for example, pressing and holding the button records and receives the user's voice));
- the above-mentioned editing operation corresponds to the second control or the second touch area in the graphical user interface (such as an editing box, an editing option bar, etc.);
- the above-mentioned playback operation corresponds to the third control or the third touch area in the graphical user interface (such as a playback area, a playback button, etc.).
- the above-mentioned voice editing tool can be inserted into the game client, and a voice editing control can be provided in the graphical user interface.
- the game player can perform a touch operation or a mouse click operation on the voice editing control to pop up the voice editing interface.
- the voice editing interface the game player can use the import control to import the original game voice and the game content text to be used, and select the part of the audio to be edited from the original game voice. For example: The original game voice used is "Welcome our great warriors here!, and the user can select the audio segment corresponding to "Warrior is here", or select the first word "Brave" and the last word "Arrival".
- the game server can perform voice masking on the original game voice "Welcome our great warriors here! to be used in the game application in real time, and obtain the masked game voice "Welcome our great [XXXX] here!. Then, the game content text to be used in the game application (Warrior player character name is here) and the masked game voice "Welcome our great [XXXX] here! are voice edited to obtain the target game voice "Welcome our great [Warrior player character name is here] here!.
- personalized audio can be played for different game players.
- the voice editing scene provided by the embodiment of the present disclosure, it is possible to interact with the user in a visual form, and generate the corresponding target audio according to the user's input operation, editing operation and playback operation, which is conducive to the application in actual scenes.
- the voice editing scenario provided according to the embodiment of the present disclosure can interact with the user in a visual form, and generate corresponding target audio according to the user's input operation, editing operation and playback operation, which is conducive to application in actual scenarios.
- FIG8 is a flow chart of a model training method according to one embodiment of the present disclosure. As shown in FIG8 , the model training method includes:
- Step S81 obtaining the training audio and training text to be processed, wherein the training text is used to determine the text content to be edited into the training audio;
- Step S82 performing voice masking on the part of the training audio to be edited to obtain the masked training audio
- Step S83 use the masked training audio and training text to train the initial speech editing model to obtain a target speech editing model, wherein the target speech editing model is used to perform speech editing on the target text and the masked original audio to obtain the target audio, and the masked original audio is obtained by speech masking the part of the audio to be edited in the original audio.
- the input requirements and output targets corresponding to the model are consistent with the model testing or model application process. That is, in the application scenario, the target speech editing model is used to perform speech editing based on the original audio and the target text to obtain the masked audio, and then the target audio is obtained. Correspondingly, in the training process, speech masking is performed based on the training audio and the training text to obtain the masked audio, and then the target training audio is obtained, so that the parameters of the initial speech editing model are optimized to obtain the target speech editing model.
- the audio masking mechanism in the process of training the speech editing model, the original audio and the first masked audio corresponding to the original audio and the target audio are used as training samples for model training, thereby enabling the trained target speech editing model to be able to be based on the first masked audio and the target audio.
- the audio prediction obtains the partial audio corresponding to the target text (i.e., the intermediate audio, i.e., the partial audio to be spliced to the audio after the first mask).
- the training phase and the testing phase (or the scenario application phase) of the target speech editing model are consistent in objectives, thus avoiding the problem of poor speech editing effect caused by the mismatch between the training objectives and the testing objectives of the model in the related art.
- the training process can also ensure smoother audio splicing.
- the method according to the above embodiment can be implemented by means of software plus a necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method.
- the technical solution of the present disclosure, or the part that contributes to the relevant technology can be embodied in the form of a software product, which is stored in a storage medium (such as a disk, an optical disk), and includes a number of instructions for a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present disclosure.
- a voice editing device is also provided, which is used to implement the above-mentioned embodiments and preferred implementation modes, and the descriptions that have been made will not be repeated.
- the term "module” can implement a combination of software and/or hardware of a predetermined function.
- the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceived.
- Figure 9 is a structural block diagram of a speech editing device according to one embodiment of the present disclosure.
- the device includes: an acquisition module 901, used to acquire the original audio and target text to be processed, wherein the target text is used to determine the text content to be edited into the original audio; a masking module 902, used to perform speech masking on the part of the audio to be edited in the original audio to obtain the first masked audio; an editing module 903, used to perform speech editing on the target text and the first masked audio to obtain the target audio.
- the masking module 902 is further used to: obtain position information of the audio portion to be edited in the original audio; and perform voice masking on the audio portion to be edited in the original audio based on the position information to obtain a first masked audio.
- the editing module 903 is further used to: perform speech conversion on the target text to obtain intermediate audio; and perform speech splicing on the intermediate audio and the first masked audio to obtain the target audio.
- the editing module 903 is further used to: perform voice editing on the target text and the first masked audio to obtain target acoustic features, wherein the target acoustic features are used to determine the audio segment corresponding to the target text; and perform voice coding conversion on the target acoustic features to obtain intermediate audio.
- the above-mentioned editing module 903 is also used to: convert the target text into graphemes to obtain a phoneme sequence; use a target speech editing model to perform speech editing on the phoneme sequence and the first masked audio to obtain target acoustic features, wherein the target speech editing model is obtained through deep learning training using multiple sets of data, and the multiple sets of data include: training audio and training text, and the training text is the text corresponding to the audio portion to be edited in the training audio.
- Figure 10 is a structural block diagram of an optional speech editing device according to one embodiment of the present disclosure.
- the device in addition to all the modules shown in Figure 9, the device also includes: an updating module 904, which is used to speech mask the part of the audio to be edited in the training audio to obtain a second masked audio; use the initial speech editing model to speech edit the second masked audio and the training text to obtain predicted acoustic features; determine the target loss by comparing the predicted acoustic features with the real acoustic features corresponding to the training text; and use the target loss to update the parameters of the initial speech editing model to obtain a target speech editing model.
- an updating module 904 is used to speech mask the part of the audio to be edited in the training audio to obtain a second masked audio
- use the initial speech editing model to speech edit the second masked audio and the training text to obtain predicted acoustic features
- determine the target loss by comparing the predicted acoustic features with the real acoustic features corresponding to the training text
- the editing module 903 is also used to: use an encoder to perform text feature space encoding on a phoneme sequence to obtain text features; use a feature adjuster to perform feature adjustment on the text features and the first masked audio to obtain a first auditory perception feature, wherein the first auditory perception feature is the auditory perception feature corresponding to the target text; use a decoder to acoustically decode the first auditory perception feature to obtain a target acoustic feature.
- the editing module 903 is further used to: extract a second auditory perception feature from the first masked audio, wherein the second auditory perception feature is an auditory perception feature corresponding to the context audio associated with the part of the audio to be edited in the original audio; and use a feature adjuster to perform feature adjustment on the text feature and the second auditory perception feature to obtain the first auditory perception feature.
- the first auditory perception feature includes at least one of the following: a pitch corresponding to the target text; an energy corresponding to the target text; a duration corresponding to the target text.
- the masking module 902 is further used to: perform voice masking on the part of the original game voice to be used in the game application to be edited to obtain the masked game voice; perform voice editing on the target text and the first masked audio to obtain the target audio, including: perform voice editing on the game content text to be used in the game application and the masked game voice to obtain the target game voice.
- the masking module 902 is further used to: perform voice masking on the part of the dubbing to be edited in the original multimedia dubbing to be used in the dubbing application to obtain the masked multimedia dubbing; perform voice editing on the target text and the first masked audio to obtain the target audio, including: perform voice editing on the dubbing content text to be used in the dubbing application and the masked multimedia dubbing to obtain the target multimedia dubbing.
- the above modules can be implemented by software or hardware. For the latter, it can be implemented in the following ways, but not limited to: the above modules are all located in the same processor; or the above modules are located in different processors in any combination.
- An embodiment of the present disclosure further provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above method embodiments when running.
- the computer-readable storage medium may include but is not limited to: a USB flash drive, a read-only memory
- Various media that can store computer programs include read-only memory (ROM), random access memory (RAM), mobile hard disk, magnetic disk or optical disk, etc.
- the computer-readable storage medium may be located in any one of the computer terminals in a computer terminal group in a computer network, or in any one of the mobile terminals in a mobile terminal group.
- the computer-readable storage medium may be configured to store a computer program for performing the following steps:
- the above-mentioned computer-readable storage medium is also configured to store a computer program for executing the following steps: obtaining position information of the audio portion to be edited in the original audio; performing voice masking on the audio portion to be edited in the original audio based on the position information to obtain a first masked audio.
- the computer-readable storage medium is further configured to store a computer program for executing the following steps: performing speech conversion on the target text to obtain intermediate audio; performing speech splicing on the intermediate audio and the first masked audio to obtain the target audio.
- the computer-readable storage medium is also configured to store a computer program for executing the following steps: performing voice editing on the target text and the first masked audio to obtain target acoustic features, wherein the target acoustic features are used to determine the audio segment corresponding to the target text; performing voice coding conversion on the target acoustic features to obtain intermediate audio.
- the computer-readable storage medium is also configured to store a computer program for executing the following steps: performing grapheme-to-phoneme conversion on a target text to obtain a phoneme sequence; performing speech editing on the phoneme sequence and the first masked audio using a target speech editing model to obtain target acoustic features, wherein the target speech editing model is obtained through deep learning training using multiple sets of data, and the multiple sets of data include: training audio and training text, and the training text is the text corresponding to the portion of the audio to be edited in the training audio.
- the computer-readable storage medium is also configured to store a computer program for executing the following steps: performing speech masking on a portion of the audio to be edited in the training audio to obtain a second masked audio; performing speech editing on the second masked audio and the training text using an initial speech editing model to obtain predicted acoustic features; determining a target loss by comparing the predicted acoustic features with actual acoustic features corresponding to the training text; and using the target loss to update the parameters of the initial speech editing model to obtain a target speech editing model.
- the computer readable storage medium is further configured to store a computer program for performing the following steps: An encoder is used to perform text feature space encoding on a phoneme sequence to obtain text features; a feature adjuster is used to perform feature adjustment on the text features and the first masked audio to obtain a first auditory perception feature, wherein the first auditory perception feature is an auditory perception feature corresponding to a target text; a decoder is used to acoustically decode the first auditory perception feature to obtain a target acoustic feature.
- the computer-readable storage medium is also configured to store a computer program for executing the following steps: extracting a second auditory perception feature from the first masked audio, wherein the second auditory perception feature is an auditory perception feature corresponding to context audio associated with the portion of audio to be edited in the original audio; and using a feature adjuster to perform feature adjustment on the text feature and the second auditory perception feature to obtain the first auditory perception feature.
- the above-mentioned computer-readable storage medium is also configured to store a computer program for performing the following steps: the first auditory perception feature includes at least one of the following: the pitch corresponding to the target text; the energy corresponding to the target text; the duration corresponding to the target text.
- the above-mentioned computer-readable storage medium is also configured to store a computer program for executing the following steps: voice masking the part of the voice to be edited in the original game voice to be used in the game application to obtain the masked game voice; voice editing the target text and the first masked audio to obtain the target audio includes: voice editing the game content text to be used in the game application and the masked game voice to obtain the target game voice.
- the above-mentioned computer-readable storage medium is also configured to store a computer program for executing the following steps: voice masking the part of the dubbing to be edited in the original multimedia dubbing to be used in the dubbing application to obtain the masked multimedia dubbing; voice editing the target text and the first masked audio to obtain the target audio includes: voice editing the dubbing content text to be used in the dubbing application and the masked multimedia dubbing to obtain the target multimedia dubbing.
- the computer-readable storage medium is also configured to store a computer program for performing the following steps: receiving original audio and target text to be processed from a client, wherein the target text is used to determine the text content to be edited into the original audio; performing voice masking on the portion of the audio to be edited in the original audio to obtain a first masked audio, and performing voice editing on the target text and the first masked audio to obtain a target audio; and feeding back the target audio to the client.
- the computer-readable storage medium is further configured to store a computer program for executing the following steps: in response to a trigger operation performed on a voice editing control, a voice editing interface pops up; in response to an input operation performed on the voice editing interface, original audio and target text are imported, wherein the target text is used to determine the text content to be edited into the original audio; in response to an editing operation performed on the voice editing interface, a portion of audio to be edited is selected from the original audio; in response to a play operation performed on the voice editing interface, the target audio is played in a game scene, wherein the target audio is obtained by voice editing the target text and the masked audio, and the masked audio is obtained by voice masking the portion of audio to be edited. get.
- the computer-readable storage medium is also configured to store a computer program for performing the following steps: obtaining training audio and training text to be processed, wherein the training text is used to determine the text content to be edited into the training audio; performing speech masking on the portion of the audio to be edited in the training audio to obtain the masked training audio; using the masked training audio and training text to train an initial speech editing model to obtain a target speech editing model, wherein the target speech editing model is used to perform speech editing on the target text and the masked original audio to obtain the target audio, and the masked original audio is obtained by performing speech masking on the portion of the audio to be edited in the original audio.
- a technical solution for implementing a voice editing method is provided.
- the target text is used to determine the text content to be edited into the original audio;
- the first masked audio is obtained by voice masking the part of the audio to be edited in the original audio;
- the target text and the first masked audio are further voice edited to obtain the target audio, thereby achieving the purpose of obtaining the target audio by voice masking the original audio to be voice edited and then voice editing, thereby achieving the technical effect of improving the fluency and realism of the voice editing result, and further solving the technical problem that the voice editing method provided in the related art has a mismatch between training and testing, resulting in low fluency and poor realism of the voice editing result.
- the technical solution according to the implementation methods of the present disclosure can be embodied in the form of a software product, which can be stored in a computer-readable storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the implementation methods of the present disclosure.
- a computer-readable storage medium which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.
- a computing device which can be a personal computer, a server, a terminal device, or a network device, etc.
- a program product capable of implementing the above method of the present embodiment is stored on a computer-readable storage medium.
- various aspects of the embodiments of the present disclosure may also be implemented in the form of a program product, which includes a program code, and when the program product is run on a terminal device, the program code is used to enable the terminal device to execute the steps according to various exemplary implementations of the present disclosure described in the above “Exemplary Method” section of the present embodiment.
- the program product for implementing the above method in the embodiment of the present disclosure, it can adopt a portable compact disk read-only memory (CD-ROM) and include program code, and can be run on a terminal device, such as a personal computer.
- a terminal device such as a personal computer.
- the program product of the embodiment of the present disclosure is not limited to this.
- the computer-readable storage medium can be any tangible medium containing or storing a program, which can be used by or in combination with an instruction execution system, an apparatus or a device.
- the program product may be in any combination of one or more computer-readable media.
- the computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination thereof. More specific examples (non-exhaustive) of computer-readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.
- program code contained in the computer-readable storage medium can be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the above.
- An embodiment of the present disclosure further provides an electronic device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
- the electronic device may further include a transmission device and an input/output device, wherein the transmission device is connected to the processor, and the input/output device is connected to the processor.
- the processor may be configured to perform the following steps through a computer program:
- the processor may also be configured to execute the following steps through a computer program: obtaining position information of the audio portion to be edited in the original audio; and performing voice masking on the audio portion to be edited in the original audio based on the position information to obtain a first masked audio.
- the processor may also be configured to perform the following steps through a computer program: performing speech conversion on the target text to obtain intermediate audio; and performing speech splicing on the intermediate audio and the first masked audio to obtain the target audio.
- the above-mentioned processor can also be configured to perform the following steps through a computer program: voice-editing the target text and the first masked audio to obtain target acoustic features, wherein the target acoustic features are used to determine the audio segment corresponding to the target text; and voice-coding the target acoustic features to obtain intermediate audio.
- the processor can also be configured to perform the following steps through a computer program: converting the target text into graphemes to obtain a phoneme sequence; performing speech editing on the phoneme sequence and the first masked audio using a target speech editing model to obtain target acoustic features, wherein the target speech editing model is obtained through deep learning training using multiple sets of data, and the multiple sets of data include: training audio and training text, and the training text is the text corresponding to the audio portion to be edited in the training audio.
- the processor can also be configured to perform the following steps through a computer program: speech masking the portion of the audio to be edited in the training audio to obtain a second masked audio; speech editing the second masked audio and the training text using an initial speech editing model to obtain predicted acoustic features; determining a target loss by comparing the predicted acoustic features with the actual acoustic features corresponding to the training text; and using the target loss to update the parameters of the initial speech editing model to obtain a target speech editing model.
- the above-mentioned processor can also be configured to perform the following steps through a computer program: use an encoder to perform text feature space encoding on the phoneme sequence to obtain text features; use a feature adjuster to perform feature adjustment on the text features and the first masked audio to obtain a first auditory perception feature, wherein the first auditory perception feature is the auditory perception feature corresponding to the target text; use a decoder to acoustically decode the first auditory perception feature to obtain a target acoustic feature.
- the processor can also be configured to perform the following steps through a computer program: extracting a second auditory perception feature from the first masked audio, wherein the second auditory perception feature is an auditory perception feature corresponding to context audio associated with the portion of audio to be edited in the original audio; and using a feature adjuster to perform feature adjustment on the text feature and the second auditory perception feature to obtain the first auditory perception feature.
- the processor may also be configured to execute the following steps through a computer program: the first auditory perception feature includes at least one of the following: the pitch corresponding to the target text; the energy corresponding to the target text; the duration corresponding to the target text.
- the above-mentioned processor can also be configured to perform the following steps through a computer program: voice masking the part of the original game voice to be used in the game application to be edited to obtain the masked game voice; voice editing the target text and the first masked audio to obtain the target audio, including: voice editing the game content text to be used in the game application and the masked game voice to obtain the target game voice.
- the processor can also be configured to perform the following steps through a computer program: voice masking the part of the dubbing to be edited in the original multimedia dubbing to be used in the dubbing application to obtain the masked multimedia dubbing; voice editing the target text and the first masked audio to obtain the target audio, including: voice editing the dubbing content text to be used in the dubbing application and the masked multimedia dubbing to obtain the target multimedia dubbing.
- the above-mentioned processor can also be configured to perform the following steps through a computer program: receiving original audio and target text to be processed from the client, wherein the target text is used to determine the text content to be edited into the original audio; performing voice masking on the audio to be edited in the original audio to obtain a first masked audio, and performing voice editing on the target text and the first masked audio to obtain a target audio; and feeding back the target audio to the client.
- the processor may be configured to perform the following steps through a computer program: in response to a trigger operation performed on a voice editing control, a voice editing interface pops up; in response to an input operation performed on the voice editing interface, the original audio and the target text are imported, wherein the target text is used to determine the text content to be edited into the original audio; in response to an input operation performed on the voice editing interface, the target text is used to determine the text content to be edited into the original audio;
- the editing operation performed on the editing interface selects a portion of the audio to be edited from the original audio; in response to the play operation performed on the voice editing interface, the target audio is played in the game scene, wherein the target audio is obtained by voice editing the target text and the masked audio, and the masked audio is obtained by voice masking the portion of the audio to be edited.
- the processor can also be configured to perform the following steps through a computer program: obtaining training audio and training text to be processed, wherein the training text is used to determine the text content to be edited into the training audio; performing voice masking on the portion of the audio to be edited in the training audio to obtain the masked training audio; using the masked training audio and training text to train an initial voice editing model to obtain a target voice editing model, wherein the target voice editing model is used to perform voice editing on the target text and the masked original audio to obtain the target audio, and the masked original audio is obtained by performing voice masking on the portion of the audio to be edited in the original audio.
- a technical solution for implementing a voice editing method is provided.
- the target text is used to determine the text content to be edited into the original audio;
- the first masked audio is obtained by voice masking the part of the audio to be edited in the original audio;
- the target text and the first masked audio are further voice edited to obtain the target audio, thereby achieving the purpose of obtaining the target audio by voice masking the original audio to be voice edited and then voice editing, thereby achieving the technical effect of improving the fluency and realism of the voice editing result, and further solving the technical problem that the voice editing method provided in the related art has a mismatch between training and testing, resulting in low fluency and poor realism of the voice editing result.
- Fig. 11 is a schematic diagram of an electronic device according to one embodiment of the present disclosure. As shown in Fig. 11, the electronic device 1100 is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
- the electronic device 1100 is presented in the form of a general-purpose computing device.
- the components of the electronic device 1100 may include, but are not limited to: the at least one processor 1110, the at least one memory 1120, a bus 1130 connecting different system components (including the memory 1120 and the processor 1110), and a display 1140.
- the memory 1120 stores program codes, which can be executed by the processor 1110, so that the processor 1110 executes the steps described in the method part of the embodiment of the present disclosure according to various exemplary embodiments of the present disclosure.
- the memory 1120 may include a readable medium in the form of a volatile storage unit, such as a random access memory unit (RAM) 11201 and/or a cache memory unit 11202, and may further include a read-only memory unit (ROM) 11203, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory.
- RAM random access memory unit
- ROM read-only memory unit
- non-volatile memory such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory.
- the memory 1120 may also include a program/utility 11204 having a set (at least one) of program modules 11205, such program modules 11205 including but not limited to: an operating system, one or more applications, Programs, other program modules and program data, each of these examples or some combination may include the implementation of a network environment.
- the memory 1120 may further include a memory remotely arranged relative to the processor 1110, and these remote memories may be connected to the electronic device 1100 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network and a combination thereof.
- the bus 1130 may represent one or more of several types of bus structures, including a memory unit bus or memory unit controller, a peripheral bus, an accelerated graphics port, a local bus of the processor 1110, or a bus using any of a variety of bus architectures.
- the display 1140 may be, for example, a touch screen liquid crystal display (LCD), which may enable a user to interact with a user interface of the electronic device 1100.
- LCD liquid crystal display
- the electronic device 1100 may also communicate with one or more external devices 1200 (e.g., keyboards, pointing devices, Bluetooth devices, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device 1100, and/or communicate with any device that enables the electronic device 1100 to communicate with one or more other computing devices (e.g., routers, modems, etc.). Such communication may be performed through an input/output (I/O) interface 1150.
- the electronic device 1100 may also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and/or a public network, such as the Internet) through a network adapter 1160. As shown in FIG.
- the network adapter 1160 communicates with other modules of the electronic device 1100 through a bus 1130.
- other hardware and/or software modules may be used in conjunction with the electronic device 1100, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, disk arrays (Redundant Arrays of Independent Disks, RAID) systems, tape drives, and data backup storage systems.
- the electronic device 1100 may further include: a keyboard, a cursor control device (such as a mouse), an input/output interface (I/O interface), a network interface, a power supply and/or a camera.
- FIG. 11 is for illustration only and does not limit the structure of the electronic device described above.
- the electronic device 1100 may also include more or fewer components than those shown in FIG. 11 , or have a configuration different from that shown in FIG. 11 .
- the memory 1120 may be used to store computer programs and corresponding data, such as the computer programs and corresponding data corresponding to the voice editing method in the embodiment of the present disclosure.
- the processor 1110 executes various functional applications and data processing by running the computer program stored in the memory 1120, that is, implements the voice editing method described above.
- the disclosed technical content can be implemented in other ways.
- the device embodiments described above are only schematic.
- the division of the units can be a logical function division. There may be other division methods in actual implementation.
- multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.
- Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
- the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
- each functional unit in each embodiment of the present disclosure may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
- the above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
- the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium.
- the technical solution of the present disclosure is essentially or the part that contributes to the relevant technology or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, server or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present disclosure.
- the aforementioned storage medium includes: U disk, read-only memory (ROM), random access memory (RAM), mobile hard disk, disk or optical disk and other media that can store program codes.
Landscapes
- Engineering & Computer Science (AREA)
- Computational Linguistics (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Human Computer Interaction (AREA)
- Physics & Mathematics (AREA)
- Acoustics & Sound (AREA)
- Multimedia (AREA)
- Signal Processing (AREA)
- Artificial Intelligence (AREA)
- Evolutionary Computation (AREA)
- Electrically Operated Instructional Devices (AREA)
Abstract
本公开公开了一种语音编辑方法、装置、存储介质及电子装置。该方法包括:获取待处理的原始音频和目标文本,其中,目标文本用于确定待编辑至原始音频的文本内容;对原始音频中的待编辑部分音频进行语音掩码,得到第一掩码后音频;对目标文本和第一掩码后音频进行语音编辑,得到目标音频。本公开解决了相关技术中提供的语音编辑方法其训练和测试不匹配导致语音编辑结果的流畅度低、真实感差的技术问题。
Description
相关申请的交叉引用
本公开要求于2023年03月20日提交的申请号为202310299825.7、名称为“语音编辑方法、装置、存储介质及电子装置”的中国专利申请的优先权,该中国专利申请的全部内容通过引用结合在本公开中。
本公开涉及计算机技术和人工智能技术领域,具体而言,涉及一种语音编辑方法、装置、存储介质及电子装置。
随着深度学习的发展,语音合成和基于文本的语音编辑技术取得了较大的进步。然而,相关技术提供的语音编辑方法中,经常出现模型训练和模型测试不匹配的问题,进而导致语音编辑结果的流畅度和真实感较差。
针对上述的问题,目前尚未提出有效的解决方案。
需要说明的是,在上述背景技术部分公开的信息仅用于加强对本公开的背景的理解,因此可以包括不构成对本领域普通技术人员已知的相关技术的信息。
发明内容
本公开至少部分实施例提供了一种语音编辑方法、装置、存储介质及电子装置,以至少解决相关技术中提供的语音编辑方法其训练和测试不匹配导致语音编辑结果的流畅度低、真实感差的技术问题。
根据本公开其中一实施例,提供了一种语音编辑方法,包括:获取待处理的原始音频和目标文本,其中,目标文本用于确定待编辑至原始音频的文本内容;对原始音频中的待编辑部分音频进行语音掩码,得到第一掩码后音频;对目标文本和第一掩码后音频进行语音编辑,得到目标音频。
根据本公开其中一实施例,提供了另一种语音编辑方法,通过终端设备提供一图形用户界面,图形用户界面所显示的内容包括一语音编辑控件,语音编辑方法包括:响应对语音编辑控件执行的触发操作,弹出语音编辑界面;响应对语音编辑界面执行的输入操作,导入原始音频和目标文本,其中,目标文本用于确定待编辑至原始音频的文本内容;响应
对语音编辑界面执行的编辑操作,从原始音频中选定待编辑部分音频;响应对语音编辑界面执行的播放操作,在游戏场景中播放目标音频,其中,目标音频通过对目标文本和掩码后音频进行语音编辑后得到,掩码后音频通过对待编辑部分音频进行语音掩码后得到。
根据本公开其中一实施例,还提供了一种模型训练方法,包括:获取待处理的训练音频和训练文本,其中,训练文本用于确定待编辑至训练音频的文本内容;对训练音频中的待编辑部分音频进行语音掩码,得到掩码后训练音频;采用掩码后训练音频和训练文本对初始语音编辑模型进行训练,得到目标语音编辑模型,其中,目标语音编辑模型用于对目标文本和掩码后原始音频进行语音编辑以得到目标音频,掩码后原始音频通过对原始音频中的待编辑部分音频进行语音掩码后得到。
根据本公开其中一实施例,还提供了一种语音编辑装置,包括:获取模块,用于获取待处理的原始音频和目标文本,其中,目标文本用于确定待编辑至原始音频的文本内容;掩码模块,用于对原始音频中的待编辑部分音频进行语音掩码,得到第一掩码后音频;编辑模块,用于对目标文本和第一掩码后音频进行语音编辑,得到目标音频。
根据本公开其中一实施例,还提供了一种计算机可读存储介质,计算机可读存储介质中存储有计算机程序,其中,计算机程序被设置为运行时执行上述语音编辑方法或者模型训练方法。
根据本公开其中一实施例,还提供了一种电子装置,包括:包括存储器和处理器,存储器中存储有计算机程序,处理器被设置为运行计算机程序以执行上述语音编辑方法或者模型训练方法。
在本申请至少部分实施例中,通过获取待处理的原始音频和目标文本,其中,目标文本用于确定待编辑至原始音频的文本内容;采用对原始音频中的待编辑部分音频进行语音掩码的方式得到第一掩码后音频;进一步对目标文本和第一掩码后音频进行语音编辑,得到目标音频,达到了通过对待执行语音编辑的原始音频先进行语音掩码再进行语音编辑得到目标音频的目的,从而实现了提高语音编辑结果的流畅度和真实感的技术效果,进而解决了相关技术中提供的语音编辑方法其训练和测试不匹配导致语音编辑结果的流畅度低、真实感差的技术问题。
此处所说明的附图用来提供对本公开的进一步理解,构成本公开的一部分,本公开的示意性实施例及其说明用于解释本公开,并不构成对本公开的不当限定。在附图中:
图1是根据本公开其中一实施例的一种语音编辑方法的移动终端的硬件结构框图;
图2是根据本公开其中一实施例的一种语音编辑方法的流程图;
图3是根据本公开其中一实施例的一种可选的语音编辑过程的示意图;
图4是根据本公开其中一实施例的一种可选的声学特征提取过程的示意图;
图5是根据本公开其中一实施例的一种可选的语音编辑方法的流程图;
图6是根据本公开其中一实施例的一种可选的在云端服务器进行语音编辑的示意图;
图7是根据本公开其中一实施例的另一种语音编辑方法的流程图;
图8是根据本公开其中一实施例的一种模型训练方法的流程图;
图9是根据本公开其中一实施例的一种语音编辑装置的结构框图;
图10是根据本公开其中一实施例的一种可选的语音编辑装置的结构框图;
图11是根据本公开其中一实施例的一种电子装置的示意图。
为了使本技术领域的人员更好地理解本公开方案,下面将结合本公开实施例中的附图,对本公开实施例中的技术方案进行清楚、完整地描述,显然,所描述的实施例仅仅是本公开一部分的实施例,而不是全部的实施例。基于本公开中的实施例,本领域普通技术人员在没有做出创造性劳动前提下所获得的所有其他实施例,都应当属于本公开保护的范围。
需要说明的是,本公开的说明书和权利要求书及上述附图中的术语“第一”、“第二”等是用于区别类似的对象,而不必用于描述特定的顺序或先后次序。应该理解这样使用的数据在适当情况下可以互换,以便这里描述的本公开的实施例能够以除了在这里图示或描述的那些以外的顺序实施。此外,术语“包括”和“具有”以及他们的任何变形,意图在于覆盖不排他的包含,例如,包含了一系列步骤或单元的过程、方法、系统、产品或设备不必限于清楚地列出的那些步骤或单元,而是可包括没有清楚地列出的或对于这些过程、方法、产品或设备固有的其它步骤或单元。
需要说明的是,在本公开的说明书中,“例如”一词用来表示“用作例子、例证或说明”。本公开中被描述为“例如”的任何实施例不一定被解释为比其它实施例更优选或更具优势。为了使本领域任何技术人员能够实现和使用本公开,给出了以下描述。在以下描述中,为了解释的目的而列出了细节。应当明白的是,本领域普通技术人员可以认识到,在不使用这些特定细节的情况下也可以实现本公开。在其它实例中,不会对公知的结构和过程进行详细阐述,以避免不必要的细节使本公开的描述变得晦涩。因此,本公开并非旨在限于所示的实施例,而是与符合本公开所公开的原理和特征的最广范围相一致。
在对本公开实施例进行描述的过程中,出现的部分名词或术语适用于如下解释:
语音合成技术(Text-to-Speech,TTS):是一种使用文本到语音的算法将文本数据转换为人声音频的技术。TTS能够用于基于现有的语音库生成带有特定声腔、情感和特征的虚拟人声。
语音编辑:是指对语音进行的处理和修改,使其变为所要求的格式和质量。语音编辑可以包括:处理语音文件(如替换、删除、插入、剪切、连接、校准等);更改声音特性(如增强低频或幅度、减少背景噪声等;设置表情标记(包括位置、注释、单元名称)。上述语音编辑能够帮助了解作品中使用的特定人声或乐器等。
随着深度学习的发展,语音合成和基于文本的语音编辑技术取得了较大的进步。然而,相关技术提供的语音编辑方法中,经常出现模型训练和模型测试不匹配的问题,进而导致语音编辑结果的流畅度和真实感较差。对此,在本公开之前相关技术领域并未提出有效的解决方法。
在本公开的一种可能的实施方式中,针对计算机技术和人工智能技术领域下涉及语音编辑的应用场景中通常所采用的语音编辑方法,发明人经过实践并仔细研究后,仍然存在语音编辑结果流畅度低、真实感差的技术问题,基于此,本公开实施例应用的游戏场景可以是计算机技术或人工智能技术领域中任何涉及语音合成或语音编辑的应用场景,所针对的游戏类型可以是动作类、冒险类、模拟类、角色扮演类和休闲类等。
本公开实施例提出了一种语音编辑方法,采用在语音编辑之前对原始音频进行语音掩码处理以统一语音编辑训练目标和测试目标的技术构思,实现了提升语音编辑结果的流畅度和真实感的技术效果,进而解决了相关技术中提供的语音编辑方法其训练和测试不匹配导致语音编辑结果的流畅度低、真实感差的技术问题。
本公开涉及到的上述方法实施例,可以在终端设备(例如,移动终端、计算机终端或者类似的运算装置)中执行。以运行在移动终端上为例,该移动终端可以是智能手机、平板电脑、掌上电脑以及移动互联网设备、游戏机等终端设备。
图1是根据本公开其中一实施例的一种语音编辑方法的移动终端的硬件结构框图。如图1所示,移动终端可以包括一个或多个(图1中仅示出一个)处理器102、存储器104、传输设备106、输入输出设备108以及显示设备110。以语音编辑方法通过该移动终端应用于电子游戏场景为例,处理器102调用并运行存储器104中存储的计算机程序以执行该语音编辑方法,所生成的目标音频通过传输设备106传输至输入输出设备108和/或显示设备110,进而将该目标音频提供给玩家。
仍然如图1所示,处理器102可以包括但不限于:中央处理器(Central Processing Unit,
CPU)、图形处理器(Graphics Processing Unit,GPU)、数字信号处理(Digital Signal Processing,DSP)芯片、微处理器(Microcontroller Unit,MCU)、可编程逻辑器件(Field Programmable Gate Array,FPGA)、神经网络处理器(Neural-Network Processing Unit,NPU)、张量处理器(Tensor Processing Unit,TPU)、人工智能(Artificial Intelligence,AI)类型处理器等的处理装置。
本领域技术人员可以理解,图1所示的结构仅为示意,其并不对上述移动终端的结构造成限定。例如,移动终端还可包括比图1中所示更多或者更少的组件,或者具有与图1所示不同的配置。
在一些以游戏场景为主的可选实施例中,上述终端设备还可以提供具有触摸触敏表面的人机交互界面,该人机交互界面可以感应手指接触和/或手势来与图形用户界面(Graphical User Interface,GUI)进行人机交互,该人机交互功能可以包括如下交互:创建网页、绘图、文字处理、制作电子文档、游戏、视频会议、即时通信、收发电子邮件、通话界面、播放数字视频、播放数字音乐和/或网络浏览等、用于执行上述人机交互功能的可执行指令被配置/存储在一个或多个处理器可执行的计算机程序产品或可读存储介质中。
本公开涉及到的上述方法实施例,还可以在服务器中执行。其中,服务器可以是独立的物理服务器,也可以是多个物理服务器构成的服务器集群或者分布式系统,还可以是提供云服务、云数据库、云计算、云函数、云存储、网络服务、云通信、中间件服务、域名服务、安全服务、内容分发网络(Content Delivery Network,CDN)、以及大数据和人工智能平台等基础云计算服务的云服务器。以语音编辑方法通过电子游戏服务器应用于电子游戏场景为例,电子游戏服务器可基于该语音编辑方法得到电子游戏场景中的目标音频,并将该目标音频提供给玩家(例如,可以渲染显示在玩家终端的显示屏上,或者,通过全息投影提供给玩家等)。
根据本公开其中一实施例,提供了一种语音编辑方法的实施例,需要说明的是,在附图的流程图示出的步骤可以在诸如一组计算机可执行指令的计算机系统中执行,并且,虽然在流程图中示出了逻辑顺序,但是在某些情况下,可以以不同于此处的顺序执行所示出或描述的步骤。
在本实施例中提供了一种运行于上述终端设备的一种语音编辑方法,图2是根据本公开其中一实施例的一种语音编辑方法的流程图,如图2所示,该方法包括如下步骤:
步骤S21,获取待处理的原始音频和目标文本,其中,目标文本用于确定待编辑至原
始音频的文本内容。
本公开提供的语音编辑方法可以但不限于适用于如下应用场景:电子游戏(例如对游戏角色或游戏场景对应的原始音频进行编辑),语音导航系统(例如应用于汽车、机器人等自动化设备中,以将文字转换为语音提示用户当前的位置和操作步骤),数字电话服务(例如在语音通话、在线聊天、智能客服等场景中将文字快速转化为语音来作为数字客户服务的一部分),智能虚拟助手,技术教育和讲座(例如将幻灯片中的文字信息快速地生成对应的声频来做讲座),有声书/新闻/广告(例如将文字内容快速地生成声频内容)。
上述原始音频为上述应用场景中待处理的音频数据。上述目标文本为上述应用场景中待编辑至原始音频的的文本内容。本公开实施例的上述语音编辑方法可以运行在客户端上。上述指定可以由用户通过客户端指定,也可以由运行于客户端的计算机应用程序根据场景需求生成控制指令来指定。此外,本公开实施例的上述语音编辑方法还可以运行在服务端上,该服务端可以是独立的服务器、服务器集群或者云服务器,特别是,上述方法运行于云服务器上时,通过软件即服务(Software as a Service,SaaS)的方式与客户端进行交互,获取客户端发送的原始音频和目标文本,进而进行对应的语音编辑,然后将语音编辑结果返回客户端。
步骤S22,对原始音频中的待编辑部分音频进行语音掩码,得到第一掩码后音频;
为解决语音合成或语音编辑问题,与相关技术中所采用的语音合成模型相比,本公开的上述语音编辑过程所采用的语音编辑模型引入了掩码训练策略,也就是说,将解决语音合成或语音编辑问题的语音合成模型转换成为语音编辑模型。上述包含掩码训练策略的语音编辑模型在模型训练阶段和模型测试阶段(或模型应用阶段)的目标保持一致,从而能够提升语音编辑模型的语音编辑效果。
具体地,上述语音掩码的作用在于:对原始音频中的待编辑部分音频进行局部修改,该局部修改的语音编辑方式可以包括以下至少之一:替换方式、删除方式和插入方式。即,上述待编辑部分音频可以是用户从原始音频中选定任意部分音频段,通过采用语音掩码技术手段将待编辑部分音频对应的原有音频内容从原始音频中抹除,并保留原始音频中除该待编辑部分音频之外的其余上下文音频内容。
步骤S23,对目标文本和第一掩码后音频进行语音编辑,得到目标音频。
上述对目标文本和第一掩码后音频进行语音编辑,可以是按照目标文本所确定的待编辑至原始音频的文本内容,对第一掩码后音频进行语音编辑。采用上述方案所得到的目标音频在语音流畅度和真实感上的表现优于相关技术的语音编辑结果。
以智能虚拟助手对用户的原始语音进行语音编辑的场景为例,对本公开实施例上述方法的技术方案进行进一步介绍。例如,上述方法运行于智能虚拟助手对应的客户端。
客户端录制用户输入的待执行语音编辑的原始语音(也即获取原始音频),以及获取用户输入的目标文本,并根据目标文本获取确定待编辑至原始音频的文本内容。在根据本公开实施例的其中一种可选的实施方式中,原始音频所录制的待执行语音编辑的原始语音的内容为“基于深度学习的语音编辑”,目标文本所确定的文本内容为“机器学习”,上述场景中的语音编辑需求为:将原始音频中的“深度学习”替换为“机器学习”。也就是说,原始音频中的待编辑部分音频为“深度学习”对应的音频。
在本公开至少部分实施例中,通过获取待处理的原始音频和目标文本,其中,目标文本用于确定待编辑至原始音频的文本内容;采用对原始音频中的待编辑部分音频进行语音掩码的方式得到第一掩码后音频;进一步对目标文本和第一掩码后音频进行语音编辑,得到目标音频,达到了通过对待执行语音编辑的原始音频先进行语音掩码再进行语音编辑得到目标音频的目的,从而实现了提高语音编辑结果的流畅度和真实感的技术效果,进而解决了相关技术中提供的语音编辑方法其训练和测试不匹配导致语音编辑结果的流畅度低、真实感差的技术问题。
在一个示例性应用场景中,上述原始音频可以为游戏应用(APP)中待使用的原始游戏语音,上述第一掩码后音频可以为掩码后游戏语音。具体地,可以对游戏应用中待使用的原始游戏语音中待编辑部分语音进行语音掩码,得到掩码后游戏语音,以及对游戏应用中待使用的游戏内容文本和掩码后游戏语音进行语音编辑,得到目标游戏语音。
在现有游戏场景中,非玩家角色(Non-Player Character,简称为NPC)在与游戏玩家操控的玩家角色进行语音交互时,通常采用预先录制的固定音频进行语音交互。例如:在游戏玩家点击NPC时,NPC通常会播放“欢迎我们伟大的勇士来到这里!”。此时,无论哪位玩家点击NPC,都会收到千篇一律的回复。为了增强游戏玩家在游戏场景内游戏交互体验,可以应用本公开所提出的上述技术方案,在上述原始游戏语音中加入玩家角色的名称。具体地,对游戏应用中待使用的原始游戏语音“欢迎我们伟大的勇士来到这里!”中待编辑部分语音“勇士来到”进行语音掩码,得到掩码后游戏语音“欢迎我们伟大的【XXXX】这里!”。然后,再对游戏应用中待使用的游戏内容文本(勇士玩家角色名称来到)和掩码后游戏语音“欢迎我们伟大的【XXXX】这里!”进行语音编辑,得到目标游戏语音“欢迎我们伟大的【勇士玩家角色名称来到】这里!”。由此,针对不同游戏玩家可以播放个性化音频。
在另一个示例性应用场景中,上述原始音频可以为配音应用中待使用的原始多媒体配音(例如:影视剧配音、动漫配音等),上述第一掩码后音频可以为掩码后多媒体配音。具体地,可以对配音应用中待使用的原始多媒体配音中待编辑部分配音进行语音掩码,得到掩码后多媒体配音,以及对配音应用中待使用的配音内容文本和掩码后多媒体配音进行语音编辑,得到目标多媒体配音。
在现有配音场景中,配音演员有时在依据影视剧的剧本完成配音之后,会发现原始多媒体配音中遗漏了部分文字没有配音。此时,如果重新录制整个配音不仅会耗费人力和时间,而且也难以确保重新录制的配音不会再次出现遗漏。为此,可以应用本公开所提出的上述技术方案,在上述原始影视剧配音中,插入剧本中的遗漏文字。具体地,对配音应用中待使用的原始影视剧配音“那好,我们一起去吧!”中待编辑部分配音“我们一起”进行语音掩码,得到掩码后多媒体配音“那好,【XXXX】去吧!”,以及对配音应用中待使用的配音内容文本(我们明天一起)和掩码后多媒体配音“那好,【XXXX】去吧!”进行语音编辑,得到目标多媒体配音“那好,我们明天一起去吧!”。由此,不仅可以有效地避免重复配音所带来的繁琐工作,而且还可以及时更正原始影视剧配音中存在的配音缺陷。以智能虚拟助手对用户的原始语音进行语音编辑的场景为例,对本公开实施例上述方法的技术方案进行进一步介绍。例如,上述方法运行于智能虚拟助手对应的客户端。
可选地,在步骤S22中,对原始音频中的待编辑部分音频进行语音掩码,得到第一掩码后音频,可以包括以下执行步骤:
步骤S221,获取待编辑部分音频在原始音频中的位置信息;
步骤S222,基于位置信息对原始音频中的待编辑部分音频进行语音掩码,得到第一掩码后音频。
在一种可选的实施方式中,客户端通过对原始音频和待编辑部分音频进行比对,得到上述位置信息,该位置信息可以是待编辑部分音频在原始音频对应的原始语音中多个音素的位置ID确定的。利用语音编辑模型中的掩码部分,基于位置信息对原始音频中的待编辑部分音频进行语音掩码,得到第一掩码后音频。
仍然以智能虚拟助手对用户的原始语音进行语音编辑的场景为例,对原始音频中的待编辑部分音频进行语音掩码的目的在于,将待编辑部分音频从原始音频中抽离。例如,将“基于深度学习的语音编辑”对应的原始音频中“深度学习”对应的音频抽离,则得到的第一掩码后音频对应的内容为“基于【XXXX】的语音编辑”。
图3是根据本公开其中一实施例的一种可选的语音编辑过程的示意图,如图3所示,
获取原始音频后,利用掩码部分对原始音频进行掩码处理,得到第一掩码后音频。
仍然以智能虚拟助手对用户的原始语音进行语音编辑的场景为例,确定“深度学习”在“基于深度学习的语音编辑”中的位置信息,该位置信息可以是待编辑部分音频对应的内容的字素ID。然后,基于将原始音频中的“深度学习”替换为“机器学习”的场景需求以及上述位置信息,对音频“基于深度学习的语音编辑”中的音频“深度学习”进行语音掩码,得到音频“基于【XXXX】的语音编辑”。
可选地,在步骤S23中,对目标文本和第一掩码后音频进行语音编辑,得到目标音频,可以包括以下执行步骤:
步骤S231,对目标文本进行语音转换,得到中间音频;
步骤S232,对中间音频和第一掩码后音频进行语音拼接,得到目标音频。
在对目标文本和第一掩码后音频进行语音编辑以得到目标音频的过程中,可以对目标文本进行语音转换以得到中间音频。例如:目标文本所确定的文本内容为“机器学习”,通过对目标文本进行语音转换所得到的中间音频即为“机器学习”对应的音频段。然后,再将上述第一掩码后音频“基于XXXX的语音编辑”与中间音频即为“机器学习”进行语音拼接以得到目标音频“基于机器学习的语音编辑”。
可选地,在步骤S231中,对目标文本进行语音转换,得到目标音频,可以包括以下执行步骤:
步骤S2311,对目标文本和第一掩码后音频进行语音编辑,得到目标声学特征,其中,目标声学特征用于确定目标文本对应的音频段;
步骤S2312,对目标声学特征进行声码转换,得到中间音频。
仍然以智能虚拟助手对用户的原始语音进行语音编辑的场景为例,对“机器学习”和音频“基于【XXXX】的语音编辑”进行语音编辑,得到“机器学习”对应的目标声学特征;对“机器学习”对应的目标声学特征进行声码转换,得到音频“机器学习”(即中间音频);对音频“机器学习”与音频“基于【XXXX】的语音编辑”进行拼接,得到目标音频,目标音频为“基于机器学习的语音编辑”。
上述音频拼接的目的为:基于“深度学习”在原始音频中对应的位置,将目标文本的内容“机器学习”与第一掩码后音频对应的内容“基于【XXXX】的语音编辑”拼接起来,得到目标音频对应的内容“基于机器学习的语音编辑”。
如图3所示,基于目标文本和第一掩码后音频进行语音编辑,得到目标声学特征,然后利用神经声码器对目标声学特征进行声码转换得到中间音频,进而对中间音频与第一掩
码后音频进行拼接,得到目标音频。
通过上述步骤S2311至步骤S2312,利用语音编辑方法生成目标文本对应的目标声学特征,从而得到中间音频。由于上述第一掩码后音频为基于待编辑部分音频在原始音频中的位置信息得到的,对中间音频与第一掩码后音频进行拼接所得到目标音频能够在语音流畅度和真实感上具有较好的表现。
可选地,在步骤S2311中,对目标文本和第一掩码后音频进行语音编辑,得到目标声学特征,可以包括以下执行步骤:
步骤S23111,对目标文本进行字素到音素转换,得到音素序列;
步骤S23112,使用目标语音编辑模型对音素序列和第一掩码后音频进行语音编辑,得到目标声学特征,其中,目标语音编辑模型采用多组数据通过深度学习训练得到,多组数据包括:训练音频和训练文本,训练文本为训练音频中的待编辑部分音频对应的文本。
仍然如图3所示,对目标文本和第一掩码后音频进行语音编辑得到目标声学特征的过程中,首先,利用转换部分对目标文本进行字素到音素转换,得到目标文本对应的音素序列。上述转换部分为图到树(Graph to Tree,G2T)转换部分,该G2T转换部分用于通过自然语言处理技术将目标文本(可以是单个句子或多个已分割的句子)表示为语法树。上述音素序列可以与语法树的多个节点相对应。然后,上述目标语音编辑模型为序列到序列语音编辑部分,利用序列到序列语音编辑部分对音素序列和第一掩码后音频进行语音编辑得到目标声学特征。
需要说明的是,上述用于训练目标语音编辑模型的多组数据可以是历史语音编辑结果,即,在应用场景中使用语音编辑模型对指定的训练音频进行语音编辑得到的数据。多组数据中每组数据包括的训练音频与训练文本相对应。
仍然以智能虚拟助手对用户的原始语音进行语音编辑的场景为例,对目标文本“机器学习”进行字素到音素转换,得到“机器学习”对应的音素序列;使用目标语音编辑模型对“机器学习”对应的音素序列和音频“基于【XXXX】的语音编辑”进行语音编辑,得到目标声学特征。
可选地,上述语音编辑方法还可以包括以下执行步骤:
步骤S241,对训练音频中的待编辑部分音频进行语音掩码,得到第二掩码后音频;
步骤S242,使用初始语音编辑模型对第二掩码后音频和训练文本进行语音编辑,得到预测声学特征;
步骤S243,通过预测声学特征与训练文本对应的真实声学特征确定目标损失;
步骤S244,利用目标损失对初始语音编辑模型的参数进行更新,得到目标语音编辑模型。
通过上述步骤S241至步骤S244,在目标语音编辑模型的训练过程中,基于训练音频和训练文本计算模型训练的目标损失,采用该目标损失对初始语音编辑模型的模型参数进行优化更新,得到目标语音编辑模型。上述目标损失对应的损失函数可以是任意常用的损失函数,本公开实施例并不对目标损失的计算方法进行限定。
可选地,目标语音编辑模型包括:编码器、特征调节器和解码器,在步骤S23112中,使用目标语音编辑模型对音素序列和第一掩码后音频进行语音编辑,得到目标声学特征,可以包括以下执行步骤:
步骤S23113,使用编码器对音素序列进行文本特征空间编码,得到文本特征;
步骤S23114,使用特征调节器对文本特征和第一掩码后音频进行特征调节,得到第一听觉感知特征,其中,第一听觉感知特征为目标文本对应的听觉感知特征;
步骤S23115,使用解码器对第一听觉感知特征进行声学解码,得到目标声学特征。
图4是根据本公开其中一实施例的一种可选的声学特征提取过程的示意图,如图4所示,目标语音编辑模型中的序列到序列语音编辑部分包括:编码器、特征调节器和解码器,其中,解码器为声学解码器。使用上述编码器对目标文本对应的因素序列进行文本特征空间编码,得到文本特征,并将文本特征同时传递给特征调节器和声学解码器。然后,使用上述特征调节器对文本特征和第一掩码后音频进行特征调节,得到第一听觉感知特征,其中,第一听觉感知特征为目标文本对应的听觉感知特征。进一步地,使用上述声学解码器对文本特征和第一听觉感知特征进行声学解码,得到上述目标声学特征。
具体地,上述编码器通过非线性变换方式将音素序列映射至高维的文本特征空间进行编码,得到上述文本特征。
具体地,上述特征调节器通过非线性变换预测方式,基于文本特征和第一掩码后音频进行特征预测和特征调节,得到上述第一听觉感知特征。同理,上述声学解码器通过非线性变换预测方式,对文本特征和第一听觉感知特征进行声学特征预测,得到上述目标声学特征。
仍然以智能虚拟助手对用户的原始语音进行语音编辑的场景为例,使用编码器对“机器学习”对应的音素序列进行文本特征空间编码,得到“机器学习”对应的文本特征;使用特征调节器对“机器学习”对应的文本特征和音频“基于【XXXX】的语音编辑”进行特征调节,得到目标文本中的“机器学习”对应的听觉感知特征;使用解码器对“机器学
习”对应的听觉感知特征进行声学解码,得到目标声学特征。
可选地,在步骤S23114中,使用特征调节器对文本特征和第一掩码后音频进行特征调节,得到第一听觉感知特征,可以包括以下执行步骤:
步骤S23116,从第一掩码后音频中提取第二听觉感知特征,其中,第二听觉感知特征为原始音频中与待编辑部分音频关联的上下文音频对应的听觉感知特征;
步骤S23117,使用特征调节器对文本特征和第二听觉感知特征进行特征调节,得到第一听觉感知特征。
仍然如图4所示,上述特征调节器还可以通过下述步骤得到第一听觉感知特征:对第一掩码后音频进行特征提取,得到第二听觉感知特征。使用特征调节器对文本特征和第二听觉感知特征进行特征调节,得到第一听觉感知特征。仍然以智能虚拟助手对用户的原始语音进行语音编辑的场景为例,上述第二听觉感知特征为“基于【XXXX】的语音编辑”对应的听觉感知特征,听觉感知特征至少包括:音高特征、能量特征和时长特征。使用特征调节器对“机器学习”的文本特征和“基于【XXXX】的语音编辑”的第二听觉感知特征进行特征调节,得到“机器学习”对应的听觉感知特征(即第一听觉感知特征)。
可选地,第一听觉感知特征包括以下至少之一:目标文本对应的音高;目标文本对应的能量;目标文本对应的时长。
仍然以智能虚拟助手对用户的原始语音进行语音编辑的场景为例,第一听觉感知特征至少包括以下之一:目标文本中文本内容“机器学习”对应的音高特征、能量特征和时长特征。
以智能虚拟助手对用户的原始语音进行语音编辑的场景为例,语音编辑场景的场景需求对应的语音编辑方式除了上述的替换方式,还可以包括插入方式和删除方式。
仍然以原始音频所录制的待执行语音编辑的原始语音的内容为“基于深度学习的语音编辑”为例,当场景需求对应的语音编辑方式为在“学习”与“的”之间插入“技术”时,上述原始音频中待编辑部分音频为插入位置的上下文内容,即“学习的”,目标文本对应的文本内容为“学习技术的”。然后利用类似于上述替换方式的处理方法,对目标文本和第一掩码后音频进行语音编辑,得到目标音频,此目标音频对应的内容为“基于深度学习技术的语音编辑”。
仍然以原始音频所录制的待执行语音编辑的原始语音的内容为“基于深度学习的语音编辑”为例,当场景需求对应的语音编辑方式为在“深度学习的”之后删除“语音”时,上述原始音频中待编辑部分音频为删除位置的上下文内容,即“的语音编辑”,目标文本
对应的文本内容为“的编辑”。然后利用类似于上述替换方式的处理方法,对目标文本和第一掩码后音频进行语音编辑,得到目标音频,此目标音频对应的内容为“基于深度学习技术的编辑”。
容易理解的是,通过本公开实施例提供的语音编辑方法,引入音频掩码处理,能够将语音编辑过程中的训练目标与测试目标相统一,使得对原始音频中的待编辑部分音频进行编辑后得到的目标音频(特别是在音频拼接位置)具备更高的流畅度和真实感。
容易理解的是,本公开实施例提供的语音编辑方法中,通过音频掩码机制,在对语音编辑模型进行训练的过程中,将原始音频和原始音频对应的第一掩码后音频和目标音频作为训练样本进行模型训练,由此,使得训练得到的目标语音编辑模型能够基于第一掩码后音频预测得到目标文本对应的部分音频(即中间音频,也即待拼接至第一掩码后音频的部分音频)。因此,上述目标语音编辑模型的训练阶段与测试阶段(或场景应用阶段)的目标一致,避免了相关技术中模型的训练目标与测试目标不匹配导致的语音编辑效果差的问题,此外,上述训练流程还能保证音频拼接更加流畅。
本公开其中一实施例还提供了一种语音编辑方法,该语音编辑方法在云端服务器上运行,图5是根据本公开其中一实施例的一种可选的语音编辑方法的流程图,如图5所示,该语音编辑方法,包括:
步骤S51,接收来自于客户端的待处理的原始音频和目标文本,其中,目标文本用于确定待编辑至原始音频的文本内容;
步骤S52,对原始音频中的待编辑部分音频进行语音掩码以得到第一掩码后音频,以及对目标文本和第一掩码后音频进行语音编辑以得到目标音频;
步骤S53,将目标音频反馈至客户端。
可选地,图6是根据本公开其中一实施例的一种可选的在云端服务器进行语音编辑的示意图,如图6所示,客户端将原始音频和目标文本上传至云端服务器,其中,目标文本用于确定待编辑至原始音频的文本内容;云端服务器对原始音频中的待编辑部分音频进行语音掩码以得到第一掩码后音频,以及对目标文本和第一掩码后音频进行语音编辑以得到目标音频。然后,云端服务器会向上述客户端反馈目标音频,最终的目标音频会通过客户端的图形用户界面提供给用户。
需要说明的是,本公开实施例所提供的上述语音编辑方法,可以但不限于适用于语音导航系统、数字电话服务、智能虚拟助手、技术教育/讲座和有声书/新闻/广告等实际应用场景,通过SaaS服务端和客户端进行交互的方式,采用客户端向服务端提供原始音频和
目标文本,服务端对原始音频中的待编辑部分音频进行语音掩码以得到第一掩码后音频,以及对目标文本和第一掩码后音频进行语音编辑的方式得到目标音频,服务端将目标音频返回客户端并提供给用户。
本公开其中一实施例还提供了又一种语音编辑方法,通过终端设备提供一图形用户界面,图形用户界面所显示的内容包括一语音编辑控件,图7是根据本公开其中一实施例的另一种语音编辑方法的流程图,如图7所示,该语音编辑方法,包括:
步骤S71,响应对语音编辑控件执行的触发操作,弹出语音编辑界面;
步骤S72,响应对语音编辑界面执行的输入操作,导入原始音频和目标文本,其中,目标文本用于确定待编辑至原始音频的文本内容;
步骤S73,响应对语音编辑界面执行的编辑操作,从原始音频中选定待编辑部分音频;
步骤S74,响应对语音编辑界面执行的播放操作,在游戏场景中播放目标音频,其中,目标音频通过对目标文本和掩码后音频进行语音编辑后得到,掩码后音频通过对待编辑部分音频进行语音掩码后得到。
上述图形用户界面中至少显示有语音编辑控件,用户通过对该语音编辑控件执行触发操作,在上述图形用户界面中弹出语音编辑界面,进一步地,用户通过该语音编辑界面导入待进行语音编辑的原始音频和目标文本,其中,目标文本用于确定待编辑至原始音频的文本内容;以及从原始音频中选定待编辑部分音频,进而对目标文本和掩码后音频进行语音编辑,并在游戏场景中播放目标音频。
上述触发操作、输入操作、编辑操作和播放操作均可以是触控操作,该触控操作可以包括单点触控、多点触控,其中,每个触控点的触控操作可以包括点击、长按、重按、划动等。上述触发操作、输入操作、编辑操作和播放操作还可以是通过鼠标、键盘等输入设备实现的操作。
上述输入操作对应于图形用户界面中的第一控件或第一触控区域(如键入框、手写输入区域、输入按钮(例如按住该按钮则录制并接收用户的语音));上述编辑操作对应于图形用户界面中的第二控件或第二触控区域(如编辑框、编辑选项栏等);上述播放操作对应于图形用户界面中的第三控件或第三触控区域(如播放区域、播放按钮等)。
在一个可选实施例中,可以在游戏客户端内插入上述语音编辑工具,并在图形用户界面内提供一个语音编辑控件。游戏玩家通过对该语音编辑控件执行触控操作或者鼠标点击操作,可以弹出语音编辑界面。在语音编辑界面内,游戏玩家可以通过导入控件导入原始游戏语音和待使用的游戏内容文本,并从原始游戏语音中选定待编辑部分音频。例如:待
使用的原始游戏语音为“欢迎我们伟大的勇士来到这里!”,用户既可以选定“勇士来到”对应的音频段,也可以选定首个文字“勇”与末尾文字“到”。然后,游戏服务器便可以实时对游戏应用中待使用的原始游戏语音“欢迎我们伟大的勇士来到这里!”中待编辑部分语音“勇士来到”进行语音掩码,得到掩码后游戏语音“欢迎我们伟大的【XXXX】这里!”。然后,再对游戏应用中待使用的游戏内容文本(勇士玩家角色名称来到)和掩码后游戏语音“欢迎我们伟大的【XXXX】这里!”进行语音编辑,得到目标游戏语音“欢迎我们伟大的【勇士玩家角色名称来到】这里!”。由此,针对不同游戏玩家可以播放个性化音频。综上,根据本公开实施例提供的语音编辑场景,能够以可视化的形式与用户进行交互,并根据用户的输入操作、编辑操作和播放操作生成对应的目标音频,有利于实际场景中的应用。
综上,根据本公开实施例提供的语音编辑场景,能够以可视化的形式与用户进行交互,并根据用户的输入操作、编辑操作和播放操作生成对应的目标音频,有利于实际场景中的应用。
本公开其中一实施例还提供了一种模型训练方法,图8是根据本公开其中一实施例的一种模型训练方法的流程图,如图8所示,该模型训练方法包括:
步骤S81,获取待处理的训练音频和训练文本,其中,训练文本用于确定待编辑至训练音频的文本内容;
步骤S82,对训练音频中的待编辑部分音频进行语音掩码,得到掩码后训练音频;
步骤S83,采用掩码后训练音频和训练文本对初始语音编辑模型进行训练,得到目标语音编辑模型,其中,目标语音编辑模型用于对目标文本和掩码后原始音频进行语音编辑以得到目标音频,掩码后原始音频通过对原始音频中的待编辑部分音频进行语音掩码后得到。
在对用于实现上述语音编辑方法的目标语音编辑模型进行训练的过程中,模型对应的输入要求和输出目标与模型测试或模型应用过程保持一致。也即,在应用场景中使用目标语音编辑模型基于原始音频和目标文本进行语音编辑,得到掩码后音频,进而得到目标音频,对应地,在训练过程中,基于训练音频和训练文本进行语音掩码得到掩码后音频,进而得到目标训练音频,从而对初始语音编辑模型进行参数优化得到目标语音编辑模型。
容易理解的是,本公开实施例提供的语音编辑方法中,通过音频掩码机制,在对语音编辑模型进行训练的过程中,将原始音频和原始音频对应的第一掩码后音频和目标音频作为训练样本进行模型训练,由此,使得训练得到的目标语音编辑模型能够基于第一掩码后
音频预测得到目标文本对应的部分音频(即中间音频,也即待拼接至第一掩码后音频的部分音频)。因此,上述目标语音编辑模型的训练阶段与测试阶段(或场景应用阶段)的目标一致,避免了相关技术中模型的训练目标与测试目标不匹配导致的语音编辑效果差的问题,此外,上述训练流程还能保证音频拼接更加流畅。
通过以上的实施方式的描述,本领域的技术人员可以清楚地了解到根据上述实施例的方法可借助软件加必需的通用硬件平台的方式来实现,当然也可以通过硬件,但很多情况下前者是更佳的实施方式。基于这样的理解,本公开的技术方案本质上或者说对相关技术做出贡献的部分可以以软件产品的形式体现出来,该计算机软件产品存储在一个存储介质(如磁碟、光盘)中,包括若干指令用以使得一台终端设备(可以是手机,计算机,服务器,或者网络设备等)执行本公开各个实施例所述的方法。
在本实施例中还提供了一种语音编辑装置,该装置用于实现上述实施例及优选实施方式,已经进行过说明的不再赘述。如以下所使用的,术语“模块”可以实现预定功能的软件和/或硬件的组合。尽管以下实施例所描述的装置较佳地以软件来实现,但是硬件,或者软件和硬件的组合的实现也是可能并被构想的。
图9是根据本公开其中一实施例的一种语音编辑装置的结构框图,如图9所示,该装置包括:获取模块901,用于获取待处理的原始音频和目标文本,其中,目标文本用于确定待编辑至原始音频的文本内容;掩码模块902,用于对原始音频中的待编辑部分音频进行语音掩码,得到第一掩码后音频;编辑模块903,用于对目标文本和第一掩码后音频进行语音编辑,得到目标音频。
可选地,上述掩码模块902,还用于:获取待编辑部分音频在原始音频中的位置信息;基于位置信息对原始音频中的待编辑部分音频进行语音掩码,得到第一掩码后音频。
可选地,上述编辑模块903,还用于:对目标文本进行语音转换,得到中间音频;对中间音频和第一掩码后音频进行语音拼接,得到目标音频。
可选地,上述编辑模块903,还用于:对目标文本和第一掩码后音频进行语音编辑,得到目标声学特征,其中,目标声学特征用于确定目标文本对应的音频段;对目标声学特征进行声码转换,得到中间音频。
可选地,上述编辑模块903,还用于:对目标文本进行字素到音素转换,得到音素序列;使用目标语音编辑模型对音素序列和第一掩码后音频进行语音编辑,得到目标声学特征,其中,目标语音编辑模型采用多组数据通过深度学习训练得到,多组数据包括:训练音频和训练文本,训练文本为训练音频中的待编辑部分音频对应的文本。
可选地,图10是根据本公开其中一实施例的一种可选的语音编辑装置的结构框图,如图10所示,该装置除包括图9所示的所有模块外,还包括:更新模块904,用于对训练音频中的待编辑部分音频进行语音掩码,得到第二掩码后音频;使用初始语音编辑模型对第二掩码后音频和训练文本进行语音编辑,得到预测声学特征;通过预测声学特征与训练文本对应的真实声学特征确定目标损失;利用目标损失对初始语音编辑模型的参数进行更新,得到目标语音编辑模型。
可选地,上述编辑模块903,还用于:使用编码器对音素序列进行文本特征空间编码,得到文本特征;使用特征调节器对文本特征和第一掩码后音频进行特征调节,得到第一听觉感知特征,其中,第一听觉感知特征为目标文本对应的听觉感知特征;使用解码器对第一听觉感知特征进行声学解码,得到目标声学特征。
可选地,上述编辑模块903,还用于:从第一掩码后音频中提取第二听觉感知特征,其中,第二听觉感知特征为原始音频中与待编辑部分音频关联的上下文音频对应的听觉感知特征;使用特征调节器对文本特征和第二听觉感知特征进行特征调节,得到第一听觉感知特征。
可选地,在上述语音编辑装置中,第一听觉感知特征包括以下至少之一:目标文本对应的音高;目标文本对应的能量;目标文本对应的时长。
可选地,上述掩码模块902,还用于:对游戏应用中待使用的原始游戏语音中待编辑部分语音进行语音掩码,得到掩码后游戏语音;对目标文本和第一掩码后音频进行语音编辑,得到目标音频包括:对游戏应用中待使用的游戏内容文本和掩码后游戏语音进行语音编辑,得到目标游戏语音。
可选地,上述掩码模块902,还用于:对配音应用中待使用的原始多媒体配音中待编辑部分配音进行语音掩码,得到掩码后多媒体配音;对目标文本和第一掩码后音频进行语音编辑,得到目标音频包括:对配音应用中待使用的配音内容文本和掩码后多媒体配音进行语音编辑,得到目标多媒体配音。
需要说明的是,上述各个模块是可以通过软件或硬件来实现的,对于后者,可以通过以下方式实现,但不限于此:上述模块均位于同一处理器中;或者,上述各个模块以任意组合的形式分别位于不同的处理器中。
本公开的实施例还提供了一种计算机可读存储介质,该计算机可读存储介质中存储有计算机程序,其中,该计算机程序被设置为运行时执行上述任一项方法实施例中的步骤。
可选地,在本实施例中,上述计算机可读存储介质可以包括但不限于:U盘、只读存
储器(Read-Only Memory,ROM)、随机存取存储器(Random Access Memory,RAM)、移动硬盘、磁碟或者光盘等各种可以存储计算机程序的介质。
可选地,在本实施例中,上述计算机可读存储介质可以位于计算机网络中计算机终端群中的任意一个计算机终端中,或者位于移动终端群中的任意一个移动终端中。
可选地,在本实施例中,上述计算机可读存储介质可以被设置为存储用于执行以下步骤的计算机程序:
S1,获取待处理的原始音频和目标文本,其中,目标文本用于确定待编辑至原始音频的文本内容;
S2,对原始音频中的待编辑部分音频进行语音掩码,得到第一掩码后音频;
S3,对目标文本和第一掩码后音频进行语音编辑,得到目标音频。
可选地,上述计算机可读存储介质还被设置为存储用于执行以下步骤的计算机程序:获取待编辑部分音频在原始音频中的位置信息;基于位置信息对原始音频中的待编辑部分音频进行语音掩码,得到第一掩码后音频。
可选地,上述计算机可读存储介质还被设置为存储用于执行以下步骤的计算机程序:对目标文本进行语音转换,得到中间音频;对中间音频和第一掩码后音频进行语音拼接,得到目标音频。
可选地,上述计算机可读存储介质还被设置为存储用于执行以下步骤的计算机程序:对目标文本和第一掩码后音频进行语音编辑,得到目标声学特征,其中,目标声学特征用于确定目标文本对应的音频段;对目标声学特征进行声码转换,得到中间音频。
可选地,上述计算机可读存储介质还被设置为存储用于执行以下步骤的计算机程序:对目标文本进行字素到音素转换,得到音素序列;使用目标语音编辑模型对音素序列和第一掩码后音频进行语音编辑,得到目标声学特征,其中,目标语音编辑模型采用多组数据通过深度学习训练得到,多组数据包括:训练音频和训练文本,训练文本为训练音频中的待编辑部分音频对应的文本。
可选地,上述计算机可读存储介质还被设置为存储用于执行以下步骤的计算机程序:对训练音频中的待编辑部分音频进行语音掩码,得到第二掩码后音频;使用初始语音编辑模型对第二掩码后音频和训练文本进行语音编辑,得到预测声学特征;通过预测声学特征与训练文本对应的真实声学特征确定目标损失;利用目标损失对初始语音编辑模型的参数进行更新,得到目标语音编辑模型。
可选地,上述计算机可读存储介质还被设置为存储用于执行以下步骤的计算机程序:
使用编码器对音素序列进行文本特征空间编码,得到文本特征;使用特征调节器对文本特征和第一掩码后音频进行特征调节,得到第一听觉感知特征,其中,第一听觉感知特征为目标文本对应的听觉感知特征;使用解码器对第一听觉感知特征进行声学解码,得到目标声学特征。
可选地,上述计算机可读存储介质还被设置为存储用于执行以下步骤的计算机程序:从第一掩码后音频中提取第二听觉感知特征,其中,第二听觉感知特征为原始音频中与待编辑部分音频关联的上下文音频对应的听觉感知特征;使用特征调节器对文本特征和第二听觉感知特征进行特征调节,得到第一听觉感知特征。
可选地,上述计算机可读存储介质还被设置为存储用于执行以下步骤的计算机程序:第一听觉感知特征包括以下至少之一:目标文本对应的音高;目标文本对应的能量;目标文本对应的时长。
可选地,上述计算机可读存储介质还被设置为存储用于执行以下步骤的计算机程序:对游戏应用中待使用的原始游戏语音中待编辑部分语音进行语音掩码,得到掩码后游戏语音;对目标文本和第一掩码后音频进行语音编辑,得到目标音频包括:对游戏应用中待使用的游戏内容文本和掩码后游戏语音进行语音编辑,得到目标游戏语音。
可选地,上述计算机可读存储介质还被设置为存储用于执行以下步骤的计算机程序:对配音应用中待使用的原始多媒体配音中待编辑部分配音进行语音掩码,得到掩码后多媒体配音;对目标文本和第一掩码后音频进行语音编辑,得到目标音频包括:对配音应用中待使用的配音内容文本和掩码后多媒体配音进行语音编辑,得到目标多媒体配音。
可选地,上述计算机可读存储介质还被设置为存储用于执行以下步骤的计算机程序:接收来自于客户端的待处理的原始音频和目标文本,其中,目标文本用于确定待编辑至原始音频的文本内容;对原始音频中的待编辑部分音频进行语音掩码以得到第一掩码后音频,以及对目标文本和第一掩码后音频进行语音编辑以得到目标音频;将目标音频反馈至客户端。
可选地,上述计算机可读存储介质还被设置为存储用于执行以下步骤的计算机程序:响应对语音编辑控件执行的触发操作,弹出语音编辑界面;响应对语音编辑界面执行的输入操作,导入原始音频和目标文本,其中,目标文本用于确定待编辑至原始音频的文本内容;响应对语音编辑界面执行的编辑操作,从原始音频中选定待编辑部分音频;响应对语音编辑界面执行的播放操作,在游戏场景中播放目标音频,其中,目标音频通过对目标文本和掩码后音频进行语音编辑后得到,掩码后音频通过对待编辑部分音频进行语音掩码后
得到。
可选地,上述计算机可读存储介质还被设置为存储用于执行以下步骤的计算机程序:获取待处理的训练音频和训练文本,其中,训练文本用于确定待编辑至训练音频的文本内容;对训练音频中的待编辑部分音频进行语音掩码,得到掩码后训练音频;采用掩码后训练音频和训练文本对初始语音编辑模型进行训练,得到目标语音编辑模型,其中,目标语音编辑模型用于对目标文本和掩码后原始音频进行语音编辑以得到目标音频,掩码后原始音频通过对原始音频中的待编辑部分音频进行语音掩码后得到。
在上述实施例的计算机可读存储介质中,提供了一种实现语音编辑方法的技术方案。通过获取待处理的原始音频和目标文本,其中,目标文本用于确定待编辑至原始音频的文本内容;采用对原始音频中的待编辑部分音频进行语音掩码的方式得到第一掩码后音频;进一步对目标文本和第一掩码后音频进行语音编辑,得到目标音频,达到了通过对待执行语音编辑的原始音频先进行语音掩码再进行语音编辑得到目标音频的目的,从而实现了提高语音编辑结果的流畅度和真实感的技术效果,进而解决了相关技术中提供的语音编辑方法其训练和测试不匹配导致语音编辑结果的流畅度低、真实感差的技术问题。
通过以上的实施方式的描述,本领域的技术人员易于理解,这里描述的示例实施方式可以通过软件实现,也可以通过软件结合必要的硬件的方式来实现。因此,根据本公开实施方式的技术方案可以以软件产品的形式体现出来,该软件产品可以存储在一个计算机可读存储介质(可以是CD-ROM,U盘,移动硬盘等)中或网络上,包括若干指令以使得一台计算设备(可以是个人计算机、服务器、终端装置、或者网络设备等)执行根据本公开实施方式的方法。
在本公开的示例性实施例中,计算机可读存储介质上存储有能够实现本实施例上述方法的程序产品。在一些可能的实施方式中,本公开实施例的各个方面还可以实现为一种程序产品的形式,其包括程序代码,当所述程序产品在终端设备上运行时,所述程序代码用于使所述终端设备执行本实施例上述“示例性方法”部分中描述的根据本公开各种示例性实施方式的步骤。
根据本公开的实施方式的用于实现上述方法的程序产品,其可以采用便携式紧凑盘只读存储器(CD-ROM)并包括程序代码,并可以在终端设备,例如个人电脑上运行。然而,本公开实施例的程序产品不限于此,在本公开实施例中,计算机可读存储介质可以是任何包含或存储程序的有形介质,该程序可以被指令执行系统、装置或者器件使用或者与其结合使用。
上述程序产品可以采用一个或多个计算机可读介质的任意组合。该计算机可读存储介质例如可以为但不限于电、磁、光、电磁、红外线、或半导体的系统、装置或器件,或者任意以上的组合。计算机可读存储介质的更具体的例子(非穷举的列举)包括:具有一个或多个导线的电连接、便携式盘、硬盘、随机存取存储器(RAM)、只读存储器(ROM)、可擦式可编程只读存储器(EPROM或闪存)、光纤、便携式紧凑盘只读存储器(CD-ROM)、光存储器件、磁存储器件、或者上述的任意合适的组合。
需要说明的是,计算机可读存储介质上包含的程序代码可以用任何适当的介质传输,包括但不限于无线、有线、光缆、RF等等,或者上述的任意合适的组合。
本公开的实施例还提供了一种电子装置,包括存储器和处理器,该存储器中存储有计算机程序,该处理器被设置为运行计算机程序以执行上述任一项方法实施例中的步骤。
可选地,上述电子装置还可以包括传输设备以及输入输出设备,其中,该传输设备和上述处理器连接,该输入输出设备和上述处理器连接。
可选地,在本实施例中,上述处理器可以被设置为通过计算机程序执行以下步骤:
S1,获取待处理的原始音频和目标文本,其中,目标文本用于确定待编辑至原始音频的文本内容;
S2,对原始音频中的待编辑部分音频进行语音掩码,得到第一掩码后音频;
S3,对目标文本和第一掩码后音频进行语音编辑,得到目标音频。
可选地,上述处理器还可以被设置为通过计算机程序执行以下步骤:获取待编辑部分音频在原始音频中的位置信息;基于位置信息对原始音频中的待编辑部分音频进行语音掩码,得到第一掩码后音频。
可选地,上述处理器还可以被设置为通过计算机程序执行以下步骤:对目标文本进行语音转换,得到中间音频;对中间音频和第一掩码后音频进行语音拼接,得到目标音频。
可选地,上述处理器还可以被设置为通过计算机程序执行以下步骤:对目标文本和第一掩码后音频进行语音编辑,得到目标声学特征,其中,目标声学特征用于确定目标文本对应的音频段;对目标声学特征进行声码转换,得到中间音频。
可选地,上述处理器还可以被设置为通过计算机程序执行以下步骤:对目标文本进行字素到音素转换,得到音素序列;使用目标语音编辑模型对音素序列和第一掩码后音频进行语音编辑,得到目标声学特征,其中,目标语音编辑模型采用多组数据通过深度学习训练得到,多组数据包括:训练音频和训练文本,训练文本为训练音频中的待编辑部分音频对应的文本。
可选地,上述处理器还可以被设置为通过计算机程序执行以下步骤:对训练音频中的待编辑部分音频进行语音掩码,得到第二掩码后音频;使用初始语音编辑模型对第二掩码后音频和训练文本进行语音编辑,得到预测声学特征;通过预测声学特征与训练文本对应的真实声学特征确定目标损失;利用目标损失对初始语音编辑模型的参数进行更新,得到目标语音编辑模型。
可选地,上述处理器还可以被设置为通过计算机程序执行以下步骤:使用编码器对音素序列进行文本特征空间编码,得到文本特征;使用特征调节器对文本特征和第一掩码后音频进行特征调节,得到第一听觉感知特征,其中,第一听觉感知特征为目标文本对应的听觉感知特征;使用解码器对第一听觉感知特征进行声学解码,得到目标声学特征。
可选地,上述处理器还可以被设置为通过计算机程序执行以下步骤:从第一掩码后音频中提取第二听觉感知特征,其中,第二听觉感知特征为原始音频中与待编辑部分音频关联的上下文音频对应的听觉感知特征;使用特征调节器对文本特征和第二听觉感知特征进行特征调节,得到第一听觉感知特征。
可选地,上述处理器还可以被设置为通过计算机程序执行以下步骤:第一听觉感知特征包括以下至少之一:目标文本对应的音高;目标文本对应的能量;目标文本对应的时长。
可选地,上述处理器还可以被设置为通过计算机程序执行以下步骤:对游戏应用中待使用的原始游戏语音中待编辑部分语音进行语音掩码,得到掩码后游戏语音;对目标文本和第一掩码后音频进行语音编辑,得到目标音频包括:对游戏应用中待使用的游戏内容文本和掩码后游戏语音进行语音编辑,得到目标游戏语音。
可选地,上述处理器还可以被设置为通过计算机程序执行以下步骤:对配音应用中待使用的原始多媒体配音中待编辑部分配音进行语音掩码,得到掩码后多媒体配音;对目标文本和第一掩码后音频进行语音编辑,得到目标音频包括:对配音应用中待使用的配音内容文本和掩码后多媒体配音进行语音编辑,得到目标多媒体配音。
可选地,上述处理器还可以被设置为通过计算机程序执行以下步骤:接收来自于客户端的待处理的原始音频和目标文本,其中,目标文本用于确定待编辑至原始音频的文本内容;对原始音频中的待编辑部分音频进行语音掩码以得到第一掩码后音频,以及对目标文本和第一掩码后音频进行语音编辑以得到目标音频;将目标音频反馈至客户端。
可选地,上述处理器还可以被设置为通过计算机程序执行以下步骤:响应对语音编辑控件执行的触发操作,弹出语音编辑界面;响应对语音编辑界面执行的输入操作,导入原始音频和目标文本,其中,目标文本用于确定待编辑至原始音频的文本内容;响应对语音
编辑界面执行的编辑操作,从原始音频中选定待编辑部分音频;响应对语音编辑界面执行的播放操作,在游戏场景中播放目标音频,其中,目标音频通过对目标文本和掩码后音频进行语音编辑后得到,掩码后音频通过对待编辑部分音频进行语音掩码后得到。
可选地,上述处理器还可以被设置为通过计算机程序执行以下步骤:获取待处理的训练音频和训练文本,其中,训练文本用于确定待编辑至训练音频的文本内容;对训练音频中的待编辑部分音频进行语音掩码,得到掩码后训练音频;采用掩码后训练音频和训练文本对初始语音编辑模型进行训练,得到目标语音编辑模型,其中,目标语音编辑模型用于对目标文本和掩码后原始音频进行语音编辑以得到目标音频,掩码后原始音频通过对原始音频中的待编辑部分音频进行语音掩码后得到。
在上述实施例的电子装置中,提供了一种实现语音编辑方法的技术方案。通过获取待处理的原始音频和目标文本,其中,目标文本用于确定待编辑至原始音频的文本内容;采用对原始音频中的待编辑部分音频进行语音掩码的方式得到第一掩码后音频;进一步对目标文本和第一掩码后音频进行语音编辑,得到目标音频,达到了通过对待执行语音编辑的原始音频先进行语音掩码再进行语音编辑得到目标音频的目的,从而实现了提高语音编辑结果的流畅度和真实感的技术效果,进而解决了相关技术中提供的语音编辑方法其训练和测试不匹配导致语音编辑结果的流畅度低、真实感差的技术问题。
图11是根据本公开其中一实施例的一种电子装置的示意图。如图11所示,电子装置1100仅仅是一个示例,不应对本公开实施例的功能和使用范围带来任何限制。
如图11所示,电子装置1100以通用计算设备的形式表现。电子装置1100的组件可以包括但不限于:上述至少一个处理器1110、上述至少一个存储器1120、连接不同系统组件(包括存储器1120和处理器1110)的总线1130和显示器1140。
其中,上述存储器1120存储有程序代码,程序代码可以被处理器1110执行,使得处理器1110执行本公开实施例的上述方法部分中描述的根据本公开各种示例性实施方式的步骤。
存储器1120可以包括易失性存储单元形式的可读介质,例如随机存取存储单元(RAM)11201和/或高速缓存存储单元11202,还可以进一步包括只读存储单元(ROM)11203,还可包括非易失性存储器,如一个或者多个磁性存储装置、闪存、或者其他非易失性固态存储器。
在一些实例中,存储器1120还可以包括具有一组(至少一个)程序模块11205的程序/实用工具11204,这样的程序模块11205包括但不限于:操作系统、一个或者多个应用
程序、其它程序模块以及程序数据,这些示例中的每一个或某种组合中可能包括网络环境的实现。存储器1120可进一步包括相对于处理器1110远程设置的存储器,这些远程存储器可以通过网络连接至电子装置1100。上述网络的实例包括但不限于互联网、企业内部网、局域网、移动通信网及其组合。
总线1130可以为表示几类总线结构中的一种或多种,包括存储单元总线或者存储单元控制器、外围总线、图形加速端口、处理器1110或者使用多种总线结构中的任意总线结构的局域总线。
显示器1140可以例如触摸屏式的液晶显示器(Liquid Crystal Display,LCD),该液晶显示器可使得用户能够与电子装置1100的用户界面进行交互。
可选地,电子装置1100也可以与一个或多个外部设备1200(例如键盘、指向设备、蓝牙设备等)通信,还可与一个或者多个使得用户能与该电子装置1100交互的设备通信,和/或与使得该电子装置1100能与一个或多个其它计算设备进行通信的任何设备(例如路由器、调制解调器等等)通信。这种通信可以通过输入/输出(I/O)接口1150进行。并且,电子装置1100还可以通过网络适配器1160与一个或者多个网络(例如局域网(Local Area Network,LAN),广域网(Wide Area Network,WAN)和/或公共网络,例如因特网)通信。如图11所示,网络适配器1160通过总线1130与电子装置1100的其它模块通信。应当明白,尽管图11中未示出,可以结合电子装置1100使用其它硬件和/或软件模块,可以包括但不限于:微代码、设备驱动器、冗余处理单元、外部磁盘驱动阵列、磁盘阵列(Redundant Arrays of Independent Disks,RAID)系统、磁带驱动器以及数据备份存储系统等。
上述电子装置1100还可以包括:键盘、光标控制设备(如鼠标)、输入/输出接口(I/O接口)、网络接口、电源和/或相机。
本领域普通技术人员可以理解,图11所示的结构仅为示意,其并不对上述电子装置的结构造成限定。例如,电子装置1100还可包括比图11中所示更多或者更少的组件,或者具有与图11所示不同的配置。存储器1120可用于存储计算机程序及对应的数据,如本公开实施例中的语音编辑方法对应的计算机程序及对应的数据。处理器1110通过运行存储在存储器1120内的计算机程序,从而执行各种功能应用以及数据处理,即实现上述的语音编辑方法。
上述本公开实施例序号仅仅为了描述,不代表实施例的优劣。
在本公开的上述实施例中,对各个实施例的描述都各有侧重,某个实施例中没有详述
的部分,可以参见其他实施例的相关描述。
在本公开所提供的几个实施例中,应该理解到,所揭露的技术内容,可通过其它的方式实现。其中,以上所描述的装置实施例仅仅是示意性的,例如所述单元的划分,可以为一种逻辑功能划分,实际实现时可以有另外的划分方式,例如多个单元或组件可以结合或者可以集成到另一个系统,或一些特征可以忽略,或不执行。另一点,所显示或讨论的相互之间的耦合或直接耦合或通信连接可以是通过一些接口,单元或模块的间接耦合或通信连接,可以是电性或其它的形式。
所述作为分离部件说明的单元可以是或者也可以不是物理上分开的,作为单元显示的部件可以是或者也可以不是物理单元,即可以位于一个地方,或者也可以分布到多个单元上。可以根据实际的需要选择其中的部分或者全部单元来实现本实施例方案的目的。
另外,在本公开各个实施例中的各功能单元可以集成在一个处理单元中,也可以是各个单元单独物理存在,也可以两个或两个以上单元集成在一个单元中。上述集成的单元既可以采用硬件的形式实现,也可以采用软件功能单元的形式实现。
所述集成的单元如果以软件功能单元的形式实现并作为独立的产品销售或使用时,可以存储在一个计算机可读取存储介质中。基于这样的理解,本公开的技术方案本质上或者说对相关技术做出贡献的部分或者该技术方案的全部或部分可以以软件产品的形式体现出来,该计算机软件产品存储在一个存储介质中,包括若干指令用以使得一台计算机设备(可为个人计算机、服务器或者网络设备等)执行本公开各个实施例所述方法的全部或部分步骤。而前述的存储介质包括:U盘、只读存储器(ROM)、随机存取存储器(RAM)、移动硬盘、磁碟或者光盘等各种可以存储程序代码的介质。
以上所述仅是本公开的优选实施方式,应当指出,对于本技术领域的普通技术人员来说,在不脱离本公开原理的前提下,还可以做出若干改进和润饰,这些改进和润饰也应视为本公开的保护范围。
Claims (16)
- 一种语音编辑方法,包括:获取待处理的原始音频和目标文本,其中,所述目标文本用于确定待编辑至所述原始音频的文本内容;对所述原始音频中的待编辑部分音频进行语音掩码,得到第一掩码后音频;对所述目标文本和所述第一掩码后音频进行语音编辑,得到目标音频。
- 根据权利要求1所述的语音编辑方法,其中,对所述原始音频中的所述待编辑部分音频进行语音掩码,得到所述第一掩码后音频包括:获取所述待编辑部分音频在所述原始音频中的位置信息;基于所述位置信息对所述原始音频中的所述待编辑部分音频进行语音掩码,得到所述第一掩码后音频。
- 根据权利要求1所述的语音编辑方法,其中,对所述目标文本和所述第一掩码后音频进行语音编辑,得到所述目标音频包括:对所述目标文本进行语音转换,得到中间音频;对所述中间音频和所述第一掩码后音频进行语音拼接,得到所述目标音频。
- 根据权利要求3所述的语音编辑方法,其中,对所述目标文本进行语音转换,得到所述中间音频包括:对所述目标文本和所述第一掩码后音频进行语音编辑,得到目标声学特征,其中,所述目标声学特征用于确定所述目标文本对应的音频段;对所述目标声学特征进行声码转换,得到所述中间音频。
- 根据权利要求4所述的语音编辑方法,其中,对所述目标文本和所述第一掩码后音频进行语音编辑,得到所述目标声学特征包括:对所述目标文本进行字素到音素转换,得到音素序列;使用目标语音编辑模型对所述音素序列和所述第一掩码后音频进行语音编辑,得到所述目标声学特征,其中,所述目标语音编辑模型采用多组数据通过深度学习训练得到,所述多组数据包括:训练音频和训练文本,所述训练文本为所述训练音频中的待编辑部分音频对应的文本。
- 根据权利要求5所述的语音编辑方法,其中,所述语音编辑方法还包括:对所述训练音频中的待编辑部分音频进行语音掩码,得到第二掩码后音频;使用初始语音编辑模型对所述第二掩码后音频和所述训练文本进行语音编辑,得到 预测声学特征;通过所述预测声学特征与所述训练文本对应的真实声学特征确定目标损失;利用所述目标损失对所述初始语音编辑模型的参数进行更新,得到所述目标语音编辑模型。
- 根据权利要求5所述的语音编辑方法,其中,所述目标语音编辑模型包括:编码器、特征调节器和解码器,使用所述目标语音编辑模型对所述音素序列和所述第一掩码后音频进行语音编辑,得到所述目标声学特征包括:使用所述编码器对所述音素序列进行文本特征空间编码,得到文本特征;使用所述特征调节器对所述文本特征和所述第一掩码后音频进行特征调节,得到第一听觉感知特征,其中,所述第一听觉感知特征为所述目标文本对应的听觉感知特征;使用所述解码器对所述第一听觉感知特征进行声学解码,得到所述目标声学特征。
- 根据权利要求7所述的语音编辑方法,其中,使用所述特征调节器对所述文本特征和所述第一掩码后音频进行特征调节,得到所述第一听觉感知特征包括:从所述第一掩码后音频中提取第二听觉感知特征,其中,所述第二听觉感知特征为所述原始音频中与所述待编辑部分音频关联的上下文音频对应的听觉感知特征;使用所述特征调节器对所述文本特征和所述第二听觉感知特征进行特征调节,得到所述第一听觉感知特征。
- 根据权利要求7所述的语音编辑方法,其中,所述第一听觉感知特征包括以下至少之一:所述目标文本对应的音高;所述目标文本对应的能量;所述目标文本对应的时长。
- 根据权利要求1所述的语音编辑方法,其中,对所述原始音频中的所述待编辑部分音频进行语音掩码,得到所述第一掩码后音频包括:对游戏应用中待使用的原始游戏语音中待编辑部分语音进行语音掩码,得到掩码后游戏语音;对所述目标文本和所述第一掩码后音频进行语音编辑,得到所述目标音频包括:对所述游戏应用中待使用的游戏内容文本和所述掩码后游戏语音进行语音编辑,得到目标游戏语音。
- 根据权利要求1所述的语音编辑方法,其中,对所述原始音频中的所述待编辑 部分音频进行语音掩码,得到所述第一掩码后音频包括:对配音应用中待使用的原始多媒体配音中待编辑部分配音进行语音掩码,得到掩码后多媒体配音;对所述目标文本和所述第一掩码后音频进行语音编辑,得到所述目标音频包括:对所述配音应用中待使用的配音内容文本和所述掩码后多媒体配音进行语音编辑,得到目标多媒体配音。
- 一种语音编辑方法,通过终端设备提供一图形用户界面,所述图形用户界面所显示的内容包括一语音编辑控件,所述语音编辑方法包括:响应对所述语音编辑控件执行的触发操作,弹出语音编辑界面;响应对所述语音编辑界面执行的输入操作,导入原始音频和目标文本,其中,所述目标文本用于确定待编辑至所述原始音频的文本内容;响应对所述语音编辑界面执行的编辑操作,从所述原始音频中选定待编辑部分音频;响应对所述语音编辑界面执行的播放操作,在游戏场景中播放目标音频,其中,所述目标音频通过对所述目标文本和掩码后音频进行语音编辑后得到,所述掩码后音频通过对所述待编辑部分音频进行语音掩码后得到。
- 一种模型训练方法,包括:获取待处理的训练音频和训练文本,其中,所述训练文本用于确定待编辑至所述训练音频的文本内容;对所述训练音频中的待编辑部分音频进行语音掩码,得到掩码后训练音频;采用所述掩码后训练音频和所述训练文本对初始语音编辑模型进行训练,得到目标语音编辑模型,其中,所述目标语音编辑模型用于对目标文本和掩码后原始音频进行语音编辑以得到目标音频,所述掩码后原始音频通过对原始音频中的待编辑部分音频进行语音掩码后得到。
- 一种语音编辑装置,包括:获取模块,用于获取待处理的原始音频和目标文本,其中,所述目标文本用于确定待编辑至所述原始音频的文本内容;掩码模块,用于对所述原始音频中的待编辑部分音频进行语音掩码,得到第一掩码后音频;编辑模块,用于对所述目标文本和所述第一掩码后音频进行语音编辑,得到目标音频。
- 一种计算机可读存储介质,所述计算机可读存储介质中存储有计算机程序,其中,所述计算机程序被设置为被处理器运行时执行权利要求1至12任一项中所述的语音编辑方法或权利要求13中所述的模型训练方法。
- 一种电子装置,包括存储器和处理器,所述存储器中存储有计算机程序,所述处理器被设置为运行所述计算机程序以执行权利要求1至12任一项中所述的语音编辑方法或权利要求13中所述的模型训练方法。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202310299825.7A CN116434731A (zh) | 2023-03-20 | 2023-03-20 | 语音编辑方法、装置、存储介质及电子装置 |
| CN202310299825.7 | 2023-03-20 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2024193227A1 true WO2024193227A1 (zh) | 2024-09-26 |
Family
ID=87084710
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2024/074070 Ceased WO2024193227A1 (zh) | 2023-03-20 | 2024-01-25 | 语音编辑方法、装置、存储介质及电子装置 |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN116434731A (zh) |
| WO (1) | WO2024193227A1 (zh) |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN120148518A (zh) * | 2025-05-14 | 2025-06-13 | 粤港澳大湾区数字经济研究院(国际先进技术应用推进中心(深圳)) | 一种音频编辑方法、系统、终端及存储介质 |
Families Citing this family (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN116434731A (zh) * | 2023-03-20 | 2023-07-14 | 网易(杭州)网络有限公司 | 语音编辑方法、装置、存储介质及电子装置 |
| CN120164482A (zh) * | 2023-12-14 | 2025-06-17 | 脸萌有限公司 | 编辑音频的方法、装置、计算设备和介质 |
| CN121640960A (zh) * | 2024-09-06 | 2026-03-10 | 北京字跳网络技术有限公司 | 用于音乐编辑的方法、装置、设备和存储介质 |
Citations (7)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US6604078B1 (en) * | 1999-08-23 | 2003-08-05 | Nec Corporation | Voice edit device and mechanically readable recording medium in which program is recorded |
| CN106971749A (zh) * | 2017-03-30 | 2017-07-21 | 联想(北京)有限公司 | 音频处理方法及电子设备 |
| CN111681641A (zh) * | 2020-05-26 | 2020-09-18 | 微软技术许可有限责任公司 | 基于短语的端对端文本到语音(tts)合成 |
| JP2020154057A (ja) * | 2019-03-19 | 2020-09-24 | 株式会社モアソンジャパン | 音声データのテキスト編集装置及び音声データのテキスト編集方法 |
| CN113421547A (zh) * | 2021-06-03 | 2021-09-21 | 华为技术有限公司 | 一种语音处理方法及相关设备 |
| CN113724686A (zh) * | 2021-11-03 | 2021-11-30 | 中国科学院自动化研究所 | 编辑音频的方法、装置、电子设备及存储介质 |
| CN116434731A (zh) * | 2023-03-20 | 2023-07-14 | 网易(杭州)网络有限公司 | 语音编辑方法、装置、存储介质及电子装置 |
Family Cites Families (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN112687259B (zh) * | 2021-03-11 | 2021-06-18 | 腾讯科技(深圳)有限公司 | 一种语音合成方法、装置以及可读存储介质 |
| CN113112995B (zh) * | 2021-05-28 | 2022-08-05 | 思必驰科技股份有限公司 | 词声学特征系统、词声学特征系统的训练方法及系统 |
| CN114187891B (zh) * | 2022-01-14 | 2025-09-02 | 百果园技术(新加坡)有限公司 | 一种语音合成模型的训练、语音合成方法及相关装置 |
| CN115620699B (zh) * | 2022-12-19 | 2023-03-31 | 深圳元象信息科技有限公司 | 语音合成方法、语音合成系统、语音合成设备及存储介质 |
-
2023
- 2023-03-20 CN CN202310299825.7A patent/CN116434731A/zh active Pending
-
2024
- 2024-01-25 WO PCT/CN2024/074070 patent/WO2024193227A1/zh not_active Ceased
Patent Citations (7)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US6604078B1 (en) * | 1999-08-23 | 2003-08-05 | Nec Corporation | Voice edit device and mechanically readable recording medium in which program is recorded |
| CN106971749A (zh) * | 2017-03-30 | 2017-07-21 | 联想(北京)有限公司 | 音频处理方法及电子设备 |
| JP2020154057A (ja) * | 2019-03-19 | 2020-09-24 | 株式会社モアソンジャパン | 音声データのテキスト編集装置及び音声データのテキスト編集方法 |
| CN111681641A (zh) * | 2020-05-26 | 2020-09-18 | 微软技术许可有限责任公司 | 基于短语的端对端文本到语音(tts)合成 |
| CN113421547A (zh) * | 2021-06-03 | 2021-09-21 | 华为技术有限公司 | 一种语音处理方法及相关设备 |
| CN113724686A (zh) * | 2021-11-03 | 2021-11-30 | 中国科学院自动化研究所 | 编辑音频的方法、装置、电子设备及存储介质 |
| CN116434731A (zh) * | 2023-03-20 | 2023-07-14 | 网易(杭州)网络有限公司 | 语音编辑方法、装置、存储介质及电子装置 |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN120148518A (zh) * | 2025-05-14 | 2025-06-13 | 粤港澳大湾区数字经济研究院(国际先进技术应用推进中心(深圳)) | 一种音频编辑方法、系统、终端及存储介质 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN116434731A (zh) | 2023-07-14 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| CN112309365B (zh) | 语音合成模型的训练方法、装置、存储介质以及电子设备 | |
| CN111653265B (zh) | 语音合成方法、装置、存储介质和电子设备 | |
| CN116434731A (zh) | 语音编辑方法、装置、存储介质及电子装置 | |
| CN115966196B (zh) | 基于文本的语音编辑方法、系统、电子设备和存储介质 | |
| CN108711423A (zh) | 智能语音交互实现方法、装置、计算机设备及存储介质 | |
| WO2024066920A1 (zh) | 虚拟场景的对话方法、装置、电子设备、计算机程序产品及计算机存储介质 | |
| CN114783410A (zh) | 语音合成方法、系统、电子设备和存储介质 | |
| KR20070020252A (ko) | 메시지를 수정하기 위한 방법 및 시스템 | |
| CN116188634A (zh) | 人脸图像预测方法、模型及装置、设备、介质 | |
| WO2021169825A1 (zh) | 语音合成方法、装置、设备和存储介质 | |
| WO2024174787A9 (zh) | 语音编辑方法、装置及相关设备 | |
| CN116072095A (zh) | 角色互动方法、装置、电子设备和存储介质 | |
| WO2025179898A1 (zh) | 语音编辑方法及装置 | |
| CN118609572A (zh) | 一种传译方法、装置、设备及其存储介质 | |
| CN109460548B (zh) | 一种面向智能机器人的故事数据处理方法及系统 | |
| CN118553229A (zh) | 语音合成方法、装置、设备、介质及程序产品 | |
| CN118762712A (zh) | 剧场音频作品的生成方法、装置、设备、介质和程序产品 | |
| CN113761268A (zh) | 音频节目内容的播放控制方法、装置、设备和存储介质 | |
| US20230410787A1 (en) | Speech processing system with encoder-decoder model and corresponding methods for synthesizing speech containing desired speaker identity and emotional style | |
| CN115938342A (zh) | 语音处理方法、装置、电子设备及存储介质 | |
| CN113223513A (zh) | 语音转换方法、装置、设备和存储介质 | |
| CN118433437A (zh) | 直播间语音直播方法、装置、直播系统、电子设备及介质 | |
| CN117409762A (zh) | 一种语音编辑及优化方法、装置、设备及存储介质 | |
| CN114694629A (zh) | 用于语音合成的语音数据扩增方法及系统 | |
| CN114242036A (zh) | 角色配音方法、装置、存储介质及电子设备 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 24773796 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 24773796 Country of ref document: EP Kind code of ref document: A1 |