WO2021189981A1 - 语音噪声的处理方法、装置、计算机设备及存储介质 - Google Patents
语音噪声的处理方法、装置、计算机设备及存储介质 Download PDFInfo
- Publication number
- WO2021189981A1 WO2021189981A1 PCT/CN2020/136367 CN2020136367W WO2021189981A1 WO 2021189981 A1 WO2021189981 A1 WO 2021189981A1 CN 2020136367 W CN2020136367 W CN 2020136367W WO 2021189981 A1 WO2021189981 A1 WO 2021189981A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- noise
- voice
- speech
- classification model
- sequences
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L21/00—Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
- G10L21/02—Speech enhancement, e.g. noise reduction or echo cancellation
- G10L21/0208—Noise filtering
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L21/00—Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
- G10L21/02—Speech enhancement, e.g. noise reduction or echo cancellation
- G10L21/0208—Noise filtering
- G10L21/0216—Noise filtering characterised by the method used for estimating noise
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L25/00—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
- G10L25/27—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the analysis technique
- G10L25/30—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the analysis technique using neural networks
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L25/00—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
- G10L25/45—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the type of analysis window
-
- Y—GENERAL TAGGING OF NEW TECHNOLOGICAL DEVELOPMENTS; GENERAL TAGGING OF CROSS-SECTIONAL TECHNOLOGIES SPANNING OVER SEVERAL SECTIONS OF THE IPC; TECHNICAL SUBJECTS COVERED BY FORMER USPC CROSS-REFERENCE ART COLLECTIONS [XRACs] AND DIGESTS
- Y02—TECHNOLOGIES OR APPLICATIONS FOR MITIGATION OR ADAPTATION AGAINST CLIMATE CHANGE
- Y02T—CLIMATE CHANGE MITIGATION TECHNOLOGIES RELATED TO TRANSPORTATION
- Y02T90/00—Enabling technologies or technologies with a potential or indirect contribution to GHG emissions mitigation
Definitions
- This application relates to the field of artificial intelligence technology, and in particular to a method, device, computer equipment, and storage medium for processing speech noise.
- the voice noise is usually recognized first, and after the voice noise is recognized, a unified noise reduction processing method is used to process the voice noise.
- this method cannot identify the types of speech noise.
- the types of speech noise in different scenarios are different. If the same noise reduction processing method is used to process the speech noise in different scenarios, The noise reduction effect that can be achieved is limited, that is, the optimal noise reduction effect cannot be achieved in different scenarios.
- This application provides a method, device, computer equipment, and storage medium for processing speech noise, mainly in that it can identify the types of speech noise in different scenarios, and adopt an appropriate noise reduction processing method according to the recognized noise type. Perform processing to achieve the optimal noise reduction processing effect.
- a method for processing speech noise including:
- the voice sequence contains voice noise, use a preset noise classification model to determine the noise category corresponding to the voice noise, wherein the noise classification model is generated from multiple noises.
- the types of speech noise generated by different noise generation models are different when the models are jointly trained;
- an optimal noise reduction processing strategy corresponding to the speech noise is determined, and the optimal noise reduction processing strategy is used to perform noise reduction processing on the speech noise.
- a speech noise processing device including:
- the acquiring unit is used to acquire the voice sequence to be recognized
- the determining unit is configured to perform noise recognition on the voice sequence, and if the voice sequence contains voice noise, use a preset noise classification model to determine the noise category corresponding to the voice noise, wherein the noise classification model is It is obtained by joint training with multiple noise generation models, and the types of speech noise generated by different noise generation models are different;
- the noise reduction unit is configured to determine an optimal noise reduction processing strategy corresponding to the speech noise based on the noise category, and use the optimal noise reduction processing strategy to perform noise reduction processing on the speech noise.
- a computer-readable storage medium on which a computer program is stored, and when the program is executed by a processor, the steps of a method for processing speech noise are realized:
- the voice sequence contains voice noise, use a preset noise classification model to determine the noise category corresponding to the voice noise, wherein the noise classification model is generated from multiple noises.
- the types of speech noise generated by different noise generation models are different when the models are jointly trained;
- an optimal noise reduction processing strategy corresponding to the speech noise is determined, and the optimal noise reduction processing strategy is used to perform noise reduction processing on the speech noise.
- a computer device including a memory, a processor, and a computer program stored in the memory and capable of running on the processor.
- the processor implements a voice noise when the program is executed. The steps of the processing method:
- the voice sequence contains voice noise, use a preset noise classification model to determine the noise category corresponding to the voice noise, wherein the noise classification model is generated from multiple noises.
- the types of speech noise generated by different noise generation models are different when the models are jointly trained;
- an optimal noise reduction processing strategy corresponding to the speech noise is determined, and the optimal noise reduction processing strategy is used to perform noise reduction processing on the speech noise.
- the speech noise processing method, device, computer equipment, and storage medium provided in this application are compared with the current method of using the same noise reduction strategy for noise reduction processing for different types of speech noise.
- This application can obtain the to-be-identified The voice sequence; and noise recognition is performed on the voice sequence.
- a preset noise classification model is used to determine the noise category corresponding to the voice noise, wherein the noise classification model is It is obtained by joint training with multiple noise generation models, and the types of speech noise generated by different noise generation models are different; at the same time, based on the noise category, determine the optimal noise reduction processing strategy corresponding to the speech noise, and use
- the optimal noise reduction processing strategy performs noise reduction processing on the speech noise, so that the noise classification model and multiple noise generation models are jointly trained, so that the noise classification model in this application can be used for speech in different scenarios.
- the type of noise is identified, and then the optimal noise reduction processing strategy can be selected to process the speech noise according to the determined noise category, and the optimal noise reduction processing effect can be achieved.
- Fig. 1 shows a flowchart of a method for processing speech noise provided by an embodiment of the present application
- FIG. 2 shows a flowchart of another method for processing voice noise according to an embodiment of the present application
- FIG. 3 shows a schematic structural diagram of a speech noise processing apparatus provided by an embodiment of the present application
- FIG. 4 shows a schematic structural diagram of another apparatus for processing speech noise according to an embodiment of the present application
- Fig. 5 shows a schematic diagram of the physical structure of a computer device provided by an embodiment of the present application.
- the voice noise is usually recognized first, and after the voice noise is recognized, a unified noise reduction processing method is used to process the voice noise.
- this method cannot identify the types of speech noise.
- the types of speech noise in different scenarios are different. If the same noise reduction processing method is used to process the speech noise in different scenarios, the reduction that can be achieved can be achieved.
- the noise effect is limited, that is, the optimal noise reduction effect cannot be achieved in different scenes.
- an embodiment of the present application provides a method for processing speech noise. As shown in FIG. 1, the method includes:
- the voice sequence to be recognized is a user voice sequence obtained from a certain scene.
- the voice sequence to be recognized is a user voice sequence collected on the side of a street, or a user voice sequence collected from a factory.
- the voice sequence may or may not contain voice noise.
- an appropriate noise reduction strategy can be selected according to the type of speech noise to process the speech noise in order to achieve the optimal
- the embodiments of the present application are mainly applicable to the processing of speech noise.
- the execution subject of the embodiments of the present application is a device or device capable of processing speech noise, which can be set on the client or server side.
- the preprocessed speech sequence is obtained, and the preprocessed speech sequence is used as the speech sequence to be recognized, so as to determine whether the speech sequence to be recognized contains speech noise, if the speech sequence to be recognized does not contain speech If there is noise, the speech sequence to be recognized is directly recognized; if the speech sequence to be recognized contains speech noise, it is necessary to further determine the type of noise required, so as to select the appropriate noise reduction process according to the determined type of speech noise Strategies for noise reduction processing, so as to achieve the best noise reduction effect.
- the noise classification model is obtained through joint training with multiple noise generation models.
- the types of speech noise generated by different noise generation models are different.
- the types of speech noise in different scenarios are different, for example, collected on the side of the street.
- the type of speech noise is different from the type of speech noise collected in the factory.
- the speech sequence to be recognized is input into a preset noise recognition model for noise recognition
- the preset noise recognition model may specifically be a first preset neural network model.
- the hidden layer in the first preset neural network model extracts the to-be-recognized According to the voice features corresponding to the voice sequence, it is determined whether the voice sequence to be recognized contains voice noise according to the extracted voice features. If the voice sequence to be recognized does not contain voice noise, then the extracted voice feature is directly subjected to voice recognition; When the recognized speech sequence contains speech noise, the extracted speech features are input into a preset noise classification model for noise classification.
- the noise classification model may be a second preset neural network model.
- the hidden layer in the second preset neural network model extracts the noise features corresponding to the voice noise, and then determines the noise type corresponding to the voice noise contained in the speech sequence to be recognized according to the extracted noise feature, so as to select according to the determined noise type
- a suitable noise reduction processing strategy performs noise reduction processing on the speech sequence to be recognized to achieve the optimal noise reduction effect in the scene.
- an optimal noise reduction processing strategy corresponding to the speech noise is determined, and the optimal noise reduction processing strategy is used to perform noise reduction processing on the speech noise.
- different types of speech noise are suitable for different optimal noise reduction processing strategies.
- speech noise from the side of the street because the noise on the side of the street is relatively random and the noise has a wide spectrum range, it can be used.
- Adaptive filter for noise reduction for the speech noise from the factory, since most of the speech noise in the factory is machine processing noise in the workshop, the randomness of the noise is small, and the noise spectrum range is narrow, so adaptive trapping can be used.
- the wave generator performs noise reduction processing.
- the noise reduction processing strategy corresponding to the noise category from the preset noise reduction strategy library, and determine it as the optimal reduction Noise processing strategy, and then use the optimal noise reduction processing strategy to reduce the noise in the speech noise in the speech sequence to be recognized, so that the optimal noise reduction processing effect can be achieved for the speech noise in different scenarios, avoiding the use of uniform Noise reduction processing strategy, noise reduction processing effect of image speech noise.
- the method for processing speech noise provided by the embodiment of the present application is compared with the current manner in which the same noise reduction strategy is used for noise reduction processing for different types of speech noise, the present application can obtain the voice sequence to be recognized; and Noise recognition is performed on the speech sequence, and if the speech sequence contains speech noise, a preset noise classification model is used to determine the noise category corresponding to the speech noise, wherein the noise classification model is related to multiple noise generation models
- the types of speech noise generated by different noise generation models are different from the joint training; at the same time, based on the noise category, the optimal noise reduction processing strategy corresponding to the speech noise is determined, and the optimal noise reduction is used
- the processing strategy performs noise reduction processing on the speech noise, so that the noise classification model and multiple noise generation models are jointly trained, so that the noise classification model in this application can recognize the types of speech noise in different scenarios. Furthermore, according to the determined noise category, the optimal noise reduction processing strategy can be selected to process the speech noise, and the optimal noise reduction processing effect can be achieved.
- an embodiment of the present application provides another method for processing speech noise. As shown in FIG. 2, the method include:
- multiple random speech sequences can obey Gaussian distribution.
- the real speech sequence is the real speech sequence of the user collected in different scenes.
- the real speech sequence is processed by noise reduction, and there is no noise, and the speech recognition can be directly performed.
- a noise recognition model and a noise classification model are constructed respectively to achieve the purpose of recognizing and classifying speech noise.
- the real voice sequence of the user in the preset sample library is obtained.
- the real voice sequence comes from different scenarios.
- step 201 specifically includes: calculating the Euclidean distance between different real speech sequences according to the preset Euclidean distance algorithm; based on the Euclidean distance, The real speech sequence is clustered to obtain real speech sequences in different clustering categories.
- the voice sequences in the preset sample library are clustered to obtain the real voice sequences under different clustering categories, and the scenes corresponding to the real voice sequences under different clustering categories are determined , And then be able to determine the real voice sequence in different scenarios.
- the Euclidean distance between different real speech sequences is calculated according to the preset Euclidean distance algorithm, and the real speech sequence is clustered according to the calculated Euclidean distance to obtain real speech sequences in different clustering categories, and then extract The voice features corresponding to the real voice sequences under different clustering categories are determined, and the scenes corresponding to the real voice sequences under different clustering categories are determined. For example, it is determined that the real voice sequences 1-10 are the voice sequences collected on the street, and the voice sequences 11- 20 is the voice sequence collected in the factory, which can determine the real voice sequence in different scenarios.
- step 202 specifically includes: constructing an initial noise classification model and multiple initial noise generation models respectively; For real speech sequences in a class category, joint iterative training is performed on the initial noise classification model and the multiple initial noise generation models to construct the noise classification model and the multiple noise generation models. Further, in order to be able to recognize speech noise, it is also necessary to construct a noise recognition model, which separately constructs an initial noise classification model and multiple initial noise generation models, including: separately constructing an initial noise recognition model, an initial noise classification model, and multiple initial noise generation models. The initial noise generation model.
- the noise classification model and the multiple initial noise generation models are jointly iteratively trained according to the multiple random voice sequences and the real voice sequences in the different clustering categories to construct the
- the noise classification model and the multiple noise generation models include: respectively inputting the multiple random speech sequences into the multiple initial noise generation models to generate different types of speech noise; and combining the generated speech noise and the real noise
- the speech sequences are respectively input to the initial noise and noise recognition model for noise recognition, and the initial noise recognition result is obtained; the speech feature corresponding to the speech noise in the initial noise recognition result is extracted, and it is input to the initial noise classification model for noise classification, Obtain the initial noise classification result; based on the initial noise recognition result and the initial noise classification result, construct the noise recognition accuracy loss function and noise classification accuracy loss function respectively; according to the noise recognition accuracy loss function and noise classification accuracy loss Function to perform joint iterative training on the initial noise recognition model, the initial noise classification model, and the multiple initial noise generation models to construct the noise recognition model, the noise classification model, and the multiple noise generation models respectively.
- the initial noise recognition results are obtained, and then the speech features corresponding to the speech noise in the initial recognition results are extracted.
- the noise recognition accuracy loss function and the noise classification accuracy loss function are constructed respectively.
- Lc is the noise recognition accuracy loss function
- Lc is the noise classification accuracy loss Korean style
- zi is speech noise
- xi is the real speech sequence
- D stands for the preset noise recognition model
- G stands for the preset noise generation model
- c stands for Noise classification model, in order to ensure that the speech noise generated by the noise generation model is closer to the real speech sequence, and increase the difficulty of the recognition of the noise recognition model.
- the optimization direction of the noise generation model and the noise recognition model is opposite, that is, the noise generation model needs to be minimized
- the accuracy of the noise recognition model is preset, so its optimization direction is to minimize Lc-Ls, and the training purpose of the noise classification model is to maximize the accuracy of the classification noise, so its optimization direction is to maximize Lc+Ls, so the optimization direction is to maximize Lc+Ls.
- the above two optimization equations can continuously train the initial noise generation model, the initial noise recognition model and the initial noise classification model to construct the noise generation model, the noise recognition model and the noise classification model.
- the voice sequence to be recognized is a user voice sequence obtained from a certain scene.
- the voice sequence may or may not contain voice noise.
- the voice sequence to be recognized contains For speech noise, it is necessary to reduce the noise of speech noise.
- the types of speech noise can be further identified so as to be based on the type of speech noise. Select the appropriate noise reduction processing strategy for the type of noise reduction.
- a preset noise classification model is used to determine the noise category corresponding to the voice noise.
- step 204 specifically includes: performing speech feature extraction on the speech sequence to obtain the speech feature corresponding to the speech sequence; and judging the voice based on the speech feature Whether the sequence contains speech noise; if it contains speech noise, based on the extracted speech features, the noise classification model is used to determine the noise category corresponding to the speech noise.
- the voice sequence to be recognized is input to the noise recognition model for noise recognition.
- the hidden layer in the preset noise recognition model will extract the voice features corresponding to the voice sequence to be recognized, based on the extracted voice
- the feature determines whether the speech sequence to be recognized contains speech noise, and if it contains speech noise, the extracted speech feature is input into the noise classification model for noise classification to determine the noise category corresponding to the speech noise.
- the noise reduction processing strategy corresponding to the noise category from the preset noise reduction strategy library, and determine it as the optimal noise reduction processing strategy, and then use the The optimal noise reduction processing strategy performs noise reduction processing on the speech noise in the speech sequence to be recognized, so that the optimal noise reduction processing effect can be achieved for the speech noise in different scenarios, and the unified noise reduction processing strategy is avoided. Noise reduction processing effect of speech noise.
- Another method for processing speech noise provided by the embodiment of the present application is compared with the current manner in which the same noise reduction strategy is used for noise reduction processing for different types of speech noise, the present application can obtain the voice sequence to be recognized; and Perform noise recognition on the voice sequence. If the voice sequence contains voice noise, use a preset noise classification model to determine the noise category corresponding to the voice noise, wherein the noise classification model is generated from multiple noises. The types of speech noise generated by different noise generation models are different from the joint training of the models; at the same time, based on the noise category, the optimal noise reduction processing strategy corresponding to the speech noise is determined, and the optimal noise reduction is used.
- the noise processing strategy performs noise reduction processing on the speech noise, so that the noise classification model and multiple noise generation models are jointly trained, so that the noise classification model in this application can identify the types of speech noise in different scenarios , And then can select the optimal noise reduction processing strategy to process the speech noise according to the determined noise category, so as to achieve the optimal noise reduction processing effect.
- an embodiment of the present application provides an apparatus for processing speech noise.
- the apparatus includes: an acquisition unit 31, a determination unit 32, and a noise reduction unit 33.
- the acquiring unit 31 may be used to acquire a voice sequence to be recognized.
- the acquiring unit 31 is the main functional module of the device for acquiring the voice sequence to be recognized.
- the determining unit 32 may be configured to perform noise recognition on the voice sequence, and if the voice sequence contains voice noise, use a preset noise classification model to determine the noise category corresponding to the voice noise, wherein the The noise classification model is jointly trained with multiple noise generation models, and the types of speech noise generated by different noise generation models are different.
- the determining unit 32 is a main functional module that performs noise recognition on the voice sequence in the device, and if the voice sequence contains voice noise, it uses a preset noise classification model to determine the main functional module of the noise category corresponding to the voice noise, It is also the core module.
- the noise reduction unit 33 may be configured to determine an optimal noise reduction processing strategy corresponding to the speech noise based on the noise category, and use the optimal noise reduction processing strategy to perform noise reduction processing on the speech noise.
- the noise reduction unit 33 determines the optimal noise reduction processing strategy corresponding to the speech noise based on the noise category in the device, and uses the optimal noise reduction processing strategy to perform noise reduction processing on the speech noise Main functional modules.
- the determination unit 32 includes an extraction module 321, a judgment module 322 and a determination module 323.
- the extraction module 321 may be used to extract voice features of the voice sequence to obtain voice features corresponding to the voice sequence to be recognized.
- the judgment module 322 may be used to judge whether the speech sequence contains speech noise based on the speech feature.
- the determining module 323 may be configured to determine the noise category corresponding to the voice noise by using the noise classification model based on the extracted voice feature if the voice noise is included.
- the device further includes a clustering unit 34 and a construction unit 35.
- the acquiring unit 31 may also be used to acquire a real voice sequence and multiple random voice sequences in a preset voice sample library.
- the clustering unit 34 may be used to perform clustering processing on the real voice sequence to obtain real voice sequences in different clustering categories.
- the construction unit 35 may be configured to construct the noise classification model and the multiple noise generation models according to the multiple random voice sequences and the real voice sequences in the different clustering categories.
- the clustering unit 34 includes: a calculation module 341 and a clustering module 342.
- the calculation module 341 may be used to calculate the Euclidean distance between different real speech sequences according to a preset Euclidean distance algorithm.
- the clustering module 342 may be used to perform clustering processing on the real speech sequence based on the Euclidean distance to obtain real speech sequences in different clustering categories.
- the construction unit 35 includes: a first construction module 351 and a second construction module 352.
- the first construction module 351 may be used to separately construct an initial noise classification model and multiple initial noise generation models.
- the second construction module 352 may be used to combine the initial noise classification model and the multiple initial noise generation models according to the multiple random voice sequences and the real voice sequences in the different clustering categories Iterative training to construct the noise classification model and the multiple noise generation models.
- the second construction module 352 includes: a generation sub-module, an identification sub-module, a classification sub-module, and a construction sub-module.
- the generating sub-module may be used to input the multiple random speech sequences into the multiple initial noise generation models to generate different types of speech noise.
- the recognition sub-module may be used to input the generated speech noise and the real speech sequence into the initial noise and noise recognition model to perform noise recognition, and obtain the initial noise recognition result.
- the classification sub-module can be used to extract the speech features corresponding to the speech noise in the initial noise recognition result, and input it into the initial noise classification model for noise classification, to obtain the initial noise classification result.
- the construction sub-module may be used to construct a noise recognition accuracy loss function and a noise classification accuracy loss function based on the initial noise recognition result and the initial noise classification result.
- the construction sub-module may also be used to combine the initial noise recognition model, the initial noise classification model, and the multiple initial noise generation models according to the noise recognition accuracy loss function and the noise classification accuracy loss function Iterative training to separately construct a noise recognition model, the noise classification model, and the multiple noise generation models.
- an embodiment of the present application also provides a computer-readable storage medium.
- the computer-readable storage medium may be non-volatile or volatile, and stored thereon.
- There is a computer program when the program is executed by the processor, the following steps are realized: obtain the speech sequence to be recognized; obtain the speech sequence to be recognized; A noise classification model is assumed to determine the noise category corresponding to the speech noise, wherein the noise classification model is jointly trained with multiple noise generation models, and the types of speech noise generated by different noise generation models are different;
- the noise category determines the optimal noise reduction processing strategy corresponding to the speech noise, and uses the optimal noise reduction processing strategy to perform noise reduction processing on the speech noise.
- the computer device includes: a processor 41, The memory 42 and a computer program that is stored on the memory 42 and can run on the processor, wherein the memory 42 and the processor 41 are both set on the bus 43, and the processor 41 implements the following steps when the program is executed: The voice sequence; perform noise recognition on the voice sequence, and if the voice sequence contains voice noise, use a preset noise classification model to determine the noise category corresponding to the voice noise, wherein the noise classification model is Multiple noise generation models are jointly trained, and the types of speech noise generated by different noise generation models are different; based on the noise category, the optimal noise reduction processing strategy corresponding to the speech noise is determined, and the optimal noise reduction is used
- the noise processing strategy performs noise reduction processing on the speech noise.
- the present application can obtain the voice sequence to be recognized; perform noise recognition on the voice sequence, and if the voice sequence contains voice noise, use a preset noise classification model to determine the voice noise
- the corresponding noise category wherein the noise classification model is jointly trained with multiple noise generation models, and the types of speech noise generated by different noise generation models are different; at the same time, based on the noise category, the The optimal noise reduction processing strategy corresponding to the speech noise, and the optimal noise reduction processing strategy is used to reduce the noise of the speech noise, so that the noise classification model and multiple noise generation models are jointly trained, so that
- the noise classification model in this application can identify the types of speech noise in different scenarios, and then can select the optimal noise reduction processing strategy to process the speech noise according to the determined noise category, and can achieve the optimal noise reduction processing effect .
Landscapes
- Engineering & Computer Science (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Computational Linguistics (AREA)
- Signal Processing (AREA)
- Health & Medical Sciences (AREA)
- Human Computer Interaction (AREA)
- Physics & Mathematics (AREA)
- Acoustics & Sound (AREA)
- Multimedia (AREA)
- Quality & Reliability (AREA)
- Evolutionary Computation (AREA)
- Artificial Intelligence (AREA)
- Circuit For Audible Band Transducer (AREA)
Abstract
一种语音噪声的处理方法、装置、计算机设备及存储介质,涉及人工智能领域。方法包括:获取待识别的语音序列(101);对语音序列进行噪音识别,若语音序列中包含语音噪声,则利用预设的噪声分类模型确定语音噪声对应的噪声类别,其中,噪声分类模型是与多个噪声生成模型联合训练得到的,不同噪声生成模型所生成的语音噪声的种类不同(102);基于噪声类别,确定语音噪声对应的最优降噪处理策略,并利用最优降噪处理策略对语音噪声进行降噪处理(103)。能够对不同场景下语音噪声的种类进行识别,并根据识别的噪声种类采用适当的降噪处理方式对语音噪声进行处理,以达到最优降噪处理效果。
Description
本申请要求于2020年10月26日提交中国专利局、申请号为202011153509.1,发明名称为“语音噪声的处理方法、装置、计算机设备及存储介质”的中国专利申请的优先权,其全部内容通过引用结合在本申请中。
本申请涉及人工智能技术领域,尤其是涉及一种语音噪声的处理方法、装置、计算机设备及存储介质。
在语音识别技术中,通常需要识别语音序列中的噪声,并对识别的噪声进行降噪处理,以提高后续语音识别的准确率,因此,有效地对语音噪声进行处理十分重要。
目前,在对语音噪声处理的过程中,通常先对语音噪声进行识别,在识别出语音噪声后采用统一的降噪处理方式对语音噪声进行处理。然而,发明人意识到这种方式无法对语音噪声的种类进行识别,不同场景下的语音噪声的种类是不同的,如果均采用相同的降噪处理方式对不同场景下的语音噪声进行处理,所能达到的降噪效果有限,即在不同场景下无法达到最优的降噪效果。
本申请提供了一种语音噪声的处理方法、装置、计算机设备及存储介质,主要在于能够对不同场景下语音噪声的种类进行识别,并根据识别的噪声种类采用适当的降噪处理方式对语音噪声进行处理,以达到最优降噪处理效果。
根据本申请的第一个方面,提供一种语音噪声的处理方法,包括:
获取待识别的语音序列;
对所述语音序列进行噪音识别,若所述语音序列中包含语音噪声,则利用预设的噪声分类模型确定所述语音噪声对应的噪声类别,其中,所述噪声分类模型是与多个噪声生成模型联合训练得到的,不同噪声生成模型所生成的语音噪声的种类不同;
基于所述噪声类别,确定所述语音噪声对应的最优降噪处理策略,并利用所述最优降噪处理策略对所述语音噪声进行降噪处理。
根据本申请的第二个方面,提供一种语音噪声的处理装置,包括:
获取单元,用于获取待识别的语音序列;
确定单元,用于对所述语音序列进行噪声识别,若所述语音序列中包含语音噪声,则利用预设的噪声分类模型确定所述语音噪声对应的噪声类别,其中,所述噪声分类模型是与多个噪声生成模型联合训练得到的,不同噪声生成模型所生成的语音噪声的种类不同;
降噪单元,用于基于所述噪声类别,确定所述语音噪声对应的最优降噪处理策略,并利用所述最优降噪处理策略对所述语音噪声进行降噪处理。
根据本申请的第三个方面,提供一种计算机可读存储介质,其上存储有计算机程序,该程序被处理器执行时实现一种语音噪声的处理方法的步骤:
获取待识别的语音序列;
对所述语音序列进行噪音识别,若所述语音序列中包含语音噪声,则利用预设的噪声分类模型确定所述语音噪声对应的噪声类别,其中,所述噪声分类模型是与多个噪声生成模型联合训练得到的,不同噪声生成模型所生成的语音噪声的种类不同;
基于所述噪声类别,确定所述语音噪声对应的最优降噪处理策略,并利用所述最优降噪处理策略对所述语音噪声进行降噪处理。
根据本申请的第四个方面,提供一种计算机设备,包括存储器、处理器及存储在存储器上并可在处理器上运行的计算机程序,所述处理器执行所述程序时实现一种语音噪声的处理方法的步骤:
获取待识别的语音序列;
对所述语音序列进行噪音识别,若所述语音序列中包含语音噪声,则利用预设的噪声分类模型确定所述语音噪声对应的噪声类别,其中,所述噪声分类模型是与多个噪声生成模型联合训练得到的,不同噪声生成模型所生成的语音噪声的种类不同;
基于所述噪声类别,确定所述语音噪声对应的最优降噪处理策略,并利用所述最优降噪处理策略对所述语音噪声进行降噪处理。
本申请提供的一种语音噪声的处理方法、装置、计算机设备及存储介质,与目前针对不同种类的语音噪声均采用同种降噪策略进行降噪处理的方式相比,本申请能够获取待识别的语音序列;并对所述语音序列进行噪音识别,若所述语音序列中包含语音噪声,则利用预设的噪声分类模型确定所述语音噪声对应的噪声类别,其中,所述噪声分类模型是与多个噪声生成模型联合训练得到的,不同噪声生成模型所生成的语音噪声的种类不同;与此同时,基于所述噪声类别,确定所述语音噪声对应的最优降噪处理策略,并利用所述最优降噪处理策略对所述语音噪声进行降噪处理,由此通过将噪声分类模型与多个噪声生成模型联合进行训练,从而使得本申请中的噪声分类模型能够对不同场景下语音噪声的种类进行识别,进而能够根据确定的噪声类别,选择最优的降噪处理策略对语音噪声进行处理,能够达到最优的降噪处理效果。
此处所说明的附图用来提供对本申请的进一步理解,构成本申请的一部分,本申请的示意性实施例及其说明用于解释本申请,并不构成对本申请的不当限定。在附图中:
图1示出了本申请实施例提供的一种语音噪声的处理方法流程图;
图2示出了本申请实施例提供的另一种语音噪声的处理方法流程图;
图3示出了本申请实施例提供的一种语音噪声的处理装置的结构示意图;
图4示出了本申请实施例提供的另一种语音噪声的处理装置的结构示意图;
图5示出了本申请实施例提供的一种计算机设备的实体结构示意图。
下文中将参考附图并结合实施例来详细说明本申请。需要说明的是,在不冲突的情况下,本申请中的实施例及实施例中的特征可以相互组合。
目前,在对语音噪声处理的过程中,通常先对语音噪声进行识别,在识别出语音噪声后采用统一的降噪处理方式对语音噪声进行处理。然而,这种方式无法对语音噪声的种类进行识别,不同场景下的语音噪声的种类是不同的,如果均采用相同的降噪处理方式对不同场景下的语音噪声进行处理,所能达到的降噪效果有限,即在不同场景下无法达到最优的降噪效果。
为了解决上述问题,本申请实施例提供了一种语音噪声的处理方法,如图1所示,所述方法包括:
101、获取待识别的语音序列。
其中,待识别的语音序列为从某场景下获取的用户语音序列,例如,待识别的语音序列为在街道旁采集的一段用户语音序列,或者从工厂中采集的一段用户语音序列,该待识别的语音序列中可能会包含语音噪声,也可能不包含语音噪声,对于本申请实施例,为了提高用户的语音识别精度,需要判断采集的用户语音序列中是否包含语音噪声,如果包含语音噪声,则需要对用户的语音序列进行降噪处理,以便提高用户的语音识别精度,具体进行降噪处理时,可以根据语音噪声的种类选择合适的降噪处理策略对语音噪声进行处理,以达到最优的降噪效果,本申请实施例主要适用于语音噪声的处理,本申请实施例的执行主体为能够对语音噪声进行处理的装置或者设备,可以设置在客户端或者服务器一侧。
具体地,获取用户在某场景下的一段语音序列,在判断该语音序列中是否包含语音噪声之前,需要对获取的用户语音序列进行预处理,具体包括预加重处理、分帧处理和加窗函数处理,由此得到预处理后的语音序列,并将预处理后的语音序列作为待识别的语音序列,以便判断待识别的语音序列中是否包含语音噪声,如果待识别的语音序列中不包含语音噪声, 则直接对待识别的语音序列进行语音识别;如果待识别的语音序列中包含语音噪声,则需要进一步确定所包含需要噪声的种类,以便根据确定的语音噪声的种类,选择合适的降噪处理策略进行降噪处理,从而达到最优的降噪效果。
102、对所述语音序列进行噪音识别,若所述语音序列中包含语音噪声,则利用预设的噪声分类模型确定所述语音噪声对应的噪声类别。
其中,所述噪声分类模型是与多个噪声生成模型联合训练得到的,不同噪声生成模型所生成的语音噪声的种类不同,此外,不同场景下语音噪声的种类不同,例如,在街道旁采集的语音噪声的种类与工厂中采集的语音噪声的种类不同,对于本申请实施例,为了判断待识别的语音序列中是否包含语音噪声,将待识别的语音序列输入至预设噪声识别模型进行噪声识别,该预设噪声识别模型具体可以为第一预设神经网络模型,在利用第一预设神经网络模型识别语音噪声的过程中,第一预设神经网络模型中的隐藏层会提取待识别的语音序列对应的语音特征,进而根据提取的语音特征判断待识别的语音序列中是否包含语音噪声,如果待识别的语音序列中不包含语音噪声,则直接对提取的语音特征进行语音识别;如果待识别的语音序列中包含语音噪声,则将提取的语音特征输入至预设的噪声分类模型进行噪声分类,所述噪声分类模型具体可以为第二预设神经网络模型,具体进行噪声分类时,利用第二预设神经网络模型中的隐藏层提取语音噪声对应的噪声特征,进而根据提取的噪声特征确定待识别的语音序列中所包含的语音噪声对应的噪声种类,以便根据确定的噪声种类,选择合适的降噪处理策略对待识别的语音序列进行降噪处理,以达到在场景下最优的降噪效果。
基于所述噪声类别,确定所述语音噪声对应的最优降噪处理策略,并利用所述最优降噪处理策略对所述语音噪声进行降噪处理。
其中,不同种类的语音噪声所适用的最优降噪处理策略不同,例如,针对来自街道旁的语音噪声,由于街道旁的噪声随机性比较大,且噪声的频谱范围较宽,因此可以采用自适应滤波器进行降噪;针对来自工厂中的语音噪声,由于工厂中的语音噪声大多是车间的机器加工噪声,噪声的随机性较小,而且噪声的频谱范围较窄,因此可以采用自适应陷波器进行降噪处理,对于本方实施例,根据确定的语音噪声对应的噪声类别,从预设降噪策略库中选择该噪声类别对应的降噪处理策略,并将其确定为最优降噪处理策略,之后利用该最优降噪处理策略对待识别的语音序列中的语音噪声进行降噪处理,从而针对不同场景下的语音噪声,均能够达到最优降噪处理效果,避免采用统一的降噪处理策略,影像语音噪声的降噪处理效果。
本申请实施例提供的一种语音噪声的处理方法,与目前针对不同种类的语音噪声均采用同种降噪策略进行降噪处理的方式相比,本申请能够获取待识别的语音序列;并对所述语音序列进行噪音识别,若所述语音序列中包含语音噪声,则利用预设的噪声分类模型确定所述语音噪声对应的噪声类别,其中,所述噪声分类模型是与多个噪声生成模型联合训练得到的,不同噪声生成模型所生成的语音噪声的种类不同;与此同时,基于所述噪声类别,确定所述语音噪声对应的最优降噪处理策略,并利用所述最优降噪处理策略对所述语音噪声进行降噪处理,由此通过将噪声分类模型与多个噪声生成模型联合进行训练,从而使得本申请中的噪声分类模型能够对不同场景下语音噪声的种类进行识别,进而能够根据确定的噪声类别,选择最优的降噪处理策略对语音噪声进行处理,能够达到最优的降噪处理效果。
进一步的,为了更好的说明上述语音噪声的处理过程,作为对上述实施例的细化和扩展,本申请实施例提供了另一种语音噪声的处理方法,如图2所示,所述方法包括:
201、获取预设语音样本库中的真实语音序列以及多个随机语音序列,并对所述真实语音序列进行聚类处理,得到不同聚类类别下的真实语音序列。
其中,多个随机语音序列可以服从高斯分布,真实语音序列为在不同场景中采集的用户的真实语音序列,该真实语音序列经过降噪处理,不存在噪声,可以直接进行语音识别,在本申请实施例中,希望利用多个随机语音序列和多个噪声生成模型,模拟用户在不同场景下的真实语音序列,由此生成不同场景下的语音噪声,进而依据生成的不同场景下的语音噪声和不同场景下的真实语音序列,分别构建噪声识别模型和噪声分类模型,以达到能够对语音噪声进行识别和分类的目的。
对于本申请实施例,获取预设样本库中用户的真实语音序列,该真实语音序列来自于不同场景,为了利用不同场景下的真实语音序列和随机语音序列,构建噪声识别模型和噪声分类模型,需要先将预设样本库中的真实语音序列进行聚类处理,基于此,步骤201具体包括: 根据预设的欧式距离算法计算不同真实语音序列之间的欧式距离;基于所述欧式距离,对所述真实语音序列进行聚类处理,得到不同聚类类别下的真实语音序列。由于不同场景下的语音序列较为相似,将预设样本库中的语音序列进行聚类处理,得到不同聚类类别下的真实语音序列,并确定不同聚类类别下的真实语音序列所对应的场景,进而能够确定不同场景下的真实语音序列。
具体地,根据预设的欧式距离算法分别计算不同真实语音序列之间的欧式距离,根据计算的欧式距离对真实语音序列进行聚类处理,得到不同聚类类别下的真实语音序列,进而通过提取不同聚类类别下的真实语音序列对应的语音特征,确定不同聚类类别下的真实语音序列对应的场景,例如,确定真实语音序列1-10为在街道旁采集的语音序列,语音序列11-20为在工厂中采集的语音序列,由此能够确定不同场景下的真实语音序列。
202、根据所述多个随机语音序列和所述不同聚类类别下的真实语音序列,构建所述噪声分类模型和所述多个噪声生成模型。
对于本申请实施例,为了构建噪声分类模型和多个噪声生成模型,步骤202具体包括:分别构建初始噪声分类模型和多个初始噪声生成模型;根据所述多个随机语音序列和所述不同聚类类别下的真实语音序列,对所述初始噪声分类模型和所述多个初始噪声生成模型进行联合迭代训练,构建所述噪声分类模型和所述多个噪声生成模型。进一步地,为了能够对语音噪声进行识别,还需要构建噪声识别模型,所述分别构建初始噪声分类模型和多个初始噪声生成模型,包括:分别构建初始噪声识别模型,初始噪声分类模型和多个初始噪声生成模型。
基于此,所述根据所述多个随机语音序列和所述不同聚类类别下的真实语音序列,对所述初始噪声分类模型和所述多个初始噪声生成模型进行联合迭代训练,构建所述噪声分类模型和所述多个噪声生成模型,包括:将所述多个随机语音序列分别输入至所述多个初始噪声生成模型,生成不同种类的语音噪声;将生成的语音噪声和所述真实语音序列分别输入至所述初始噪声噪声识别模型进行噪声识别,得到初始噪声识别结果;提取初始噪声识别结果中语音噪声对应的语音特征,并将其输入至所述初始噪声分类模型进行噪声分类,得到初始噪声分类结果;基于所述初始噪声识别结果和所述初始噪声分类结果,分别构建噪声识别准确度损失函数和噪声分类准确度损失函数;根据噪声识别准确度损失函数和噪声分类准确度损失函数,对所述初始噪声识别模型、所述初始噪声分类模型和所述多个初始噪声生成模型进行联合迭代训练,分别构建噪声识别模型、所述噪声分类模型和所述多个噪声生成模型。其中,预设噪声生成模型采用卷积神经网络。
具体地,通过将不同种类的语音噪声和不同聚类类别下的真实语音序列分别输入至初始噪声识别模型进行噪声识别,得到初始噪声识别结果,之后提取初始识别结果中语音噪声对应的语音特征,将其输入至预设初始噪声分类模型进行噪声分类,得到噪声分类结果,并根据噪声分类结果和噪声识别结果,分别构建噪声识别准确度损失函数和噪声分类准确度损失函数,具体公式如下:
其中,Lc为噪声识别准确度损失函数,Lc为噪声分类准确度损失韩式,zi为语音噪声,xi为真实语音序列,D代表预设噪声识别模型,G代表预设噪声生成模型,c代表噪声分类模型,为了保证噪声生成模型所生成的语音噪声与真实语音序列更接近,增加噪声识别模型的识别难度,噪声生成模型与噪声识别模型的优化方向是相反的,即噪声生成模型需要最小化预设噪声识别模型的准确率,因此其优化方向是最小化Lc-Ls,而噪声分类模型的训练目的是最大化分类噪声的准确率,因此其优化方向是最大化Lc+Ls,由此通过上述两个优化方程,能够不断对初始噪声生成模型、初始噪声识别模型和初始噪声分类模型进行联合训练,构建噪声生成模型,噪声识别模型和噪声分类模型。
203、获取待识别的语音序列。
其中,待识别的语音序列为从某场景下获取的用户语音序列,该语音序列中可能包含语音噪声,也可能不包含语音噪声,为了确保后续的语音识别结果,如果待识别的语音序列中包含语音噪声,需要对语音噪声进行降噪进行降噪处理,在对语音噪声进行降噪处理时,为提高语音噪声的降噪处理效果,可以进一步对语音噪声的种类进行识别,以便根据语音噪声的种类选择合适的降噪处理策略对其进行降噪处理。
对所述语音序列进行噪音识别,若所述语音序列中包含语音噪声,则利用预设的噪声分类模型确定所述语音噪声对应的噪声类别。
其中,所述噪声分类模型是与多个噪声生成模型联合训练得到的,不同噪声生成模型所生成的语音噪声的种类不同。对于本申请实施例,为了确定语音噪声对应的噪声种类,步骤204具体包括:对所述语音序列进行语音特征提取,得到所述语音序列对应的语音特征;基于所述语音特征,判断所述语音序列中是否包含语音噪声;若包含语音噪声,则基于提取的语音特征,利用所述噪声分类模型确定所述语音噪声对应的噪声类别。
具体地,将待识别的语音序列输入至噪声识别模型进行噪声识别,在噪声识别的过程中,预设噪声识别模型中的隐藏层会提取待识别的语音序列对应的语音特征,基于提取的语音特征判定待识别的语音序列中是否包含语音噪声,若包含语音噪声,则将提取的语音特征输入至所述噪声分类模型进行噪声分类,以确定语音噪声对应的噪声类别。
205、基于所述噪声类别,确定所述语音噪声对应的最优降噪处理策略,并利用所述最优降噪处理策略对所述语音噪声进行降噪处理。
对于本方实施例,根据确定的语音噪声对应的噪声类别,从预设降噪策略库中选择该噪声类别对应的降噪处理策略,并将其确定为最优降噪处理策略,之后利用该最优降噪处理策略对待识别的语音序列中的语音噪声进行降噪处理,从而能够针对不同场景下的语音噪声,均能够达到最优降噪处理效果,避免采用统一的降噪处理策略,影像语音噪声的降噪处理效果。
本申请实施例提供的另一种语音噪声的处理方法,与目前针对不同种类的语音噪声均采用同种降噪策略进行降噪处理的方式相比,本申请能够获取待识别的语音序列;并对所述语音序列进行噪音识别,若所述语音序列中包含语音噪声,则利用预设的噪声分类模型确定所述语音噪声对应的噪声类别,其中,所述噪声分类模型是与多个噪声生成模型联合训练得到的,不同噪声生成模型所生成的语音噪声的种类不同;与此同时,基于所述噪声类别,确定所述语音噪声对应的最优降噪处理策略,并利用所述最优降噪处理策略对所述语音噪声进行降噪处理,由此通过将噪声分类模型与多个噪声生成模型联合进行训练,从而使得本申请中的噪声分类模型能够对不同场景下语音噪声的种类进行识别,进而能够根据确定的噪声类别,选择最优的降噪处理策略对语音噪声进行处理,能够达到最优的降噪处理效果。
进一步地,作为图1的具体实现,本申请实施例提供了一种语音噪声的处理装置,如图3所示,所述装置包括:获取单元31、确定单元32和降噪单元33。
所述获取单元31,可以用于获取待识别的语音序列。所述获取单元31是本装置中获取待识别的语音序列的主要功能模块。
所述确定单元32,可以用于对所述语音序列进行噪音识别,若所述语音序列中包含语音噪声,则利用预设的噪声分类模型确定所述语音噪声对应的噪声类别,其中,所述噪声分类模型是与多个噪声生成模型联合训练得到的,不同噪声生成模型所生成的语音噪声的种类不同。所述确定单元32是本装置中对所述语音序列进行噪音识别,若所述语音序列中包含语音噪声,则利用预设的噪声分类模型确定所述语音噪声对应的噪声类别的主要功能模块,也是核心模块。
所述降噪单元33,可以用于基于所述噪声类别,确定所述语音噪声对应的最优降噪处理策略,并利用所述最优降噪处理策略对所述语音噪声进行降噪处理。所述降噪单元33是本装置中基于所述噪声类别,确定所述语音噪声对应的最优降噪处理策略,并利用所述最优降噪处理策略对所述语音噪声进行降噪处理的主要功能模块。
进一步地,为了确定所述语音噪声对应的噪声类别,如图4所示,所述确定单元32,包括提取模块321、判断模块322和确定模块323。
所述提取模块321,可以用于对所述语音序列进行语音特征提取,得到所述待识别语音序列对应的语音特征。
所述判断模块322,可以用于基于所述语音特征,判断所述语音序列中是否包含语音噪声。
所述确定模块323,可以用于若包含语音噪声,则基于提取的语音特征,利用所述噪声分类模型确定所述语音噪声对应的噪声类别。
进一步地,为了构建预设噪声分类模型和多个噪声生成模型,所述装置还包括:聚类单元34和构建单元35。
所述获取单元31,还可以用于获取预设语音样本库中的真实语音序列以及多个随机语音序列。
所述聚类单元34,可以用于对所述真实语音序列进行聚类处理,得到不同聚类类别下的真实语音序列。
所述构建单元35,可以用于根据所述多个随机语音序列和所述不同聚类类别下的真实语音序列,构建所述噪声分类模型和所述多个噪声生成模型。
进一步地,为了对真实语音序列进行聚类处理,所述聚类单元34,包括:计算模块341和聚类模块342。
所述计算模块341,可以用于根据预设的欧式距离算法计算不同真实语音序列之间的欧式距离。
所述聚类模块342,可以用于基于所述欧式距离,对所述真实语音序列进行聚类处理,得到不同聚类类别下的真实语音序列。
进一步地,为了构建噪声分类模型和多个噪声生成模型,所述构建单元35,包括:第一构建模块351和第二构建模块352。
所述第一构建模块351,可以用于分别构建初始噪声分类模型和多个初始噪声生成模型。
所述第二构建模块352,可以用于根据所述多个随机语音序列和所述不同聚类类别下的真实语音序列,对所述初始噪声分类模型和所述多个初始噪声生成模型进行联合迭代训练,构建所述噪声分类模型和所述多个噪声生成模型。
进一步地,所述第二构建模块352,包括:生成子模块、识别子模块、分类子模块和构建子模块。
所述生成子模块,可以用于将所述多个随机语音序列分别输入至所述多个初始噪声生成模型,生成不同种类的语音噪声。
所述识别子模块,可以用于将生成的的语音噪声和所述真实语音序列分别输入至所述初始噪声噪声识别模型进行噪声识别,得到初始噪声识别结果。
所述分类子模块,可以用于提取初始噪声识别结果中语音噪声对应的语音特征,并将其输入至所述初始噪声分类模型进行噪声分类,得到初始噪声分类结果。
所述构建子模块,可以用于基于所述初始噪声识别结果和所述初始噪声分类结果,分别构建噪声识别准确度损失函数和噪声分类准确度损失函数。
所述构建子模块,还可以用于根据噪声识别准确度损失函数和噪声分类准确度损失函数,对所述初始噪声识别模型、所述初始噪声分类模型和所述多个初始噪声生成模型进行联合迭代训练,分别构建噪声识别模型、所述噪声分类模型和所述多个噪声生成模型。
需要说明的是,本申请实施例提供的一种语音噪声的处理装置所涉及各功能模块的其他相应描述,可以参考图1所示方法的对应描述,在此不再赘述。
基于上述如图1所示方法,相应的,本申请实施例还提供了一种计算机可读存储介质,所述计算机可读存储介质可以是非易失性,也可以是易失性,其上存储有计算机程序,该程序被处理器执行时实现以下步骤:获取待识别语音序列;获取待识别的语音序列;对所述语音序列进行噪音识别,若所述语音序列中包含语音噪声,则利用预设的噪声分类模型确定所述语音噪声对应的噪声类别,其中,所述噪声分类模型是与多个噪声生成模型联合训练得到的,不同噪声生成模型所生成的语音噪声的种类不同;基于所述噪声类别,确定所述语音噪声对应的最优降噪处理策略,并利用所述最优降噪处理策略对所述语音噪声进行降噪处理。
基于上述如图1所示方法和如图3所示装置的实施例,本申请实施例还提供了一种计算机设备的实体结构图,如图5所示,该计算机设备包括:处理器41、存储器42、及存储在存储器42上并可在处理器上运行的计算机程序,其中存储器42和处理器41均设置在总线43上所述处理器41执行所述程序时实现以下步骤:获取待识别的语音序列;对所述语音序列进行噪音识别,若所述语音序列中包含语音噪声,则利用预设的噪声分类模型确定所述语 音噪声对应的噪声类别,其中,所述噪声分类模型是与多个噪声生成模型联合训练得到的,不同噪声生成模型所生成的语音噪声的种类不同;基于所述噪声类别,确定所述语音噪声对应的最优降噪处理策略,并利用所述最优降噪处理策略对所述语音噪声进行降噪处理。
通过本申请的技术方案,本申请能获取待识别的语音序列;并对所述语音序列进行噪音识别,若所述语音序列中包含语音噪声,则利用预设的噪声分类模型确定所述语音噪声对应的噪声类别,其中,所述噪声分类模型是与多个噪声生成模型联合训练得到的,不同噪声生成模型所生成的语音噪声的种类不同;与此同时,基于所述噪声类别,确定所述语音噪声对应的最优降噪处理策略,并利用所述最优降噪处理策略对所述语音噪声进行降噪处理,由此通过将噪声分类模型与多个噪声生成模型联合进行训练,从而使得本申请中的噪声分类模型能够对不同场景下语音噪声的种类进行识别,进而能够根据确定的噪声类别,选择最优的降噪处理策略对语音噪声进行处理,能够达到最优的降噪处理效果。
Claims (20)
- 一种语音噪声的处理方法,包括:获取待识别的语音序列;对所述语音序列进行噪音识别,若所述语音序列中包含语音噪声,则利用预设的噪声分类模型确定所述语音噪声对应的噪声类别,其中,所述噪声分类模型是与多个噪声生成模型联合训练得到的,不同噪声生成模型所生成的语音噪声的种类不同;基于所述噪声类别,确定所述语音噪声对应的最优降噪处理策略,并利用所述最优降噪处理策略对所述语音噪声进行降噪处理。
- 根据权利要求1所述的方法,其中,所述若所述语音序列中包含语音噪声,则利用预设的噪声分类模型确定所述语音噪声对应的噪声类别,包括:对所述语音序列进行语音特征提取,得到所述语音序列对应的语音特征;基于所述语音特征,判断所述语音序列中是否包含语音噪声;若包含语音噪声,则基于提取的语音特征,利用所述噪声分类模型确定所述语音噪声对应的噪声类别。
- 根据权利要求1所述的方法,其中,在所述获取待识别的语音序列之前,所述方法还包括:获取预设语音样本库中的真实语音序列以及多个随机语音序列;对所述真实语音序列进行聚类处理,得到不同聚类类别下的真实语音序列;根据所述多个随机语音序列和所述不同聚类类别下的真实语音序列,构建所述噪声分类模型和所述多个噪声生成模型。
- 根据权利要求3所述的方法,其中,所述对所述真实语音序列进行聚类处理,得到不同聚类类别下的真实语音序列,包括:根据预设的欧式距离算法计算不同真实语音序列之间的欧式距离;基于所述欧式距离,对所述真实语音序列进行聚类处理,得到不同聚类类别下的真实语音序列。
- 根据权利要求3所述的方法,其中,所述根据所述多个随机语音序列和所述不同聚类类别下的真实语音序列,建预所述噪声分类模型和所述多个噪声生成模型,包括:分别构建初始噪声分类模型和多个初始噪声生成模型;根据所述多个随机语音序列和所述不同聚类类别下的真实语音序列,对所述初始噪声分类模型和所述多个初始噪声生成模型进行联合迭代训练,构建所述噪声分类模型和所述多个噪声生成模型。
- 根据权利要求5所述的方法,其中,所述分别构建初始噪声分类模型和多个初始噪声生成模型,包括:分别构建初始噪声识别模型,初始噪声分类模型和多个初始噪声生成模型;所述根据所述多个随机语音序列和所述不同聚类类别下的真实语音序列,对所述初始噪声分类模型和所述多个初始噪声生成模型进行联合迭代训练,构建所述噪声分类模型和所述多个噪声生成模型,包括:将所述多个随机语音序列分别输入至所述多个初始噪声生成模型,生 成不同种类的语音噪声;将生成的的语音噪声和所述真实语音序列分别输入至所述初始噪声噪声识别模型进行噪声识别,得到初始噪声识别结果;提取初始噪声识别结果中语音噪声对应的语音特征,并将其输入至所述初始噪声分类模型进行噪声分类,得到初始噪声分类结果;基于所述初始噪声识别结果和所述初始噪声分类结果,分别构建噪声识别准确度损失函数和噪声分类准确度损失函数;根据噪声识别准确度损失函数和噪声分类准确度损失函数,对所述初始噪声识别模型、所述初始噪声分类模型和所述多个初始噪声生成模型进行联合迭代训练,分别构建噪声识别模型、所述噪声分类模型和所述多个噪声生成模型。
- 根据权利要求3-6任一项所述的方法,其中,所述多个随机语音序列服从高斯分布。
- 一种语音噪声的处理装置,包括:获取单元,用于获取待识别的语音序列;确定单元,用于对所述语音序列进行噪声识别,若所述语音序列中包含语音噪声,则利用预设的噪声分类模型确定所述语音噪声对应的噪声类别,其中,所述噪声分类模型是与多个噪声生成模型联合训练得到的,不同噪声生成模型所生成的语音噪声的种类不同;降噪单元,用于基于所述噪声类别,确定所述语音噪声对应的最优降噪处理策略,并利用所述最优降噪处理策略对所述语音噪声进行降噪处理。
- 一种计算机可读存储介质,其上存储有计算机程序,所述计算机程序被处理器执行时实现一种语音噪声的处理方法的步骤:获取待识别的语音序列;对所述语音序列进行噪音识别,若所述语音序列中包含语音噪声,则利用预设的噪声分类模型确定所述语音噪声对应的噪声类别,其中,所述噪声分类模型是与多个噪声生成模型联合训练得到的,不同噪声生成模型所生成的语音噪声的种类不同;基于所述噪声类别,确定所述语音噪声对应的最优降噪处理策略,并利用所述最优降噪处理策略对所述语音噪声进行降噪处理。
- 根据权利要求9所述的计算机可读存储介质,其中,所述若所述语音序列中包含语音噪声,则利用预设的噪声分类模型确定所述语音噪声对应的噪声类别,包括:对所述语音序列进行语音特征提取,得到所述语音序列对应的语音特征;基于所述语音特征,判断所述语音序列中是否包含语音噪声;若包含语音噪声,则基于提取的语音特征,利用所述噪声分类模型确定所述语音噪声对应的噪声类别。
- 根据权利要求9所述的计算机可读存储介质,其中,在所述获取待识别的语音序列之前,所述方法还包括:获取预设语音样本库中的真实语音序列以及多个随机语音序列;对所述真实语音序列进行聚类处理,得到不同聚类类别下的真实语音序列;根据所述多个随机语音序列和所述不同聚类类别下的真实语音序列,构建所述噪声分类模型和所述多个噪声生成模型。
- 根据权利要求11所述的计算机可读存储介质,其中,所述对所述真实语音序列进行聚类处理,得到不同聚类类别下的真实语音序列,包括:根据预设的欧式距离算法计算不同真实语音序列之间的欧式距离;基于所述欧式距离,对所述真实语音序列进行聚类处理,得到不同聚类类别下的真实语音序列。
- 根据权利要求11所述的计算机可读存储介质,其中,所述根据所述多个随机语音序列和所述不同聚类类别下的真实语音序列,建预所述噪声分类模型和所述多个噪声生成模型,包括:分别构建初始噪声分类模型和多个初始噪声生成模型;根据所述多个随机语音序列和所述不同聚类类别下的真实语音序列,对所述初始噪声分类模型和所述多个初始噪声生成模型进行联合迭代训练,构建所述噪声分类模型和所述多个噪声生成模型。
- 根据权利要求13所述的计算机可读存储介质,其中,所述分别构建初始噪声分类模型和多个初始噪声生成模型,包括:分别构建初始噪声识别模型,初始噪声分类模型和多个初始噪声生成模型;所述根据所述多个随机语音序列和所述不同聚类类别下的真实语音序列,对所述初始噪声分类模型和所述多个初始噪声生成模型进行联合迭代训练,构建所述噪声分类模型和所述多个噪声生成模型,包括:将所述多个随机语音序列分别输入至所述多个初始噪声生成模型,生成不同种类的语音噪声;将生成的的语音噪声和所述真实语音序列分别输入至所述初始噪声噪声识别模型进行噪声识别,得到初始噪声识别结果;提取初始噪声识别结果中语音噪声对应的语音特征,并将其输入至所述初始噪声分类模型进行噪声分类,得到初始噪声分类结果;基于所述初始噪声识别结果和所述初始噪声分类结果,分别构建噪声识别准确度损失函数和噪声分类准确度损失函数;根据噪声识别准确度损失函数和噪声分类准确度损失函数,对所述初始噪声识别模型、所述初始噪声分类模型和所述多个初始噪声生成模型进行联合迭代训练,分别构建噪声识别模型、所述噪声分类模型和所述多个噪声生成模型。
- 根据权利要求11-14任一项所述的计算机可读存储介质,其中,所述多个随机语音序列服从高斯分布。
- 一种计算机设备,包括存储器、处理器及存储在存储器上并可在处理器上运行的计算机程序,所述计算机程序被处理器执行时实现一种语音噪声的处理方法的步骤:获取待识别的语音序列;对所述语音序列进行噪音识别,若所述语音序列中包含语音噪声,则利用预设的噪声分类模型确定所述语音噪声对应的噪声类别,其中,所述噪声分类模型是与多个噪声生成模型联合训练得到的,不同噪声生成模型所生成的语音噪声的种类不同;基于所述噪声类别,确定所述语音噪声对应的最优降噪处理策略,并利用所述最优降噪处理策略对所述语音噪声进行降噪处理。
- 根据权利要求16所述的计算机设备,其中,所述若所述语音序列中包含语音噪声,则利用预设的噪声分类模型确定所述语音噪声对应的噪声类别,包括:对所述语音序列进行语音特征提取,得到所述语音序列对应的语音特征;基于所述语音特征,判断所述语音序列中是否包含语音噪声;若包含语音噪声,则基于提取的语音特征,利用所述噪声分类模型确定所述语音噪声对应的噪声类别。
- 根据权利要求16所述的计算机设备,其中,在所述获取待识别的语音序列之前,所述方法还包括:获取预设语音样本库中的真实语音序列以及多个随机语音序列;对所述真实语音序列进行聚类处理,得到不同聚类类别下的真实语音序列;根据所述多个随机语音序列和所述不同聚类类别下的真实语音序列,构建所述噪声分类模型和所述多个噪声生成模型。
- 根据权利要求18所述的计算机设备,其中,所述对所述真实语音序列进行聚类处理,得到不同聚类类别下的真实语音序列,包括:根据预设的欧式距离算法计算不同真实语音序列之间的欧式距离;基于所述欧式距离,对所述真实语音序列进行聚类处理,得到不同聚类类别下的真实语音序列。
- 根据权利要求18所述的计算机设备,其中,所述根据所述多个随机语音序列和所述不同聚类类别下的真实语音序列,建预所述噪声分类模型和所述多个噪声生成模型,包括:分别构建初始噪声分类模型和多个初始噪声生成模型;根据所述多个随机语音序列和所述不同聚类类别下的真实语音序列,对所述初始噪声分类模型和所述多个初始噪声生成模型进行联合迭代训练,构建所述噪声分类模型和所述多个噪声生成模型。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202011153509.1 | 2020-10-26 | ||
| CN202011153509.1A CN112201270B (zh) | 2020-10-26 | 2020-10-26 | 语音噪声的处理方法、装置、计算机设备及存储介质 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2021189981A1 true WO2021189981A1 (zh) | 2021-09-30 |
Family
ID=74011358
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2020/136367 Ceased WO2021189981A1 (zh) | 2020-10-26 | 2020-12-15 | 语音噪声的处理方法、装置、计算机设备及存储介质 |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN112201270B (zh) |
| WO (1) | WO2021189981A1 (zh) |
Families Citing this family (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN113869107A (zh) * | 2021-08-20 | 2021-12-31 | 杭州回车电子科技有限公司 | 信号去噪方法、装置、电子装置和存储介质 |
| CN117409767A (zh) * | 2023-09-27 | 2024-01-16 | 杭州阿里云飞天信息技术有限公司 | 语音数据、会议语音的处理方法及服务器 |
| CN118571241B (zh) * | 2024-08-02 | 2024-09-27 | 深圳波洛斯科技有限公司 | 一种基于dnn降噪技术的窗口对讲系统 |
| CN119296560B (zh) * | 2024-12-11 | 2025-03-14 | 杭州华亭科技有限公司 | 一种多噪声环境下的语音降噪系统 |
Citations (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN101710490A (zh) * | 2009-11-20 | 2010-05-19 | 安徽科大讯飞信息科技股份有限公司 | 语音评测的噪声补偿方法及装置 |
| US20120185246A1 (en) * | 2011-01-19 | 2012-07-19 | Broadcom Corporation | Noise suppression using multiple sensors of a communication device |
| CN102693724A (zh) * | 2011-03-22 | 2012-09-26 | 张燕 | 一种基于神经网络的高斯混合模型的噪声分类方法 |
| CN103065631A (zh) * | 2013-01-24 | 2013-04-24 | 华为终端有限公司 | 一种语音识别的方法、装置 |
| CN103219011A (zh) * | 2012-01-18 | 2013-07-24 | 联想移动通信科技有限公司 | 降噪方法、装置与通信终端 |
| CN104575510A (zh) * | 2015-02-04 | 2015-04-29 | 深圳酷派技术有限公司 | 降噪方法、降噪装置和终端 |
Family Cites Families (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CA2636684C (en) * | 1997-12-24 | 2009-08-18 | Mitsubishi Denki Kabushiki Kaisha | A method for speech coding, method for speech decoding and their apparatuses |
| JP4033299B2 (ja) * | 2003-03-12 | 2008-01-16 | 株式会社エヌ・ティ・ティ・ドコモ | 音声モデルの雑音適応化システム、雑音適応化方法、及び、音声認識雑音適応化プログラム |
| EP1732063A4 (en) * | 2004-03-31 | 2007-07-04 | Pioneer Corp | LANGUAGE RECOGNITION AND LANGUAGE RECOGNITION METHOD |
| US9313585B2 (en) * | 2008-12-22 | 2016-04-12 | Oticon A/S | Method of operating a hearing instrument based on an estimation of present cognitive load of a user and a hearing aid system |
| CN109471853B (zh) * | 2018-09-18 | 2023-06-16 | 平安科技(深圳)有限公司 | 数据降噪方法、装置、计算机设备和存储介质 |
-
2020
- 2020-10-26 CN CN202011153509.1A patent/CN112201270B/zh active Active
- 2020-12-15 WO PCT/CN2020/136367 patent/WO2021189981A1/zh not_active Ceased
Patent Citations (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN101710490A (zh) * | 2009-11-20 | 2010-05-19 | 安徽科大讯飞信息科技股份有限公司 | 语音评测的噪声补偿方法及装置 |
| US20120185246A1 (en) * | 2011-01-19 | 2012-07-19 | Broadcom Corporation | Noise suppression using multiple sensors of a communication device |
| CN102693724A (zh) * | 2011-03-22 | 2012-09-26 | 张燕 | 一种基于神经网络的高斯混合模型的噪声分类方法 |
| CN103219011A (zh) * | 2012-01-18 | 2013-07-24 | 联想移动通信科技有限公司 | 降噪方法、装置与通信终端 |
| CN103065631A (zh) * | 2013-01-24 | 2013-04-24 | 华为终端有限公司 | 一种语音识别的方法、装置 |
| CN104575510A (zh) * | 2015-02-04 | 2015-04-29 | 深圳酷派技术有限公司 | 降噪方法、降噪装置和终端 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN112201270B (zh) | 2023-05-23 |
| CN112201270A (zh) | 2021-01-08 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| CN112088402B (zh) | 用于说话者识别的联合神经网络 | |
| US10776470B2 (en) | Verifying identity based on facial dynamics | |
| CN110147721B (zh) | 一种三维人脸识别方法、模型训练方法和装置 | |
| WO2021189981A1 (zh) | 语音噪声的处理方法、装置、计算机设备及存储介质 | |
| CN111723679A (zh) | 基于深度迁移学习的人脸和声纹认证系统及方法 | |
| CN112951258B (zh) | 一种音视频语音增强处理方法及装置 | |
| CN112132847A (zh) | 模型训练方法、图像分割方法、装置、电子设备和介质 | |
| WO2019015466A1 (zh) | 人证核实的方法及装置 | |
| WO2018176894A1 (zh) | 一种说话人确认方法及装置 | |
| CN111563422A (zh) | 基于双模态情绪识别网络的服务评价获取方法及其装置 | |
| CN111108508B (zh) | 脸部情感识别方法、智能装置和计算机可读存储介质 | |
| CN113361636A (zh) | 一种图像分类方法、系统、介质及电子设备 | |
| CN113948105B (zh) | 基于语音的图像生成方法、装置、设备及介质 | |
| CN118378029B (zh) | 一种基于机器翻译的多模态数据预处理方法 | |
| CN105989849A (zh) | 一种语音增强方法、语音识别方法、聚类方法及装置 | |
| WO2022268183A1 (zh) | 一种基于视频的随机手势认证方法及系统 | |
| CN112199976A (zh) | 证件图片生成方法及装置 | |
| CN114218428A (zh) | 音频数据聚类方法、装置、设备及存储介质 | |
| CN114399803A (zh) | 人脸关键点检测方法及装置 | |
| CN119517057A (zh) | 一种基于时频图卷积网络的语音增强方法及系统 | |
| CN119516054A (zh) | 一种基于大模型可学习文本潜码的说话数字人生成方法 | |
| CN114663685A (zh) | 一种行人重识别模型训练的方法、装置和设备 | |
| CN111368763A (zh) | 基于头像的图像处理方法、装置及计算机可读存储介质 | |
| US20200394289A1 (en) | Biometric verification framework that utilizes a convolutional neural network for feature matching | |
| CN119963703A (zh) | 高保真高同步的说话人脸生成模型训练方法及系统 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 20926754 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 20926754 Country of ref document: EP Kind code of ref document: A1 |

