WO2023045779A1 - 一种音频降噪方法、装置、设备及存储介质 - Google Patents
一种音频降噪方法、装置、设备及存储介质 Download PDFInfo
- Publication number
- WO2023045779A1 WO2023045779A1 PCT/CN2022/118040 CN2022118040W WO2023045779A1 WO 2023045779 A1 WO2023045779 A1 WO 2023045779A1 CN 2022118040 W CN2022118040 W CN 2022118040W WO 2023045779 A1 WO2023045779 A1 WO 2023045779A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- audio data
- audio
- spectrum
- complex
- denoised
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/08—Learning methods
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L21/00—Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
- G10L21/02—Speech enhancement, e.g. noise reduction or echo cancellation
- G10L21/0208—Noise filtering
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L21/00—Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
- G10L21/02—Speech enhancement, e.g. noise reduction or echo cancellation
- G10L21/0208—Noise filtering
- G10L21/0216—Noise filtering characterised by the method used for estimating noise
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L21/00—Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
- G10L21/02—Speech enhancement, e.g. noise reduction or echo cancellation
- G10L21/0208—Noise filtering
- G10L21/0216—Noise filtering characterised by the method used for estimating noise
- G10L21/0224—Processing in the time domain
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L21/00—Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
- G10L21/02—Speech enhancement, e.g. noise reduction or echo cancellation
- G10L21/0208—Noise filtering
- G10L21/0216—Noise filtering characterised by the method used for estimating noise
- G10L21/0232—Processing in the frequency domain
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L21/00—Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
- G10L21/02—Speech enhancement, e.g. noise reduction or echo cancellation
- G10L21/0316—Speech enhancement, e.g. noise reduction or echo cancellation by changing the amplitude
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04R—LOUDSPEAKERS, MICROPHONES, GRAMOPHONE PICK-UPS OR LIKE ACOUSTIC ELECTROMECHANICAL TRANSDUCERS; ELECTRIC HEARING AIDS; PUBLIC ADDRESS SYSTEMS
- H04R3/00—Circuits for transducers
- H04R3/04—Circuits for transducers for correcting frequency response
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04R—LOUDSPEAKERS, MICROPHONES, GRAMOPHONE PICK-UPS OR LIKE ACOUSTIC ELECTROMECHANICAL TRANSDUCERS; ELECTRIC HEARING AIDS; PUBLIC ADDRESS SYSTEMS
- H04R2430/00—Signal processing covered by H04R, not provided for in its groups
- H04R2430/01—Aspects of volume control, not necessarily automatic, in sound systems
Definitions
- the present disclosure relates to the field of data processing, and in particular to an audio noise reduction method, device, equipment and storage medium.
- an embodiment of the present disclosure provides an audio noise reduction method, which can implement audio noise reduction, thereby better improving the sound quality of the audio.
- the present disclosure provides an audio noise reduction method, the method comprising:
- Denoising result audio data corresponding to the audio data to be reduced is determined based on the first-order enhanced amplitude spectrum corresponding to the audio data to be reduced and the complex time-frequency mask.
- the estimating the complex time-frequency mask of the audio data to be denoised by using the preset complex network model includes:
- the complex spectrum to be denoised includes a complex spectrum determined based on the first-order enhanced amplitude spectrum corresponding to the audio data to be denoised and the original phase spectrum of the audio data to be denoised, or , a complex spectrum determined based on the original spectrum and the original phase spectrum of the audio data to be denoised;
- the determining the noise reduction result audio data corresponding to the audio data to be reduced based on the first-order enhanced amplitude spectrum corresponding to the audio data to be reduced and the complex time-frequency mask includes:
- phase enhancement spectrum corresponding to the audio data to be reduced based on the phase gain and the original phase spectrum corresponding to the audio data to be reduced
- the preset real network model and the preset complex network model are used to form a two-stage time-domain convolutional network TCN model.
- the preset real number network model before estimating the amplitude-time-frequency masking of the audio data to be denoised by using the preset real number network model, it further includes:
- the two-stage TCN model is trained by using audio training samples whose sampling rate is higher than a preset sampling rate threshold.
- the audio training samples whose sampling rate is higher than the preset sampling rate threshold before training the two-stage TCN model, it also includes:
- using the audio training samples whose sampling rate is higher than the preset sampling rate threshold to train the two-stage TCN model includes:
- the two-stage TCN model is trained by using the augmented audio training samples; wherein, the sampling rate of the augmented audio training samples is higher than a preset sampling rate threshold.
- the present disclosure provides an audio noise reduction device, the device comprising:
- An acquisition module configured to acquire audio data to be denoised
- the first estimation module is configured to estimate the amplitude time-frequency mask of the audio data to be reduced by using a preset real number network model; wherein the amplitude time-frequency mask is used to determine the first-order enhancement corresponding to the audio data to be reduced amplitude spectrum;
- the second estimation module is used to estimate the complex time-frequency mask of the audio data to be denoised by using a preset complex network model
- the first determining module is configured to determine the noise reduction result audio data corresponding to the audio data to be reduced based on the first-order enhanced amplitude spectrum corresponding to the audio data to be reduced and the complex time-frequency mask.
- the present disclosure provides a computer-readable storage medium, where instructions are stored in the computer-readable storage medium, and when the instructions are run on a terminal device, the terminal device is made to implement the above method.
- the present disclosure provides a device, including: a memory, a processor, and a computer program stored on the memory and operable on the processor, when the processor executes the computer program, Implement the above method.
- the present disclosure provides a computer program product, where the computer program product includes a computer program/instruction, and when the computer program/instruction is executed by a processor, the above method is implemented.
- An embodiment of the present disclosure provides an audio noise reduction method.
- the audio data to be reduced is obtained, and then the amplitude time-frequency mask of the audio data to be reduced is estimated by using a preset real number network model, and the corresponding frequency of the audio data to be reduced can be obtained.
- First order enhanced magnitude spectrum is estimated.
- the complex time-frequency mask of the audio data to be denoised is estimated by using the preset complex network model, and the denoising result audio data corresponding to the audio data to be denoised is determined by combining the first-order enhanced amplitude spectrum and the complex time-frequency mask.
- the embodiments of the present disclosure use the preset real number network model to enhance the amplitude spectrum of the audio data to be denoised, and use the preset complex number network model to simultaneously enhance the amplitude spectrum and phase spectrum of the audio data to be denoised. It can be seen that the embodiments of the present disclosure can realize the Noise reduction Noise reduction processing of audio data, so as to better improve the sound quality of audio.
- FIG. 1 is a flowchart of an audio noise reduction method provided by an embodiment of the present disclosure
- FIG. 2 is a schematic diagram of a two-stage TCN model provided by an embodiment of the present disclosure
- FIG. 3 is a schematic structural diagram of an audio noise reduction device provided by an embodiment of the present disclosure.
- Fig. 4 is a schematic structural diagram of an audio noise reduction device provided by an embodiment of the present disclosure.
- noise in audio can be divided into at least two types: stationary noise and non-stationary noise.
- Stationary noise means that the statistical characteristics of noise will not change with time, and common ones include white noise and pink noise.
- Non-stationary noise refers to statistical noise Characteristics change over time, such as keyboard sound, mouse click sound, etc.
- audio noise reduction tools often use a single network model to achieve audio noise reduction.
- the complexity of the network model is low, it is difficult to guarantee the noise reduction effect on audio, for example, it is especially difficult to guarantee the non-stationary noise in audio. inhibitory effect.
- an embodiment of the present disclosure provides an audio noise reduction method, which uses a preset real network model and a preset complex network model to perform noise reduction processing on the audio data to be reduced, and then synthesizes the noise reduction results of the two to determine the The audio data of the noise reduction result corresponding to the noise reduction audio data, it can be seen that compared with using a single network model for audio noise reduction, the embodiments of the present disclosure can have a better suppression effect on non-stationary noise, thereby ensuring the overall sound quality of the audio Noise reduction effect, and then better improve the sound quality of the audio.
- the embodiment of the present disclosure obtains the audio data to be denoised, and then uses a preset real number network model to estimate the amplitude time-frequency mask of the audio data to be denoised, so as to obtain the first-order enhanced amplitude spectrum corresponding to the audio data to be denoised. Furthermore, the complex time-frequency mask of the audio data to be denoised is estimated by using the preset complex network model, and the denoising result audio data corresponding to the audio data to be denoised is determined by combining the first-order enhanced amplitude spectrum and the complex time-frequency mask.
- the embodiments of the present disclosure use the preset real number network model to enhance the amplitude spectrum of the audio data to be denoised, and use the preset complex number network model to simultaneously enhance the amplitude spectrum and phase spectrum of the audio data to be denoised. It can be seen that the embodiments of the present disclosure can realize the Noise reduction Noise reduction processing of audio data, while ensuring the noise reduction effect, thereby better improving the sound quality of the audio.
- an embodiment of the present disclosure provides an audio noise reduction method.
- FIG. 1 it is a flow chart of an audio noise reduction method provided by an embodiment of the present disclosure. The method includes:
- S101 Acquire audio data to be denoised.
- the audio data to be denoised in the embodiments of the present disclosure may be any audio segment, where the audio segment may also be an audio segment extracted from a video, or the like.
- the embodiment of the present disclosure does not limit the audio data to be denoised.
- the embodiment of the present disclosure may perform real-time noise reduction processing on the audio data to be reduced for noise during the audio recording stage, or may perform noise reduction processing on the audio data to be reduced during the audio editing stage.
- Embodiments of the present disclosure do not limit noise reduction scenarios.
- S102 Estimate an amplitude-time-frequency mask of the audio data to be denoised by using a preset real number network model.
- the amplitude time-frequency mask is used to determine the first-order enhanced amplitude spectrum corresponding to the audio data to be denoised.
- the preset real number network model is trained by using the audio training samples, and the trained preset real number network model is obtained, which is used to perform amplitude enhancement processing on the audio data to be denoised.
- the preset real number network model can be realized based on any AI model.
- the preset real number network model can be realized by Temporal Convolutional Network (TCN) or by Recurrent Neural Network (RNN). ) to achieve and so on.
- the audio data to be denoised can be input into the preset real number network model for processing, and the preset real number network model outputs the amplitude of the audio data to be denoised time-frequency masking.
- the magnitude time-frequency mask is used to represent the proportional relationship between the enhanced magnitude spectrum and the original magnitude spectrum.
- the amplitude time-frequency mask is used to determine the first-order enhanced amplitude spectrum corresponding to the audio data to be reduced, including: the amplitude time-frequency mask is used to match the original frequency spectrum of the audio data to be reduced. The amplitude spectra are multiplied to obtain the first-order enhanced amplitude spectrum corresponding to the audio data to be denoised.
- the amplitude time-frequency mask of the audio data to be denoised is obtained, and then, by comparing the amplitude time-frequency mask with the original amplitude spectrum of the audio data to be denoised Multiply, get the enhanced amplitude spectrum of the audio data to be denoised, as the first-order enhanced amplitude spectrum.
- the first-order enhanced amplitude spectrum is the amplitude spectrum after the frequency spectrum of the audio data to be denoised is enhanced through a preset real number network model.
- S103 Estimate a complex time-frequency mask of the audio data to be denoised by using a preset complex network model.
- the preset complex network model is trained by using the audio training data to obtain the trained preset complex network model, which is used to simultaneously enhance the amplitude and phase of the audio data to be denoised.
- the preset complex network model can be realized based on any AI model.
- the preset complex network model can be realized by Temporal Convolutional Network (TCN), or by Recurrent Neural Network (RNN). ) to achieve and so on.
- TCN Temporal Convolutional Network
- RNN Recurrent Neural Network
- the complex spectrum determined based on the original spectrum and the original phase spectrum of the audio data to be denoised is first determined as the Noise reduction complex spectrum. Then, the complex frequency spectrum to be denoised is input into a preset complex network model for processing, and the preset complex network model outputs a complex time-frequency mask corresponding to the audio data to be denoised.
- the complex time-frequency mask is used to represent the proportional relationship between the enhanced spectrum and the original spectrum, and the complex time-frequency mask includes a real part and an imaginary part.
- the embodiments of the present disclosure can also determine the complex spectrum determined based on the first-order enhanced amplitude spectrum and the original phase spectrum corresponding to the audio data to be denoised as the complex spectrum to be denoised, so as to preset the complex network model
- the amplitude and phase of the frequency spectrum of the audio data to be reduced can be further enhanced, thereby further improving the effect of noise reduction.
- the original phase spectrum of the audio data to be denoised is first obtained, and then the frequency spectrum determined based on the first-order enhanced amplitude spectrum and the original phase spectrum corresponding to the audio data to be denoised is determined as the complex number to be denoised spectrum. Furthermore, the complex frequency spectrum to be denoised is input into a preset complex network model for processing, and the preset complex network model outputs a complex time-frequency mask corresponding to the audio data to be denoised.
- S104 Determine noise reduction result audio data corresponding to the audio data to be reduced based on the first-order enhanced magnitude spectrum corresponding to the audio data to be reduced and the complex time-frequency mask.
- the corresponding First-order enhanced magnitude spectrum and complex time-frequency masking After the amplitude enhancement of the audio data to be denoised by the preset real number network model, and the simultaneous enhancement of the amplitude and phase of the audio data to be denoised by the preset complex number network model, the corresponding First-order enhanced magnitude spectrum and complex time-frequency masking. Then, based on the first-order enhanced amplitude spectrum and the complex time-frequency mask corresponding to the audio data to be reduced, the noise reduction result audio data corresponding to the audio data to be reduced is determined, and the noise reduction processing of the frequency to be reduced is realized.
- the amplitude gain and the phase gain are determined based on the complex time-frequency mask.
- the amplitude gain is used to characterize the amplitude enhancement of the spectrum of the audio data to be denoised by the preset complex network model
- the phase gain is used to characterize the phase enhancement of the frequency spectrum of the audio data to be denoised by the preset complex network model.
- the phase enhancement spectrum corresponding to the audio data to be reduced is determined.
- the second-order enhanced amplitude spectrum corresponding to the audio data to be reduced determine the second-order enhanced amplitude spectrum corresponding to the audio data to be reduced.
- the second-order enhanced amplitude spectrum is an amplitude spectrum obtained by performing amplitude enhancement on the audio data to be denoised through a preset real number network model and a preset complex number network model. Furthermore, based on the second-order enhanced amplitude spectrum and phase enhanced spectrum, the enhanced spectrum corresponding to the audio data to be reduced is determined, and the noise reduction result audio data corresponding to the audio data to be reduced is determined based on the enhanced spectrum.
- formulas (1) and (2) can be used to calculate the amplitude gain and phase gain, respectively.
- formula (3) can be used to calculate the enhanced spectrum corresponding to the audio data to be denoised, the following formula (3):
- Y phase is used to represent the original phase spectrum, is used to represent the phase-enhanced spectrum, is used to represent the first-order enhanced magnitude spectrum, Used to represent the second-order enhanced magnitude spectrum.
- the denoising result audio data corresponding to the audio data to be denoised is obtained through processing such as inverse Fourier transform.
- the audio data to be reduced is obtained, and then the amplitude time-frequency mask of the audio data to be reduced is estimated by using the preset real number network model, and the corresponding frequency of the audio data to be reduced can be obtained.
- the complex time-frequency mask of the audio data to be denoised is estimated by using the preset complex network model, and the denoising result audio data corresponding to the audio data to be denoised is determined by combining the first-order enhanced amplitude spectrum and the complex time-frequency mask.
- the embodiments of the present disclosure use the preset real number network model to enhance the amplitude spectrum of the audio data to be denoised, and use the preset complex number network model to simultaneously enhance the amplitude spectrum and phase spectrum of the audio data to be denoised. It can be seen that the embodiments of the present disclosure can realize the Noise reduction Noise reduction processing of audio data, so as to better improve the sound quality of audio.
- the embodiments of the present disclosure can implement a preset real number network model and a preset complex number network model based on the TCN model.
- the embodiments of the present disclosure can use the two-stage temporal convolution network TCN model to perform noise reduction processing on audio, thereby improving the sound quality of audio to a large extent.
- the two-stage TCN model includes a real TCN model and a complex TCN model, and Y(n) is used to represent the audio data to be denoised.
- the complex time-frequency mask corresponding to Y(n) includes the real part and the imaginary part
- the two-stage TCN model before using the two-stage TCN model to denoise the audio, the two-stage TCN model is first trained. Specifically, the two-stage TCN model can be trained by using audio training data whose sampling rate is higher than the preset sampling rate threshold, so that the trained two-stage TCN model can perform better noise reduction on audio data with a higher sampling rate Effect.
- the preset sampling rate threshold may be a value greater than 16K.
- the two-stage TCN model can be trained by using the time domain loss function SISNR.
- the time domain loss function SISNR will not be introduced here.
- preset data augmentation processing can be performed on the audio training samples to enrich the diversity of the audio training samples.
- the preset data augmentation processing includes performing high-pass, low-pass, band-pass, setting different volumes and/or equalizing the audio training samples according to preset probabilities
- the preset data augmentation processing may include processing operations such as high-pass, low-pass, band-pass, setting different volumes and/or equalization on the audio training samples with a certain probability.
- the augmented audio training samples can be used to train the two-stage TCN model.
- the sampling rate of the augmented audio training samples may be higher than a preset sampling rate threshold, so as to ensure the robustness of the two-stage TCN model to high sampling rate audio data noise reduction processing.
- the audio noise reduction method provided by the embodiments of the present disclosure can use the two-stage TCN model to achieve audio noise reduction, especially for the suppression of non-stable noise in the audio, which further improves the noise reduction effect and improves the audio quality. Sound quality improves user experience.
- the present disclosure also provides an audio noise reduction device.
- FIG. 3 it is a schematic structural diagram of an audio noise reduction device provided by an embodiment of the present disclosure.
- the device includes:
- An acquisition module 301 configured to acquire audio data to be denoised
- the first estimation module 302 is configured to estimate the amplitude-time-frequency mask of the audio data to be reduced by using a preset real number network model; wherein, the amplitude-time-frequency mask is used to determine the first-order corresponding to the audio data to be reduced Enhanced amplitude spectrum;
- the second estimation module 303 is configured to estimate the complex time-frequency mask of the audio data to be denoised by using a preset complex network model
- a determining module 304 configured to determine noise reduction result audio data corresponding to the audio data to be reduced based on the first-order enhanced magnitude spectrum corresponding to the audio data to be reduced and the complex time-frequency mask.
- the second estimation module includes:
- the first determining submodule is used to determine the complex spectrum to be denoised; wherein, the complex spectrum to be denoised includes the first-order enhanced magnitude spectrum corresponding to the audio data to be denoised and the original audio data to be denoised A complex spectrum determined by the phase spectrum, or a complex spectrum determined based on the original spectrum of the audio data to be denoised and the original phase spectrum;
- the first processing submodule is configured to input the complex frequency spectrum to be denoised into a preset complex network model, and output the complex time-frequency mask corresponding to the audio data to be denoised after being processed by the preset complex network model .
- the determination module includes:
- a second determining submodule configured to determine an amplitude gain and a phase gain based on the complex time-frequency mask
- the third determination submodule is used to determine the phase enhancement spectrum corresponding to the audio data to be reduced based on the phase gain and the original phase spectrum corresponding to the audio data to be reduced;
- the fourth determining submodule is used to determine the second-order enhanced amplitude spectrum corresponding to the audio data to be reduced based on the amplitude gain and the first-order enhanced amplitude spectrum corresponding to the audio data to be reduced;
- the fifth determining submodule is configured to determine the noise reduction result audio data corresponding to the audio data to be reduced based on the second-order enhanced magnitude spectrum and the phase enhanced spectrum.
- the preset real network model and the preset complex network model are used to form a two-stage time-domain convolutional network TCN model.
- the device further includes:
- the training module is used to train the two-stage TCN model by using audio training samples whose sampling rate is higher than a preset sampling rate threshold.
- the device further includes:
- An augmentation module configured to perform preset data augmentation processing on the audio training samples to obtain augmented audio training samples
- the training module is specifically used for:
- the two-stage TCN model is trained; wherein, the augmented audio training samples have a sampling rate higher than a preset sampling rate threshold.
- the preset data augmentation processing includes performing high-pass, low-pass, band-pass, setting different volumes and/or equalizing the audio training samples according to preset probabilities.
- the training module is specifically used to train the two-stage TCN model by using the time domain loss function SISNR.
- the amplitude time-frequency mask is used to determine the first-order enhanced amplitude spectrum corresponding to the audio data to be reduced, including: the amplitude time-frequency mask is used to match the original frequency spectrum of the audio data to be reduced. The amplitude spectra are multiplied to obtain the first-order enhanced amplitude spectrum corresponding to the audio data to be denoised.
- the audio data to be reduced is obtained, and then the amplitude-time-frequency mask of the audio data to be reduced is estimated by using a preset real number network model, and a value corresponding to the audio data to be reduced can be obtained.
- Order Enhanced Magnitude Spectrum Furthermore, the complex time-frequency mask of the audio data to be denoised is estimated by using the preset complex network model, and the denoising result audio data corresponding to the audio data to be denoised is determined by combining the first-order enhanced amplitude spectrum and the complex time-frequency mask.
- the embodiments of the present disclosure use the preset real number network model to enhance the amplitude spectrum of the audio data to be denoised, and use the preset complex number network model to simultaneously enhance the amplitude spectrum and phase spectrum of the audio data to be denoised. It can be seen that the embodiments of the present disclosure can realize the Noise reduction Noise reduction processing of audio data, so as to better improve the sound quality of audio.
- an embodiment of the present disclosure also provides a computer-readable storage medium, where instructions are stored in the computer-readable storage medium, and when the instructions are run on a terminal device, the terminal device realizes this
- the audio noise reduction method described in the embodiment is disclosed.
- the embodiment of the present disclosure further provides a computer program product, the computer program product includes a computer program/instruction, and when the computer program/instruction is executed by a processor, the audio noise reduction method described in the embodiment of the present disclosure is implemented.
- an embodiment of the present disclosure also provides an audio noise reduction device, as shown in FIG. 4 , which may include:
- Processor 401 memory 402 , input device 403 and output device 404 .
- the number of processors 401 in the audio noise reduction device may be one or more, and one processor is taken as an example in FIG. 4 .
- the processor 401 , the memory 402 , the input device 43 and the output device 404 may be connected through a bus or in other ways, wherein connection through a bus is taken as an example in FIG. 4 .
- the memory 402 can be used to store software programs and modules, and the processor 401 executes various functional applications and data processing of the audio noise reduction device by running the software programs and modules stored in the memory 402 .
- the memory 402 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application program required by at least one function, and the like.
- the memory 402 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one magnetic disk storage device, flash memory device, or other volatile solid-state storage devices.
- the input device 403 can be used to receive input digital or character information, and generate signal input related to user settings and function control of the audio noise reduction device.
- the processor 401 loads the executable files corresponding to the processes of one or more application programs into the memory 402 according to the following instructions, and the processor 401 runs the executable files stored in the memory 402. Application program, so as to realize various functions of the above-mentioned audio noise reduction equipment.
Landscapes
- Engineering & Computer Science (AREA)
- Physics & Mathematics (AREA)
- Acoustics & Sound (AREA)
- Signal Processing (AREA)
- Health & Medical Sciences (AREA)
- Computational Linguistics (AREA)
- Multimedia (AREA)
- Human Computer Interaction (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Quality & Reliability (AREA)
- Theoretical Computer Science (AREA)
- General Health & Medical Sciences (AREA)
- Software Systems (AREA)
- Molecular Biology (AREA)
- Computing Systems (AREA)
- General Engineering & Computer Science (AREA)
- General Physics & Mathematics (AREA)
- Mathematical Physics (AREA)
- Evolutionary Computation (AREA)
- Data Mining & Analysis (AREA)
- Biophysics (AREA)
- Biomedical Technology (AREA)
- Artificial Intelligence (AREA)
- Life Sciences & Earth Sciences (AREA)
- Soundproofing, Sound Blocking, And Sound Damping (AREA)
Abstract
Description
Claims (13)
- 一种音频降噪方法,其中,所述方法包括:获取待降噪音频数据;利用预设实数网络模型估计所述待降噪音频数据的幅度时频掩蔽;其中,所述幅度时频掩蔽用于确定所述待降噪音频数据对应的一阶增强幅度谱;利用预设复数网络模型估计所述待降噪音频数据的复数时频掩蔽;基于所述待降噪音频数据对应的一阶增强幅度谱和所述复数时频掩蔽,确定所述待降噪音频数据对应的降噪结果音频数据。
- 根据权利要求1所述的方法,其中,所述利用预设复数网络模型估计所述待降噪音频数据的复数时频掩蔽,包括:确定待降噪复数频谱;其中,所述待降噪复数频谱包括基于所述待降噪音频数据对应的一阶增强幅度谱和所述待降噪音频数据的原始相位谱确定的复数频谱,或者,基于所述待降噪音频数据的原始频谱和原始相位谱确定的复数频谱;将所述待降噪复数频谱输入至预设复数网络模型,经过所述预设复数网络模型的处理后,输出所述待降噪音频数据对应的复数时频掩蔽。
- 根据权利要求1或2所述的方法,其中,所述基于所述待降噪音频数据对应的一阶增强幅度谱和所述复数时频掩蔽,确定所述待降噪音频数据对应的降噪结果音频数据,包括:基于所述复数时频掩蔽,确定幅度增益和相位增益;基于所述相位增益和所述待降噪音频数据对应的原始相位谱,确定所述待降噪音频数据对应的相位增强谱;以及,基于所述幅度增益和所述待降噪音频数据对应的一阶增强幅度谱,确定所述待降噪音频数据对应的二阶增强幅度谱;基于所述二阶增强幅度谱和所述相位增强谱,确定所述待降噪音频数据对应的降噪结果音频数据。
- 根据权利要求1所述的方法,其中,所述预设实数网络模型和所述预设复数网络模型用于构成双阶段时域卷积网络TCN模型。
- 根据权利要求4所述的方法,其中,所述利用预设实数网络模型估计所述待降噪音频数据的幅度时频掩蔽之前,还包括:利用采样率高于预设采样率阈值的音频训练样本,对所述双阶段TCN模型进行训练。
- 根据权利要求5所述的方法,其中,所述利用采样率高于预设采样率阈值的音频训练样本,对所述双阶段TCN模型进行训练之前,还包括:对所述音频训练样本进行预设数据增广处理,得到增广后音频训练样本;相应的,所述利用采样率高于预设采样率阈值的音频训练样本,对所述双阶段TCN模型进行训练,包括:利用所述增广后音频训练样本,对所述双阶段TCN模型进行训练;其中,所述增广后音频训练样本的采样率高于预设采样率阈值。
- 根据权利要求6所述的方法,其中,所述预设数据增广处理包括按预设的概率对所述音频训练样本进行高通、低通、带通、设置不同音量和/或均衡。
- 根据权利要求5所述的方法,其中,所述对所述双阶段TCN模型进行训练,包括:采用时域损失函数SISNR对双阶段TCN模型进行训练。
- 根据权利要求1所述的方法,其中,所述幅度时频掩蔽用于确定所述待降噪音频数据对应的一阶增强幅度谱,包括:所述幅度时频掩蔽用于与所述待降噪音频数据的原始幅度谱相乘,得到所述待降噪音频数据对应的一阶增强幅度谱。
- 一种音频降噪装置,其中,所述装置包括:获取模块,用于获取待降噪音频数据;第一估计模块,用于利用预设实数网络模型估计所述待降噪音频数据的幅度时频掩蔽;其中,所述幅度时频掩蔽用于确定所述待降噪音频数据对应的一阶增强幅度谱;第二估计模块,用于利用预设复数网络模型估计所述待降噪音频数据的复数时频掩蔽;确定模块,用于基于所述待降噪音频数据对应的一阶增强幅度谱和所述复数时频掩蔽,确定所述待降噪音频数据对应的降噪结果音频数据。
- 一种计算机可读存储介质,其中,所述计算机可读存储介质中存储有指令,当所述指令在终端设备上运行时,使得所述终端设备实现如权利要求1-9任一项所述的方法。
- 一种设备,其包括:存储器,处理器,及存储在所述存储器上并可在所述处理器上运行的计算机程序,所述处理器执行所述计算机程序时,实现如权利要求1-9任一项所述的方法。
- 一种计算机程序产品,其中,所述计算机程序产品包括计算机程序/指令,所述计算机程序/指令被处理器执行时实现如权利要求1-9任一项所述的方法。
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US18/571,119 US20240284100A1 (en) | 2021-09-24 | 2022-09-09 | Audio denoising method and device, apparatus and storage medium |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202111124158.6A CN115862649B (zh) | 2021-09-24 | 2021-09-24 | 一种音频降噪方法、装置、设备及存储介质 |
| CN202111124158.6 | 2021-09-24 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2023045779A1 true WO2023045779A1 (zh) | 2023-03-30 |
Family
ID=85652626
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2022/118040 Ceased WO2023045779A1 (zh) | 2021-09-24 | 2022-09-09 | 一种音频降噪方法、装置、设备及存储介质 |
Country Status (3)
| Country | Link |
|---|---|
| US (1) | US20240284100A1 (zh) |
| CN (1) | CN115862649B (zh) |
| WO (1) | WO2023045779A1 (zh) |
Families Citing this family (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN116469402B (zh) * | 2023-04-23 | 2026-04-24 | 百果园技术(新加坡)有限公司 | 一种音频降噪方法、装置、设备、存储介质及产品 |
| CN117953911B (zh) * | 2024-03-26 | 2024-07-02 | 北京航空航天大学 | 一种飞机模拟器声音降噪方法、系统、设备及介质 |
| CN120529234B (zh) * | 2025-07-24 | 2025-11-14 | 歌尔股份有限公司 | 基于扬声器的抑噪音频生成方法、抑噪音频生成设备及存储介质 |
Citations (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN108735213A (zh) * | 2018-05-29 | 2018-11-02 | 太原理工大学 | 一种基于相位补偿的语音增强方法及系统 |
| CN110739002A (zh) * | 2019-10-16 | 2020-01-31 | 中山大学 | 基于生成对抗网络的复数域语音增强方法、系统及介质 |
| CN110808063A (zh) * | 2019-11-29 | 2020-02-18 | 北京搜狗科技发展有限公司 | 一种语音处理方法、装置和用于处理语音的装置 |
| CN111508514A (zh) * | 2020-04-10 | 2020-08-07 | 江苏科技大学 | 基于补偿相位谱的单通道语音增强算法 |
| US20210012767A1 (en) * | 2020-09-25 | 2021-01-14 | Intel Corporation | Real-time dynamic noise reduction using convolutional networks |
| CN112567458A (zh) * | 2018-08-16 | 2021-03-26 | 三菱电机株式会社 | 音频信号处理系统、音频信号处理方法及计算机可读存储介质 |
Family Cites Families (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US11373672B2 (en) * | 2016-06-14 | 2022-06-28 | The Trustees Of Columbia University In The City Of New York | Systems and methods for speech separation and neural decoding of attentional selection in multi-speaker environments |
| CN108615535B (zh) * | 2018-05-07 | 2020-08-11 | 腾讯科技(深圳)有限公司 | 语音增强方法、装置、智能语音设备和计算机设备 |
| CN109448751B (zh) * | 2018-12-29 | 2021-03-23 | 中国科学院声学研究所 | 一种基于深度学习的双耳语音增强方法 |
| KR20210105688A (ko) * | 2020-02-19 | 2021-08-27 | 라인플러스 주식회사 | 머신러닝 모델을 사용하여 노이즈를 포함하는 입력 음성 신호로부터 노이즈가 제거된 음성 신호를 복원하는 방법 및 장치 |
| CN113314147B (zh) * | 2021-05-26 | 2023-07-25 | 北京达佳互联信息技术有限公司 | 音频处理模型的训练方法及装置、音频处理方法及装置 |
| CN113241088B (zh) * | 2021-07-09 | 2021-10-22 | 北京达佳互联信息技术有限公司 | 语音增强模型的训练方法及装置、语音增强方法及装置 |
-
2021
- 2021-09-24 CN CN202111124158.6A patent/CN115862649B/zh active Active
-
2022
- 2022-09-09 US US18/571,119 patent/US20240284100A1/en active Pending
- 2022-09-09 WO PCT/CN2022/118040 patent/WO2023045779A1/zh not_active Ceased
Patent Citations (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN108735213A (zh) * | 2018-05-29 | 2018-11-02 | 太原理工大学 | 一种基于相位补偿的语音增强方法及系统 |
| CN112567458A (zh) * | 2018-08-16 | 2021-03-26 | 三菱电机株式会社 | 音频信号处理系统、音频信号处理方法及计算机可读存储介质 |
| CN110739002A (zh) * | 2019-10-16 | 2020-01-31 | 中山大学 | 基于生成对抗网络的复数域语音增强方法、系统及介质 |
| CN110808063A (zh) * | 2019-11-29 | 2020-02-18 | 北京搜狗科技发展有限公司 | 一种语音处理方法、装置和用于处理语音的装置 |
| CN111508514A (zh) * | 2020-04-10 | 2020-08-07 | 江苏科技大学 | 基于补偿相位谱的单通道语音增强算法 |
| US20210012767A1 (en) * | 2020-09-25 | 2021-01-14 | Intel Corporation | Real-time dynamic noise reduction using convolutional networks |
Non-Patent Citations (3)
| Title |
|---|
| "Master's Thesis", 27 March 2020, ZHEJIANG UNIVERSITY, China, article LI, BIN: "Single Channel Speech Enhancement Based on Deep Neural Network", pages: 1 - 60, XP009544831, DOI: 10.27461/d.cnki.gzjdx.2020.003246 * |
| LI, WANLING, ZHANG QIU-JU: "Speech Enhancement Based on Joint Maximum A Posteriori Probability", COMPUTER SYSTEMS AND APPLICATIONS, ZHONGGUO KEXUEYUAN RUANJIAN YANJIUSUO, CN, vol. 27, no. 12, 1 January 2018 (2018-01-01), CN , pages 163 - 168, XP093053996, ISSN: 1003-3254, DOI: 10.15888/j.cnki.csa.006670 * |
| ZHENG NAIJUN: "SIGNAL ENHANCEMENT BASED ON COMPLEX-VALUED NEURAL NETWORKS", XIDIAN UNIVERSITY MASTER'S THESES, no. 05, 1 January 2018 (2018-01-01), XP055827314 * |
Also Published As
| Publication number | Publication date |
|---|---|
| CN115862649A (zh) | 2023-03-28 |
| US20240284100A1 (en) | 2024-08-22 |
| CN115862649B (zh) | 2025-07-22 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US10511908B1 (en) | Audio denoising and normalization using image transforming neural network | |
| WO2023045779A1 (zh) | 一种音频降噪方法、装置、设备及存储介质 | |
| CN104685562B (zh) | 用于从嘈杂输入信号中重构目标信号的方法和设备 | |
| CN110164465B (zh) | 一种基于深层循环神经网络的语音增强方法及装置 | |
| JP7667247B2 (ja) | 機械学習を用いたノイズ削減 | |
| CN107113521A (zh) | 用辅助键座麦克风来检测和抑制音频流中的键盘瞬态噪声 | |
| CN113299308B (zh) | 一种语音增强方法、装置、电子设备及存储介质 | |
| CN106558315B (zh) | 异质麦克风自动增益校准方法及系统 | |
| CN116469402B (zh) | 一种音频降噪方法、装置、设备、存储介质及产品 | |
| CN109102821B (zh) | 时延估计方法、系统、存储介质及电子设备 | |
| CN118800268B (zh) | 语音信号处理方法、语音信号处理设备及存储介质 | |
| CN113990343B (zh) | 语音降噪模型的训练方法和装置及语音降噪方法和装置 | |
| CN115171714A (zh) | 一种语音增强方法、装置、电子设备及存储介质 | |
| CN113314147A (zh) | 音频处理模型的训练方法及装置、音频处理方法及装置 | |
| CN114220451A (zh) | 音频消噪方法、电子设备和存储介质 | |
| CN118899005A (zh) | 一种音频信号处理方法、装置、计算机设备及存储介质 | |
| CN107045874B (zh) | 一种基于相关性的非线性语音增强方法 | |
| Steinmetz et al. | High-fidelity noise reduction with differentiable signal processing | |
| Hendriks et al. | MAP estimators for speech enhancement under normal and Rayleigh inverse Gaussian distributions | |
| WO2014132499A1 (ja) | 信号処理装置および方法 | |
| WO2016197629A1 (en) | System and method for frequency estimation | |
| Thiem et al. | Reducing artifacts in GAN audio synthesis | |
| CN113611320A (zh) | 风噪抑制方法、装置、音频设备及系统 | |
| JP7722467B2 (ja) | 信号処理装置、信号処理方法及び信号処理プログラム | |
| CN116312592A (zh) | 语音广播扩声系统的啸叫处理方法和装置 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 22871827 Country of ref document: EP Kind code of ref document: A1 |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 18571119 Country of ref document: US |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 32PN | Ep: public notification in the ep bulletin as address of the adressee cannot be established |
Free format text: NOTING OF LOSS OF RIGHTS PURSUANT TO RULE 112(1) EPC (EPO FORM 1205A DATED 09.07.2024) |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 22871827 Country of ref document: EP Kind code of ref document: A1 |
