WO2021114733A1 - 一种分频段进行处理的噪声抑制方法及其系统 - Google Patents
一种分频段进行处理的噪声抑制方法及其系统 Download PDFInfo
- Publication number
- WO2021114733A1 WO2021114733A1 PCT/CN2020/111672 CN2020111672W WO2021114733A1 WO 2021114733 A1 WO2021114733 A1 WO 2021114733A1 CN 2020111672 W CN2020111672 W CN 2020111672W WO 2021114733 A1 WO2021114733 A1 WO 2021114733A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- frequency
- power spectrum
- value
- frequency point
- signal
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L21/00—Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
- G10L21/02—Speech enhancement, e.g. noise reduction or echo cancellation
- G10L21/0208—Noise filtering
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L21/00—Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
- G10L21/02—Speech enhancement, e.g. noise reduction or echo cancellation
- G10L21/0208—Noise filtering
- G10L21/0216—Noise filtering characterised by the method used for estimating noise
- G10L21/0232—Processing in the frequency domain
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04M—TELEPHONIC COMMUNICATION
- H04M1/00—Substation equipment, e.g. for use by subscribers
- H04M1/02—Constructional features of telephone sets
- H04M1/19—Arrangements of transmitters, receivers, or complete sets to prevent eavesdropping, to attenuate local noise or to prevent undesired transmission; Mouthpieces or receivers specially adapted therefor
Definitions
- This application relates to the technical field of voice calls, and in particular to a noise suppression method and system for sub-band processing.
- voice calls have gradually replaced handwritten letters. Compared with text communication, voice calls can convey information faster and convey more accurate information.
- the communicator can use the tone, intonation, and tone of the other party. Perceive the emotions of the other party through wording, etc., thereby improving the efficiency of communication.
- a microphone is generally used to collect voice data.
- the voice data is an analog signal, and then the voice data of the analog signal is converted into a digital signal, which is transmitted to the other party in a wired or wireless manner.
- the other party receives the digital signal After that, the digital signal is converted into an analog signal for playback.
- noise data is generally the environmental noise of the user during a call.
- the noise data will be transmitted to the other end of the call along with the user's voice data. ;
- noise will affect the quality of the call; in severe cases, the noise will cause the call obstacles and make the voice conveyed during the call deviate; how to accurately distinguish the voice and noise during the call, and effectively suppress the spread of noise, and improve the call Quality becomes a big problem.
- the embodiments of the present application disclose a noise suppression method and system for sub-band processing, which are used for accurately distinguishing voice and noise during a call, and effectively suppress the propagation of noise, and improve the quality of the call.
- the embodiment of the present application discloses a noise suppression method for sub-band processing, which is used to suppress the noise in the call sound.
- the method includes the following steps: Step S1, the signal type of the collected sound data is changed from the time domain The signal is transformed into a frequency domain signal; step S2, the frequency domain signal is divided into a low frequency signal and a medium and high frequency signal, and the low frequency noise power spectrum estimated value is calculated according to the amplitude of the low frequency signal, and the medium and high frequency signal is calculated according to the amplitude of the medium and high frequency signal Noise power spectrum estimated value; Step S3, calculate according to the low-frequency noise power spectrum estimated value, the mid-high frequency noise power spectrum estimated value and the frequency domain signal to obtain the gain frequency domain signal; Step S4, transform the gain frequency domain signal It is the time domain signal after gain.
- the embodiment of the present application also discloses a noise suppression system for sub-band processing.
- the sub-band noise suppression system includes: a signal conversion module for converting the signal type of the collected sound data from a time-domain signal to Frequency domain signal; noise power spectrum estimation module, used to divide the frequency domain signal into low frequency signal and medium and high frequency signal, and calculate the low frequency noise power spectrum estimation value according to the amplitude of the low frequency signal, and calculate according to the amplitude of the medium and high frequency signal Mid and high frequency noise power spectrum estimation value; gain frequency domain calculation module, used to calculate according to the low frequency noise power spectrum estimation value, mid and high frequency noise power spectrum estimation value and frequency domain signal to obtain the frequency domain signal after gain; signal inverse transformation module , Used to transform the frequency domain signal after gain into the time domain signal after gain.
- the noise suppression method and system for sub-band processing of the present application divide the frequency domain signal into a low-frequency signal and a medium-high frequency signal, and process the low-frequency signal to obtain a low-frequency noise power spectrum estimated value, and process the medium and high-frequency signal to obtain Mid-high frequency noise power spectrum estimation value, the frequency domain signal is calculated according to the low frequency noise power spectrum estimation value and mid-high frequency noise power spectrum estimation value to obtain the frequency domain signal after gain; according to the signal characteristics of different frequencies, different noise processing methods are adopted , Divide noise and voice more accurately, effectively suppress the spread of noise, improve the efficiency of voice calls, and enhance the effect of voice calls.
- Fig. 1 is a flowchart of a noise suppression method for sub-band processing in an embodiment of the present application.
- Fig. 2 is a sub-flow chart of step S1 in Fig. 1.
- Fig. 3 is a sub-flow chart of step S2 in Fig. 1.
- Fig. 4 is a sub-flow chart of step S22 in Fig. 3.
- Fig. 5 is a sub-flow chart of step S223 in Fig. 4.
- Fig. 6 is a sub-flow chart of step S23 in Fig. 3.
- Fig. 7 is a sub-flow chart of step S3 in Fig. 1.
- Fig. 8 is a sub-flow chart of step S33 in Fig. 7.
- FIG. 9 is a structural relationship diagram of a noise suppression system that performs processing by frequency bands in an embodiment of the present application.
- FIG. 10 is a schematic diagram of the structure of the signal conversion module in FIG. 9.
- FIG. 11 is a schematic diagram of the structure of the noise power spectrum estimation module in FIG. 9.
- FIG. 12 is a schematic diagram of the structure of the low-frequency power spectrum estimation module in FIG. 11.
- FIG. 13 is a schematic diagram of the structure of the power spectrum estimation module for non-fundamental frequency points in FIG. 12.
- FIG. 14 is a schematic diagram of the structure of the middle and high frequency power spectrum estimation module in FIG. 11.
- FIG. 15 is a schematic diagram of the structure of the gain frequency domain calculation module in FIG. 9.
- FIG. 16 is a schematic diagram of the structure of the gain signal calculation module in FIG. 15.
- FIG. 1 is a flowchart of a noise suppression method for sub-band processing in an embodiment of the present application.
- the present application provides a noise suppression method for sub-band processing, which is used to suppress noise in the voice of a call.
- the method includes the following steps: Step S1: The type is transformed from the time domain signal to the frequency domain signal; step S2, the frequency domain signal is divided into low frequency signal and medium and high frequency signal, and the low frequency noise power spectrum estimation value is calculated according to the amplitude of the low frequency signal, and according to the amplitude of the medium and high frequency signal Calculate the estimated value of the mid and high frequency noise power spectrum; step S3, calculate according to the estimated value of the low frequency noise power spectrum, the estimated value of the mid and high frequency noise power spectrum, and the frequency domain signal to obtain the frequency domain signal after gain; step S4, the gain The frequency domain signal is transformed into a time domain signal after gain.
- the noise suppression method for sub-band processing divides the frequency domain signal into a low-frequency signal and a medium-high frequency signal, and processes the low-frequency signal to obtain a low-frequency noise power spectrum estimation value, and performs processing on the medium and high-frequency signal.
- Process to obtain the estimated value of the mid and high frequency noise power spectrum calculate the frequency domain signal according to the estimated value of the low frequency noise power spectrum and the mid and high frequency noise power spectrum estimation value to obtain the frequency domain signal after gain; use different noise according to the signal characteristics of different frequencies
- the processing method divides noise and voice more accurately, effectively suppresses the spread of noise, improves the efficiency of voice calls, and enhances the effect of voice calls.
- FIG. 2 is a schematic diagram of a processing flow of a full-frequency signal in an embodiment of the present application.
- the step S1 includes: step S11, acquiring the collected sound data whose signal type is a time-domain signal; step S12, framing the time-domain signal at a preset time interval Obtain multi-frame time-domain signals; step S13, obtain frequency-domain signals by performing time-frequency transformation on each frame of time-domain signals.
- the sound data includes the voice data and environmental sound data sent by the user when the user is using the call device; the voice data is generally collected and converted through a microphone, and the microphone collects the user's voice data.
- the analog signal After the analog signal of the emitted voice and environmental sound, the analog signal is converted into a digital signal through analog-to-digital conversion; the sound data here is a digital signal after the analog-to-digital conversion.
- the time domain signal is a signal obtained by arranging the sound data in the order of collection time; the time domain is used to describe the relationship between the physical quantity and time.
- the time-domain waveform of the time-domain signal is used to express the change of the signal over time.
- the frequency domain signal is a signal obtained by sorting the sound data in order of frequency.
- the frequency domain is used to describe the relationship between physical quantity and frequency.
- the frequency domain waveform of the frequency domain signal is used to express the change of the amplitude of the signal with frequency.
- the frequency domain signal is composed of multiple frequency points, and each frequency point contains corresponding amplitude information and phase information.
- the framing refers to combining the collected sound data in the time domain signal according to a preset time interval, and combining the time domain signals within the preset time interval as a frame, and the time domain signals are combined
- the latter is a multi-frame time-domain signal, and the time-frequency signal will be subjected to subsequent transformations in units of frames.
- the time-frequency transformation in this embodiment refers to the Fourier transformation of each frame in the time domain signal, and the amplitude and phase information of the corresponding multiple frequency points are obtained after each frame is Fourier transformed.
- the frequency domain signal is obtained after summarizing all frequency points; in other embodiments, other transformation methods may also be used to convert the time domain signal into a frequency domain signal.
- Fig. 3 is a sub-flow chart of step S2 in Fig. 1.
- the step S2 includes: step S21, dividing the frequency domain signal into a low-frequency signal and a medium-high frequency signal according to a preset frequency division standard; step S22: The single-channel low-frequency signal collected by the single-channel microphone is used to estimate the noise power spectrum density to obtain the noise power spectrum estimation value of all low-frequency points in the single-channel low-frequency signal; step S23, the double-channel mid-high collected by the two-channel microphone in the frequency domain signal The noise power spectrum density is estimated for the frequency signal, and the noise power spectrum estimation value of all the middle and high frequency points in the two-way middle and high frequency signal is obtained.
- the preset frequency division standard refers to a preset standard for dividing frequency domain signals; in this embodiment, the frequency domain signal can be divided into low-frequency signals according to the preset frequency division standard , Medium and high frequency signals; in this embodiment, the preset frequency division standard may be a certain fixed frequency; for example, the preset frequency division standard may be to distinguish low frequency, medium and high frequency signals based on a preset frequency, for example, 1000 Hz Standards.
- the low-frequency signal refers to a signal segment lower than the preset frequency in the frequency domain signal.
- a signal segment lower than the preset frequency in the frequency domain signal belongs to a low frequency signal.
- the frequency points in the low-frequency signal are low-frequency frequency points.
- the medium and high frequency signal refers to a signal segment of the frequency domain signal that is not lower than the preset frequency.
- a signal segment equal to or higher than the preset frequency in the frequency domain signal belongs to a medium and high frequency signal.
- the frequency points in the medium and high frequency signal are medium and high frequency frequency points.
- the microphone refers to a component used to collect sound data in the call device.
- a frequency domain signal of a microphone with a high signal-to-noise ratio is selected for noise suppression processing.
- the power spectral density is sometimes called spectral power distribution (spectral power distribution, SPD), which is the Fourier transform of the autocorrelation function of the signal (noise), that is, the signal per unit frequency (noise) )
- SPD spectral power distribution
- Power spectral density is a statistical method of probability and a measure of the mean square value of random variables.
- the estimated value of the noise power spectrum that is, the estimated value of the power spectrum density of the noise, is also called power spectrum estimation.
- the power spectral density of noise is used to describe the relationship between the energy characteristics of noise and the frequency.
- the signal segment of the frequency domain signal whose frequency is lower than the preset frequency division standard is regarded as a low frequency band according to the preset frequency division standard, which is equal to or higher than the preset frequency division
- the standard signal segment is used as the middle and high frequency band.
- the low frequency band is used as a low frequency signal, and noise power spectrum density estimation is performed on the low frequency signal to obtain noise power spectrum estimation values of all low frequency points in the low frequency signal.
- the mid- and high-frequency bands of the two frequency domain signals will be selected as mid- and high-frequency signals at the same time, and the noise power spectral density of the mid- and high-frequency signals will be estimated to obtain all mid- and high-frequency frequencies in the mid- and high-frequency signals.
- the estimated value of the noise power spectrum of the point is the value of the noise power spectrum of the point.
- the preset frequency division standard may be a fixed frequency band; the two frequency domain signals are divided more carefully according to the frequency band; for example, 0-1000 Hz is the low frequency band, and 1000-3000 Hz is the mid frequency band. Above 3000 Hz is the high frequency band.
- two channels of sound data are collected through two microphones, and then the two channels of sound data are respectively time-frequency transformed to obtain two channels of frequency domain signals.
- the low-frequency signal and the medium-high frequency signal are selected respectively, and then Perform noise power spectral density estimation on low-frequency signals and mid-high-frequency signals, use the accuracy of noise estimation in low-frequency signals and the correlation of noise estimation in mid- and high-frequency signals, and process the frequency domain signals in segments, making full use of the two collected channels
- the sound data improves the accuracy of the estimated value of the noise power spectrum and removes the residual noise as much as possible.
- FIG. 4 is a sub-flow chart of step S22 in FIG. 3.
- the step S22 includes: step S221, squaring the amplitude of each low-frequency frequency point in the single-channel low-frequency signal to obtain the square value of the amplitude of each low-frequency frequency point; step S222, by detecting the pitch The low-frequency frequency points are divided into fundamental frequency points and non-fundamental frequency points; step S223: calculate the noise power spectrum estimation value of the non-fundamental frequency points according to the square value of the amplitude of the non-fundamental frequency points; step S224, according to the fundamental frequency point The square value of the amplitude and the estimated value of the noise power spectrum of the non-fundamental frequency point are calculated to obtain the estimated value of the noise power spectrum of the fundamental frequency point; step S225, the noise power spectrum estimated value of the non-fundamental frequency point is compared with the noise power of the fundamental frequency point The spectrum estimation value combination obtains the noise power spectrum estimation value of all low frequency frequency points.
- the pitch refers to the period of vocal cord vibration when a person emits a voiced sound.
- the estimation of the pitch period is called pitch detection. Its purpose is to extract the trajectory curve of the pitch period change that is consistent with or as close as possible to the vibration frequency of the human vocal cord. That is, the pitch frequency, which is one of the most important characteristic parameters in speech signal processing.
- the pitch detection method is a method for detecting the pitch signal. Since the voice signal can be regarded as a dynamic non-stationary random process, the frequency variation range of the voice waveform and vocal cord vibration is large and very complicated, so the pitch detection method includes many kinds Algorithms, such as overtone inner product spectrum method, cepstrum analysis method, maximum likelihood estimation method, etc.; in this embodiment, the cepstrum method is used to detect the low-frequency frequency points, and the pitch frequency in the pitch detection can be calculated by the following calculation method Make sure:
- FT and FT -1 denote Fourier transform and inverse Fourier transform, respectively. Since the time domain signal x(n) is obtained by the glottal pulse excitation u(n) filtered by the channel response v(n), that is
- the fundamental frequency point refers to the frequency point at which the square of the amplitude in the low-frequency signal conforms to the fundamental frequency.
- the non-fundamental frequency point refers to a frequency point where the square of the amplitude of the low-frequency signal does not meet the pitch frequency.
- the estimated value of the noise power spectrum at the fundamental frequency point refers to an estimated value of the power spectrum density of the noise at the fundamental frequency point.
- the estimated value of the noise power spectrum at the non-fundamental frequency point refers to an estimated value of the power spectrum density of the noise at the non-fundamental frequency point.
- the estimated value of the noise power spectrum can be calculated by the following formula:
- M represents the set of fundamental frequency points.
- 2 is the original input, that is, the estimated value of the noise power spectrum of the non-fundamental frequency point is the square of the amplitude of the non-fundamental frequency point;
- the estimated value of the noise power spectrum at the fundamental frequency point is to obtain the estimated value of the noise power spectrum at the fundamental frequency point through interpolation of the estimated values of the noise power spectrum at two adjacent non-fundamental frequency points.
- the amplitude of each low-frequency frequency point in the single-channel low-frequency signal is squared to obtain the square value of the amplitude of each low-frequency frequency point.
- All frequency points in the low-frequency signal are screened according to the pitch detection method, and the square of the amplitude of the frequency points in the low-frequency signal is selected as the fundamental frequency point, and the frequency points that do not meet the pitch frequency are regarded as the non-fundamental frequency. point.
- the fundamental frequency point noise power spectrum estimation value and the non-fundamental frequency point noise power spectrum estimation value are calculated and obtained .
- FIG. 5 is a sub-flow chart of step S223 in FIG. 4.
- the step S223 includes: step S2231, calculating according to the square value of the amplitude of each low-frequency frequency point to obtain the speech existence probability value of each low-frequency frequency point; step S2232, according to each non-fundamental frequency point The square value of the amplitude of, is calculated to obtain the preliminary estimated value of the noise power spectrum of each non-fundamental frequency point; step S2233, in the speech existence probability value of all low-frequency frequency points, find the non-fundamental frequency point corresponding to the non-fundamental frequency point Speech existence probability value; Step S2234, calculate the noise power spectrum estimation value of each non-fundamental frequency point according to the speech existence probability value of each non-fundamental frequency point and the corresponding preliminary noise power spectrum estimation value of the corresponding non-fundamental frequency point.
- the voice refers to the voice of the user's speech in the voice data, the voice data includes voice data and noise data, and the voice data is data that needs to be communicated during a call.
- the speech existence probability value of the low frequency frequency point refers to the possibility of speech existence in the low frequency frequency point.
- the speech existence probability value can be calculated by the following formula:
- q(k, ⁇ ) represents the probability of no speech, which can be calculated by comparing the square of the amplitude spectrum of the corresponding frequency point with a preset threshold; ⁇ (k, ⁇ ) is the prior signal-to-noise ratio; ⁇ (k, ⁇ ) can be calculated by the definition of a posteriori SNR and a priori SNR.
- the a priori SNR is the power of the pure speech signal divided by the power of the noise signal.
- the posterior signal-to-noise ratio is the power of the noisy speech signal divided by the power of the noise signal.
- the preliminary estimated value of the noise power spectrum at the non-fundamental frequency point is a value obtained by preliminary calculation of the estimated value of the noise power spectrum at the non-fundamental frequency point by using the above formula.
- the square value of the amplitude of each low-frequency frequency point in the low-frequency signal is brought into the above-mentioned speech existence probability formula for calculation, and the speech existence probability values of all low-frequency frequency points in the low-frequency signal are obtained.
- the above-mentioned formula of noise power spectrum estimation value is brought into calculation to obtain a preliminary estimation value of the noise power spectrum of each non-fundamental frequency point.
- the speech existence probability values of all low-frequency frequency points and the non-fundamental frequency points in the aforementioned fundamental tone detection results are correspondingly searched for the speech existence probability value of each non-fundamental frequency point.
- the speech existence probability value of each non-fundamental frequency point is multiplied by the corresponding preliminary noise power spectrum estimation value of the corresponding non-fundamental frequency point, and the noise power spectrum estimation value of each non-fundamental frequency point is calculated.
- the low-frequency frequency points in the low-frequency signal are divided into fundamental frequency points and non-fundamental frequency points, and then different calculation methods are used for the fundamental frequency points and non-fundamental frequency points to calculate the noise power spectrum of the fundamental frequency points in the low-frequency signal.
- the estimated value and the estimated value of the noise power spectrum at the non-fundamental frequency point; the introduction of the pitch detection method ensures the quality of the call voice, improves the voice call effect in a stable and non-stationary noise environment, and makes the distinction of noise more accurate , Improve the efficiency of noise suppression.
- FIG. 6 is a sub-flow chart of step S23 in FIG. 3.
- step S231 is to obtain the square value of the amplitude of each medium and high frequency point by squaring the amplitude of each frequency point in the dual-channel medium and high frequency signal; step S232, according to the amplitude of each medium and high frequency point The square value of the value is calculated to obtain the self-power spectrum value of each middle and high frequency frequency point, and the cross power spectrum value of each middle and high frequency frequency point; step S233, according to the square value of the amplitude of each middle and high frequency frequency point, the self-power spectrum value and The cross power spectrum value is calculated to obtain the correlation value of each medium and high frequency frequency point; step S234, the noise power spectrum of each medium and high frequency frequency point is calculated according to the correlation value, the self power spectrum value and the cross power spectrum value of each medium and high frequency frequency point Step S235, calculate the speech existence probability of each medium and high frequency point according to the correlation value of each medium and high frequency point; Step S236, by comparing the preliminary estimated value of the noise power spectrum of each medium and high frequency point with
- the self-power spectrum value and the cross-power spectrum value are used to reflect the internal relationship between the random signal itself and other signals at different moments expressed by the correlation function in the time domain, and to understand the similarity of waveforms between the same random sample at different moments.
- the calculation formulas of the self-power spectrum and the cross-power spectrum value are as follows:
- ⁇ represents the number of frames
- ⁇ represents the frequency point
- ⁇ s is the smoothing coefficient
- Is the self-power spectrum and cross-power spectrum value of the two microphones
- Xi, Xj are the amplitude of the frequency domain signal
- * represents the complex conjugate.
- the correlation value includes the correlation value of the voice signal in the two microphones and the correlation value of the noise signal in the two microphone signals.
- the correlation value can be calculated according to the following formula:
- ⁇ x, ⁇ is the correlation of the signals received by the two microphones; Is the estimation of noise correlation; ⁇ s, cor is the preliminary estimation of speech correlation; Is the estimation of the smoothed speech correlation; ⁇ is the posterior signal-to-noise ratio; ⁇ ⁇ is the smoothing coefficient.
- the speech existence probability of the middle and high frequency points refers to the possibility of the existence of a speech signal in each middle and high frequency frequency point.
- the speech existence probability of the medium and high frequency points can be calculated by the following formula:
- the smoothing processing refers to bringing the estimated value of the noise power spectrum of each middle and high frequency point and the speech existence probability of the corresponding frequency point into a formula to calculate the smoothed noise power spectrum estimation value of the middle and high frequency point;
- the preliminary estimated value of the noise power spectrum of the mid- and high-frequency frequency points refers to the estimated value of the noise power spectrum of the mid- and high-frequency frequency points obtained by preliminary calculation according to the formula; the calculation formula of the preliminary estimated value of the noise power spectrum of the mid- and high-frequency frequency points is as follows:
- the mid- and high-frequency signals in the dual frequency domain signals are simultaneously acquired, and the amplitude of each frequency point in the mid- and high-frequency signal is squared to obtain the square value of the amplitude of each mid- and high-frequency frequency point.
- the square value of the amplitude, the self-power spectrum value, and the cross-power spectrum value of each middle and high frequency frequency point are brought into the above-mentioned correlation value calculation formula, and the correlation value of each middle and high frequency frequency point is calculated.
- the calculation formula of the preliminary estimated value of the noise power spectrum of the mid-high frequency point is brought into the calculation formula, and the noise power of each mid-high frequency point is calculated. Preliminary estimate of the spectrum.
- the calculation formula of the speech existence probability of the medium and high frequency points is brought into the calculation formula, and the speech existence probability of each medium and high frequency point is calculated.
- the noise suppression processing is performed on the frequency points of the mid- and high-frequency signals to obtain the The estimated value of the noise power spectrum at each frequency point in the mid- and high-frequency signal.
- FIG. 7 is a sub-flow chart of step S3 in FIG. 1.
- the step S3 includes: step S31, combining the low-frequency noise power spectrum estimation value and the mid-high frequency noise power spectrum estimation value to obtain the full-band noise power spectrum estimation value; step S32, according to The noise power spectrum estimation value of each frequency point in the full-band noise power spectrum estimation value is gain calculation to obtain the amplitude spectrum gain value of each frequency point; step S33, according to the amplitude spectrum gain value of each frequency point and the frequency domain signal Calculate the frequency domain signal after gain.
- the estimated value of the full-band noise power spectrum refers to the estimated value of the noise power spectrum of all frequency points in the frequency domain signal.
- the amplitude spectrum gain value refers to the gain ratio of the amplitude of each frequency point obtained by bringing the estimated value of the noise power spectrum of the frequency point into the amplitude gain function.
- the frequency domain signal after gain refers to the amplitude of each frequency point obtained by gaining each frequency point in the original frequency domain signal according to the amplitude spectrum gain value of the corresponding frequency point.
- the original frequency domain signal refers to the frequency domain signal obtained after the above-mentioned time-frequency transformation.
- the low-frequency noise power spectrum and the mid-high frequency noise power spectrum are combined and summarized to obtain the noise power spectrum estimation values of all frequency points in the frequency domain signal.
- the noise power spectrum estimation value of each frequency point in the noise power spectrum estimation value of all frequency points is brought into the gain function for calculation, and the amplitude spectrum gain value of each frequency point is obtained.
- FIG. 8 is a sub-flow chart of step S33 in FIG. 7.
- the step S33 includes: step S331, obtaining the amplitude spectrum gain value of each frequency point and the amplitude value of the corresponding frequency point in the frequency domain signal; step S332, combining each frequency The amplitude spectrum gain value of a point is multiplied by the amplitude value of the corresponding frequency point to obtain the gain amplitude value of each frequency point, and the gain amplitude values of all frequency points in the frequency domain signal are combined to obtain the gain frequency domain signal.
- the amplitude value of each frequency point corresponding to the original frequency domain signal is obtained.
- the amplitude spectrum gain value of each frequency point is calculated according to the noise power spectrum estimation value of each frequency point in the frequency domain signal, and then the amplitude spectrum gain value of each amplitude spectrum gain value is compared with the amplitude value of the corresponding frequency point in the original frequency domain signal. Multiply, calculate the gain frequency domain signal, improve the recognizability of speech, and carry out noise suppression processing for each frequency point, improve the efficiency of noise suppression, make the voice data in the call clearer and suppress The dissemination of noise data improves the quality of calls.
- the step S4 includes: performing an inverse time-frequency transform on the gain amplitude of each frequency point in the gain frequency domain signal to obtain the gain time domain signal.
- the inverse time-frequency transform refers to transforming a frequency domain signal into a time domain signal, and transforming the entire frequency domain signal into a time domain signal by performing an inverse Fourier transform on each frequency point.
- the frequency domain signal after gain is subjected to inverse Fourier transform according to the frame to obtain the time domain signal of the corresponding frame, and the time domain signal of all frames is combined to obtain the time domain signal after gain;
- the frequency domain signal is transformed into a time domain signal after gain.
- the noise suppression method for sub-band processing can be implemented in hardware or firmware, or can be stored in computer-readable storage media such as CD, ROM, RAM, floppy disk, hard disk, or magneto-optical disk.
- the suppression method can be represented by a general-purpose computer or a special processor or in programmable or dedicated hardware such as ASIC or FPGA as software stored on a recording medium.
- a computer, processor, microprocessor, controller, or programmable hardware includes memory components, such as RAM, ROM, flash memory, etc., when the computer, processor, or hardware implements a sub-band performance described herein
- the memory component can store or receive the software or computer code.
- a general-purpose computer accesses code for implementing the processing shown here, the execution of the code converts the general-purpose computer into a dedicated computer for executing the processing shown here.
- the computer-readable storage medium may be a solid-state memory, a memory card, an optical disc, and the like.
- the computer-readable storage medium stores program instructions and is called by the computer to execute the noise suppression method for sub-band processing shown in FIGS. 1 to 8.
- FIG. 9 is a structural relationship diagram of a noise suppression system 100 for sub-band processing in an embodiment of the present application.
- the noise suppression system 100 for sub-band processing includes: a signal conversion module 10 for converting the signal type of the collected sound data from a time domain signal to a frequency domain signal; a noise power spectrum estimation module 20. It is used to divide the frequency domain signal into low frequency signal and medium and high frequency signal, and calculate the low frequency noise power spectrum estimation value according to the amplitude of the low frequency signal, and calculate the medium and high frequency noise power spectrum estimation value according to the amplitude of the medium and high frequency signal;
- the gain frequency domain calculation module 30 is used for calculating according to the low frequency noise power spectrum estimation value, the mid and high frequency noise power spectrum estimation value and the frequency domain signal to obtain the frequency domain signal after gain;
- the signal inverse transform module 40 is used for converting the gain The frequency domain signal is transformed into a time domain signal after gain.
- FIG. 10 is a schematic structural diagram of the signal conversion module 10 in FIG. 9.
- the signal conversion module 10 includes: a signal acquisition module 11, which is used to acquire the collected sound data whose signal type is a time-domain signal; a signal framing module 12, which is used to set the time-domain signal according to a preset The time interval framing is performed to obtain a multi-frame time domain signal; the time-frequency transformation module 13 is used to obtain a frequency domain signal by performing time-frequency transformation on each frame of the time domain signal.
- FIG. 11 is a schematic structural diagram of the noise power spectrum estimation module 20 in FIG. 9.
- the noise power spectrum estimation module 20 includes: a signal division module 21 for dividing the frequency domain signal into low-frequency signals and medium-high frequency signals according to a preset frequency division standard; a low-frequency power spectrum estimation module 22. It is used to estimate the noise power spectrum density of a single low frequency signal collected by a single microphone in the frequency domain signal to obtain the noise power spectrum estimation value of all low frequency points in the single low frequency signal; the middle and high frequency power spectrum estimation module 23, It is used to estimate the noise power spectrum density of the two-channel medium and high frequency signals collected by the two-channel microphone in the frequency domain signal, and obtain the noise power spectrum estimation value of all the medium and high frequency points in the two-channel medium and high frequency signal.
- FIG. 12 is a schematic structural diagram of the low-frequency power spectrum estimation module 22 in FIG. 11.
- the low-frequency power spectrum estimation module 22 includes: a low-frequency amplitude squaring module 221, configured to square the amplitude of each low-frequency frequency point in a single low-frequency signal to obtain the amplitude of each low-frequency frequency point Pitch detection module 222, used to divide low-frequency frequency points into fundamental frequency points and non-fundamental frequency points through the pitch detection method; non-fundamental frequency point power spectrum estimation module 223, used to base the amplitude of non-fundamental frequency points The square value of is calculated to obtain the estimated value of the noise power spectrum of the non-fundamental frequency point; the fundamental frequency point power spectrum estimation module 224 is used to calculate the estimated value of the noise power spectrum of the non-fundamental frequency point according to the square value of the amplitude of the fundamental frequency point The estimated value of the noise power spectrum at the fundamental frequency point; the low-frequency power spectrum combination module 225 is used to combine the estimated value of the noise power spectrum at the non-fundamental frequency point with the estimated value of the noise power spectrum at the
- FIG. 13 is a schematic structural diagram of the non-fundamental frequency point power spectrum estimation module 223 in FIG. 12.
- the non-fundamental frequency point power spectrum estimation module 223 includes: a low-frequency speech probability calculation module 2231, configured to calculate the presence of speech at each low-frequency frequency point according to the square value of the amplitude of each low-frequency frequency point Probability value; non-fundamental frequency point power spectrum preliminary estimation module 2232, used to calculate the preliminary noise power spectrum estimation value of each non-fundamental frequency point according to the square value of the amplitude of each non-fundamental frequency point; non-fundamental frequency point speech
- the probability search module 2233 is used to find the voice existence probability value of the non-fundamental frequency point corresponding to the non-fundamental frequency point among the speech existence probability values of all low-frequency frequency points;
- the non-fundamental frequency point power spectrum calculation module 2234 is used to calculate according to The speech existence probability value of each non-fundamental frequency point and the preliminary estimation value of the noise power spectrum of the corresponding non-fundamental frequency point are calculated to obtain the noise power spectrum estimation value of each non-fundamental frequency point.
- FIG. 14 is a schematic diagram of the structure of the middle and high frequency power spectrum estimation module 23 in FIG. 11.
- the mid and high frequency power spectrum estimation module 23 includes: a mid and high frequency amplitude squaring module 231, configured to obtain each mid and high frequency frequency point by squaring the amplitude of each frequency point in the dual-channel mid and high frequency signal The square value of the amplitude; the power spectrum value calculation module 232 is used to calculate the self-power spectrum value of each medium and high frequency frequency point and the mutual power of each medium and high frequency frequency point according to the square value of the amplitude value of each medium and high frequency frequency point Spectrum value; correlation value calculation module 233, used to calculate the correlation value of each middle and high frequency frequency point according to the square value of the amplitude of each middle and high frequency frequency point, the self-power spectrum value and the cross power spectrum value; The module 234 is used to calculate the preliminary estimated value of the noise power spectrum of each middle and high frequency frequency point according to the correlation value, the self-power spectrum value and the cross power spectrum value of each middle and high frequency frequency point; the middle and high frequency speech probability calculation module 235 uses In accordance
- FIG. 15 is a schematic structural diagram of the gain frequency domain calculation module 30 in FIG. 9.
- the gain frequency domain calculation module 30 includes: a full-band power spectrum combining module 31, configured to combine the low-frequency noise power spectrum estimation value with the mid-high frequency noise power spectrum estimation value to obtain the full-band noise power spectrum estimation value ; Amplitude spectrum gain calculation module 32, used for gain calculation according to the noise power spectrum estimated value of each frequency point in the full-band noise power spectrum estimated value, to obtain the amplitude spectrum gain value of each frequency point; gain signal calculation module 33, It is used to calculate the frequency domain signal after gain according to the amplitude spectrum gain value of each frequency point and the frequency domain signal.
- FIG. 16 is a schematic structural diagram of the gain signal calculation module 33 in FIG. 15.
- the gain signal calculation module 33 includes: an amplitude spectrum gain value obtaining module 331, configured to obtain the amplitude spectrum gain value of each frequency point and the amplitude value of the corresponding frequency point in the frequency domain signal; the gain amplitude value The calculation module 332 is configured to multiply the amplitude spectrum gain value of each frequency point by the amplitude value of the corresponding frequency point to obtain the gain amplitude value of each frequency point, and calculate the gain amplitude value of all frequency points in the frequency domain signal. Value combination obtains the frequency domain signal after gain.
- the signal inverse transform module 40 is further configured to perform inverse time-frequency transform on the gain amplitude value of each frequency point in the gain frequency domain signal to obtain the gain time domain signal.
- the noise suppression system 100 for sub-band processing further includes a central processing unit (Central Processing Unit, CPU), and may also be other general-purpose processors, digital signal processors (Digital Signal Processors, DSPs), and application specific integrated circuits (Application Specific Integrated Circuits). Integrated Circuit, ASIC), Field-Programmable Gate Array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc.
- the general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.
- the processing unit is the data processing center of the noise suppression system that performs processing in the sub-bands, and connects the entire plant by wired or wireless lines. Each module of the noise suppression system 100 that performs processing in sub-bands is described. Used to process the data sent from each module.
- the noise suppression system 100 for sub-band processing further includes a storage module 50, and the storage module 50 is configured to store the time domain signal and the frequency domain signal.
- the storage module 50 may include a high-speed random access memory, and may also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), and a secure digital (Secure Digital). , SD) card, flash memory card (Flash Card), multiple disk storage devices, flash memory devices, or other volatile solid-state storage devices.
- a non-volatile memory such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), and a secure digital (Secure Digital).
- SD Secure Digital
- flash memory card Flash Card
- the storage module 50 is located in the communication device, and is mainly connected to the above-mentioned processing unit, and is used to store the sound data collected by the microphone and the frequency domain signal, time domain signal, and noise power spectrum estimation processed by the processing unit. Value etc.
- the storage module 50 may also be connected to each module of the noise suppression system 100 that performs processing by frequency bands, and is used to store data generated by each module in the parameter calculation process.
- the noise suppression system for sub-band processing includes a call device, the call device may be a mobile terminal, the signal conversion module 10 and the noise power spectrum estimation module 20 , The gain frequency domain calculation module 30, the signal inverse transform module 40, the storage module 50, etc., are all set in the host device; the signal transform module 10 and the noise power spectrum estimation module 20 are connected in a wired or wireless manner, The frequency domain signal after the time-frequency transformation is transmitted to the noise power spectrum estimation module 20; the noise power spectrum estimation module 20 and the gain frequency domain calculation module 30 are connected in a wired or wireless manner, and the calculated The low-frequency noise power spectrum estimated value and the medium-high frequency noise power spectrum estimated value are transmitted to the gain frequency domain calculation module 30; the gain frequency domain calculation module 30 and the signal inverse transform module 40 are connected in a wired or wireless manner.
- the gain frequency domain signal is transmitted to the signal inverse transform module 40, and the signal inverse transform module 40 inversely transforms the gain frequency domain signal into a time domain signal; the storage module 50 and the signal transform module 10, noise
- the power spectrum estimation module 20, the gain frequency domain calculation module 30, and the signal inverse transform module 40 are all connected in a wired or wireless manner, and can store time domain signals, frequency domain signals, low-frequency noise power spectrum estimates, and mid- and high-frequency noise power spectrum estimates. Value and other data.
- the present application provides a noise suppression method and system for sub-band processing.
- the frequency domain signal is divided into a low-frequency signal and a medium-high frequency signal, and the low-frequency signal is processed to obtain a low-frequency noise power spectrum estimated value.
- Process to obtain the estimated value of the medium and high frequency noise power spectrum and calculate the frequency domain signal according to the estimated value of the low frequency noise power spectrum and the estimated value of the medium and high frequency noise power spectrum to obtain the frequency domain signal after gain; according to the signal characteristics of different frequencies, different
- the noise processing method divides noise and voice more accurately, effectively suppresses the spread of noise, improves the efficiency of voice calls, and enhances the effect of voice calls.
Landscapes
- Engineering & Computer Science (AREA)
- Signal Processing (AREA)
- Computational Linguistics (AREA)
- Quality & Reliability (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Human Computer Interaction (AREA)
- Physics & Mathematics (AREA)
- Acoustics & Sound (AREA)
- Multimedia (AREA)
- Telephone Function (AREA)
Abstract
一种分频段进行处理的噪声抑制方法,用于对通话声音中的噪声进行抑制处理,方法包括以下步骤:将采集到的声音数据的信号类型从时域信号变换为频域信号(S1);将频域信号分为低频信号与中高频信号,并根据低频信号的幅值计算出低频噪声功率谱估计值,根据中高频信号的幅值计算出中高频噪声功率谱估计值(S2);根据低频噪声功率谱估计值、中高频噪声功率谱估计值以及频域信号进行计算,得到增益后的频域信号(S3);将增益后的频域信号变换为增益后的时域信号(S4)。一种分频段进行处理的噪声抑制系统;抑制了噪声传播,提高了语音通话的质量,有效的对通话中的噪声进行抑制。
Description
本申请涉及语音通话技术领域,尤其涉及一种分频段进行处理的噪声抑制方法及其系统。
随着科技的发展,语音通话逐渐取代了手写信件,语音通话相比于文字交流,信息传达的速度更快,传达的信息也更加准确;沟通过程中,沟通者可以通过对方的语气、语调、措辞等感知对方的情绪,从而提高交流的效率。对于现阶段的语音通话一般是采用麦克风采集语音数据,所述语音数据为模拟信号,再将模拟信号的语音数据转换为数字信号,通过有线或者无线的方式传送给对方,当对方接收到数字信号后,再将数字信号转换为模拟信号进行播放。
目前,一般的麦克风在采集语音数据时,无法避免的采集到一些噪声数据,所述噪声数据一般是用户在通话时的环境噪声,噪声数据会伴随用户的语音数据一同被传送到通话的另一端;一般情况下,噪声会影响通话质量;严重时,噪声会导致通话障碍,使通话时传达的语音产生偏差;如何在通话中精准的区分语音和噪声,并有效的抑制噪声的传播,提高通话质量成为一大难题。
发明内容
有鉴于此,本申请实施例公开了一种分频段进行处理的噪声抑制方法及其系统,用于在通话中精准的区分语音和噪声,并有效的抑制噪声的传播,提高通话质量。
本申请实施例公开一种分频段进行处理的噪声抑制方法,用于对通话声音中的噪声进行抑制处理,所述方法包括以下步骤:步骤S1,将采集到的声音数据的信号类型从时域信号变换为频域信号;步骤S2,将频域信号分为低频信号与中高频信号,并根据低频信号的幅值计算出低频噪声功率谱估计值,根据中高频信号的幅值计算出中高频噪声功率谱估计值;步骤S3,根据低频噪声功率谱估计值、中高频噪声功率谱估计值以及频域信号进行计算,得到增益 后的频域信号;步骤S4,将增益后的频域信号变换为增益后的时域信号。
本申请实施例还公开一种分频段进行处理的噪声抑制系统,所述分频段进行处理的噪声抑制系统包括:信号变换模块,用于将采集到的声音数据的信号类型从时域信号变换为频域信号;噪声功率谱估计模块,用于将频域信号分为低频信号与中高频信号,并根据低频信号的幅值计算出低频噪声功率谱估计值,根据中高频信号的幅值计算出中高频噪声功率谱估计值;增益频域计算模块,用于根据低频噪声功率谱估计值、中高频噪声功率谱估计值以及频域信号进行计算,得到增益后的频域信号;信号逆变换模块,用于将增益后的频域信号变换为增益后的时域信号。
本申请的分频段进行处理的噪声抑制方法及其系统,通过将频域信号划分为低频信号及中高频信号,并对低频信号进行处理得到低频噪声功率谱估计值,对中高频信号进行处理得到中高频噪声功率谱估计值,根据低频噪声功率谱估计值以及中高频噪声功率谱估计值对频域信号进行计算得到增益后的频域信号;根据不同频率的信号特征,采用不同的噪声处理方式,更加精确的划分噪声和语音,有效的抑制了噪声的传播,提高了语音通话的效率,起到增强语音通话的效果。
为了更清楚地说明本申请实施例中的技术方案,下面将对实施例中所需要使用的附图作简单地介绍,显而易见地,下面描述中的附图仅仅是本申请的一些实施例,对于本领域普通技术人员来讲,在不付出创造性劳动的前提下,还可以根据这些附图获得其他的附图。
图1是本申请一实施例中的一种分频段进行处理的噪声抑制方法的流程图。
图2是图1中步骤S1的子流程图。
图3是图1中步骤S2的子流程图。
图4是图3中步骤S22的子流程图。
图5是图4中步骤S223的子流程图。
图6是图3中步骤S23的子流程图。
图7是图1中步骤S3的子流程图。
图8是图7中步骤S33的子流程图。
图9是本申请一实施例中的一种分频段进行处理的噪声抑制系统的结构关系图。
图10是图9中信号变换模块的结构示意图。
图11是图9中噪声功率谱估计模块的结构示意图。
图12是图11中低频功率谱估计模块的结构示意图。
图13是图12中非基频点功率谱估计模块的结构示意图。
图14是图11中中高频功率谱估计模块的结构示意图。
图15是图9中增益频域计算模块的结构示意图。
图16是图15中增益信号计算模块的结构示意图。
下面将结合本申请实施例中的附图,对本申请实施例中的技术方案进行清楚、完整地描述,显然,所描述的实施例仅是本申请一部分实施例,而不是全部的实施例。基于本申请中的实施例,本领域普通技术人员在没有做出创造性劳动前提下所获得的所有其他实施例,都属于本申请保护的范围。
请参阅图1,图1是本申请一实施例中的一种分频段进行处理的噪声抑制方法的流程图。
如图1所示,本申请提供一种分频段进行处理的噪声抑制方法,用于对通话声音中的噪声进行抑制处理,所述方法包括以下步骤:步骤S1,将采集到的声音数据的信号类型从时域信号变换为频域信号;步骤S2,将频域信号分为低频信号与中高频信号,并根据低频信号的幅值计算出低频噪声功率谱估计值,根据中高频信号的幅值计算出中高频噪声功率谱估计值;步骤S3,根据低频噪声功率谱估计值、中高频噪声功率谱估计值以及频域信号进行计算,得到增益后的频域信号;步骤S4,将增益后的频域信号变换为增益后的时域信号。
从而,本申请提供的一种分频段进行处理的噪声抑制方法,通过将频域信号划分为低频信号及中高频信号,并对低频信号进行处理得到低频噪声功率谱 估计值,对中高频信号进行处理得到中高频噪声功率谱估计值,根据低频噪声功率谱估计值以及中高频噪声功率谱估计值对频域信号进行计算得到增益后的频域信号;根据不同频率的信号特征,采用不同的噪声处理方式,更加精确的划分噪声和语音,有效的抑制了噪声的传播,提高了语音通话的效率,起到增强语音通话的效果。
请参阅图2,图2是本申请一实施例中的全频信号的处理流程示意图。
如图2所示,在一些实施例中,所述步骤S1包括:步骤S11,获取采集到的信号类型为时域信号的声音数据;步骤S12,将时域信号按照预设时间间隔进行组帧得到多帧时域信号;步骤S13,通过对每一帧时域信号进行时频变换得到频域信号。
其中,所述声音数据包括用户在使用通话设备的过程中,所述通话设备采集用户发出的语音数据及环境声音数据;所述声音数据一般通过麦克风进行采集及转换得到,所述麦克风采集到用户发出的语音和环境声音的模拟信号后,经过模数转换,将所述模拟信号转换为数字信号;这里的声音数据是经模数转换后的数字信号。
所述时域信号是将所述声音数据按照采集的时间的顺序进行排列得到的信号;所述时域用来描述物理量与时间的变化关系。所述时域信号的时域波形用于表达信号随时间的变化情况。
所述频域信号是将所述声音数据按照频率大小的顺序进行排序得到的信号。所述频域用来描述物理量与频率的变化关系。所述频域信号的频域波形用于表达信号的幅值随着频率的变化情况,所述频域信号由多个频点构成,每个频点均包含对应的幅值信息和相位信息。
所述组帧是指在时域信号中按照预先设置的时间间隔,将采集到的声音数据进行组合,将在预设时间间隔内的时域信号组合作为一帧,所述时域信号在组合后为多帧时域信号,所述时频信号将以帧为单位进行后续变换。
所述时频变换在本实施例中是指针对时域信号中的每一帧进行傅里叶变换,每一帧在经过傅里叶变换后得到对应的多个频点的幅值、相位信息,所有频点汇总后得到频域信号;在其他实施例中,也可以采用其他变换方式将所述时域信号转换为频域信号。
具体的,将麦克风采集声音数据按照采集的时间顺序进行排列,生成时域信号,在时域信号中按照预设时间间隔进行组合,将在预设时间间隔内的时域信号组合为一帧,再对组合后的每一帧进行傅里叶变换,将每一帧经傅里叶变换得到的频点汇总得到频域信号。
请参阅图3,图3是图1中步骤S2的子流程图。
如图3所示,在一实施例中,所述步骤S2包括:步骤S21,按照预设频率划分标准将所述频域信号化分为低频信号及中高频信号;步骤S22,对频域信号中单路麦克风采集的单路低频信号进行噪声功率谱密度估计,得到单路低频信号中所有低频频点的噪声功率谱估计值;步骤S23,对频域信号中双路麦克风采集的双路中高频信号进行噪声功率谱密度估计,得到双路中高频信号中所有中高频频点的噪声功率谱估计值。
其中,所述预设频率划分标准是指预先设定的用于对频域信号进行划分的标准;本实施例中,按照所述预设频率划分标准可以将所述频域信号划分为低频信号、中高频信号;本实施例中,所述预设频率划分标准可以为某一固定频率;例如,所述预设频率划分标准可为以预设频率,例如1000Hz为界线区分低频、中高频信号的标准。
所述低频信号是指所述频域信号中低于所述预设频率的信号段。例如,频域信号中低于所述预设频率的信号段属于低频信号。所述低频信号中的频点为低频频点。
所述中高频信号是指所述频域信号中不低于所述预设频率的信号段。例如,频域信号中等于或高于所述预设频率的信号段属于中高频信号。所述中高频信号中的频点为中高频频点。
所述麦克风是指所述通话设备中用于采集声音数据的部件,所述麦克风一般为两个,分别设置于通话设备的不同部位;在用户进行通话时,两路麦克风同时采集声音数据,并将采集到的声音数据分别进行时频变换,得到两路频域信号;在对不同频段的信号进行噪声抑制处理时,采用不同路的麦克风的频域信号。在本实施例中,对低频信号进行噪声抑制处理时,选择信噪比较高的一路麦克风的频域信号进行噪声抑制处理。对中高频信号进行噪声抑制处理时,同时选择两路麦克风的双路频域信号进行噪声抑制处理。
所述功率谱密度(power spectral density,PSD)有时亦称为谱功率分布(spectral power distribution,SPD),是信号(噪声)的自相关函数的傅里叶变换,即每单位频率的信号(噪声)所携带的功率。功率谱密度是一种概率统计方法,是对随机变量均方值的度量。
所述噪声功率谱估计值,即对噪声的功率谱密度的估计值,又称功率谱估计。噪声的功率谱密度是用来描述噪声的能量特征随频率的变化关系。
具体的,当所述预设频率划分标准确定后,按照预设频率划分标准将所述频域信号中频率低于预设频率划分标准的信号段作为低频段,等于或高于预设频率划分标准的信号段作为中高频段。
将两路麦克风采集的声音数据进行时频变换,得到两路频域信号,选择其中信噪比较高的一路频域信号,并根据所述预设频率划分标准选取该路频域信号中的低频段作为低频信号,并对所述低频信号进行噪声功率谱密度估计,得到所述低频信号中所有低频频点的噪声功率谱估计值。
将根据所述预设频率划分标准,同时选取两路频域信号的中高频段作为中高频信号,并对所述中高频信号进行噪声功率谱密度估计,得到所述中高频信号中所有中高频频点的噪声功率谱估计值。
在其他实施例中,所述预设频率划分标准可以为固定的频率段;根据频率段对两路频域信号进行更细致的划分;例如0-1000Hz为低频段,1000-3000Hz为中频段,3000Hz以上为高频段。
从而,通过两路麦克风分别采集两路声音数据,再对两路声音数据分别进行时频变换,得到两路频域信号,在两路频域信号中分别选定低频信号及中高频信号,再对低频信号及中高频信号进行噪声功率谱密度估计,利用低频信号中噪声估计的准确性以及中高频信号中噪声估计的相关性,对频域信号分段处理,充分利用了采集到的两路声音数据,提高了噪声功率谱估计值的精准度,尽可能地去除残留的噪声。
请参阅图4,图4是图3中步骤S22的子流程图。
在一些实施例中,所述步骤S22包括:步骤S221,将单路低频信号中每个低频频点的幅值进行平方得到每个低频频点的幅值的平方值;步骤S222,通过基音检测法将低频频点分为基频点和非基频点;步骤S223,根据非基频 点的幅值的平方值计算得到非基频点的噪声功率谱估计值;步骤S224,根据基频点的幅值的平方值以及非基频点的噪声功率谱估计值计算得到基频点的噪声功率谱估计值;步骤S225,将非基频点的噪声功率谱估计值与基频点的噪声功率谱估计值组合得到所有低频频点的噪声功率谱估计值。
其中,所述基音是指人发出浊音时声带振动的周期,基音周期的估计称为基音检测,其目的是提取出与人的声带振动频率一致或尽可能相吻合的基音周期变化的轨迹曲线,即基音频率,所述基音频率是语音信号处理中最重要的特征参数之一。
所述基音检测法是用于检测基音信号的方法,由于语音信号可视为一个动态非平稳随机过程,语音波形和声带振动的频率变化范围大且十分复杂,所以所述基音检测法包括很多种算法,例如泛音内积频谱法、倒频谱分析法、最大似然估计法等;本实施例中,采用倒谱法对所述低频频点进行检测,基音检测中的基音频率可以通过以下计算方法进行确定:
X(ω)=FT[x(n)]
其中,FT和FT
-1分别表示傅里叶变换和傅里叶逆变换。由于时域信号x(n)是由声门脉冲激励u(n)经声道响应v(n)滤波而得,即
x(n)=u(n)*v(n)
所述基频点是指低频信号中幅值的平方符合基音频率的频点。
所述非基频点是指低频信号中幅值的平方不符合基音频率的频点。
所述基频点的噪声功率谱估计值是指对基频点的噪声的功率谱密度的估计值。所述非基频点的噪声功率谱估计值是指对非基频点的噪声的功率谱密度 的估计值。所述噪声功率谱估计值可以通过如下公式进行计算:
其中,M={f
0,2f
0,3f
0,...}。f
0为基音频率,2f
0,3f
0,…表示谐波频率。M表示基频点的集合。|X(λ,μ)|
2为原输入,即非基频点的噪声功率谱估计值为该非基频点幅值的平方;|X
inter(λ,μ)|
2是利用插值法计算基频点的噪声功率谱估计值,即通过相邻两个非基频点的噪声功率谱估计值进行插值得到所述基频点的噪声功率谱估计值。
具体的,将所述单路低频信号中每个低频频点的幅值进行平方,得到每个低频频点的幅值的平方值。
根据基音检测法对所述低频信号中所有的频点进行筛选,筛选出低频信号中频点的幅值的平方符合基音频率的频点作为基频点,不符合基音频率的频点作为非基频点。
根据上述的所述基频点噪声功率谱估计值和非基频点噪声功率谱估计值的计算公式,计算求得所述基频点噪声功率谱估计值和非基频点噪声功率谱估计值。
将计算得到的非基频点的噪声功率谱估计值与基频点的噪声功率谱估计值进行组合汇总,即得到所有低频频点的噪声功率谱估计值。
请参阅图5,图5是图4中步骤S223的子流程图。
在一些实施例中,所述步骤S223包括:步骤S2231,根据每个低频频点的幅值的平方值计算得到每个低频频点的语音存在概率值;步骤S2232,根据每个非基频点的幅值的平方值计算得到每个非基频点的噪声功率谱初步估计值;步骤S2233,在所有低频频点的语音存在概率值中查找出与非基频点对应的非基频点的语音存在概率值;步骤S2234,根据每个非基频点的语音存在概率值以及对应的非基频点的噪声功率谱初步估计值计算得到每个非基频点的噪声功率谱估计值。
其中,所述语音是指所述声音数据中用户言语的声音,所述声音数据包括语音数据及噪声数据,所述语音数据是通话中需要传达的数据。
所述低频频点的语音存在概率值是指低频频点中语音存在的可能性。所述 语音存在概率值可以通过以下公式计算得到:
该计算公式中,q(k,λ)代表语音不存在概率,该值可以通过对应频点的幅度谱平方与预设阈值进行比较计算得到;ξ(k,λ)为先验信噪比;υ(k,λ)可以由后验信噪比和先验信噪比定义计算得到。其中,先验信噪比是纯净语音信号的功率除以噪声信号的功率。后验信噪比是含噪语音信号的功率除以噪声信号的功率。
所述非基频点的噪声功率谱初步估计值是利用上述公式,对所述非基频点的噪声功率谱估计值进行初步计算得到的值。
具体的,将低频信号中的每个低频频点的幅值的平方值带入上述语音存在概率的公式进行计算,得到低频信号中所有低频频点的语音存在概率值。
根据每个非基频点的幅值的平方值带入上述噪声功率谱估计值的公式计算得到每个非基频点的噪声功率谱初步估计值。
根据所有低频频点的语音存在概率值以及上述基音检测结果中的非基频点,在所有低频频点的语音存在概率值对应查找出每个非基频点的语音存在概率值。
将每个非基频点的语音存在概率值与对应的非基频点的噪声功率谱初步估计值相乘,计算得到每个非基频点的噪声功率谱估计值。
从而,将低频信号中的低频频点划分为基频点与非基频点,再分别针对基频点和非基频点采用不同的计算方式,计算出低频信号中基频点的噪声功率谱估计值以及非基频点的噪声功率谱估计值;通过引入基音检测法,保证了通话语音的质量,提高了在平稳及非平稳的噪声环境下的语音通话效果,使得对噪声的区分更加精准,提高了噪声抑制的效率。
请参阅图6,图6是图3中步骤S23的子流程图。
在一些实施例中,步骤S231,通过对双路中高频信号中每个频点的幅值进行平方得到每个中高频频点的幅值的平方值;步骤S232,根据每个中高频频点的幅值的平方值计算得到每个中高频频点的自功率谱值,以及每个中高频频 点的互功率谱值;步骤S233,根据每个中高频频点的幅值的平方值、自功率谱值及互功率谱值计算得到每个中高频频点的相关性值;步骤S234,根据每个中高频频点的相关性值、自功率谱值及互功率谱值计算得到每个中高频频点的噪声功率谱的初步估计值;步骤S235,根据每个中高频频点的相关性值计算得到每个中高频频点的语音存在概率;步骤S236,通过将每个中高频频点的噪声功率谱的初步估计值与对应的中高频频点的语音存在概率进行平滑处理得到每个中高频频点的噪声功率谱估计值。
其中,所述自功率谱值以及互功率谱值用于反映相关函数在时域内表达随机信号自身与其他信号在不同时刻的内在联系,了解不同时刻同一随机样本间的波形相似程度。所述自功率谱以及互功率谱值的计算公式如下:
所述相关性值包括两个麦克风中语音信号的相关性值以及两个麦克风信号中噪声信号的相关性值。所述相关性值可以根据以下公式计算得到:
所述中高频频点的语音存在概率是指每个中高频频点中语音信号存在的可能性。所述中高频频点的语音存在概率可以通过以下公式计算得到:
所述平滑处理是指将每个中高频频点的噪声功率谱估计值以及对应频点的语音存在概率带入公式计算得到平滑后的中高频频点的噪声功率谱估计值;
所述平滑处理的计算公式如下:
所述中高频频点的噪声功率谱的初步估计值是指根据公式初步计算得到的中高频频点的噪声功率谱的估计值;所述中高频频点的噪声功率谱的初步估计值的计算公式如下:
具体的,首先,同时获取双路频域信号中的中高频信号,对中高频信号中的每个频点的幅值进行平方,得到每个中高频频点的幅值的平方值。
再将每个中高频频点的幅值的平方值带入上述自功率谱值以及互功率谱值的计算公式,计算得到每个中高频频点的自功率谱值以及互功率谱值。
再将每个中高频频点的幅值的平方值、自功率谱值以及互功率谱值带入上述相关性值的计算公式,计算得到每个中高频频点的相关性值。
再根据每个中高频频点的相关性值、自功率谱值及互功率谱值带入上述的中高频频点的噪声功率谱的初步估计值的计算公式,计算得到每个中高频频点的噪声功率谱的初步估计值。
同时,根据每个中高频频点的相关性值带入上述中高频频点的语音存在概率的计算公式,计算得到每个中高频频点的语音存在概率。
最后,通过将每个中高频频点的噪声功率谱的初步估计值与对应的中高频频点的语音存在概率带入上述的平滑处理的公式进行计算,得到每个中高频频点的噪声功率谱估计值。
从而,通过将所述中高频频点的幅值带入一系列的公式进行计算,利用语 音在中高频段相关性相差明显的特征,对中高频信号的频点进行噪声的抑制处理,得到所述中高频信号中每个频点的噪声功率谱估计值。
请参阅图7,图7是图1中步骤S3的子流程图。
如图7所示,在一些实施例中,所述步骤S3包括:步骤S31,将低频噪声功率谱估计值与中高频噪声功率谱估计值组合得到全频带噪声功率谱估计值;步骤S32,根据全频带噪声功率谱估计值中每个频点的噪声功率谱估计值进行增益计算,得到每个频点的幅度谱增益值;步骤S33,根据每个频点的幅度谱增益值及频域信号计算增益后的频域信号。
其中,所述全频带噪声功率谱估计值是指频域信号中所有频点的噪声功率谱估计值。
所述幅度谱增益值是指将所述频点的噪声功率谱估计值带入幅度增益函数求出的每个频点的幅值的增益比例。
所述增益后的频域信号是指对原频域信号中每个频点按照对应频点的幅度谱增益值进行增益后得到的每个频点的幅值。所述原频域信号是指上述时频变换后得到的频域信号。
具体的,将低频噪声功率谱与中高频噪声功率谱进行组合汇总得到频域信号中所有频点的噪声功率谱估计值。
将所有频点的噪声功率谱估计值中每个频点的噪声功率谱估计值带入增益函数进行计算,得到每个频点的幅度谱增益值。
根据每个频点的幅度谱增益值与原频域信号中对应频点的幅值进行计算,得到每个频点增益后的幅值,将所有频点增益后的幅值进行组合汇总后,得到增益后的频域信号。
请参阅图8,图8是图7中步骤S33的子流程图。
如图8所示,在一些实施例中,所述步骤S33包括:步骤S331,获取每个频点的幅度谱增益值以及频域信号中对应频点的幅值;步骤S332,将每个频点的幅度谱增益值与对应频点的幅值相乘得到每个频点的增益后的幅值,将频域信号中所有频点的增益后的幅值组合得到增益后的频域信号。
具体的,根据每个频点的幅度谱增益值,在原频域信号中获取到对应的每个频点的幅值。
将每个频点的幅度谱增益值与对应频点的幅值相乘,得到每个频点的增益后的幅值;将所有频点的增益后的幅值组合汇总后,得到增益后的频域信号。
从而,根据频域信号中每个频点的噪声功率谱估计值计算得到每个频点的幅度谱增益值,再根据每个幅度谱增益值与原频域信号中对应的频点的幅值相乘,计算得到增益后的频域信号,提高了语音的可识别性,并针对每个频点进行了噪声抑制处理,提高了噪声抑制的效率,使得通话中的语音数据更为清晰,抑制噪声数据的传播,提高了通话的质量。
在一些实施例中,所述步骤S4包括:将增益后的频域信号中每个频点的增益后的幅值进行逆时频变换,得到增益后的时域信号。
其中,所述逆时频变换是指将频域信号变换为时域信号,通过对每个频点进行傅里叶逆变换,从而实现将整个频域信号变换为时域信号。
具体的,将增益后的频域信号按照帧进行傅里叶逆变换,得到对应帧的时域信号,所有帧的时域信号组合后得到增益后的时域信号;从而,将所述增益后的频域信号变换为增益后的时域信号。
本申请提供的一种分频段进行处理的噪声抑制方法可以在硬件、固件中实施,或者可以作为可以存储在例如CD、ROM、RAM、软盘、硬盘或磁光盘的等计算机可读存储介质中的软件或计算机代码,或者可以作为原始存储在远程记录介质或非瞬时的机器可读介质上、通过网络下载并且存储在本地记录介质中的计算机代码,从而这里描述的一种分频段进行处理的噪声抑制方法可以利用通用计算机或特殊处理器或在诸如ASIC或FPGA之类的可编程或专用硬件中以存储在记录介质上的软件来呈现。如本领能够理解的,计算机、处理器、微处理器、控制器或可编程硬件包括存储器组件,例如,RAM、ROM、闪存等,当计算机、处理器或硬件实施这里描述的一种分频段进行处理的噪声抑制方法而存取和执行软件或计算机代码时,存储器组件可以存储或接收软件或计算机代码。另外,当通用计算机存取用于实施这里示出的处理的代码时,代码的执行将通用计算机转换为用于执行这里示出的处理的专用计算机。
其中,所述计算机可读存储介质可为固态存储器、存储卡、光碟等。所述计算机可读存储介质存储有程序指令而供计算机调用后执行图1至图8所示的一种分频段进行处理的噪声抑制方法。
请参阅图9,图9是本申请一实施例中的一种分频段进行处理的噪声抑制系统100的结构关系图。
在一些实施例中,所述分频段进行处理的噪声抑制系统100包括:信号变换模块10,用于将采集到的声音数据的信号类型从时域信号变换为频域信号;噪声功率谱估计模块20,用于将频域信号分为低频信号与中高频信号,并根据低频信号的幅值计算出低频噪声功率谱估计值,根据中高频信号的幅值计算出中高频噪声功率谱估计值;增益频域计算模块30,用于根据低频噪声功率谱估计值、中高频噪声功率谱估计值以及频域信号进行计算,得到增益后的频域信号;信号逆变换模块40,用于将增益后的频域信号变换为增益后的时域信号。
请参阅图10,图10是图9中信号变换模块10的结构示意图。
在一些实施例中,所述信号变换模块10包括:信号获取模块11,用于获取采集到的信号类型为时域信号的声音数据;信号组帧模块12,用于将时域信号按照预设时间间隔进行组帧得到多帧时域信号;时频变换模块13,用于通过对每一帧时域信号进行时频变换得到频域信号。
请参阅图11,图11是图9中噪声功率谱估计模块20的结构示意图。
在一些实施例中,所述噪声功率谱估计模块20包括:信号划分模块21,用于按照预设频率划分标准将所述频域信号化分为低频信号及中高频信号;低频功率谱估计模块22,用于对频域信号中单路麦克风采集的单路低频信号进行噪声功率谱密度估计,得到单路低频信号中所有低频频点的噪声功率谱估计值;中高频功率谱估计模块23,用于对频域信号中双路麦克风采集的双路中高频信号进行噪声功率谱密度估计,得到双路中高频信号中所有中高频频点的噪声功率谱估计值。
请参阅图12,图12是图11中低频功率谱估计模块22的结构示意图。
在一些实施例中,所述低频功率谱估计模块22包括:低频幅值平方模块221,用于将单路低频信号中每个低频频点的幅值进行平方得到每个低频频点的幅值的平方值;基音检测模块222,用于通过基音检测法将低频频点分为基频点和非基频点;非基频点功率谱估计模块223,用于根据非基频点的幅值的平方值计算得到非基频点的噪声功率谱估计值;基频点功率谱估计模块224, 用于根据基频点的幅值的平方值以及非基频点的噪声功率谱估计值计算得到基频点的噪声功率谱估计值;低频功率谱组合模块225,用于将非基频点的噪声功率谱估计值与基频点的噪声功率谱估计值组合得到所有低频频点的噪声功率谱估计值。
请参阅图13,图13是图12中非基频点功率谱估计模块223的结构示意图。
在一些实施例中,所述非基频点功率谱估计模块223包括:低频语音概率计算模块2231,用于根据每个低频频点的幅值的平方值计算得到每个低频频点的语音存在概率值;非基频点功率谱初步估计模块2232,用于根据每个非基频点的幅值的平方值计算得到每个非基频点的噪声功率谱初步估计值;非基频点语音概率查找模块2233,用于在所有低频频点的语音存在概率值中查找出与非基频点对应的非基频点的语音存在概率值;非基频点功率谱计算模块2234,用于根据每个非基频点的语音存在概率值以及对应的非基频点的噪声功率谱初步估计值计算得到每个非基频点的噪声功率谱估计值。
请参阅图14,图14是图11中中高频功率谱估计模块23的结构示意图。
在一些实施例中,所述中高频功率谱估计模块23包括:中高频幅值平方模块231,用于通过对双路中高频信号中每个频点的幅值进行平方得到每个中高频频点的幅值的平方值;功率谱值计算模块232,用于根据每个中高频频点的幅值的平方值计算得到每个中高频频点的自功率谱值,以及每个中高频频点的互功率谱值;相关性值计算模块233,用于根据每个中高频频点的幅值的平方值、自功率谱值及互功率谱值计算得到每个中高频频点的相关性值;中高频初步估计模块234,用于根据每个中高频频点的相关性值、自功率谱值及互功率谱值计算得到每个中高频频点的噪声功率谱的初步估计值;中高频语音概率计算模块235,用于根据每个中高频频点的相关性值计算得到每个中高频频点的语音存在概率;中高频平滑处理模块236,用于通过将每个中高频频点的噪声功率谱的初步估计值与对应的中高频频点的语音存在概率进行平滑处理得到每个中高频频点的噪声功率谱估计值。
请参阅图15,图15是图9中增益频域计算模块30的结构示意图。
在一些实施例中,所述增益频域计算模块30包括:全频带功率谱组合模块31,用于将低频噪声功率谱估计值与中高频噪声功率谱估计值组合得到全频带 噪声功率谱估计值;幅度谱增益计算模块32,用于根据全频带噪声功率谱估计值中每个频点的噪声功率谱估计值进行增益计算,得到每个频点的幅度谱增益值;增益信号计算模块33,用于根据每个频点的幅度谱增益值及频域信号计算增益后的频域信号。
请参阅图16,图16是图15中增益信号计算模块33的结构示意图。
在一些实施例中,所述增益信号计算模块33包括:幅度谱增益值获取模块331,用于获取每个频点的幅度谱增益值以及频域信号中对应频点的幅值;增益幅值计算模块332,用于将每个频点的幅度谱增益值与对应频点的幅值相乘得到每个频点的增益后的幅值,将频域信号中所有频点的增益后的幅值组合得到增益后的频域信号。
在一些实施例中,所述信号逆变换模块40还用于将增益后的频域信号中每个频点的增益后的幅值进行逆时频变换,得到增益后的时域信号。
所述分频段进行处理的噪声抑制系统100还包括中央处理单元(Central Processing Unit,CPU),还可以是其他通用处理器、数字信号处理器(Digital Signal Processor,DSP)、专用集成电路(Application Specific Integrated Circuit,ASIC)、现场可编程门阵列(Field-Programmable Gate Array,FPGA)或者其他可编程逻辑器件、分立门或者晶体管逻辑器件、分立硬件组件等。通用处理器可以是微处理器或者该处理器也可以是任何常规的处理器等,所述处理单元是所述分频段进行处理的噪声抑制系统的数据处理中心,利用有线或者无线线路连接整个所述分频段进行处理的噪声抑制系统100的各个模块。用于对各个模块传送来的数据进行处理。
如图9所示,在一些实施例中,所述分频段进行处理的噪声抑制系统100还包括存储模块50,所述存储模块50用于存储所述时域信号以及所述频域信号
其中,所述存储模块50可以包括高速随机存取存储器,还可以包括非易失性存储器,例如硬盘、内存、插接式硬盘,智能存储卡(Smart Media Card,SMC),安全数字(Secure Digital,SD)卡,闪存卡(Flash Card)、多个磁盘存储器件、闪存器件、或其他易失性固态存储器件。
具体的,所述存储模块50位于所述通话设备内,主要与上述处理单元相 连接,用于存储麦克风采集到的声音数据以及处理单元处理后的频域信号、时域信号、噪声功率谱估计值等。
在其他实施例中,所述存储模块50还可以与所述分频段进行处理的噪声抑制系统100的各个模块相连接,用于存储各个模块在参数计算过程中产生的数据。
如图1所示,在一些实施例中,所述一种分频段进行处理的噪声抑制系统包括通话设备,所述通话设备可以是移动终端,所述信号变换模块10、噪声功率谱估计模块20、增益频域计算模块30、信号逆变换模块40、存储模块50等,均设置于所述主机设备内;所述信号变换模块10与噪声功率谱估计模块20通过有线或者无线的方式连接,能够将时频变换后的频域信号传送给所述噪声功率谱估计模块20;所述噪声功率谱估计模块20与所述增益频域计算模块30通过有线或者无线的方式连接,能够将计算得到的低频噪声功率谱估计值、中高频噪声功率谱估计值传送给所述增益频域计算模块30;所述增益频域计算模块30与所述信号逆变换模块40通过有线或者无线的方式连接,能够将增益后的频域信号传送给信号逆变换模块40,用于信号逆变换模块40将增益后的频域信号逆时频变换为时域信号;所述存储模块50与信号变换模块10、噪声功率谱估计模块20、增益频域计算模块30、信号逆变换模块40均通过有线或者无线的方式连接,能够存储时域信号、频域信号、低频噪声功率谱估计值、中高频噪声功率谱估计值等数据。
本申请提供的一种分频段进行处理的噪声抑制方法及其系统,通过将频域信号划分为低频信号及中高频信号,并对低频信号进行处理得到低频噪声功率谱估计值,对中高频信号进行处理得到中高频噪声功率谱估计值,根据低频噪声功率谱估计值以及中高频噪声功率谱估计值对频域信号进行计算得到增益后的频域信号;根据不同频率的信号特征,采用不同的噪声处理方式,更加精确的划分噪声和语音,有效的抑制了噪声的传播,提高了语音通话的效率,起到增强语音通话的效果。
以上所述是本申请的优选实施例,应当指出,对于本技术领域的普通技术人员来说,在不脱离本申请原理的前提下,还可以做出若干改进和润饰,这些改进和润饰也视为本申请的保护范围。
Claims (20)
- 一种分频段进行处理的噪声抑制方法,用于对通话声音中的噪声进行抑制处理,其特征在于,所述方法包括以下步骤:步骤S1,将采集到的声音数据的信号类型从时域信号变换为频域信号;步骤S2,将频域信号分为低频信号与中高频信号,并根据低频信号的幅值计算出低频噪声功率谱估计值,根据中高频信号的幅值计算出中高频噪声功率谱估计值;步骤S3,根据低频噪声功率谱估计值、中高频噪声功率谱估计值以及频域信号进行计算,得到增益后的频域信号;步骤S4,将增益后的频域信号变换为增益后的时域信号。
- 如权利要求1所述的一种分频段进行处理的噪声抑制方法,其特征在于,所述步骤S1包括:步骤S11,获取采集到的信号类型为时域信号的声音数据;步骤S12,将时域信号按照预设时间间隔进行组帧得到多帧时域信号;步骤S13,通过对每一帧时域信号进行时频变换得到频域信号。
- 如权利要求1或2任一项所述的一种分频段进行处理的噪声抑制方法,其特征在于,所述步骤S2包括:步骤S21,按照预设频率划分标准将所述频域信号化分为低频信号及中高频信号;步骤S22,对频域信号中单路麦克风采集的单路低频信号进行噪声功率谱密度估计,得到单路低频信号中所有低频频点的噪声功率谱估计值;步骤S23,对频域信号中双路麦克风采集的双路中高频信号进行噪声功率谱密度估计,得到双路中高频信号中所有中高频频点的噪声功率谱估计值。
- 如权利要求3所述的一种分频段进行处理的噪声抑制方法,其特征在于,所述步骤S22包括:步骤S221,将单路低频信号中每个低频频点的幅值进行平方得到每个低频频点的幅值的平方值;步骤S222,通过基音检测法将低频频点分为基频点和非基频点;步骤S223,根据非基频点的幅值的平方值计算得到非基频点的噪声功率谱估计值;步骤S224,根据基频点的幅值的平方值以及非基频点的噪声功率谱估计值计算得到基频点的噪声功率谱估计值;步骤S225,将非基频点的噪声功率谱估计值与基频点的噪声功率谱估计值组合得到所有低频频点的噪声功率谱估计值。
- 如权利要求4所述的一种分频段进行处理的噪声抑制方法,其特征在于,所述步骤S223包括:步骤S2231,根据每个低频频点的幅值的平方值计算得到每个低频频点的语音存在概率值;步骤S2232,根据每个非基频点的幅值的平方值计算得到每个非基频点的噪声功率谱初步估计值;步骤S2233,在所有低频频点的语音存在概率值中查找出与非基频点对应的非基频点的语音存在概率值;步骤S2234,根据每个非基频点的语音存在概率值以及对应的非基频点的噪声功率谱初步估计值计算得到每个非基频点的噪声功率谱估计值。
- 如权利要求3所述的一种分频段进行处理的噪声抑制方法,其特征在于,所述步骤S23包括:步骤S231,通过对双路中高频信号中每个频点的幅值进行平方得到每个中高频频点的幅值的平方值;步骤S232,根据每个中高频频点的幅值的平方值计算得到每个中高频频点的自功率谱值,以及每个中高频频点的互功率谱值;步骤S233,根据每个中高频频点的幅值的平方值、自功率谱值以及互功率谱值计算得到每个中高频频点的相关性值;步骤S234,根据每个中高频频点的相关性值、自功率谱值以及互功率谱值计算得到每个中高频频点的噪声功率谱的初步估计值;步骤S235,根据每个中高频频点的相关性值计算得到每个中高频频点的语音存在概率;步骤S236,通过将每个中高频频点的噪声功率谱的初步估计值与对应的 中高频频点的语音存在概率进行平滑处理得到每个中高频频点的噪声功率谱估计值。
- 如权利要求1所述的一种分频段进行处理的噪声抑制方法,其特征在于,所述步骤S3包括:步骤S31,将低频噪声功率谱估计值与中高频噪声功率谱估计值组合得到全频带噪声功率谱估计值;步骤S32,根据全频带噪声功率谱估计值中每个频点的噪声功率谱估计值进行增益计算,得到每个频点的幅度谱增益值;步骤S33,根据每个频点的幅度谱增益值及频域信号计算增益后的频域信号。
- 如权利要求7所述的一种分频段进行处理的噪声抑制方法,其特征在于,所述步骤S33包括:步骤S331,获取每个频点的幅度谱增益值以及频域信号中对应频点的幅值;步骤S332,将每个频点的幅度谱增益值与对应频点的幅值相乘得到每个频点的增益后的幅值,将频域信号中所有频点的增益后的幅值组合得到增益后的频域信号。
- 如权利要求8所述的一种分频段进行处理的噪声抑制方法,其特征在于,所述步骤S4包括:将增益后的频域信号中每个频点的增益后的幅值进行逆时频变换,得到增益后的时域信号。
- 如权利要求2或9所述的一种分频段进行处理的噪声抑制方法,其特征在于,所述时频变换为傅里叶变换,所述逆时频变换为傅里叶逆变换。
- 一种分频段进行处理的噪声抑制系统,其特征在于,所述分频段进行处理的噪声抑制系统包括:信号变换模块,用于将采集到的声音数据的信号类型从时域信号变换为频域信号;噪声功率谱估计模块,用于将频域信号分为低频信号与中高频信号,并根据低频信号的幅值计算出低频噪声功率谱估计值,根据中高频信号的幅值计算出中高频噪声功率谱估计值;增益频域计算模块,用于根据低频噪声功率谱估计值、中高频噪声功率谱估计值以及频域信号进行计算,得到增益后的频域信号;信号逆变换模块,用于将增益后的频域信号变换为增益后的时域信号。
- 如权利要11所述的一种分频段进行处理的噪声抑制系统,其特征在于,所述信号变换模块包括:信号获取模块,用于获取采集到的信号类型为时域信号的声音数据;信号组帧模块,用于将时域信号按照预设时间间隔进行组帧得到多帧时域信号;时频变换模块,用于通过对每一帧时域信号进行时频变换得到频域信号。
- 如权利要求11或12任一项所述的一种分频段进行处理的噪声抑制系统,其特征在于,所述噪声功率谱估计模块包括:信号划分模块,用于按照预设频率划分标准将所述频域信号化分为低频信号及中高频信号;低频功率谱估计模块,用于对频域信号中单路麦克风采集的单路低频信号进行噪声功率谱密度估计,得到单路低频信号中所有低频频点的噪声功率谱估计值;中高频功率谱估计模块,用于对频域信号中双路麦克风采集的双路中高频信号进行噪声功率谱密度估计,得到双路中高频信号中所有中高频频点的噪声功率谱估计值。
- 如权利要求13所述的一种分频段进行处理的噪声抑制系统,其特征在于,所述低频功率谱估计模块包括:低频幅值平方模块,用于将单路低频信号中每个低频频点的幅值进行平方得到每个低频频点的幅值的平方值;基音检测模块,用于通过基音检测法将低频频点分为基频点和非基频点;非基频点功率谱估计模块,用于根据非基频点的幅值的平方值计算得到非基频点的噪声功率谱估计值;基频点功率谱估计模块,用于根据基频点的幅值的平方值以及非基频点的噪声功率谱估计值计算得到基频点的噪声功率谱估计值;低频功率谱组合模块,用于将非基频点的噪声功率谱估计值与基频点的噪 声功率谱估计值组合得到所有低频频点的噪声功率谱估计值。
- 如权利要求14所述的一种分频段进行处理的噪声抑制系统,其特征在于,所述非基频点功率谱估计模块包括:低频语音概率计算模块,用于根据每个低频频点的幅值的平方值计算得到每个低频频点的语音存在概率值;非基频点功率谱初步估计模块,用于根据每个非基频点的幅值的平方值计算得到每个非基频点的噪声功率谱初步估计值;非基频点语音概率查找模块,用于在所有低频频点的语音存在概率值中查找出与非基频点对应的非基频点的语音存在概率值;非基频点功率谱计算模块,用于根据每个非基频点的语音存在概率值以及对应的非基频点的噪声功率谱初步估计值计算得到每个非基频点的噪声功率谱估计值。
- 如权利要求13所述的一种分频段进行处理的噪声抑制系统,其特征在于,所述中高频功率谱估计模块包括:中高频幅值平方模块,用于通过对双路中高频信号中每个频点的幅值进行平方得到每个中高频频点的幅值的平方值;功率谱值计算模块,用于根据每个中高频频点的幅值的平方值计算得到每个中高频频点的自功率谱值,以及每个中高频频点的互功率谱值;相关性值计算模块,用于根据每个中高频频点的幅值的平方值、自功率谱值及互功率谱值计算得到每个中高频频点的相关性值;中高频初步估计模块,用于根据每个中高频频点的相关性值、自功率谱值及互功率谱值计算得到每个中高频频点的噪声功率谱的初步估计值;中高频语音概率计算模块,用于根据每个中高频频点的相关性值计算得到每个中高频频点的语音存在概率;中高频平滑处理模块,用于通过将每个中高频频点的噪声功率谱的初步估计值与对应的中高频频点的语音存在概率进行平滑处理得到每个中高频频点的噪声功率谱估计值。
- 如权利要求11所述的一种分频段进行处理的噪声抑制系统,其特征在于,所述增益频域计算模块包括:全频带功率谱组合模块,用于将低频噪声功率谱估计值与中高频噪声功率谱估计值组合得到全频带噪声功率谱估计值;幅度谱增益计算模块,用于根据全频带噪声功率谱估计值中每个频点的噪声功率谱估计值进行增益计算,得到每个频点的幅度谱增益值;增益信号计算模块,用于根据每个频点的幅度谱增益值及频域信号计算增益后的频域信号。
- 如权利要求17所述的一种分频段进行处理的噪声抑制系统,其特征在于,所述增益信号计算模块包括:幅度谱增益值获取模块,用于获取每个频点的幅度谱增益值以及频域信号中对应频点的幅值;增益幅值计算模块,用于将每个频点的幅度谱增益值与对应频点的幅值相乘得到每个频点的增益后的幅值,将频域信号中所有频点的增益后的幅值组合得到增益后的频域信号。
- 如权利要求18所述的一种分频段进行处理的噪声抑制系统,其特征在于,所述信号逆变换模块还用于将增益后的频域信号中每个频点的增益后的幅值进行逆时频变换,得到增益后的时域信号。
- 如权利要求12或19所述的一种分频段进行处理的噪声抑制系统,其特征在于,所述时频变换为傅里叶变换,所述逆时频变换为傅里叶逆变换。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN201911278646.5 | 2019-12-10 | ||
| CN201911278646.5A CN111128213B (zh) | 2019-12-10 | 2019-12-10 | 一种分频段进行处理的噪声抑制方法及其系统 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2021114733A1 true WO2021114733A1 (zh) | 2021-06-17 |
Family
ID=70498955
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2020/111672 Ceased WO2021114733A1 (zh) | 2019-12-10 | 2020-08-27 | 一种分频段进行处理的噪声抑制方法及其系统 |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN111128213B (zh) |
| WO (1) | WO2021114733A1 (zh) |
Cited By (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN113613112A (zh) * | 2021-09-23 | 2021-11-05 | 三星半导体(中国)研究开发有限公司 | 抑制麦克风的风噪的方法和电子装置 |
| CN113851151A (zh) * | 2021-10-26 | 2021-12-28 | 北京融讯科创技术有限公司 | 掩蔽阈值估计方法、装置、电子设备和存储介质 |
Families Citing this family (13)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN111128213B (zh) * | 2019-12-10 | 2022-09-27 | 展讯通信(上海)有限公司 | 一种分频段进行处理的噪声抑制方法及其系统 |
| CN114242074B (zh) * | 2020-09-08 | 2026-04-17 | 华为技术有限公司 | 人声检测方法和设备 |
| CN112309419B (zh) * | 2020-10-30 | 2023-05-02 | 浙江蓝鸽科技有限公司 | 多路音频的降噪、输出方法及其系统 |
| CN112331192B (zh) * | 2020-11-06 | 2025-06-03 | 深圳Tcl新技术有限公司 | 音频数据的处理方法、终端设备及计算机可读存储介质 |
| CN113516988B (zh) * | 2020-12-30 | 2024-02-23 | 腾讯科技(深圳)有限公司 | 一种音频处理方法、装置、智能设备及存储介质 |
| CN113571078B (zh) * | 2021-01-29 | 2024-04-26 | 腾讯科技(深圳)有限公司 | 噪声抑制方法、装置、介质以及电子设备 |
| CN112951262B (zh) * | 2021-02-24 | 2023-03-10 | 北京小米松果电子有限公司 | 音频录制方法及装置、电子设备及存储介质 |
| CN112700787B (zh) * | 2021-03-24 | 2021-06-25 | 深圳市中科蓝讯科技股份有限公司 | 一种降噪方法、非易失性可读存储介质及电子设备 |
| CN113393857B (zh) * | 2021-06-10 | 2024-06-14 | 腾讯音乐娱乐科技(深圳)有限公司 | 一种音乐信号的人声消除方法、设备及介质 |
| CN113778226A (zh) * | 2021-08-26 | 2021-12-10 | 江西恒必达实业有限公司 | 一种基于语音识别技术控制智能家居的红外ai智能眼镜 |
| JP7815786B2 (ja) * | 2022-01-21 | 2026-02-18 | ヤマハ株式会社 | 音声処理装置および音声処理方法 |
| CN116608938A (zh) * | 2023-05-22 | 2023-08-18 | 国网重庆市电力公司电力科学研究院 | 一种变电站噪声确定方法、装置、电子设备及介质 |
| CN116896706B (zh) * | 2023-07-28 | 2025-09-16 | 歌尔科技有限公司 | 信号处理方法、装置、设备及计算机可读存储介质 |
Citations (8)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US6999920B1 (en) * | 1999-11-27 | 2006-02-14 | Alcatel | Exponential echo and noise reduction in silence intervals |
| JP2013037174A (ja) * | 2011-08-08 | 2013-02-21 | Nippon Telegr & Teleph Corp <Ntt> | 雑音/残響除去装置とその方法とプログラム |
| CN103871421A (zh) * | 2014-03-21 | 2014-06-18 | 厦门莱亚特医疗器械有限公司 | 一种基于子带噪声分析的自适应降噪方法与系统 |
| CN106297817A (zh) * | 2015-06-09 | 2017-01-04 | 中国科学院声学研究所 | 一种基于双耳信息的语音增强方法 |
| CN108831500A (zh) * | 2018-05-29 | 2018-11-16 | 平安科技(深圳)有限公司 | 语音增强方法、装置、计算机设备及存储介质 |
| CN108877826A (zh) * | 2018-08-29 | 2018-11-23 | 昆明理工大学 | 一种基于多窗谱的语音减噪方法 |
| CN108986832A (zh) * | 2018-07-12 | 2018-12-11 | 北京大学深圳研究生院 | 基于语音出现概率和一致性的双耳语音去混响方法和装置 |
| CN111128213A (zh) * | 2019-12-10 | 2020-05-08 | 展讯通信(上海)有限公司 | 一种分频段进行处理的噪声抑制方法及其系统 |
Family Cites Families (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN101582264A (zh) * | 2009-06-12 | 2009-11-18 | 瑞声声学科技(深圳)有限公司 | 语音增强的方法及语音增加的声音采集系统 |
| US8433564B2 (en) * | 2009-07-02 | 2013-04-30 | Alon Konchitsky | Method for wind noise reduction |
| CN104103278A (zh) * | 2013-04-02 | 2014-10-15 | 北京千橡网景科技发展有限公司 | 一种实时语音去噪的方法和设备 |
| CN104867499A (zh) * | 2014-12-26 | 2015-08-26 | 深圳市微纳集成电路与系统应用研究院 | 一种用于助听器的分频段维纳滤波去噪方法和系统 |
| CN105280193B (zh) * | 2015-07-20 | 2022-11-08 | 广东顺德中山大学卡内基梅隆大学国际联合研究院 | 基于mmse误差准则的先验信噪比估计方法 |
-
2019
- 2019-12-10 CN CN201911278646.5A patent/CN111128213B/zh active Active
-
2020
- 2020-08-27 WO PCT/CN2020/111672 patent/WO2021114733A1/zh not_active Ceased
Patent Citations (8)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US6999920B1 (en) * | 1999-11-27 | 2006-02-14 | Alcatel | Exponential echo and noise reduction in silence intervals |
| JP2013037174A (ja) * | 2011-08-08 | 2013-02-21 | Nippon Telegr & Teleph Corp <Ntt> | 雑音/残響除去装置とその方法とプログラム |
| CN103871421A (zh) * | 2014-03-21 | 2014-06-18 | 厦门莱亚特医疗器械有限公司 | 一种基于子带噪声分析的自适应降噪方法与系统 |
| CN106297817A (zh) * | 2015-06-09 | 2017-01-04 | 中国科学院声学研究所 | 一种基于双耳信息的语音增强方法 |
| CN108831500A (zh) * | 2018-05-29 | 2018-11-16 | 平安科技(深圳)有限公司 | 语音增强方法、装置、计算机设备及存储介质 |
| CN108986832A (zh) * | 2018-07-12 | 2018-12-11 | 北京大学深圳研究生院 | 基于语音出现概率和一致性的双耳语音去混响方法和装置 |
| CN108877826A (zh) * | 2018-08-29 | 2018-11-23 | 昆明理工大学 | 一种基于多窗谱的语音减噪方法 |
| CN111128213A (zh) * | 2019-12-10 | 2020-05-08 | 展讯通信(上海)有限公司 | 一种分频段进行处理的噪声抑制方法及其系统 |
Cited By (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN113613112A (zh) * | 2021-09-23 | 2021-11-05 | 三星半导体(中国)研究开发有限公司 | 抑制麦克风的风噪的方法和电子装置 |
| CN113613112B (zh) * | 2021-09-23 | 2024-03-29 | 三星半导体(中国)研究开发有限公司 | 抑制麦克风的风噪的方法和电子装置 |
| CN113851151A (zh) * | 2021-10-26 | 2021-12-28 | 北京融讯科创技术有限公司 | 掩蔽阈值估计方法、装置、电子设备和存储介质 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN111128213B (zh) | 2022-09-27 |
| CN111128213A (zh) | 2020-05-08 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2021114733A1 (zh) | 一种分频段进行处理的噪声抑制方法及其系统 | |
| CN106486131B (zh) | 一种语音去噪的方法及装置 | |
| CN103854662B (zh) | 基于多域联合估计的自适应语音检测方法 | |
| KR101266894B1 (ko) | 특성 추출을 사용하여 음성 향상을 위한 오디오 신호를 프로세싱하기 위한 장치 및 방법 | |
| US8484020B2 (en) | Determining an upperband signal from a narrowband signal | |
| US20190172480A1 (en) | Voice activity detection systems and methods | |
| US20150081287A1 (en) | Adaptive noise reduction for high noise environments | |
| US20210193149A1 (en) | Method, apparatus and device for voiceprint recognition, and medium | |
| CN106653062A (zh) | 一种低信噪比环境下基于谱熵改进的语音端点检测方法 | |
| WO2014153800A1 (zh) | 语音识别系统 | |
| CN108428456A (zh) | 语音降噪算法 | |
| CN105144290A (zh) | 信号处理装置、信号处理方法和信号处理程序 | |
| CN110085246A (zh) | 语音增强方法、装置、设备和存储介质 | |
| CN103021405A (zh) | 基于music和调制谱滤波的语音信号动态特征提取方法 | |
| CN103295582B (zh) | 噪声抑制方法及其系统 | |
| CN105103230B (zh) | 信号处理装置、信号处理方法、信号处理程序 | |
| US9076446B2 (en) | Method and apparatus for robust speaker and speech recognition | |
| CN103971697B (zh) | 基于非局部均值滤波的语音增强方法 | |
| CN106024017A (zh) | 语音检测方法及装置 | |
| TWI749547B (zh) | 應用深度學習的語音增強系統 | |
| CN112397087A (zh) | 共振峰包络估计、语音处理方法及装置、存储介质、终端 | |
| CN113593604B (zh) | 检测音频质量方法、装置及存储介质 | |
| CN101154383A (zh) | 噪声抑制、提取语音特征、语音识别及训练语音模型的方法和装置 | |
| CN112216285A (zh) | 多人会话检测方法、系统、移动终端及存储介质 | |
| CN107919136B (zh) | 一种基于高斯混合模型的数字语音采样频率估计方法 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 20898229 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 20898229 Country of ref document: EP Kind code of ref document: A1 |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 20898229 Country of ref document: EP Kind code of ref document: A1 |
