WO2020143670A1 - 语音去混响的方法及装置 - Google Patents
语音去混响的方法及装置 Download PDFInfo
- Publication number
- WO2020143670A1 WO2020143670A1 PCT/CN2020/070922 CN2020070922W WO2020143670A1 WO 2020143670 A1 WO2020143670 A1 WO 2020143670A1 CN 2020070922 W CN2020070922 W CN 2020070922W WO 2020143670 A1 WO2020143670 A1 WO 2020143670A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- speech
- reverberation
- spectrum
- signal
- frequency response
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L21/00—Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
- G10L21/02—Speech enhancement, e.g. noise reduction or echo cancellation
- G10L21/0208—Noise filtering
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L21/00—Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
- G10L21/02—Speech enhancement, e.g. noise reduction or echo cancellation
- G10L21/0208—Noise filtering
- G10L21/0216—Noise filtering characterised by the method used for estimating noise
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L21/00—Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
- G10L21/02—Speech enhancement, e.g. noise reduction or echo cancellation
- G10L21/0208—Noise filtering
- G10L21/0216—Noise filtering characterised by the method used for estimating noise
- G10L21/0224—Processing in the time domain
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L21/00—Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
- G10L21/02—Speech enhancement, e.g. noise reduction or echo cancellation
- G10L21/0208—Noise filtering
- G10L2021/02082—Noise filtering the noise being echo, reverberation of the speech
Definitions
- the present disclosure relates to the technical field of speech signal processing, and in particular, to a method and device for speech dereverberation.
- voice dereverberation technology has been widely used in hands-free phones, hearing aids, teleconferencing systems, high-fidelity voice control systems and automatic voice recognition systems.
- the method of speech dereverberation also known as reverberation elimination method, is generally divided into three categories, but the methods of speech dereverberation in related technologies have the following problems respectively:
- the first type of de-reverberation technology based on microphone array processing its performance is limited by the number of microphones in the array. To obtain satisfactory de-reverberation results, a large number of microphones are inevitably required, which leads to increased cost and structure of the actual product The difficulty of design increases.
- the second type of dereverberation technology that suppresses the post-reverberation signal in the frequency domain needs to first estimate the reverberation time parameter (RT60) of the working environment, but due to the lack of frequency-related reverberation time in the working environment Parameter (RT60) high-precision real-time estimation algorithm, so the dereverberation performance of this technology is limited.
- the third type of WPE method that can be practically used in the dereverberation technology based on the inverse filtering idea involves a pseudo-inverse operation of the correlation matrix of high-order observation data, so it usually consumes more computing resources when implemented on a commercial DSP.
- Some embodiments of the present disclosure provide a method and device for speech dereverberation, to solve the related art speech dereverberation technology has higher cost, more complicated structural design, limited dereverberation capability, and more expensive calculation The problem of resources.
- some embodiments of the present disclosure provide a method for speech dereverberation, including:
- the time domain dereverberation speech signal is acquired.
- the normalizing the spectrum signal of the reverberation speech to obtain a normalized spectrum diagram of the reverberation speech includes:
- X norm (k,t) is the normalized spectrum of reverberation speech
- X(k,t) is the spectral signal of reverberation speech
- X max is the maximum value of X(k,t)
- M Is a preset positive integer, and M is the number of gray level levels
- k is the index of the discrete frequency
- t is the index of the time-domain signal frame.
- performing two-dimensional filtering in the time-frequency domain on the normalized spectrogram of the reverberation speech to obtain the frequency response of the reverb-like system includes:
- H(k,t) X norm (k,t)*G(k,t), obtain the frequency response of the reverb-like system
- H (k, t) is the frequency response of a reverb-like system
- X norm (k, t) is the normalized spectral spectrum of reverb speech
- * is the linear convolution operator
- ⁇ is the filter radius, and the larger the ⁇ , the sharper the image
- k is the index of the discrete frequency
- t is the index of the time-domain signal frame.
- the post-processing the frequency response of the reverb-like system to obtain the frequency response estimate of the reverberation system includes:
- H(k,t) is the frequency response of the reverberation-like system
- H min is the minimum value of H(k, t)
- H max is the maximum value of H(k, t)
- ⁇ is the decreasing function
- k is the index of the discrete frequency
- t is the time domain signal Frame index.
- the acquiring the speech spectrum of the dereverberated speech signal according to the speech signal of the reverberated speech and the frequency response estimation of the reverberation system includes:
- X(k,t) is the spectrum signal of the reverberation speech
- It is the frequency response estimation of the reverberation system
- k is the index of the discrete frequency
- t is the index of the time-domain signal frame.
- Some embodiments of the present disclosure also provide a voice dereverberation device, including a memory, a processor, and a computer program stored on the memory and executable on the processor; wherein, the processor executes the The computer program implements the following steps:
- the time domain dereverberation speech signal is acquired.
- the processor implements the following steps when executing the computer program:
- X norm (k,t) is the normalized spectrum of reverberation speech
- X(k,t) is the spectral signal of reverberation speech
- X max is the maximum value of X(k,t)
- M Is a preset positive integer, and M is the number of gray level levels
- k is the index of the discrete frequency
- t is the index of the time-domain signal frame.
- the processor implements the following steps when executing the computer program:
- H(k,t) X norm (K,t)*G(k,t), obtain the frequency response of the reverb-like system
- H (k, t) is the frequency response of a reverb-like system
- X norm (k, t) is the normalized spectral spectrum of reverb speech
- * is the linear convolution operator
- ⁇ is the filter radius, and the larger the ⁇ , the sharper the image
- k is the index of the discrete frequency
- t is the index of the time-domain signal frame.
- the processor implements the following steps when executing the computer program:
- H(k,t) is the frequency response of the reverberation-like system
- H min is the minimum value of H(k, t)
- H max is the maximum value of H(k, t)
- ⁇ is the decreasing function
- k is the index of the discrete frequency
- t is the time domain signal Frame index.
- the processor implements the following steps when executing the computer program:
- X(k,t) is the spectrum signal of the reverberation speech
- It is the frequency response estimation of the reverberation system
- k is the index of the discrete frequency
- t is the index of the time-domain signal frame.
- Some embodiments of the present disclosure also provide a computer-readable storage medium on which a computer program is stored, wherein, when the computer program is executed by a processor, the steps in the above-mentioned method for de-reverberation of speech are implemented.
- Some embodiments of the present disclosure also provide a voice dereverberation device, including:
- the first obtaining module is used to obtain the spectral signal of the reverberation speech
- a second acquisition module configured to normalize the spectrum signal of the reverberation speech to obtain a normalized spectrum diagram of the reverberation speech
- a third acquiring module configured to perform two-dimensional filtering in the time-frequency domain on the normalized spectrogram of the reverberation speech to obtain the frequency response of the reverb-like system
- a fourth obtaining module which is used to post-process the frequency response of the reverb-like system to obtain the frequency response estimate of the reverberation system;
- a fifth acquisition module configured to acquire the spectrum of the dereverberated speech signal based on the spectrum signal of the reverberated speech and the frequency response estimate of the reverberation system;
- the sixth obtaining module is configured to obtain a time-domain dereverberation speech signal according to the language spectrum of the dereverberation speech signal.
- the second obtaining module is used to:
- X norm (k,t) is the normalized spectrum of reverberation speech
- X(k,t) is the spectral signal of reverberation speech
- X max is the maximum value of X(k,t)
- M Is a preset positive integer, and M is the number of gray level levels
- k is the index of the discrete frequency
- t is the index of the time-domain signal frame.
- the third obtaining module is configured to:
- H(K,t) X norm (k,t)*G(k,t), obtain the frequency response of the reverb-like system
- H (k, t) is the frequency response of a reverb-like system
- X norm (k, t) is the normalized spectral spectrum of reverb speech
- * is the linear convolution operator
- ⁇ is the filter radius, and the larger the ⁇ , the sharper the image
- k is the index of the discrete frequency
- t is the index of the time-domain signal frame.
- the fourth obtaining module is used to:
- H(k,t) is the frequency response of the reverberation-like system
- H min is the minimum value of H(k, t)
- H max is the maximum value of H(k, t)
- ⁇ is the decreasing function
- k is the index of the discrete frequency
- t is the time domain signal Frame index.
- the fifth obtaining module is configured to:
- X(k,t) is the spectrum signal of the reverberation speech
- It is the frequency response estimation of the reverberation system
- k is the index of the discrete frequency
- t is the index of the time-domain signal frame.
- FIG. 1 is a schematic flowchart of a method for speech dereverberation according to some embodiments of the present disclosure
- Figure 2 shows a schematic diagram of the principle of applying the SR model to the dereverberation processing of speech signals
- Figure 3 shows a schematic diagram of the first implementation of the transformation function ⁇
- FIG. 4 shows a schematic diagram of a second implementation manner of the transformation function ⁇
- FIG. 5 is a schematic block diagram of a device for speech dereverberation according to some embodiments of the present disclosure
- FIG. 6 shows a schematic structural diagram of a voice dereverberation device according to some embodiments of the present disclosure.
- Speech dereverberation methods also known as reverberation elimination methods, are generally divided into three categories:
- the first type is the use of microphone array processing technology, which first estimates the orientation of the sound source relative to the microphone array (Direction of Arrival, DOA), by controlling the directionality of the microphone array to enhance the direct signal component from the direction of the sound source, and reduce And to eliminate the signal components reflected from the sound source from other directions, so as to achieve the purpose of dereverberation, in order to obtain a satisfactory dereverberation effect, the technology usually requires a large number of microphones, so that the array can obtain sufficient directional gain.
- DOA Direction of Arrival
- the second type of dereverberation technology is a method of suppressing the post-reverberation signal in the frequency domain.
- This method first estimates the reverberation time parameter (RT60) of the working environment, and estimates the power of the post-reverberation signal based on this Spectrum, and then apply the spectral subtraction in noise suppression to the post-reverberation signal.
- RT60 reverberation time parameter
- the technology does not involve the phase information of the signal and its processing performance is relatively robust, because of the lack of work environment
- the high-precision real-time estimation algorithm of the reverberation time parameter (RT60) associated with frequency so the dereverberation performance of this technology is limited.
- the third type of dereverberation technology is based on the idea of inverse filtering. Its goal is to estimate the inverse filter of the room impulse response (RIR) that causes reverberation, and use it to filter the reverberation speech signal.
- RTF room transfer function
- RTF (or its equivalent inverse filter) is time-varying and unknown, and needs to be estimated from the obtained observation data.
- DLP Delayed Linear Prediction
- the present disclosure addresses the problems of voice de-reverberation technology in the related art, such as higher cost, more complicated structural design, limited de-reverberation capability, and more computational resources for implementation, and provides a voice de-reverberation technology Method and device.
- the present disclosure proposes a novel and practical dereverberation method based on the Surrounding Retinex (SR) model. Compared with the above traditional dereverberation technology, this method has more reasonable computational complexity and can achieve effective dereverberation. Even in a severe reverberation environment, the performance of Automatic Speech Recognition (ASR) It has also improved.
- SR Surrounding Retinex
- the present disclosure proposes a novel and practical dereverberation technology based on the Surrounding Retinex (SR) model. Compared with the above traditional dereverberation technology, it has more reasonable calculation complexity and can achieve effective dereverberation. Even in a severe reverberation environment, the performance of Automatic Speech Recognition (ASR) also has some performance. improve.
- SR Surrounding Retinex
- the main idea of the present disclosure is that since the surrounding retinal cortex (SR) model has been proven to be an effective image enhancement tool, which can estimate the image illumination source from the degraded image, given the time-frequency domain as a reverberant speech signal An effectively characterized "Spectrogram" can be regarded as similar to a contaminated scene lighting image.
- SR retinal cortex
- a natural idea is to use the SR model to estimate the reverb from the "Spectragram” of the reverberated speech signal "Speech map" of the speech signal, and then obtain the time domain dereverberation speech signal.
- the "retinal cortex” model is an image enhancement theory based on the human visual system proposed by famous scholars EHLand and JJ McConn in 1971. This theory states that although the amount of visual light reaching the eye depends on reflectance and illumination, it is a natural The perceived image in the scene has a strong correlation with the reflectivity. In other words, even under difficult lighting conditions, the human visual system can perceive colors by relying on the reflectivity of the scene and ignoring the scene lighting. This theory is based on the reflectivity image model, which can be expressed mathematically as:
- F (x, y) represents a perceived image of a natural scene
- R (x, y) represents a reflected image, which depends only on the reflectivity of the scene surface, corresponding to the reflected brightness of high frequency
- I (x, y ) Represents an illumination image, which is determined by the illumination light source and is related to the amount of illumination, corresponding to the low-frequency brightness.
- the key technology of the SR model is to estimate the illumination image I(x,y) based on the perceived image F(x,y).
- D.J. Jobson et al. suggested that the illumination image I(x, y) can be estimated as a blurring scheme for the perceptual image F(x, y), which is expressed as:
- * is a linear convolution operator
- G(x, y) is a smooth kernel
- G(x, y) is usually taken as the Gauss kernel form of the following formula 3:
- ⁇ is the filter radius, the larger the ⁇ , the sharper the image;
- ⁇ is the normalization coefficient constant, so that the total integral of the second half of the formula is equal to 1. Therefore, the estimation of the reflection image R(x, y) can be expressed as:
- a reverberation time-domain digital voice signal x(n) can be mathematically characterized as:
- * is a linear convolution operator
- s(n) is the source voice digital signal
- h(n) is the channel impulse response between the signal source and the microphone. It can be seen that the reverberated speech signal x(n) is a linear convolution of the “clean” speech signal s(n) and the impulse response h(n).
- STDFT short-time discrete Fourier change
- X(k,t), S(k,t) and H(k,t) are STDFT of signals x(n), s(n) and h(n) respectively
- k is the index of discrete frequency
- t is Time domain signal frame index.
- the functional block diagram is shown in Figure 1, where the "STDFT” module converts the time-domain reverberant speech x(n) into a spectral signal X(k ,t);
- the “Normal Spectrum Processing” module normalizes the spectral signal X(k,t) to a signal X norm (k,t) with M gray level levels, and then infers using formula three
- X(k,t) For the reverberation speech signal spectrogram X(k,t), let max ⁇
- X norm (k,t) is defined as:
- the "post-processing transformer ⁇ " module is designed to achieve this conversion, here the conversion function ⁇ is defined as:
- the realization part of the left half of the parabolic curve in FIG. 2 is used as the transformation function ⁇ , where the specific definition function of the parabolic curve in FIG. 2 is: A1 is the preset constant; B1 is the preset constant, and H 0 is the minimum value of the parabolic curve.
- a section of straight line in FIG. 3 is used as the transformation function ⁇ , wherein the specific definition function of the straight line in FIG. 3 is: A2 is the preset constant; B2 is the preset constant.
- the “Dereverberation Speech Spectrum Estimator” module calculates the spectrum of the dereverberation speech signal according to Equation 9.
- ISDFT Inverse Short-Time Discrete Fourier Transform
- some embodiments of the present disclosure provide a method for speech dereverberation, including:
- Step 41 Acquire the spectral signal of the reverberant speech
- Step 42 Perform normalization processing on the spectrum signal of the reverberation speech to obtain a normalized spectrum diagram of the reverberation speech;
- Step 43 Perform two-dimensional filtering in the time-frequency domain on the normalized spectrogram of the reverberation speech to obtain the frequency response of the reverb-like system;
- Step 44 Post-process the frequency response of the reverberation system to obtain an estimate of the frequency response of the reverberation system
- Step 45 Acquire the speech spectrum of the dereverberated speech signal according to the speech signal of the reverberated speech and the frequency response estimate of the reverberation system;
- Step 46 Acquire a time-domain dereverberation speech signal according to the language spectrum of the dereverberation speech signal.
- step 42 is:
- X norm (k,t) is the normalized spectrum of reverberation speech
- X(k,t) is the spectral signal of reverberation speech
- X max is the maximum value of X(k,t)
- M Is a preset positive integer, and M is the number of gray level levels
- k is the index of the discrete frequency
- t is the index of the time-domain signal frame.
- step 43 is:
- H(k,t) X norm (k,t)*G(k,t), obtain the frequency response of the reverb-like system
- H (k, t) is the frequency response of a reverb-like system
- X norm (k, t) is the normalized spectral spectrum of reverb speech
- * is the linear convolution operator
- ⁇ is the filter radius, and the larger the ⁇ , the sharper the image
- k is the index of the discrete frequency
- t is the index of the time-domain signal frame.
- step 44 is:
- H(k,t) is the frequency response of the reverberation-like system
- H min is the minimum value of H(k, t)
- H max is the maximum value of H(k, t)
- ⁇ is the decreasing function
- k is the index of the discrete frequency
- t is the time domain signal Frame index.
- step 45 is:
- X(k,t) is the spectrum signal of the reverberation speech
- It is the frequency response estimation of the reverberation system
- k is the index of the discrete frequency
- t is the index of the time-domain signal frame.
- step 46 applying inverse short-time discrete Fourier transform (ISTDFT) to de-reverberate the speech spectrum Transform back to the time domain, you can get the time domain demixed speech signal.
- ISDFT inverse short-time discrete Fourier transform
- Some embodiments of the present disclosure provide a corresponding relationship between the SR model in the image enhancement technology and the spectrogram model of the reverberant speech signal, and accordingly apply the relevant algorithms of the SR model to the reverberant speech language Spectral diagrams to complete the dereverberation task of reverberated speech signals; in order to facilitate the application of the SR model, some embodiments of the present disclosure first convert the spectrogram X(k,t) of the reverberated speech signals into an M-level grayscale language Spectral graph Xnorm(k,t), and then use Gauss smoothing kernel function to perform two-dimensional filtering in time-frequency domain to obtain the "illumination image"H(k,t); considering the important information representation and speech spectrum in the image
- the inverse correspondence of the important information in the picture ie: the important information in the image is characterized by low-gray-scale pixels, while the important information in the speech spectrum is characterized by high-gray-level pixels), using post-processing
- Some embodiments of the present disclosure can save de-reverberation computing resources, reduce de-reverberation costs, and thereby achieve effective de-reverberation. Even in a severe reverberation environment, the performance of automatic speech recognition is significantly improved.
- some embodiments of the present disclosure also provide a device for de-reverberation of speech, including:
- the first obtaining module 51 is used to obtain the spectral signal of the reverberation speech
- the second obtaining module 52 is configured to normalize the spectrum signal of the reverberation speech to obtain a normalized spectrum diagram of the reverberation speech;
- the third obtaining module 53 is configured to perform two-dimensional filtering in the time-frequency domain on the normalized spectrogram of the reverberation speech to obtain the frequency response of the reverb-like system;
- the fourth obtaining module 54 is used to post-process the frequency response of the reverb-like system to obtain the frequency response estimate of the reverberation system;
- the fifth obtaining module 55 is configured to obtain the spectrum of the dereverberated speech signal according to the spectrum signal of the reverberated speech and the frequency response estimate of the reverberation system;
- the sixth obtaining module 56 is configured to obtain a time-domain dereverberation voice signal according to the language spectrum of the dereverberation voice signal.
- the second obtaining module 52 is used to:
- X norm (k,t) is the normalized spectrum of reverberation speech
- X(k,t) is the spectral signal of reverberation speech
- X max is the maximum value of X(k,t)
- M Is a preset positive integer, and M is the number of gray level levels
- k is the index of the discrete frequency
- t is the index of the time-domain signal frame.
- the third obtaining module 53 is used to:
- H(k,t) X norm (k,t)*G(k,t), obtain the frequency response of the reverb-like system
- H (k, t) is the frequency response of a reverb-like system
- X norm (k, t) is the normalized spectral spectrum of reverb speech
- * is the linear convolution operator
- ⁇ is the filter radius, and the larger the ⁇ , the sharper the image
- k is the index of the discrete frequency
- t is the index of the time-domain signal frame.
- the fourth obtaining module 54 is used to:
- H(k,t) is the frequency response of the reverberation-like system
- H min is the minimum value of H(k, t)
- H max is the maximum value of H(k, t)
- ⁇ is the decreasing function
- k is the index of the discrete frequency
- t is the time domain signal Frame index.
- the fifth obtaining module 55 is used to:
- X(k,t) is the spectrum signal of the reverberation speech
- It is the frequency response estimation of the reverberation system
- k is the index of the discrete frequency
- t is the index of the time-domain signal frame.
- the embodiment of the device is one-to-one corresponding to the above method embodiment. All the implementations in the above method embodiment are applicable to the embodiment of the device, and the same technical effect can also be achieved.
- some embodiments of the present disclosure also provide a device for speech dereverberation, including a processor 61, a memory 62, and a computer program stored on the memory 62 and executable on the processor 61 ;
- the processor 61 is used to read the program in the memory, perform the following process:
- the time domain dereverberation speech signal is acquired.
- the bus architecture may include any number of interconnected buses and bridges. Specifically, one or more processors represented by the processor 61 and various circuits of the memory represented by the memory 62 are linked together.
- the bus architecture can also link various other circuits such as peripheral devices, voltage regulators, and power management circuits, etc., which are well known in the art, and therefore, they will not be further described in this article.
- the bus interface provides an interface. For different devices, the processor 61 is responsible for managing the bus architecture and general processing, and the memory 62 can store data used by the processor 61 when performing operations.
- the processor 61 implements the following steps when executing the computer program:
- X norm (k,t) is the normalized spectrum of reverberation speech
- X(k,t) is the spectral signal of reverberation speech
- X max is the maximum value of X(k,t)
- M Is a preset positive integer, and M is the number of gray level levels
- k is the index of the discrete frequency
- t is the index of the time-domain signal frame.
- the processor 61 implements the following steps when executing the computer program:
- H(k,t) X norm (k,t)*G(k,t), obtain the frequency response of the reverb-like system
- H (k, t) is the frequency response of a reverb-like system
- X norm (k, t) is the normalized spectral spectrum of reverb speech
- * is the linear convolution operator
- ⁇ is the filter radius, and the larger the ⁇ , the sharper the image
- k is the index of the discrete frequency
- t is the index of the time-domain signal frame.
- the processor 61 implements the following steps when executing the computer program:
- H(k,t) is the frequency response of the reverberation-like system
- H min is the minimum value of H(k, t)
- H max is the maximum value of H(k, t)
- ⁇ is the decreasing function
- k is the index of the discrete frequency
- t is the time domain signal Frame index.
- the processor 61 implements the following steps when executing the computer program:
- X(k,t) is the spectrum signal of the reverberation speech
- It is the frequency response estimation of the reverberation system
- k is the index of the discrete frequency
- t is the index of the time-domain signal frame.
- Some embodiments of the present disclosure also provide a computer-readable storage medium on which a computer program is stored, and when the computer program is executed by a processor, the above-described method of speech dereverberation is implemented.
Landscapes
- Engineering & Computer Science (AREA)
- Computational Linguistics (AREA)
- Quality & Reliability (AREA)
- Signal Processing (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Human Computer Interaction (AREA)
- Physics & Mathematics (AREA)
- Acoustics & Sound (AREA)
- Multimedia (AREA)
- Circuit For Audible Band Transducer (AREA)
Abstract
一种语音去混响的方法,包括:获取混响语音的语谱信号(41);对混响语音的语谱信号进行归一化处理,得到混响语音的归一化语谱图(42);对混响语音的归一化语谱图进行时频域二维滤波,获取类混响系统频响(43);对类混响系统频响进行后处理,获取混响系统频响估计(44);根据混响语音的语谱信号和混响系统频响估计,获取去混响语音信号的语谱(45);根据去混响语音信号的语谱,获取时域去混响语音信号(46)。
Description
相关申请的交叉引用
本申请主张在2019年1月8日在中国提交的中国专利申请号No.201910016620.7的优先权,其全部内容通过引用包含于此。
本公开涉及语音信号处理技术领域,特别涉及一种语音去混响的方法及装置。
众所周知,语音去混响技术已广泛应用于免提电话、助听器、电话会议系统、高保真语音控制系统和自动语音识别系统。语音去混响方法又称混响消除法,通常大致分为有三大类,但是相关技术中的语音去混响方法又分别存在下列问题:
第一类基于麦克风阵列处理的去混响技术,其性能受限于阵列的麦克风数目,要获得令人满意的去混响结果,势必需要大量的麦克风,这便导致实际产品的成本提高和结构设计的困难增加。第二类在频域对后混响信号进行抑制处理的去混响技术需要首先估计出工作环境的混响时间参数(RT60),但由于目前尚缺乏关于工作环境中与频率关联的混响时间参数(RT60)的高精度实时估计算法,故该技术的去混响性能受限。第三类基于逆滤波思想的去混响技术中能实际应用的WPE方法涉及一个高阶观测数据相关矩阵的伪逆运算,因而在商用DSP上实现时通常耗费较多的计算资源。
发明内容
本公开一些实施例提供一种语音去混响的方法及装置,以解决相关技术中的语音去混响技术存在成本较高、结构设计较为复杂、去混响能力受限、实现耗费较多计算资源的问题。
为了解决上述技术问题,本公开一些实施例提供一种语音去混响的方法, 包括:
获取混响语音的语谱信号;
对所述混响语音的语谱信号进行归一化处理,得到混响语音的归一化语谱图;
对所述混响语音的归一化语谱图进行时频域二维滤波,获取类混响系统频响;
对所述类混响系统频响进行后处理,获取混响系统频响估计;
根据所述混响语音的语谱信号和所述混响系统频响估计,获取去混响语音信号的语谱;
根据所述去混响语音信号的语谱,获取时域去混响语音信号。
可选地,所述对所述混响语音的语谱信号进行归一化处理,得到混响语音归一化语谱图,包括:
其中,X
norm(k,t)为混响语音的归一化语谱图;X(k,t)为混响语音的语谱信号;X
max为X(k,t)的最大值;M为预设的正整数,且M为灰度电平等级的个数;
为向下取整函数;k为离散频率的索引;t为时域信号帧索引。
可选地,所述对所述混响语音的归一化语谱图进行时频域二维滤波,获取类混响系统频响,包括:
根据公式:H(k,t)=X
norm(k,t)*G(k,t),获取类混响系统频响;
其中,H(k,t)为类混响系统频响;X
norm(k,t)为混响语音的归一化语谱图;*为线性卷积算子;G(k,t)为平滑核,且
∫∫G(k,t)dtdk=1,β为归一化系数常数,且β使得
的全积分等于1;α为滤波半径,且α越大,图像越锐化;k为离散频率的索引;t为时域信号帧索引。
可选地,所述对所述类混响系统频响进行后处理,获取混响系统频响估计,包括:
根据公式:
其中,
为混响系统频响估计;H(k,t)为类混响系统频响;
为
的最小值;
为
的最大值;H
min为H(k,t)的最小值;H
max为H(k,t)的最大值;Ф{·}为递减函数;k为离散频率的索引;t为时域信号帧索引。
可选地,所述根据所述混响语音的语谱信号和所述混响系统频响估计,获取去混响语音信号的语谱,包括:
本公开一些实施例还提供一种语音去混响的装置,包括存储器、处理器及存储在所述存储器上并可在所述处理器上运行的计算机程序;其中,所述处理器执行所述计算机程序时实现以下步骤:
获取混响语音的语谱信号;
对所述混响语音的语谱信号进行归一化处理,得到混响语音的归一化语谱图;
对所述混响语音的归一化语谱图进行时频域二维滤波,获取类混响系统频响;
对所述类混响系统频响进行后处理,获取混响系统频响估计;
根据所述混响语音的语谱信号和所述混响系统频响估计,获取去混响语音信号的语谱;
根据所述去混响语音信号的语谱,获取时域去混响语音信号。
可选地,所述处理器执行所述计算机程序时实现以下步骤:
其中,X
norm(k,t)为混响语音的归一化语谱图;X(k,t)为混响语音的语谱信号;X
max为X(k,t)的最大值;M为预设的正整数,且M为灰度电平等级的 个数;
为向下取整函数;k为离散频率的索引;t为时域信号帧索引。
可选地,所述处理器执行所述计算机程序时实现以下步骤:
根据公式:H(k,t)=X
norm(K,t)*G(k,t),获取类混响系统频响;
其中,H(k,t)为类混响系统频响;X
norm(k,t)为混响语音的归一化语谱图;*为线性卷积算子;G(k,t)为平滑核,且
∫∫G(k,t)dtdk=1,β为归一化系数常数,且β使得
的全积分等于1;α为滤波半径,且α越大,图像越锐化;k为离散频率的索引;t为时域信号帧索引。
可选地,所述处理器执行所述计算机程序时实现以下步骤:
根据公式:
其中,
为混响系统频响估计;H(k,t)为类混响系统频响;
为
的最小值;
为
的最大值;H
min为H(k,t)的最小值;H
max为H(k,t)的最大值;Ф{·}为递减函数;k为离散频率的索引;t为时域信号帧索引。
可选地,所述处理器执行所述计算机程序时实现以下步骤:
本公开一些实施例还提供一种计算机可读存储介质,其上存储有计算机程序,其中,所述计算机程序被处理器执行时实现上述的语音去混响的方法中的步骤。
本公开一些实施例还提供一种语音去混响的装置,包括:
第一获取模块,用于获取混响语音的语谱信号;
第二获取模块,用于对所述混响语音的语谱信号进行归一化处理,得到混响语音的归一化语谱图;
第三获取模块,用于对所述混响语音的归一化语谱图进行时频域二维滤 波,获取类混响系统频响;
第四获取模块,用于对所述类混响系统频响进行后处理,获取混响系统频响估计;
第五获取模块,用于根据所述混响语音的语谱信号和所述混响系统频响估计,获取去混响语音信号的语谱;
第六获取模块,用于根据所述去混响语音信号的语谱,获取时域去混响语音信号。
可选地,所述第二获取模块,用于:
其中,X
norm(k,t)为混响语音的归一化语谱图;X(k,t)为混响语音的语谱信号;X
max为X(k,t)的最大值;M为预设的正整数,且M为灰度电平等级的个数;
为向下取整函数;k为离散频率的索引;t为时域信号帧索引。
可选地,所述第三获取模块,用于:
根据公式:H(K,t)=X
norm(k,t)*G(k,t),获取类混响系统频响;
其中,H(k,t)为类混响系统频响;X
norm(k,t)为混响语音的归一化语谱图;*为线性卷积算子;G(k,t)为平滑核,且
∫∫G(k,t)dtdk=1,β为归一化系数常数,且β使得
的全积分等于1;α为滤波半径,且α越大,图像越锐化;k为离散频率的索引;t为时域信号帧索引。
可选地,所述第四获取模块,用于:
其中,
为混响系统频响估计;H(k,t)为类混响系统频响;
为
的最小值;
为
的最大值;H
min为H(k,t)的最小值;H
max为H(k,t)的最大值;Ф{·}为递减函数;k为离散频率的索引;t为时域信号帧索引。
可选地,所述第五获取模块,用于:
本公开的有益效果是:
上述方案,通过对混响语音的语谱信号进行归一化处理,得到混响语音的归一化语谱图,在对混响语音的归一化语谱图进行时频域二维滤波,获取类混响系统频响,然后对类混响系统频响进行后处理,获取混响系统频响估计,最后得到时域去混响语音信号,此种方式,可以节省去混响的计算资源,降低了去混响成本,可以实现有效的去混响,即使在严重混响的环境中,自动语音识别的性能也有明显的提高。
图1表示本公开一些实施例的语音去混响的方法的流程示意图;
图2表示应用SR模型来进行语音信号去混响处理的原理示意图;
图3表示变换函数ψ的第一种实现方式示意图;
图4表示变换函数ψ的第二种实现方式示意图;
图5表示本公开一些实施例的语音去混响的装置的模块示意图;
图6表示本公开一些实施例的语音去混响的装置的结构示意图。
为使本公开的目的、技术方案和优点更加清楚,下面将结合附图及具体实施例对本公开进行详细描述。
在进行本公开一些实施例的说明时,首先对下面描述中所用到的一些概念进行解释说明。
语音去混响方法又称混响消除法,通常大致分为有三大类:
第一类是采用麦克风阵列处理技术,该技术首先估计声源相对麦克风阵列的方位(Direction of Arrival,DOA),通过控制麦克风阵列的方向性来增强来自声源方向的直达信号成分,并减小和消除来自其它方向的声源反射信号成分,从而达到去混响的目的,为了获得满意的去混响效果,该技术通 常需要大量数目的麦克风,以便阵列获得充分的方向性增益。
第二类去混响技术则是在频域对后混响信号进行抑制处理的方法,该方法首先估计出工作环境的混响时间参数(RT60),并据此估计出后混响信号的功率谱,然后应用噪声抑制中的谱减法对后混响信号进行抑制处理,尽管该技术不涉及信号的相位信息而使其处理性能具有较好的鲁棒性,但由于目前尚缺乏关于工作环境中与频率关联的混响时间参数(RT60)的高精度实时估计算法,故该技术的去混响性能受限。
第三类去混响技术则是基于逆滤波的思想,其目标是估计出引发混响的室内冲激响应(Room Impulse Response,RIR)的逆滤波器,用其对混响语音信号进行滤波处理以恢复源信号,在声源到麦克风的室内传递函数(Room Transfer Function,RTF)已知的情况下,用RTF的逆滤波器可以从观测的混响信号中精确地恢复出其源信号,业已证明:在麦克风数目大于已激活的声源数目、并且每个声源到每个麦克风的RTF不存在共同的零点的条件下,上述功能的逆滤波器解是存在的。然而在实际应用中,RTF(或其等效的逆滤波器)是时变的、未知的,需要从已获的观测数据中估计出。为此,大量学者致力于该领域的探索和研究,提出了许多方法,最为引人注目的便是基于延时的线性预测(Delayed Linear Prediction,DLP)的后混响抑制技术,该技术能有效地抑制后混响成分而未明显地损伤语音的短时相关性,但它要求DLP的滤波器阶数很高(滤波器通常有数千个系数),因而需要很长的观测数据,由此导致该技术具有很高的计算负荷,难以在商用的数字信号处理器(Digital Signal Processor,DSP)芯片上实时实现。
此外,人们还提出将时变语音信号源模型与多声道线性预测相结合来进行去混响的方法,该方法可以基于较短的观测数据有效地抑制后混响,而且对前混响也有抑制的效果,但它固有的计算复杂度致使其无法在实际中应用。最近,人们将基于DLP的去混响技术拓展到处理时变语音信号的场景,提出了一种称之为方差归一化延时的线性预测(NDLP)去混响技术,NDLP的频域实现即为著名的加权预测误差(Weighted Prediction Error,WPE)去混响算法,尽管WPE性能具有较好的鲁棒性,但它涉及一个高阶观测数据相关矩阵的伪逆运算,因而在商用DSP上实现时通常耗费较多的计算资源。
正如上面所述,本公开针对相关技术中的语音去混响技术存在成本较高、结构设计较为复杂、去混响能力受限、实现耗费较多计算资源的问题,提供一种语音去混响的方法及装置。
具体地,本公开基于环绕视网膜皮层(Surround Retinex,SR)模型,提出了一种新颖实用的去混响方法。该方法与上述传统的去混响技术相比,具有更合理的计算复杂度,可以实现有效的去混响,即使在严重混响的环境中,自动语音识别(Automatic Speech Recognition,ASR)的性能也有所提高。
下面对本公开一些实施例的实现原理进行说明如下。
本公开基于环绕视网膜皮层(Surround Retinex,SR)模型,提出了一种新颖实用的去混响技术。与上述传统的去混响技术相比,具有更合理的计算复杂度,可以实现有效的去混响,即使在严重混响的环境中,自动语音识别(Automatic Speech Recognition,ASR)的性能也有所提高。
本公开的主要思想是:既然环绕视网膜皮层(SR)模型已被证明是一种有效的图像增强工具,它能够从退化图像中估计出图像照明源,那么鉴于作为混响语音信号时-频域有效表征的“语谱图”(Spectrogram)可以看成类似于具有被污染的场景照明图像,一种自然的想法便是应用SR模型从混响语音信号的“语谱图”中估计出去混响的语音信号“语谱图”,进而获得时域去混响语音信号。
一、SR模型简介:
“视网膜皮层”模型是著名学者E.H.Land和J.J.McCann于1971年提出的一种基于人视觉系统的图像增强理论,这一理论指出:尽管到达眼睛的视觉光量取决于反射率和光照,但是一个自然场景中感知的图像与反射率有很强的相关性。换言之,即使在困难的光照情况下,人类视觉系统也能通过依靠场景的反射率和忽视场景照明的方式能感知颜色。这个理论是基于反射率图像模型,在数学上可以表述为:
公式一、F(x,y)=R(x,y)·I(x,y)
其中,F(x,y)表示一个自然场景的感知图像;R(x,y)表示一个反射图像,它仅取决于场景表面的反射率,对应于高频的反射亮度;I(x,y)表示一个光照图像,它由照明光源决定并与照明量有关,对应于低频的亮度。
SR模型的关键技术是基于感知图像F(x,y)来估计光照图像I(x,y)。D.J.Jobson等人建议:光照图像I(x,y)可以估计为感知图像F(x,y)的一种模糊方案,即利用公式二表示为:
其中,*为线性卷积算子,G(x,y)为平滑核,G(x,y)通常取为下述公式三的Gauss核形式:
其中,α是滤波半径,α越大,图像越锐化;β是为归一化系数常数,使得公式后半部分的全积分等于1。因此,反射图像R(x,y)的估计可用公式四表达为:
二、基于SR模型的语音信号去混响技术
通常情况下,一个混响时域数字语音信号x(n)数学上可表征为:
公式五:x(n)=s(n)*h(n)
其中,*为线性卷积算子,s(n)为源语音数字信号,h(n)为信号源与麦克风间的信道冲激响应。可见,混响语音信号x(n)是“干净”语音信号s(n)和冲击响应h(n)的线性卷积。
对公式五两边进行短时离散傅里叶变化(STDFT)便得:
公式六:X(k,t)=S(k,t)·H(k,t)
其中X(k,t)、S(k,t)和H(k,t)分别为信号x(n)、s(n)和h(n)的STDFT,k为离散频率的索引,t为时域信号帧索引。
比较公式一和公式六,我们可以看出SR模型和混响语音模型之间存在表1所示的对应关系。
| 数学模型 | 环绕视网膜皮层模型 | 混响语音模型 |
| 信号类别 | 图像 | 语音 |
| 模型 | F(x,y)=R(x,y)·I(x,y) | X(k,t)=S(k,t)·H(k,t) |
| 获取的信号 | 退化图像:F(x,y) | 混响语音语谱:X(k,t) |
| 源信号 | “干净”图像:R(x,y) | “干净”语音语谱:S(k,t) |
| 退化源 | 光照图像:I(x,y) | 混响系统频响:H(k,t) |
表1 SR模型和混响语音模型之间的对应关系
由此我们提出应用SR模型来进行语音信号去混响处理的算法,其原理框图如图1所示,其中“STDFT”模块将时域混响语音x(n)转化为语谱信号X(k,t);“语谱归一化处理”模块将语谱信号X(k,t)归一化为具有M个灰度电平等级的信号X
norm(k,t),然后用公式三推理得到的Gauss平滑核G(k,t)(注:这里取x=k,y=t)对X
norm(k,t)进行时-频域二维滤波获得类似于SR模型中“光照图像”的类混响系统频响H(k,t)。
作为一个实施例,我们给出混响语音语谱图的一种归一化方法如下:
对混响语音信号语谱图X(k,t)而言,记max{|X(k,t)|}为Xmax,那么X(k,t)对应的具有M个灰度电平的归一化语谱图X
norm(k,t)定义为:
考虑到光照图像中的重要信息(例如人脸图像中眼和嘴)通常由低灰度值像素来表征,而语音信号语谱图中的高灰度值像素则代表了重要的语音信息,那么我们需要对估计的“光照图像”H(k,t)进行后处理以便完成这种对应关系的转换,“后处理变换器ψ”模块便是为实现这种转化而设计的,这里将转换函数Ψ定义为:
如图2所示,为公式八的一种可选地实现形式,采用图2中的抛物线曲线的左半部实现部作为变换函数ψ,其中,图2中的抛物线曲线的具体定义函数为:
A1为预设常数;B1为预设常数, H
0为抛物线曲线的最小值。
下面对本公开一些实施例的具体实现过程说明如下。
如图4所示,本公开一些实施例提供一种语音去混响的方法,包括:
步骤41,获取混响语音的语谱信号;
步骤42,对所述混响语音的语谱信号进行归一化处理,得到混响语音的归一化语谱图;
步骤43,对所述混响语音的归一化语谱图进行时频域二维滤波,获取类混响系统频响;
步骤44,对所述类混响系统频响进行后处理,获取混响系统频响估计;
步骤45,根据所述混响语音的语谱信号和所述混响系统频响估计,获取去混响语音信号的语谱;
步骤46,根据所述去混响语音信号的语谱,获取时域去混响语音信号。
具体地,所述步骤42的实现方式为:
其中,X
norm(k,t)为混响语音的归一化语谱图;X(k,t)为混响语音的语谱信号;X
max为X(k,t)的最大值;M为预设的正整数,且M为灰度电平等级的个数;
为向下取整函数;k为离散频率的索引;t为时域信号帧索引。
具体地,所述步骤43的实现方式为:
根据公式十:H(k,t)=X
norm(k,t)*G(k,t),获取类混响系统频响;
其中,H(k,t)为类混响系统频响;X
norm(k,t)为混响语音的归一化语谱 图;*为线性卷积算子;G(k,t)为平滑核,且
∫∫G(k,t)dtdk=1,β为归一化系数常数,且β使得
的全积分等于1;α为滤波半径,且α越大,图像越锐化;k为离散频率的索引;t为时域信号帧索引。
具体地,所述步骤44的实现方式为:
根据上述公式八:
其中,
为混响系统频响估计;H(k,t)为类混响系统频响;
为
的最小值;
为
的最大值;H
min为H(k,t)的最小值;H
max为H(k,t)的最大值;Ф{·}为递减函数;k为离散频率的索引;t为时域信号帧索引。
具体地,所述步骤45的实现方式为:
本公开一些实施例通过对比图像增强技术中的SR模型与混响语音信号的语谱图模型,给出了两者之间的对应关系,据此将SR模型的相关算法应用于混响语音语谱图以便完成混响语音信号的去混响任务;为便于应用SR模型,本公开一些实施例首先将混响语音信号的语谱图X(k,t)转化为一个M级灰度的语谱图Xnorm(k,t),然后用Gauss平滑核函数对之进行时-频域二维滤波处理而获得“光照图像”H(k,t);考虑到图像中重要信息表征与语音语谱图中重要信息表征的逆对应关系(即:图像中重要信息由低灰度级的像素来表征,而语音语谱图中的重要信息由高灰度级的像素来表征),采用后处理变换的方式将H(k,t)转化为所需的混响系统频响
并根据混响系统频响
和混响语音的语谱信号来计算去混响语音信号的语谱,从而获得去混响的时域去混响语音信号。
本公开一些实施例可以节省去混响的计算资源,降低去混响成本,进而实现有效的去混响,即使在严重混响的环境中,自动语音识别的性能也有明显的提高。
如图5所示,本公开一些实施例还提供一种语音去混响的装置,包括:
第一获取模块51,用于获取混响语音的语谱信号;
第二获取模块52,用于对所述混响语音的语谱信号进行归一化处理,得到混响语音的归一化语谱图;
第三获取模块53,用于对所述混响语音的归一化语谱图进行时频域二维滤波,获取类混响系统频响;
第四获取模块54,用于对所述类混响系统频响进行后处理,获取混响系统频响估计;
第五获取模块55,用于根据所述混响语音的语谱信号和所述混响系统频响估计,获取去混响语音信号的语谱;
第六获取模块56,用于根据所述去混响语音信号的语谱,获取时域去混响语音信号。
进一步地,所述第二获取模块52,用于:
其中,X
norm(k,t)为混响语音的归一化语谱图;X(k,t)为混响语音的语谱信号;X
max为X(k,t)的最大值;M为预设的正整数,且M为灰度电平等级的个数;
为向下取整函数;k为离散频率的索引;t为时域信号帧索引。
进一步地,所述第三获取模块53,用于:
根据公式:H(k,t)=X
norm(k,t)*G(k,t),获取类混响系统频响;
其中,H(k,t)为类混响系统频响;X
norm(k,t)为混响语音的归一化语谱图;*为线性卷积算子;G(k,t)为平滑核,且
∫∫G(k,t)dtdk=1,β为归一化系数常数,且β使得
的全积分等于1;α为滤波半径,且α越大,图像越锐化;k为离散频率的索引;t为时域信 号帧索引。
进一步地,所述第四获取模块54,用于:
其中,
为混响系统频响估计;H(k,t)为类混响系统频响;
为
的最小值;
为
的最大值;H
min为H(k,t)的最小值;H
max为H(k,t)的最大值;Ф{·}为递减函数;k为离散频率的索引;t为时域信号帧索引。
进一步地,所述第五获取模块55,用于:
需要说明的是,该装置的实施例是与上述方法实施例一一对应的装置,上述方法实施例中所有实现方式均适用于该装置的实施例中,也能达到相同的技术效果。
如图6所示,本公开一些实施例还提供一种语音去混响的装置,包括处理器61、存储器62及存储在所述存储器62上并可在所述处理器61上运行的计算机程序;其中,所述处理器61用于读取存储器中的程序,执行下列过程:
获取混响语音的语谱信号;
对所述混响语音的语谱信号进行归一化处理,得到混响语音的归一化语谱图;
对所述混响语音的归一化语谱图进行时频域二维滤波,获取类混响系统频响;
对所述类混响系统频响进行后处理,获取混响系统频响估计;
根据所述混响语音的语谱信号和所述混响系统频响估计,获取去混响语音信号的语谱;
根据所述去混响语音信号的语谱,获取时域去混响语音信号。
需要说明的是,在图6中,总线架构可以包括任意数量的互联的总线和桥,具体由处理器61代表的一个或多个处理器和存储器62代表的存储器的各种电路链接在一起。总线架构还可以将诸如外围设备、稳压器和功率管理电路等之类的各种其他电路链接在一起,这些都是本领域所公知的,因此,本文不再对其进行进一步描述。总线接口提供接口。针对不同的装置,处理器61负责管理总线架构和通常的处理,存储器62可以存储处理器61在执行操作时所使用的数据。
可选地,所述处理器61执行所述计算机程序时实现以下步骤:
其中,X
norm(k,t)为混响语音的归一化语谱图;X(k,t)为混响语音的语谱信号;X
max为X(k,t)的最大值;M为预设的正整数,且M为灰度电平等级的个数;
为向下取整函数;k为离散频率的索引;t为时域信号帧索引。
可选地,所述处理器61执行所述计算机程序时实现以下步骤:
根据公式:H(k,t)=X
norm(k,t)*G(k,t),获取类混响系统频响;
其中,H(k,t)为类混响系统频响;X
norm(k,t)为混响语音的归一化语谱图;*为线性卷积算子;G(k,t)为平滑核,且
∫∫G(k,t)dtdk=1,β为归一化系数常数,且β使得
的全积分等于1;α为滤波半径,且α越大,图像越锐化;k为离散频率的索引;t为时域信号帧索引。
可选地,所述处理器61执行所述计算机程序时实现以下步骤:
其中,
为混响系统频响估计;H(k,t)为类混响系统频响;
为
的最小值;
为
的最大值;H
min为H(k,t)的最小值;H
max为H(k,t)的最大值;Ф{·}为递减函数;k为离散频率的索引;t为时域信号帧索引。
可选地,所述处理器61执行所述计算机程序时实现以下步骤:
本公开一些实施例还提供一种计算机可读存储介质,其上存储有计算机程序,所述计算机程序被处理器执行时实现上述的语音去混响的方法。
以上所述的是本公开的一些实施方式,应当指出对于本技术领域的普通人员来说,在不脱离本公开所述的原理前提下还可以作出若干改进和润饰,这些改进和润饰也在本公开的保护范围内。
Claims (16)
- 一种语音去混响的方法,包括:获取混响语音的语谱信号;对所述混响语音的语谱信号进行归一化处理,得到混响语音的归一化语谱图;对所述混响语音的归一化语谱图进行时频域二维滤波,获取类混响系统频响;对所述类混响系统频响进行后处理,获取混响系统频响估计;根据所述混响语音的语谱信号和所述混响系统频响估计,获取去混响语音信号的语谱;根据所述去混响语音信号的语谱,获取时域去混响语音信号。
- 一种语音去混响的装置,包括存储器、处理器及存储在所述存储器上并可在所述处理器上运行的计算机程序;其中,所述处理器执行所述计算机程序时实现以下步骤:获取混响语音的语谱信号;对所述混响语音的语谱信号进行归一化处理,得到混响语音的归一化语谱图;对所述混响语音的归一化语谱图进行时频域二维滤波,获取类混响系统频响;对所述类混响系统频响进行后处理,获取混响系统频响估计;根据所述混响语音的语谱信号和所述混响系统频响估计,获取去混响语音信号的语谱;根据所述去混响语音信号的语谱,获取时域去混响语音信号。
- 一种计算机可读存储介质,其上存储有计算机程序,其中,所述计算机程序被处理器执行时实现如权利要求1至5任一项所述的语音去混响的方法中的步骤。
- 一种语音去混响的装置,包括:第一获取模块,用于获取混响语音的语谱信号;第二获取模块,用于对所述混响语音的语谱信号进行归一化处理,得到混响语音的归一化语谱图;第三获取模块,用于对所述混响语音的归一化语谱图进行时频域二维滤波,获取类混响系统频响;第四获取模块,用于对所述类混响系统频响进行后处理,获取混响系统频响估计;第五获取模块,用于根据所述混响语音的语谱信号和所述混响系统频响估计,获取去混响语音信号的语谱;第六获取模块,用于根据所述去混响语音信号的语谱,获取时域去混响语音信号。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN201910016620.7 | 2019-01-08 | ||
| CN201910016620.7A CN109637553A (zh) | 2019-01-08 | 2019-01-08 | 一种语音去混响的方法及装置 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2020143670A1 true WO2020143670A1 (zh) | 2020-07-16 |
Family
ID=66060276
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2020/070922 Ceased WO2020143670A1 (zh) | 2019-01-08 | 2020-01-08 | 语音去混响的方法及装置 |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN109637553A (zh) |
| WO (1) | WO2020143670A1 (zh) |
Families Citing this family (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN109637553A (zh) * | 2019-01-08 | 2019-04-16 | 电信科学技术研究院有限公司 | 一种语音去混响的方法及装置 |
| CN111785292B (zh) * | 2020-05-19 | 2023-03-31 | 厦门快商通科技股份有限公司 | 一种基于图像识别的语音混响强度估计方法、装置及存储介质 |
| CN114283827B (zh) * | 2021-08-19 | 2024-03-29 | 腾讯科技(深圳)有限公司 | 音频去混响方法、装置、设备和存储介质 |
| CN117995193B (zh) * | 2024-04-02 | 2024-06-18 | 山东天意装配式建筑装备研究院有限公司 | 一种基于自然语言处理的智能机器人语音交互方法 |
Citations (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN106485674A (zh) * | 2016-09-20 | 2017-03-08 | 天津大学 | 一种基于融合技术的弱光图像增强方法 |
| CN109637553A (zh) * | 2019-01-08 | 2019-04-16 | 电信科学技术研究院有限公司 | 一种语音去混响的方法及装置 |
-
2019
- 2019-01-08 CN CN201910016620.7A patent/CN109637553A/zh active Pending
-
2020
- 2020-01-08 WO PCT/CN2020/070922 patent/WO2020143670A1/zh not_active Ceased
Patent Citations (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN106485674A (zh) * | 2016-09-20 | 2017-03-08 | 天津大学 | 一种基于融合技术的弱光图像增强方法 |
| CN109637553A (zh) * | 2019-01-08 | 2019-04-16 | 电信科学技术研究院有限公司 | 一种语音去混响的方法及装置 |
Non-Patent Citations (2)
| Title |
|---|
| MINGMING ZHANG ; WEIFENG LI ; LONGBIAO WANG ; JIANGUO WEI ; ZHIYONG WU ; QINGMIN LIAO: "Frequency-domain Dereverberation on Speech Signal using Surround Retinex", 2013 ASIA-PACIFIC SIGNAL AND INFORMATION PROCESSING ASSOCIATION ANNUAL SUMMIT AND CONFERENCE, 1 November 2013 (2013-11-01), pages 1 - 5, XP032549705, DOI: 10.1109/APSIPA.2013.6694195 * |
| XIAO, CHUNZHI ET AL: "A Speech Enhancement Algorithm Based on Speech Spectrogram", AUDIO ENGINEERING, vol. 36, no. 9, 30 September 2012 (2012-09-30), XP009521903, ISSN: 1002-8684 * |
Also Published As
| Publication number | Publication date |
|---|---|
| CN109637553A (zh) | 2019-04-16 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| CN108986838B (zh) | 一种基于声源定位的自适应语音分离方法 | |
| CN109767783B (zh) | 语音增强方法、装置、设备及存储介质 | |
| CN109643554B (zh) | 自适应语音增强方法和电子设备 | |
| CN108172231B (zh) | 一种基于卡尔曼滤波的去混响方法及系统 | |
| CN102739886B (zh) | 基于回声频谱估计和语音存在概率的立体声回声抵消方法 | |
| CN111445919B (zh) | 结合ai模型的语音增强方法、系统、电子设备和介质 | |
| JP2021128328A (ja) | 畳み込みニューラルネットワークに基づく電話音声信号の強調のための方法 | |
| US20130294611A1 (en) | Source separation by independent component analysis in conjuction with optimization of acoustic echo cancellation | |
| GB2577824A (en) | Earbud speech estimation | |
| CN111768796A (zh) | 一种声学回波消除与去混响方法及装置 | |
| CN109637553A (zh) | 一种语音去混响的方法及装置 | |
| CN113345460B (zh) | 音频信号处理方法、装置、设备及存储介质 | |
| WO2022218254A1 (zh) | 语音信号增强方法、装置及电子设备 | |
| CN108986832A (zh) | 基于语音出现概率和一致性的双耳语音去混响方法和装置 | |
| WO2020168981A1 (zh) | 风噪声抑制方法及装置 | |
| CN106384588B (zh) | 基于矢量泰勒级数的加性噪声与短时混响的联合补偿方法 | |
| WO2021007841A1 (zh) | 噪声估计方法、噪声估计装置、语音处理芯片以及电子设备 | |
| JP2025503325A (ja) | レイテンシを減少させた状態での音声信号強調のための方法およびシステム | |
| WO2020124325A1 (zh) | 一种回声消除中的自适应滤波方法、装置、设备及存储介质 | |
| WO2024139120A1 (zh) | 一种用于带噪语音信号的处理恢复方法和控制系统 | |
| CN114882898A (zh) | 多通道语音信号增强方法和装置及计算机设备和存储介质 | |
| CN107045874B (zh) | 一种基于相关性的非线性语音增强方法 | |
| Chen | Noise reduction of bird calls based on a combination of spectral subtraction, Wiener filtering, and Kalman filtering | |
| Xu et al. | Two-stage unet with channel and temporal-frequency attention for multi-channel speech enhancement | |
| Van Compernolle | DSP techniques for speech enhancement |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 20738334 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 20738334 Country of ref document: EP Kind code of ref document: A1 |









