EP2312579A1 - Speech from noise separation with reference information - Google Patents
Speech from noise separation with reference information Download PDFInfo
- Publication number
- EP2312579A1 EP2312579A1 EP09173163A EP09173163A EP2312579A1 EP 2312579 A1 EP2312579 A1 EP 2312579A1 EP 09173163 A EP09173163 A EP 09173163A EP 09173163 A EP09173163 A EP 09173163A EP 2312579 A1 EP2312579 A1 EP 2312579A1
- Authority
- EP
- European Patent Office
- Prior art keywords
- signal
- mixture
- reference signal
- cues
- information
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
- 238000000926 separation method Methods 0.000 title claims description 6
- 239000000203 mixture Substances 0.000 claims abstract description 87
- 230000002452 interceptive effect Effects 0.000 claims abstract description 29
- 238000000034 method Methods 0.000 claims abstract description 25
- 238000012545 processing Methods 0.000 claims abstract description 15
- 238000011156 evaluation Methods 0.000 claims description 3
- 230000009467 reduction Effects 0.000 description 5
- 238000000605 extraction Methods 0.000 description 3
- 238000013459 approach Methods 0.000 description 2
- 230000005540 biological transmission Effects 0.000 description 2
- 230000001419 dependent effect Effects 0.000 description 2
- 230000003044 adaptive effect Effects 0.000 description 1
- 238000013528 artificial neural network Methods 0.000 description 1
- 230000002238 attenuated effect Effects 0.000 description 1
- 210000000988 bone and bone Anatomy 0.000 description 1
- 210000004556 brain Anatomy 0.000 description 1
- 230000015556 catabolic process Effects 0.000 description 1
- 238000006731 degradation reaction Methods 0.000 description 1
- 230000000694 effects Effects 0.000 description 1
- 230000008030 elimination Effects 0.000 description 1
- 238000003379 elimination reaction Methods 0.000 description 1
- 230000003203 everyday effect Effects 0.000 description 1
- 238000001914 filtration Methods 0.000 description 1
- 230000006870 function Effects 0.000 description 1
- 230000003993 interaction Effects 0.000 description 1
- 230000007246 mechanism Effects 0.000 description 1
- 230000011218 segmentation Effects 0.000 description 1
- 238000005204 segregation Methods 0.000 description 1
- 230000003595 spectral effect Effects 0.000 description 1
- 238000001228 spectrum Methods 0.000 description 1
- 238000012546 transfer Methods 0.000 description 1
- 230000009466 transformation Effects 0.000 description 1
- 230000001131 transforming effect Effects 0.000 description 1
Images
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L21/00—Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
- G10L21/02—Speech enhancement, e.g. noise reduction or echo cancellation
- G10L21/0208—Noise filtering
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L21/00—Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
- G10L21/02—Speech enhancement, e.g. noise reduction or echo cancellation
- G10L21/0208—Noise filtering
- G10L2021/02085—Periodic noise
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L21/00—Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
- G10L21/02—Speech enhancement, e.g. noise reduction or echo cancellation
- G10L21/0208—Noise filtering
- G10L21/0216—Noise filtering characterised by the method used for estimating noise
- G10L2021/02161—Number of inputs available containing the signal or the noise to be suppressed
- G10L2021/02165—Two microphones, one receiving mainly the noise signal and the other one mainly the speech signal
Definitions
- the invention generally refers to the processing of acoustically sensed signals.
- the present invention relates to a system and a method for separating a mixture signal containing a mixture of acoustical target information ("speech") and interfering information (“noise").
- speech acoustical target information
- noise interfering information
- the European patent application EP 1 879 180 A1 shows a method to reduce the background noise in speech signals with the help of a reference microphone.
- the main idea behind this method is to estimate the spectrum of the noise based on a reference microphone which captures only the noise and then subtract these spectral components from the microphone signal which captures the mixture of the speech and noise signal.
- the main disadvantage, however, of this method is that the acoustic environment where the noise is captured is normally different from that where the mixture of the noise and the speech signal are captured (e.g. the engine compartment and the passenger compartment if one wants to reduce the engine noise in a car).
- the present invention relates to a technique for reducing noise in a mixture signal containing a mixture of noise and speech by means of additional reference information captures e.g. by a second microphone.
- additional reference information captures e.g. by a second microphone.
- the present invention does not try to reduce the noise directly in the signal domain but uses techniques inspired by Computational Auditory Scene Analysis (CASA) to reduce the noise.
- CASA Computational Auditory Scene Analysis
- a system for the separation of a mixture signal containing a mixture of target information and interfering information comprises means for receiving the mixture signal, means for receiving a reference signal and a signal processing unit configured to extract cues from the reference signal and to separate the target information from the mixture signal using these cues.
- a method for separating a mixture signal containing a mixture of target information and interfering information comprises the steps of receiving the mixture signal, receiving a reference signal and extracting cues from the reference signal and separating the target information from the mixture signal using these cues.
- the means for receiving the mixture signal and the means for receiving the reference signal may comprise a microphone and a recording unit each, wherein the microphone for the mixture signal may be positioned close to the origin of the target information and the microphone for the reference signal may be positioned close the origin of the interfering information, when the means for receiving the reference signal are configured to receive interfering information.
- the interfering information can be also extracted from the speed of an engine.
- the means for receiving the reference signal may be configured to receive target information, wherein this information can be extracted from a video signal, in particular from the movement of a speaker's body or the speaker's lip movement in the video signal.
- the signal processing unit of the system may comprise means for splitting the reference signal and the mixture signal in a multitude of frequency channels, means for extracting grouping cues from the reference signal and evaluating the grouping cues in the mixture signal for each frequency channel at each instant in time and means for allocating each frequency channel of the mixture signal at each instant in time to either the target information or the interfering information and separating the mixture signal into the target information and the interfering information.
- the signal processing unit of the system comprises means for splitting the mixture signal in a multitude of frequency channels, means for extracting grouping cues from the reference signal and evaluating the grouping cues in the mixture signal at each instant in time and means for allocating each frequency channel of the mixture signal at each instant in time to either the target information or the interfering information and separating the mixture signal into the target information and the interfering information.
- the grouping cues may be the fundamental frequency or on- or off-sets.
- the target information may be speech and the interfering information may be noise.
- the system for separating a mixture signal may be included in a motorcycle helmet, wherein the means for receiving the mixture signal are positioned inside the helmet and the means for receiving the reference signal are positioned partly inside the helmet and partly close to the engine of a motorcycle, wherein the means for receiving the reference signal are connected via a cable or wireless.
- Fig. 1 shows a motorcyclist 7 driving a motorcycle 6 and wearing a helmet 4.
- the helmet 4 includes a system according to the invention.
- the system comprises a signal processing unit 1, means for receiving a mixture signal, here a microphone 2, and means for receiving a reference signal, here a microphone 3a, and a receiving unit 3b.
- the microphone 2 for receiving the mixture signal and the receiving unit 3b are connected to the signal processing unit 1 via a cable.
- the microphone 3a for receiving the reference signal is, however, not positioned in the helmet 4, but close to the engine 5 of the motorcycle 6 to be at the origin of the interfering signal, which in the shown example may be the harmonic noise generated by the engine 5.
- the transmission of the reference signal of the microphone 3a to the receiving unit 3b can be for example accomplished via a wireless transmission.
- the microphone 2 for receiving the mixture signal is positioned to the front of the helmet 4 close to the mouth of the motorcyclist 7.
- the microphone 2 is therefore positioned close to the origin of the target signal, here the acoustically sensed speech signal of the motorcyclist 7.
- the microphone 2 also receives noise of the engine 5, due to the fact that the engine 5 of the motorcycle 6 is quite loud and the engine noise is only slightly attenuated by the helmet 4. Therefore the mixture signal received by the microphone 2 contains a mixture of speech and noise.
- the signal processing unit 1 is configured to extract cues from the reference signal received by the microphone 3 and to separate the speech from the mixture signal received by the microphone 2 using the cues. A detailed description of the extraction and separation will be given in combination with the method and Fig. 3 .
- the system according to the invention is therefore able to significantly reduce the engine noise in the mixture signal. As a result of the reduction telecommunication while riding will be improved.
- FIG. 2 Another application area for the system according to the invention is shown in Fig. 2 , where a car 8 is shown including a similar system to that in Fig. 1 .
- the signal processing unit 1 and the microphone 2 for receiving the mixture signal are not positioned in a helmet, but inside the car 8.
- the microphone 2 for receiving the mixture signal is positioned in the passenger compartment to be near to the mouth of the driver 7.
- the microphone 3a for receiving the reference signal is again positioned close to the engine 5.
- the signal processing unit 1 can be positioned anywhere in the car and has connections to the microphone 2 for receiving the mixture signal and the microphone 3a for receiving the reference signal. Therefore a receiving unit 3b is not needed here.
- the system according to the invention does not only improve the headset free telecommunication but also speech based operation of devices in a car. In particular the reduction of the harmonic noise generated by the engine is here helpful.
- Fig. 3 shows a method according to an embodiment of the invention. At the beginning an acoustically sensed mixture signal and a reference signal are received (100, 101).
- the reference signal is preferably directly sensed from the origin of the noise (e.g. an engine, fan, ). In an ideal setup only the reference signal without any additional signals would be sensed such that the reference signal is available without distortions. This can best be achieved by sensing the reference signal close to its source.
- the target signal will also preferably be sensed close to its source.
- the target signal In the case of a speech signal of the driver of a car or a motorcycle sensing close to the drivers mouth would be best.
- the target signal is commonly sensed at a certain distance from its source.
- this mixture signal would be a mixture of noise generated by the engine and a speech signal of the driver.
- the mixture signal and reference signal can for example be sensed by microphones.
- both signals are split into a multitude of adjacent frequency channels 102, 103.
- auditory scene analysis cues are extracted from the reference signal ("noise signal") 105.
- These cues which are typically used in Computational Auditory Scene Analysis (CASA) systems for the separation of sources can be e.g. one or more of:
- These auditory cues provide information on the reference signal. Knowing these cues allows identifying the reference signal in the mixture signal. For doing so, these cues are extracted in the reference signal, where the reference signal is mostly undistorted and these cues can easily be extracted, and then evaluated in the mixture signal. As a result of this evaluation parts, i.e. frequency channels at each instant in time, can be identified in which the reference signal is dominating the mixture signal.
- the reference signal is transformed into the frequency domain and they are determined for each frequency channel at each instance in time.
- these cues as e.g. the fundamental frequency it is also possible to extract these cues directly in the time domain and then calculate their effect on the different frequency channels (in the case of the fundamental frequency signal parts will be present at the fundamental frequency and at its harmonics which can easily be calculated from the fundamental frequency).
- the auditory cues After extracting the auditory cues from the reference signal they are evaluated in the mixture signal (comprising e.g. noise and speech).
- the mixture signal comprising e.g. noise and speech
- the mixture signal is separated, in discrete time steps, into frequency channels "speech" and frequency channels "noise” (107).
- Figs. 1 and 2 the system according to the invention is included in a motorcycle helmet and in a car.
- a system is for example included in a robot. This would help to improve speech recognition systems in robots. Therefore robots or any other technical systems which are controlled by speech or interpret speech can be even used in loud and noisy environments.
- Another application area where the system according to the invention can be used is the field of hearing devices.
- the elimination of a noise in a mixture signal that the hearing device is receiving helps the person who uses the hearing device to even better understand the speech of other persons.
- Figs. 1 and 2 are showing a system where the reference signal uses noise from an engine.
- the reference signal information on the speech signal can be obtained e. g. by using a bone conductive microphone.
- the necessary grouping information is extracted from the speech signal and then used to separate the speech signal from the noise signal.
Landscapes
- Engineering & Computer Science (AREA)
- Computational Linguistics (AREA)
- Quality & Reliability (AREA)
- Signal Processing (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Human Computer Interaction (AREA)
- Physics & Mathematics (AREA)
- Acoustics & Sound (AREA)
- Multimedia (AREA)
- Soundproofing, Sound Blocking, And Sound Damping (AREA)
- Fittings On The Vehicle Exterior For Carrying Loads, And Devices For Holding Or Mounting Articles (AREA)
Abstract
System and method for separating a mixture signal containing a mixture of target information and interfering information comprising means (2) for receiving the mixture signal, means (3) for receiving a reference signal and a signal processing unit (1) configured to extract cues from the reference signal and to separate the target information from the mixture signal using the cues.
Description
- The invention generally refers to the processing of acoustically sensed signals.
- The present invention relates to a system and a method for separating a mixture signal containing a mixture of acoustical target information ("speech") and interfering information ("noise").
- In many everyday situations different background noise sources are present while we talk on the phone or try to operate a device via speech. This noise, however, makes speech recognition more difficult for humans and especially for machines. For instance, headset free telecommunication in a car or the operation of different devices in the car (e.g. navigation systems, radio) via speech are often interfered by the driving noise (e.g. engine noise) of the car. A similar situation is found when riding a motorcycle, because the noise generated by the motorcycle engine significantly impairs the quality of telecommunication during a ride. Another area where speech recognition is more and more used is robotics. Background noise (e.g. caused by a fan to cool the robot's CPU) is here also reducing the quality of speech recognition.
- Therefore, different approaches have already been proposed to reduce the noise and thereby improve the speech signal. The European
patent application EP 1 879 180 A1 for example shows a method to reduce the background noise in speech signals with the help of a reference microphone. The main idea behind this method is to estimate the spectrum of the noise based on a reference microphone which captures only the noise and then subtract these spectral components from the microphone signal which captures the mixture of the speech and noise signal. The main disadvantage, however, of this method is that the acoustic environment where the noise is captured is normally different from that where the mixture of the noise and the speech signal are captured (e.g. the engine compartment and the passenger compartment if one wants to reduce the engine noise in a car). As a consequence not the noise at it was present in the passenger compartment is subtracted from the corresponding mixture signal but the noise as it was recorded in the engine compartment. To remedy this, a filter is used which replicates the transformation of the signal it underwent on its way from the engine compartment to the passenger compartment. - However, this filtering operation is highly dependent on the position of the speaker and has hence to be adaptive. This, on the other hand, is problematic as no error signal is available to adapt the filter when the speaker is talking.
- There also exist approaches to enhance speech signals in a way inspired by the processing in the human brain. They are commonly referred to as Computational Auditory Scene Analysis (CASA). For humans it was observed that they are able to separate different concurrent sound sources and focus on one source. The underlying mechanisms are able to separate signals based on cues which can serve to bind or group different time frequency regions to one acoustic source. Such cues are e.g. fundamental frequency, location in space, or common on- and off-sets.
- Different systems have already been developed which are able to separate a speech signal from an interfering signal based on such cues, e.g. fundamental frequency or common on- and off-sets. For doing so these cues are estimated from the signal containing the mixture of the different sources present. This step usually comprises the split of this mixture signal into different frequency channels. Based on the mentioned cues it is determined next for each of the frequency channels at each instant in time to which of the detected sources the signal belongs. The results obtained are usually quite good but also entail significant degradations of the speech signal. One reason for this is that the cues used for separating the sound sources (e.g. fundamental frequency, on- and off-sets) have to be extracted from the mixture of the sources and hence this extraction process is error prone. Similar the number of present sources also has to be estimated from the mixture signal.
- It is therefore the object of the present invention to propose a system and a method to improve noise reduction in a signal that contains a mixture of a noise and a speech signal.
- This object is achieved by means of the features of the independent claims. The dependent claims develop further the central idea of the present invention.
- The present invention relates to a technique for reducing noise in a mixture signal containing a mixture of noise and speech by means of additional reference information captures e.g. by a second microphone. In contrast to previous methods the present invention, however, does not try to reduce the noise directly in the signal domain but uses techniques inspired by Computational Auditory Scene Analysis (CASA) to reduce the noise.
- Therefore, a system for the separation of a mixture signal containing a mixture of target information and interfering information is proposed, that comprises means for receiving the mixture signal, means for receiving a reference signal and a signal processing unit configured to extract cues from the reference signal and to separate the target information from the mixture signal using these cues.
- In addition, a method for separating a mixture signal containing a mixture of target information and interfering information is proposed, said comprises the steps of receiving the mixture signal, receiving a reference signal and extracting cues from the reference signal and separating the target information from the mixture signal using these cues.
- The means for receiving the mixture signal and the means for receiving the reference signal may comprise a microphone and a recording unit each, wherein the microphone for the mixture signal may be positioned close to the origin of the target information and the microphone for the reference signal may be positioned close the origin of the interfering information, when the means for receiving the reference signal are configured to receive interfering information. The interfering information can be also extracted from the speed of an engine.
- Furthermore the means for receiving the reference signal may be configured to receive target information, wherein this information can be extracted from a video signal, in particular from the movement of a speaker's body or the speaker's lip movement in the video signal.
- The signal processing unit of the system may comprise means for splitting the reference signal and the mixture signal in a multitude of frequency channels, means for extracting grouping cues from the reference signal and evaluating the grouping cues in the mixture signal for each frequency channel at each instant in time and means for allocating each frequency channel of the mixture signal at each instant in time to either the target information or the interfering information and separating the mixture signal into the target information and the interfering information.
- In another embodiment the signal processing unit of the system comprises means for splitting the mixture signal in a multitude of frequency channels, means for extracting grouping cues from the reference signal and evaluating the grouping cues in the mixture signal at each instant in time and means for allocating each frequency channel of the mixture signal at each instant in time to either the target information or the interfering information and separating the mixture signal into the target information and the interfering information.
- The grouping cues may be the fundamental frequency or on- or off-sets. The target information may be speech and the interfering information may be noise.
- The system for separating a mixture signal may be included in a motorcycle helmet, wherein the means for receiving the mixture signal are positioned inside the helmet and the means for receiving the reference signal are positioned partly inside the helmet and partly close to the engine of a motorcycle, wherein the means for receiving the reference signal are connected via a cable or wireless.
- These and other aspects and advantages of the present invention will become more apparent when studying the following detailed description, in connection with the figures, in which
- Fig. 1
- shows a motorcyclist with a helmet that includes a system according to the invention driving a motorcycle;
- Fig. 2
- shows a car with a driver including a system according to the invention;
- Fig. 3
- shows a method according to the invention.
-
Fig. 1 shows amotorcyclist 7 driving a motorcycle 6 and wearing ahelmet 4. Thehelmet 4 includes a system according to the invention. The system comprises asignal processing unit 1, means for receiving a mixture signal, here amicrophone 2, and means for receiving a reference signal, here a microphone 3a, and a receiving unit 3b. Themicrophone 2 for receiving the mixture signal and the receiving unit 3b are connected to thesignal processing unit 1 via a cable. The microphone 3a for receiving the reference signal is, however, not positioned in thehelmet 4, but close to the engine 5 of the motorcycle 6 to be at the origin of the interfering signal, which in the shown example may be the harmonic noise generated by the engine 5. The transmission of the reference signal of the microphone 3a to the receiving unit 3b can be for example accomplished via a wireless transmission. - The
microphone 2 for receiving the mixture signal is positioned to the front of thehelmet 4 close to the mouth of themotorcyclist 7. Themicrophone 2 is therefore positioned close to the origin of the target signal, here the acoustically sensed speech signal of themotorcyclist 7. However, themicrophone 2 also receives noise of the engine 5, due to the fact that the engine 5 of the motorcycle 6 is quite loud and the engine noise is only slightly attenuated by thehelmet 4. Therefore the mixture signal received by themicrophone 2 contains a mixture of speech and noise. - The
signal processing unit 1 is configured to extract cues from the reference signal received by themicrophone 3 and to separate the speech from the mixture signal received by themicrophone 2 using the cues. A detailed description of the extraction and separation will be given in combination with the method andFig. 3 . - The system according to the invention is therefore able to significantly reduce the engine noise in the mixture signal. As a result of the reduction telecommunication while riding will be improved.
- Another application area for the system according to the invention is shown in
Fig. 2 , where acar 8 is shown including a similar system to that inFig. 1 . However, thesignal processing unit 1 and themicrophone 2 for receiving the mixture signal are not positioned in a helmet, but inside thecar 8. In particular themicrophone 2 for receiving the mixture signal is positioned in the passenger compartment to be near to the mouth of thedriver 7. The microphone 3a for receiving the reference signal is again positioned close to the engine 5. Thesignal processing unit 1 can be positioned anywhere in the car and has connections to themicrophone 2 for receiving the mixture signal and the microphone 3a for receiving the reference signal. Therefore a receiving unit 3b is not needed here. The system according to the invention does not only improve the headset free telecommunication but also speech based operation of devices in a car. In particular the reduction of the harmonic noise generated by the engine is here helpful. -
Fig. 3 shows a method according to an embodiment of the invention. At the beginning an acoustically sensed mixture signal and a reference signal are received (100, 101). - The reference signal is preferably directly sensed from the origin of the noise (e.g. an engine, fan, ...). In an ideal setup only the reference signal without any additional signals would be sensed such that the reference signal is available without distortions. This can best be achieved by sensing the reference signal close to its source.
- In a same way the target signal will also preferably be sensed close to its source. In the case of a speech signal of the driver of a car or a motorcycle sensing close to the drivers mouth would be best. However, to allow a speech interaction where the driver does not need to wear any special device, i.e. headset free, the target signal is commonly sensed at a certain distance from its source. As a consequence only a mixture of the target signal and other sound sources is sensed. Hence, in one application of the present invention this mixture signal would be a mixture of noise generated by the engine and a speech signal of the driver. As already described above, the mixture signal and reference signal can for example be sensed by microphones.
- After receiving the reference signal and the mixture signal both signals are split into a multitude of
adjacent frequency channels 102, 103. For each frequency channel at each instant in time grouping auditory scene analysis cues are extracted from the reference signal ("noise signal") 105. These cues which are typically used in Computational Auditory Scene Analysis (CASA) systems for the separation of sources can be e.g. one or more of: - fundamental frequency
- common on- and off-sets
- common modulation / fate
- spatial cues (ITD/ILD i.e. perceived origin)
- continuity
- sequential similarity.
- These auditory cues provide information on the reference signal. Knowing these cues allows identifying the reference signal in the mixture signal. For doing so, these cues are extracted in the reference signal, where the reference signal is mostly undistorted and these cues can easily be extracted, and then evaluated in the mixture signal. As a result of this evaluation parts, i.e. frequency channels at each instant in time, can be identified in which the reference signal is dominating the mixture signal.
- For the extraction of these auditory cues the reference signal is transformed into the frequency domain and they are determined for each frequency channel at each instance in time. For some of these cues as e.g. the fundamental frequency it is also possible to extract these cues directly in the time domain and then calculate their effect on the different frequency channels (in the case of the fundamental frequency signal parts will be present at the fundamental frequency and at its harmonics which can easily be calculated from the fundamental frequency). After extracting the auditory cues from the reference signal they are evaluated in the mixture signal (comprising e.g. noise and speech). After transforming the mixture signal into the frequency domain and the cues are evaluated for each frequency channel at each instant in time (104). Then it is possible to allocate each frequency channel at each instant in time to either the speech or the noise (106). Based on this allocation the mixture signal is separated, in discrete time steps, into frequency channels "speech" and frequency channels "noise" (107).
- With this method it is now possible to measure the Computational Auditory Scene Analysis (CASA) cues from an undistorted noise signal and can then use this information to eliminate the noise in the mixture of speech and noise without the need to estimate the transfer function between the site of the recording of the noise and the mixture signal (e. g. from the engine compartment to the passenger compartment).
- In
Figs. 1 and 2 the system according to the invention is included in a motorcycle helmet and in a car. However, it is also possible that such a system is for example included in a robot. This would help to improve speech recognition systems in robots. Therefore robots or any other technical systems which are controlled by speech or interpret speech can be even used in loud and noisy environments. - Another application area where the system according to the invention can be used is the field of hearing devices. The elimination of a noise in a mixture signal that the hearing device is receiving helps the person who uses the hearing device to even better understand the speech of other persons.
- The examples in
Figs. 1 and 2 are showing a system where the reference signal uses noise from an engine. However it is also possible that for the reference signal information on the speech signal can be obtained e. g. by using a bone conductive microphone. In this case the necessary grouping information is extracted from the speech signal and then used to separate the speech signal from the noise signal. -
- [1]
- H. Puder, F. Steffens. Improved Noise Reduction for HandsFree Car Phones Utilizing Information on Vehicle and Engine Speeds. EUSIPCO, 2000
- [2]
-
EP1879180 - Reduction of background noise in hands- free Systems - [3]
- Bregman, A. Auditory Scene Analysis MIT Press, 1990
- [4]
- Brown, G. J. & Cooke, M. P. Computational Auditory Scene Analysis Computer Speech and Language, 1994, 1, 297-336
- [5]
- Heckmann, M.; Joublin, F. & Körner, E. Sound Source Separation for a Robot Based on Pitch Proc IEEE/RSJ Int . 1 Conf. on Robots and Intell. Syst., 2005, 203- 208
- [6]
- Hu, G. & Wang, D. L. Monaural Speech Segregation Based on Pitch Tracking and Amplitude Modulation IEEE Trans. Neural Networks, 2004, 15, 1135-1150
- [7]
- Hu, G. & Wang, D. Auditory segmentation based on onset and offset analysis IEEE Transactions on Audio, Speech, and Language Processing, 2007, 15, 396- 405
Claims (21)
- System for the separation of an acoustical target signal from an acoustically sensed mixture signal containing a mixture of said target signal and interfering signal, such as noise, the system comprising• means (2) for acoustically sensing the mixture signal,• means (3) for sensing a reference signal, preferably close to a known source of noise, , and• a signal processing unit (1) configured to- extract cues from the reference signal for each instance in time at for each of a plurality of preferably adjacent and continuous frequency channels, and- to separate the target signal from the mixture signal for each of said frequency channels and each instance in time, using the cues.
- System according to claim 1,
characterized in that,
the means for receiving the mixture signal comprising a microphone (2), wherein the microphone (2) is positioned close to the origin of the target information. - System according to any of claims 1 to 2,
characterized in that,
the means (3) for receiving the reference signal are configured to receive interfering information. - System according to claim 3,
characterized in that,
the means (3) for receiving the reference signal comprising a microphone (3a), wherein the microphone (3a) is positioned close to the origin of the interfering information. - System according to any of claims 1 to 2,
characterized in that,
the means (3) for receiving the reference signal are configured to receive target information. - System according to any of claims 1 to 5,
characterized in that,
the signal processing unit (1) comprises• means for splitting the reference signal and the mixture signal in a multitude of frequency channels,• means for extracting grouping cues from the reference signal and evaluating the grouping cues in the mixture signal for each frequency channel at each instant in time and• means for allocating each frequency channel of the mixture signal at each instant in time to either the target information or the interfering information and separating the mixture signal into the target information and the interfering information. - System according to any of claims 1 to 5,
characterized in that,
the signal processing unit (1) comprises• means for splitting the mixture signal in a multitude of frequency channels,• means for extracting grouping cues from the reference signal and evaluating the grouping cues in the mixture signal at each instant in time and• means for allocating each frequency channel of the mixture signal at each instant in time to either the target information or the interfering information and separating the mixture signal into the target information and the interfering information. - A robot, an air/land/sea vehicle, a voice-controlled system or an artificial hearing aid, comprising a system according to any of the preceding claims.
- A method for separating a mixture signal containing a mixture of target information and interfering information, comprising the steps• Receiving the mixture signal (100),• Receiving a reference signal (101) and• Extracting cues from the reference signal and separating the target information from the mixture signal using the cues (102-107).
- The method according to claim 9,
characterized in that,
the target information is speech and the interfering information is noise. - Method according to any of claims 9 to 10,
characterized in that,
the step of extracting cues from the reference signal and separating the target information from the mixture signal using the cues comprises the steps• Splitting the mixture signal in a multitude of frequency channels (102),• Splitting the reference signal in a multitude of frequency channels (103),• Extracting grouping cues from the reference signal for each frequency channel at each instant in time (105),• Evaluating grouping cues in the mixture signal for each frequency channel at each instant in time (104),• Allocating each frequency channel at each instant in time to either the target information or the interfering information based on the evaluation of the grouping cues (106) and• Separating the mixture of the target information and the interfering information based on the previous allocation (107). - Method according to any of claims 9 or 10,
characterized in that,
the step of extracting cues from the reference signal and separating the target information from the mixture signal using the cues comprises the steps• Splitting the mixture signal in a multitude of frequency channels,• Extracting grouping cues from the reference signal at each instant in time,• Evaluating grouping cues in the mixture signal for each frequency channel at each instant in time ,• Allocating each frequency channel at each instant in time to either the target information or the interfering information based on the evaluation of the grouping cues and• Separating the mixture of the target information and the interfering information based on the previous allocation. - The method according to any of claims 11 to 12,
characterized in that,
the fundamental frequency is used as grouping cue. - The method according to any of claims 11 to 12,
characterized in that,
on- or off-sets are used as grouping cue. - The method according to any of claims 9 to 14,
characterized in that,
the reference signal contains interfering information. - The method according to claim 15,
characterized in that,
the interfering information of the reference signal is extracted from the speed of an engine (5) . - The method according to any of claims 9 to 14,
characterized in that,
the reference signal contains target information. - The method according to claim 17,
characterized in that,
the target information of the reference signal is extracted from the movements of a speaker's body or lip movements of a speaker in a video signal. - A computer software program product,
performing a method according to any of claims 9 to 18 when run on a computing unit. - A motorcycle helmet (4), being provided with a system according to any of the claims 1 to 7.
- A motorcycle helmet according to claim 20,
characterized in that,
the means (2) for receiving the mixture signal are positioned inside the helmet (4) and the means (3) for receiving the reference signal are positioned partly (3b) inside the helmet (4) and partly (3a) close to a engine (5) of a motorcycle (6), wherein the means (3a, 3b) for receiving the reference signal are connected via a cable or wireless.
Priority Applications (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| EP09173163A EP2312579A1 (en) | 2009-10-15 | 2009-10-15 | Speech from noise separation with reference information |
| JP2010182876A JP5377442B2 (en) | 2009-10-15 | 2010-08-18 | System that separates speech from noise by reference information |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| EP09173163A EP2312579A1 (en) | 2009-10-15 | 2009-10-15 | Speech from noise separation with reference information |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP2312579A1 true EP2312579A1 (en) | 2011-04-20 |
Family
ID=41694574
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP09173163A Ceased EP2312579A1 (en) | 2009-10-15 | 2009-10-15 | Speech from noise separation with reference information |
Country Status (2)
| Country | Link |
|---|---|
| EP (1) | EP2312579A1 (en) |
| JP (1) | JP5377442B2 (en) |
Cited By (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2013041942A1 (en) * | 2011-09-20 | 2013-03-28 | Toyota Jidosha Kabushiki Kaisha | Sound source detecting system and sound source detecting method |
| EP3116236A1 (en) | 2015-07-06 | 2017-01-11 | Sivantos Pte. Ltd. | Method for processing signals for a hearing aid, hearing aid, hearing aid system and interference transmitter for a hearing aid system |
| WO2020012229A1 (en) * | 2018-07-12 | 2020-01-16 | Bosch Car Multimedia Portugal, S.A. | Selective active noise cancelling system |
| CN112614501A (en) * | 2020-12-08 | 2021-04-06 | 深圳创维-Rgb电子有限公司 | Noise reduction method, noise reduction apparatus, noise canceller, microphone, and readable storage medium |
Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US4932063A (en) * | 1987-11-01 | 1990-06-05 | Ricoh Company, Ltd. | Noise suppression apparatus |
| WO1998001956A2 (en) * | 1996-07-08 | 1998-01-15 | Chiefs Voice Incorporated | Microphone noise rejection system |
| GB2377805A (en) * | 2001-07-10 | 2003-01-22 | 20 20 Speech Ltd | Localisation of a person in a conveyance |
| EP1879180A1 (en) | 2006-07-10 | 2008-01-16 | Harman Becker Automotive Systems GmbH | Reduction of background noise in hands-free systems |
Family Cites Families (9)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP3384540B2 (en) * | 1997-03-13 | 2003-03-10 | 日本電信電話株式会社 | Receiving method, apparatus and recording medium |
| JPH11265199A (en) * | 1998-03-18 | 1999-09-28 | Nippon Telegr & Teleph Corp <Ntt> | Transmitter |
| JP4119112B2 (en) * | 2001-11-05 | 2008-07-16 | 本田技研工業株式会社 | Mixed sound separator |
| EP1605439B1 (en) * | 2004-06-04 | 2007-06-27 | Honda Research Institute Europe GmbH | Unified treatment of resolved and unresolved harmonics |
| JP2007079389A (en) * | 2005-09-16 | 2007-03-29 | Yamaha Motor Co Ltd | Speech analysis method and speech analysis apparatus |
| CN101512374B (en) * | 2006-11-09 | 2012-04-11 | 松下电器产业株式会社 | Sound source position detection device |
| JP4336378B2 (en) * | 2007-04-26 | 2009-09-30 | 株式会社神戸製鋼所 | Objective sound extraction device, objective sound extraction program, objective sound extraction method |
| JP4519901B2 (en) * | 2007-04-26 | 2010-08-04 | 株式会社神戸製鋼所 | Objective sound extraction device, objective sound extraction program, objective sound extraction method |
| JP4493690B2 (en) * | 2007-11-30 | 2010-06-30 | 株式会社神戸製鋼所 | Objective sound extraction device, objective sound extraction program, objective sound extraction method |
-
2009
- 2009-10-15 EP EP09173163A patent/EP2312579A1/en not_active Ceased
-
2010
- 2010-08-18 JP JP2010182876A patent/JP5377442B2/en not_active Expired - Fee Related
Patent Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US4932063A (en) * | 1987-11-01 | 1990-06-05 | Ricoh Company, Ltd. | Noise suppression apparatus |
| WO1998001956A2 (en) * | 1996-07-08 | 1998-01-15 | Chiefs Voice Incorporated | Microphone noise rejection system |
| GB2377805A (en) * | 2001-07-10 | 2003-01-22 | 20 20 Speech Ltd | Localisation of a person in a conveyance |
| EP1879180A1 (en) | 2006-07-10 | 2008-01-16 | Harman Becker Automotive Systems GmbH | Reduction of background noise in hands-free systems |
Non-Patent Citations (7)
| Title |
|---|
| BREGMAN, A: "Auditory Scene Analysis", 1990, MIT PRESS |
| BROWN, G. J.; COOKE, M. P., COMPUTATIONAL AUDITORY SCENE ANALYSIS COMPUTER SPEECH AND LANGUAGE, no. 1, 1994, pages 297 - 336 |
| HECKMANN, M.; JOUBLIN, F.; KORNER, E.: "Sound Source Separation for a Robot Based on Pitch Proc IEEE/RSJ Int.l Conf. on Robots and Intell", SYST, 2005, pages 203 - 208 |
| HU, G.; WANG, D. L.: "Monaural Speech Segregation Based on Pitch Tracking and Amplitude Modulation IEEE Trans", NEURAL NETWORKS, vol. 15, 2004, pages 1135 - 1150 |
| HU, G.; WANG, D., AUDITORY SEGMENTATION BASED ON ONSET AND OFFSET ANALYSIS IEEE TRANSACTIONS ON AUDIO, SPEECH, AND LANGUAGE PROCESSING, vol. 15, 2007, pages 396 - 405 |
| KUO S M ET AL: "ACTIVE NOISE CONTROL: A TUTORIAL REVIEW", PROCEEDINGS OF THE IEEE, IEEE. NEW YORK, US, vol. 87, no. 6, 1 June 1999 (1999-06-01), pages 943 - 973, XP011044219, ISSN: 0018-9219 * |
| PUDER H ET AL: "Improved noise reduction for hands-free car phones utilizing information on vehicle and engine speeds", SIGNAL PROCESSING : THEORIES AND APPLICATIONS, PROCEEDINGS OFEUSIPCO, XX, XX, vol. 3, 1 January 2000 (2000-01-01), pages 1851 - 1854, XP009030255 * |
Cited By (8)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2013041942A1 (en) * | 2011-09-20 | 2013-03-28 | Toyota Jidosha Kabushiki Kaisha | Sound source detecting system and sound source detecting method |
| CN103930791A (en) * | 2011-09-20 | 2014-07-16 | 丰田自动车株式会社 | Sound source detecting system and sound source detecting method |
| CN103930791B (en) * | 2011-09-20 | 2016-03-02 | 丰田自动车株式会社 | Sound Sources Detection system and sound source detection method |
| US9299334B2 (en) | 2011-09-20 | 2016-03-29 | Toyota Jidosha Kabushiki Kaisha | Sound source detecting system and sound source detecting method |
| EP3116236A1 (en) | 2015-07-06 | 2017-01-11 | Sivantos Pte. Ltd. | Method for processing signals for a hearing aid, hearing aid, hearing aid system and interference transmitter for a hearing aid system |
| WO2020012229A1 (en) * | 2018-07-12 | 2020-01-16 | Bosch Car Multimedia Portugal, S.A. | Selective active noise cancelling system |
| CN112614501A (en) * | 2020-12-08 | 2021-04-06 | 深圳创维-Rgb电子有限公司 | Noise reduction method, noise reduction apparatus, noise canceller, microphone, and readable storage medium |
| CN112614501B (en) * | 2020-12-08 | 2024-07-12 | 深圳创维-Rgb电子有限公司 | Noise reduction method, device, noise canceller, microphone, and readable storage medium |
Also Published As
| Publication number | Publication date |
|---|---|
| JP5377442B2 (en) | 2013-12-25 |
| JP2011085904A (en) | 2011-04-28 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| CN110070868B (en) | Voice interaction method, device, automobile and machine-readable medium for in-vehicle system | |
| US11856379B2 (en) | Method, device and electronic device for controlling audio playback of multiple loudspeakers | |
| JP6464449B2 (en) | Sound source separation apparatus and sound source separation method | |
| US20100185308A1 (en) | Sound Signal Processing Device And Playback Device | |
| EP2207168A3 (en) | Robust two microphone noise suppression system | |
| JPWO2001095314A1 (en) | Robot hearing device and robot hearing system | |
| EP2312579A1 (en) | Speech from noise separation with reference information | |
| CN110931027A (en) | Audio processing method and device, electronic equipment and computer readable storage medium | |
| JP2012189907A (en) | Voice discrimination device, voice discrimination method and voice discrimination program | |
| JP2004198656A (en) | Robot audiovisual system | |
| US20250184665A1 (en) | Ear-worn device and reproduction method | |
| US20120197635A1 (en) | Method for generating an audio signal | |
| CN113707156A (en) | Vehicle-mounted voice recognition method and system | |
| CN118506805A (en) | A method and device for transparently transmitting ambient sound in an intelligent car cabin | |
| CN106328154B (en) | A kind of front audio processing system | |
| WO2017000774A1 (en) | System for robot to eliminate own sound source | |
| KR102208536B1 (en) | Speech recognition device and operating method thereof | |
| CN120091256B (en) | Method and system for ensuring smooth communication in large private car | |
| CN110012391B (en) | Operation consultation system and operating room audio acquisition method | |
| KR20030010432A (en) | Apparatus for speech recognition in noisy environment | |
| CN118942491B (en) | Data processing method, electronic device, storage medium, and computer program product | |
| Nakadai et al. | Humanoid active audition system improved by the cover acoustics | |
| KR101081972B1 (en) | Processing Method For Hybrid Feature Vector and Speaker Recognition Method And Apparatus Using the Same | |
| Marquardt et al. | A natural acoustic front-end for Interactive TV in the EU-Project DICIT | |
| EP4738351A1 (en) | Real-time vocal removal from an audio source |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| 17P | Request for examination filed |
Effective date: 20100818 |
|
| AK | Designated contracting states |
Kind code of ref document: A1 Designated state(s): AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC MK MT NL NO PL PT RO SE SI SK SM TR |
|
| AX | Request for extension of the european patent |
Extension state: AL BA RS |
|
| 17Q | First examination report despatched |
Effective date: 20111012 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE APPLICATION HAS BEEN REFUSED |
|
| 18R | Application refused |
Effective date: 20121116 |