EP2312579A1 - Speech from noise separation with reference information - Google Patents

Speech from noise separation with reference information Download PDF

Info

Publication number
EP2312579A1
EP2312579A1 EP09173163A EP09173163A EP2312579A1 EP 2312579 A1 EP2312579 A1 EP 2312579A1 EP 09173163 A EP09173163 A EP 09173163A EP 09173163 A EP09173163 A EP 09173163A EP 2312579 A1 EP2312579 A1 EP 2312579A1
Authority
EP
European Patent Office
Prior art keywords
signal
mixture
reference signal
cues
information
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
EP09173163A
Other languages
German (de)
French (fr)
Inventor
Martin Heckmann
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Honda Research Institute Europe GmbH
Original Assignee
Honda Research Institute Europe GmbH
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Honda Research Institute Europe GmbH filed Critical Honda Research Institute Europe GmbH
Priority to EP09173163A priority Critical patent/EP2312579A1/en
Priority to JP2010182876A priority patent/JP5377442B2/en
Publication of EP2312579A1 publication Critical patent/EP2312579A1/en
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L21/00Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
    • G10L21/02Speech enhancement, e.g. noise reduction or echo cancellation
    • G10L21/0208Noise filtering
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L21/00Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
    • G10L21/02Speech enhancement, e.g. noise reduction or echo cancellation
    • G10L21/0208Noise filtering
    • G10L2021/02085Periodic noise
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L21/00Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
    • G10L21/02Speech enhancement, e.g. noise reduction or echo cancellation
    • G10L21/0208Noise filtering
    • G10L21/0216Noise filtering characterised by the method used for estimating noise
    • G10L2021/02161Number of inputs available containing the signal or the noise to be suppressed
    • G10L2021/02165Two microphones, one receiving mainly the noise signal and the other one mainly the speech signal

Definitions

  • the invention generally refers to the processing of acoustically sensed signals.
  • the present invention relates to a system and a method for separating a mixture signal containing a mixture of acoustical target information ("speech") and interfering information (“noise").
  • speech acoustical target information
  • noise interfering information
  • the European patent application EP 1 879 180 A1 shows a method to reduce the background noise in speech signals with the help of a reference microphone.
  • the main idea behind this method is to estimate the spectrum of the noise based on a reference microphone which captures only the noise and then subtract these spectral components from the microphone signal which captures the mixture of the speech and noise signal.
  • the main disadvantage, however, of this method is that the acoustic environment where the noise is captured is normally different from that where the mixture of the noise and the speech signal are captured (e.g. the engine compartment and the passenger compartment if one wants to reduce the engine noise in a car).
  • the present invention relates to a technique for reducing noise in a mixture signal containing a mixture of noise and speech by means of additional reference information captures e.g. by a second microphone.
  • additional reference information captures e.g. by a second microphone.
  • the present invention does not try to reduce the noise directly in the signal domain but uses techniques inspired by Computational Auditory Scene Analysis (CASA) to reduce the noise.
  • CASA Computational Auditory Scene Analysis
  • a system for the separation of a mixture signal containing a mixture of target information and interfering information comprises means for receiving the mixture signal, means for receiving a reference signal and a signal processing unit configured to extract cues from the reference signal and to separate the target information from the mixture signal using these cues.
  • a method for separating a mixture signal containing a mixture of target information and interfering information comprises the steps of receiving the mixture signal, receiving a reference signal and extracting cues from the reference signal and separating the target information from the mixture signal using these cues.
  • the means for receiving the mixture signal and the means for receiving the reference signal may comprise a microphone and a recording unit each, wherein the microphone for the mixture signal may be positioned close to the origin of the target information and the microphone for the reference signal may be positioned close the origin of the interfering information, when the means for receiving the reference signal are configured to receive interfering information.
  • the interfering information can be also extracted from the speed of an engine.
  • the means for receiving the reference signal may be configured to receive target information, wherein this information can be extracted from a video signal, in particular from the movement of a speaker's body or the speaker's lip movement in the video signal.
  • the signal processing unit of the system may comprise means for splitting the reference signal and the mixture signal in a multitude of frequency channels, means for extracting grouping cues from the reference signal and evaluating the grouping cues in the mixture signal for each frequency channel at each instant in time and means for allocating each frequency channel of the mixture signal at each instant in time to either the target information or the interfering information and separating the mixture signal into the target information and the interfering information.
  • the signal processing unit of the system comprises means for splitting the mixture signal in a multitude of frequency channels, means for extracting grouping cues from the reference signal and evaluating the grouping cues in the mixture signal at each instant in time and means for allocating each frequency channel of the mixture signal at each instant in time to either the target information or the interfering information and separating the mixture signal into the target information and the interfering information.
  • the grouping cues may be the fundamental frequency or on- or off-sets.
  • the target information may be speech and the interfering information may be noise.
  • the system for separating a mixture signal may be included in a motorcycle helmet, wherein the means for receiving the mixture signal are positioned inside the helmet and the means for receiving the reference signal are positioned partly inside the helmet and partly close to the engine of a motorcycle, wherein the means for receiving the reference signal are connected via a cable or wireless.
  • Fig. 1 shows a motorcyclist 7 driving a motorcycle 6 and wearing a helmet 4.
  • the helmet 4 includes a system according to the invention.
  • the system comprises a signal processing unit 1, means for receiving a mixture signal, here a microphone 2, and means for receiving a reference signal, here a microphone 3a, and a receiving unit 3b.
  • the microphone 2 for receiving the mixture signal and the receiving unit 3b are connected to the signal processing unit 1 via a cable.
  • the microphone 3a for receiving the reference signal is, however, not positioned in the helmet 4, but close to the engine 5 of the motorcycle 6 to be at the origin of the interfering signal, which in the shown example may be the harmonic noise generated by the engine 5.
  • the transmission of the reference signal of the microphone 3a to the receiving unit 3b can be for example accomplished via a wireless transmission.
  • the microphone 2 for receiving the mixture signal is positioned to the front of the helmet 4 close to the mouth of the motorcyclist 7.
  • the microphone 2 is therefore positioned close to the origin of the target signal, here the acoustically sensed speech signal of the motorcyclist 7.
  • the microphone 2 also receives noise of the engine 5, due to the fact that the engine 5 of the motorcycle 6 is quite loud and the engine noise is only slightly attenuated by the helmet 4. Therefore the mixture signal received by the microphone 2 contains a mixture of speech and noise.
  • the signal processing unit 1 is configured to extract cues from the reference signal received by the microphone 3 and to separate the speech from the mixture signal received by the microphone 2 using the cues. A detailed description of the extraction and separation will be given in combination with the method and Fig. 3 .
  • the system according to the invention is therefore able to significantly reduce the engine noise in the mixture signal. As a result of the reduction telecommunication while riding will be improved.
  • FIG. 2 Another application area for the system according to the invention is shown in Fig. 2 , where a car 8 is shown including a similar system to that in Fig. 1 .
  • the signal processing unit 1 and the microphone 2 for receiving the mixture signal are not positioned in a helmet, but inside the car 8.
  • the microphone 2 for receiving the mixture signal is positioned in the passenger compartment to be near to the mouth of the driver 7.
  • the microphone 3a for receiving the reference signal is again positioned close to the engine 5.
  • the signal processing unit 1 can be positioned anywhere in the car and has connections to the microphone 2 for receiving the mixture signal and the microphone 3a for receiving the reference signal. Therefore a receiving unit 3b is not needed here.
  • the system according to the invention does not only improve the headset free telecommunication but also speech based operation of devices in a car. In particular the reduction of the harmonic noise generated by the engine is here helpful.
  • Fig. 3 shows a method according to an embodiment of the invention. At the beginning an acoustically sensed mixture signal and a reference signal are received (100, 101).
  • the reference signal is preferably directly sensed from the origin of the noise (e.g. an engine, fan, ). In an ideal setup only the reference signal without any additional signals would be sensed such that the reference signal is available without distortions. This can best be achieved by sensing the reference signal close to its source.
  • the target signal will also preferably be sensed close to its source.
  • the target signal In the case of a speech signal of the driver of a car or a motorcycle sensing close to the drivers mouth would be best.
  • the target signal is commonly sensed at a certain distance from its source.
  • this mixture signal would be a mixture of noise generated by the engine and a speech signal of the driver.
  • the mixture signal and reference signal can for example be sensed by microphones.
  • both signals are split into a multitude of adjacent frequency channels 102, 103.
  • auditory scene analysis cues are extracted from the reference signal ("noise signal") 105.
  • These cues which are typically used in Computational Auditory Scene Analysis (CASA) systems for the separation of sources can be e.g. one or more of:
  • These auditory cues provide information on the reference signal. Knowing these cues allows identifying the reference signal in the mixture signal. For doing so, these cues are extracted in the reference signal, where the reference signal is mostly undistorted and these cues can easily be extracted, and then evaluated in the mixture signal. As a result of this evaluation parts, i.e. frequency channels at each instant in time, can be identified in which the reference signal is dominating the mixture signal.
  • the reference signal is transformed into the frequency domain and they are determined for each frequency channel at each instance in time.
  • these cues as e.g. the fundamental frequency it is also possible to extract these cues directly in the time domain and then calculate their effect on the different frequency channels (in the case of the fundamental frequency signal parts will be present at the fundamental frequency and at its harmonics which can easily be calculated from the fundamental frequency).
  • the auditory cues After extracting the auditory cues from the reference signal they are evaluated in the mixture signal (comprising e.g. noise and speech).
  • the mixture signal comprising e.g. noise and speech
  • the mixture signal is separated, in discrete time steps, into frequency channels "speech" and frequency channels "noise” (107).
  • Figs. 1 and 2 the system according to the invention is included in a motorcycle helmet and in a car.
  • a system is for example included in a robot. This would help to improve speech recognition systems in robots. Therefore robots or any other technical systems which are controlled by speech or interpret speech can be even used in loud and noisy environments.
  • Another application area where the system according to the invention can be used is the field of hearing devices.
  • the elimination of a noise in a mixture signal that the hearing device is receiving helps the person who uses the hearing device to even better understand the speech of other persons.
  • Figs. 1 and 2 are showing a system where the reference signal uses noise from an engine.
  • the reference signal information on the speech signal can be obtained e. g. by using a bone conductive microphone.
  • the necessary grouping information is extracted from the speech signal and then used to separate the speech signal from the noise signal.

Landscapes

  • Engineering & Computer Science (AREA)
  • Computational Linguistics (AREA)
  • Quality & Reliability (AREA)
  • Signal Processing (AREA)
  • Health & Medical Sciences (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Human Computer Interaction (AREA)
  • Physics & Mathematics (AREA)
  • Acoustics & Sound (AREA)
  • Multimedia (AREA)
  • Soundproofing, Sound Blocking, And Sound Damping (AREA)
  • Fittings On The Vehicle Exterior For Carrying Loads, And Devices For Holding Or Mounting Articles (AREA)

Abstract

System and method for separating a mixture signal containing a mixture of target information and interfering information comprising means (2) for receiving the mixture signal, means (3) for receiving a reference signal and a signal processing unit (1) configured to extract cues from the reference signal and to separate the target information from the mixture signal using the cues.

Description

  • The invention generally refers to the processing of acoustically sensed signals.
  • The present invention relates to a system and a method for separating a mixture signal containing a mixture of acoustical target information ("speech") and interfering information ("noise").
  • In many everyday situations different background noise sources are present while we talk on the phone or try to operate a device via speech. This noise, however, makes speech recognition more difficult for humans and especially for machines. For instance, headset free telecommunication in a car or the operation of different devices in the car (e.g. navigation systems, radio) via speech are often interfered by the driving noise (e.g. engine noise) of the car. A similar situation is found when riding a motorcycle, because the noise generated by the motorcycle engine significantly impairs the quality of telecommunication during a ride. Another area where speech recognition is more and more used is robotics. Background noise (e.g. caused by a fan to cool the robot's CPU) is here also reducing the quality of speech recognition.
  • Therefore, different approaches have already been proposed to reduce the noise and thereby improve the speech signal. The European patent application EP 1 879 180 A1 for example shows a method to reduce the background noise in speech signals with the help of a reference microphone. The main idea behind this method is to estimate the spectrum of the noise based on a reference microphone which captures only the noise and then subtract these spectral components from the microphone signal which captures the mixture of the speech and noise signal. The main disadvantage, however, of this method is that the acoustic environment where the noise is captured is normally different from that where the mixture of the noise and the speech signal are captured (e.g. the engine compartment and the passenger compartment if one wants to reduce the engine noise in a car). As a consequence not the noise at it was present in the passenger compartment is subtracted from the corresponding mixture signal but the noise as it was recorded in the engine compartment. To remedy this, a filter is used which replicates the transformation of the signal it underwent on its way from the engine compartment to the passenger compartment.
  • However, this filtering operation is highly dependent on the position of the speaker and has hence to be adaptive. This, on the other hand, is problematic as no error signal is available to adapt the filter when the speaker is talking.
  • There also exist approaches to enhance speech signals in a way inspired by the processing in the human brain. They are commonly referred to as Computational Auditory Scene Analysis (CASA). For humans it was observed that they are able to separate different concurrent sound sources and focus on one source. The underlying mechanisms are able to separate signals based on cues which can serve to bind or group different time frequency regions to one acoustic source. Such cues are e.g. fundamental frequency, location in space, or common on- and off-sets.
  • Different systems have already been developed which are able to separate a speech signal from an interfering signal based on such cues, e.g. fundamental frequency or common on- and off-sets. For doing so these cues are estimated from the signal containing the mixture of the different sources present. This step usually comprises the split of this mixture signal into different frequency channels. Based on the mentioned cues it is determined next for each of the frequency channels at each instant in time to which of the detected sources the signal belongs. The results obtained are usually quite good but also entail significant degradations of the speech signal. One reason for this is that the cues used for separating the sound sources (e.g. fundamental frequency, on- and off-sets) have to be extracted from the mixture of the sources and hence this extraction process is error prone. Similar the number of present sources also has to be estimated from the mixture signal.
  • It is therefore the object of the present invention to propose a system and a method to improve noise reduction in a signal that contains a mixture of a noise and a speech signal.
  • This object is achieved by means of the features of the independent claims. The dependent claims develop further the central idea of the present invention.
  • The present invention relates to a technique for reducing noise in a mixture signal containing a mixture of noise and speech by means of additional reference information captures e.g. by a second microphone. In contrast to previous methods the present invention, however, does not try to reduce the noise directly in the signal domain but uses techniques inspired by Computational Auditory Scene Analysis (CASA) to reduce the noise.
  • Therefore, a system for the separation of a mixture signal containing a mixture of target information and interfering information is proposed, that comprises means for receiving the mixture signal, means for receiving a reference signal and a signal processing unit configured to extract cues from the reference signal and to separate the target information from the mixture signal using these cues.
  • In addition, a method for separating a mixture signal containing a mixture of target information and interfering information is proposed, said comprises the steps of receiving the mixture signal, receiving a reference signal and extracting cues from the reference signal and separating the target information from the mixture signal using these cues.
  • The means for receiving the mixture signal and the means for receiving the reference signal may comprise a microphone and a recording unit each, wherein the microphone for the mixture signal may be positioned close to the origin of the target information and the microphone for the reference signal may be positioned close the origin of the interfering information, when the means for receiving the reference signal are configured to receive interfering information. The interfering information can be also extracted from the speed of an engine.
  • Furthermore the means for receiving the reference signal may be configured to receive target information, wherein this information can be extracted from a video signal, in particular from the movement of a speaker's body or the speaker's lip movement in the video signal.
  • The signal processing unit of the system may comprise means for splitting the reference signal and the mixture signal in a multitude of frequency channels, means for extracting grouping cues from the reference signal and evaluating the grouping cues in the mixture signal for each frequency channel at each instant in time and means for allocating each frequency channel of the mixture signal at each instant in time to either the target information or the interfering information and separating the mixture signal into the target information and the interfering information.
  • In another embodiment the signal processing unit of the system comprises means for splitting the mixture signal in a multitude of frequency channels, means for extracting grouping cues from the reference signal and evaluating the grouping cues in the mixture signal at each instant in time and means for allocating each frequency channel of the mixture signal at each instant in time to either the target information or the interfering information and separating the mixture signal into the target information and the interfering information.
  • The grouping cues may be the fundamental frequency or on- or off-sets. The target information may be speech and the interfering information may be noise.
  • The system for separating a mixture signal may be included in a motorcycle helmet, wherein the means for receiving the mixture signal are positioned inside the helmet and the means for receiving the reference signal are positioned partly inside the helmet and partly close to the engine of a motorcycle, wherein the means for receiving the reference signal are connected via a cable or wireless.
  • These and other aspects and advantages of the present invention will become more apparent when studying the following detailed description, in connection with the figures, in which
  • Fig. 1
    shows a motorcyclist with a helmet that includes a system according to the invention driving a motorcycle;
    Fig. 2
    shows a car with a driver including a system according to the invention;
    Fig. 3
    shows a method according to the invention.
  • Fig. 1 shows a motorcyclist 7 driving a motorcycle 6 and wearing a helmet 4. The helmet 4 includes a system according to the invention. The system comprises a signal processing unit 1, means for receiving a mixture signal, here a microphone 2, and means for receiving a reference signal, here a microphone 3a, and a receiving unit 3b. The microphone 2 for receiving the mixture signal and the receiving unit 3b are connected to the signal processing unit 1 via a cable. The microphone 3a for receiving the reference signal is, however, not positioned in the helmet 4, but close to the engine 5 of the motorcycle 6 to be at the origin of the interfering signal, which in the shown example may be the harmonic noise generated by the engine 5. The transmission of the reference signal of the microphone 3a to the receiving unit 3b can be for example accomplished via a wireless transmission.
  • The microphone 2 for receiving the mixture signal is positioned to the front of the helmet 4 close to the mouth of the motorcyclist 7. The microphone 2 is therefore positioned close to the origin of the target signal, here the acoustically sensed speech signal of the motorcyclist 7. However, the microphone 2 also receives noise of the engine 5, due to the fact that the engine 5 of the motorcycle 6 is quite loud and the engine noise is only slightly attenuated by the helmet 4. Therefore the mixture signal received by the microphone 2 contains a mixture of speech and noise.
  • The signal processing unit 1 is configured to extract cues from the reference signal received by the microphone 3 and to separate the speech from the mixture signal received by the microphone 2 using the cues. A detailed description of the extraction and separation will be given in combination with the method and Fig. 3.
  • The system according to the invention is therefore able to significantly reduce the engine noise in the mixture signal. As a result of the reduction telecommunication while riding will be improved.
  • Another application area for the system according to the invention is shown in Fig. 2, where a car 8 is shown including a similar system to that in Fig. 1. However, the signal processing unit 1 and the microphone 2 for receiving the mixture signal are not positioned in a helmet, but inside the car 8. In particular the microphone 2 for receiving the mixture signal is positioned in the passenger compartment to be near to the mouth of the driver 7. The microphone 3a for receiving the reference signal is again positioned close to the engine 5. The signal processing unit 1 can be positioned anywhere in the car and has connections to the microphone 2 for receiving the mixture signal and the microphone 3a for receiving the reference signal. Therefore a receiving unit 3b is not needed here. The system according to the invention does not only improve the headset free telecommunication but also speech based operation of devices in a car. In particular the reduction of the harmonic noise generated by the engine is here helpful.
  • Fig. 3 shows a method according to an embodiment of the invention. At the beginning an acoustically sensed mixture signal and a reference signal are received (100, 101).
  • The reference signal is preferably directly sensed from the origin of the noise (e.g. an engine, fan, ...). In an ideal setup only the reference signal without any additional signals would be sensed such that the reference signal is available without distortions. This can best be achieved by sensing the reference signal close to its source.
  • In a same way the target signal will also preferably be sensed close to its source. In the case of a speech signal of the driver of a car or a motorcycle sensing close to the drivers mouth would be best. However, to allow a speech interaction where the driver does not need to wear any special device, i.e. headset free, the target signal is commonly sensed at a certain distance from its source. As a consequence only a mixture of the target signal and other sound sources is sensed. Hence, in one application of the present invention this mixture signal would be a mixture of noise generated by the engine and a speech signal of the driver. As already described above, the mixture signal and reference signal can for example be sensed by microphones.
  • After receiving the reference signal and the mixture signal both signals are split into a multitude of adjacent frequency channels 102, 103. For each frequency channel at each instant in time grouping auditory scene analysis cues are extracted from the reference signal ("noise signal") 105. These cues which are typically used in Computational Auditory Scene Analysis (CASA) systems for the separation of sources can be e.g. one or more of:
    • fundamental frequency
    • common on- and off-sets
    • common modulation / fate
    • spatial cues (ITD/ILD i.e. perceived origin)
    • continuity
    • sequential similarity.
  • These auditory cues provide information on the reference signal. Knowing these cues allows identifying the reference signal in the mixture signal. For doing so, these cues are extracted in the reference signal, where the reference signal is mostly undistorted and these cues can easily be extracted, and then evaluated in the mixture signal. As a result of this evaluation parts, i.e. frequency channels at each instant in time, can be identified in which the reference signal is dominating the mixture signal.
  • For the extraction of these auditory cues the reference signal is transformed into the frequency domain and they are determined for each frequency channel at each instance in time. For some of these cues as e.g. the fundamental frequency it is also possible to extract these cues directly in the time domain and then calculate their effect on the different frequency channels (in the case of the fundamental frequency signal parts will be present at the fundamental frequency and at its harmonics which can easily be calculated from the fundamental frequency). After extracting the auditory cues from the reference signal they are evaluated in the mixture signal (comprising e.g. noise and speech). After transforming the mixture signal into the frequency domain and the cues are evaluated for each frequency channel at each instant in time (104). Then it is possible to allocate each frequency channel at each instant in time to either the speech or the noise (106). Based on this allocation the mixture signal is separated, in discrete time steps, into frequency channels "speech" and frequency channels "noise" (107).
  • With this method it is now possible to measure the Computational Auditory Scene Analysis (CASA) cues from an undistorted noise signal and can then use this information to eliminate the noise in the mixture of speech and noise without the need to estimate the transfer function between the site of the recording of the noise and the mixture signal (e. g. from the engine compartment to the passenger compartment).
  • In Figs. 1 and 2 the system according to the invention is included in a motorcycle helmet and in a car. However, it is also possible that such a system is for example included in a robot. This would help to improve speech recognition systems in robots. Therefore robots or any other technical systems which are controlled by speech or interpret speech can be even used in loud and noisy environments.
  • Another application area where the system according to the invention can be used is the field of hearing devices. The elimination of a noise in a mixture signal that the hearing device is receiving helps the person who uses the hearing device to even better understand the speech of other persons.
  • The examples in Figs. 1 and 2 are showing a system where the reference signal uses noise from an engine. However it is also possible that for the reference signal information on the speech signal can be obtained e. g. by using a bone conductive microphone. In this case the necessary grouping information is extracted from the speech signal and then used to separate the speech signal from the noise signal.
  • Prior Art References
  • [1]
    H. Puder, F. Steffens. Improved Noise Reduction for HandsFree Car Phones Utilizing Information on Vehicle and Engine Speeds. EUSIPCO, 2000
    [2]
    EP1879180 - Reduction of background noise in hands- free Systems
    [3]
    Bregman, A. Auditory Scene Analysis MIT Press, 1990
    [4]
    Brown, G. J. & Cooke, M. P. Computational Auditory Scene Analysis Computer Speech and Language, 1994, 1, 297-336
    [5]
    Heckmann, M.; Joublin, F. & Körner, E. Sound Source Separation for a Robot Based on Pitch Proc IEEE/RSJ Int . 1 Conf. on Robots and Intell. Syst., 2005, 203- 208
    [6]
    Hu, G. & Wang, D. L. Monaural Speech Segregation Based on Pitch Tracking and Amplitude Modulation IEEE Trans. Neural Networks, 2004, 15, 1135-1150
    [7]
    Hu, G. & Wang, D. Auditory segmentation based on onset and offset analysis IEEE Transactions on Audio, Speech, and Language Processing, 2007, 15, 396- 405

Claims (21)

  1. System for the separation of an acoustical target signal from an acoustically sensed mixture signal containing a mixture of said target signal and interfering signal, such as noise, the system comprising
    • means (2) for acoustically sensing the mixture signal,
    • means (3) for sensing a reference signal, preferably close to a known source of noise, , and
    • a signal processing unit (1) configured to
    - extract cues from the reference signal for each instance in time at for each of a plurality of preferably adjacent and continuous frequency channels, and
    - to separate the target signal from the mixture signal for each of said frequency channels and each instance in time, using the cues.
  2. System according to claim 1,
    characterized in that,
    the means for receiving the mixture signal comprising a microphone (2), wherein the microphone (2) is positioned close to the origin of the target information.
  3. System according to any of claims 1 to 2,
    characterized in that,
    the means (3) for receiving the reference signal are configured to receive interfering information.
  4. System according to claim 3,
    characterized in that,
    the means (3) for receiving the reference signal comprising a microphone (3a), wherein the microphone (3a) is positioned close to the origin of the interfering information.
  5. System according to any of claims 1 to 2,
    characterized in that,
    the means (3) for receiving the reference signal are configured to receive target information.
  6. System according to any of claims 1 to 5,
    characterized in that,
    the signal processing unit (1) comprises
    • means for splitting the reference signal and the mixture signal in a multitude of frequency channels,
    • means for extracting grouping cues from the reference signal and evaluating the grouping cues in the mixture signal for each frequency channel at each instant in time and
    • means for allocating each frequency channel of the mixture signal at each instant in time to either the target information or the interfering information and separating the mixture signal into the target information and the interfering information.
  7. System according to any of claims 1 to 5,
    characterized in that,

    the signal processing unit (1) comprises
    • means for splitting the mixture signal in a multitude of frequency channels,
    • means for extracting grouping cues from the reference signal and evaluating the grouping cues in the mixture signal at each instant in time and
    • means for allocating each frequency channel of the mixture signal at each instant in time to either the target information or the interfering information and separating the mixture signal into the target information and the interfering information.
  8. A robot, an air/land/sea vehicle, a voice-controlled system or an artificial hearing aid, comprising a system according to any of the preceding claims.
  9. A method for separating a mixture signal containing a mixture of target information and interfering information, comprising the steps
    • Receiving the mixture signal (100),
    • Receiving a reference signal (101) and
    • Extracting cues from the reference signal and separating the target information from the mixture signal using the cues (102-107).
  10. The method according to claim 9,
    characterized in that,
    the target information is speech and the interfering information is noise.
  11. Method according to any of claims 9 to 10,
    characterized in that,
    the step of extracting cues from the reference signal and separating the target information from the mixture signal using the cues comprises the steps
    • Splitting the mixture signal in a multitude of frequency channels (102),
    • Splitting the reference signal in a multitude of frequency channels (103),
    • Extracting grouping cues from the reference signal for each frequency channel at each instant in time (105),
    • Evaluating grouping cues in the mixture signal for each frequency channel at each instant in time (104),
    • Allocating each frequency channel at each instant in time to either the target information or the interfering information based on the evaluation of the grouping cues (106) and
    • Separating the mixture of the target information and the interfering information based on the previous allocation (107).
  12. Method according to any of claims 9 or 10,
    characterized in that,
    the step of extracting cues from the reference signal and separating the target information from the mixture signal using the cues comprises the steps
    • Splitting the mixture signal in a multitude of frequency channels,
    • Extracting grouping cues from the reference signal at each instant in time,
    • Evaluating grouping cues in the mixture signal for each frequency channel at each instant in time ,
    • Allocating each frequency channel at each instant in time to either the target information or the interfering information based on the evaluation of the grouping cues and
    • Separating the mixture of the target information and the interfering information based on the previous allocation.
  13. The method according to any of claims 11 to 12,
    characterized in that,
    the fundamental frequency is used as grouping cue.
  14. The method according to any of claims 11 to 12,
    characterized in that,
    on- or off-sets are used as grouping cue.
  15. The method according to any of claims 9 to 14,
    characterized in that,
    the reference signal contains interfering information.
  16. The method according to claim 15,
    characterized in that,
    the interfering information of the reference signal is extracted from the speed of an engine (5) .
  17. The method according to any of claims 9 to 14,
    characterized in that,
    the reference signal contains target information.
  18. The method according to claim 17,
    characterized in that,
    the target information of the reference signal is extracted from the movements of a speaker's body or lip movements of a speaker in a video signal.
  19. A computer software program product,
    performing a method according to any of claims 9 to 18 when run on a computing unit.
  20. A motorcycle helmet (4), being provided with a system according to any of the claims 1 to 7.
  21. A motorcycle helmet according to claim 20,
    characterized in that,
    the means (2) for receiving the mixture signal are positioned inside the helmet (4) and the means (3) for receiving the reference signal are positioned partly (3b) inside the helmet (4) and partly (3a) close to a engine (5) of a motorcycle (6), wherein the means (3a, 3b) for receiving the reference signal are connected via a cable or wireless.
EP09173163A 2009-10-15 2009-10-15 Speech from noise separation with reference information Ceased EP2312579A1 (en)

Priority Applications (2)

Application Number Priority Date Filing Date Title
EP09173163A EP2312579A1 (en) 2009-10-15 2009-10-15 Speech from noise separation with reference information
JP2010182876A JP5377442B2 (en) 2009-10-15 2010-08-18 System that separates speech from noise by reference information

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
EP09173163A EP2312579A1 (en) 2009-10-15 2009-10-15 Speech from noise separation with reference information

Publications (1)

Publication Number Publication Date
EP2312579A1 true EP2312579A1 (en) 2011-04-20

Family

ID=41694574

Family Applications (1)

Application Number Title Priority Date Filing Date
EP09173163A Ceased EP2312579A1 (en) 2009-10-15 2009-10-15 Speech from noise separation with reference information

Country Status (2)

Country Link
EP (1) EP2312579A1 (en)
JP (1) JP5377442B2 (en)

Cited By (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2013041942A1 (en) * 2011-09-20 2013-03-28 Toyota Jidosha Kabushiki Kaisha Sound source detecting system and sound source detecting method
EP3116236A1 (en) 2015-07-06 2017-01-11 Sivantos Pte. Ltd. Method for processing signals for a hearing aid, hearing aid, hearing aid system and interference transmitter for a hearing aid system
WO2020012229A1 (en) * 2018-07-12 2020-01-16 Bosch Car Multimedia Portugal, S.A. Selective active noise cancelling system
CN112614501A (en) * 2020-12-08 2021-04-06 深圳创维-Rgb电子有限公司 Noise reduction method, noise reduction apparatus, noise canceller, microphone, and readable storage medium

Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US4932063A (en) * 1987-11-01 1990-06-05 Ricoh Company, Ltd. Noise suppression apparatus
WO1998001956A2 (en) * 1996-07-08 1998-01-15 Chiefs Voice Incorporated Microphone noise rejection system
GB2377805A (en) * 2001-07-10 2003-01-22 20 20 Speech Ltd Localisation of a person in a conveyance
EP1879180A1 (en) 2006-07-10 2008-01-16 Harman Becker Automotive Systems GmbH Reduction of background noise in hands-free systems

Family Cites Families (9)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP3384540B2 (en) * 1997-03-13 2003-03-10 日本電信電話株式会社 Receiving method, apparatus and recording medium
JPH11265199A (en) * 1998-03-18 1999-09-28 Nippon Telegr & Teleph Corp <Ntt> Transmitter
JP4119112B2 (en) * 2001-11-05 2008-07-16 本田技研工業株式会社 Mixed sound separator
EP1605439B1 (en) * 2004-06-04 2007-06-27 Honda Research Institute Europe GmbH Unified treatment of resolved and unresolved harmonics
JP2007079389A (en) * 2005-09-16 2007-03-29 Yamaha Motor Co Ltd Speech analysis method and speech analysis apparatus
CN101512374B (en) * 2006-11-09 2012-04-11 松下电器产业株式会社 Sound source position detection device
JP4336378B2 (en) * 2007-04-26 2009-09-30 株式会社神戸製鋼所 Objective sound extraction device, objective sound extraction program, objective sound extraction method
JP4519901B2 (en) * 2007-04-26 2010-08-04 株式会社神戸製鋼所 Objective sound extraction device, objective sound extraction program, objective sound extraction method
JP4493690B2 (en) * 2007-11-30 2010-06-30 株式会社神戸製鋼所 Objective sound extraction device, objective sound extraction program, objective sound extraction method

Patent Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US4932063A (en) * 1987-11-01 1990-06-05 Ricoh Company, Ltd. Noise suppression apparatus
WO1998001956A2 (en) * 1996-07-08 1998-01-15 Chiefs Voice Incorporated Microphone noise rejection system
GB2377805A (en) * 2001-07-10 2003-01-22 20 20 Speech Ltd Localisation of a person in a conveyance
EP1879180A1 (en) 2006-07-10 2008-01-16 Harman Becker Automotive Systems GmbH Reduction of background noise in hands-free systems

Non-Patent Citations (7)

* Cited by examiner, † Cited by third party
Title
BREGMAN, A: "Auditory Scene Analysis", 1990, MIT PRESS
BROWN, G. J.; COOKE, M. P., COMPUTATIONAL AUDITORY SCENE ANALYSIS COMPUTER SPEECH AND LANGUAGE, no. 1, 1994, pages 297 - 336
HECKMANN, M.; JOUBLIN, F.; KORNER, E.: "Sound Source Separation for a Robot Based on Pitch Proc IEEE/RSJ Int.l Conf. on Robots and Intell", SYST, 2005, pages 203 - 208
HU, G.; WANG, D. L.: "Monaural Speech Segregation Based on Pitch Tracking and Amplitude Modulation IEEE Trans", NEURAL NETWORKS, vol. 15, 2004, pages 1135 - 1150
HU, G.; WANG, D., AUDITORY SEGMENTATION BASED ON ONSET AND OFFSET ANALYSIS IEEE TRANSACTIONS ON AUDIO, SPEECH, AND LANGUAGE PROCESSING, vol. 15, 2007, pages 396 - 405
KUO S M ET AL: "ACTIVE NOISE CONTROL: A TUTORIAL REVIEW", PROCEEDINGS OF THE IEEE, IEEE. NEW YORK, US, vol. 87, no. 6, 1 June 1999 (1999-06-01), pages 943 - 973, XP011044219, ISSN: 0018-9219 *
PUDER H ET AL: "Improved noise reduction for hands-free car phones utilizing information on vehicle and engine speeds", SIGNAL PROCESSING : THEORIES AND APPLICATIONS, PROCEEDINGS OFEUSIPCO, XX, XX, vol. 3, 1 January 2000 (2000-01-01), pages 1851 - 1854, XP009030255 *

Cited By (8)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2013041942A1 (en) * 2011-09-20 2013-03-28 Toyota Jidosha Kabushiki Kaisha Sound source detecting system and sound source detecting method
CN103930791A (en) * 2011-09-20 2014-07-16 丰田自动车株式会社 Sound source detecting system and sound source detecting method
CN103930791B (en) * 2011-09-20 2016-03-02 丰田自动车株式会社 Sound Sources Detection system and sound source detection method
US9299334B2 (en) 2011-09-20 2016-03-29 Toyota Jidosha Kabushiki Kaisha Sound source detecting system and sound source detecting method
EP3116236A1 (en) 2015-07-06 2017-01-11 Sivantos Pte. Ltd. Method for processing signals for a hearing aid, hearing aid, hearing aid system and interference transmitter for a hearing aid system
WO2020012229A1 (en) * 2018-07-12 2020-01-16 Bosch Car Multimedia Portugal, S.A. Selective active noise cancelling system
CN112614501A (en) * 2020-12-08 2021-04-06 深圳创维-Rgb电子有限公司 Noise reduction method, noise reduction apparatus, noise canceller, microphone, and readable storage medium
CN112614501B (en) * 2020-12-08 2024-07-12 深圳创维-Rgb电子有限公司 Noise reduction method, device, noise canceller, microphone, and readable storage medium

Also Published As

Publication number Publication date
JP5377442B2 (en) 2013-12-25
JP2011085904A (en) 2011-04-28

Similar Documents

Publication Publication Date Title
CN110070868B (en) Voice interaction method, device, automobile and machine-readable medium for in-vehicle system
US11856379B2 (en) Method, device and electronic device for controlling audio playback of multiple loudspeakers
JP6464449B2 (en) Sound source separation apparatus and sound source separation method
US20100185308A1 (en) Sound Signal Processing Device And Playback Device
EP2207168A3 (en) Robust two microphone noise suppression system
JPWO2001095314A1 (en) Robot hearing device and robot hearing system
EP2312579A1 (en) Speech from noise separation with reference information
CN110931027A (en) Audio processing method and device, electronic equipment and computer readable storage medium
JP2012189907A (en) Voice discrimination device, voice discrimination method and voice discrimination program
JP2004198656A (en) Robot audiovisual system
US20250184665A1 (en) Ear-worn device and reproduction method
US20120197635A1 (en) Method for generating an audio signal
CN113707156A (en) Vehicle-mounted voice recognition method and system
CN118506805A (en) A method and device for transparently transmitting ambient sound in an intelligent car cabin
CN106328154B (en) A kind of front audio processing system
WO2017000774A1 (en) System for robot to eliminate own sound source
KR102208536B1 (en) Speech recognition device and operating method thereof
CN120091256B (en) Method and system for ensuring smooth communication in large private car
CN110012391B (en) Operation consultation system and operating room audio acquisition method
KR20030010432A (en) Apparatus for speech recognition in noisy environment
CN118942491B (en) Data processing method, electronic device, storage medium, and computer program product
Nakadai et al. Humanoid active audition system improved by the cover acoustics
KR101081972B1 (en) Processing Method For Hybrid Feature Vector and Speaker Recognition Method And Apparatus Using the Same
Marquardt et al. A natural acoustic front-end for Interactive TV in the EU-Project DICIT
EP4738351A1 (en) Real-time vocal removal from an audio source

Legal Events

Date Code Title Description
PUAI Public reference made under article 153(3) epc to a published international application that has entered the european phase

Free format text: ORIGINAL CODE: 0009012

17P Request for examination filed

Effective date: 20100818

AK Designated contracting states

Kind code of ref document: A1

Designated state(s): AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC MK MT NL NO PL PT RO SE SI SK SM TR

AX Request for extension of the european patent

Extension state: AL BA RS

17Q First examination report despatched

Effective date: 20111012

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: THE APPLICATION HAS BEEN REFUSED

18R Application refused

Effective date: 20121116