WO2020044166A1 - Integrated noise reduction - Google Patents

Integrated noise reduction Download PDF

Info

Publication number
WO2020044166A1
WO2020044166A1 PCT/IB2019/057011 IB2019057011W WO2020044166A1 WO 2020044166 A1 WO2020044166 A1 WO 2020044166A1 IB 2019057011 W IB2019057011 W IB 2019057011W WO 2020044166 A1 WO2020044166 A1 WO 2020044166A1
Authority
WO
WIPO (PCT)
Prior art keywords
estimate
sound
priori
signals
target sound
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/IB2019/057011
Other languages
French (fr)
Inventor
Randall ALI
Toon Van Waterschoot
Marc Moonen
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Cochlear Ltd
Original Assignee
Cochlear Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Cochlear Ltd filed Critical Cochlear Ltd
Priority to US17/261,778 priority Critical patent/US11943590B2/en
Publication of WO2020044166A1 publication Critical patent/WO2020044166A1/en
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04RLOUDSPEAKERS, MICROPHONES, GRAMOPHONE PICK-UPS OR LIKE ACOUSTIC ELECTROMECHANICAL TRANSDUCERS; ELECTRIC HEARING AIDS; PUBLIC ADDRESS SYSTEMS
    • H04R3/00Circuits for transducers
    • H04R3/005Circuits for transducers for combining the signals of two or more microphones
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04RLOUDSPEAKERS, MICROPHONES, GRAMOPHONE PICK-UPS OR LIKE ACOUSTIC ELECTROMECHANICAL TRANSDUCERS; ELECTRIC HEARING AIDS; PUBLIC ADDRESS SYSTEMS
    • H04R25/00Electric hearing aids
    • H04R25/40Arrangements for obtaining a desired directivity characteristic
    • H04R25/405Arrangements for obtaining a desired directivity characteristic by combining a plurality of transducers
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04RLOUDSPEAKERS, MICROPHONES, GRAMOPHONE PICK-UPS OR LIKE ACOUSTIC ELECTROMECHANICAL TRANSDUCERS; ELECTRIC HEARING AIDS; PUBLIC ADDRESS SYSTEMS
    • H04R5/00Stereophonic arrangements
    • H04R5/04Circuit arrangements, e.g. for selective connection of amplifier inputs/outputs to loudspeakers, for loudspeaker detection, or for adaptation of settings to personal preferences or hearing impairments
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04SSTEREOPHONIC SYSTEMS 
    • H04S7/00Indicating arrangements; Control arrangements, e.g. balance control
    • H04S7/30Control circuits for electronic adaptation of the sound field
    • H04S7/302Electronic adaptation of stereophonic sound system to listener position or orientation
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04RLOUDSPEAKERS, MICROPHONES, GRAMOPHONE PICK-UPS OR LIKE ACOUSTIC ELECTROMECHANICAL TRANSDUCERS; ELECTRIC HEARING AIDS; PUBLIC ADDRESS SYSTEMS
    • H04R2225/00Details of deaf aids covered by H04R25/00, not provided for in any of its subgroups
    • H04R2225/43Signal processing in hearing aids to enhance the speech intelligibility
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04RLOUDSPEAKERS, MICROPHONES, GRAMOPHONE PICK-UPS OR LIKE ACOUSTIC ELECTROMECHANICAL TRANSDUCERS; ELECTRIC HEARING AIDS; PUBLIC ADDRESS SYSTEMS
    • H04R2225/00Details of deaf aids covered by H04R25/00, not provided for in any of its subgroups
    • H04R2225/67Implantable hearing aids or parts thereof not covered by H04R25/606
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04RLOUDSPEAKERS, MICROPHONES, GRAMOPHONE PICK-UPS OR LIKE ACOUSTIC ELECTROMECHANICAL TRANSDUCERS; ELECTRIC HEARING AIDS; PUBLIC ADDRESS SYSTEMS
    • H04R2430/00Signal processing covered by H04R, not provided for in its groups
    • H04R2430/20Processing of the output signals of the acoustic transducers of an array for obtaining a desired directivity characteristic
    • H04R2430/25Array processing for suppression of unwanted side-lobes in directivity characteristics, e.g. a blocking matrix
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04SSTEREOPHONIC SYSTEMS 
    • H04S2420/00Techniques used stereophonic systems covered by H04S but not provided for in its groups
    • H04S2420/01Enhancing the perception of the sound image or of the spatial distribution using head related transfer functions [HRTF's] or equivalents thereof, e.g. interaural time difference [ITD] or interaural level difference [ILD]

Definitions

  • the present invention generally relates to integrated noise reduction for devices having at least one local microphone array.
  • Hearing loss is a type of sensory impairment that is generally of two types, namely conductive and/or sensorineural.
  • Conductive hearing loss occurs when the normal mechanical pathways of the outer and/or middle ear are impeded, for example, by damage to the ossicular chain or ear canal.
  • Sensorineural hearing loss occurs when there is damage to the inner ear, or to the nerve pathways from the inner ear to the brain.
  • auditory prostheses include, for example, acoustic hearing aids, bone conduction devices, and direct acoustic stimulators.
  • a method comprises: receiving sound signals with at least a local microphone array of a device, wherein the sound signals comprise at least one target sound; generating an a priori estimate of the at least one target sound in the received sound signals based on a predetermined location of a source of the at least one target sound; generating a direct estimate of the at least one target sound in the received sound signals based on a real-time estimate of a location of a source of the at least one target sound; and generating a weighted combination of the a priori estimate and the direct estimate, wherein the weighted combination is an integrated estimate of the target sound.
  • a device comprising: a local microphone array configured to receive sound signals, wherein the sound signals comprise at least one target sound; and one or more processors configured to: generate an a priori estimate of the at least one target sound in the received sound signals using only an a priori relative transfer function (RTF) vector generated from the received sound signals, generate a direct estimate of the at least one target sound in the received sound signals using only an a priori relative transfer function (RTF) vector generated from the received sound signals, and generate a weighted combination of the a priori estimate and the direct estimate, wherein the weighted combination is an integrated estimate of the target sound.
  • RTF priori relative transfer function
  • FIG. 1 is a functional block diagram illustrating the generation of pre-whitened transformed signals
  • FIG. 2 is a functional block diagram illustrating the generation of an a priori estimate of at least one target sound in sound signals received at a local microphone array
  • FIG. 3 is a functional block diagram illustrating the generation of a direct estimate of at least one target sound in sound signals received at a local microphone array
  • FIG. 4 is a functional block diagram illustrating the generation of an integrated estimate of at least one target sound in sound signals received at a local microphone array
  • FIG. 5 is a functional block diagram illustrating the generation of an a priori estimate of at least one target sound in sound signals received at a local microphone array and at least one external microphone
  • FIG. 6 is a functional block diagram illustrating the generation of a direct estimate of at least one target sound in sound signals received at a local microphone array and at least one external microphone
  • FIG. 7 is a functional block diagram illustrating the generation of an integrated estimate of at least one target sound in sound signals received at a local microphone array and at least one external microphone;
  • FIG. 8 is flowchart of a two stage process, in accordance with embodiments presented herein;
  • FIG. 9 is a table summarizing the various noise reduction strategies, in accordance with embodiments presented herein;
  • FIG. 10A is a schematic diagram illustrating a cochlear implant, in accordance with certain embodiments presented herein;
  • FIG. 10B is a block diagram of the cochlear implant of FIG. 10A;
  • FIG. 11 is a block diagram of a totally implantable cochlear implant, in accordance with certain embodiments presented herein;
  • FIG. 12 is a block diagram of a bone conduction device that includes a spatial pre-filter, in accordance with embodiments presented herein.
  • FIG. 13 is a flowchart of a method, in accordance with embodiments presented herein.
  • multi -microphone noise reduction systems are used to preserve desired sounds (e.g., speech), while rejecting unwanted sounds (e.g., noise).
  • desired sounds e.g., speech
  • unwanted sounds e.g., noise
  • a local microphone array (LMA) worn on the recipient i.e., part of the device
  • LMA local microphone array
  • a sound source e.g., speaker
  • the integrated noise reduction techniques presented herein improve upon these existing noise reduction systems in several distinct ways: (i) by including the ability to focus on a target sound source (e.g., speaker) that is not in the predefined direction and, in certain arrangements, (ii) by including external microphones (XMs) that operate together with the LMA, resulting in further noise reduction as opposed to using only the LMA.
  • a target sound source e.g., speaker
  • XMs external microphones
  • integrated noise reduction techniques will utilize two separate tuning parameters, one for controlling the sound received from the predefined direction, and the other for the sound received from an estimated direction where the target sound source may be located.
  • each of these directions can be defined using the LMA and the XMs.
  • a modified version of the improved method of estimation of a transfer function for the XM is used, where the input signals have to undergo a specific series of transformations.
  • Using one or several XMs along with the LMA can provide significant speech intelligibility improvement, for instance in the case where XMs may be quite close to the desired speaker, or even if it provides a relevant noise reference. Additionally, the integrated noise reduction techniques presented herein are flexible in that they encompass a wide range of noise reduction options according to the tuning of the system.
  • section II describes a data model, which considers the general case of a local microphone array (LMA) in conjunction with one or several external microphones (XMs), which can be reduced to a single external microphone without compromising the equations provided herein.
  • LMA local microphone array
  • XMs external microphones
  • a transformed domain, as well as a pre-whitened-transformed domain is also introduced in order to simplify the flow of signal processing operations and realize distinct digital signal processing (DSP) block schemes.
  • DSP digital signal processing
  • section III an integrated minimum variance distortionless response (MVDR) beamformer is discussed as applied to a local microphone array.
  • section III describes an integrated MVDR beamformer, which leverages the use of a priori assumptions and the use of estimated quantities.
  • section IV an integrated MVDR beamformer as applied to a local microphone array together with one or more external microphones is described.
  • an integrated MVDR beamformer for application to a local microphone array together with one or more external microphones which leverages the use of a priori assumptions and the use of estimated quantities is described.
  • y e [y e,i y e,2 - y e, M e ] T are the external microphone signals
  • n [n a n ] T represents the noise component, which consists of a combination of correlated and uncorrelated noises.
  • Variables with the subscript“a” refer to the LMA signals and variables with the subscript“e” refer to the XM signals.
  • the dependencies on k and l will be introduced herein, as needed, for mathematical derivations.
  • the speech component (target sound), x can be represented in terms of a relative transfer function (RTF) vector such that:
  • s x a a l s
  • h the RTF vector defined as:
  • the noise reduction system will aim to produce an estimate for the speech component in the reference microphone
  • the speech-plus-noise and noise-only correlation matrices are estimated from the received microphone signals during speech-plus-noise and noise-only periods, using a voice activity detector (VAD).
  • VAD voice activity detector
  • B a B a l (Ma -i ) and in general denotes the ⁇ x ⁇ identity matric, and b a can be interpreted as a scaled matched filter. W.l.o.g, b a will simply be referred to as a matched filter in the following derivations.
  • an (M a + M e ) X (M a + M e ) unitary transformation matrix, T can be subsequently defined:
  • the transformed noise signals can also be similarly defined:
  • this transformation domain is the LMA signals that pass through a blocking matrix and a matched filter, as in the first stage of a generalized sidelobe canceller (GSC) (i.e., the adaptive implementation of an MVDR beamformer), along with the XM signals.
  • GSC generalized sidelobe canceller
  • a spatial pre-whitening operation can be defined from the noise-only correlation matrix in the previously described transform domain by using the Cholesky decomposition:
  • L is an (M a + M e ) X (M a + M e ) lower triangular matrix.
  • L can be realized as:
  • L a and L x are lower triangular matrices. It should be noted that L a corresponds to the LMA signals and are from a Cholesky decomposition of the noise correlation matrix from the LMA signals in the transformed domain, hence:
  • a signal vector in the transformed domain can be consequently pre-whitened by pre-multiplying it with LT 1 .
  • Such signal quantities will be denoted with the underbar Q notation.
  • the signal y in this so-called pre-whitened-transformed domain is given by:
  • FIG. 1 is a block diagram illustrating the flow of the previously described transformations on the unprocessed signals.
  • Transformation block 102 is a processing block that represents the first transformation of section II-B, in which the LMA signals pass through a blocking matrix 104 and a matched filter 106, analogous to the first stage of a GSC.
  • the XM signals are unaltered.
  • the pre-whitening block 108 is a processing block that represents the pre-whitening operation of section II-C, yielding signals 109 in the pre-whitened-transformed domain.
  • the noise reduction filters that will be developed below will then be directly applied to these pre-whitened-transformed signals (i.e., the output of pre-whitening block 108) in order to yield the desired speech estimate.
  • the MVDR beamformer minimizes the total noise power (minimum variance), while preserving the received signal in a particular direction (distortionless response). This direction is specified by defining the appropriate RTF vector for the MVDR beamformer.
  • the MVDR problem can be formulated as follows (which will be referred to as the MVDRa): where h a is the RTF vector from (4), which in practice is unknown and hence will be replaced either by a priori assumptions or estimated from the speech-plus-noise correlation matrices.
  • the optimal noise reduction filter is then given by:
  • the speech estimate, z a l , from this MVDR a beamformer is obtained through the linear filtering of the microphone signals with the complex -valued filter w a :
  • Section III-A and III-B strategies for designing an MVDR a beamformer using an RTF vector based either on a priori assumptions or estimated from the speech-plus-noise correlation matrices are discussed.
  • Section III-C illustrates an integrated beamformer that integrates the use of priori assumptions with estimates.
  • This h a can be based on a priori assumptions regarding microphone characteristics, position, speaker location and room acoustics (e.g., no reverberation).
  • the optimal noise reduction filter is then given by: n h aha h a
  • FIG. 2 illustrates transformation block 102 and pre-whitening block 108, as described above with reference to FIG. 1.
  • a priori filter 110 which produces pp and processing block 112 which applies pp to y a Ma .
  • the application of pp to y a Ma produces an a priori speech estimate z a 1 .
  • the a priori speech estimate, z a l is an estimate of the target sound (e.g., speech) in the received sound signals, based solely on an a priori RTF vector.
  • the RTF vector is generated uses assumptions regarding, for example, location of the source of the target sound, characteristics of the microphones (e.g., microphone calibration in regards to gains, phases, etc.), reverberant characteristics of the target sound source, etc.
  • the a priori speech estimate z a 1 is an example of an a priori estimate of at least one target sound in the received sound signals.
  • the RTF vector may also be estimated without reliance on any a priori assumptions and can be used to enhance the speech regardless of the speech source location.
  • One such method is a method of covariance whitening or equivalently that which involves a Generalized Eigenvalue Decomposition (GEVD).
  • GSVD Generalized Eigenvalue Decomposition
  • a rank- 1 matrix approximation problem can be formulated to estimate the RTF vector for a given set of LMA signals such that: iP in ll ( R yaya - R n a n a ) - Rxa,rl
  • this filter based on estimated quantities can also be reformulated in the transformed, pre-whitened-transformed domain. Leaving the derivations once again to Appendix B, the corresponding speech estimate using the estimated RTF vector is:
  • PpP max can be considered as the pre-whitened-transformed filter (where ⁇ . ⁇ * is the complex conjugate), which can be used to directly filter the pre-whitened, transformed signals, y a .
  • These operations can also be realized in a distinct set of signal processing blocks, as illustrated in FIG. 3.
  • FIG. 3 illustrates transformation block 102 and pre-whitening block 108, as described above with reference to FIG. 1, which produce pre-whitened-transformed signals. Also shown is block 114, which filters the pre-whitened-transformed signals in accordance with PpP max (he., 114 represents the hermitian transposed pre-whitened- transformed filter). The output of the pre-whitened-transformed filter 114 is a direct speech estimate, z a 1 (i.e., (32), above).
  • the direct speech estimate, z a 1 is an estimate of the target sound (e.g., speech) in the received sound signals, based solely on an estimated RTF vector.
  • the estimated RTF vector is generated using real-time estimates of, for example, the location of the source of the target sound, reverberant characteristics of the target sound source, etc.
  • the direct speech estimate, z a l is an example of a direct estimate of at least one target sound in the received sound signals.
  • the integrated MVDR a beamformer provides for integrated tunings which allow different“weights” to be applied to each of (1) an a priori assumed representation of target sound within received sound signals (e.g., an a priori estimate of at least one target sound in the received sound signals), and (2) an estimated representation of the target sound within received sound signals (e.g., a direct estimate of at least one target sound in the received sound signal).
  • the weights applied to each of the a priori assumed representation of the target sound and the estimated representation of the target sound are selected based on“confidence measures” associated with each of the a priori assumed representation of the target sound and the estimated representation of the target sound, respectively.
  • the integrated MVDR a beamformer if the speech source moves outside of the direction defined by an a priori assumed RTF vector, more weight can be given to an estimated RTF vector to account for the loss in performance that would otherwise result from using the a priori assumed RTF vector alone.
  • the estimated RTF vector becomes unreliable, less weight can be given thereto and the system can revert to using the a priori assumed RTF vector, which may have an improved performance if the speech source is indeed in the direction defined by the a priori assumed RTF vector.
  • Combination/mixing of the a priori assumed RTF vector and the estimated RTF vector is also possible. That is, the tuning parameters can achieve multiple beamformers, i.e.
  • One particular tuning of interest may be to place a large weight on an a priori assumed RTF vector, but weighting an estimated RTF vector only when appropriate. This represents a mechanism for reverting to an a priori assumed RTF vector when the estimated RTF vector was unreliable.
  • an integrated MVDRa cost function can be given as: where a e [0, ⁇ ] and b E [0, ⁇ ] are tuning parameters that control how much of the respective RTF vectors (i.e., the a priori assumed RTF vector and the estimated RTF vector) are weighted.
  • This cost function is the combination of that of an MVDRa (as in (22)) defined by h a and another defined by h a , except that the constraints have been softened by a and b.
  • This integrated MVDR beamformer reveals that the MVDRa beamformer based on a priori assumptions from (25) and that which is based on estimated quantities from (31) can be combined according to the functions f pr (a, b) and f est ( a > b) respectively.
  • this integrated beamformer can also be expressed in the pre-whitened-transformed domain as follows:
  • FIG. 4 is a block diagram of an integrated MVDR a beamformer 125 in accordance with embodiments presented herein.
  • the integrated MVDR a beamformer 125 comprises a plurality of processing blocks, which include transformation block 102 and pre- whitening block 108. As described above with reference to FIG. 1 transformation block 102 and pre-whitening block 108 produce signals 109 in the pre-whitened-transformed domain (pre-whitened-transformed signal s) .
  • FIG. 4 Also shown in FIG. 4 are two processing branches 113(1) and 113(2) that each operate based on all or part of the pre-whitened-transformed signals 109.
  • 113(1) includes an a priori filter 110, which produces and a processing block 112 which
  • a priori RTF vector i.e., an estimate of the speech in the received sound signals, based solely on a priori assumptions such as microphone characteristics, source location, and reverberant characteristics of the target sound (e.g., speech) source.
  • application of to y a Ma generates an a priori estimate of at least one target sound in the received sound signals.
  • the first branch 113(1) also comprises a first weighting block 116.
  • the first weighting block 116 is configured to weight the speech estimate, z a l , in accordance with the complex conjugate of the function f pr (a, b) (i.e., (35) and (40), above). More generally, the first weighting block 116 is configured to weight the speech estimate, z a 1 , in accordance with a cost function controlled by a plurality of tuning parameters (e.g., (a, /?)). The tuning parameters of the cost function (e.g., / pr (a, ?)), are set based on one or more confidence measures 118 generated for the speech estimate, z a l .
  • the tuning parameters of the cost function e.g., / pr (a, ?)
  • the one or more confidence measures 118 represent an assessment or estimate of the accuracy/reliability of the a priori speech estimate, z a 1 , and the hence the accuracy of the a priori RTF vector used to generate the speech estimate, z a 1 .
  • the first weighting block 116 generates a weighted a priori speech estimate, shown in FIG. 5 by arrow 119.
  • the second branch 113(2) includes a pre-whitened-transformed filter 114, which filters the pre-whitened-transformed signals in accordance with (32).
  • the output of the pre-whitened- transformed filter 114 is a direct speech estimate, z a l , that is generated based solely on an estimated RTF vector (i.e., an estimate of the speech in the received sound signals, which takes into consideration microphone characteristics and may contain information such as the location and some reverberant characteristics of the speech source).
  • the direct speech estimate z a 1 is an example of a direct estimate of at least one target sound in the received sound signals.
  • the second branch 113(2) also comprises a second weighting block 120.
  • the second weighting block 120 is configured to weight the speech estimate, z a 1 , in accordance with complex conjugate of the function f est ( ⁇ x, ) (i.e., (36) and (40), above). More generally, the second weighting block 120 is configured to weight the direct speech estimate, z a 1 , in accordance with a cost function controlled by a plurality of tuning parameters (e.g., (a, b)).
  • the tuning parameters of the cost function e.g., f est a > b
  • the tuning parameters of the cost function are set based on one or more confidence measures 122 generated for the speech estimate, z a l .
  • the one or more confidence measures 122 represent an assessment or estimate of the accuracy/reliability of the speech estimate, z a 1 , and the hence the accuracy of the estimated RTF vector used to generate the speech estimate, z a l .
  • the second weighting block 120 generates a weighted direct speech estimate, shown in FIG. 5 by arrow 123.
  • FIG. 4 also illustrates processing block 124 which integrates/combines the weighted a priori speech estimate 119 and the weighted direct speech estimate 123.
  • the combination of the weighted a priori speech estimate 119 and the weighted direct speech estimate 123 is referred to as an integrated speech estimate, z a int (i.e., (40), above).
  • the integrated speech estimate may be used for subsequent processing in the device (e.g., auditory prosthesis).
  • Section III illustrates an embodiment in which the integrated beamformer operates based on local microphone array (LMA) signals.
  • LMA signals are generated by a local microphone array (LMA) that are part of the device that performs the integrated noise reduction techniques.
  • LMA is worn on the recipient.
  • the integrated noise reduction techniques described herein can be extended to include external microphone (XM) signals, in addition to the LMA signals.
  • XM signals are generated by one or more external microphones (XMs) that are not part of the device that performs the integrated noise reduction techniques, but that can nevertheless communicate with the device (e.g., via a wireless connection).
  • the external microphones may be any type of microphone (e.g., microphones in a wireless microphone device, microphones in a separate computing device (e.g., phone laptop, tablet, etc.), microphones in another auditory prosthesis, microphones in a conference phone system, microphones in hands-free system, etc) for which the location of the microphone(s) is unknown relative to the microphones of the LMA.
  • an external microphone may be any microphone that has an unknown location, which may change over time, with respect to the local microphone array.
  • the integrated beamformer is referred to as the MVDR a L
  • h is the RTF vector ((4), above) that includes M a components corresponding to the LMA, h a , and M e components corresponding to the XMs, h e , and R nn is the (M a + M e ) X (M a + M e ) noise correlation matrix:
  • h for the MVDR a,e is such that the a priori RTF vector for the LMA signals, h a , is preserved and only the RTF vector for the XM signals is estimated.
  • RTF will therefore be defined as follows:
  • R x rl is a rank-l approximation to R xx (recall (8)).
  • the a priori assumed RTF vector for the LMA signals can also be included for the definition of R x rl and hence is given by:
  • the estimation problem of (45) can be equivalently formulated in the pre-whitened-transformed domain.
  • Appendix C it is shown that the estimated RTF vector could be found from a GEVD on the matrix pencil
  • this GEVD can consequently be computed from the EVD of J T R yy J, which is a lower order correlation matrix, of dimensions (M e + 1) X (M e + 1) that could be constructed from the last (M e + 1) elements of the pre-whitened-transformed signals, namely that in relation to the last element of the LMA - y a Ma , ar
  • the resulting RTF vector for the XM signals is then defined from the corresponding principal (first in this case) eigenvector, v max : where the selection matrix, J
  • FIG. 5 is a block diagram illustrating a transformation block 502 representing the first transformation of section II-B, in which the LMA signals pass through a blocking matrix 504 and a matched filter 506, analogous to the first stage of a GSC.
  • the XM signals are unaltered.
  • the pre-whitening block 508 represents the pre-whitening operation.
  • the output of the pre-whitening block 508 is signals in the pre-whitened-transformed domain, referred to as pre-whitened-transformed signals 509.
  • filter 530 (i.e., (50), above), which uses the whitened- transformed signals 509 to generate an a priori speech estimate, z c .
  • the a priori speech estimate, z c is a speech estimate using a partial a priori assumed RTF vector and partial estimated RTF vector (i.e., using a priori assumptions for the definition of the RTF vector for the LMA signals, while estimating only the RTF vector for the XM signals).
  • the a priori speech estimate, z c is generated from assumptions such as microphone characteristics, location and reverberant characteristics of the speech within the sound signals detected by the LMA, and based on a real-time estimate of speech within the sound signals detected by the XM, which adhere to the same assumptions used for the LMA.
  • the a priori speech estimate z 1 is an example of an a priori estimate of at least one target sound in the received sound signals.
  • R x rl is a rank-l approximation to R xx (without any a priori information):
  • the estimated RTF vector can therefore be used as an alternative to h for the MVDR a,e :
  • h h q max can be considered as a pre-whitened-transformed filter, which can be used to directly filter the pre-whitened-transformed signals, y.
  • FIG. 6 is a block diagram illustrating a transformation block 502 representing the first transformation of section II-B, in which the LMA signals pass through a blocking matrix 504 and a matched filter 506, analogous to the first stage of a GSC.
  • the XM signals are unaltered.
  • the pre-whitening block 508 represents the pre-whitening operation.
  • the output of the pre-whitening block 508 is signals in the pre-whitened-transformed domain, referred to as pre-whitened-transformed signals 509.
  • filter 532 (i.e., (55), above), which uses the whitened- transformed signals 509 to generate a direct speech estimate, z L .
  • the direct speech estimate, z c is a speech estimate using an estimated RTF vector including both the LMA and XM signals.
  • the speech estimate, z c is generated from a real-time estimate of the speech within the sound signals detected by both the LMA and XM, which takes into consideration microphone characteristics and may contain information such as the location and some reverberant characteristics of the target sound.
  • the speech estimate z c is an example of a direct estimate of at least one target sound in the received sound signals.
  • Wint 9 r La > b w + # est O ⁇ £)w (57) where w x and w x are given (48) and (54) respectively.
  • this integrated MVDR a,e beamformer also reveals that the MVDR a c beamformer based on a priori assumptions from (48) and that which is based on estimated quantities from (54) can be combined according to the functions g pr (a, b) and g est ( a > b) respectively.
  • This integrated beamformer can also be expressed in the pre-whitened-transformed domain as follows:
  • the transformed, pre-whitened signals can be directly filtered accordingly, and then combined with the appropriate weightings as defined by the functions g pr (a, ?) and g est ( a ,/?), to yield the respective speech estimate.
  • These functions $ rG (a, ?) and g est ( a > b) can be tuned such as to emphasize the result from an MVDR beamformer that uses either an a priori assumed RTF vector or an estimated RTF vector. This results in a digital signal processing scheme as depicted in FIG. 7.
  • FIG. 7 is a block diagram of an integrated MVDR a,e beamformer 525 in accordance with embodiments presented herein.
  • the integrated MVDR a,e beamformer 525 comprises a plurality of processing blocks, which include transformation block 502 and pre-whitening block 508.
  • the transformation block 502 represent the first transformation of section II-B, in which the LMA signals pass through a blocking matrix 504 and a matched filter 506, while the XM signals are unaltered.
  • the pre-whitening block 508 represents the pre-whitening operation.
  • the output of the pre-whitening block 508 is signals in the pre-whitened-transformed domain, referred to as pre-whitened-transformed signals 509.
  • the first processing branch 513(1) includes a filter 530 which, as described above with reference to FIG. 5, uses the whitened-transformed signals 509 to generate an a priori speech estimate, z x (i.e., an estimate of the speech in the received sound signals, based on a priori assumptions for the definition of the RTF vector for the LMA signals, while estimating only the RTF vector for the XM signals).
  • the speech estimate z 1 is an example of an a priori estimate of at least one target sound in the received sound signals.
  • the first branch 513(1) also comprises a first weighting block 516.
  • the first weighting block 516 is configured to weight the speech estimate, z c , in accordance with the complex conjugate of the function g pr (a, b) (i.e., (58) and (63), above). More generally, the first weighting block 516 is configured to weight the speech estimate, z c , in accordance with a cost function controlled by a plurality of tuning parameters (e.g., (a, /?)).
  • the tuning parameters of the cost function e.g., g pr ( ,b)
  • the one or more confidence measures 518 represent an assessment or estimate of the accuracy/reliability of the speech estimate, z c , and the hence the accuracy of the partial a priori assumed RTF vector and partial estimated RTF vector used to generate the speech estimate (i.e., using a priori assumptions for the definition of the RTF vector for the LMA signals, while estimating only the RTF vector for the XM signals).
  • the first weighting block 518 generates a weighted a priori speech estimate, shown in FIG. 5 by arrow 519.
  • the second branch 513(2) includes the filter 532 (i.e., (55), above), which uses the whitened-transformed signals 509 to generate a direct speech estimate, z x (i.e., a speech estimate generated using an estimated RTF vector including both the LMA and XM signals).
  • the second branch 513(2) also comprises a second weighting block 520.
  • the second weighting block 520 is configured to weight the direct speech estimate, z c , in accordance with the complex conjugate of the function g est (i.e., (59) and (63), above).
  • the second weighting block 120 is configured to weight the direct speech estimate, 3 ⁇ 4, in accordance with a cost function controlled by a plurality of tuning parameters (e.g., ( a, b )).
  • the tuning parameters of the cost function e.g., g est ( a > b) are set based on one or more confidence measures 522 generated for the speech estimate, z 1.
  • the one or more confidence measures 522 represent an assessment or estimate of the accuracy/reliability of the speech estimate, z c , and the hence the accuracy of the estimated RTF vector including both the LMA and XM signals.
  • the second weighting block 520 generates a weighted direct speech estimate, shown in FIG. 5 by arrow 123.
  • FIG. 7 also illustrates processing block 524 which integrates/combines the weighted a priori speech estimate 519 and the weighted direct speech estimate 523.
  • the combination of the weighted a priori speech estimate 519 and the weighted direct speech estimate 523 is referred to as an integrated speech estimate, z int (i.e., (63), above).
  • the integrated speech estimate, z int may be used for subsequent processing in the device (e.g., auditory prosthesis).
  • the process 840 is comprised of two main decisions, referred to as decisions 842 and 844.
  • decisions 842 and 844 it can be determined whether or not the XM signals are reliable (i.e., decide whether or not to use the XM signals). If the XM signals are not reliable, the system uses MVDR with LMA only (i.e., MVDR a ). If the XM signals are reliable, the system uses MVDR with LMA and XMs (i.e., MVDR a e ).
  • a decision is made as to whether or not estimated RTF vector is reliable. In other words, a decision can then be made on how much to weight the a priori assumed RTF vector and the estimated RTF vector. This decision is controlled by a and b in the same manner as for the Integrated MVDR a Beamformer from section III-C.
  • the a priori assumed RTF vector consists of an a priori assumed RTF vector for the LMA signals and an estimated RTF vector for the XM signals, the estimated RTF vector is for both the LMA and XM signals.
  • a and b could be made inversely proportional, and can even be tuned such that g pr ( ⁇ x ) and g est ( a fl) form a convex combination.
  • a— > ⁇ if it is imposed that a— > ⁇ , then this preserves the a priori constraint and it is only b that remains to be tuned, which would be that of a contingency noise reduction strategy.
  • LCMV linearly constrained minimum variance
  • FIG. 9 includes a table, referred to as Table I, which illustrates limiting cases of a , b for the various MVDR beamformers.
  • the integrated noise reduction techniques presented herein may be implemented in a number of devices/systems that include a local microphone array (LMA) to capture sound signals.
  • LMA local microphone array
  • These devices/ systems include, for example, auditory prostheses (e.g., cochlear implant, acoustic hearing aids, auditory brainstem stimulators, bone conduction devices, middle ear auditory prostheses, direct acoustic stimulators, bimodal auditory prosthesis, bilateral auditory prostheses, etc.), computing devices (e.g., mobile phones, tablet computers, etc.), conference phones, hands-free telephone systems, etc.
  • FIGs. 10A, 10B, 11, and 12 are schematic block diagrams of example devices configured to implement the integrated noise reduction techniques presented herein. It is to be appreciated that these examples are illustrative and that, as noted, the integrated noise reduction techniques presented herein may be implemented in a number of different devices/systems.
  • FIG. 10A shown is a schematic diagram of an exemplary cochlear implant 1000 configured to implement aspects of the techniques presented herein, while FIG. 10B is a block diagram of the cochlear implant 1000.
  • FIGs. 10A and 10B will be described together.
  • the cochlear implant 1000 comprises an external component 1002 and an intemal/implantable component 1004.
  • the external component 1002 includes a sound processing unit 1012 that is directly or indirectly attached to the body of the recipient, an external coil 1006 and, generally, a magnet (not shown in FIG. 10 A) fixed relative to the external coil 1006.
  • the sound processing unit 1012 comprises a local microphone array (LMA) 1013, comprised of microphones 1008(1) and 1008(2), configured to receive sound input signals.
  • the sound processing unit 1012 may also include one or more auxiliary input devices 1009, such as one or more telecoils, audio ports, data ports, cable ports, etc., and a wireless transmitter/receiver (transceiver) 1011.
  • the sound processing unit 1012 also includes, for example, at least one battery 1007, a radio-frequency (RF) transceiver 1021, and a processing block 1050.
  • the processing block 1050 comprises a number of elements, including an integrated noise reduction module 1025 and a sound processor 1033.
  • the processing block 1050 may also include other elements that, have for ease of illustration, been omitted from FIG. 10B.
  • Each of the integrated noise reduction module 1025 and a sound processor 1033 may be formed by one or more processors (e.g., one or more Digital Signal Processors (DSPs), one or more uC cores, etc), firmware, software, etc. arranged to perform operations described herein. That is, the integrated noise reduction module 1025 and a sound processor 1033 may each be implemented as firmware elements, partially or fully implemented with digital logic gates in one or more application- specific integrated circuits (ASICs), partially or fully implemented in software, etc.
  • DSPs Digital Signal Processors
  • ASICs application-specific integrated circuits
  • the integrated noise reduction module 1025 is configured to perform the integrated noise reduction techniques described elsewhere herein.
  • the integrated noise reduction module 1025 corresponds to the integrated MVDR a beamformer 125 and the MVDR a,e beamformer 525, described above.
  • the integrated noise reduction module 1025 may include the processing blocks described above with reference to FIGs. 4 and 7, as well as other combinations of processing blocks configured to perform the integrated noise reduction techniques described elsewhere herein.
  • the integrated noise reduction techniques and thus the integrated noise reduction module 1025, generates an integrated speech estimate from sound signals received via at least the LMA 1013.
  • Shown in FIG. 10 is at least one optional external microphone (XM) which may also be in communication with the sound processing unit 1012. If present, the XM 1017 is configured to capture sound signals and provide XM signals to the sound processing unit 1012. These XM signals may also be used to generate the integrated speech estimate.
  • the sound processor 1033 is configured to use the integrated speech estimate (generated from one or both of the LMA signals and the XM signals) to generate stimulation signals for delivery to the recipient.
  • the implantable component 1004 comprises an implant body (main module) 1014, a lead region 1016, and an intra-cochlear stimulating assembly 1018, all configured to be implanted under the skin/tissue (tissue) 1005 of the recipient.
  • the implant body 1014 generally comprises a hermetically- sealed housing 1015 in which RF interface circuitry 1024 and a stimulator unit 1020 are disposed.
  • the implant body 1014 also includes an internal/implantable coil 1022 that is generally external to the housing 1015, but which is connected to the RF interface circuitry 1024 via a hermetic feedthrough (not shown in FIG. 10B).
  • stimulating assembly 1018 is configured to be at least partially implanted in the recipient’s cochlea 1037.
  • Stimulating assembly 1018 includes a plurality of longitudinally spaced intra-cochlear electrical stimulating contacts (electrodes) 1026 that collectively form a contact or electrode array 1028 for delivery of electrical stimulation (current) to the recipient’s cochlea.
  • Stimulating assembly 1018 extends through an opening in the recipient’s cochlea (e.g., cochleostomy, the round window, etc) and has a proximal end connected to stimulator unit 1020 via lead region 1016 and a hermetic feedthrough (not shown in FIG. 10B).
  • Lead region 1016 includes a plurality of conductors (wires) that electrically couple the electrodes 1026 to the stimulator unit 1020.
  • the cochlear implant 1000 includes the external coil 1006 and the implantable coil 1022.
  • the coils 1006 and 1022 are typically wire antenna coils each comprised of multiple turns of electrically insulated single-strand or multi-strand platinum or gold wire.
  • a magnet is fixed relative to each of the external coil 1006 and the implantable coil 1022.
  • the magnets fixed relative to the external coil 1006 and the implantable coil 1022 facilitate the operational alignment of the external coil with the implantable coil.
  • This operational alignment of the coils 1006 and 1022 enables the external component 1002 to transmit data, as well as possibly power, to the implantable component 1004 via a closely-coupled wireless link formed between the external coil 1006 with the implantable coil 1022.
  • the closely- coupled wireless link is a radio frequency (RF) link.
  • RF radio frequency
  • various other types of energy transfer such as infrared (IR), electromagnetic, capacitive and inductive transfer, may be used to transfer the power and/or data from an external component to an implantable component and, as such, FIG. 10B illustrates only one example arrangement.
  • the integrated noise reduction module 1025 is configured to generate an integrated speech estimate
  • the sound processor 1033 is configured to use the integrated speech estimate to generate stimulation signals for delivery to the recipient. More specifically, the sound processor 1033 (e.g., one or more processing elements implementing firmware, software, etc) is configured to use the integrated speech estimate to generate stimulation control signals 1036 that represent electrical stimulation for delivery to the recipient.
  • the stimulation control signals 1036 are provided to the RF transceiver 1021, which transcutaneously transfers the stimulation control signals 1036 (e.g., in an encoded manner) to the implantable component 1004 via external coil 1006 and implantable coil 1022.
  • the stimulation control signals 1036 are received at the RF interface circuitry 1024 via implantable coil 1022 and provided to the stimulator unit 1020.
  • the stimulator unit 1020 is configured to utilize the stimulation control signals 1036 to generate electrical stimulation signals (e.g., current signals) for delivery to the recipient’s cochlea via one or more stimulating contacts 1026.
  • electrical stimulation signals e.g., current signals
  • cochlear implant 1000 electrically stimulates the recipient’s auditory nerve cells, bypassing absent or defective hair cells that normally transduce acoustic vibrations into neural activity, in a manner that causes the recipient to perceive one or more components of the input audio signals.
  • FIGs. 10A and 10B illustrate an arrangement in which the cochlear implant 1000 includes an external component.
  • embodiments of the present invention may be implemented in cochlear implants having alternative arrangements.
  • the techniques presented herein could also be implemented in a totally implantable or mostly implantable auditory prosthesis where components shown in sound processing unit 1012, such as processing block 1050, could instead be implanted in the recipient.
  • FIG. 11 is a functional block diagram of one example arrangement for a bone conduction device 1100 in accordance with embodiments presented herein.
  • Bone conduction device 1100 is configured to be positioned at (e.g., behind) a recipient’s ear.
  • the bone conduction device 1100 comprises a microphone array 1113, an electronics module 1170, a transducer 1171, a user interface 1172, and a power source 1173.
  • the local microphone array (LMA) 1113 comprises microphones 1108(1) and 1108(2) that are configured to convert received sound signals 1116 into LMA signals.
  • bone conduction device 1100 may also comprise other sound inputs, such as ports, telecoils, etc.
  • the LMA signals are provided to electronics module 1170 for further processing.
  • electronics module 1170 is configured to convert the LMA signals into one or more transducer drive signals 1180 that active transducer 1171. More specifically, electronics module 1170 includes, among other elements, a processing block 1150 and transducer drive components 1176.
  • the processing block 1174 comprises a number of elements, including an integrated noise reduction module 1125 and sound processor 1133.
  • Each of the integrated noise reduction module 1125 and the sound processor 1133 may be formed by one or more processors (e.g., one or more Digital Signal Processors (DSPs), one or more uC cores, etc.), firmware, software, etc. arranged to perform operations described herein. That is, the integrated noise reduction module 1125 and the sound processor 1133 may each be implemented as firmware elements, partially or fully implemented with digital logic gates in one or more application-specific integrated circuits (ASICs), partially or fully in software, etc.
  • DSPs Digital Signal Processors
  • ASICs application-specific integrated circuits
  • the integrated noise reduction module 1125 is configured to perform the integrated noise reduction techniques described elsewhere herein.
  • the integrated noise reduction module 1125 corresponds to the integrated MVDR a beamformer 125 and the MVDR a,e beamformer 525, described above.
  • the integrated noise reduction module 1125 may include the processing blocks described above with reference to FIGs. 4 and 7, as well as other combinations of processing blocks configured to perform the integrated noise reduction techniques described elsewhere herein.
  • at least one optional external microphone (XM) may be in communication with the bone conduction device 1100. If present, the XM is configured to capture sound signals and provide XM signals to the conduction device 1100 for processing by the integrated noise reduction module 1125 (i.e., the XM signals may also be used to generate the integrated speech estimate).
  • XM external microphone
  • the sound processor 1133 is configured to process the integrated speech estimate (generated from one or both of the LMA signals and the XM signals) for use by the transducer drive components 1176.
  • the transducer drive components 1176 generate transducer drive signal(s) 1180 which are provided to the transducer 1171.
  • the transducer 1171 illustrates an example of a stimulation unit that receives the transducer drive signal(s) 1180 and generates vibrations for delivery to the skull of the recipient via a transcutaneous or percutaneous anchor system (not shown) that is coupled to bone conduction device 1100. Delivery of the vibration causes motion of the cochlea fluid in the recipient’s contralateral functional ear, thereby activating the hair cells in the functional ear.
  • FIG. 11 also illustrates the power source 1173 that provides electrical power to one or more components of bone conduction device 1300.
  • Power source 1173 may comprise, for example, one or more batteries.
  • power source 1173 has been shown connected only to user interface 1172 and electronics module 1170. However, it should be appreciated that power source 1173 may be used to supply power to any electrically powered circuits/components of bone conduction device 1100.
  • User interface 1172 allows the recipient to interact with bone conduction device 1100.
  • user interface 1172 may allow the recipient to adjust the volume, alter the speech processing strategies, power on/off the device, etc.
  • bone conduction device 1100 may further include an external interface that may be used to connect electronics module 1170 to an external device, such as a fitting system.
  • FIG. 12 is a block diagram of an arrangement of a mobile computing device 1200, such as a smartphone, configured to be implemented the integrated noise reduction techniques presented herein. It is to be appreciated that FIG.1 2 is merely illustrative.
  • Mobile computing device 1200 first comprises an antenna 1236 and a telecommunications interface 1238 that are configured for communication on a telecommunications network.
  • the telecommunications network over which the radio antenna 1236 and the radio interface 1238 communicate may be, for example, a Global System for Mobile Communications (GSM) network, code division multiple access (CDMA) network, time division multiple access (TDMA), or other kinds of networks.
  • GSM Global System for Mobile Communications
  • CDMA code division multiple access
  • TDMA time division multiple access
  • the mobile computing device 1200 also includes a wireless local area network interface 1240 and a short-range wireless interface/transceiver 1242 (e.g., an infrared (IR) or Bluetooth® transceiver).
  • IR infrared
  • Bluetooth® is a registered trademark owned by the Bluetooth® SIG.
  • the wireless local area network interface 1240 allows the mobile computing device 1200 to connect to the Internet, while the short-range wireless transceiver 1242 enables the external device 1206 to wirelessly communicate (i.e., directly receive and transmit data to/from another device via a wireless connection), such as over a 2.4 Gigahertz (GHz) link.
  • a wireless local area network interface 1240 allows the mobile computing device 1200 to connect to the Internet, while the short-range wireless transceiver 1242 enables the external device 1206 to wirelessly communicate (i.e., directly receive and transmit data to/from another device via a wireless connection), such as over a 2.4 Gigahertz (GHz) link.
  • GHz
  • any other interfaces now known or later developed including, but not limited to, Institute of Electrical and Electronics Engineers (IEEE) 802.11, IEEE 802.16 (WiMAX), fixed line, Long Term Evolution (LTE), etc., may also or alternatively form part of the mobile computing device 1200.
  • IEEE Institute of Electrical and Electronics Engineers
  • WiMAX IEEE 802.16
  • LTE Long Term Evolution
  • mobile computing device 1200 also comprises an audio port 1244, a local microphone array (LMA) 1213, a speaker 1248, a display screen 1258, a subscriber identity module or subscriber identification module (SIM) card 1252, a battery 1254, a user interface 1256, one or more processors 1250, and a memory 1260.
  • LMA 1213 includes microphones 1208(1) and 1208(2).
  • Stored in memory 1260 is integrated noise reduction logic 1225 and sound processing logic 1233.
  • the display screen 1258 is an output device, such as a liquid crystal display (LCD), for presentation of visual information to the cochlear implant recipient.
  • the user interface 1256 may take many different forms and may include, for example, a keypad, keyboard, mouse, touchscreen, display screen, etc.
  • Memory 1260 may comprise any one or more of read only memory (ROM), random access memory (RAM), magnetic disk storage media devices, optical storage media devices, flash memory devices, electrical, optical, or other physical/tangible memory storage devices.
  • the one or more processors 1258 are, for example, microprocessors or microcontrollers that execute instructions for the integrated noise reduction logic 1225 and sound processing logic 1233.
  • the integrated noise reduction logic 1225 When executed by the one or more processors 1250, the integrated noise reduction logic 1225 is configured to perform the integrated noise reduction techniques described elsewhere herein.
  • the integrated noise reduction logic 1225 corresponds to the integrated MVDR a beamformer 125 and the MVDR a,e beamformer 525, described above.
  • the integrated noise logic 1225 may include software forming the processing blocks described above with reference to FIGs. 4 and 7, as well as other combinations of processing blocks configured to perform the integrated noise reduction techniques described elsewhere herein to generate an integrated noise estimate.
  • the sound processing logic 1233 When executed by the one or more processors 1250, the sound processing logic 1233 is configured to perform sound processing operations using the integrated noise estimate.
  • FIG. 13 is a flowchart of a method 1390 performed/executed by a device comprising at least a local microphone array (LMA), in accordance with embodiments presented herein.
  • Method 1390 begins at 1392 where sound signals are received with at least the local microphone array of the device.
  • the received sound signals comprise/include at least one target sound.
  • an a priori estimate of the at least one target sound in the received sound signals is generated, wherein the a priori estimate is based at least on a predetermined location of a source of the at least one target sound.
  • a direct estimate of the at least one target sound in the received sound signals is generated, wherein the direct estimate is based at least on a real-time estimate of a location of a source of the at least one target sound.
  • a weighted combination of the a priori estimate and the direct estimate is generated, where the weighted combination is an integrated estimate of the target sound. Subsequent sound processing operations may be performed in the device using the integrated estimate of the target sound.
  • the a priori estimate of the at least one target sound is generated using only an a priori relative transfer function (RTF) vector generated from the received sound signals.
  • the direct estimate of the at least one target sound is generated using only an estimated relative transfer function (RTF) vector for the received sound signals.
  • the weighted combination of the a priori estimate and the direct estimate is generated by weighting the a priori estimate in accordance with a first cost function controlled by a first set of tuning parameters to generate a weighted a priori estimate; and weighting the direct estimate in accordance with a second cost function controlled by a second set of tuning parameters to generate a weighted direct estimate.
  • the weighted direct estimate with the weighted a priori estimate are then mixed with one another.
  • the first set of tuning parameters may be set based on one or more confidence measures associated with the a priori estimate of the of the at least one target sound, wherein the one or more confidence measures represent an estimate of a reliability of the a priori estimate.
  • the second set of tuning parameters may be set based on one or more confidence measures associated with the direct estimate of the of the at least one target sound, wherein the one or more confidence measures represent an estimate of a reliability of the direct estimate.
  • integrated noise reduction techniques sometimes referred to as an integrated beamformer (e.g., an integrated MVDR a beamformer or an integrated MVDR a,e beamformer).
  • the integrated noise reduction techniques combine the use of an a priori (i.e., predetermined, assumed, or pre-defmed) location of a target sound source with a real-time estimated location of the sound source.
  • a pre-whitened-transformed version of the a priori assumed RTF vector can be considered where:
  • Tlris estimated RTF vector can now be used as an alternative to h a for the MVDR a defined in (25), and is given by:
  • This filter based on estimated quantities can also be reformulated in the pre-whitened- transformed domain.
  • the pre-whitening operation can also be included in the optimisation problem:
  • K A is an (ikf a— 1) x (M a— 1) matrix
  • K B an (M a -- 1) x M I 1) matrix a (M e -f- 1) x (M a --- 1) matrix
  • K x rl and K X are ( M e -f 1) x ( e -;- 1) matrices realised as:
  • this estimate is then used to compute the corresponding MVDR 3 e filter with an a priori assumed RTF vector and a partially estimated RTF vector, along with the penalty term as:
  • This filter can also be realised in the pre- whitened- transformed domain.
  • the pre-whitened-transfbnned version of h can firstly be considered where:
  • This filter based on estimated quantities can also be reformulated in the pre-whitened-transformed domain.
  • This filter based on estimated quantities can also be reformulated in the pre-whitened-transformed domain.

Landscapes

  • Physics & Mathematics (AREA)
  • Engineering & Computer Science (AREA)
  • Acoustics & Sound (AREA)
  • Signal Processing (AREA)
  • Health & Medical Sciences (AREA)
  • General Health & Medical Sciences (AREA)
  • Otolaryngology (AREA)
  • Neurosurgery (AREA)
  • Circuit For Audible Band Transducer (AREA)

Abstract

Presented herein are techniques for generated an integrated estimate of a target sound (e.g., speech) in sound signals received by at least a local microphone array of a device. In embodiments, the integrated estimate may be generated based on sound signals received by the at least a local microphone array of a device and at least one external microphone.

Description

INTEGRATED NOISE REDUCTION
BACKGROUND
Field of the Invention
[oooi] The present invention generally relates to integrated noise reduction for devices having at least one local microphone array.
Related Art
[0002] Hearing loss is a type of sensory impairment that is generally of two types, namely conductive and/or sensorineural. Conductive hearing loss occurs when the normal mechanical pathways of the outer and/or middle ear are impeded, for example, by damage to the ossicular chain or ear canal. Sensorineural hearing loss occurs when there is damage to the inner ear, or to the nerve pathways from the inner ear to the brain.
[0003] Individuals who suffer from conductive hearing loss typically have some form of residual hearing because the hair cells in the cochlea are undamaged. As such, individuals suffering from conductive hearing loss typically receive an auditory prosthesis that generates motion of the cochlea fluid. Such auditory prostheses include, for example, acoustic hearing aids, bone conduction devices, and direct acoustic stimulators.
[0004] In many people who are profoundly deaf, however, the reason for their deafness is sensorineural hearing loss. Those suffering from some forms of sensorineural hearing loss are unable to derive suitable benefit from auditory prostheses that generate mechanical motion of the cochlea fluid. Such individuals can benefit from implantable auditory prostheses that stimulate nerve cells of the recipient’s auditory system in other ways (e.g., electrical, optical and the like). Cochlear implants are often proposed when the sensorineural hearing loss is due to the absence or destruction of the cochlea hair cells, which transduce acoustic signals into nerve impulses. An auditory brainstem stimulator is another type of stimulating auditory prosthesis that might also be proposed when a recipient experiences sensorineural hearing loss due to damage to the auditory nerve.
SUMMARY
[0005] In one aspect, a method is provided. The method comprises: receiving sound signals with at least a local microphone array of a device, wherein the sound signals comprise at least one target sound; generating an a priori estimate of the at least one target sound in the received sound signals based on a predetermined location of a source of the at least one target sound; generating a direct estimate of the at least one target sound in the received sound signals based on a real-time estimate of a location of a source of the at least one target sound; and generating a weighted combination of the a priori estimate and the direct estimate, wherein the weighted combination is an integrated estimate of the target sound.
[0006] In another aspect, a device is provided. The device comprises: a local microphone array configured to receive sound signals, wherein the sound signals comprise at least one target sound; and one or more processors configured to: generate an a priori estimate of the at least one target sound in the received sound signals using only an a priori relative transfer function (RTF) vector generated from the received sound signals, generate a direct estimate of the at least one target sound in the received sound signals using only an a priori relative transfer function (RTF) vector generated from the received sound signals, and generate a weighted combination of the a priori estimate and the direct estimate, wherein the weighted combination is an integrated estimate of the target sound.
BRIEF DESCRIPTION OF THE DRAWINGS
[0007] Embodiments of the present invention are described herein in conjunction with the accompanying drawings, in which:
[0008] FIG. 1 is a functional block diagram illustrating the generation of pre-whitened transformed signals;
[0009] FIG. 2 is a functional block diagram illustrating the generation of an a priori estimate of at least one target sound in sound signals received at a local microphone array;
[ooio] FIG. 3 is a functional block diagram illustrating the generation of a direct estimate of at least one target sound in sound signals received at a local microphone array;
[ooii] FIG. 4 is a functional block diagram illustrating the generation of an integrated estimate of at least one target sound in sound signals received at a local microphone array;
[0012] FIG. 5 is a functional block diagram illustrating the generation of an a priori estimate of at least one target sound in sound signals received at a local microphone array and at least one external microphone; [0013] FIG. 6 is a functional block diagram illustrating the generation of a direct estimate of at least one target sound in sound signals received at a local microphone array and at least one external microphone;
[0014] FIG. 7 is a functional block diagram illustrating the generation of an integrated estimate of at least one target sound in sound signals received at a local microphone array and at least one external microphone;
[0015] FIG. 8 is flowchart of a two stage process, in accordance with embodiments presented herein;
[0016] FIG. 9 is a table summarizing the various noise reduction strategies, in accordance with embodiments presented herein;
[0017] FIG. 10A is a schematic diagram illustrating a cochlear implant, in accordance with certain embodiments presented herein;
[0018] FIG. 10B is a block diagram of the cochlear implant of FIG. 10A;
[0019] FIG. 11 is a block diagram of a totally implantable cochlear implant, in accordance with certain embodiments presented herein;
[0020] FIG. 12 is a block diagram of a bone conduction device that includes a spatial pre-filter, in accordance with embodiments presented herein.
[0021] FIG. 13 is a flowchart of a method, in accordance with embodiments presented herein.
DETAILED DESCRIPTION
I. Introduction
[0022] In devices having one or more microphone arrays, such as auditory prostheses (e.g., hearing aids, cochlear implants, bone conduction devices, etc.), multi -microphone noise reduction systems are used to preserve desired sounds (e.g., speech), while rejecting unwanted sounds (e.g., noise). In certain conventional noise reduction systems, a local microphone array (LMA) worn on the recipient (i.e., part of the device) is used to focus on a sound source (e.g., speaker) that is in a predefined direction, such as directly in front of recipient. While such a noise reduction system may be robust, it is also prone to poor performance in situations where the desired speaker is not in the predefined direction. Examples of such situations may be found in classroom environments or while a recipient is travelling in a motor vehicle. The integrated noise reduction techniques presented herein improve upon these existing noise reduction systems in several distinct ways: (i) by including the ability to focus on a target sound source (e.g., speaker) that is not in the predefined direction and, in certain arrangements, (ii) by including external microphones (XMs) that operate together with the LMA, resulting in further noise reduction as opposed to using only the LMA.
[0023] In certain embodiments presented herein, integrated noise reduction techniques will utilize two separate tuning parameters, one for controlling the sound received from the predefined direction, and the other for the sound received from an estimated direction where the target sound source may be located. In these embodiments, each of these directions can be defined using the LMA and the XMs. In order to define the predefined direction with the LMA and the XMs, a modified version of the improved method of estimation of a transfer function for the XM is used, where the input signals have to undergo a specific series of transformations.
[0024] Using one or several XMs along with the LMA can provide significant speech intelligibility improvement, for instance in the case where XMs may be quite close to the desired speaker, or even if it provides a relevant noise reference. Additionally, the integrated noise reduction techniques presented herein are flexible in that they encompass a wide range of noise reduction options according to the tuning of the system.
[0025] For ease of understanding, the following description is organized into several sections. In particular, section II describes a data model, which considers the general case of a local microphone array (LMA) in conjunction with one or several external microphones (XMs), which can be reduced to a single external microphone without compromising the equations provided herein. A transformed domain, as well as a pre-whitened-transformed domain is also introduced in order to simplify the flow of signal processing operations and realize distinct digital signal processing (DSP) block schemes.
[0026] In section III, an integrated minimum variance distortionless response (MVDR) beamformer is discussed as applied to a local microphone array. In particular, section III describes an integrated MVDR beamformer, which leverages the use of a priori assumptions and the use of estimated quantities. In section IV, an integrated MVDR beamformer as applied to a local microphone array together with one or more external microphones is described. Again, an integrated MVDR beamformer for application to a local microphone array together with one or more external microphones, which leverages the use of a priori assumptions and the use of estimated quantities is described.
II. Data Model
A. Unprocessed Signals
[0027] Consider a noise reduction system that consists of a local microphone array (LMA) of Ma microphones and Me external microphones, providing a total of Ma + Me number of microphones. Also consider a scenario where there is only one desired/target sound source, such as a target speech source, in a noisy environment. Proceeding to formulate the problem in the short-time Fourier transform (STFT) domain, the received signal can be represented at one particular frequency, k , and one time frame, l as: y(k, l) = x(k, l) + n(k, l) (1)
= a (k, l)s(k, l ) + n (k, l ) (2)
where (dropping the dependency on k and l for brevity), y = [y ye]T <ya = \y a,i Ya.2 - y a,M„]T are the local microphone signals, ye = [ye,i ye,2 - ye,Me]T are the external microphone signals, x is the speech component consisting of a =
Figure imgf000006_0001
the acoustic transfer function (ATF) from the speech source to all Ma + Me microphones and s, the speech source signal. Finally, n = [na n ]T represents the noise component, which consists of a combination of correlated and uncorrelated noises. Variables with the subscript“a” refer to the LMA signals and variables with the subscript“e” refer to the XM signals. The dependencies on k and l will be introduced herein, as needed, for mathematical derivations.
[0028] In general, the speech component (target sound), x, can be represented in terms of a relative transfer function (RTF) vector such that:
x = as = hs-L (3) where sx = aa ls, is the speech in a reference microphone of the LMA (w.l.o.g the first microphone is chosen as the reference microphone) and h is the RTF vector defined as:
Figure imgf000007_0001
consisting of an RTF vector corresponding to the LMA signals, ha and an RTF vector corresponding to the XM signals, he. With such a formulation, the noise reduction system will aim to produce an estimate for the speech component in the reference microphone,
Figure imgf000007_0002
[0029] The (Ma + Me) X (Ma + Me) speech-plus-noise, noise-only, and speech-only spatial correlation matrices are given respectively as:
Ryy = E{yyH) (5) Rnn = E{nnH) (6) Rxx = E{xxH) (7)
where E{. } is the expectation operator and H is the Hermitian transpose. It is assumed that the speech components are uncorrelated with the noise components, and hence the speech-only correlation matrix can be found from the difference of the speech-plus-noise correlation matrix and the noise-only correlation matrix:
R lvxx = R lvyy— R lvnn (8) 7
The speech-plus-noise and noise-only correlation matrices are estimated from the received microphone signals during speech-plus-noise and noise-only periods, using a voice activity detector (VAD). The correlation matrices can also be calculated solely for the LMA signals respectively as Ryaya =
Figure imgf000007_0003
(which can be realized by the top left (Ma X Ma) block of the corresponding entire correlation matrices in (5)-
(7))· [0030] The estimate of the speech component in the reference microphone, zc, is then obtained through the linear filtering of the microphone signals, such that:
Zi = wHy (9)
Where w = [wjwe ]T is the complex- valued filter to be designed.
B. Transformed Domain
[0031] As will be described later, working with the signals in a transformed domain will result in convenient relations to be made and an overall simplification of the flow of signal processing operations. The transformation will be based on an a priori assumed RTF vector for the LMA signals, ha (which may or may not be equal to ha). Firstly, an Ma X (Ma— 1) unitary blocking matrix Ba for ha and an Ma x 1 vector ba are defined such that:
Figure imgf000008_0001
where Ba Ba = l(Ma-i) and in general denotes the ϋ x ϋ identity matric, and ba can be interpreted as a scaled matched filter. W.l.o.g, ba will simply be referred to as a matched filter in the following derivations. Using Ba and ba, an (Ma + Me) X (Ma + Me) unitary transformation matrix, T, can be subsequently defined:
Figure imgf000008_0002
where Ta = [ Ba ba], T Ta = IMa, and hence indeed THT = fMa+Mey Consequently, the transformed input signals, y, become:
Figure imgf000008_0003
The transformed noise signals can also be similarly defined:
Figure imgf000008_0004
[0032] It should be understood that this transformation domain is the LMA signals that pass through a blocking matrix and a matched filter, as in the first stage of a generalized sidelobe canceller (GSC) (i.e., the adaptive implementation of an MVDR beamformer), along with the XM signals.
C. Pre-Whitened-Transformed Domain
[0033] A spatial pre-whitening operation can be defined from the noise-only correlation matrix in the previously described transform domain by using the Cholesky decomposition:
Έ{(THh (T)H} = LIP (14) where L is an (Ma + Me) X (Ma + Me) lower triangular matrix. In block form, L can be realized as:
L = (15)
Figure imgf000009_0004
[0034] Where La and Lx are lower triangular matrices. It should be noted that La corresponds to the LMA signals and are from a Cholesky decomposition of the noise correlation matrix from the LMA signals in the transformed domain, hence:
E{(7’i a)(ra¾a)ff} = LaL2 (16)
[0035] A signal vector in the transformed domain can be consequently pre-whitened by pre-multiplying it with LT1. Such signal quantities will be denoted with the underbar Q notation. Hence, the signal y in this so-called pre-whitened-transformed domain is given by:
(17)
Figure imgf000009_0003
and similarly for n:
Figure imgf000009_0001
The respective correlation matrices are also given by:
Figure imgf000009_0002
The spatial correlation matrices for the speech and noise and the noise-only, and the speech- only can also be calculated solely for the LMA signals respectively as Ryaya =
Figure imgf000010_0001
D. Summary of symbols and realization
[0036] FIG. 1 is a block diagram illustrating the flow of the previously described transformations on the unprocessed signals. Transformation block 102 is a processing block that represents the first transformation of section II-B, in which the LMA signals pass through a blocking matrix 104 and a matched filter 106, analogous to the first stage of a GSC. The XM signals are unaltered. The pre-whitening block 108 is a processing block that represents the pre-whitening operation of section II-C, yielding signals 109 in the pre-whitened-transformed domain. The noise reduction filters that will be developed below will then be directly applied to these pre-whitened-transformed signals (i.e., the output of pre-whitening block 108) in order to yield the desired speech estimate.
[0037] The following is also a summary of how the symbolic notation should be interpreted throughout this document:
• (. )a refer to quantities associated with the LMA signals, e.g., ya.
• (. )e refer to quantities associated with the XM signals, e.g., ye .
• (. ) refer to a priori assumed quantities, e.g., h.
• (. ) refer to estimated quantities, e.g., h.
• (. ) refer to quantities in the pre-whitened-transformed domain, e.g., ya.
III. MVDR using a LMA (MVDRA
[0038] The MVDR beamformer minimizes the total noise power (minimum variance), while preserving the received signal in a particular direction (distortionless response). This direction is specified by defining the appropriate RTF vector for the MVDR beamformer. Considering only the LMA, the MVDR problem can be formulated as follows (which will be referred to as the MVDRa):
Figure imgf000010_0002
where ha is the RTF vector from (4), which in practice is unknown and hence will be replaced either by a priori assumptions or estimated from the speech-plus-noise correlation matrices. The optimal noise reduction filter is then given by:
Figure imgf000011_0001
Finally, the speech estimate, za l, from this MVDRa beamformer is obtained through the linear filtering of the microphone signals with the complex -valued filter wa :
za,l = wa Ja (24)
[0039] In sections III-A and III-B, strategies for designing an MVDRa beamformer using an RTF vector based either on a priori assumptions or estimated from the speech-plus-noise correlation matrices are discussed. Section III-C illustrates an integrated beamformer that integrates the use of priori assumptions with estimates.
A. Using an a priori assumed RTF Vector
[0040] The MVDRa problem can be formulated as in (22), except with using an a priori assumed RFT vector, ha = [l ha 2 --- ha M] instead of ha. This ha can be based on a priori assumptions regarding microphone characteristics, position, speaker location and room acoustics (e.g., no reverberation). Similar to (23), the optimal noise reduction filter is then given by: n haha ha
wn
h p-1 h
lla lxnana lla (25)
The speech estimate, za l, from this MVDRa with an a priori assumed RTF vector is then:
¾, i = w ya (26)
[0041] This conventional formulation of the MVDRa can also be equivalently posed in the pre- whitened-transformed domain (section II-C). As derived in Appendix A, the speech estimate in this domain is given by:
(27)
Figure imgf000011_0002
Where lMa is the bottom-right element in La, and y a Ma is the last component of the pre- whitened-transformed signals, ya. In other words, the speech estimate for an MVDRa filter that uses an a priori assumed RTF vector results in a simple scaling of the last component of the pre-whitened-transformed signals. With such a formulation in this domain, this beamforming algorithm can be realized in a distinct set of signal processing blocks as illustrated in FIG. 2.
[0042] More specifically, FIG. 2 illustrates transformation block 102 and pre-whitening block 108, as described above with reference to FIG. 1. However, in the example of FIG. 2, in - whitening block 108, the only the last row of L^1 is used, (16), thus the resulting in the signal ya Ma. Also shown is an a priori filter 110, which produces pp and processing block 112 which applies pp to ya Ma. The application of pp to ya Ma produces an a priori speech estimate za 1. The a priori speech estimate, za l, is an estimate of the target sound (e.g., speech) in the received sound signals, based solely on an a priori RTF vector. The RTF vector is generated uses assumptions regarding, for example, location of the source of the target sound, characteristics of the microphones (e.g., microphone calibration in regards to gains, phases, etc.), reverberant characteristics of the target sound source, etc. The a priori speech estimate za 1, is an example of an a priori estimate of at least one target sound in the received sound signals.
B. Using an estimated RTF vector
[0043] The RTF vector may also be estimated without reliance on any a priori assumptions and can be used to enhance the speech regardless of the speech source location. One such method is a method of covariance whitening or equivalently that which involves a Generalized Eigenvalue Decomposition (GEVD).
[0044] In such examples, a rank- 1 matrix approximation problem can be formulated to estimate the RTF vector for a given set of LMA signals such that: iPin ll (Ryaya - Rnana) - Rxa,rl ||J (28)
Kx,ri where ||. || is the Frobenius norm, and Rxa I-i is a rank-l approximation to (Ryaya— Rnana) defined as:
Figure imgf000012_0001
7
Where ha = [lfia,2 - h a,Ma] is the estimated RTF vector. [0045] As opposed to using the raw signal correlation matrices, the estimation problem of (28) can be equivalently formulated in the pre-whitened-transformed domain. In appendix B, it is shown that the estimated RTF vector is then:
Figure imgf000013_0001
where pmax is a generalized eigenvector of the matrix pencil {Ryaya, Rnana}, which as a result of the pre-whitening (Rnana
Figure imgf000013_0002
corresponds to the principal (first in this case) eigenvector of Ryaya, the scaling hr = ealTaLaPmax and the M x 1 vector eal = [1 0 ... 0]T. The resulting MVDRa using this estimated RTF vector is now given by:
Figure imgf000013_0003
[0046] As was done in section III-A, this filter based on estimated quantities can also be reformulated in the transformed, pre-whitened-transformed domain. Leaving the derivations once again to Appendix B, the corresponding speech estimate using the estimated RTF vector is:
Za,i— Pp Pmax
Figure imgf000013_0004
where PpPmax can be considered as the pre-whitened-transformed filter (where {. }* is the complex conjugate), which can be used to directly filter the pre-whitened, transformed signals, ya. These operations can also be realized in a distinct set of signal processing blocks, as illustrated in FIG. 3.
[0047] More specifically, FIG. 3 illustrates transformation block 102 and pre-whitening block 108, as described above with reference to FIG. 1, which produce pre-whitened-transformed signals. Also shown is block 114, which filters the pre-whitened-transformed signals in accordance with PpPmax (he., 114 represents the hermitian transposed pre-whitened- transformed filter). The output of the pre-whitened-transformed filter 114 is a direct speech estimate, za 1 (i.e., (32), above).
[0048] The direct speech estimate, za 1, is an estimate of the target sound (e.g., speech) in the received sound signals, based solely on an estimated RTF vector. The estimated RTF vector is generated using real-time estimates of, for example, the location of the source of the target sound, reverberant characteristics of the target sound source, etc. The direct speech estimate, za l, is an example of a direct estimate of at least one target sound in the received sound signals.
C. Integrated MVDRa Beamformer
[0049] Described above are two general MVDR approaches, one that imposes a priori assumptions for the definition of the RTF vector in the MVDR filter, and another that involves an estimation of this RTF vector. In conventional arrangements, a choice typically has to be made between one of these approaches with an acceptance of their inevitable drawbacks. However, in accordance the integrated noise reduction techniques presented herein, both approaches are integrated into one global filter, referred to herein as an“integrated MVDRa beamformer” that exploits the benefits of each approach.
[0050] In general, the integrated MVDRa beamformer provides for integrated tunings which allow different“weights” to be applied to each of (1) an a priori assumed representation of target sound within received sound signals (e.g., an a priori estimate of at least one target sound in the received sound signals), and (2) an estimated representation of the target sound within received sound signals (e.g., a direct estimate of at least one target sound in the received sound signal). The weights applied to each of the a priori assumed representation of the target sound and the estimated representation of the target sound are selected based on“confidence measures” associated with each of the a priori assumed representation of the target sound and the estimated representation of the target sound, respectively.
[0051] For instance, with the integrated MVDRa beamformer, if the speech source moves outside of the direction defined by an a priori assumed RTF vector, more weight can be given to an estimated RTF vector to account for the loss in performance that would otherwise result from using the a priori assumed RTF vector alone. On the other hand, if the estimated RTF vector becomes unreliable, less weight can be given thereto and the system can revert to using the a priori assumed RTF vector, which may have an improved performance if the speech source is indeed in the direction defined by the a priori assumed RTF vector. Combination/mixing of the a priori assumed RTF vector and the estimated RTF vector is also possible. That is, the tuning parameters can achieve multiple beamformers, i.e. one that relies on a priori assumptions alone, one that relies on estimated quantities alone, or the mixture of both. [0052] One particular tuning of interest may be to place a large weight on an a priori assumed RTF vector, but weighting an estimated RTF vector only when appropriate. This represents a mechanism for reverting to an a priori assumed RTF vector when the estimated RTF vector was unreliable.
[0053] In the following, the integrated MVDRa beamformer is briefly derived. If the case is considered where ha is defined according to a priori assumptions and ha is estimated from (86), an integrated MVDRa cost function can be given as:
Figure imgf000015_0001
where a e [0, ¥] and b E [0, ¥] are tuning parameters that control how much of the respective RTF vectors (i.e., the a priori assumed RTF vector and the estimated RTF vector) are weighted. This cost function is the combination of that of an MVDRa (as in (22)) defined by ha and another defined by ha, except that the constraints have been softened by a and b.
[0054] The solution to (33) is given by: a,int
Figure imgf000015_0002
where wa and wa are defined in (25) and (31) respectively.
Figure imgf000015_0003
with the constants:
Figure imgf000015_0004
[0055] This integrated MVDR beamformer reveals that the MVDRa beamformer based on a priori assumptions from (25) and that which is based on estimated quantities from (31) can be combined according to the functions fpr(a, b) and fest(a > b) respectively.
[0056] As in the previous sections, this integrated beamformer can also be expressed in the pre-whitened-transformed domain as follows:
(38)
Figure imgf000016_0001
and with the constants equivalently, but alternatively defined as:
Figure imgf000016_0002
where ha and ha are given in (79) and (88) respectively.
[0057] The resulting speech estimate from this integrated beamformer is then given by:
Figure imgf000016_0006
[0058] The benefit of this pre-whitened-transformed domain is apparent where, with such an integrated beamformer of (38), wa Afa and wa can be directly used to filter the pre-whitened- transformed signals, and then combined with the appropriate weightings as defined by the functions fpr(a, b) and
Figure imgf000016_0003
to yield the respective speech estimate. These functions brn(a, b) and fest(a >b) can be tuned such as to emphasize the result from an MVDR beamformer that uses either an a priori assumed RTF vector or an estimated RTF vector. This results in a digital signal scheme as depicted in FIG. 4.
[0059] More specifically, FIG. 4 is a block diagram of an integrated MVDRa beamformer 125 in accordance with embodiments presented herein. The integrated MVDRa beamformer 125 comprises a plurality of processing blocks, which include transformation block 102 and pre- whitening block 108. As described above with reference to FIG. 1 transformation block 102 and pre-whitening block 108 produce signals 109 in the pre-whitened-transformed domain (pre-whitened-transformed signal s) .
[0060] Also shown in FIG. 4 are two processing branches 113(1) and 113(2) that each operate based on all or part of the pre-whitened-transformed signals 109. The first processing branch
113(1) includes an a priori filter 110, which produces
Figure imgf000016_0004
and a processing block 112 which
||ha||
applies to ya Ma · The application of to ya Ma generates the a priori speech estimate za l,
Figure imgf000016_0005
that is generated based solely on an a priori RTF vector (i.e., an estimate of the speech in the received sound signals, based solely on a priori assumptions such as microphone characteristics, source location, and reverberant characteristics of the target sound (e.g., speech) source. In other words, application of to ya Ma generates an a priori estimate of at least one target sound in the received sound signals.
[0061] The first branch 113(1) also comprises a first weighting block 116. The first weighting block 116 is configured to weight the speech estimate, za l, in accordance with the complex conjugate of the function fpr(a, b) (i.e., (35) and (40), above). More generally, the first weighting block 116 is configured to weight the speech estimate, za 1, in accordance with a cost function controlled by a plurality of tuning parameters (e.g., (a, /?)). The tuning parameters of the cost function (e.g., /pr(a, ?)), are set based on one or more confidence measures 118 generated for the speech estimate, za l. The one or more confidence measures 118 represent an assessment or estimate of the accuracy/reliability of the a priori speech estimate, za 1, and the hence the accuracy of the a priori RTF vector used to generate the speech estimate, za 1. The first weighting block 116 generates a weighted a priori speech estimate, shown in FIG. 5 by arrow 119.
[0062] The second branch 113(2) includes a pre-whitened-transformed filter 114, which filters the pre-whitened-transformed signals in accordance with (32). The output of the pre-whitened- transformed filter 114 is a direct speech estimate, za l, that is generated based solely on an estimated RTF vector (i.e., an estimate of the speech in the received sound signals, which takes into consideration microphone characteristics and may contain information such as the location and some reverberant characteristics of the speech source). In other words, the direct speech estimate za 1, is an example of a direct estimate of at least one target sound in the received sound signals.
[0063] The second branch 113(2) also comprises a second weighting block 120. The second weighting block 120 is configured to weight the speech estimate, za 1, in accordance with complex conjugate of the function fest(<x, ) (i.e., (36) and (40), above). More generally, the second weighting block 120 is configured to weight the direct speech estimate, za 1, in accordance with a cost function controlled by a plurality of tuning parameters (e.g., (a, b)). The tuning parameters of the cost function (e.g., fest a > b) are set based on one or more confidence measures 122 generated for the speech estimate, za l. The one or more confidence measures 122 represent an assessment or estimate of the accuracy/reliability of the speech estimate, za 1, and the hence the accuracy of the estimated RTF vector used to generate the speech estimate, za l . The second weighting block 120 generates a weighted direct speech estimate, shown in FIG. 5 by arrow 123.
[0064] FIG. 4 also illustrates processing block 124 which integrates/combines the weighted a priori speech estimate 119 and the weighted direct speech estimate 123. The combination of the weighted a priori speech estimate 119 and the weighted direct speech estimate 123 is referred to as an integrated speech estimate, za int (i.e., (40), above). The integrated speech estimate may be used for subsequent processing in the device (e.g., auditory prosthesis).
IV. MVDR with a LMA and XM signals (MVDRa el
[0065] Section III, above, illustrates an embodiment in which the integrated beamformer operates based on local microphone array (LMA) signals. As noted above, LMA signals are generated by a local microphone array (LMA) that are part of the device that performs the integrated noise reduction techniques. In the case of auditory prostheses, such as cochlear implants, the LMA is worn on the recipient.
[0066] As described further below, the integrated noise reduction techniques described herein can be extended to include external microphone (XM) signals, in addition to the LMA signals. These XM signals are generated by one or more external microphones (XMs) that are not part of the device that performs the integrated noise reduction techniques, but that can nevertheless communicate with the device (e.g., via a wireless connection). The external microphones may be any type of microphone (e.g., microphones in a wireless microphone device, microphones in a separate computing device (e.g., phone laptop, tablet, etc.), microphones in another auditory prosthesis, microphones in a conference phone system, microphones in hands-free system, etc) for which the location of the microphone(s) is unknown relative to the microphones of the LMA. In other words, as used herein, an external microphone may be any microphone that has an unknown location, which may change over time, with respect to the local microphone array.
[0067] Extending the techniques herein to the use of LMA signals and XM signals, the integrated beamformer is referred to as the MVDRaL
min wHRnn w
W
s. t wHh = 1 where h is the RTF vector ((4), above) that includes Ma components corresponding to the LMA, ha, and Me components corresponding to the XMs, he, and Rnn is the (Ma + Me) X (Ma + Me) noise correlation matrix:
Rn n —
(42)
Figure imgf000019_0003
where the upper left block is the noise correlation matrix from the LMA signals, Rnane, is the noise cross-correlation between the LMA signals and the XM signals and Rnene is the noise correlation of the XM signals. Similar to (23), the solution to (41) is given by:
Figure imgf000019_0001
with the speech estimate, z = wHy. Since, as noted above, the XMs have an unknown location, which may change over time, with respect to the local microphone array, generally no a priori assumptions can be made about the location of the XMs. Consequently, there are two potential approaches that can be taken in order to find /?, namely: (i) only the missing component of the RTF vector corresponding to that of the XM signals needs to be estimated, while the a priori assumed RTF vector for the LMA signals is preserved; or (ii) the entire RTF vector is estimated for the LMA signals and the XM signals. In sections, IV-A and IV-B strategies for both approaches are briefly described.
A. Using a partial a priori assumed RTF vector and partial estimated RTF vector
[0068] As previously mentioned, one option for the definition of h for the MVDRa,e is such that the a priori RTF vector for the LMA signals, ha, is preserved and only the RTF vector for the XM signals is estimated. Such an RTF will therefore be defined as follows:
Figure imgf000019_0002
[0069] It should be noted that although h partially contains an estimated RTF vector, this is done with respect to the a priori assumptions set by ha, and hence the notation for h is kept to be that of an a priori RTF vector (this is further elaborated upon in section IV-E). A method to compute he in the case of one XM using the cross-correlation between the external microphone and a speech reference provided by (26) using a GEVD is outlined below [0070] As in (28) a rank-l matrix approximation problem can be formulated to estimate an entire RTF vector for a given set of microphone signals such that:
Figure imgf000020_0001
where Rx rl is a rank-l approximation to Rxx (recall (8)). The a priori assumed RTF vector for the LMA signals can also be included for the definition of Rx rl and hence is given by:
[ha he ] (46)
Figure imgf000020_0002
[0071] As opposed to using the raw signal correlation matrices, the estimation problem of (45) can be equivalently formulated in the pre-whitened-transformed domain. In Appendix C, it is shown that the estimated RTF vector could be found from a GEVD on the matrix pencil
0TRyy J’JTRnnx J), where the selection matrix, J = [0(Me+1)x(Ma-1) \ Me+1] . As a result of the pre-whitening (Rnn = 1 Ma+Me ), this GEVD can consequently be computed from the EVD of JTRyy J, which is a lower order correlation matrix, of dimensions (Me + 1) X (Me + 1) that could be constructed from the last (Me + 1) elements of the pre-whitened-transformed signals, namely that in relation to the last element of the LMA - ya Ma, ar|d those in relation to the XM signals - ye. The resulting RTF vector for the XM signals is then defined from the corresponding principal (first in this case) eigenvector, vmax:
Figure imgf000020_0003
where the selection matrix, J
Figure imgf000020_0004
[0072] Finally, this estimate is then used to compute the corresponding MVDRa c filter with an a priori assumed RTF vector and a partially estimated RTF vector as:
Figure imgf000020_0005
where h as defined in (53) can be equivalently represented as:
Figure imgf000020_0006
[0073] As was done in section III, this filter can also be reformulated in the pre-whitened- transformed domain. Leaving the derivations once again to Appendix C, the corresponding speech estimate was then found to be:
(50)
Figure imgf000021_0002
where ¾§-y-vmax can be considered as a pre-whitened-transformed filter, which can be used to
||ha||
directly filter the last (Me + 1) elements of the pre-whitened-transformed signals, i.e. ya Ma ar|d ye-
[0074] More specifically, FIG. 5 is a block diagram illustrating a transformation block 502 representing the first transformation of section II-B, in which the LMA signals pass through a blocking matrix 504 and a matched filter 506, analogous to the first stage of a GSC. The XM signals are unaltered. The pre-whitening block 508 represents the pre-whitening operation. The output of the pre-whitening block 508 is signals in the pre-whitened-transformed domain, referred to as pre-whitened-transformed signals 509.
[0075] Also shown in FIG. 5 is filter 530 (i.e., (50), above), which uses the whitened- transformed signals 509 to generate an a priori speech estimate, zc. As such, the a priori speech estimate, zc, is a speech estimate using a partial a priori assumed RTF vector and partial estimated RTF vector (i.e., using a priori assumptions for the definition of the RTF vector for the LMA signals, while estimating only the RTF vector for the XM signals). Stated differently, the a priori speech estimate, zc, is generated from assumptions such as microphone characteristics, location and reverberant characteristics of the speech within the sound signals detected by the LMA, and based on a real-time estimate of speech within the sound signals detected by the XM, which adhere to the same assumptions used for the LMA. The a priori speech estimate z1 is an example of an a priori estimate of at least one target sound in the received sound signals.
[0076] In the case where the RTF vector for both the LMA and XM signals is to be estimated, a variation of (45) is considered:
Figure imgf000021_0001
where Rx rl is a rank-l approximation to Rxx (without any a priori information):
Figure imgf000022_0001
with qa the estimated RTF vector for the LMA signals and qe the RTF vector for the XM signals.
[0077] Once again, it will be convenient to re-frame the problem in the pre-whitened- transformed domain. From the derivations in Appendix D, the estimated RTF vector is given by: Qmax
Figure imgf000022_0002
q (53) where qmax is a generalized eigenvector of the matrix pencil {Ryy, Rnn}. which as a result of the pre-whitening (Rnn = 1 Ma+Me) corresponds to the principal (first in this case) eigenvector of Ryy, ?7q = exlTL qmax and exl = [10 ... 010 ... 0]T. The estimated RTF vector can therefore be used as an alternative to h for the MVDRa,e:
Figure imgf000022_0003
[0078] As derived in Appendix D, the corresponding speech estimate in the pre-whitened- transformed domain is given by:
Figure imgf000022_0004
where hh qmax can be considered as a pre-whitened-transformed filter, which can be used to directly filter the pre-whitened-transformed signals, y.
[0079] More specifically, FIG. 6 is a block diagram illustrating a transformation block 502 representing the first transformation of section II-B, in which the LMA signals pass through a blocking matrix 504 and a matched filter 506, analogous to the first stage of a GSC. The XM signals are unaltered. The pre-whitening block 508 represents the pre-whitening operation. The output of the pre-whitening block 508 is signals in the pre-whitened-transformed domain, referred to as pre-whitened-transformed signals 509.
[0080] Also shown in FIG. 6 is filter 532 (i.e., (55), above), which uses the whitened- transformed signals 509 to generate a direct speech estimate, zL . As such, the direct speech estimate, zc, is a speech estimate using an estimated RTF vector including both the LMA and XM signals. Stated differently, the speech estimate, zc, is generated from a real-time estimate of the speech within the sound signals detected by both the LMA and XM, which takes into consideration microphone characteristics and may contain information such as the location and some reverberant characteristics of the target sound. The speech estimate zc, is an example of a direct estimate of at least one target sound in the received sound signals.
B. Integrated Beamformer
[0081] In the case of the integrated MVDRa for the LMA signals in section III-C, two general approaches for designing the beamformer were considered: one that imposes a priori assumptions for the definition of the RTF vector in the MVDR filter, and another that involves an estimation of this RTF vector. For the MVDRa,e, two analogous approaches can also be considered: one that imposes a priori assumptions for the definition of the RTF vector for the LMA signals, while estimating only the RTF vector for the XM signals or an estimation of the entire RTF vector including both the LMA and XM signals. Although in both approaches there is an estimation; for the approach where only the RTF vector for the XM signals is estimated, it is done so in accordance with the a priori assumptions set by the LMA. Therefore, just as in the integrated MVDRa, two general approaches to designing the MVDRa,e according to either a priori assumptions or full estimation can be considered. Consequently, an integrated MVDRa e beamformer can also be derived in order to integrate the two general approaches. The resulting cost function, is:
Figure imgf000023_0001
W
where h is defined from (49) and h from (53). The solution is then:
Wint = 9rLa> b w + #estO< £)w (57) where wx and wx are given (48) and (54) respectively.
Figure imgf000023_0002
with the constants:
(60)
Figure imgf000024_0001
[0082] As in section III-C, this integrated MVDRa,e beamformer also reveals that the MVDRa c beamformer based on a priori assumptions from (48) and that which is based on estimated quantities from (54) can be combined according to the functions gpr(a, b) and gest(a >b) respectively.
[0083] This integrated beamformer can also be expressed in the pre-whitened-transformed domain as follows:
Figure imgf000024_0002
and the constants equivalently, but alternatively defined as: khh = hHh; kqq = h¾ khq = hHh; kqh = hHh (62) where h and h are given in (88) from Appendix C and (97) from Appendix D respectively.
[0084] The resulting speech estimate from this integrated beamformer is then given by:
Figure imgf000024_0003
[0085] The benefit of the pre-whitened-transformed domain is once again apparent. With such an integrated beamformer, the transformed, pre-whitened signals can be directly filtered accordingly, and then combined with the appropriate weightings as defined by the functions gpr(a, ?) and gest(a ,/?), to yield the respective speech estimate. These functions $rG(a, ?) and gest (a > b) can be tuned such as to emphasize the result from an MVDR beamformer that uses either an a priori assumed RTF vector or an estimated RTF vector. This results in a digital signal processing scheme as depicted in FIG. 7.
[0086] More specifically, FIG. 7 is a block diagram of an integrated MVDRa,e beamformer 525 in accordance with embodiments presented herein. The integrated MVDRa,e beamformer 525 comprises a plurality of processing blocks, which include transformation block 502 and pre-whitening block 508. As described above with reference to FIGs. 5 and 6, the transformation block 502 represent the first transformation of section II-B, in which the LMA signals pass through a blocking matrix 504 and a matched filter 506, while the XM signals are unaltered. The pre-whitening block 508 represents the pre-whitening operation. The output of the pre-whitening block 508 is signals in the pre-whitened-transformed domain, referred to as pre-whitened-transformed signals 509.
[0087] Also shown in FIG. 7 are two processing branches 513(1) and 513(2) that each operate based on all or part of the pre-whitened-transformed signals 509. The first processing branch 513(1) includes a filter 530 which, as described above with reference to FIG. 5, uses the whitened-transformed signals 509 to generate an a priori speech estimate, zx (i.e., an estimate of the speech in the received sound signals, based on a priori assumptions for the definition of the RTF vector for the LMA signals, while estimating only the RTF vector for the XM signals). The speech estimate z1 is an example of an a priori estimate of at least one target sound in the received sound signals.
[0088] The first branch 513(1) also comprises a first weighting block 516. The first weighting block 516 is configured to weight the speech estimate, zc, in accordance with the complex conjugate of the function gpr(a, b) (i.e., (58) and (63), above). More generally, the first weighting block 516 is configured to weight the speech estimate, zc, in accordance with a cost function controlled by a plurality of tuning parameters (e.g., (a, /?)). The tuning parameters of the cost function (e.g., gpr ( ,b)), are set based on one or more confidence measures 518 generated for the speech estimate, zc. The one or more confidence measures 518 represent an assessment or estimate of the accuracy/reliability of the speech estimate, zc, and the hence the accuracy of the partial a priori assumed RTF vector and partial estimated RTF vector used to generate the speech estimate (i.e., using a priori assumptions for the definition of the RTF vector for the LMA signals, while estimating only the RTF vector for the XM signals). The first weighting block 518 generates a weighted a priori speech estimate, shown in FIG. 5 by arrow 519.
[0089] The second branch 513(2) includes the filter 532 (i.e., (55), above), which uses the whitened-transformed signals 509 to generate a direct speech estimate, zx (i.e., a speech estimate generated using an estimated RTF vector including both the LMA and XM signals). The second branch 513(2) also comprises a second weighting block 520. The second weighting block 520 is configured to weight the direct speech estimate, zc, in accordance with the complex conjugate of the function gest
Figure imgf000025_0001
(i.e., (59) and (63), above). More generally, the second weighting block 120 is configured to weight the direct speech estimate, ¾, in accordance with a cost function controlled by a plurality of tuning parameters (e.g., ( a, b )). The tuning parameters of the cost function (e.g., gest(a > b) are set based on one or more confidence measures 522 generated for the speech estimate, z1. The one or more confidence measures 522 represent an assessment or estimate of the accuracy/reliability of the speech estimate, zc, and the hence the accuracy of the estimated RTF vector including both the LMA and XM signals. The second weighting block 520 generates a weighted direct speech estimate, shown in FIG. 5 by arrow 123.
[0090] FIG. 7 also illustrates processing block 524 which integrates/combines the weighted a priori speech estimate 519 and the weighted direct speech estimate 523. The combination of the weighted a priori speech estimate 519 and the weighted direct speech estimate 523 is referred to as an integrated speech estimate, zint (i.e., (63), above). The integrated speech estimate, zint, may be used for subsequent processing in the device (e.g., auditory prosthesis).
[0091] With this integrated beamformer for both the LMA and XMs, the decision process is now, as shown in the flowchart of FIG. 8, a two stage process 840. More specifically, the process 840 is comprised of two main decisions, referred to as decisions 842 and 844. Referring first to 842, it can be determined whether or not the XM signals are reliable (i.e., decide whether or not to use the XM signals). If the XM signals are not reliable, the system uses MVDR with LMA only (i.e., MVDRa). If the XM signals are reliable, the system uses MVDR with LMA and XMs (i.e., MVDRa e).
[0092] At 844, after determining whether or not the XM signals should be used, a decision is made as to whether or not estimated RTF vector is reliable. In other words, a decision can then be made on how much to weight the a priori assumed RTF vector and the estimated RTF vector. This decision is controlled by a and b in the same manner as for the Integrated MVDRa Beamformer from section III-C. In the case where the XM is used, the a priori assumed RTF vector consists of an a priori assumed RTF vector for the LMA signals and an estimated RTF vector for the XM signals, the estimated RTF vector is for both the LMA and XM signals.
[0093] In the second stage of the decision process, it should be noted that in order to simplify the tuning, a and b could be made inversely proportional, and can even be tuned such that gpr(<x ) and gest(afl) form a convex combination. Alternatively, if it is imposed that a— > ¥, then this preserves the a priori constraint and it is only b that remains to be tuned, which would be that of a contingency noise reduction strategy. In the case where both a— > oo and b— > oo3 this corresponds to two hard constraints imposed upon the noise minimization, and is then considered as a linearly constrained minimum variance (LCMV) beamformer . It is also noted for the case of the MVDRa where a— > oo, b = 0, that the original MVDRa with a priori constraints is achieved. Hence, the original beamformer has not been compromised and can be reverted to at anytime with this particular tuning.
[0094] A summary of the various noise reduction strategies encompassed by this integrated beamformer is summarized in FIG. 9. More specifically, FIG. 9 includes a table, referred to as Table I, which illustrates limiting cases of a , b for the various MVDR beamformers.
[0095] The integrated noise reduction techniques presented herein may be implemented in a number of devices/systems that include a local microphone array (LMA) to capture sound signals. These devices/ systems include, for example, auditory prostheses (e.g., cochlear implant, acoustic hearing aids, auditory brainstem stimulators, bone conduction devices, middle ear auditory prostheses, direct acoustic stimulators, bimodal auditory prosthesis, bilateral auditory prostheses, etc.), computing devices (e.g., mobile phones, tablet computers, etc.), conference phones, hands-free telephone systems, etc. FIGs. 10A, 10B, 11, and 12 are schematic block diagrams of example devices configured to implement the integrated noise reduction techniques presented herein. It is to be appreciated that these examples are illustrative and that, as noted, the integrated noise reduction techniques presented herein may be implemented in a number of different devices/systems.
[0096] Referring first to FIG. 10A, shown is a schematic diagram of an exemplary cochlear implant 1000 configured to implement aspects of the techniques presented herein, while FIG. 10B is a block diagram of the cochlear implant 1000. For ease of illustration, FIGs. 10A and 10B will be described together.
[0097] The cochlear implant 1000 comprises an external component 1002 and an intemal/implantable component 1004. The external component 1002 includes a sound processing unit 1012 that is directly or indirectly attached to the body of the recipient, an external coil 1006 and, generally, a magnet (not shown in FIG. 10 A) fixed relative to the external coil 1006.
[0098] The sound processing unit 1012 comprises a local microphone array (LMA) 1013, comprised of microphones 1008(1) and 1008(2), configured to receive sound input signals. In this example, the sound processing unit 1012 may also include one or more auxiliary input devices 1009, such as one or more telecoils, audio ports, data ports, cable ports, etc., and a wireless transmitter/receiver (transceiver) 1011.
[0099] The sound processing unit 1012 also includes, for example, at least one battery 1007, a radio-frequency (RF) transceiver 1021, and a processing block 1050. The processing block 1050 comprises a number of elements, including an integrated noise reduction module 1025 and a sound processor 1033. The processing block 1050 may also include other elements that, have for ease of illustration, been omitted from FIG. 10B. Each of the integrated noise reduction module 1025 and a sound processor 1033 may be formed by one or more processors (e.g., one or more Digital Signal Processors (DSPs), one or more uC cores, etc), firmware, software, etc. arranged to perform operations described herein. That is, the integrated noise reduction module 1025 and a sound processor 1033 may each be implemented as firmware elements, partially or fully implemented with digital logic gates in one or more application- specific integrated circuits (ASICs), partially or fully implemented in software, etc.
[ooioo] The integrated noise reduction module 1025 is configured to perform the integrated noise reduction techniques described elsewhere herein. For example, the integrated noise reduction module 1025 corresponds to the integrated MVDRa beamformer 125 and the MVDRa,e beamformer 525, described above. As such, in different embodiments, the integrated noise reduction module 1025 may include the processing blocks described above with reference to FIGs. 4 and 7, as well as other combinations of processing blocks configured to perform the integrated noise reduction techniques described elsewhere herein.
[ooioi] As noted above, the integrated noise reduction techniques, and thus the integrated noise reduction module 1025, generates an integrated speech estimate from sound signals received via at least the LMA 1013. Shown in FIG. 10 is at least one optional external microphone (XM) which may also be in communication with the sound processing unit 1012. If present, the XM 1017 is configured to capture sound signals and provide XM signals to the sound processing unit 1012. These XM signals may also be used to generate the integrated speech estimate. The sound processor 1033 is configured to use the integrated speech estimate (generated from one or both of the LMA signals and the XM signals) to generate stimulation signals for delivery to the recipient.
[00102] Returning to the example embodiment of FIGs. 10A and 10B, the implantable component 1004 comprises an implant body (main module) 1014, a lead region 1016, and an intra-cochlear stimulating assembly 1018, all configured to be implanted under the skin/tissue (tissue) 1005 of the recipient. The implant body 1014 generally comprises a hermetically- sealed housing 1015 in which RF interface circuitry 1024 and a stimulator unit 1020 are disposed. The implant body 1014 also includes an internal/implantable coil 1022 that is generally external to the housing 1015, but which is connected to the RF interface circuitry 1024 via a hermetic feedthrough (not shown in FIG. 10B). [00103] As noted, stimulating assembly 1018 is configured to be at least partially implanted in the recipient’s cochlea 1037. Stimulating assembly 1018 includes a plurality of longitudinally spaced intra-cochlear electrical stimulating contacts (electrodes) 1026 that collectively form a contact or electrode array 1028 for delivery of electrical stimulation (current) to the recipient’s cochlea. Stimulating assembly 1018 extends through an opening in the recipient’s cochlea (e.g., cochleostomy, the round window, etc) and has a proximal end connected to stimulator unit 1020 via lead region 1016 and a hermetic feedthrough (not shown in FIG. 10B). Lead region 1016 includes a plurality of conductors (wires) that electrically couple the electrodes 1026 to the stimulator unit 1020.
[00104] As noted, the cochlear implant 1000 includes the external coil 1006 and the implantable coil 1022. The coils 1006 and 1022 are typically wire antenna coils each comprised of multiple turns of electrically insulated single-strand or multi-strand platinum or gold wire. Generally, a magnet is fixed relative to each of the external coil 1006 and the implantable coil 1022. The magnets fixed relative to the external coil 1006 and the implantable coil 1022 facilitate the operational alignment of the external coil with the implantable coil. This operational alignment of the coils 1006 and 1022 enables the external component 1002 to transmit data, as well as possibly power, to the implantable component 1004 via a closely-coupled wireless link formed between the external coil 1006 with the implantable coil 1022. In certain examples, the closely- coupled wireless link is a radio frequency (RF) link. However, various other types of energy transfer, such as infrared (IR), electromagnetic, capacitive and inductive transfer, may be used to transfer the power and/or data from an external component to an implantable component and, as such, FIG. 10B illustrates only one example arrangement.
[00105] As noted above, the integrated noise reduction module 1025 is configured to generate an integrated speech estimate, and the sound processor 1033 is configured to use the integrated speech estimate to generate stimulation signals for delivery to the recipient. More specifically, the sound processor 1033 (e.g., one or more processing elements implementing firmware, software, etc) is configured to use the integrated speech estimate to generate stimulation control signals 1036 that represent electrical stimulation for delivery to the recipient. In the embodiment of FIG. 10B, the stimulation control signals 1036 are provided to the RF transceiver 1021, which transcutaneously transfers the stimulation control signals 1036 (e.g., in an encoded manner) to the implantable component 1004 via external coil 1006 and implantable coil 1022. That is, the stimulation control signals 1036 are received at the RF interface circuitry 1024 via implantable coil 1022 and provided to the stimulator unit 1020. The stimulator unit 1020 is configured to utilize the stimulation control signals 1036 to generate electrical stimulation signals (e.g., current signals) for delivery to the recipient’s cochlea via one or more stimulating contacts 1026. In this way, cochlear implant 1000 electrically stimulates the recipient’s auditory nerve cells, bypassing absent or defective hair cells that normally transduce acoustic vibrations into neural activity, in a manner that causes the recipient to perceive one or more components of the input audio signals.
[00106] FIGs. 10A and 10B illustrate an arrangement in which the cochlear implant 1000 includes an external component. However, it is to be appreciated that embodiments of the present invention may be implemented in cochlear implants having alternative arrangements. For example, the techniques presented herein could also be implemented in a totally implantable or mostly implantable auditory prosthesis where components shown in sound processing unit 1012, such as processing block 1050, could instead be implanted in the recipient.
[00107] FIG. 11 is a functional block diagram of one example arrangement for a bone conduction device 1100 in accordance with embodiments presented herein. Bone conduction device 1100 is configured to be positioned at (e.g., behind) a recipient’s ear. The bone conduction device 1100 comprises a microphone array 1113, an electronics module 1170, a transducer 1171, a user interface 1172, and a power source 1173.
[00108] The local microphone array (LMA) 1113 comprises microphones 1108(1) and 1108(2) that are configured to convert received sound signals 1116 into LMA signals. Although not shown in FIG. 11, bone conduction device 1100 may also comprise other sound inputs, such as ports, telecoils, etc.
[00109] The LMA signals are provided to electronics module 1170 for further processing. In general, electronics module 1170 is configured to convert the LMA signals into one or more transducer drive signals 1180 that active transducer 1171. More specifically, electronics module 1170 includes, among other elements, a processing block 1150 and transducer drive components 1176.
[ooiio] The processing block 1174 comprises a number of elements, including an integrated noise reduction module 1125 and sound processor 1133. Each of the integrated noise reduction module 1125 and the sound processor 1133 may be formed by one or more processors (e.g., one or more Digital Signal Processors (DSPs), one or more uC cores, etc.), firmware, software, etc. arranged to perform operations described herein. That is, the integrated noise reduction module 1125 and the sound processor 1133 may each be implemented as firmware elements, partially or fully implemented with digital logic gates in one or more application-specific integrated circuits (ASICs), partially or fully in software, etc.
[ooiii] The integrated noise reduction module 1125 is configured to perform the integrated noise reduction techniques described elsewhere herein. For example, the integrated noise reduction module 1125 corresponds to the integrated MVDRa beamformer 125 and the MVDRa,e beamformer 525, described above. As such, in different embodiments, the integrated noise reduction module 1125 may include the processing blocks described above with reference to FIGs. 4 and 7, as well as other combinations of processing blocks configured to perform the integrated noise reduction techniques described elsewhere herein. Although not shown in FIG. 11 is at least one optional external microphone (XM) may be in communication with the bone conduction device 1100. If present, the XM is configured to capture sound signals and provide XM signals to the conduction device 1100 for processing by the integrated noise reduction module 1125 (i.e., the XM signals may also be used to generate the integrated speech estimate).
[00112] The sound processor 1133 is configured to process the integrated speech estimate (generated from one or both of the LMA signals and the XM signals) for use by the transducer drive components 1176. The transducer drive components 1176 generate transducer drive signal(s) 1180 which are provided to the transducer 1171. The transducer 1171 illustrates an example of a stimulation unit that receives the transducer drive signal(s) 1180 and generates vibrations for delivery to the skull of the recipient via a transcutaneous or percutaneous anchor system (not shown) that is coupled to bone conduction device 1100. Delivery of the vibration causes motion of the cochlea fluid in the recipient’s contralateral functional ear, thereby activating the hair cells in the functional ear.
[00113] FIG. 11 also illustrates the power source 1173 that provides electrical power to one or more components of bone conduction device 1300. Power source 1173 may comprise, for example, one or more batteries. For ease of illustration, power source 1173 has been shown connected only to user interface 1172 and electronics module 1170. However, it should be appreciated that power source 1173 may be used to supply power to any electrically powered circuits/components of bone conduction device 1100.
[00114] User interface 1172 allows the recipient to interact with bone conduction device 1100. For example, user interface 1172 may allow the recipient to adjust the volume, alter the speech processing strategies, power on/off the device, etc. Although not shown in FIG. 11, bone conduction device 1100 may further include an external interface that may be used to connect electronics module 1170 to an external device, such as a fitting system.
[00115] FIG. 12 is a block diagram of an arrangement of a mobile computing device 1200, such as a smartphone, configured to be implemented the integrated noise reduction techniques presented herein. It is to be appreciated that FIG.1 2 is merely illustrative.
[00116] Mobile computing device 1200 first comprises an antenna 1236 and a telecommunications interface 1238 that are configured for communication on a telecommunications network. The telecommunications network over which the radio antenna 1236 and the radio interface 1238 communicate may be, for example, a Global System for Mobile Communications (GSM) network, code division multiple access (CDMA) network, time division multiple access (TDMA), or other kinds of networks.
[00117] The mobile computing device 1200 also includes a wireless local area network interface 1240 and a short-range wireless interface/transceiver 1242 (e.g., an infrared (IR) or Bluetooth® transceiver). Bluetooth® is a registered trademark owned by the Bluetooth® SIG. The wireless local area network interface 1240 allows the mobile computing device 1200 to connect to the Internet, while the short-range wireless transceiver 1242 enables the external device 1206 to wirelessly communicate (i.e., directly receive and transmit data to/from another device via a wireless connection), such as over a 2.4 Gigahertz (GHz) link. It is to be appreciated that that any other interfaces now known or later developed including, but not limited to, Institute of Electrical and Electronics Engineers (IEEE) 802.11, IEEE 802.16 (WiMAX), fixed line, Long Term Evolution (LTE), etc., may also or alternatively form part of the mobile computing device 1200.
[00118] In the example of FIG. 12, mobile computing device 1200 also comprises an audio port 1244, a local microphone array (LMA) 1213, a speaker 1248, a display screen 1258, a subscriber identity module or subscriber identification module (SIM) card 1252, a battery 1254, a user interface 1256, one or more processors 1250, and a memory 1260. The LMA 1213 includes microphones 1208(1) and 1208(2). Stored in memory 1260 is integrated noise reduction logic 1225 and sound processing logic 1233.
[00119] The display screen 1258 is an output device, such as a liquid crystal display (LCD), for presentation of visual information to the cochlear implant recipient. The user interface 1256 may take many different forms and may include, for example, a keypad, keyboard, mouse, touchscreen, display screen, etc. Memory 1260 may comprise any one or more of read only memory (ROM), random access memory (RAM), magnetic disk storage media devices, optical storage media devices, flash memory devices, electrical, optical, or other physical/tangible memory storage devices. The one or more processors 1258 are, for example, microprocessors or microcontrollers that execute instructions for the integrated noise reduction logic 1225 and sound processing logic 1233.
[00120] When executed by the one or more processors 1250, the integrated noise reduction logic 1225 is configured to perform the integrated noise reduction techniques described elsewhere herein. For example, the integrated noise reduction logic 1225 corresponds to the integrated MVDRa beamformer 125 and the MVDRa,e beamformer 525, described above. As such, in different embodiments, the integrated noise logic 1225 may include software forming the processing blocks described above with reference to FIGs. 4 and 7, as well as other combinations of processing blocks configured to perform the integrated noise reduction techniques described elsewhere herein to generate an integrated noise estimate. When executed by the one or more processors 1250, the sound processing logic 1233 is configured to perform sound processing operations using the integrated noise estimate.
[00121] FIG. 13 is a flowchart of a method 1390 performed/executed by a device comprising at least a local microphone array (LMA), in accordance with embodiments presented herein. Method 1390 begins at 1392 where sound signals are received with at least the local microphone array of the device. The received sound signals comprise/include at least one target sound.
[00122] At 1394, an a priori estimate of the at least one target sound in the received sound signals is generated, wherein the a priori estimate is based at least on a predetermined location of a source of the at least one target sound. At 1396, a direct estimate of the at least one target sound in the received sound signals is generated, wherein the direct estimate is based at least on a real-time estimate of a location of a source of the at least one target sound. At 1398, a weighted combination of the a priori estimate and the direct estimate is generated, where the weighted combination is an integrated estimate of the target sound. Subsequent sound processing operations may be performed in the device using the integrated estimate of the target sound.
[00123] In certain embodiments, the a priori estimate of the at least one target sound is generated using only an a priori relative transfer function (RTF) vector generated from the received sound signals. In certain embodiments, the direct estimate of the at least one target sound is generated using only an estimated relative transfer function (RTF) vector for the received sound signals.
[00124] In certain embodiments, the weighted combination of the a priori estimate and the direct estimate is generated by weighting the a priori estimate in accordance with a first cost function controlled by a first set of tuning parameters to generate a weighted a priori estimate; and weighting the direct estimate in accordance with a second cost function controlled by a second set of tuning parameters to generate a weighted direct estimate. The weighted direct estimate with the weighted a priori estimate are then mixed with one another. The first set of tuning parameters may be set based on one or more confidence measures associated with the a priori estimate of the of the at least one target sound, wherein the one or more confidence measures represent an estimate of a reliability of the a priori estimate. The second set of tuning parameters may be set based on one or more confidence measures associated with the direct estimate of the of the at least one target sound, wherein the one or more confidence measures represent an estimate of a reliability of the direct estimate.
[00125] As detailed above, presented herein are integrated noise reduction techniques, sometimes referred to as an integrated beamformer (e.g., an integrated MVDRa beamformer or an integrated MVDRa,e beamformer). In general, the integrated noise reduction techniques combine the use of an a priori (i.e., predetermined, assumed, or pre-defmed) location of a target sound source with a real-time estimated location of the sound source.
[00126] It is to be appreciated that the above described embodiments are not mutually exclusive and that the various embodiments can be combined in various manners and arrangements.
[00127] The invention described and claimed herein is not to be limited in scope by the specific preferred embodiments herein disclosed, since these embodiments are intended as illustrations, and not limitations, of several aspects of the invention. Any equivalent embodiments are intended to be within the scope of this invention. Indeed, various modifications of the invention in addition to those shown and described herein will become apparent to those skilled in the art from the foregoing description. Such modifications are also intended to fall within the scope of the appended claims.
Figure imgf000035_0001
I. APPENDIX A - MVDRa WITH A PRIORI ASSUMED RTF VECTOR
A pre-whitened-transformed version of the a priori assumed RTF vector can be considered where:
Figure imgf000035_0002
where 1ML is the bottom-right element in La. Using the definition from (16), i.e , R 1 = (TaLaLa T^ )- 1
Figure imgf000035_0003
the MVDRa filter of (25) can then be re-written as:
wa = TaLa-¾ (65) where
Figure imgf000035_0004
Substitution of (65) into (26) yields the speech estimate as:
Figure imgf000035_0005
II. APPENDIX B - MVDRa WITH ESTIMATED RTF VECTOR
AS opposed to using the raw signal correlation matrices, the estimation problem of (28) can be equivalently formulated first in the transformed domain since the Frobenius norm is invariant under a unitary transformation , therefore:
Figure imgf000035_0006
Furthermore, it is argued in that spatial pre-whitening should also be included in the optimisation problem. Consequently, the estimation problem can be re-framed in the pre-whitened-transformed domain as follows:
Figure imgf000035_0007
where Ryay = La 1 Ta Ryaya TaLa H, and Rn n = La 1Ta Rnar.a TaLa L = lMs . The solution then follows from the GEVD on the matrix pencil {R^ , RnaJla}, and hence reduces to an EVD
Figure imgf000035_0008
where P is a unitary matrix of eigenvectors and A is a diagonal matrix with the associated eigenvalues in descending order. The estimated RTF vector is then defined using the principal (first in this case) eigenvector, pmax:
Figure imgf000035_0009
where the scaling hr = ea] TaLapr and the M X 1 vector eai ~ [1 0 . . 011 .
Tlris estimated RTF vector can now be used as an alternative to ha for the MVDRa defined in (25), and is given by:
Figure imgf000035_0010
This filter based on estimated quantities can also be reformulated in the pre-whitened- transformed domain. Starting with the definition of the pre-whitened- transformed version of ha:
Figure imgf000036_0001
Hence (72) becomes:
Wa = TaL- ¾ (74) where
Figure imgf000036_0002
Substitution of (74) into (32) yields the speech estimate as:
Figure imgf000036_0003
III. APPENDIX C - MVDRa e WITH PARTIAL A PRIORI ASSUMED RTF VECTOR AND PARTIAL ESTIMATED RTF VECTOR
Following the procedure as in (68), the transformation is firstly applied, also including the penalty term:
Figure imgf000036_0004
after which, the pre-whitening operation can also be included in the optimisation problem:
. min. | j (Ryy - ! 1 2
B„„) - L - ! TH (Fc,Gi ~ a | [h h? ] ) TL - H
I (78) where Ryv = L 1T ri RyyTL ri and Rnn = L-1Tri RnnATL~ i‘ = I(M.H-M.)· Expansion of (78) then results in:
Figure imgf000036_0005
where the block dimensions are such thal KA is an (ikfa— 1) x (Ma— 1) matrix, KB an (Ma -- 1) x M I 1) matrix a (Me -f- 1) x (Ma --- 1) matrix and Kx rl and KX are ( Me -f 1) x ( e -;- 1) matrices realised as:
Figure imgf000036_0006
where Bc,p = L ~ 1 T b Rx ,„ i T L ~ H and J = [ 0(¾+1! x (¾-:1) I ,M +] ) is a selection matrix. It is then evident that kc ; can essentially be constructed from the last (Me + 1) elements of the pre-whitened -transformed signals, namely that in relation to the last element of the LMA - ya M and those in relation to the XM signals - equivalently:
Figure imgf000036_0007
and similarly for the second term of K , . It follows that (79) then reduces to the following (ikfe -;- 1) x (Me 4- 1) matrix approximation problem:
Figure imgf000036_0008
The solution then follows from the GEVD on the matrix pencil
Figure imgf000036_0009
J, 31 Rnn Jl and hence reduces to an EVD of
J ¾y J:
J TRyy J = VrV'¾r (84) where V is a (Me + 1) x (Me 4- 1) unitary matrix of eigenvectors and G is a diagonal matrix with the associated eigenvalues in descending order. The estimated RTF vector for the XM signals is then defined from the corresponding principal (first in this case) eigenvector, vma :
where the selection matrix,
Figure imgf000037_0001
Finally, this estimate is then used to compute the corresponding MVDR3 e filter with an a priori assumed RTF vector and a partially estimated RTF vector, along with the penalty term as:
Figure imgf000037_0002
where h as defined in (44) can be equivalently represented as:
i jhal l
h
Figure imgf000037_0003
hi* vi
This filter can also be realised in the pre- whitened- transformed domain. The pre-whitened-transfbnned version of h can firstly be considered where:
Figure imgf000037_0004
Therefore, (86) can be re-written as:
W TL W (89) where:
Figure imgf000037_0005
Therefore, the corresponding speech estimate will be:
Figure imgf000037_0006
IV. APPENDIX D - MVDR3 e WITH ESTIMATED RTF VECTOR
Once again, it will be convenient to re- frame the problem in the pre-wbitened-transformed domain similarly to (78):
Figure imgf000037_0007
In this case however, the problem cannot be reduced to a lower order as tire entire RTF vector is being estimated. Hence the solution follows from
Figure imgf000037_0008
Figure imgf000037_0009
where Q is a (Ma+ Me) x (Ma + Me) unitary matrix of eigenvectors and å is a diagonal matrix with the associated eigenvalues in descending order. The estimated RTF vector is then given by the principal (first in this case) eigenvector, qmax:
Figure imgf000037_0010
Use estimated RTF vector can therefore be used as an alternative to h for the MVDRa e:
(96)
Figure imgf000038_0001
This filter based on estimated quantities can also be reformulated in the pre-whitened-transformed domain. Starting with the definition for the pre-whitened-transformed version of this estimated RTF:
Figure imgf000038_0002
Hence ( 96) becomes:
w = TL w (98) when
Figure imgf000038_0003
The corresponding speech estimate using the estimated RTF vector is therefore:
Figure imgf000038_0004

Claims

CLAIMS What is claimed is:
1. A method, comprising:
receiving sound signals with at least a local microphone array of a device, wherein the sound signals comprise at least one target sound;
generating an a priori estimate of the at least one target sound in the received sound signals, wherein the a priori estimate is based at least on a predetermined location of a source of the at least one target sound;
generating a direct estimate of the at least one target sound in the received sound signals, wherein the direct estimate is based at least on a real-time estimate of a location of a source of the at least one target sound; and
generating a weighted combination of the a priori estimate and the direct estimate, wherein the weighted combination is an integrated estimate of the target sound.
2. The method of claim 1, wherein generating the a priori estimate of the at least one target sound in the received sound signal, comprises:
generating the a priori estimate using only an a priori relative transfer function (RTF) vector generated from the received sound signals.
3. The method of claim 1, wherein generating the direct estimate of the at least one target sound in the received sound signals, comprises:
generating the direct estimate using only an estimated relative transfer function (RTF) vector for the received sound signals.
4. The method of claim 1, wherein generating the weighted combination of the a priori estimate of the at least one target sound and the direct estimate of the at least one target sound, comprises:
weighting the a priori estimate in accordance with a first cost function controlled by a first set of tuning parameters to generate a weighted a priori estimate;
weighting the direct estimate in accordance with a second cost function controlled by a second set of tuning parameters to generate a weighted direct estimate; and
mixing the weighted direct estimate with the weighted a priori estimate.
5. The method of claim 4, further comprising:
setting the first set of tuning parameters based on one or more confidence measures associated with the a priori estimate of the of the at least one target sound, wherein the one or more confidence measures represent an estimate of a reliability of the a priori estimate.
6. The method of claim 4, further comprising:
setting the second set of tuning parameters based on one or more confidence measures associated with the direct estimate of the of the at least one target sound, wherein the one or more confidence measures represent an estimate of a reliability of the direct estimate.
7. The method of claim 1, wherein generating the a priori estimate of the at least one target sound in the received sound signal, comprises:
generating the a priori estimate based at least on the predetermined location of a source of the at least one target sound, one or more assumptions regarding characteristics of the local microphone array, and one or more assumptions regarding reverberant
characteristics of the at least one target sound.
8. The method of claim 1, wherein generating the direct estimate of the at least one target sound in the received sound signals, comprises:
generating the direct estimate based at least on a real-time estimate of a location of a source of the at least one target sound, estimated characteristics of the local microphone array, and estimated reverberant characteristics of the at least one target sound.
9. The method of claim 1, further comprising:
performing subsequent sound processing operations in the device using the integrated estimate of the target sound.
10. The method of claim 1, wherein receiving the sound signals with at least a local microphone array of a device, comprises:
receiving the first portion of the sound signals with the local microphone array of the device; and
receiving a second portion of the sound signals with at least one external microphone.
11. The method of claim 10, wherein generating the a priori estimate of the at least one target sound in the received sound signals, comprises:
generating the a priori estimate using both the first portion of the sound signals and the second portion of the sound signals in accordance with at least the predetermined location of the source of the at least one target sound.
12. The method of claim 10, wherein generating the direct estimate of the at least one target sound in the received sound signals, comprises:
generating the direct estimate using both the first portion of the sound signals and the second portion of the sound signals in accordance with at least the real-time estimate of the location of the source of the at least one target sound.
13. A device, comprising:
a local microphone array configured to receive sound signals, wherein the sound signals comprise at least one target sound; and
one or more processors configured to:
generate an a priori estimate of the at least one target sound in the received sound signals using only an a priori relative transfer function (RTF) vector generated from the received sound signals,
generate a direct estimate of the at least one target sound in the received sound signals using only an a priori relative transfer function (RTF) vector generated from the received sound signals, and
generate a weighted combination of the a priori estimate and the direct estimate, wherein the weighted combination is an integrated estimate of the target sound.
14. The device of claim 13, wherein to generate the weighted combination of the a priori estimate of the at least one target sound and the direct estimate of the at least one target sound, the one or more processors are configured to:
weight the a priori estimate in accordance with a first cost function controlled by a first set of tuning parameters to generate a weighted a priori estimate; weight the direct estimate in accordance with a second cost function controlled by a second set of tuning parameters to generate a weighted direct estimate; and
mix the weighted direct estimate with the weighted a priori estimate.
15. The device of claim 14, wherein the one or more processors are configured to:
set the first set of tuning parameters based on one or more confidence measures associated with the a priori estimate of the of the at least one target sound, wherein the one or more confidence measures represent an estimate of a reliability of the a priori estimate.
16. The device of claim 14, wherein the one or more processors are configured to:
set the second set of tuning parameters based on one or more confidence measures associated with the direct estimate of the of the at least one target sound, wherein the one or more confidence measures represent an estimate of a reliability of the direct estimate.
17. The device of claim 13, wherein to generate the a priori estimate of the at least one target sound in the received sound signal, the one or more processors are configured to:
generate the a priori estimate based at least on the predetermined location of a source of the at least one target sound, one or more assumptions regarding characteristics of the local microphone array, and one or more assumptions regarding reverberant characteristics of the at least one target sound.
18. The device of claim 13, wherein to generate the direct estimate of the at least one target sound in the received sound signals, the one or more processors are configured to: generate the direct estimate based at least on a real-time estimate of a location of a source of the at least one target sound, estimated characteristics of the local microphone array, and estimated reverberant characteristics of the at least one target sound.
19. The device of claim 13, wherein the one or more processors are configured to:
perform subsequent sound processing operations in the device using the integrated estimate of the target sound.
20. A system including the device of claim 13, wherein the local microphone array is configured to receive a first portion of the sound signals, and wherein the system comprises: at least one external microphone configured to receive a second portion of the sound signals.
21. The system of claim 20, wherein to generate the a priori estimate of the at least one target sound in the received sound signals, the one or more processors are configured to: generate the a priori estimate using both the first portion of the sound signals and the second portion of the sound signals in accordance with at least the predetermined location of the source of the at least one target sound.
22. The system of claim 20, wherein to generate the direct estimate of the at least one target sound in the received sound signals, the one or more processors are configured to: generate the direct estimate using both the first portion of the sound signals and the second portion of the sound signals in accordance with at least the real-time estimate of the location of the source of the at least one target sound.
PCT/IB2019/057011 2018-08-27 2019-08-20 Integrated noise reduction Ceased WO2020044166A1 (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
US17/261,778 US11943590B2 (en) 2018-08-27 2019-08-20 Integrated noise reduction

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
US201862723157P 2018-08-27 2018-08-27
US62/723,157 2018-08-27

Publications (1)

Publication Number Publication Date
WO2020044166A1 true WO2020044166A1 (en) 2020-03-05

Family

ID=69645124

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/IB2019/057011 Ceased WO2020044166A1 (en) 2018-08-27 2019-08-20 Integrated noise reduction

Country Status (2)

Country Link
US (1) US11943590B2 (en)
WO (1) WO2020044166A1 (en)

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN115856813A (en) * 2022-11-17 2023-03-28 中国人民解放军海军航空大学 Radar target sidelobe suppression method based on APC and IARFT cascade processing

Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20040175006A1 (en) * 2003-03-06 2004-09-09 Samsung Electronics Co., Ltd. Microphone array, method and apparatus for forming constant directivity beams using the same, and method and apparatus for estimating acoustic source direction using the same
US20070003071A1 (en) * 1997-08-14 2007-01-04 Alon Slapak Active noise control system and method
US20090202091A1 (en) * 2008-02-07 2009-08-13 Oticon A/S Method of estimating weighting function of audio signals in a hearing aid
US20110103626A1 (en) * 2006-06-23 2011-05-05 Gn Resound A/S Hearing Instrument with Adaptive Directional Signal Processing
US20120239385A1 (en) * 2011-03-14 2012-09-20 Hersbach Adam A Sound processing based on a confidence measure

Patent Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20070003071A1 (en) * 1997-08-14 2007-01-04 Alon Slapak Active noise control system and method
US20040175006A1 (en) * 2003-03-06 2004-09-09 Samsung Electronics Co., Ltd. Microphone array, method and apparatus for forming constant directivity beams using the same, and method and apparatus for estimating acoustic source direction using the same
US20110103626A1 (en) * 2006-06-23 2011-05-05 Gn Resound A/S Hearing Instrument with Adaptive Directional Signal Processing
US20090202091A1 (en) * 2008-02-07 2009-08-13 Oticon A/S Method of estimating weighting function of audio signals in a hearing aid
US20120239385A1 (en) * 2011-03-14 2012-09-20 Hersbach Adam A Sound processing based on a confidence measure

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN115856813A (en) * 2022-11-17 2023-03-28 中国人民解放军海军航空大学 Radar target sidelobe suppression method based on APC and IARFT cascade processing

Also Published As

Publication number Publication date
US20210306743A1 (en) 2021-09-30
US11943590B2 (en) 2024-03-26

Similar Documents

Publication Publication Date Title
US11917370B2 (en) Hearing device and a hearing system comprising a multitude of adaptive two channel beamformers
EP3694229B1 (en) A hearing device comprising a noise reduction system
US11503414B2 (en) Hearing device comprising a speech presence probability estimator
US12574689B2 (en) Hearing device comprising a noise reduction system
US7657038B2 (en) Method and device for noise reduction
US10728677B2 (en) Hearing device and a binaural hearing system comprising a binaural noise reduction system
EP3471440B1 (en) A hearing device comprising a speech intelligibilty estimator for influencing a processing algorithm
US20240260887A1 (en) System for capturing electrooculography signals
EP3236672B1 (en) A hearing device comprising a beamformer filtering unit
CN104469643B (en) Hearing aid device comprising an input transducer system
EP3185589B1 (en) A hearing device comprising a microphone control system
EP2993915B1 (en) A hearing device comprising a directional system
CN107071674B (en) Hearing device and hearing system configured to locate a sound source
EP2876900A1 (en) Spatial filter bank for hearing system
EP3404935A2 (en) A hearing aid for placement at an ear of a user
US11943590B2 (en) Integrated noise reduction
US20240015449A1 (en) Magnified binaural cues in a binaural hearing system
US20210243533A1 (en) Combinatory directional processing of sound signals
US20250203299A1 (en) Multi-band channel coordination
EP3972290A1 (en) Battery contacting system in a hearing device

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 19856116

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 19856116

Country of ref document: EP

Kind code of ref document: A1