EP4681198A1 - Far-field noise reduction via spatial filtering using a microphone array - Google Patents
Far-field noise reduction via spatial filtering using a microphone arrayInfo
- Publication number
- EP4681198A1 EP4681198A1 EP24775389.0A EP24775389A EP4681198A1 EP 4681198 A1 EP4681198 A1 EP 4681198A1 EP 24775389 A EP24775389 A EP 24775389A EP 4681198 A1 EP4681198 A1 EP 4681198A1
- Authority
- EP
- European Patent Office
- Prior art keywords
- field
- audio
- frequency band
- audio signal
- audio frequency
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S7/00—Indicating arrangements; Control arrangements, e.g. balance control
- H04S7/30—Control circuits for electronic adaptation of the sound field
- H04S7/302—Electronic adaptation of stereophonic sound system to listener position or orientation
- H04S7/303—Tracking of listener position or orientation
- H04S7/304—For headphones
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L21/00—Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
- G10L21/02—Speech enhancement, e.g. noise reduction or echo cancellation
- G10L21/0208—Noise filtering
- G10L21/0216—Noise filtering characterised by the method used for estimating noise
- G10L21/0232—Processing in the frequency domain
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04R—LOUDSPEAKERS, MICROPHONES, GRAMOPHONE PICK-UPS OR LIKE ACOUSTIC ELECTROMECHANICAL TRANSDUCERS; ELECTRIC HEARING AIDS; PUBLIC ADDRESS SYSTEMS
- H04R3/00—Circuits for transducers
- H04R3/005—Circuits for transducers for combining the signals of two or more microphones
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L21/00—Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
- G10L21/02—Speech enhancement, e.g. noise reduction or echo cancellation
- G10L21/0208—Noise filtering
- G10L21/0216—Noise filtering characterised by the method used for estimating noise
- G10L2021/02161—Number of inputs available containing the signal or the noise to be suppressed
- G10L2021/02166—Microphone arrays; Beamforming
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04R—LOUDSPEAKERS, MICROPHONES, GRAMOPHONE PICK-UPS OR LIKE ACOUSTIC ELECTROMECHANICAL TRANSDUCERS; ELECTRIC HEARING AIDS; PUBLIC ADDRESS SYSTEMS
- H04R2499/00—Aspects covered by H04R or H04S not otherwise provided for in their subgroups
- H04R2499/10—General applications
- H04R2499/15—Transducers incorporated in visual displaying devices, e.g. televisions, computer displays, laptops
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S2400/00—Details of stereophonic systems covered by H04S but not provided for in its groups
- H04S2400/15—Aspects of sound capture and related signal processing for recording or reproduction
Definitions
- the present invention generally relates to systems and methods for reducing far-field noise through spatial filtering using a microphone array, which may be utilized in spatialized audio systems within virtual reality, augmented reality and/or mixed reality systems.
- AR mixed reality systems
- VR virtual reality
- AR augmented reality
- a virtual reality, or “VR” typically involves presentation of digital or virtual image information without transparency to actual real-world visual input.
- An augmented reality, or “AR”, scenario typically involves presentation of digital or virtual image information as an augmentation to visualization of the actual world around the user (i.e., transparency to other actual real-world visual input).
- AR scenarios involve presentation of digital or virtual image information with transparency to other actual real-world visual input.
- extended reality and XR are used to refer collectively to any of VR, AR and/or MR.
- AR means either, or both, AR and MR.
- VR and AR systems typically employ head-worn displays (or helmet-mounted displays, or smart glasses) that are at least loosely coupled to a user’s head, and thus move when the end user’s head moves. If the end user’s head motions are detected by the display system, the data being displayed can be updated to take into account the change in head pose (i.e., the orientation and/or location of the user’s head).
- head-worn displays or helmet-mounted displays, or smart glasses
- the virtual object can be rendered for each viewpoint (corresponding to a position and/or orientation of the head-worn display device), giving the user the perception that they are walking around an object that occupies real space.
- the head-worn display device is used to present multiple virtual objects at different depths, measurements of head pose can be used to render the scene to match the user’s dynamically changing head pose and provide an increased sense of immersion.
- Head-worn displays that enable AR (i.e., the concurrent viewing of virtual and real objects) can have several different types of configurations.
- a camera captures elements of a real scene
- a computing system superimposes virtual elements onto the captured real scene
- a non-transparent display presents the composite image to the eyes.
- Another configuration is often referred to as an “optical see- through” display, in which the end user can see through transparent (or semi-transparent) elements in the display system to directly view the light from real objects in the environment.
- the transparent element often referred to as a “combiner,” superimposes light from the display over the end user’s view of the real world.
- a camera may be mounted onto the head-worn display device to capture images or videos of the scene being viewed by the user.
- XR systems typically also include a microphone arrangement including one or more microphones for sensing audio (i.e., sound), such as the speech (i.e., voice) of the user, and ambient/environment sound in the real world surroundings of the user, and generating audio signals corresponding to the audio.
- audio i.e., sound
- the user may wish to communicate using various speech based transmission protocols (e.g., group chat, IP based speech communication platforms, etc.) in real world environments where the ambient noise level may be relatively high.
- the user may be on a factory floor, in other commercial environments, in the vicinity of children playing, in the vicinity of media such as television or music playing in the background, etc.
- the relatively high level of ambient noise hinders effective speech communication because the person(s) to whom the user is communicating cannot clearly hear the user’s speech through the XR system, and/or may negatively impact speech recognition performance by speech recognition systems such as systems configured to recognize and process voice commands.
- the present disclosure is directed to systems and methods for reducing unwanted noise from audio signals generated by a microphone array.
- the systems and methods disclosed herein use an innovative spatial filtering approach which filters far-field ambient noise from the near- field audio and reduces the far-field noise in the near-field audio signal thereby improving and isolating the near-field audio from the far-field interference.
- the systems and methods are useful for effectively filtering speech (i.e., near- field audio) from ambient noise (i.e., far-field noise).
- speech i.e., near- field audio
- ambient noise i.e., far-field noise
- such systems and methods may be implemented on XR systems to improve and isolate a user’s speech using a microphone array on a headset of the XR system.
- the systems and methods disclosed herein may be implemented and carried out on any suitable hardware system having a microphone array, a computer processor and software (which may include firmware) configured to program the system to perform a process for reducing far- field noise in a near-field audio signal.
- one embodiment disclosed herein is directed to a computer-implemented method for reducing far-field noise in a near-field audio signal using a microphone array.
- the method includes acquiring a near-field audio signal from at least one primary microphone of the microphone array, and acquiring one or more reference audio signals from one or more reference microphones of a microphone array.
- a far-field signal is determined from the one or more reference audio signals.
- at least one primary microphone may include a speech/voice microphone positioned near the mouth of a user, while the one or more reference microphones are positioned further away from the user’s mouth, such as at the side of the user’s head.
- a far-field signal is determined from the one or more reference audio signals.
- the far-field signal may be calculated by combining the reference audio signals from the one or more reference microphones, such as calculating a difference between the one or more reference microphones.
- the near-field audio signal and the far-field signal are partitioned into a plurality of audio frequency bands.
- the near-field audio signal and the far-field audio signal may be partitioned using a weighted overlap-add (WOLA) analysis.
- the number of audio frequency bands may be any suitable number, such as from 20-60 bands, 40-50 bands, greater than 20 bands, greater than 40 bands, or greater than 50 bands.
- a near-field energy for the near-field audio signal in each audio frequency band and a far-field energy for the far-field audio signal in each audio frequency band is calculated.
- the resultant near-field energy and far-field energy in each audio frequency band are used to calculate an energy ratio of the near-field energy to the far-field energy for each respective audio frequency band.
- the ratio of the near-field energy in a first audio frequency band to the far- field energy in the same first audio frequency band is calculated
- the ratio of the near-field energy in a second audio frequency band to the far-field energy in the same second audio frequency band is calculated, and so-on for each of the audio frequency bands.
- the method determines whether the energy ratio for each audio frequency band is below a predetermined threshold for the respective audio frequency band.
- the predetermined threshold may be frequency band dependent, such that each audio frequency band has its own predetermined threshold, which may be different for each audio frequency band.
- the predetermined threshold may be the same for all of the audio frequency bands.
- a respective time-varying masking gain is calculated based upon the energy ratio for the respective audio frequency band.
- the time-varying masking gain for each audio frequency band may be proportional to the amount the respective energy ratio is below the respective predetermined threshold. In such case, if the energy ratio is just slightly below (i.e., relatively close to) the predetermined threshold, then the time-varying masking gain is small as compared to the time-varying masking gain for a frequency band having an energy ratio that is far below its respective threshold and having a greater time-varying masking gain (i.e., greater amount of reduction).
- the time-varying masking gain for all of the audio frequency bands can be used to generate a spatial filtering mask.
- the time-varying masking gain for each audio frequency band is applied to the near-field audio signal for each respective audio frequency band to generate a filtered near-field audio signal for each audio frequency band. This results in a filtered near-field audio signal in which the audio signal in audio frequency bands determined to include a significant far-field audio signal is reduced by the respective time-varying masking gain.
- the filtered near-field audio signal for each audio frequency band are synthesized (i.e., combined) to produce a synthesized near-field audio signal.
- the synthesized near-field audio signal corresponds to filtered near-field audio in which the near-field audio is isolated from the far-field noise. In other words, the background noise and interference has been decreased, while the near-field sound remains prominent, thereby improving the perception of the near-field sounds.
- the near-field audio signal for such audio frequency band may be passed through unchanged, thereby forming the filtered near field audio signal for such audio frequency band.
- the near-field audio signal in which the audio signal in audio frequency bands determined to include lower levels of far-field audio signal are passed through unchanged as the filtered near field audio signal for such respective audio frequency bands.
- the step of synthesizing the filtered near-field audio signal for each audio frequency band may comprise performing a weighted overlap-add (WOLA) synthesis of the filtered near-field audio signal for each audio frequency band to produce the synthesized near-field audio signal.
- WOLA weighted overlap-add
- the step of partitioning the near-field audio signal and the far-field signal may comprise performing a weighted overlap-add (WOLA) analysis and subband partitioning of each of the near-field audio signal and far-field audio signal to partition the near-field audio signal and the far-field signal into the plurality of audio frequency bands.
- WOLA weighted overlap-add
- the method may further comprise performing an adaptive frequency smoothing of the time varying gain for each audio frequency band based on an estimate of a signal to noise ratio of the near-field audio signal, wherein for low signal to noise ratio signals the weighting of neighboring sub-bands are weighted more heavily than the weighting of neighboring sub-bands for high signal to noise ratio signals.
- This step may be performed after the step of calculating a time-varying masking gain for each audio frequency band using the energy ratio for the respective audio frequency band.
- the far-field audio signal may be generated from combining a plurality of audio signals from a plurality of reference microphones.
- the microphone array may include two symmetric reference microphones. Then, the far-field audio signal may be generated by taking the difference between the respective audio signals of the two or more reference microphones.
- the time-varying masking gain for each audio frequency band may be linearly related to the amount the respective energy ratio is below the predetermined threshold, such that the further below the predetermined threshold the greater the amount of time-varying masking gain.
- the relationship may be non-linear instead of linear.
- the method may be specifically implemented for filtering the vocal speech of a user from background noise.
- the near-field audio signal is a near-field speech signal of a user and the primary microphone is positioned proximate to a user’s mouth.
- the far-field signal is a background signal and the one or more reference microphones are positioned further away from the user’s mouth than the primary microphone.
- the method isolates the user’s speech and attenuates ambient noise and improves an isolation of the user’s speech from background noise by increasing the signal to noise ratio of the user’s speech.
- Another embodiment disclosed herein is directed to an XR system which implements any of the methods disclosed herein for reducing far-field noise in a near-field audio signal using a microphone array.
- the noise reducing methods are implemented on an XR system to improve and isolate a user’s speech from background noise.
- the XR system comprises an XR computer system having a computer processor, memory, a storage device, and software stored on the storage device and executable to program the computer to perform operations enabling the XR system.
- the XR system also has a wearable support structure, such as a headset, configured to be worn on the head of the subject and a display system for displaying 3D virtual images (i.e., XR images) in an XR field of view to a user.
- the display system is carried by the support structure, such as eyepieces on a headset.
- the display may include a pair of light projectors, panel displays, or the like, and optic elements to project the 3D virtual images in the XR field of view into the eyes of the user.
- the XR system is configured to present 3D virtual images in an XR field of view to the user which simulate accurate locations of virtual objects in a world coordinate system.
- the headset may also allow a degree of transparency to the real-world surrounding the user such that the XR images augment the visualization of the real-world.
- the 3D virtual images may simulate accurate locations of virtual objects in a world coordinate system.
- the XR system also includes a microphone array.
- the microphone array includes a primary microphone configured to be positioned proximate to a user’s mouth to sense near field audio and generate a near-field audio signal corresponding to the near-field audio sound.
- the microphone array also includes one or more reference microphones configured to be positioned further away from the user’s mouth than the primary microphone to sense far-field audio sound and to generate a far-field audio signal corresponding to the far-field audio sound.
- the primary microphone may be carried on an XR headset (the support structure) in a front-side of the headset to position the primary microphone in front of the user’s face
- the one or more reference microphones may be carried on the XR headset on either side of the headset to position the one or more reference microphones toward a side of the user’s head.
- the XR system also has an audio processor operably coupled to the primary microphone and the one or more reference microphones to acquire the near-field audio signal and the far-field audio signal.
- the audio processor includes a microprocessor and software configured to program the audio processor to perform a process for reducing far-field noise from a microphone array including any of the methods disclosed herein. For instance, the process may comprise:
- the audio processor may be housed in a body-pack configured to be removably attached to the user’s body.
- the body-pack may be a beltpack for wearing on the user’s hip, or any other suitable structure.
- the audio processor may be integrated with the XR computer system, or it may be a separate system or module.
- the XR system may be configured such that the process for reducing far-field noise includes any combination of the one or more aspects of the method embodiments for reducing far-field noise in a near-field audio signal using a microphone array.
- Another disclosed embodiment disclosed herein is directed to a non-transitory computer- readable medium having stored thereon a sequence of instructions that, when stored in memory and executed by a processor programs the processor to perform a process for reducing far-field noise in a near-field audio signal using a microphone array. Accordingly, in one embodiment, the process includes:
- Fig. l is a picture of a three-dimensional extended reality scene that can be displayed to an end user by an extended reality system, according to some embodiments;
- FIG. 2 is a perspective view and block diagram of an augmented reality system constructed in accordance with one embodiment of the present inventions
- FIG. 3 illustrates the extended reality system of Fig. 2 having a body-pack configured to be removably attached to the user’s body;
- Fig. 4 is a flow chart illustrating a process for reducing far-field noise in a speech audio signal using the microphone array of the extended reality system of Fig. 2;
- Fig. 5 shows an example of the energy ratios for frequency bands from an test case of speech sound captured by a one or more primary microphones and ambient sound captured by two reference microphones ;
- Fig. 6 is a graph representing the spatial filter mask of Fig. 5 after thresholding and warping;
- Fig. 7 is a graph of estimates of the signal to noise ratio for the near-field audio signal in the test case of Fig. 5;
- Fig. 8A shows an example of a graph of adaptive frequency smoothing curves for a low signal to noise near-field audio signal
- Fig. 8B shows an example of a graph of adaptive frequency smoothing curves for a high signal to noise near-field audio signal
- Fig. 9 shows a near-field audio signal before noise reduction
- Fig. 10 shows the near-field audio signal after noise reduction.
- the description that follows discloses the technology for reducing far-field noise from near-field audio using a microphone array as implemented in an illustrative XR system 200 (see Figs. 2-3).
- the disclosed systems and methods for reducing far-field noise from near-field audio using microphone array utilizes a spatial filtering approach which filters far- field ambient noise from the near-field audio and reduces the far-field noise in the near-field audio.
- This filtering technique improves the near-field audio by isolating the near-field audio from the far-field interference.
- this filtering technique is used to effectively filter speech (i.e., near-field audio) from ambient noise (i.e., far-field noise).
- speech i.e., near-field audio
- ambient noise i.e., far-field noise
- the systems and methods for reducing far-field noise from near-field audio using a microphone array are not limited to use in XR systems, but may be used in any suitable application, device, or system to improve and isolate a target near-field audio signal from far-field audio interference.
- AR scenarios typically include presentation of virtual content (e.g., images and sound) corresponding to virtual objects in relationship to real -world objects.
- Fig. 1 depicts an illustration of an XR scenario (specifically, an AR scenario) with certain virtual reality objects, and certain physical, real-world objects, as viewed by a user on a 3D display system of the XR system 200 (see Fig. 2).
- an XR scene 100 is depicted wherein the user of XR system 200 sees a real-world, physical, park-like setting 102 featuring people, trees, buildings in the background, and a real -world, physical concrete platform 104.
- the user 252 of the XR system 200 also perceives that they “see” a virtual robot statue 106 standing upon the physical concrete platform 104, and a virtual cartoon-like avatar character 108 flying by which seems to be a personification of a bumblebee, even though these virtual objects 106, 108 do not exist in the real -world.
- Fig. 2 illustrates an XR system 200, according to some embodiments disclosed herein.
- the XR system 200 is a wearable system which comprises a display-mounted headset 205 which is worn on the head 250 of the user 252.
- the XR headset 205 includes a wearable support structure comprising a frame structure 206 configured to be worn on the head 250 of the user 252, similar to an eyeglasses frame.
- the XR system 200 is not required to be a wearable system, but instead may include a separate display which may be a portable monitor, table-top monitor, tablet computer, smartphone or the like.
- a wearable system has the advantage of allowing the user to keep his/her hands free while using the XR system 200, and in the case of a headset, provides an immersive XR experience.
- the display screen 204 is a partially transparent display screen through which real objects in the ambient environment can be seen by the end user 252 and onto which images of virtual objects may be displayed.
- the frame structure 206 carries the partially transparent display screen 204, such that the display screen 204 is positioned in front of the eyes 248 of the end user 50, and in particular in the end user’s 252 field of view between the eyes 248 of the end user 252 and the ambient environment.
- the display subsystem 208 is designed to present the eyes 248 of the end user 252 with photo-based radiation patterns that can be comfortably perceived as augmentations to physical reality, with high-levels of image quality and three-dimensional perception, as well as being capable of presenting two-dimensional content.
- the display subsystem 208 presents a sequence of frames at high frequency that provides the perception of a single coherent scene.
- the XR system 200 may employ one or more imagers (e.g., cameras) to capture and transform images of the ambient environment into video data, which can then be inter-mixed with video data representing the virtual objects, in which case, the XR system 200 may display images representative the intermixed video data to the end user 252 on an opaque display surface.
- imagers e.g., cameras
- the XR system 200 may display images representative the intermixed video data to the end user 252 on an opaque display surface.
- Further details describing display subsystems are provided in U.S. Provisional Patent Application Ser. No. 14/212,961, entitled “Display Subsystem and Method,” and U.S. Provisional Patent Application Ser. No. 14/331,216, entitled “Planar Waveguide Apparatus With Diffraction Element(s) and Subsystem Employing Same,” which are expressly incorporated herein by reference.
- the XR system 200 further comprises one or more speaker(s) 210 for presenting sound only from virtual objects to the end user 252, while allowing the end user 252 to directly hear sound from real objects.
- the speaker(s) 210 are carried by the frame structure 206, such that the speaker(s) 210 are positioned adjacent (in or around) the ear canals of the end user 252, e.g., earbuds or headphones.
- the speaker(s) 210 may provide for stereo/shapeable sound control. Although the speaker(s) 210 are described as being positioned adjacent the ear canals, other types of speakers that are not located adjacent the ear canals can be used to convey sound to the end user 252.
- speakers 210 may be placed at a distance from the ear canals, e.g., using bone conduction technology.
- multiple spatialized speakers 210 may be located about the head 250 of the end user 252 (e.g., four speakers) and be configured for exhibiting sound from the left, right, front, and rear of the head 250 and pointed towards the left and right ears 254 of the end user 252. Further details on spatialized speakers that can be used for augmented reality systems are described in U.S. Provisional Patent Application Ser. No. 62/369,561, entitled “Mixed Reality System with Spatialized Audio,” which is expressly incorporated herein by reference.
- the XR system 200 further comprises a microphone array 220 configured for capturing and converting real sound, including speech of the user 252 and sounds originating from the ambient environment surrounding the user 252 into audio signals which are output to an audio processor 246.
- the microphone array 220 includes a plurality of microphones 222, 224 carried on the frame structure 206 of the headset 205. In the illustrated embodiment of Fig. 2, the microphone array 220 includes 4 microphones, including 2 front/primary microphones 222a-222b and 2 side/reference microphones 224a-224b.
- the microphone array 220 may include any suitable number of microphones 222, 224 to effectively capture sound and process the sound for use by the XR system 200, as disclosed herein.
- the sound captured by the microphones 222 can be used for communication by the user, and/or inter-mixed with the audio data from virtual sound, in which case, the speaker(s) 210 may convey sound representative of the intermixed audio data to the end user 252.
- the front microphones 222 also referred to as primary microphones 222 are configured and positioned to capture speech/voice of the user and ambient sound in front of the user 252.
- the side microphones 224 also referred to as reference microphones 224) are configured and positioned to capture ambient sound originating from the ambient environment.
- the microphones 224 may be positioned anywhere relative to the user including on the front and back of the user, but are positioned and configured to capture minimal speech/voice sound, or at least significantly less speech/voice sound than the front microphones 222. Accordingly, the side microphones 224 are typically located further from the user’ mouth than the front microphones 222 and/or are directed to capture sounds originating from the ambient environment and not the user’s mouth.
- the microphone array 220 includes a first front microphone 222a (also referred to as a first primary microphone 222a) which is located on the bottom part of the right (from the perspective of the user 252), front side of the headset 205 to position the first primary microphone 222a proximate the user’s mouth in order to effectively capture the speech/voice of the user 252.
- the microphone array 220 includes a second front microphone 222b located on the left, front side of the headset 205. Hence, the second microphone 222b is also located relatively close to the user’s mouth such that it can be used as a second primary microphone 222b in the far-field noise reduction process described herein.
- the microphone array 220 also includes a right side microphone 224a (also referred to as a first reference microphone 224a) located on the right temple piece of the frame structure 206 to position the right side microphone 224a on the right side of the user’s head.
- a left side microphone 224b (also referred to as a second reference microphone 224b) is located on the left temple piece of the frame structure 206 to position the left side microphone 224b on the left side of the user’s head.
- the XR system 200 may include additional front microphones 222 (i.e., primary microphones 222) and/or additional side microphones 224 (i.e., reference microphones 224, configured and positioned accordingly.
- Each of the microphones 222 and 224 are configured to sense sound and to output an audio signal corresponding to the sensed sound.
- the microphones 222 and 224 may be digital microphones which output a digital audio signal or analog microphones which output an analog audio signal.
- the XR system 200 may optionally employ a spatialized audio system that renders and presents spatialized audio corresponding to virtual objects with the known virtual locations and orientations in real and physical three-dimensional (3D) space, making it appear to the end user 252 that the sounds are originating from the virtual locations of the real objects, so as to affect clarity or realism of the sound.
- the XR system 200 tracks a position of the end user 252 to more accurately render spatialized audio, such that audio associated with various virtual objects appear to originate from their virtual positions.
- the XR system 200 may track a head pose of the end user 252 to more accurately render spatialized audio, such that directional audio associated with various virtual objects appears to propagate in virtual directions appropriate for the respective virtual objects (e.g., out of the mouth of a virtual character, and not out of the back of the virtual characters’ head). Moreover, the XR system 200 may take into account other real physical and virtual objects in rendering the spatialized audio, such that audio associated with various virtual objects appear to appropriately reflect off of, or occluded or obstructed by, the real physical and virtual objects.
- the sensor(s) may include one or more image capture devices (such as visible and infrared light cameras), inertial measurement units (including accelerometers and gyroscopes), compasses, microphones, GPS units, or radio devices.
- the sensor(s) comprises the forward-facing camera(s) 230.
- the forward-facing camera(s) 230 are particularly suited to capture information indicative of distance and angular position (i.e., the direction in which the head is pointed) of the head 250 of the end user 252 with respect to the environment in which the end user 250 is located. Head orientation may be detected in any direction (e.g., up/down, left, right with respect to the reference frame of the end user 252).
- the forward-facing camera(s) 230 are also configured for acquiring video data of real objects in the ambient environment to facilitate the video recording function of the XR system 200. Cameras may also be provided for tracking real objects in the ambient environment.
- the frame structure 206 may be designed, such that the cameras may be mounted on the front and back of the frame structure 106. In this manner, the array of cameras may encircle the head 250 of the end user 252 to cover all directions of relevant objects.
- the XR system 200 may also optionally include one or more rearward-facing camera(s) 232 and a corresponding processor that track the eyes 248 of the end user 252, and in particular the direction and/or distance at which the end user 252 is focused.
- the rearward-facing camera(s) 232 may track angular position (the direction in which the eye or eyes are pointing), blinking, and depth of focus (by detecting eye convergence) of the eyes 248 of the end user 252. Further details discussing eye tracking devices are provided in U.S. Provisional Patent Application Ser. No. 14/212,961, entitled “Display Subsystem and Method,” U.S. Patent Application Ser. No.
- the XR system 200 further comprises a three-dimensional database 242 configured for storing a virtual three-dimensional scene, which comprises virtual objects (both content data of the virtual objects, as well as absolute metadata associated with these virtual objects, e.g., the absolute position and orientation of these virtual objects in the 3D scene) and virtual objects (both content data of the virtual objects, as well as absolute metadata associated with these virtual objects, e.g., the volume and absolute position and orientation of these virtual objects in the 3D scene, as well as space acoustics surrounding each virtual object, including any virtual or real objects in the vicinity of the virtual source, room dimensions, wall/floor materials, etc.).
- virtual objects both content data of the virtual objects, as well as absolute metadata associated with these virtual objects, e.g., the absolute position and orientation of these virtual objects in the 3D scene
- virtual objects both content data of the virtual objects, as well as absolute metadata associated with these virtual objects, e.g., the volume and absolute position and orientation of these virtual objects in the 3D scene, as well as space acou
- the augmented reality system 200 further comprises a control subsystem 202 that, in addition to recording video data originating from virtual objects and real objects that appear in the field of view and audio data captured by the microphone array 220.
- the XR system 200 may also record metadata associated with the video data and audio data, so that synchronized video and audio may be accurately re-rendered during playback.
- control subsystem 202 comprises a video processor 244 configured for acquiring the video content and absolute metadata associated with the virtual objects from the three-dimensional database 242 and acquiring head pose data of the end user 252 (which can be used to localize the absolute metadata for the video to the head 250 of the end user 252) from the head/object tracking subsystem 240, and rendering video therefrom, which is then conveyed to the display subsystem 208 for transformation into images that are intermixed with images originating from real objects in the ambient environment in the field of view of the end user 252.
- the video processor 244 is also configured for acquiring video data originating from real objects of the ambient environment from the forward-facing camera(s) 230, which along with video data originating from the virtual objects, will be subsequently recorded, as will be further described below.
- the audio processor 246 is configured for acquiring audio content and metadata associated with the virtual objects from the three-dimensional database 242 and acquiring head pose data of the end user 252 (which can be used to localize the absolute metadata for the audio to the head 250 of the end user 252) from the head/object tracking subsystem 240, and rendering spatialized audio therefrom, which is then conveyed to the speaker(s) 210 for transformation into spatialized sound that is intermixed with the sounds originating from the real objects in the ambient environment.
- the audio processor 246 is also configured for acquiring audio data captured by the microphones 222, 224 of the microphone array 220, including sound originating from the user’s mouth (speech/voice) and from the ambient environment. This audio data, along with the spatialized audio data from the selected virtual objects, along with any resulting metadata localized to the head 250 of the end user 252 (e.g., position, orientation, and volume data) for each virtual object, as well as global metadata (e.g., volume data globally set by the XR system 200 or end user 252), may be subsequently recorded, as will be further described below.
- the audio processor 246 and XR system 200 are also configured to enable voice communication between the user 252 and other parties, and/or to use the voice sound for voice-activated functions and commands.
- the audio processor 246 is further configured to process the speech audio signal (i.e., a near-field audio signal) captured by the primary microphone(s) 222 and to filter far-field noise from the speech audio signal using the microphone array 220 to produce a filtered speech audio signal (i.e., filtered near-field audio signal) having improved isolation of the speech sound with decreased background noise.
- the audio processor 246 includes a noise filtering software program (which may be in the form of software and/or firmware) which programs a microprocessor to perform a process for reducing far-field noise in the speech audio signal (i.e., a near-field audio signal) captured by the primary microphone(s) 222 using the microphone array 220.
- a process 400 for reducing far-field noise in the speech audio signal i.e., a near-field audio signal
- the audio processor acquires the respective near-field audio signals from one or more of the primary microphones 222.
- the process 400 may utilize only the first near-field audio signal from just the first primary microphone 222a, or the process may utilize both the first and second near-field audio signals from both the first and second primary microphones 222a, 222b and combine them into a near-field audio signal.
- the near-field audio signals can be combined to provide some directionality to the near-field audio signal, such as a dipole pattern.
- the audio processor 246 acquires the respective reference audio signals from one or more of the reference microphones 224.
- the audio processor 246 may acquire and utilize in the process 400 only one of the reference audio signals from either the first or second reference microphones 224a, 224b, or both the first and second reference audio signals from both the first and second reference microphones 224a, 224b.
- the reference audio signals can be combined to provide some directionality to the far-field audio signal, such as a dipole audio signal.
- the reference microphones 224 may be configured to utilize beamforming techniques such as Delay and Sum, Frost, or Minimum Variance Distortionless Response (MVDR) to determine a far-field audio signal.
- beamforming techniques such as Delay and Sum, Frost, or Minimum Variance Distortionless Response (MVDR)
- the audio processor 246 calculates a near-field energy for the near-field audio signal in each audio frequency band and a far-field energy for the far-field audio signal in each audio frequency band. This results in a respective near-field energy value for each audio frequency band and a respective far-field energy value for each audio frequency band.
- the resultant near-field energy and far-field energy in each audio frequency band are used to calculate an energy ratio of the near-field energy to the far-field energy for each respective audio frequency band.
- the ratio of the near-field energy in a first audio frequency band to the far-field energy in the same first audio frequency band is calculated
- the ratio of the near-field energy in a second audio frequency band to the far-field energy in the same second audio frequency band is calculated, and so-on for each of the audio frequency bands.
- the audio processor 246 determines whether the energy ratio for each audio frequency band is below a predetermined threshold for the respective audio frequency band.
- the predetermined threshold sets a limit of the ratio of the near-field energy to the far-field energy for a respective audio frequency band.
- This step of the process determines the audio frequency bands having a relatively high energy ratio of the near-field energy to the far-field energy such that the near-field audio signal in that frequency band is considered to be significantly speech sound (i.e., target, desirable sound), and the audio frequency bands having a relatively low energy ratio of the near-field energy such that the near-field audio signal in that frequency band is considered to be mostly background noise or interference and not speech.
- the predetermined threshold may be unity, or near unity.
- the predetermined thresholds may be frequency band dependent, such that each audio frequency band has its own predetermined threshold, which may be different for each audio frequency band. In an alternative embodiment, the predetermined threshold may be the same for all of the audio frequency bands.
- the predetermined threshold(s) are tunable through empirical experimentation to determine predetermined threshold(s) providing the desired, or best, far-field filtering effect.
- Fig. 5 shows an example of the energy ratios for 42 frequency bands from a test case of speech sound captured by one or more primary microphones and ambient sound captured by two reference microphones.
- the graph of Figs. 5 represents a spatial filtering mask prior to thresholding and warping, as described below.
- the audio processor 246 calculates a respective time-varying masking gain based upon the energy ratio for the respective audio frequency band. This is also called warping the spatial filtering mask.
- the time-varying masking gain for each audio frequency band may be proportional to the amount the respective energy ratio is below the respective predetermined threshold. Hence, if the energy ratio is just slightly below (i.e., relatively close to) the predetermined threshold, then the time-varying masking gain is small as compared to the time- varying masking gain for a frequency band having an energy ratio that is far below its respective threshold and having a greater time-varying masking gain (i.e., greater amount of reduction).
- the time-varying masking gain for all of the audio frequency bands can be used to generate a spatial filtering mask.
- the near-field audio signal for such audio frequency band is passed through unchanged, as the filtered near field audio signal for such respective audio frequency bands.
- the graph of Fig. 6 represents the spatial filter mask of Fig. 5 after thresholding and warping. It can be seen in Fig. 6 that the audio frequency bands having low energy ratios in Fig. 5 are significantly decreased in Fig. 6.
- Step 420 is an optional step, and is included in this embodiment of the method 400.
- the audio processor 246 performs an adaptive frequency smoothing of the time-varying masking gain for each audio frequency band based on an estimate of a signal to noise ratio of the near-field audio signal.
- Fig. 7 shows a graph of estimates of the signal to noise ratio for the near- field audio signal in the test case of Fig. 5.
- the weighting of neighboring sub-bands are weighted more heavily than the weighting of neighboring sub-bands for high signal to noise ratio signals.
- the adaptive frequency smoothing smooths in frequency the spatial masking gain thereby reducing musical noise artifacts.
- Figs. 8A and 8B show an example of a graph of adaptive frequency smoothing curves for a low signal to noise near-field audio signal. Each curve in the graph shows the smoothing curve for a respective frequency band.
- Fig. 8B shows an example of a graph of adaptive frequency smoothing curves for a high signal to noise near-field audio signal. Again, each curve in the graph shows the smoothing curve for a respective frequency band. It can be seen by comparing the adaptive frequency curves of Fig. 8A to Fig. 8B that for a low signal to noise ratio near-field audio signal, the neighboring frequency bands are more heavily weighted as compared to a high signal to noise ratio near-field audio signal.
- the audio processor 246 applies the time-varying masking gain for each audio frequency band to the near-field audio signal for each respective audio frequency band to generate a filtered near-field audio signal for each audio frequency band. This results in a filtered near- field audio signal in which the audio signal in audio frequency bands determined to include a significant far-field audio signal is reduced by the respective time-varying masking gain.
- Figs. 9 and 10 illustrate the filtering of the near-field audio signal in the test case of Fig. 5.
- the audio processor 246 synthesizes (i.e., combines) the filtered near-field audio signal for each audio frequency band thereby producing a synthesized near-field audio signal.
- the synthesized near-field audio signal corresponds to filtered near-field audio in which the near-field audio is isolated from the far-field noise.
- the background noise and interference have been decreased, while the near-field sounds remain prominent, thereby improving the perception of the near-field speech sound.
- Fig. 9 shows the near-field audio signal before noise reduction.
- Fig. 10 shows the near-field audio signal after noise reduction using the method 400. It can be seen in Fig.
- the XR system 200 further comprises memory 260, and a recorder 262 configured for storing video and audio in the memory 260, which may be accessed for playback.
- the control subsystem that performs the functions of the video processor 1244, audio processor 246, recorder 262, and an audio/video player may take any of a large variety of forms, and may include a number of controllers, for instance one or more microcontrollers, microprocessors or central processing units (CPUs), digital signal processors, graphics processing units (GPUs), other integrated circuit controllers, such as application specific integrated circuits (ASICs), programmable gate arrays (PGAs), for instance, field PGAs (FPGAs), and/or programmable logic controllers (PLUs).
- controllers for instance one or more microcontrollers, microprocessors or central processing units (CPUs), digital signal processors, graphics processing units (GPUs), other integrated circuit controllers, such as application specific integrated circuits (ASICs), programmable gate arrays (PGAs), for instance, field PGAs (FPGAs), and/or programmable logic controllers (PLUs).
- the functions of the video processor 244, audio processor 246, and recorder 262 may be respectively performed by single integrated devices, at least some of the functions of the video processor 244, audio processor 246, and recorder 262 may be combined into a single integrated device, or the functions of each of the video processor 244, audio processor 246 and recorder 262 may be distributed amongst several devices.
- the video processor 244 may comprise a graphics processing unit (GPU) that acquires the video data of virtual objects from the three- dimensional database 242 and renders the synthetic video frames therefrom, and a central processing unit (CPU) that acquires the video frames of real objects from the forward-facing camera(s) 230.
- the audio processor 246 may comprise a digital signal processor (DSP) that processes the audio data acquired from a microphone subsystem and microphone array 220, and the CPU that processes the audio data.
- DSP digital signal processor
- the various processing components of the XR system 200 may be physically contained in a distributed subsystem.
- the XR system 200 may further comprises a remote processing module 300 and remote data repository 302 operatively coupled, such as by a wired lead or wireless connectivity 304, 306, to a local processing and data module 308, such that these remote modules 300, 302 are operatively coupled to each other and available as resources to the local processing and data module 308.
- a remote processing module 300 and remote data repository 302 operatively coupled, such as by a wired lead or wireless connectivity 304, 306, to a local processing and data module 308, such that these remote modules 300, 302 are operatively coupled to each other and available as resources to the local processing and data module 308.
- the XR system 200 may comprise a local processing and data module 308 operatively coupled, such as by a wired lead or wireless connectivity , to components carried by the headset 205 worn on the head 250 of the end user 252 (e.g., the projection subsystem of the display subsystem 208, microphone array 220, audio processor 246, speakers 210, and cameras 230, 232).
- the local processing and data module 308 may include any one or more of the components and subsystems included in the control subsystem 201 of Fig. 2, such as the display subsystem 208, audio processor 246, etc.
- any one or more of the GPU of the video processor 244 and CPU of the video processor 244 and/or audio processor 246 may be contained in the remote processing module 308, in order to limit the power consumption of the local processing and data module 308 which may be battery-powered, to improve portability, although in alternative embodiments, these components, or portions thereof may be contained in the local processing and data module 308.
- local processing and data module 308 may be mounted in a body -pack configured to be removably attached to the user’s body, such as a belt-pack having a belt-coupling for wearing on the user’s hip 310.
- the local processing and data module 308 may be fixedly attached to the frame structure 206, fixedly attached to a helmet or other headpiece worn by the user 252, embedded in headphones, removably attachable to the user’s torso, etc.
- the local processing and data module 308 may comprise a power-efficient processor or controller, as well as digital memory, such as flash memory, both of which may be utilized to assist in the processing, caching, and storage of data captured from the sensors and/or acquired and/or processed using the remote processing module 300 and/or remote data repository 302, possibly for passage to the display subsystem 208 after such processing or retrieval.
- the remote processing module 300 may comprise one or more relatively powerful processors or controllers configured to analyze and process data and/or image information.
- the remote data repository 302 may comprise a relatively large-scale digital data storage facility, which may be available through the internet or other networking configuration in a “cloud” resource configuration. In an alternative embodiment, all data is stored and all computation is performed in the local processing and data module 308, allowing fully autonomous use from any remote modules.
- the couplings 304, 306 between the various components described above may include one or more wired interfaces or ports for providing wires or optical communications, or one or more wireless interfaces or ports, such as via RF, microwave, and IR for providing wireless communications.
- all communications may be wired, while in other implementations all communications may be wireless, with the exception of optical fiber(s) used in the display subsystem 208.
- the choice of wired and wireless communications may be different from that illustrated in Fig. 3. Thus, the particular choice of wired or wireless communications should not be considered limiting.
Landscapes
- Engineering & Computer Science (AREA)
- Signal Processing (AREA)
- Physics & Mathematics (AREA)
- Acoustics & Sound (AREA)
- Health & Medical Sciences (AREA)
- Otolaryngology (AREA)
- General Health & Medical Sciences (AREA)
- Computational Linguistics (AREA)
- Quality & Reliability (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Human Computer Interaction (AREA)
- Multimedia (AREA)
- Circuit For Audible Band Transducer (AREA)
- Soundproofing, Sound Blocking, And Sound Damping (AREA)
Abstract
Systems and methods for reducing far-field noise (background noise) from near-field audio signals generated by a microphone array. The systems and methods disclosed herein use an innovative spatial filtering approach which filters far-field ambient noise from the near-field audio and reduces the far-field noise in the near-field audio signal thereby improving and isolating the near-field audio from the far-field interference.
Description
FAR-FIELD NOISE REDUCTION VIA SPATIAL FILTERING USING A
MICROPHONE ARRAY
CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of priority of U.S. Provisional Patent Application Serial No. 63/453,022, entitled “FAR-FIELD NOISE REDUCTION VIA SPATIAL FILTERING USING A MICROPHONE ARRAY,” and filed on March 17, 2023. The contents of the foregoing application is hereby expressly incorporated by reference for all purposes.
FIELD OF THE INVENTION
[0002] The present invention generally relates to systems and methods for reducing far-field noise through spatial filtering using a microphone array, which may be utilized in spatialized audio systems within virtual reality, augmented reality and/or mixed reality systems.
BACKGROUND
[0003] Modern computing and display technologies have facilitated the development of mixed reality systems (“MR”) for so called “virtual reality” or “augmented reality” experiences, wherein digitally reproduced images or portions thereof are presented to a user in a manner wherein they seem to be, or may be perceived as, real. A virtual reality, or “VR”, scenario typically involves presentation of digital or virtual image information without transparency to actual real-world visual input. An augmented reality, or “AR”, scenario typically involves presentation of digital or virtual image information as an augmentation to visualization of the actual world around the user (i.e., transparency to other actual real-world visual input). Accordingly, AR scenarios involve presentation of digital or virtual image information with transparency to other actual real-world
visual input. As used herein, the terms “extended reality” and “XR” are used to refer collectively to any of VR, AR and/or MR. In addition, the term “AR” means either, or both, AR and MR.
[0004] VR and AR systems typically employ head-worn displays (or helmet-mounted displays, or smart glasses) that are at least loosely coupled to a user’s head, and thus move when the end user’s head moves. If the end user’s head motions are detected by the display system, the data being displayed can be updated to take into account the change in head pose (i.e., the orientation and/or location of the user’s head).
[0005] As an example, if a user wearing a head-worn display device views a virtual representation of a virtual object on the display device and walks around an area where the virtual object appears, the virtual object can be rendered for each viewpoint (corresponding to a position and/or orientation of the head-worn display device), giving the user the perception that they are walking around an object that occupies real space. If the head-worn display device is used to present multiple virtual objects at different depths, measurements of head pose can be used to render the scene to match the user’s dynamically changing head pose and provide an increased sense of immersion. However, there is an inevitable lag between rendering a scene and displaying/projecting the rendered scene.
[0006] Head-worn displays that enable AR (i.e., the concurrent viewing of virtual and real objects) can have several different types of configurations. In one such configuration, often referred to as a “video see-through” display, a camera captures elements of a real scene, a computing system superimposes virtual elements onto the captured real scene, and a non-transparent display presents the composite image to the eyes. Another configuration is often referred to as an “optical see- through” display, in which the end user can see through transparent (or semi-transparent) elements in the display system to directly view the light from real objects in the environment. The
transparent element, often referred to as a “combiner,” superimposes light from the display over the end user’s view of the real world. A camera may be mounted onto the head-worn display device to capture images or videos of the scene being viewed by the user.
[0007] XR systems typically also include a microphone arrangement including one or more microphones for sensing audio (i.e., sound), such as the speech (i.e., voice) of the user, and ambient/environment sound in the real world surroundings of the user, and generating audio signals corresponding to the audio. For instance, in many instances, the user may wish to communicate using various speech based transmission protocols (e.g., group chat, IP based speech communication platforms, etc.) in real world environments where the ambient noise level may be relatively high. For example, the user may be on a factory floor, in other commercial environments, in the vicinity of children playing, in the vicinity of media such as television or music playing in the background, etc. The relatively high level of ambient noise hinders effective speech communication because the person(s) to whom the user is communicating cannot clearly hear the user’s speech through the XR system, and/or may negatively impact speech recognition performance by speech recognition systems such as systems configured to recognize and process voice commands.
[0008] Thus, there remains a need for improved means for effectively discriminating between user speech and ambient noise in the audio signals generated by a microphone arrangement, which can, for example, improve speech communication and speech recognition.
SUMMARY
[0009] The present disclosure is directed to systems and methods for reducing unwanted noise from audio signals generated by a microphone array. The systems and methods disclosed herein use an innovative spatial filtering approach which filters far-field ambient noise from the near-
field audio and reduces the far-field noise in the near-field audio signal thereby improving and isolating the near-field audio from the far-field interference. Although not limited to only this use, as described herein, the systems and methods are useful for effectively filtering speech (i.e., near- field audio) from ambient noise (i.e., far-field noise). As described herein, such systems and methods may be implemented on XR systems to improve and isolate a user’s speech using a microphone array on a headset of the XR system. It should be understood that the systems and methods disclosed herein may be utilized to filter undesired far-field audio noise from a target near-field audio signal in any suitable scenario, and not solely for isolating a near-field speech signal from far-field noise, or solely within an XR system. For example, the systems and methods could be used to improve and isolate a near-field signal of a target audio (e.g., sound generated by sliding parts of a machine) from ambient audio noise (e.g., noise on a factory floor). This may be desirable to use the improved and isolated target audio signal for testing the machine, detecting defects, or other purposes.
[0010] The systems and methods disclosed herein may be implemented and carried out on any suitable hardware system having a microphone array, a computer processor and software (which may include firmware) configured to program the system to perform a process for reducing far- field noise in a near-field audio signal.
[0011] Hence, one embodiment disclosed herein is directed to a computer-implemented method for reducing far-field noise in a near-field audio signal using a microphone array. The method includes acquiring a near-field audio signal from at least one primary microphone of the microphone array, and acquiring one or more reference audio signals from one or more reference microphones of a microphone array. A far-field signal is determined from the one or more reference audio signals. For example, at least one primary microphone may include a speech/voice
microphone positioned near the mouth of a user, while the one or more reference microphones are positioned further away from the user’s mouth, such as at the side of the user’s head. A far-field signal is determined from the one or more reference audio signals. For example, the far-field signal may be calculated by combining the reference audio signals from the one or more reference microphones, such as calculating a difference between the one or more reference microphones.
[0012] Next, the near-field audio signal and the far-field signal are partitioned into a plurality of audio frequency bands. For example, in one aspect, the near-field audio signal and the far-field audio signal may be partitioned using a weighted overlap-add (WOLA) analysis. The number of audio frequency bands may be any suitable number, such as from 20-60 bands, 40-50 bands, greater than 20 bands, greater than 40 bands, or greater than 50 bands.
[0013] Next, a near-field energy for the near-field audio signal in each audio frequency band and a far-field energy for the far-field audio signal in each audio frequency band is calculated. The resultant near-field energy and far-field energy in each audio frequency band are used to calculate an energy ratio of the near-field energy to the far-field energy for each respective audio frequency band. In other words, the ratio of the near-field energy in a first audio frequency band to the far- field energy in the same first audio frequency band is calculated, the ratio of the near-field energy in a second audio frequency band to the far-field energy in the same second audio frequency band is calculated, and so-on for each of the audio frequency bands.
[0014] The method then determines whether the energy ratio for each audio frequency band is below a predetermined threshold for the respective audio frequency band. For example, in another aspect, the predetermined threshold may be frequency band dependent, such that each audio frequency band has its own predetermined threshold, which may be different for each audio
frequency band. Alternatively, the predetermined threshold may be the same for all of the audio frequency bands.
[0015] For each audio frequency band having an energy ratio below the predetermined threshold, a respective time-varying masking gain is calculated based upon the energy ratio for the respective audio frequency band. For example, in one aspect, the time-varying masking gain for each audio frequency band may be proportional to the amount the respective energy ratio is below the respective predetermined threshold. In such case, if the energy ratio is just slightly below (i.e., relatively close to) the predetermined threshold, then the time-varying masking gain is small as compared to the time-varying masking gain for a frequency band having an energy ratio that is far below its respective threshold and having a greater time-varying masking gain (i.e., greater amount of reduction). In another aspect, the time-varying masking gain for all of the audio frequency bands can be used to generate a spatial filtering mask.
[0016] The time-varying masking gain for each audio frequency band is applied to the near-field audio signal for each respective audio frequency band to generate a filtered near-field audio signal for each audio frequency band. This results in a filtered near-field audio signal in which the audio signal in audio frequency bands determined to include a significant far-field audio signal is reduced by the respective time-varying masking gain.
[0017] Finally, the filtered near-field audio signal for each audio frequency band are synthesized (i.e., combined) to produce a synthesized near-field audio signal. The synthesized near-field audio signal corresponds to filtered near-field audio in which the near-field audio is isolated from the far-field noise. In other words, the background noise and interference has been decreased, while the near-field sound remains prominent, thereby improving the perception of the near-field sounds.
[0018] In another aspect of the method, for each audio frequency band having an energy ratio equal to or greater than the predetermined threshold, the near-field audio signal for such audio frequency band may be passed through unchanged, thereby forming the filtered near field audio signal for such audio frequency band. In other words, the near-field audio signal in which the audio signal in audio frequency bands determined to include lower levels of far-field audio signal are passed through unchanged as the filtered near field audio signal for such respective audio frequency bands.
[0019] In yet another aspect of the method, the step of synthesizing the filtered near-field audio signal for each audio frequency band may comprise performing a weighted overlap-add (WOLA) synthesis of the filtered near-field audio signal for each audio frequency band to produce the synthesized near-field audio signal.
[0020] In still another aspect of the method, the step of partitioning the near-field audio signal and the far-field signal may comprise performing a weighted overlap-add (WOLA) analysis and subband partitioning of each of the near-field audio signal and far-field audio signal to partition the near-field audio signal and the far-field signal into the plurality of audio frequency bands.
[0021] In another aspect, the method may further comprise performing an adaptive frequency smoothing of the time varying gain for each audio frequency band based on an estimate of a signal to noise ratio of the near-field audio signal, wherein for low signal to noise ratio signals the weighting of neighboring sub-bands are weighted more heavily than the weighting of neighboring sub-bands for high signal to noise ratio signals. This step may be performed after the step of calculating a time-varying masking gain for each audio frequency band using the energy ratio for the respective audio frequency band.
[0022] In another aspect of the method, the far-field audio signal may be generated from combining a plurality of audio signals from a plurality of reference microphones. As one example, the microphone array may include two symmetric reference microphones. Then, the far-field audio signal may be generated by taking the difference between the respective audio signals of the two or more reference microphones.
[0023] In yet another aspect of the method, the time-varying masking gain for each audio frequency band may be linearly related to the amount the respective energy ratio is below the predetermined threshold, such that the further below the predetermined threshold the greater the amount of time-varying masking gain. In still another aspect, the relationship may be non-linear instead of linear.
[0024] In still another aspect, the method may be specifically implemented for filtering the vocal speech of a user from background noise. Accordingly, the near-field audio signal is a near-field speech signal of a user and the primary microphone is positioned proximate to a user’s mouth. The far-field signal is a background signal and the one or more reference microphones are positioned further away from the user’s mouth than the primary microphone. Thus, the method isolates the user’s speech and attenuates ambient noise and improves an isolation of the user’s speech from background noise by increasing the signal to noise ratio of the user’s speech.
[0025] Another embodiment disclosed herein is directed to an XR system which implements any of the methods disclosed herein for reducing far-field noise in a near-field audio signal using a microphone array. In particular, the noise reducing methods are implemented on an XR system to improve and isolate a user’s speech from background noise. In one embodiment, the XR system comprises an XR computer system having a computer processor, memory, a storage device, and software stored on the storage device and executable to program the computer to perform
operations enabling the XR system. The XR system also has a wearable support structure, such as a headset, configured to be worn on the head of the subject and a display system for displaying 3D virtual images (i.e., XR images) in an XR field of view to a user. In one aspect, the display system is carried by the support structure, such as eyepieces on a headset. For example, the display may include a pair of light projectors, panel displays, or the like, and optic elements to project the 3D virtual images in the XR field of view into the eyes of the user. The XR system is configured to present 3D virtual images in an XR field of view to the user which simulate accurate locations of virtual objects in a world coordinate system. In the case that the XR system provides an AR, and/or MR experience, the headset may also allow a degree of transparency to the real-world surrounding the user such that the XR images augment the visualization of the real-world. The 3D virtual images may simulate accurate locations of virtual objects in a world coordinate system.
[0026] The XR system also includes a microphone array. The microphone array includes a primary microphone configured to be positioned proximate to a user’s mouth to sense near field audio and generate a near-field audio signal corresponding to the near-field audio sound. The microphone array also includes one or more reference microphones configured to be positioned further away from the user’s mouth than the primary microphone to sense far-field audio sound and to generate a far-field audio signal corresponding to the far-field audio sound. For example, in one aspect, the primary microphone may be carried on an XR headset (the support structure) in a front-side of the headset to position the primary microphone in front of the user’s face, and the one or more reference microphones may be carried on the XR headset on either side of the headset to position the one or more reference microphones toward a side of the user’s head.
[0027] The XR system also has an audio processor operably coupled to the primary microphone and the one or more reference microphones to acquire the near-field audio signal and the far-field
audio signal. The audio processor includes a microprocessor and software configured to program the audio processor to perform a process for reducing far-field noise from a microphone array including any of the methods disclosed herein. For instance, the process may comprise:
[0028] a) acquiring a near-field audio signal from a primary microphone of the microphone array;
[0029] b) acquiring one or more reference audio signals from one or more respective reference microphones of the microphone array, and determining [an estimate for] a far-field audio signal from the one or more reference audio signals;
[0030] c) partitioning the near-field audio signal and the far-field signal into a plurality of audio frequency bands;
[0031] d) calculating a near-field energy for the near-field audio signal in each audio frequency band and a far-field energy for the far-field audio signal in each audio frequency band; [0032] e) calculating an energy ratio of the near-field energy to the far-field energy for each audio frequency band;
[0033] f) determining whether the energy ratio for each audio frequency band is below a predetermined threshold;
[0034] g) for each audio frequency band having an energy ratio below the predetermined threshold, using the energy ratio for each audio frequency band to produce a timevarying masking gain for each audio frequency band;
[0035] h) applying the time-varying masking gain for each audio frequency band to the near-field audio signal for each respective audio frequency band to generate a filtered near- field audio signal for each audio frequency band; and
[0036] i) synthesizing the filtered near-field audio signal for each audio frequency band to produce a synthesized near-field audio signal.
[0037] In another aspect of the XR system, the audio processor may be housed in a body-pack configured to be removably attached to the user’s body. For example, the body-pack may be a beltpack for wearing on the user’s hip, or any other suitable structure.
[0038] In additional aspects of the XR system, the audio processor may be integrated with the XR computer system, or it may be a separate system or module.
[0039] In additional aspects, the XR system may be configured such that the process for reducing far-field noise includes any combination of the one or more aspects of the method embodiments for reducing far-field noise in a near-field audio signal using a microphone array.
[0040] Another disclosed embodiment disclosed herein is directed to a non-transitory computer- readable medium having stored thereon a sequence of instructions that, when stored in memory and executed by a processor programs the processor to perform a process for reducing far-field noise in a near-field audio signal using a microphone array. Accordingly, in one embodiment, the process includes:
[0041] a) acquiring a near-field audio signal from a primary microphone of the microphone array;
[0042] b) acquiring one or more reference audio signals from one or more respective reference microphones of the microphone array, and determining [an estimate for] a far-field audio signal from the one or more reference audio signals;
[0043] c) partitioning the near-field audio signal and the far-field signal into a plurality of audio frequency bands;
[0044] d) calculating a near-field energy for the near-field audio signal in each audio frequency band and a far-field energy for the far-field audio signal in each audio frequency band; [0045] e) calculating an energy ratio of the near-field energy to the far-field energy for each audio frequency band;
[0046] f) determining whether the energy ratio for each audio frequency band is below a predetermined threshold;
[0047] g) for each audio frequency band having an energy ratio below the predetermined threshold, using the energy ratio for each audio frequency band to produce a timevarying masking gain for each audio frequency band;
[0048] h) applying the time-varying masking gain for each audio frequency band to the near-field audio signal for each respective audio frequency band to generate a filtered near- field audio signal for each audio frequency band; and
[0049] i) synthesizing the filtered near-field audio signal for each audio frequency band to produce a synthesized near-field audio signal.
[0050] Additional and other objects, features, and advantages of the technology set forth herein are described in the detailed description, figures and claims.
BRIEF DESCRIPTION OF THE DRAWINGS
[0051] The Provisional Patent Application Serial No. 63/453,022 contained several drawings executed in color. Such color drawings have been converted to gray-scale drawings for the present application and are intended to convey all of the same information as the color drawings. The color drawings are expressly incorporated by reference herein for all purposes. Copies of color drawings
will be provided by the U.S. Patent and Trademark Office upon request and payment of the necessary fee.
[0052] The drawings illustrate the design and utility of preferred embodiments of the present invention, in which similar elements are referred to by common reference numerals. In order to better appreciate how the above-recited and other advantages and objects of the present inventions are obtained, a more particular description of the present inventions briefly described above will be rendered by reference to specific embodiments thereof, which are illustrated in the accompanying drawings. Understanding that these drawings depict only typical embodiments of the invention and are not therefore to be considered limiting of its scope, the invention will be described and explained with additional specificity and detail through the use of the accompanying drawings in which:
[0053] Fig. l is a picture of a three-dimensional extended reality scene that can be displayed to an end user by an extended reality system, according to some embodiments;
[0054] Fig. 2 is a perspective view and block diagram of an augmented reality system constructed in accordance with one embodiment of the present inventions;
[0055] Fig. 3 illustrates the extended reality system of Fig. 2 having a body-pack configured to be removably attached to the user’s body;
[0056] Fig. 4 is a flow chart illustrating a process for reducing far-field noise in a speech audio signal using the microphone array of the extended reality system of Fig. 2;
[0057] Fig. 5 shows an example of the energy ratios for frequency bands from an test case of speech sound captured by a one or more primary microphones and ambient sound captured by two reference microphones ;
[0058] Fig. 6 is a graph representing the spatial filter mask of Fig. 5 after thresholding and warping;
[0059] Fig. 7 is a graph of estimates of the signal to noise ratio for the near-field audio signal in the test case of Fig. 5;
[0060] Fig. 8A shows an example of a graph of adaptive frequency smoothing curves for a low signal to noise near-field audio signal;
[0061] Fig. 8B shows an example of a graph of adaptive frequency smoothing curves for a high signal to noise near-field audio signal;
[0062] Fig. 9 shows a near-field audio signal before noise reduction;
[0063] Fig. 10 shows the near-field audio signal after noise reduction.
DETAILED DESCRIPTION
[0064] Various embodiments will now be described in detail with reference to the drawings, which are provided as illustrative examples of the disclosure so as to enable those skilled in the art to practice the disclosure. Notably, the figures and the examples below are not meant to limit the scope of the present disclosure. Where certain elements of the present disclosure may be partially or fully implemented using known components (or methods or processes), only those portions of such known components (or methods or processes) that are necessary for an understanding of the present disclosure will be described, and the detailed descriptions of other portions of such known components (or methods or processes) will be omitted so as not to obscure the disclosure. Further, various embodiments encompass present and future known equivalents to the components referred to herein by way of illustration.
[0065] The description that follows discloses the technology for reducing far-field noise from near-field audio using a microphone array as implemented in an illustrative XR system 200 (see Figs. 2-3). As described herein, the disclosed systems and methods for reducing far-field noise from near-field audio using microphone array utilizes a spatial filtering approach which filters far-
field ambient noise from the near-field audio and reduces the far-field noise in the near-field audio.
This filtering technique improves the near-field audio by isolating the near-field audio from the far-field interference. In the non-limiting implementation in an XR system, this filtering technique is used to effectively filter speech (i.e., near-field audio) from ambient noise (i.e., far-field noise). However, it is to be understood that the embodiments also lend themselves to applications in other types of display systems (including other types of VR, AR, and/or MR systems), and therefore the embodiments are not to be limited to only the illustrative system disclosed herein. Moreover, the systems and methods for reducing far-field noise from near-field audio using a microphone array are not limited to use in XR systems, but may be used in any suitable application, device, or system to improve and isolate a target near-field audio signal from far-field audio interference.
[0066] Referring to Fig. 1, AR scenarios typically include presentation of virtual content (e.g., images and sound) corresponding to virtual objects in relationship to real -world objects. For example, Fig. 1 depicts an illustration of an XR scenario (specifically, an AR scenario) with certain virtual reality objects, and certain physical, real-world objects, as viewed by a user on a 3D display system of the XR system 200 (see Fig. 2). As shown in Fig. 1, an XR scene 100 is depicted wherein the user of XR system 200 sees a real-world, physical, park-like setting 102 featuring people, trees, buildings in the background, and a real -world, physical concrete platform 104. In addition to these items, the user 252 of the XR system 200 also perceives that they “see” a virtual robot statue 106 standing upon the physical concrete platform 104, and a virtual cartoon-like avatar character 108 flying by which seems to be a personification of a bumblebee, even though these virtual objects 106, 108 do not exist in the real -world.
[0067] Fig. 2 illustrates an XR system 200, according to some embodiments disclosed herein. The XR system 200 is a wearable system which comprises a display-mounted headset 205 which is
worn on the head 250 of the user 252. The XR headset 205 includes a wearable support structure comprising a frame structure 206 configured to be worn on the head 250 of the user 252, similar to an eyeglasses frame. The XR system 200 is not required to be a wearable system, but instead may include a separate display which may be a portable monitor, table-top monitor, tablet computer, smartphone or the like. However, a wearable system has the advantage of allowing the user to keep his/her hands free while using the XR system 200, and in the case of a headset, provides an immersive XR experience.
[0068] Referring to Fig. 2, in the illustrated embodiment, the display screen 204 is a partially transparent display screen through which real objects in the ambient environment can be seen by the end user 252 and onto which images of virtual objects may be displayed. The frame structure 206 carries the partially transparent display screen 204, such that the display screen 204 is positioned in front of the eyes 248 of the end user 50, and in particular in the end user’s 252 field of view between the eyes 248 of the end user 252 and the ambient environment.
[0069] The display subsystem 208 is designed to present the eyes 248 of the end user 252 with photo-based radiation patterns that can be comfortably perceived as augmentations to physical reality, with high-levels of image quality and three-dimensional perception, as well as being capable of presenting two-dimensional content. The display subsystem 208 presents a sequence of frames at high frequency that provides the perception of a single coherent scene.
[0070] In alternative embodiments, the XR system 200 may employ one or more imagers (e.g., cameras) to capture and transform images of the ambient environment into video data, which can then be inter-mixed with video data representing the virtual objects, in which case, the XR system 200 may display images representative the intermixed video data to the end user 252 on an opaque display surface.
[0071] Further details describing display subsystems are provided in U.S. Provisional Patent Application Ser. No. 14/212,961, entitled “Display Subsystem and Method,” and U.S. Provisional Patent Application Ser. No. 14/331,216, entitled “Planar Waveguide Apparatus With Diffraction Element(s) and Subsystem Employing Same,” which are expressly incorporated herein by reference.
[0072] The XR system 200 further comprises one or more speaker(s) 210 for presenting sound only from virtual objects to the end user 252, while allowing the end user 252 to directly hear sound from real objects. The speaker(s) 210 are carried by the frame structure 206, such that the speaker(s) 210 are positioned adjacent (in or around) the ear canals of the end user 252, e.g., earbuds or headphones. The speaker(s) 210 may provide for stereo/shapeable sound control. Although the speaker(s) 210 are described as being positioned adjacent the ear canals, other types of speakers that are not located adjacent the ear canals can be used to convey sound to the end user 252. For example, speakers 210 may be placed at a distance from the ear canals, e.g., using bone conduction technology. In an optional embodiment, multiple spatialized speakers 210 may be located about the head 250 of the end user 252 (e.g., four speakers) and be configured for exhibiting sound from the left, right, front, and rear of the head 250 and pointed towards the left and right ears 254 of the end user 252. Further details on spatialized speakers that can be used for augmented reality systems are described in U.S. Provisional Patent Application Ser. No. 62/369,561, entitled “Mixed Reality System with Spatialized Audio,” which is expressly incorporated herein by reference.
[0073] The XR system 200 further comprises a microphone array 220 configured for capturing and converting real sound, including speech of the user 252 and sounds originating from the ambient environment surrounding the user 252 into audio signals which are output to an audio
processor 246. The microphone array 220 includes a plurality of microphones 222, 224 carried on the frame structure 206 of the headset 205. In the illustrated embodiment of Fig. 2, the microphone array 220 includes 4 microphones, including 2 front/primary microphones 222a-222b and 2 side/reference microphones 224a-224b. The microphone array 220 may include any suitable number of microphones 222, 224 to effectively capture sound and process the sound for use by the XR system 200, as disclosed herein. The sound captured by the microphones 222 can be used for communication by the user, and/or inter-mixed with the audio data from virtual sound, in which case, the speaker(s) 210 may convey sound representative of the intermixed audio data to the end user 252. The front microphones 222 (also referred to as primary microphones 222 are configured and positioned to capture speech/voice of the user and ambient sound in front of the user 252. The side microphones 224 (also referred to as reference microphones 224) are configured and positioned to capture ambient sound originating from the ambient environment. Although referred to as “side” microphones, the microphones 224 may be positioned anywhere relative to the user including on the front and back of the user, but are positioned and configured to capture minimal speech/voice sound, or at least significantly less speech/voice sound than the front microphones 222. Accordingly, the side microphones 224 are typically located further from the user’ mouth than the front microphones 222 and/or are directed to capture sounds originating from the ambient environment and not the user’s mouth.
[0074] The microphone array 220 includes a first front microphone 222a (also referred to as a first primary microphone 222a) which is located on the bottom part of the right (from the perspective of the user 252), front side of the headset 205 to position the first primary microphone 222a proximate the user’s mouth in order to effectively capture the speech/voice of the user 252. The microphone array 220 includes a second front microphone 222b located on the left, front side of
the headset 205. Hence, the second microphone 222b is also located relatively close to the user’s mouth such that it can be used as a second primary microphone 222b in the far-field noise reduction process described herein. The microphone array 220 also includes a right side microphone 224a (also referred to as a first reference microphone 224a) located on the right temple piece of the frame structure 206 to position the right side microphone 224a on the right side of the user’s head. A left side microphone 224b (also referred to as a second reference microphone 224b) is located on the left temple piece of the frame structure 206 to position the left side microphone 224b on the left side of the user’s head. The XR system 200 may include additional front microphones 222 (i.e., primary microphones 222) and/or additional side microphones 224 (i.e., reference microphones 224, configured and positioned accordingly.
[0075] Each of the microphones 222 and 224 are configured to sense sound and to output an audio signal corresponding to the sensed sound. The microphones 222 and 224 may be digital microphones which output a digital audio signal or analog microphones which output an analog audio signal. Accordingly, the first primary microphone 222a generates and outputs a first nearfield audio signal (may also be referred to as “first primary audio signal”), the second primary microphone 222b generates and outputs a second near-field audio signal (may also be referred to as “second first primary audio signal”), the first reference microphone 224a generates and outputs a first reference audio signal, and the second reference microphone 224b generates and outputs a second reference audio signal, and so on, in the case that there are additional primary microphones 222 and/or reference microphones 224.
[0076] In the illustrated embodiment, the XR system 200 may optionally employ a spatialized audio system that renders and presents spatialized audio corresponding to virtual objects with the known virtual locations and orientations in real and physical three-dimensional (3D) space, making
it appear to the end user 252 that the sounds are originating from the virtual locations of the real objects, so as to affect clarity or realism of the sound. The XR system 200 tracks a position of the end user 252 to more accurately render spatialized audio, such that audio associated with various virtual objects appear to originate from their virtual positions. Further, the XR system 200 may track a head pose of the end user 252 to more accurately render spatialized audio, such that directional audio associated with various virtual objects appears to propagate in virtual directions appropriate for the respective virtual objects (e.g., out of the mouth of a virtual character, and not out of the back of the virtual characters’ head). Moreover, the XR system 200 may take into account other real physical and virtual objects in rendering the spatialized audio, such that audio associated with various virtual objects appear to appropriately reflect off of, or occluded or obstructed by, the real physical and virtual objects.
[0077] To this end, the XR system 200 may optionally further comprise a head/object tracking subsystem 240 for tracking the position and orientation of the head 250 of the end user 252 relative to the virtual three-dimensional scene, as well as tracking the position and orientation of real objects relative to the head 250 of the end user 252. For example, the head/object tracking subsystem 240 may comprise one or more sensors configured for collecting head pose data (position and orientation) of the end user 252, and a processor (not shown) configured for determining the head pose of the end user 252 in the known coordinate system based on the head pose data collected by the sensor(s). The sensor(s) may include one or more image capture devices (such as visible and infrared light cameras), inertial measurement units (including accelerometers and gyroscopes), compasses, microphones, GPS units, or radio devices. In the illustrated embodiment, the sensor(s) comprises the forward-facing camera(s) 230. When head worn in this manner, the forward-facing camera(s) 230 are particularly suited to capture information indicative
of distance and angular position (i.e., the direction in which the head is pointed) of the head 250 of the end user 252 with respect to the environment in which the end user 250 is located. Head orientation may be detected in any direction (e.g., up/down, left, right with respect to the reference frame of the end user 252). The forward-facing camera(s) 230 are also configured for acquiring video data of real objects in the ambient environment to facilitate the video recording function of the XR system 200. Cameras may also be provided for tracking real objects in the ambient environment. The frame structure 206 may be designed, such that the cameras may be mounted on the front and back of the frame structure 106. In this manner, the array of cameras may encircle the head 250 of the end user 252 to cover all directions of relevant objects.
[0078] The XR system 200 may also optionally include one or more rearward-facing camera(s) 232 and a corresponding processor that track the eyes 248 of the end user 252, and in particular the direction and/or distance at which the end user 252 is focused. The rearward-facing camera(s) 232 may track angular position (the direction in which the eye or eyes are pointing), blinking, and depth of focus (by detecting eye convergence) of the eyes 248 of the end user 252. Further details discussing eye tracking devices are provided in U.S. Provisional Patent Application Ser. No. 14/212,961, entitled “Display Subsystem and Method,” U.S. Patent Application Ser. No. 14/726,429, entitled “Methods and Subsystem for Creating Focal Planes in Virtual and Augmented Reality,” and U.S. Patent Application Ser. No. 14/205,126, entitled “Subsystem and Method for Augmented and Virtual Reality,” which are expressly incorporated herein by reference.
[0079] The XR system 200 further comprises a three-dimensional database 242 configured for storing a virtual three-dimensional scene, which comprises virtual objects (both content data of the virtual objects, as well as absolute metadata associated with these virtual objects, e.g., the absolute position and orientation of these virtual objects in the 3D scene) and virtual objects (both content
data of the virtual objects, as well as absolute metadata associated with these virtual objects, e.g., the volume and absolute position and orientation of these virtual objects in the 3D scene, as well as space acoustics surrounding each virtual object, including any virtual or real objects in the vicinity of the virtual source, room dimensions, wall/floor materials, etc.).
[0080] The augmented reality system 200 further comprises a control subsystem 202 that, in addition to recording video data originating from virtual objects and real objects that appear in the field of view and audio data captured by the microphone array 220. The XR system 200 may also record metadata associated with the video data and audio data, so that synchronized video and audio may be accurately re-rendered during playback.
[0081] To this end, the control subsystem 202 comprises a video processor 244 configured for acquiring the video content and absolute metadata associated with the virtual objects from the three-dimensional database 242 and acquiring head pose data of the end user 252 (which can be used to localize the absolute metadata for the video to the head 250 of the end user 252) from the head/object tracking subsystem 240, and rendering video therefrom, which is then conveyed to the display subsystem 208 for transformation into images that are intermixed with images originating from real objects in the ambient environment in the field of view of the end user 252. The video processor 244 is also configured for acquiring video data originating from real objects of the ambient environment from the forward-facing camera(s) 230, which along with video data originating from the virtual objects, will be subsequently recorded, as will be further described below.
[0082] The audio processor 246 is configured for acquiring audio content and metadata associated with the virtual objects from the three-dimensional database 242 and acquiring head pose data of the end user 252 (which can be used to localize the absolute metadata for the audio to the head 250
of the end user 252) from the head/object tracking subsystem 240, and rendering spatialized audio therefrom, which is then conveyed to the speaker(s) 210 for transformation into spatialized sound that is intermixed with the sounds originating from the real objects in the ambient environment.
[0083] The audio processor 246 is also configured for acquiring audio data captured by the microphones 222, 224 of the microphone array 220, including sound originating from the user’s mouth (speech/voice) and from the ambient environment. This audio data, along with the spatialized audio data from the selected virtual objects, along with any resulting metadata localized to the head 250 of the end user 252 (e.g., position, orientation, and volume data) for each virtual object, as well as global metadata (e.g., volume data globally set by the XR system 200 or end user 252), may be subsequently recorded, as will be further described below. The audio processor 246 and XR system 200 are also configured to enable voice communication between the user 252 and other parties, and/or to use the voice sound for voice-activated functions and commands.
[0084] As described herein, the audio processor 246 is further configured to process the speech audio signal (i.e., a near-field audio signal) captured by the primary microphone(s) 222 and to filter far-field noise from the speech audio signal using the microphone array 220 to produce a filtered speech audio signal (i.e., filtered near-field audio signal) having improved isolation of the speech sound with decreased background noise. The audio processor 246 includes a noise filtering software program (which may be in the form of software and/or firmware) which programs a microprocessor to perform a process for reducing far-field noise in the speech audio signal (i.e., a near-field audio signal) captured by the primary microphone(s) 222 using the microphone array 220.
[0085] Referring now to Fig. 4, one embodiment of a process 400 for reducing far-field noise in the speech audio signal (i.e., a near-field audio signal) captured by the primary microphone(s) 222
using the microphone array 220 is illustrated. At step 402, the audio processor acquires the respective near-field audio signals from one or more of the primary microphones 222. The process 400 may utilize only the first near-field audio signal from just the first primary microphone 222a, or the process may utilize both the first and second near-field audio signals from both the first and second primary microphones 222a, 222b and combine them into a near-field audio signal. By using near-field audio signals from multiple primary microphones 222, the near- field audio signals can be combined to provide some directionality to the near-field audio signal, such as a dipole pattern.
[0086] At step 404, the audio processor 246 acquires the respective reference audio signals from one or more of the reference microphones 224. The audio processor 246 may acquire and utilize in the process 400 only one of the reference audio signals from either the first or second reference microphones 224a, 224b, or both the first and second reference audio signals from both the first and second reference microphones 224a, 224b.
[0087] At step 406, the audio processor 246 uses the one or more reference audio signals to determine a far-field audio signal. Typically, a far-field is defined as being at least 1 meter away from the near-field, so the process 400 determines a far-field audio signal by making an estimate. In the case of using just a single reference audio signal, the far-field audio signal is simply the single reference audio signal. When using multiple reference audio signals, the audio processor 246 combines the reference audio signals. As an example, in the embodiment of Fig. 2, the far-field audio signal may be the difference between the first reference audio signal and the second reference audio signal, which assumes a dipole pattern in the arrangement of the first and second reference microphones 224a, 224b. Again, by using reference audio signals from multiple reference microphones 224, the reference audio signals can be combined to provide some
directionality to the far-field audio signal, such as a dipole audio signal. In an embodiment utilizing more than two reference microphones 224, the reference microphones 224 may be configured to utilize beamforming techniques such as Delay and Sum, Frost, or Minimum Variance Distortionless Response (MVDR) to determine a far-field audio signal.
[0088] At step 408, the audio processor 246 partitions the near-field audio signal and the far-field signal into a plurality of audio frequency bands. This may be accomplished by any suitable process, including using a WOLA analysis. In the example described below with respect to Figs. 5-10, the near-field audio signal and the far-field signal are partitioned into 42 frequency bands within a range of 0 - 25000 Hz. The number of audio frequency bands may be any suitable number, such as from 20-60 bands, 40-50 bands, greater than 20 bands, greater than 40 bands, or greater than 50 bands.
[0089] At step 410, the audio processor 246 calculates a near-field energy for the near-field audio signal in each audio frequency band and a far-field energy for the far-field audio signal in each audio frequency band. This results in a respective near-field energy value for each audio frequency band and a respective far-field energy value for each audio frequency band.
[0090] At step 412, the resultant near-field energy and far-field energy in each audio frequency band are used to calculate an energy ratio of the near-field energy to the far-field energy for each respective audio frequency band. In other words, the ratio of the near-field energy in a first audio frequency band to the far-field energy in the same first audio frequency band is calculated, the ratio of the near-field energy in a second audio frequency band to the far-field energy in the same second audio frequency band is calculated, and so-on for each of the audio frequency bands.
[0091] At step 414, the audio processor 246 determines whether the energy ratio for each audio frequency band is below a predetermined threshold for the respective audio frequency band. The
predetermined threshold sets a limit of the ratio of the near-field energy to the far-field energy for a respective audio frequency band. This step of the process determines the audio frequency bands having a relatively high energy ratio of the near-field energy to the far-field energy such that the near-field audio signal in that frequency band is considered to be significantly speech sound (i.e., target, desirable sound), and the audio frequency bands having a relatively low energy ratio of the near-field energy such that the near-field audio signal in that frequency band is considered to be mostly background noise or interference and not speech. As an example, the predetermined threshold may be unity, or near unity. The predetermined thresholds may be frequency band dependent, such that each audio frequency band has its own predetermined threshold, which may be different for each audio frequency band. In an alternative embodiment, the predetermined threshold may be the same for all of the audio frequency bands. The predetermined threshold(s) are tunable through empirical experimentation to determine predetermined threshold(s) providing the desired, or best, far-field filtering effect. Fig. 5 shows an example of the energy ratios for 42 frequency bands from a test case of speech sound captured by one or more primary microphones and ambient sound captured by two reference microphones. The graph of Figs. 5 represents a spatial filtering mask prior to thresholding and warping, as described below.
[0092] At step 416, for each audio frequency band having an energy ratio below the predetermined threshold, the audio processor 246 calculates a respective time-varying masking gain based upon the energy ratio for the respective audio frequency band. This is also called warping the spatial filtering mask. In one embodiment, the time-varying masking gain for each audio frequency band may be proportional to the amount the respective energy ratio is below the respective predetermined threshold. Hence, if the energy ratio is just slightly below (i.e., relatively close to) the predetermined threshold, then the time-varying masking gain is small as compared to the time-
varying masking gain for a frequency band having an energy ratio that is far below its respective threshold and having a greater time-varying masking gain (i.e., greater amount of reduction). In another embodiment, the time-varying masking gain for all of the audio frequency bands can be used to generate a spatial filtering mask.
[0093] The following equations represent one embodiment of the thresholding steps 414 and the warping step 416:
If Energy RatioThresholded > 0.5
[00 6] Energy RatioFinal(k) = EnergyRatioThresholded(k')-SupressionE:xp>
If Energy RatioThresholded <0.5
[0097] At step 418, for each audio frequency band having an energy ratio equal to or greater than the predetermined threshold, the near-field audio signal for such audio frequency band is passed through unchanged, as the filtered near field audio signal for such respective audio frequency bands.
[0098] The graph of Fig. 6 represents the spatial filter mask of Fig. 5 after thresholding and warping. It can be seen in Fig. 6 that the audio frequency bands having low energy ratios in Fig. 5 are significantly decreased in Fig. 6.
[0099] Step 420 is an optional step, and is included in this embodiment of the method 400. At step 420, the audio processor 246 performs an adaptive frequency smoothing of the time-varying masking gain for each audio frequency band based on an estimate of a signal to noise ratio of the near-field audio signal. Fig. 7 shows a graph of estimates of the signal to noise ratio for the near-
field audio signal in the test case of Fig. 5. In this embodiment, for low signal to noise ratio signals the weighting of neighboring sub-bands are weighted more heavily than the weighting of neighboring sub-bands for high signal to noise ratio signals. The adaptive frequency smoothing smooths in frequency the spatial masking gain thereby reducing musical noise artifacts. During periods of low signal to noise ratio, a higher degree of smoothing is utilized, as illustrated in Figs. 8A and 8B. Fig. 8A shows an example of a graph of adaptive frequency smoothing curves for a low signal to noise near-field audio signal. Each curve in the graph shows the smoothing curve for a respective frequency band. Fig. 8B shows an example of a graph of adaptive frequency smoothing curves for a high signal to noise near-field audio signal. Again, each curve in the graph shows the smoothing curve for a respective frequency band. It can be seen by comparing the adaptive frequency curves of Fig. 8A to Fig. 8B that for a low signal to noise ratio near-field audio signal, the neighboring frequency bands are more heavily weighted as compared to a high signal to noise ratio near-field audio signal.
[00100] At step 422, the audio processor 246 applies the time-varying masking gain for each audio frequency band to the near-field audio signal for each respective audio frequency band to generate a filtered near-field audio signal for each audio frequency band. This results in a filtered near- field audio signal in which the audio signal in audio frequency bands determined to include a significant far-field audio signal is reduced by the respective time-varying masking gain. Figs. 9 and 10 illustrate the filtering of the near-field audio signal in the test case of Fig. 5.
[00101] At step 424, the audio processor 246 synthesizes (i.e., combines) the filtered near-field audio signal for each audio frequency band thereby producing a synthesized near-field audio signal. The synthesized near-field audio signal corresponds to filtered near-field audio in which the near-field audio is isolated from the far-field noise. As illustrated in the example of Figs. 9 and
10, the background noise and interference have been decreased, while the near-field sounds remain prominent, thereby improving the perception of the near-field speech sound. Fig. 9 shows the near-field audio signal before noise reduction. Fig. 10 shows the near-field audio signal after noise reduction using the method 400. It can be seen in Fig. 9 that there is a broad-band interference or background noise at about 75dBA, while Fig. 10 shows that this broad-band interference at 75dBA is significantly reduced and the user speech signal is extracted with minimal spectral artifacts. [00102] Turning back to Fig. 2, the XR system 200 further comprises memory 260, and a recorder 262 configured for storing video and audio in the memory 260, which may be accessed for playback.
[00103] The control subsystem that performs the functions of the video processor 1244, audio processor 246, recorder 262, and an audio/video player (not shown) may take any of a large variety of forms, and may include a number of controllers, for instance one or more microcontrollers, microprocessors or central processing units (CPUs), digital signal processors, graphics processing units (GPUs), other integrated circuit controllers, such as application specific integrated circuits (ASICs), programmable gate arrays (PGAs), for instance, field PGAs (FPGAs), and/or programmable logic controllers (PLUs).
[00104] The functions of the video processor 244, audio processor 246, and recorder 262 may be respectively performed by single integrated devices, at least some of the functions of the video processor 244, audio processor 246, and recorder 262 may be combined into a single integrated device, or the functions of each of the video processor 244, audio processor 246 and recorder 262 may be distributed amongst several devices. For example, the video processor 244 may comprise a graphics processing unit (GPU) that acquires the video data of virtual objects from the three- dimensional database 242 and renders the synthetic video frames therefrom, and a central
processing unit (CPU) that acquires the video frames of real objects from the forward-facing camera(s) 230. Similarly, the audio processor 246 may comprise a digital signal processor (DSP) that processes the audio data acquired from a microphone subsystem and microphone array 220, and the CPU that processes the audio data. The recording functions of the recorder 262 may also be performed by the CPU.
[00105] Furthermore, the various processing components of the XR system 200 may be physically contained in a distributed subsystem. The XR system 200 may further comprises a remote processing module 300 and remote data repository 302 operatively coupled, such as by a wired lead or wireless connectivity 304, 306, to a local processing and data module 308, such that these remote modules 300, 302 are operatively coupled to each other and available as resources to the local processing and data module 308. For example, as illustrated in Fig. 3, the XR system 200 may comprise a local processing and data module 308 operatively coupled, such as by a wired lead or wireless connectivity , to components carried by the headset 205 worn on the head 250 of the end user 252 (e.g., the projection subsystem of the display subsystem 208, microphone array 220, audio processor 246, speakers 210, and cameras 230, 232). The local processing and data module 308 may include any one or more of the components and subsystems included in the control subsystem 201 of Fig. 2, such as the display subsystem 208, audio processor 246, etc. Any one or more of the GPU of the video processor 244 and CPU of the video processor 244 and/or audio processor 246 may be contained in the remote processing module 308, in order to limit the power consumption of the local processing and data module 308 which may be battery-powered, to improve portability, although in alternative embodiments, these components, or portions thereof may be contained in the local processing and data module 308. As shown in Fig. 3, local processing and data module 308 may be mounted in a body -pack configured to be removably
attached to the user’s body, such as a belt-pack having a belt-coupling for wearing on the user’s hip 310. Alternatively, the local processing and data module 308 may be fixedly attached to the frame structure 206, fixedly attached to a helmet or other headpiece worn by the user 252, embedded in headphones, removably attachable to the user’s torso, etc.
[00106] The local processing and data module 308 may comprise a power-efficient processor or controller, as well as digital memory, such as flash memory, both of which may be utilized to assist in the processing, caching, and storage of data captured from the sensors and/or acquired and/or processed using the remote processing module 300 and/or remote data repository 302, possibly for passage to the display subsystem 208 after such processing or retrieval. The remote processing module 300 may comprise one or more relatively powerful processors or controllers configured to analyze and process data and/or image information. The remote data repository 302 may comprise a relatively large-scale digital data storage facility, which may be available through the internet or other networking configuration in a “cloud” resource configuration. In an alternative embodiment, all data is stored and all computation is performed in the local processing and data module 308, allowing fully autonomous use from any remote modules.
[00107] The couplings 304, 306 between the various components described above may include one or more wired interfaces or ports for providing wires or optical communications, or one or more wireless interfaces or ports, such as via RF, microwave, and IR for providing wireless communications. In some implementations, all communications may be wired, while in other implementations all communications may be wireless, with the exception of optical fiber(s) used in the display subsystem 208. In still further implementations, the choice of wired and wireless communications may be different from that illustrated in Fig. 3. Thus, the particular choice of wired or wireless communications should not be considered limiting.
[00108] In the foregoing specification, the invention has been described with reference to specific embodiments thereof. It will, however, be evident that various modifications and changes may be made thereto without departing from the broader spirit and scope of the invention. For example, the above-described process flows are described with reference to a particular ordering of process actions. However, the ordering of many of the described process actions may be changed without affecting the scope or operation of the invention. The specification and drawings are, accordingly, to be regarded in an illustrative rather than restrictive sense.
Claims
1. A method for reducing far-field noise in a near-field audio signal using a microphone array, comprising: a) acquiring a near-field audio signal from a primary microphone of the microphone array; b) acquiring one or more reference audio signals from one or more respective reference microphones of the microphone array, and determining a far-field audio signal from the one or more reference audio signals; c) partitioning the near-field audio signal and the far-field signal into a plurality of audio frequency bands; d) calculating a near-field energy for the near-field audio signal in each audio frequency band and a far-field energy for the far-field audio signal in each audio frequency band; e) calculating an energy ratio of the near-field energy to the far-field energy for each audio frequency band; f) determining whether the energy ratio for each audio frequency band is below a predetermined threshold; g) for each audio frequency band having an energy ratio below the predetermined threshold, using the energy ratio for each audio frequency band to produce a time-varying masking gain for each audio frequency band; h) applying the time-varying masking gain for each audio frequency band to the near-field audio signal for each respective audio frequency band to generate a filtered near-field audio signal for each audio frequency band; and i) synthesizing the filtered near-field audio signal for each audio frequency band to produce a synthesized near-field audio signal.
2. The method of claim 1, further comprising: after step f), for each audio frequency band having an energy ratio equal to or greater than the predetermined threshold, passing through the near-field audio signal for such audio frequency band unchanged as the filtered near field audio signal for such audio frequency band.
3. The method of any of claims 1-2, wherein step i) comprises a weighted overlap-add (WOLA) synthesis of the filtered near-field audio signal for each audio frequency band to produce the synthesized near-field audio signal.
4. The method of any of claims 1-3, wherein step c) comprises: performing a weighted overlap-add (WOLA) analysis and sub-band partitioning of each of the near-field audio signal and far-field audio signal to partition the near-field audio signal and the far-field signal into the plurality of audio frequency bands.
5. The method of any of claims 1-4, further comprising: after step g), performing an adaptive frequency smoothing of the time varying gain for each audio frequency band based on an estimate of a signal to noise ratio of the near-field audio signal, wherein for low signal to noise ratio signals the weighting of neighboring sub-bands are weighted more heavily than the weighting of neighboring sub-bands for high signal to noise ratio signals.
6. The method of any of claims 1-5, wherein the far-field audio signal is generated from a combining a plurality of reference audio signals from a plurality of reference microphones.
7. The method of claim 6, wherein the far-field audio signal is generated by taking the difference between the respective audio signals of the two or more reference microphones.
8. The method of any of claims 1-7, wherein the time-varying masking gain for each audio frequency band is linearly related to the amount the respective energy ratio is below the predetermined threshold, such that the further below the predetermined threshold the greater the amount of time-varying masking gain.
9. The method of any of claims 1-7, wherein the time-varying masking gain for each audio frequency band is non-linearly related to the amount the respective energy ratio is below the
predetermined threshold, such that the further below the predetermined threshold the greater the amount of time-varying masking gain.
10. The method of any of claims 1-9, wherein: the near-field audio signal is a near-field speech signal of a user and the primary microphone is positioned proximate a user’s mouth; the far-field signal is a background signal and the one or more reference microphones are positioned further away from the user’s mouth than the primary microphone; and the method isolates the user’s speech by attenuating the far-field ambient noise on an audio frequency band by audio frequency band basis.
11. The method of any of claims 1-10, wherein the predetermined threshold is frequency band dependent, such that each audio frequency band has a predetermined threshold, and the predetermined threshold are different for at least two of the audio frequency bands.
11. An extended reality (XR) system, comprising: an XR headset configured to be worn on a head of a subject, the XR headset comprising: a support structure configured to be worn on the head of the subject; a display system carried by the support structure, the display system for displaying virtual images generated by a video processor of the XR system; a microphone array, including: a primary microphone configured to be positioned proximate a user’s mouth, the primary microphone configured to sense nearfield audio sound and generate a near-field audio signal corresponding to the near-field audio sound; and one or more reference microphones configured to be positioned further away from the user’s mouth than the primary microphone to sense far-field audio sound, the reference microphones configured to generate a far-field audio signal corresponding to the far-field audio sound; an audio processor, the audio processor operably coupled to the primary microphone and the one or more reference microphones to acquire the near-field audio signal and the far-field
audio signal, the audio processor configured to perform a process for reducing far-field noise from microphone array, the process comprising: a) acquiring the near-field audio signal from the primary microphone of the microphone array; b) acquiring one or more reference audio signals from the one or more reference microphones of the microphone array, and determining a far-field audio signal from the one or more reference audio signals; c) partitioning the near-field audio signal and the far-field signal into a plurality of audio frequency bands; d) calculating a near-field energy for the near-field audio signal in each audio frequency band and a far-field energy for the far-field audio signal in each audio frequency band; e) calculating an energy ratio of the near-field energy to the far-field energy for each audio frequency band; f) determining whether the energy ratio for each audio frequency band is below a predetermined threshold; g) for each audio frequency band having an energy ratio below the predetermined threshold, using the energy ratio for each audio frequency band to produce a timevarying masking gain for each audio frequency band; h) applying the time-varying masking gain for each audio frequency band to the near-field audio signal for each respective audio frequency band to generate a filtered near- field audio signal for each audio frequency band; and i) synthesizing the filtered near-field audio signal for each audio frequency band to produce a synthesized near-field audio signal.
12. The system of claim 11, wherein: wherein the support structure is an XR headset; the primary microphone is carried on the XR headset in a front-side of the headset to position the primary microphone in front of the user’s face; and the one or more reference microphones are carried on the headset on either side of the headset to position the one or more reference microphones toward a side of the user’s head.
13. The system of any of claims 11-12, wherein the audio processor is housed in a body-pack configured to be removably attached to a user’s body.
14. The system of any of claim 11-13, wherein the process further comprises: after step f), for each audio frequency band having an energy ratio equal to or greater than the predetermined threshold, passing through the near-field audio signal for such audio frequency band unchanged as the filtered near field audio signal for such audio frequency band.
15. The system of any of claims 11-14, wherein step i) comprises a WOLA synthesis of the filtered near-field audio signal for each audio frequency band to produce the synthesized near- field audio signal.
16. The system of any of claims 11-15, wherein step c) comprises: performing a weighted overlap-add analysis and sub-band partitioning of each of the near-field audio signal and far-field audio signal to partition the near-field audio signal and the far-field signal into the plurality of audio frequency bands.
17. The system of any of claims 11-16, further comprising: after step g), performing an adaptive frequency smoothing of the time varying gain for each audio frequency band based on an estimate of a signal to noise ratio of the near-field audio signal, wherein for low signal to noise ratio signals the weighting of neighboring sub-bands are weighted more heavily than the weighting of neighboring sub-bands for high signal to noise ratio signals.
18. The system of any of claims 11-17, wherein the far-field audio signal is generated from a combining a plurality of reference audio signals from a plurality of reference microphones.
19. The system of claim 18, wherein the far-field audio signal is generated by taking the difference between the respective audio signals of the two or more reference microphones.
20. The system of any of claims 11-19, wherein the time-varying masking gain for each audio frequency band is linearly related to the amount the respective energy ratio is below the predetermined threshold, such that the further below the predetermined threshold the greater the amount of time-varying masking gain.
21. The system of any of claims 11-20, wherein the time-varying masking gain for each audio frequency band is non-linearly related to the amount the respective energy ratio is below the predetermined threshold, such that the further below the predetermined threshold the greater the amount of time-varying masking gain.
22. The system of any of claims 11-21, wherein: the near-field audio signal is a near-field speech signal of a user and the primary microphone is positioned proximate a user’s mouth; the far-field signal is a background signal and the one or more microphones are positioned further away from the user’s mouth than the primary microphone; and the method isolates the user’s speech by attenuating the far-field ambient noise on an audio frequency band by audio frequency band basis.
23. A non-transitory computer-readable medium having software instructions stored thereon, the software instructions executable by a computer processor of an audio processor to perform a process for reducing far-field noise using a microphone array, the process comprising: a) acquiring a near-field audio signal from a primary microphone of the microphone array; b) acquiring one or more reference audio signals from one or more reference microphones of the microphone array, and determining a far-field audio signal from the one or more reference audio signals; c) partitioning the near-field audio signal and the far-field signal into a plurality of audio frequency bands;
d) calculating a near-field energy for the near-field audio signal in each audio frequency band and a far-field energy for the far-field audio signal in each audio frequency band; e) calculating an energy ratio of the near-field energy to the far-field energy for each audio frequency band; f) determining whether the energy ratio for each audio frequency band is below a predetermined threshold; g) for each audio frequency band having an energy ratio below the predetermined threshold, using the energy ratio for each audio frequency band to produce a time-varying masking gain for each audio frequency band; h) applying the time-varying masking gain for each audio frequency band to the near-field audio signal for each respective audio frequency band to generate a filtered near-field audio signal for each audio frequency band; and i) synthesizing the filtered near-field audio signal for each audio frequency band to produce a synthesized near-field audio signal.
24. The computer-readable medium of claim 23, further comprising: after step f), for each audio frequency band having an energy ratio equal to or greater than the predetermined threshold, passing through the near-field audio signal for such audio frequency band unchanged as the filtered near field audio signal for such audio frequency band.
25. The computer-readable medium of any of claims 23-24, wherein step i) comprises a WOLA synthesis of the filtered near-field audio signal for each audio frequency band to produce the synthesized near-field audio signal.
26. The computer-readable medium of any of claims 23-25, wherein step c) comprises: performing a weighted overlap-add analysis and sub-band partitioning of each of the near-field audio signal and far-field audio signal to partition the near-field audio signal and the far-field signal into the plurality of audio frequency bands.
27. The computer-readable medium of any of claims 23-26, further comprising:
after step g), performing an adaptive frequency smoothing of the time varying gain for each audio frequency band based on an estimate of a signal to noise ratio of the near-field audio signal, wherein for low signal to noise ratio signals the weighting of neighboring sub-bands are weighted more heavily than the weighting of neighboring sub-bands for high signal to noise ratio signals.
28. The computer-readable medium of any of claims 23-27, wherein the far-field audio signal is generated from a combining a plurality of reference audio signals from a plurality of reference microphones.
29. The computer-readable medium of claim 28, wherein the far-field audio signal is generated by taking the difference between the respective audio signals of the two or more reference microphones.
30. The computer-readable medium of any of claims 23-29, wherein the time-varying masking gain for each audio frequency band is linearly related to the amount the respective energy ratio is below the predetermined threshold, such that the further below the predetermined threshold the greater the amount of time-varying masking gain.
31. The computer-readable medium of any of claims 23-30, wherein the time-varying masking gain for each audio frequency band is non-linearly related to the amount the respective energy ratio is below the predetermined threshold, such that the further below the predetermined threshold the greater the amount of time-varying masking gain.
32. The computer-readable medium of any of claims 23-31, wherein: the near-field audio signal is a near-field speech signal of a user and the primary microphone is positioned proximate a user’s mouth; the far-field signal is a background signal and the one or more microphones are positioned further away from the user’s mouth than the primary microphone; and
the method isolates the user’s speech by attenuating the far-field ambient noise on an audio frequency band by audio frequency band basis.
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US202363453022P | 2023-03-17 | 2023-03-17 | |
| PCT/US2024/019790 WO2024196673A1 (en) | 2023-03-17 | 2024-03-13 | Far-field noise reduction via spatial filtering using a microphone array |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP4681198A1 true EP4681198A1 (en) | 2026-01-21 |
Family
ID=92842574
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP24775389.0A Pending EP4681198A1 (en) | 2023-03-17 | 2024-03-13 | Far-field noise reduction via spatial filtering using a microphone array |
Country Status (4)
| Country | Link |
|---|---|
| EP (1) | EP4681198A1 (en) |
| JP (1) | JP2026510853A (en) |
| CN (1) | CN120917513A (en) |
| WO (1) | WO2024196673A1 (en) |
Family Cites Families (7)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| EP3267697A1 (en) * | 2016-07-06 | 2018-01-10 | Oticon A/s | Direction of arrival estimation in miniature devices using a sound sensor array |
| GB201709846D0 (en) * | 2017-06-20 | 2017-08-02 | Nokia Technologies Oy | Processing audio signals |
| EP4035418A2 (en) * | 2019-09-23 | 2022-08-03 | Dolby Laboratories Licensing Corporation | Hybrid near/far-field speaker virtualization |
| CN112889109B (en) * | 2019-09-30 | 2023-09-29 | 深圳市韶音科技有限公司 | Systems and methods for noise reduction using subband noise reduction technology |
| GB2598960A (en) * | 2020-09-22 | 2022-03-23 | Nokia Technologies Oy | Parametric spatial audio rendering with near-field effect |
| US20220180885A1 (en) * | 2020-12-04 | 2022-06-09 | Facebook Technologies, Llc | Audio system including for near field and far field enhancement that uses a contact transducer |
| US12126971B2 (en) * | 2020-12-23 | 2024-10-22 | Intel Corporation | Acoustic signal processing adaptive to user-to-microphone distances |
-
2024
- 2024-03-13 EP EP24775389.0A patent/EP4681198A1/en active Pending
- 2024-03-13 WO PCT/US2024/019790 patent/WO2024196673A1/en not_active Ceased
- 2024-03-13 CN CN202480019374.XA patent/CN120917513A/en active Pending
- 2024-03-13 JP JP2025553719A patent/JP2026510853A/en active Pending
Also Published As
| Publication number | Publication date |
|---|---|
| WO2024196673A1 (en) | 2024-09-26 |
| JP2026510853A (en) | 2026-04-10 |
| CN120917513A (en) | 2025-11-07 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| JP7275227B2 (en) | Recording virtual and real objects in mixed reality devices | |
| JP7165215B2 (en) | Virtual Reality, Augmented Reality, and Mixed Reality Systems with Spatialized Audio | |
| US20250225971A1 (en) | Microphone array geometry | |
| JP2009536406A (en) | How to give emotional features to computer-generated avatars during gameplay | |
| EP4681198A1 (en) | Far-field noise reduction via spatial filtering using a microphone array | |
| JP7402185B2 (en) | Low frequency interchannel coherence control | |
| US20240406666A1 (en) | Sound field capture with headpose compensation |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE |
|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE |
|
| 17P | Request for examination filed |
Effective date: 20251017 |
|
| AK | Designated contracting states |
Kind code of ref document: A1 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR |