EP4599601A1 - Internal-noise-source filtering for active acoustic sensing - Google Patents

Internal-noise-source filtering for active acoustic sensing

Info

Publication number
EP4599601A1
EP4599601A1 EP23848366.3A EP23848366A EP4599601A1 EP 4599601 A1 EP4599601 A1 EP 4599601A1 EP 23848366 A EP23848366 A EP 23848366A EP 4599601 A1 EP4599601 A1 EP 4599601A1
Authority
EP
European Patent Office
Prior art keywords
signal
hearable
noise
ultrasound
user
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Granted
Application number
EP23848366.3A
Other languages
German (de)
French (fr)
Other versions
EP4599601B1 (en
Inventor
Xiaoran FAN
Trausti MUNDSSON
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Google LLC
Original Assignee
Google Technology Holdings LLC
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Google Technology Holdings LLC filed Critical Google Technology Holdings LLC
Publication of EP4599601A1 publication Critical patent/EP4599601A1/en
Application granted granted Critical
Publication of EP4599601B1 publication Critical patent/EP4599601B1/en
Active legal-status Critical Current
Anticipated expiration legal-status Critical

Links

Classifications

    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04RLOUDSPEAKERS, MICROPHONES, GRAMOPHONE PICK-UPS OR LIKE ACOUSTIC ELECTROMECHANICAL TRANSDUCERS; ELECTRIC HEARING AIDS; PUBLIC ADDRESS SYSTEMS
    • H04R1/00Details of transducers, loudspeakers or microphones
    • H04R1/10Earpieces; Attachments therefor ; Earphones; Monophonic headphones
    • H04R1/1083Reduction of ambient noise
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L25/00Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
    • G10L25/78Detection of presence or absence of voice signals
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04RLOUDSPEAKERS, MICROPHONES, GRAMOPHONE PICK-UPS OR LIKE ACOUSTIC ELECTROMECHANICAL TRANSDUCERS; ELECTRIC HEARING AIDS; PUBLIC ADDRESS SYSTEMS
    • H04R1/00Details of transducers, loudspeakers or microphones
    • H04R1/10Earpieces; Attachments therefor ; Earphones; Monophonic headphones
    • H04R1/1016Earpieces of the intra-aural type
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04RLOUDSPEAKERS, MICROPHONES, GRAMOPHONE PICK-UPS OR LIKE ACOUSTIC ELECTROMECHANICAL TRANSDUCERS; ELECTRIC HEARING AIDS; PUBLIC ADDRESS SYSTEMS
    • H04R1/00Details of transducers, loudspeakers or microphones
    • H04R1/10Earpieces; Attachments therefor ; Earphones; Monophonic headphones
    • H04R1/1041Mechanical or electronic switches, or control elements
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04RLOUDSPEAKERS, MICROPHONES, GRAMOPHONE PICK-UPS OR LIKE ACOUSTIC ELECTROMECHANICAL TRANSDUCERS; ELECTRIC HEARING AIDS; PUBLIC ADDRESS SYSTEMS
    • H04R2460/00Details of hearing devices, i.e. of ear- or headphones covered by H04R1/10 or H04R5/033 but not provided for in any of their subgroups, or of hearing aids covered by H04R25/00 but not provided for in any of its subgroups
    • H04R2460/01Hearing devices using active noise cancellation

Definitions

  • Wireless technology has become prevalent in everyday life, making communication and data readily accessible to users.
  • wireless hearables examples of which include wireless earbuds and wireless headphones.
  • Wireless hearables have allowed users freedom of movement while listening to audio content from music, audio books, podcasts, and videos.
  • wireless hearables With the prevalence of wireless hearables, there is a market for adding additional features to existing hearables without introducing hardware changes.
  • intemal-noise-source filtering for active acoustic sensing.
  • interference caused by a hearable performing other operations e.g., rendering audio content, performing active-noise cancellation, or operating in accordance with a transparency mode
  • This performance improvement enables audioplethysmography to be performed while the hearable performs these other operations.
  • it improves the ability of audioplethysmography to be used for voice processing, which can include voice activity detection, speech recognition, and/or conversation detection.
  • the method includes transmitting, during a first time period, an audible signal that propagates within at least a portion of an ear canal of a user.
  • the method also includes transmitting, during the first time period, an ultrasound transmit signal that propagates within at least a portion of the ear canal of the user.
  • the method additionally includes receiving, during the first time period, an ultrasound receive signal.
  • the ultrasound receive signal represents a version of the ultrasound transmit signal with one or more characteristics modified based on the propagation within the ear canal and based on a vocalization made by the user during the first time period.
  • the received ultrasound receive signal includes an internal noise component caused by interference generated by the rendering of the audible signal.
  • the method further includes generating a denoised signal by filtering the internal noise component within the received ultrasound receive signal based on a version of the audible signal.
  • the method also includes detecting the vocalization based on the denoised signal.
  • the version of the audible signal can represent an electrical version of the audible signal and/or a digital version of the audible signal.
  • aspects described below include a device with at least one transducer and at least one processor.
  • the device is configured to perform, using the at least one transducer and the at least one processor, any one of the methods described herein.
  • FIG. 10 illustrates an example flow diagram for operating a hearable
  • FIG. 11 illustrates an example scheme implemented by a calibration module of a hearable
  • FIG. 13 illustrates an example implementation of a measurement module for performing intemal-noise-source filtering
  • FIG. 16 illustrates an example method for performing intemal-noise-source filtering
  • FIG. 17 illustrates another example method for performing intemal-noise-source filtering
  • FIG. 18 illustrates an example computing system embodying, or in which techniques maybe implemented that enable use of, intemal-noise-source filtering for active acoustic sensing.
  • a user may use an electronic device to get daily weather and traffic information, control a temperature of a home, answer a doorbell, turn on or off a light, and/or play background music. Interacting with some electronic devices, however, can be cumbersome and inefficient.
  • An electronic device for instance, can have a physical user interface that may require a user to navigate through one or more prompts by physically touching the electronic device. In this case, the user has to devote attention away from other pnmary tasks to interact with the electronic device, which can be inconvenient and disruptive.
  • some electronic devices support voice control, which enables a user to interact with the electronic device in a non-physical and less cognitively demanding way- compared to other interfaces that require physical touch and/or the user’s visual attention.
  • voice control the electronic device seamlessly exists in the surrounding environment and provides the user access to information and services while the user performs a primary task, such as cooking, cleaning, driving, talking with people, or reading a book.
  • Audioplethysmography is an active acoustic method capable of sensing subtle physiologically -related changes observable at a user’s outer and middle ear. Instead of relying on other auxiliary sensors, such as optical or electrical sensors, audioplethysmography involves transmitting and receiving acoustic signals that at least partially propagate within auser’s ear canal. To perform audioplethysmography, the hearable forms at least a partial seal in or around the user’s outer ear.
  • This seal enables formation of an acoustic circuit, which includes the seal, the hearable, the ear canal, and an ear drum of the ear.
  • the hearable can recognize changes in the acoustic circuit to detect and/or recognize vocalizations made by the user.
  • Vocalizations can include any sound that is produced using the user’s lung’s, vocal cords, and/or mouth.
  • Example types of vocalizations can involve the user speaking, whispering, shouting, humming, whistling, singing, or making other utterances.
  • the hearable can utilize audioplethysmography to support a variety of different features, including a voice user interface (VUI) and/or multi-factor voice authentication.
  • VUI voice user interface
  • signals generated for rendering audio content, providing active noise cancellation, and/or for operating in accordance with a transparency mode can interfere with an ultrasound signal that is received for audioplethysmography. This interference can make it challenging for the hearable to use audioplethysmography for voice processing.
  • intemal-noise-source filtering for active acoustic sensing.
  • interference caused by the hearable performing other operations e.g., rendering audio content, performing active-noise cancellation, or operating in accordance with a transparency mode
  • This performance improvement enables audioplethysmography to be performed while the hearable performs these other operations.
  • it improves the ability of audioplethysmography to be used for voice processing, which can include voice activity detection, speech recognition, and/or conversation detection.
  • FIG. 1-1 is an illustration of an example environment 100 in which active acoustic sensing can be implemented.
  • a hearable 102 is connected to a computing device 104 using a physical or wireless interface.
  • the hearable 102 is a device that can play audible content provided by the computing device 104 and direct the audible content into a user 106’s ear 108.
  • the hearable 102 operates together with the computing device 104.
  • the hearable 102 can operate or be implemented as a stand-alone device.
  • the computing device 104 can include other types of devices, including those described with respect to FIG. 6.
  • the hearable 102 is capable of performing audioplethysmography 110, which is an active acoustic method of sensing that occurs at the ear 108.
  • the hearable 102 can perform this sensing without the use of other auxiliary sensors, such as an optical sensor or an electrical sensor.
  • the hearable 102 can utilize voice processing to detect and/or recognize vocalizations made by the user 106.
  • Example types of voice processing can include voice activity’ detection 112, speech recognition 114, and/or conversation detection 116.
  • conversation detection 116 can be used to control an operation of the hearable 102 and/orthe computing device 104.
  • audio content can be paused or resumed based on whether a conversation is detected or not.
  • a volume of the audio content is decreased or increased based on whether a conversation is detected or not.
  • active noise cancellation can be temporarily halted and the hearable 102 can operate in accordance with a transparency mode based on the conversation detection 116 detecting a conversation.
  • active noise cancellation can resume and the hearable 102 can exit out of the transparency mode based on the conversation detection 116 detecting the absence of a conversation.
  • Audioplethysmography 110 To use audioplethysmography 110, the user 106 positions the hearable 102 in a manner that creates at least a partial seal 120 around or in the ear 108. Some parts of the ear 108 are shown in FIG. 1-1, including the ear canal 118 and an ear drum 122 (or tympanic membrane). Due to the seal 120, the hearable 102, the ear canal 118, and the ear drum 122 couple together to form an acoustic circuit. Audioplethysmography 110 involves, at least in part, measuring properties associated with this acoustic circuit. The properties of the acoustic circuit can change due to a variety of different situations or actions.
  • Example changes to the physical structure include a change in a geometric shape of the ear canal 118 and/or a change in a volume of the ear canal 118. This change can be caused, at least in part, by bone conduction and/or a pressure wave associated with a vocalization made by the user 106.
  • FIG. 2 illustrates example environments 200-1, 200-2, and 200-3 in which internal -noise- source filtering 202 can be implemented for active acoustic sensing.
  • the hearable 102 renders audio content 204 for the user 106.
  • the audio content 204 represents any type of sound that is produced by the hearable 102, such as music, a ringtone, an alarm, a caller’s voice, and so forth.
  • the active-noise-cancellation circuitry 316 performs active noise cancellation 206, as shown in the environment 200-2. To perform active noise cancellation 206. the active-noise- cancellation circuitry 316 generates an anti -noise signal 324. The anti-noise signal 324 represents another type of noise signal 308, which can interfere with the received signal 306 and contribute to the internal noise component 310.
  • the transparency mode circuitry 318 enables the hearable 102 to operate in accordance with the transparency mode, as shown in the environment 200-3. For the transparency mode, the transparency mode circuitry 7 318 generates a transparency -mode signal 326, which includes sounds from the external environment. The transparency-mode signal 326 represents another type of noise signal 308, which can interfere with the receive signal 306 and contribute to the rntemal noise component 310.
  • the hearable 102 receives at least one bone-conduction signal 408.
  • the bone-conduction signal 408 represents sound that travels, via bone conduction, to the user 106’s ear 108.
  • the bone-conduction signal 408 also includes the voice component 404. In most circumstances, the bone-conduction signal 408 does not include the noise component 406.
  • the hearable 102 may be designed to minimize interference between the over-the-air signal 402 and the ultrasound signal 410.
  • the hearable 102 can utilize different microphones to receive these signals.
  • the hearable 102 can directly process the ultrasound signal 410 to detect the vocalization.
  • the hearable 102 utilizes a same microphone to receive the over-the- air signal 402 and the ultrasound signal 410. While this may be beneficial for meeting size constraints of the hearable 102, it can cause a version of the noise component 406 to be present within the ultrasound signal 410.
  • FIG. 5 illustrates an example operation of the microphone 302 of the hearable 102.
  • the microphone 302 receives the over-the-air signal 402, the bone-conduction signal 408, and the ultrasound signal 410.
  • the microphone 302 can include a filter module, which can generate separate signals associated with different frequency ranges.
  • the microphone 302 generates a received audible signal 504 and a received ultrasound signal 506, which are further depicted in a graph 508 at the bottom of FIG. 5. These signals can be downconverted to baseband frequencies.
  • the received audible signal 504 and the received ultrasound signal 506 represent electrical signals that can be processed by other components of the hearable 102.
  • the received audible signal 504 represents a convolution of the over-the-air signal 402 and the bone-conduction signal 408. As such, the received audible signal 504 includes the voice component 404 associated with the vocalization and the noise component 406. The voice component 404 is provided by both the over-the-air signal 402 and the bone-conduction signal 408.
  • the received audible signal 504 can be represented by Equation 1:
  • the computing device 104 can also include a network interface 612 for communicating data over wired, wireless, or optical networks.
  • the network interface 612 may communicate data over a local-area-network (LAN), a wireless local-area-network (WLAN), a personal -area-network (PAN), a wire-area-network (WAN), an intranet, the Internet, a peer-to- peer network, point-to-point network, a mesh network, Bluetooth®, and the like.
  • the computing device 104 may also include the display 614.
  • the hearable 102 can be integrated within the computing device 104, or can connect physically or wirelessly to the computing device 104. The hearable 102 is further described with respect to FIG. 7.
  • FIG. 7 illustrates an example hearable 102.
  • the hearable 102 is illustrated with various non-limiting example devices, including wireless earbuds 702-1, wired earbuds 702-2, and headphones 702-3.
  • the earbuds 702-1 and 702-2 are a type of in-ear device that fits into the ear canal 118.
  • Each earbud 702-1 or 702-2 can represent a hearable 102.
  • Headphones 702-3 can rest on top of or over the ears 108.
  • the headphones 702-3 can represent closed-back headphones, open-back headphones, on-ear headphones, or over-ear headphones.
  • Each headphone 702-2 includes two hearables 102, which are physically packaged together. In general, there is one hearable 102 for each ear 108.
  • the hearable 102 includes at least one transducer 706 that can convert electrical signals into sound waves.
  • the transducer 706 can also detect and convert sound waves into electrical signals.
  • These sound waves may include ultrasonic frequencies and/or audible frequencies, either of which may be used for audioplethysmography 110.
  • a frequency spectrum e.g., range of frequencies
  • a frequency spectrum that the transducer 706 uses to generate an acoustic signal can include frequencies from a low-end of the audible range to ahigh-end of the ultrasonic range, e.g., between 20 hertz (Hz) to 2 megahertz (MHz).
  • frequency spectrums for audioplethysmography 110 can encompass frequencies between 20 Hz and 20 kilohertz (kHz), between 20 kHz and 2 MHz, between 20 and 96 kHz, between 20 and 60 kHz, or betw een 30 and 40 kHz.
  • the transducer 706 can be implemented with a bistatic topology, which includes multiple transducers that are physically separate.
  • a first transducer converts the electrical signal into sound waves (e.g., transmits acoustic signals)
  • a second transducer converts sound waves into an electrical signal (e.g., receives the acoustic signals).
  • An example bistatic topology can be implemented using at least one speaker 708 and at least one microphone 710.
  • the speaker 708 can direct ultrasound signals towards the ear canal 118, and the microphone 710 is responsive to receiving ultrasound signals from the direction associated with the ear canal 118.
  • the hearable 102 includes another microphone 710 that is directed away from the ear canal 118 towards an external environment (e.g., oriented away from the ear canal 118). This other microphone can be used to receive the over-the-air signal 402.
  • the hearable 102 includes at least one analog circuit 712. which includes circuitry and logic for conditioning electrical signals in an analog domain.
  • the analog circuit 712 can include analog-to-digital converters, digital-to-analog converters, amplifiers, filters, mixers, and swatches for generating and modifying electrical signals.
  • the analog circuit 712 includes other hardware circuitry associated with the speaker 708 or microphone 710.
  • the hearable 102 also includes at least one system processor 714 and at least one system medium 716 (e.g., one or more computer-readable storage media).
  • the system medium 716 includes a pre-processing module 718 and a measurement module 720.
  • the system medium 716 also optionally includes a calibration module 722.
  • the pre-processing module 718, the measurement module 720, and the calibration module 722 can be implemented using hardware, software, firmware, or a combination thereof.
  • the system processor 714 implements the pre-processing module 718, the measurement module 720, and the calibration module 722.
  • the computer processor 602 of the computing device 104 can implement at least a portion of the pre-processing module 718, the measurement module 720, and/or the calibration module 722.
  • the hearable 102 can communicate digital samples of the acoustic signals to the computing device 104 using the communication interface 704.
  • Operations of the pre-processing module 718, the measurement module 720, and the calibration module 722 are further described with respect to FIGs. 11 to 12.
  • Aspects of intemal- noise-source filtering 202 can be performed, at least partially, by the measurement module 720, as further described with respect to FIGs. 13 and 14.
  • the measurement module 720 can also perform aspects of voice activity detection 112, speech recognition 114, and/or conversation detection 1 16 using active acoustic sensing, as further described with respect to FIG. 13.
  • Some hearables 102 include the active-noise-cancellation circuitry' 316, which enables the hearables 102 to reduce background or environmental noise.
  • the microphone 710 used for audioplethysmography 110 can be implemented using a feedback microphone of the active-noise-cancellation circuitry 316.
  • the feedback microphone provides feedback information regarding the performance of the active noise cancellation 206.
  • the feedback microphone receives an ultrasound signal 410, which is provided to the pre-processing module 718. In some situations, active noise cancellation 206 and audioplethysmography 110 are performed simultaneously using the feedback microphone.
  • the ultrasound signal 410 received by the feedback microphone can be provided to the pre-processing module 718 and the feedback signal for active noise cancellation 206 can be provided to the active-noise-cancellation circuitry' 316.
  • the microphone 710 is implemented using a feedforward microphone of the active-noise-cancellation circuitry 316.
  • some hearables 102 include the transparency -mode circuitry' 318, which enables the hearables 102 to amplify background or environmental sounds.
  • the transparency -mode circuitry 318 can include at least one amplifier in some implementations.
  • the microphone 710 that is used for audioplethysmography 110 can also be used for the transparency mode 214. More specifically, the ultrasound signal 410 received by the microphone 710 can be provided to the preprocessing module 718 and the sensed signal for the transparency mode 214 can be provided to the transparency-mode circuitry 318.
  • the active-noise-cancellation circuitry 316, the transparency-mode circuitry 318, and/or the speaker 708 represent example types of circuitry 304 that can unintentionally act as an internal noise source 312 during audioplethysmography 110.
  • the system medium 716 can also include a voice user interface 608 and/or a voice authenticator 610.
  • the voice user interface 608 enables the user 1 6 to use voice controls to control an operation of the hearable 102.
  • the voice authenticator 610 can authenticate the user 106 and enable the voice user interface 608 for the hearable 102.
  • Different types of audioplethysmography 110 are further described with respect to FIG. 8.
  • the first hearable 102-1 uses the speaker 708 to transmit a first ultrasound transmit 802-1, which propagates within at least a portion of the user 106’s right ear canal 118.
  • the first hearable 102-1 uses the microphone 710 to receive a first ultrasound receive signal 804-1.
  • the first ultrasound receive signal 804-1 represents a version of the first ultrasound transmit signal 802-1 that is modified, at least in part, by the acoustic circuit associated with the right ear canal 118. This modification can change an amplitude, phase, and/or frequency of the first ultrasound receive signal 804-1 relative to the first ultrasound transmit signal 802-1.
  • the two hearables 102-1 and 102-2 perform two-ear audioplethysmography 110.
  • at least one of the hearables 102 e.g., the first hearable 102-1
  • the other hearables 102 e.g., the second hearable 102-2
  • the hearables 102-1 and 102-2 operate together in a bistatic manner during the same time period.
  • the third ultrasound receive signal 804-3 represents a version of the third ultrasound transmit signal 802-3 that is modified by the acoustic circuit associated with the right ear canal 118, modified by the acoustic channel associated with the user 106's face, and modified by the acoustic circuit associated with the left ear canal 118.
  • This modification can change an amplitude, phase, and/or frequency of the third ultrasound receive signal 804-3 relative to the third ultrasound transmit signal 802-3.
  • the hearable 102-2 measures the time-of-flight (ToF) associated with the propagation from the first hearable 102-1 to the second hearable 102-2.
  • ToF time-of-flight
  • a combination of single-ear and two-ear audioplethysmography 110 are applied to further improve measurement confidence.
  • the pre-processing module 718 performs frequency downconversion and demodulation to generate at least one pre-processed signal 910 based on the digital transmit signal 906 and the digital receive signal 908.
  • the pre-processing module 718 can also apply filtering to generate the pre-processed signal 910.
  • the calibration module 722 processes the pre- processed signal 910 to detennine the selected tones 904-1 to 904-N.
  • the selected tones 904-1 to 904-N can improve performance of audioplethysmography 110 during the measurement procedure.
  • the calibration module 722 communicates the selected tones 904-1 to 904-N to the speaker 708 using a control signal.
  • the speaker 708 accepts the control signal that identifies the selected tones 904-1 to 904-N and can transmit a subsequent ultrasound transmit signal 802 for the measurement procedure using the selected tones 904-1 to 904-N.
  • the measurement module 720 can perform aspects of intemal-noise-source filtering 202 using the pre-processed signal 910 and the noise signal 308.
  • the measurement module 720 can also perform aspects of voice activity 7 detection 112, speech recognition 114, and/or conversation detection 116 to generate voice data 912.
  • the voice data 912 can also be referred to as speech data or audioplethysmography data.
  • the measurement module 720 can also utilize the received audible signal 504 provided by the microphone 710 to further process the pre-processed signal 910 for voice activity 7 detection 112, speech recognition 114, and/or conversation detection 116.
  • FIG. 10 illustrates an example flow diagram 1000 for operating a hearable 102.
  • the hearable 102 can optionally perform a calibration procedure at 1002 using the calibration module 722.
  • the calibration procedure can determine appropriate characteristics (e.g., waveform or signal characteristics) of ultrasound transmit signals 802 to improve audioplethysmography 110 (e.g., to enhance the performance of voice activity' detection 112).
  • the calibration procedure enables audioplethysmography 110 to take into account the wear of the hearable 102 (e.g., the position of the hearable 102 relative to the ear canal 118) and the physical structure of the ear canal 118 to determine a transmission frequency that can increase sensitivity 7 .
  • the hearable 102 can dynamically adjust the transmission frequency (e.g., one or more carrier frequencies) each time the seal 120 is formed (e.g., based on the wear of the hearable 102) and based on the unique physical structure of the ear 108.
  • the hearables 102 on different ears 108 may operate with one or more different ultrasound frequencies. Steps of the calibration procedure are further described below.
  • the ultrasound transmit signal 902 has seven tones 902 (e.g., AT equals 7). In some cases, the tones 902 are evenly distributed across an interval. For example, the tones 902 can be in 1 kHz increments between 32 kHz and 38 kHz (e.g., at approximately 32, 33, 34, 35, 36, 37. and 38 kHz). The term “approximately” means that the tones 902 can be within 5% of a given value or less (e.g., within 3%, 2%, or 1% of the given value).
  • An amplitude of the ultrasound transmit signal 802 can be approximately the same across the tones 902-1 to 902-M. In this manner, power is evenly distributed across each tone 902.
  • the quantity of tones 902 (e.g., M) can be determined based on an output power of the speaker 708. Increasing the quantity of tones 902 can increase a likelihood that the hearable 102 can support voice activity 7 detection 112 across various conditions including user wear and a physical structure of the user 106’s ear canal 118. However, an amplitude of the ultrasound transmit signal 802 can be limited across these tones 902 based on the output power of the speaker 708. Thus, the quantity of tones 902 can be optimized based on an amount of output power that is available for audioplethysmography 110.
  • the calibration procedure selects one or more tones 904-1 to 904-N to be used for a measurement procedure based on one or more modified characteristics of the ultrasound receive signal 804.
  • the process for selecting the tones 904 is further described with respect to FIG. 11.
  • the calibration procedure determines that the selected tones 904 improve a signal-to- noise ratio for audioplethysmography 110 (or more specifically for voice activity detection 1 12).
  • the hearable 102 performs a measurement procedure using the measurement module 720.
  • the hearable 102 transmits a second ultrasound transmit signal 802 that propagates within at least the portion of the ear canal 118 of the user 106. If the calibration procedure was performed, the second ultrasound transmit signal 802 can have the selected tones 904-1 to 904-M that were determined by the calibration procedure.
  • the selected tones 904 can be transmitted in parallel or in series over a given time interval.
  • An amplitude of the second ultrasound transmit signal 802 can be approximately the same across the selected tones 904-1 to 904-N. In this manner, power is evenly distributed across each selected tone.
  • the amplitude of the second ultrasound transmit signal 802 can be higher than the amplitude of the first ultrasound transmit signal 802 because the available output power is distributed across fewer tones.
  • a duration of each of the selected tones 904 of the second ultrasound transmit signal 802 can be longer than the duration of the tones 902 of the first ultrasound transmit signal 802. The higher amplitude and/or the longer duration can further improve the signal -to-noise ratio performance of the hearable 102 for audioplethysmography 110. By using a few selected tones 904 that were determined to improve signal-to-noise ratio performance, the measurement procedure can achieve a higher accuracy for voice activity detection 112.
  • the hearable 102 performs audioplethysmography (e.g., voice processing) using the second ultrasound signal (e.g., the second ultrasound receive signal 804).
  • audioplethysmography 110 can include performing internal -noise-source filtering 202. as further described with respect to FIGs. 12 and 13.
  • the calibration module 722 is further described with respect to FIG. 11.
  • the quality detector 1106 measures quality metrics 1114-1 to 1114-2M for each of the tones 902-1 to 902-M and for each of the characteristics (e g., amplitude 1110 and phase 1112).
  • the quality metrics 1114 can represent a variety of different metrics, including peak- to-average ratios and/or signal-to-noise ratios.
  • the peak-to-average ratio represents a peak intensity w ithin a frequency range of interest divided by an average intensity 7 within this frequency range.
  • a higher quality metric 1114 indicates a higher-quality signal, or more generally, better performance for audioplethysmography 110.
  • the filter 1204 generates a filtered signal 1210 based on the down-converted signal 1208.
  • the filter 1204 filters the down-converted signal 1208 to attenuate spurious or undesired frequencies (e g., intermodulation products), some of which can be associated with an operation of the in-phase and quadrature mixer 1202.
  • the filtered signal 1210 represents a combination of the in-phase and quadrature components of the down-converted signal 1208.
  • the filtered signal 1210 can represent separate or distinct in-phase and quadrature components, which are individually passed to the frequency selector 1206, the calibration module 722, or the measurement module 720.
  • the pre-processing module 718 can optionally apply the frequency selector 1206.
  • the frequency selector 1206 passes tones that meet a quality threshold level of performance for audioplethysmography 110.
  • the frequency selector 1206 passes tones 904 having an amplitude 1110 and/or phase 1112 with a quality metric 1114 that is greater than or equal to a threshold 1116.
  • the resulting signal outputted by the frequency selector 1206 is represented by signal 1212.
  • this signal 1212 is passed to the measurement module 720 as the pre-processed signal 910.
  • the filtered signal 1210 can be passed to the measurement module 720 and/or the calibration module 722 as the pre- processed signal 910.
  • the measurement module 720 can generate the voice data 912 based on the pre- processed signal 910. In this case, the measurement module 720 can analyze the changes in the amplitude 1110 and/or phase 1112 of the voice component 404 of the pre-processed signal 910 for voice processing.
  • This processing technique can be utilized in implementations of the hearable 102 that have minimal (if any) interference betw een the over-the-air signal 402 and the ultrasound signal 410, or in situations in which audioplethysmography 110 is performed in a relatively quiet environment.
  • the measurement module 720 can generate the voice data 912 based on the pre-processed signal 910 and the received audible signal 504. More specifically, the measurement module 720 can utilize the received audible signal 504 as a reference to attenuate the modulation component 510 within the pre-processed signal 910 and improve the signal-to-noise-ratio associated with the voice component 404.
  • the voice processor 1304 can perform aspects of voice activity detection 112, speech recognition 114, and/or conversation detection 116.
  • the voice processor 1304 can be implemented using a machine-learned model or another module that performs signal and/or data processing.
  • the voice processor 1304 can analyze changes in the amplitude 1110 and/or phase 1112 of the voice component 404 of the pre-processed signal 910 for voice processing.
  • the multi-stage filter 1302 filters the pre-processed signal 910 to improve the signal-to-noise ratio associated with the voice component 404.
  • the multi-stage filter 1302 includes at least one intemal-noise-source filter stage 1306 and at least one extemal-noise-source filter stage 1308.
  • the multi-stage filter 1302 includes the intemal-noise-source filter stage 1306 and does not include the extemal-noise-source filter stage 1308. This implementation is possible in situations in which audioplethysmography is performed in a relatively quiet environment and/or for implementations in which a design of the hearable 102 minimizes the impact of the modulation component 510.
  • the voice processor 1304 analyzes the denoised signal 1314 to generate the voice data 912.
  • the voice processor 1304 can generate the voice data 912 associated with voice activity detection 112, speech recognition 114. and/or conversation detection 116.
  • the voice data 912 can indicate whether or not the voice component 404 is detected.
  • the voice processor 1304 can perform a signal-to-noise ratio detection process, which determines whether or not an amplitude of an input signal exceeds a detection threshold. If the amplitude exceeds the detection threshold, the voice processor 1304 generates the voice data 912 to indicate that the vocalization is detected.
  • FIG. 14 illustrates an example implementation of the intemal-noise-source filter stage 1306.
  • the intemal-noise-source filter stage 1306 includes two adaptive filters 1310-1 and 1310-2.
  • the adaptive filters 1310-1 and 1310-2 can apply various adaptive filtering techniques, including techniques based on least mean squares (LMS) or recursive least squares (RLS).
  • the adaptive filter 1310-1 performs adaptive filtering using the noise signal 308 as a noise reference to attenuate the internal noise component 310 within the pre- processed signal 910.
  • the adaptive filter 1310-2 performs adaptive filtering using the noise signal 308 as a noise reference to attenuate the internal noise component 310 within the received audible signal 504.
  • the noise signal 308 is not correlated with the desired voice component 404 within the pre-processed signal 910 and the received audible signal 504.
  • the noise signal 308 can be used as a reference signal to attenuate the internal noise component 310 within the pre-processed signal 910 and the received audible signal 504 using adaptive filtering techniques.
  • the denoised pre-processed signal 1402 and the denoised audible signal 1404 do not include the internal noise component 310 or the internal noise component 310 is substantially attenuated relative to the corresponding input signal.
  • the filtering process continues with the extemal-noise-source filter stage 1306, which is further described with respect to FIG. 15.
  • the vocalization enhancer 1504 enhances (e.g., amplifies relative to a noise level) the voice component 404 within an output signal provided by the filter module 1502. In this way, the vocalization enhancer 1504 can increase sensitivity for voice processing.
  • the vocalization enhancer 1504 is implemented using a wiener filter 1510.
  • FIGs. 16 and 17 depict example methods 1600 and 1700 for implementing aspects of voice activity detection 112 using active acoustic sensing.
  • Methods 1600 and 1700 are shown as sets of operations (or acts) performed but not necessarily limited to the order or combinations in which the operations are shown herein. Further, any of one or more of the operations may be repeated, combined, reorganized, or linked to provide a wide array of additional and/or alternate methods.
  • the techniques are not limited to performance by one entity or multiple entities operating on one device.
  • an audible signal is transmitted during a first time period.
  • the audible signal is intended to propagate within at least a portion of an ear canal of a user.
  • the circuitry 7 304 transmits (or renders) the audible signal during the first time period.
  • the audible signal is intended to propagate within at least a portion of an ear canal 118 of the user 106.
  • the circuitry 304 can include a speaker 314 or 708, the active-noise-cancellation circuitry 316, or the transparency-mode circuitry 318.
  • an ultrasound receive signal is received.
  • the ultrasound receive signal represents a version of the ultrasound transmit signal with one or more waveform characteristics modified based on the propagation within the ear canal and based on a vocalization made by the user during the first time period.
  • the ultrasound receive signal comprises an internal noise component caused by interference generated by the rendering of the audible signal.
  • the transducer 706 (or the microphone 710) of the hearable 102 receives the ultrasound receive signal 804.
  • the ultrasound receive signal 804 represents a version of the ultrasound transmit signal 802 with one or more waveform characteristics modified based on the propagation within the ear canal 118 and based on a vocalization made by the user 106 during the first time period.
  • Example waveform characteristics include amplitude, phase, and/or frequency.
  • the ultrasound receive signal 804 includes the internal noise component 310, which is caused by interference generated by the rendering of the audible signal.
  • the internal noise component 310 can represent a portion of the audible signal that is modulated onto or mixed with the ultrasound receive signal 804.
  • the internal noise component 310 represents a portion of the ultrasound receive signal 804 (e.g.. amplitude, phase, and/or frequency) that is modified by the interference associated with the audible signal.
  • a denoised signal is generated by filtering the internal noise component within the received ultrasound receive signal based on a version of the audible signal.
  • the measurement module 720 e.g., the multi-stage filter 1302
  • the denoised signal 1314 by filtering the internal noise component 310 within the ultrasound receive signal 804 based on a version of the audible signal, which is represented by the noise signal 308 as shown in FIG. 13.
  • the vocalization is detected based on the denoised signal.
  • the hearable 102 uses the measurement module 720 (e.g., the voice processor 1304) to detect the vocalization based on the denoised signal 1314. More specifically, the measurement module 720 analyzes the voice component 404 that is present within the denoised signal 1314 to perform voice activity detection 112, speech recognition 114, and/or conversation detection 116. The measurement module 720 can generate voice data 912, which can be used to control the hearable 102 and/or the computing device 104.
  • active acoustic sensing is performed to detect a pressure wave that propagates within an ear canal of a user and is associated with a vocalization of the user.
  • the hearable 102 performs active acoustic sensing to detect a pressure wave that propagates within an ear canal 118 of a user 106 and is associated with a vocalization of the user 106.
  • the hearable 102 transmits and receives an ultrasound signal 410 (e.g., the ultrasound transmit signal 802 and the ultrasound receive signal 804).
  • the received ultrasound signal 410 includes the voice component 404, which enables audioplethysmography 1 10 to be used for voice processing.
  • an audible signal that interferes with active acoustic sensing is rendered.
  • circuitry 304 renders an audible signal that interferes with the active acoustic sensing.
  • the circuitry 304 can include a speaker 314 or 708, active-noise-cancellation circuitry 316, and/or transparency-mode circuitry 318.
  • the audible signal can be an audible signal 320 that includes audio content 204, the anti-noise signal 324, and/or the transparency-mode signal 326.
  • intemal-noise-source filtering is performed to attenuate the interference caused by the rendering of the audible signal.
  • the measurement module 720 performs intemal-noise-source filtering 202 to attenuate the interference (e.g., the internal noise component 310) caused by the rendering of the audible signal.
  • This interference can be attenuated within the pre-processed signal 910 and optionally the received audible signal 504, as shown in FIG. 14.
  • voice processing is performed based on the active acoustic sensing and the intemal-noise-source filtering.
  • the voice processor 1304 performs voice processing (e.g., voice activity detection 112, speech recognition 114, and/or conversation detection 116) based on the active acoustic sensing and the intemal-noise-source filtering (e.g., based on the denoised pre-processed signal 1402).
  • Example 7 The method of example 6, wherein the audio content comprises music.
  • Example 10 The method of any previous example, wherein the generating of the denoised signal comprises filtering a version of the received ultrasound receive signal based on a version of the audible signal using an adaptive filter.
  • Example 20 The device of any one of examples 17 to 19, wherein the device comprises: at least one earbud; or headphones.

Landscapes

  • Physics & Mathematics (AREA)
  • Engineering & Computer Science (AREA)
  • Acoustics & Sound (AREA)
  • Signal Processing (AREA)
  • Soundproofing, Sound Blocking, And Sound Damping (AREA)
  • Headphones And Earphones (AREA)
  • Transducers For Ultrasonic Waves (AREA)
  • Circuit For Audible Band Transducer (AREA)

Abstract

Techniques and apparatuses are described for performing internal-noise-source filtering for active acoustic sensing. With internal-noise-source filtering (202), interference caused by a hearable (102) performing other operations (e.g., rendering audio content (204), performing active noise cancellation (206), or operating in accordance with a transparency mode (214)) can be attenuated within a received ultrasound signal to improve sensitivity and accuracy for audioplethysmography. This performance improvement enables audioplethysmography to be performed while the hearable (102) performs these other operations. Furthermore, it improves the ability of audioplethysmography to be used for voice processing, which can include voice activity detection, speech recognition, and/or conversation detection.

Description

INTERNAL-NOISE-SOURCE FILTERING FOR ACTIVE ACOUSTIC SENSING
BACKGROUND
[0001] Wireless technology has become prevalent in everyday life, making communication and data readily accessible to users. One type of wireless technology are wireless hearables, examples of which include wireless earbuds and wireless headphones. Wireless hearables have allowed users freedom of movement while listening to audio content from music, audio books, podcasts, and videos. With the prevalence of wireless hearables, there is a market for adding additional features to existing hearables without introducing hardware changes.
SUMMARY
[0002] Techniques and apparatuses are described for performing intemal-noise-source filtering for active acoustic sensing. With intemal-noise-source filtering, interference caused by a hearable performing other operations (e.g., rendering audio content, performing active-noise cancellation, or operating in accordance with a transparency mode) can be attenuated within a received ultrasound signal to improve sensitivity and accuracy for audioplethysmography. This performance improvement enables audioplethysmography to be performed while the hearable performs these other operations. Furthermore, it improves the ability of audioplethysmography to be used for voice processing, which can include voice activity detection, speech recognition, and/or conversation detection.
[0003] Aspects described below include a method for performing intemal-noise-source filtering for active acoustic sensing. The method includes transmitting, during a first time period, an audible signal that propagates within at least a portion of an ear canal of a user. The method also includes transmitting, during the first time period, an ultrasound transmit signal that propagates within at least a portion of the ear canal of the user. The method additionally includes receiving, during the first time period, an ultrasound receive signal. The ultrasound receive signal represents a version of the ultrasound transmit signal with one or more characteristics modified based on the propagation within the ear canal and based on a vocalization made by the user during the first time period. The received ultrasound receive signal includes an internal noise component caused by interference generated by the rendering of the audible signal. The method further includes generating a denoised signal by filtering the internal noise component within the received ultrasound receive signal based on a version of the audible signal. The method also includes detecting the vocalization based on the denoised signal. In an example, the version of the audible signal can represent an electrical version of the audible signal and/or a digital version of the audible signal. [0004] Aspects described below include a computer-readable storage medium comprising instructions that, responsive to execution by a processor, cause a hearable to perform any one of the methods described herein.
[0005] Aspects described below include a device with at least one transducer and at least one processor. The device is configured to perform, using the at least one transducer and the at least one processor, any one of the methods described herein.
[0006] Aspects described below include a system with means for performing intemal-noise- source filtering for active acoustic sensing.
BRIEF DESCRIPTION OF DRAWINGS
[0007] Apparatuses for and techniques that perform intemal-noise-source filtering for active acoustic sensing are described with reference to the following drawings. The same numbers are used throughout the drawings to reference like features and components:
FIG. 1-1 illustrates an example environment in which active acoustic sensing can be implemented;
FIG. 1-2 illustrates an example geometric change in an ear canal, which can be detected using active acoustic sensing;
FIG. 2 illustrates an example environment in which intemal-noise-source filtering for active acoustic sensing can be implemented;
FIG. 3 illustrates example internal noise sources that can impact audioplethysmography;
FIG. 4 illustrates example signals that can be detected using a hearable;
FIG. 5 illustrates an example operation of a microphone of a hearable;
FIG. 6 illustrates example components of a computing device;
FIG. 7 illustrates example components of a hearable;
FIG. 8 illustrates example operations of two hearables;
FIG. 9 illustrates an example implementation of a hearable capable of performing intemal- noise-source filtering for active acoustic sensing;
FIG. 10 illustrates an example flow diagram for operating a hearable;
FIG. 11 illustrates an example scheme implemented by a calibration module of a hearable;
FIG. 12 illustrates an example implementation of a pre-processing module for performing aspects of audioplethysmography;
FIG. 13 illustrates an example implementation of a measurement module for performing intemal-noise-source filtering;
FIG. 14 illustrates an example implementation of an intemal-noise-source filter stage for performing intemal-noise-source filtering; FIG. 15 illustrates an example implementation of an external -noise-source filter stage;
FIG. 16 illustrates an example method for performing intemal-noise-source filtering;
FIG. 17 illustrates another example method for performing intemal-noise-source filtering; and
FIG. 18 illustrates an example computing system embodying, or in which techniques maybe implemented that enable use of, intemal-noise-source filtering for active acoustic sensing.
DETAILED DESCRIPTION
[0008] As electronic devices become more ubiquitous, users incorporate them into everyday life. A user, for example, may use an electronic device to get daily weather and traffic information, control a temperature of a home, answer a doorbell, turn on or off a light, and/or play background music. Interacting with some electronic devices, however, can be cumbersome and inefficient. An electronic device, for instance, can have a physical user interface that may require a user to navigate through one or more prompts by physically touching the electronic device. In this case, the user has to devote attention away from other pnmary tasks to interact with the electronic device, which can be inconvenient and disruptive.
[0009] To address this problem, some electronic devices support voice control, which enables a user to interact with the electronic device in a non-physical and less cognitively demanding way- compared to other interfaces that require physical touch and/or the user’s visual attention. With voice control, the electronic device seamlessly exists in the surrounding environment and provides the user access to information and services while the user performs a primary task, such as cooking, cleaning, driving, talking with people, or reading a book.
[0010] While voice control can provide a convenient means of interacting with an electronic device, there are several challenges associated with voice control. In a noisy environment, for instance, the user’s voice can be made imperceptible by the other external noise. Consequently, it can be challenging for voice control to detect and/or recognize voice commands spoken by the user. Also, sometimes the noisy environment can cause the voice control to incorrectly respond to a voice of another person who is not authorized to use the electronic device.
[0011] Some devices address these challenges by integrating a voice accelerometer (VA) into earbuds. The voice accelerometer can detect a user speaking based on sound that travels by means of bone conduction. The voice accelerometer, however, can be bulky and expensive. To improve aesthetics and reduce encumbrance, it can be desirable to design hearables with smaller sizes. As space becomes limited, it can be challenging to integrate additional components, such as the voice accelerometer, within the hearables. With the prevalence of hearables, there is a market for adding additional features to existing hearables without introducing hardware changes. [0012] Provided according to one or more preferred embodiments is a hearable, such as an earbud, that is capable of performing a novel physiological monitoring process termed herein audioplethysmography. Audioplethysmography is an active acoustic method capable of sensing subtle physiologically -related changes observable at a user’s outer and middle ear. Instead of relying on other auxiliary sensors, such as optical or electrical sensors, audioplethysmography involves transmitting and receiving acoustic signals that at least partially propagate within auser’s ear canal. To perform audioplethysmography, the hearable forms at least a partial seal in or around the user’s outer ear. This seal enables formation of an acoustic circuit, which includes the seal, the hearable, the ear canal, and an ear drum of the ear. By transmitting and receiving acoustic signals, the hearable can recognize changes in the acoustic circuit to detect and/or recognize vocalizations made by the user. Vocalizations can include any sound that is produced using the user’s lung’s, vocal cords, and/or mouth. Example types of vocalizations can involve the user speaking, whispering, shouting, humming, whistling, singing, or making other utterances. The hearable can utilize audioplethysmography to support a variety of different features, including a voice user interface (VUI) and/or multi-factor voice authentication.
[0013] Some hearables can provide other features in addition to audioplethysmography. Example features can include rendering (e.g., playing) audible content, providing active noise cancellation, and/or operating in accordance with a transparency mode. The transparency mode enables sounds from the external environment to be rendered for the user. In this manner, the transparency mode enables the hearable to function as a hearing aid. While active noise cancellation can significantly attenuate sounds from the external environment, the transparency mode allows sounds from the external environment to pass to the user’s ear. In some cases, the transparency mode can further amplify these sounds to assist with a hearing impairment.
[0014] It can be challenging, however, to provide these other features while also performing audioplethysmography. More specifically, signals generated for rendering audio content, providing active noise cancellation, and/or for operating in accordance with a transparency mode can interfere with an ultrasound signal that is received for audioplethysmography. This interference can make it challenging for the hearable to use audioplethysmography for voice processing.
[0015] To address this challenge, techniques and apparatuses are described for performing intemal-noise-source filtering for active acoustic sensing. With intemal-noise-source filtering, interference caused by the hearable performing other operations (e.g., rendering audio content, performing active-noise cancellation, or operating in accordance with a transparency mode) can be attenuated within a received ultrasound signal to improve sensitivity and accuracy for audioplethysmography. This performance improvement enables audioplethysmography to be performed while the hearable performs these other operations. Furthermore, it improves the ability of audioplethysmography to be used for voice processing, which can include voice activity detection, speech recognition, and/or conversation detection.
Operating Environment
[0016] FIG. 1-1 is an illustration of an example environment 100 in which active acoustic sensing can be implemented. In the example environment 100, a hearable 102 is connected to a computing device 104 using a physical or wireless interface. The hearable 102 is a device that can play audible content provided by the computing device 104 and direct the audible content into a user 106’s ear 108. In this example, the hearable 102 operates together with the computing device 104. In other examples, the hearable 102 can operate or be implemented as a stand-alone device. Although depicted as a smartphone, the computing device 104 can include other types of devices, including those described with respect to FIG. 6.
[0017] The hearable 102 is capable of performing audioplethysmography 110, which is an active acoustic method of sensing that occurs at the ear 108. The hearable 102 can perform this sensing without the use of other auxiliary sensors, such as an optical sensor or an electrical sensor. Through audioplethysmography 110, the hearable 102 can utilize voice processing to detect and/or recognize vocalizations made by the user 106. Example types of voice processing can include voice activity’ detection 112, speech recognition 114, and/or conversation detection 116.
[0018] For voice activity detection 112, the hearable 102 detects vocalizations made by the user 106. To perform voice activity detection 112, the hearable 102 uses audioplethysmography 1 10 to detect subtle pressure waves that propagate to the user 106’s ear canal 1 18. The pressure waves can occur due to movement associated with the user 106’s jaw, a vibration associated with the vocalization, or some combination thereof. Generally speaking, these pressure waves modify characteristics of ultrasound signals that are transmitted and received by the hearable 102 and propagate through the ear canal 118. With voice activity detection 112, the hearable 102 can enhance performance of a voice user interface or provide multi-factor authentication.
[0019] For speech recognition 114, the hearable 102 recognizes words spoken by the user 106. The hearable 102 can perfonn speech recognition 114 using audioplethysmography 110 or using audioplethysmography 1 10 to augment information provided by another signal, such as an audible signal captured by the hearable 102. Speech recognition 114 can also be used to enhance performance of the voice user interface or to enhance multi-factor authentication. [0020] For conversation detection 116, the hearable 102 determines whether or not the user 106 is conversing with another person. To perform conversation detection 116, the hearable 102 uses audioplethysmography 110 to filter an audible signal that is captured by the hearable 102. As a signal associated with audioplethysmography 110 is uncorrelated with the speech from another person, audioplethysmography 110 can be used to filter and/or attenuate signal components that are common between the audible signal and the signal used for audioplethysmography 1 10. This filtering improves the signal-to-noise ratio for detecting vocalizations made by a person other than the user 106.
[0021] Responsive to detecting occurrence of a conversation, conversation detection 116 can be used to control an operation of the hearable 102 and/orthe computing device 104. In one example, audio content can be paused or resumed based on whether a conversation is detected or not. In another example, a volume of the audio content is decreased or increased based on whether a conversation is detected or not. In yet another example, active noise cancellation can be temporarily halted and the hearable 102 can operate in accordance with a transparency mode based on the conversation detection 116 detecting a conversation. Alternatively, active noise cancellation can resume and the hearable 102 can exit out of the transparency mode based on the conversation detection 116 detecting the absence of a conversation.
[0022] To use audioplethysmography 110, the user 106 positions the hearable 102 in a manner that creates at least a partial seal 120 around or in the ear 108. Some parts of the ear 108 are shown in FIG. 1-1, including the ear canal 118 and an ear drum 122 (or tympanic membrane). Due to the seal 120, the hearable 102, the ear canal 118, and the ear drum 122 couple together to form an acoustic circuit. Audioplethysmography 110 involves, at least in part, measuring properties associated with this acoustic circuit. The properties of the acoustic circuit can change due to a variety of different situations or actions.
[0023] For example, consider FIG. 1-2 in which a change occurs in a physical structure of the ear 108. Example changes to the physical structure include a change in a geometric shape of the ear canal 118 and/or a change in a volume of the ear canal 118. This change can be caused, at least in part, by bone conduction and/or a pressure wave associated with a vocalization made by the user 106.
[0024] At 124, for instance, the tissue around the ear canal 118 and the ear drum 122 itself are slightly "squeezed" due to the bone conduction and/or the pressure wave. This squeeze causes a volume of the ear canal 1 18 to be slightly reduced at 124. At 126, however, the squeezing subsides and the volume of the ear canal 118 is slightly increased relative to 124. The physical changes within the ear 108 can modulate an amplitude and/or phase of an ultrasound signal that propagates through the ear canal 1 18, as further described with respect to FIG. 8. The techniques for audioplethysmography 110 can be performed while the hearable 102 is playing audible content to the user 106 as further described with respect to FIGs. 2, 3-1, and 3-2.
[0025] FIG. 2 illustrates example environments 200-1, 200-2, and 200-3 in which internal -noise- source filtering 202 can be implemented for active acoustic sensing. The environments 200-1, 200-2, and 200-3 depict example features of the hearable 102, which may be active during audioplethysmography 110. In the environment 200-1, the hearable 102 renders audio content 204 for the user 106. The audio content 204 represents any type of sound that is produced by the hearable 102, such as music, a ringtone, an alarm, a caller’s voice, and so forth.
[0026] In the environment 200-2, the hearable 102 provides active noise cancellation 206. With active noise cancellation 206, the hearable 102 significantly attenuates other sounds that are present within the external environment, such as speech 208, music 210, and/or external noise 212. Active noise cancellation 206 can create a quiet environment for the user 106 to concentrate. Additionally or alternatively, active noise cancellation 206 can make it easier for the user 106 to hear the audio content 204, which can also be played for the user 106 in the environment 200-2.
[0027] In the environment 200-3, the hearable 102 operates in accordance with a transparency mode 214. The transparency mode enables sounds from the external environment to be rendered for the user 106. The rendering of these sounds can also include amplifying the sounds to assist with a hearing impairment. In the environment 200-3, the transparency mode 214 enables the user 106 to hear another person’s speech 216 (e.g., to hear the other person talking).
[0028] While the hearable 102 renders audio content 204, provides active noise cancellation 206, and/or operates in accordance with a transparency mode 214, internal signals generated by the hearable 102 during any of these operations can interfere with a received ultrasound signal. This interference can make it challenging for audioplethysmography 110 to be used for voice processing. To address this problem, the hearable 102 performs intemal-noise-source filtering 202. Example types of internal noise sources that can be filtered using intemal-noise- source filtering 202 are further described with respect to FIG. 3.
[0029] FIG. 3 illustrates example components of a hearable 102. In the depicted configuration, the hearable 102 includes at least one microphone 302 and circuitry 304. The microphone 302 can be used to receive ultrasound signals for audioplethysmography 110 and/or audible signals for other operations of the hearable 102. The circuitry’ 304 can provide any of the features described with respect to FIG. 2. For example, the circuitry 304 can render the audio content 204, perform active noise cancellation 206, and/or support the transparency mode 214. [0030] During operation, the microphone 302 receives a signal, which is represented by received signal 306. The circuitry 304 can generate one or more signals while the microphone 302 is generating the received signal 306 and/or while the receive signal 306 is being processed by the hearable 102. The hearable 102 renders (e.g., plays, emits, or transmits) the signal that is generated by the circuitry 304. The rendered signal propagates within the ear canal 118 of the user 106.
[0031] From the perspective of any techniques that process the received signal 306, the signal generated by the circuitry' 304 represents a noise signal 308 (or a self-generated noise signal 308), which can interfere with the received signal 306. For example, the noise signal 308 can impact waveform characteristics (e.g., an amplitude, frequency and/or phase) of the received signal 306. This modification to the waveform characteristics of the receive signal 306 is represented by an internal noise component 310. In some cases, the internal noise component 310 is caused by the noise signal 308 being modulated onto or mixed with the received signal 306. This interference between the noise signal 308 and the received signal 306 can be due to non-linearities in the microphone 302, a mixing operation performed by the hearable 102, or some other component and/or operation of the hearable 102. This internal noise component 310 can make it challenging for audioplethysmography 110 to be used for voice processing.
[0032] Generally speaking, the circuitry’ 304 represents an internal noise source 312. Example internal noise sources 312 can include any type of circuit that is implemented within the hearable 102 (e.g., within a housing of the hearable 102) and whose operation can interfere yvith signals that are received for audioplethysmography 110. In FIG. 3, the internal noise source 312 can include a speaker 314, active-noise-cancellation (ANC) circuitry 316 (ANC circuitry 316), transparency-mode circuitry 318, or some combination thereof.
[0033] The speaker 314 can be used to render the audio content 204 in the environment 200-1. To render the audio content 204, the speaker 314 generates an audible signal 320, which can represent the noise signal 308. The audible signal 320 can be associated with music 322, a human voice, or some other type of sound. The audible signal 320 can also be referred to as an audible frequency spectrum signal, and generally includes frequencies between 20 hertz (Hz) and 20 kilohertz (kHz).
[0034] The active-noise-cancellation circuitry 316 performs active noise cancellation 206, as shown in the environment 200-2. To perform active noise cancellation 206. the active-noise- cancellation circuitry 316 generates an anti -noise signal 324. The anti-noise signal 324 represents another type of noise signal 308, which can interfere with the received signal 306 and contribute to the internal noise component 310. [0035] The transparency mode circuitry 318 enables the hearable 102 to operate in accordance with the transparency mode, as shown in the environment 200-3. For the transparency mode, the transparency mode circuitry7 318 generates a transparency -mode signal 326, which includes sounds from the external environment. The transparency-mode signal 326 represents another type of noise signal 308, which can interfere with the receive signal 306 and contribute to the rntemal noise component 310.
[0036] It can be challenging to design the hearable 102 in a manner that mitigates the interference caused by the noise signal 308. Cost and/or size constraints of the hearable 102 can make it challenging to incorporate some interference-shielding techniques and/or utilize more expensive components that are less likely to generate this interference. To address this challenge, techniques for intemal-noise-source filtering 202 filter (e.g., attenuate) the internal noise component 310 within the received signal 306 to enhance audioplethysmography 110 and/or voice processing. Example techniques for utilizing audioplethysmography 110 for speech processing are further described with respect to FIGs. 4 and 5.
[0037] FIG. 4 illustrates example signals that can be detected by the hearable 102. In the environment 400, the hearable 102 is being worn by the user 106. During operation, the hearable 102 can receive a variety of different signals. In one aspect, the hearable 102 receives at least one over-the-air (OTA) signal 402 (OTA signal 402). The over-the-air signal 402 can include a voice component 404 and/or a noise component 406. The voice component 404 can include a vocalization made by the user 106, such as a voiceprint phrase for activating a voice user interface. The noise component 406 can represent any undesired audible sound that can mask or interfere with detection of the vocalization. The noise component 406 can represent the environmental noise 212, the music 210, and/or the speech 208 of FIG. 2. In general, the over-the-air signal 402 includes audible frequencies.
[0038] In another aspect, the hearable 102 receives at least one bone-conduction signal 408. The bone-conduction signal 408 represents sound that travels, via bone conduction, to the user 106’s ear 108. The bone-conduction signal 408 also includes the voice component 404. In most circumstances, the bone-conduction signal 408 does not include the noise component 406.
[0039] To perform active acoustic sensing, the hearable 102 transmits and receives at least one ultrasound signal 410. The ultrasound signal 410 propagates within the ear canal 118. A vocalization of the user 106 can cause a physical structure of the ear 108 to change. As such, the ultrasound signal 410 can also include the voice component 404.
[0040] The ultrasound signal 410 may not directly include the noise component 406. However, some designs of the hearable 102 can cause the noise component 406 associated with the over- the-air signal 402 to interfere with the detection of the voice component 404 within the ultrasound signal 410, as further explained below.
[0041] In some example implementations, the hearable 102 may be designed to minimize interference between the over-the-air signal 402 and the ultrasound signal 410. The hearable 102, for instance, can utilize different microphones to receive these signals. In this case, the hearable 102 can directly process the ultrasound signal 410 to detect the vocalization. In other example implementations, the hearable 102 utilizes a same microphone to receive the over-the- air signal 402 and the ultrasound signal 410. While this may be beneficial for meeting size constraints of the hearable 102, it can cause a version of the noise component 406 to be present within the ultrasound signal 410. As the hearable 102 receives the over-the-air signal 402, the bone-conduction signal 408, and the ultrasound signal 410, these signals can interact with each other and make it challenging to utilize audioplethysmography 110 to detect the voice component 404. In particular, the over-the-air signal 402 can be modulated onto or mixed with the ultrasound signal 410, as further described with respect to FIG. 5.
[0042] FIG. 5 illustrates an example operation of the microphone 302 of the hearable 102. During an operation, the microphone 302 receives the over-the-air signal 402, the bone-conduction signal 408, and the ultrasound signal 410. The microphone 302 can include a filter module, which can generate separate signals associated with different frequency ranges. In this example, the microphone 302 generates a received audible signal 504 and a received ultrasound signal 506, which are further depicted in a graph 508 at the bottom of FIG. 5. These signals can be downconverted to baseband frequencies. In general, the received audible signal 504 and the received ultrasound signal 506 represent electrical signals that can be processed by other components of the hearable 102.
[0043] In the graph 508, the received audible signal 504 and the received ultrasound signal 506 are shown to include different frequencies. The received audible signal 504 can include frequencies associated with the audible frequency spectrum (e.g., frequencies between approximately 20 Hz and 20 kHz). As such, the received audible signal 504 can also be referred to as a received audible frequency spectrum signal. In contrast, the received ultrasound signal 506 can include frequencies associated with the ultrasound frequency spectrum (e.g., frequencies between approximately 20 kHz and 2 megahertz (MHz)).
[0044] The received audible signal 504 represents a convolution of the over-the-air signal 402 and the bone-conduction signal 408. As such, the received audible signal 504 includes the voice component 404 associated with the vocalization and the noise component 406. The voice component 404 is provided by both the over-the-air signal 402 and the bone-conduction signal 408. The received audible signal 504 can be represented by Equation 1:
YRAS = hBC - S + hOTA(S + N) Equation 1 where YRAS represents the received audible signal 504, IIBC represents a bone-conduction channel, 5 represents the vocalization, hoTA represents an over-the-air channel, and N represents noise (e.g., the external noise 212, the music 210, and/or the speech 208). The hBC ■ S term represents a boneconduction component 512. The h0TA ■ S term represents the voice component 404 of the received audible signal 504. The h0TA ■ N term represents the noise component 406 of the received audible signal 504.
[0045] The received ultrasound signal 506 represents a convolution of the ultrasound signal 410 and a modulated version of the received audible signal 504, which is represented by a modulation component 510. The modulation component 510 can be caused by an interaction of the over-the- air signal 402 and the bone-conduction signal 408 with the ultrasound signal 410 as the microphone 302 receives these signals. More specifically, the modulation component 510 can be due to interference, intermodulation distortion, and/or harmonics generated based on the design of the hearable 102, non-linearities within one or more components of the hearable 102, and/or an operation of the hearable 102 (e.g.. a mixing operation). Generally speaking, the modulation component 510 represents a version of the received audible signal 504 that is shifted to the ultrasound frequencies. The modulation component 510 is linearly modulated onto the ultrasound frequency spectrum. The received ultrasound signal 506 can be represented by Equation 2:
YRUS = us ' + ^-MC ‘ YRAS Equation 2 where YRUS represents the received ultrasound signal 506, hi s represents an ultrasound channel, S represents the vocalization, IIMC represents a modulation channel, and YR \S represents the received audible signal 504. The hus ■ S term represents the voice component 404 of the received ultrasound signal 506. The hMC ■ YRAS term represents the modulation component 510. Due to the modulation channel, the received ultrasound signal 506 includes a linearly modulated version of the received audible signal 504. As such, the noise component 406 within the modulated version of the received audible signal 504 can make it challenging to directly detect the voice component 404 within the received ultrasound signal 506.
[0046] Generally speaking, the voice component 404 is superimposed onto the ultrasound signal 410 and is correlated with the vocalization. The voice component 404 is not correlated with the noise component 406. The voice component 404 modulates the received ultrasound signal 506 in a different manner than the modulation component 510 due to the bone conduction and change in the physical structure within the ear 108. [0047] The voice component 404 of the received ultrasound signal 506 is frequency dependent. In other words, different ultrasound frequencies modulate the vocalization differently. The boneconduction component 512, however, does not have frequency selectivity7. In other words, the bone conduction component 512 is associated with a fixed channel. The techniques for voice activity detection 112, speech recognition 114, and/or conversation detection 1 16 utilize the received audible signal 504 to extract the voice component 404 from the received ultrasound signal 506, as further described with respect to FIGs. 13 and 14.
[0048] FIG. 6 illustrates an example implementation of the computing device 104. The computing device 104 is illustrated with various non-limiting example devices including a desktop computer 104-1, a tablet 104-2, a laptop 104-3, a television 104-4, a computing watch 104-5, computing glasses 104-6, a gaming system 104-7, a microwave 104-8, and a vehicle 104-9. Other devices may also be used, such as an augmented and/or virtual reality7 headset, a home service device, a smart speaker, a smart thermostat, a baby monitor, a Wi-Fi™ router, a drone, a trackpad, a drawing pad, a netbook, an e-reader, a home automation and control system, a wall display, and another home appliance. Note that the computing device 104 can be wearable, non-wearable but mobile, or relatively immobile (e.g., desktops and appliances).
[0049] The computing device 104 includes one or more computer processors 602 and at least one computer-readable medium 604. which includes memory media and storage media. Applications and/or an operating system (not shown) embodied as computer-readable instructions on the computer-readable medium 604 can be executed by the computer processor 602 to provide some of the functionalities described herein. The computer-readable medium 604 can optionally include an application 606, a voice user interface 608, and/or a voice authenticator 610. The application 606 can use information provided by the hearable 102 to perform an action. Example actions can include displaying data associated with audioplethysmography 110 to the user 106. For voice activity7 detection 112, the application 606 can indicate whether or not the vocalization is detected. The voice user interface 608 can enable the user 106 to control the computing device 104 via voice commands. The voice authenticator 610 can authenticate the user 106 and enable use of the voice user interface 608 upon successful authentication. The application 606, the voice user interface 608, and/or the voice authenticator 610 can utilize aspects of voice activity' detection 112 and/or speech recognition 114 to improve performance and/or enhance security of the computing device 104.
[0050] The computing device 104 can also include a network interface 612 for communicating data over wired, wireless, or optical networks. For example, the network interface 612 may communicate data over a local-area-network (LAN), a wireless local-area-network (WLAN), a personal -area-network (PAN), a wire-area-network (WAN), an intranet, the Internet, a peer-to- peer network, point-to-point network, a mesh network, Bluetooth®, and the like. The computing device 104 may also include the display 614. Although not explicitly shown, the hearable 102 can be integrated within the computing device 104, or can connect physically or wirelessly to the computing device 104. The hearable 102 is further described with respect to FIG. 7.
[0051] FIG. 7 illustrates an example hearable 102. The hearable 102 is illustrated with various non-limiting example devices, including wireless earbuds 702-1, wired earbuds 702-2, and headphones 702-3. The earbuds 702-1 and 702-2 are a type of in-ear device that fits into the ear canal 118. Each earbud 702-1 or 702-2 can represent a hearable 102. Headphones 702-3 can rest on top of or over the ears 108. The headphones 702-3 can represent closed-back headphones, open-back headphones, on-ear headphones, or over-ear headphones. Each headphone 702-2 includes two hearables 102, which are physically packaged together. In general, there is one hearable 102 for each ear 108.
[0052] The hearable 102 includes a communication interface 704 to communicate with the computing device 104, though this need not be used when the hearable 102 is integrated within the computing device 104. The communication interface 704 can be a wired interface or a wireless interface, in which audio content is passed from the computing device 104 to the hearable 102. The hearable 102 can also use the communication interface 704 to pass data associated with audioplethysmography 1 10 to the computing device 104. In general, the data provided by the communication interface 704 is in a format usable by the application 606, the voice user interface 608, and/or the voice authenticator 610.
[0053] The communication interface 704 also enables the hearable 102 to communicate with another hearable 102. During bistatic sensing, for instance, the hearable 102 can use the communication interface 704 to coordinate with the other hearable 102 to support two-ear audioplethysmography 110, as further described with respect to FIG. 8. In particular, the transmitting hearable 102 can communicate timing and wavefonn information to the receiving hearable 102 to enable the receiving hearable 102 to appropriately demodulate a received ultrasound signal 506.
[0054] The hearable 102 includes at least one transducer 706 that can convert electrical signals into sound waves. The transducer 706 can also detect and convert sound waves into electrical signals. These sound waves may include ultrasonic frequencies and/or audible frequencies, either of which may be used for audioplethysmography 110. In particular, a frequency spectrum (e.g., range of frequencies) that the transducer 706 uses to generate an acoustic signal can include frequencies from a low-end of the audible range to ahigh-end of the ultrasonic range, e.g., between 20 hertz (Hz) to 2 megahertz (MHz). Other example frequency spectrums for audioplethysmography 110 can encompass frequencies between 20 Hz and 20 kilohertz (kHz), between 20 kHz and 2 MHz, between 20 and 96 kHz, between 20 and 60 kHz, or betw een 30 and 40 kHz.
[0055] In an example implementation, the transducer 706 has a monostatic topology. With this topology, the transducer 706 can convert the electrical signals into sound waves and convert sound waves into electrical signals (e.g., can transmit or receive acoustic signals). Example monostatic transducers may include piezoelectric transducers, capacitive transducers, and micro-machined ultrasonic transducers (MUTs) that use microelectromechanical systems (MEMS) technology.
[0056] Alternatively, the transducer 706 can be implemented with a bistatic topology, which includes multiple transducers that are physically separate. In this case, a first transducer converts the electrical signal into sound waves (e.g., transmits acoustic signals), and a second transducer converts sound waves into an electrical signal (e.g., receives the acoustic signals). An example bistatic topology can be implemented using at least one speaker 708 and at least one microphone 710. The speaker 708 and the microphone 710 can be dedicated for audioplethysmography 110 or can be used for both audioplethysmography 110 and other functions of the computing device 104 (e.g., presenting audible content to the user 106, capturing the user 106’s voice for a phone call, or for voice control). The speaker 708 can represent the speaker 314 of FIG. 3. The microphone 710 can represent the microphone 302 of FIGs. 3 and 5. [0057] In general, the speaker 708 and the microphone 710 are directed tow ards the ear canal 118 (e.g., oriented towards the ear canal 118). Accordingly, the speaker 708 can direct ultrasound signals towards the ear canal 118, and the microphone 710 is responsive to receiving ultrasound signals from the direction associated with the ear canal 118. In some cases, the hearable 102 includes another microphone 710 that is directed away from the ear canal 118 towards an external environment (e.g., oriented away from the ear canal 118). This other microphone can be used to receive the over-the-air signal 402.
[0058] The hearable 102 includes at least one analog circuit 712. which includes circuitry and logic for conditioning electrical signals in an analog domain. The analog circuit 712 can include analog-to-digital converters, digital-to-analog converters, amplifiers, filters, mixers, and swatches for generating and modifying electrical signals. In some implementations, the analog circuit 712 includes other hardware circuitry associated with the speaker 708 or microphone 710.
[0059] The hearable 102 also includes at least one system processor 714 and at least one system medium 716 (e.g., one or more computer-readable storage media). In the depicted configuration, the system medium 716 includes a pre-processing module 718 and a measurement module 720. The system medium 716 also optionally includes a calibration module 722. The pre-processing module 718, the measurement module 720, and the calibration module 722 can be implemented using hardware, software, firmware, or a combination thereof. In this example, the system processor 714 implements the pre-processing module 718, the measurement module 720, and the calibration module 722. In an alternative example, the computer processor 602 of the computing device 104 can implement at least a portion of the pre-processing module 718, the measurement module 720, and/or the calibration module 722. In this case, the hearable 102 can communicate digital samples of the acoustic signals to the computing device 104 using the communication interface 704.
[0060] Operations of the pre-processing module 718, the measurement module 720, and the calibration module 722 are further described with respect to FIGs. 11 to 12. Aspects of intemal- noise-source filtering 202 can be performed, at least partially, by the measurement module 720, as further described with respect to FIGs. 13 and 14. The measurement module 720 can also perform aspects of voice activity detection 112, speech recognition 114, and/or conversation detection 1 16 using active acoustic sensing, as further described with respect to FIG. 13.
[0061] Some hearables 102 include the active-noise-cancellation circuitry' 316, which enables the hearables 102 to reduce background or environmental noise. In this case, the microphone 710 used for audioplethysmography 110 can be implemented using a feedback microphone of the active-noise-cancellation circuitry 316. During active noise cancellation, the feedback microphone provides feedback information regarding the performance of the active noise cancellation 206. During audioplethysmography 110, the feedback microphone receives an ultrasound signal 410, which is provided to the pre-processing module 718. In some situations, active noise cancellation 206 and audioplethysmography 110 are performed simultaneously using the feedback microphone. In this case, the ultrasound signal 410 received by the feedback microphone can be provided to the pre-processing module 718 and the feedback signal for active noise cancellation 206 can be provided to the active-noise-cancellation circuitry' 316. Other implementations are also possible in which the microphone 710 is implemented using a feedforward microphone of the active-noise-cancellation circuitry 316.
[0062] Also, some hearables 102 include the transparency -mode circuitry' 318, which enables the hearables 102 to amplify background or environmental sounds. The transparency -mode circuitry 318 can include at least one amplifier in some implementations. The microphone 710 that is used for audioplethysmography 110 can also be used for the transparency mode 214. More specifically, the ultrasound signal 410 received by the microphone 710 can be provided to the preprocessing module 718 and the sensed signal for the transparency mode 214 can be provided to the transparency-mode circuitry 318. The active-noise-cancellation circuitry 316, the transparency-mode circuitry 318, and/or the speaker 708 represent example types of circuitry 304 that can unintentionally act as an internal noise source 312 during audioplethysmography 110.
[0063] Although not explicitly shown in FIG. 7. the system medium 716 can also include a voice user interface 608 and/or a voice authenticator 610. In this case, the voice user interface 608 enables the user 1 6 to use voice controls to control an operation of the hearable 102. The voice authenticator 610 can authenticate the user 106 and enable the voice user interface 608 for the hearable 102. Different types of audioplethysmography 110 are further described with respect to FIG. 8.
Active Acoustic Sensing
[0064] FIG. 8 illustrates example operations of two hearables 102-1 and 102-2. In a first example operation, the hearables 102-1 and 102-2 perform single-ear audioplethysmography 110. This means that the hearables 102-1 and 102-2 independently perform audioplethysmography 110 on different ears 108 of the user 106. In this case, the first hearable 102-1 is proximate to the user 106’s right ear 108, and the second hearable 102-2 is proximate to the user 106’s left ear 108. Each hearable 102-1 and 102-2 includes a speaker 708 and a microphone 710. The hearables 102-1 and 102-2 can operate in a monostatic manner during the same time period or during different time periods. In other words, each hearable 102-1 and 102-2 can independently transmit and receive ultrasound signals.
[0065] For example, the first hearable 102-1 uses the speaker 708 to transmit a first ultrasound transmit 802-1, which propagates within at least a portion of the user 106’s right ear canal 118. The first hearable 102-1 uses the microphone 710 to receive a first ultrasound receive signal 804-1. The first ultrasound receive signal 804-1 represents a version of the first ultrasound transmit signal 802-1 that is modified, at least in part, by the acoustic circuit associated with the right ear canal 118. This modification can change an amplitude, phase, and/or frequency of the first ultrasound receive signal 804-1 relative to the first ultrasound transmit signal 802-1.
[0066] Similarly, the second hearable 102-2 uses the speaker 708 to transmit a second ultrasound transmit signal 802-2, which propagates within at least a portion of the user 106’s left ear canal 118. The second hearable 102-2 uses the microphone 710 to receive a second ultrasound receive signal 804-2. The second ultrasound receive signal 804-2 represents a version of the second ultrasound transmit signal 802-2 that is modified by the acoustic circuit associated with the left ear canal 118. This modification can change an amplitude, phase, and/or frequency of the second ultrasound receive signal 804-2 relative to the second ultrasound transmit signal 802-2. [0067] The techniques of single-ear audioplethysmography 110 can be particularly beneficial as it enables the computing device 104 to compile information from both hearables 102-1 and 102-2, which can further improve measurement confidence. For some aspects of audioplethysmography 110, it can be beneficial to analyze the acoustic channel between two ears 108, as further described below.
[0068] In a second example operation, the two hearables 102-1 and 102-2 perform two-ear audioplethysmography 110. This means that the hearables 102-1 and 102-2 jointly perform audioplethysmography 110 across two ears 108 of the user 106. In this case, at least one of the hearables 102 (e.g., the first hearable 102-1) includes the speaker 708. and at least one of the other hearables 102 (e.g., the second hearable 102-2) includes the microphone 710. The hearables 102-1 and 102-2 operate together in a bistatic manner during the same time period.
[0069] During operation, the first hearable 102- 1 transmits a third ultrasound transmit 802-3 using the speaker 708. The third ultrasound transmit signal 802-3 propagates through the user IO6's right ear canal 118. The third ultrasound transmit signal 802-3 also propagates through an acoustic channel that exists between the right and left ears 108. In the left ear 108, the third ultrasound transmit signal 802-3 propagates through the user 106’s left ear canal 118 and is represented as a third ultrasound receive signal 804-3. The second hearable 102-2 receives the third ultrasound receive signal 804-3 using the microphone 710. The third ultrasound receive signal 804-3 represents a version of the third ultrasound transmit signal 802-3 that is modified by the acoustic circuit associated with the right ear canal 118, modified by the acoustic channel associated with the user 106's face, and modified by the acoustic circuit associated with the left ear canal 118. This modification can change an amplitude, phase, and/or frequency of the third ultrasound receive signal 804-3 relative to the third ultrasound transmit signal 802-3. In some cases, the hearable 102-2 measures the time-of-flight (ToF) associated with the propagation from the first hearable 102-1 to the second hearable 102-2. Sometimes a combination of single-ear and two-ear audioplethysmography 110 are applied to further improve measurement confidence.
[0070] The ultrasound transmit signals 802 of FIG. 8 can represent a variety of different types of signals as described above with respect to FIG. 7. In example implementations, the ultrasound transmit signal 802 can be the ultrasound signal 410 of FIGs. 4 and 5. Also, the ultrasound transmit signal 802 can be a continuous-wave signal (e.g., a sinusoidal signal) or a pulsed signal. Some ultrasound transmit signals 802 can have a particular tone (or frequency). Other ultrasound transmit signals 802 can have multiple tones (or multiple frequencies). A variety of modulations can be applied to generate the ultrasound transmit signal 802. Example modulations include linear frequency modulations, triangular frequency modulations, stepped frequency modulations, phase modulations, or amplitude modulations. The ultrasound transmit signal 802 can be transmitted as part of a calibration procedure or a measurement procedure, as further described as part of FIG. 9. [0071] FIG. 9 illustrates an example implementation of the hearable 102 for performing voice activity detection 112. In the depicted configuration, the hearable 102 includes the speaker 708, the microphone 710, the analog circuit 712, the pre-processing module 718, the measurement module 720, the calibration module 722, and the circuitry' 304. Other implementations of the hearable 102, however, are also possible in which the hearable 102 does not include the calibration module 722 to reduce processing power requirements. In this case, the pre-processing module 718 can perform aspects of frequency selection as further described with respect to FIG. 12 to improve the signal -to-noise ratio for audioplethysmography 110.
[0072] Outputs of the speaker 708 and the microphone 710 are coupled to inputs of the analog circuit 712. The pre-processing module 718 has inputs that are coupled to outputs of the analog circuit 712. The pre-processing module 718 also has outputs that are coupled to inputs of the measurement module 720 and the calibration module 722. The measurement module 720 has other inputs that are respectively coupled to the microphone 710 and the circuitry 304. The calibration module 722 has an output that is coupled to the speaker 708.
[0073] Consider an example operation of the hearable 102 in accordance with single-ear audioplethysmography 110. In the case that the hearable 102 includes the calibration module 722, the hearable 102 can perform a calibration process prior to performing a measurement process. The calibration process and the measurement process are further described with respect to FIG. 10. [0074] During both the calibration process and the measurement process, the speaker 708 transmits the ultrasound transmit signal 802 and the microphone 710 receives the ultrasound receive signal 804. During the calibration process, the ultrasound transmit signal 802 and the ultrasound receive signal 804 can have tones 902-1 to 902-M, where M represents a positive integer. During the measurement process, the ultrasound transmit signal 802 and the ultrasound receive signal 804 can have selected tones 904-1 to 904-N, where N represents a positive integer that is less than or equal to M. The selected tones 904-1 to 904-N can represent a subset (sometimes a proper subset) of the tones 902-1 to 902-M. The microphone 710 can also receive the over-the-air signal 402 and the bone-conduction signal 408 during the measurement process. [0075] The analog circuit 712 performs analog-to-digital conversion to generate a digital transmit signal 906 and a digital receive signal 908 based on the ultrasound transmit signal 802 and the received ultrasound signal 506, respectively. The pre-processing module 718 performs frequency downconversion and demodulation to generate at least one pre-processed signal 910 based on the digital transmit signal 906 and the digital receive signal 908. The pre-processing module 718 can also apply filtering to generate the pre-processed signal 910.
[0076] As part of the calibration procedure, the calibration module 722 processes the pre- processed signal 910 to detennine the selected tones 904-1 to 904-N. The selected tones 904-1 to 904-N can improve performance of audioplethysmography 110 during the measurement procedure. The calibration module 722 communicates the selected tones 904-1 to 904-N to the speaker 708 using a control signal. The speaker 708 accepts the control signal that identifies the selected tones 904-1 to 904-N and can transmit a subsequent ultrasound transmit signal 802 for the measurement procedure using the selected tones 904-1 to 904-N.
[0077] As part of the measurement procedure, the measurement module 720 can perform aspects of intemal-noise-source filtering 202 using the pre-processed signal 910 and the noise signal 308. The measurement module 720 can also perform aspects of voice activity7 detection 112, speech recognition 114, and/or conversation detection 116 to generate voice data 912. The voice data 912 can also be referred to as speech data or audioplethysmography data. In cases in which the environment is noisy, the measurement module 720 can also utilize the received audible signal 504 provided by the microphone 710 to further process the pre-processed signal 910 for voice activity7 detection 112, speech recognition 114, and/or conversation detection 116. The voice data 912 can be communicated to the application 606, the voice user interface 608, and/or the voice authenticator 610. The voice data 912 can include an indication of whether or not the user 106’s voice is detected, a phrase that is recognized based on the voice component 404 of the pre-processed signal 910, a signal including the voice component 404, and so forth. Additionally or alternatively, the voice data 912 can include a control signal for controlling operation of the hearable 102 and/or the computing device 104. The calibration procedure and the measurement procedure are further described with respect to FIG. 10.
[0078] FIG. 10 illustrates an example flow diagram 1000 for operating a hearable 102. In FIG. 10, the hearable 102 can optionally perform a calibration procedure at 1002 using the calibration module 722. The calibration procedure can determine appropriate characteristics (e.g., waveform or signal characteristics) of ultrasound transmit signals 802 to improve audioplethysmography 110 (e.g., to enhance the performance of voice activity' detection 112). The calibration procedure enables audioplethysmography 110 to take into account the wear of the hearable 102 (e.g., the position of the hearable 102 relative to the ear canal 118) and the physical structure of the ear canal 118 to determine a transmission frequency that can increase sensitivity7. With the calibration procedure, the hearable 102 can dynamically adjust the transmission frequency (e.g., one or more carrier frequencies) each time the seal 120 is formed (e.g., based on the wear of the hearable 102) and based on the unique physical structure of the ear 108. Through this calibration procedure, the hearables 102 on different ears 108 may operate with one or more different ultrasound frequencies. Steps of the calibration procedure are further described below.
[0079] In some circumstances, the hearable 102 can perform on-head detection (or in-ear detection) by detecting the presence of the seal 120 and initiating the calibration procedure based on a determination that on-head detection is “true.” In other circumstances, the hearable 102 can initiate the calibration procedure based on a specified schedule or a timer, which can be controlled by the user 106 via the computing device 104.
[0080] At 1004, the hearable 102 executes the calibration procedure by transmitting and receiving a first ultrasound signal. The first ultrasound signal propagates within at least a portion of the ear canal 118 of the user 106 and has multiple tones 902-1 to 902-M (or multiple carrier frequencies). The multiple tones 902-1 to 902-M are transmitted in parallel or in series over a given time interv al. The first ultrasound transmit signal 802 can have a particular bandwidth on the order of several kilohertz. For example, the ultrasound transmit signal 802 can have a bandwidth of approximately 4, 5, 6, 8, 10, 16, or 20 kHz. In example implementations, the first ultrasound transmit signal 802 is transmitted over multiple seconds, such as 2, 3, 4, 6, or more seconds. A duration of each tone 902 can be evenly divided over a total duration of the first ultrasound transmit signal 802.
[0081] In an example implementation, the ultrasound transmit signal 902 has seven tones 902 (e.g., AT equals 7). In some cases, the tones 902 are evenly distributed across an interval. For example, the tones 902 can be in 1 kHz increments between 32 kHz and 38 kHz (e.g., at approximately 32, 33, 34, 35, 36, 37. and 38 kHz). The term “approximately” means that the tones 902 can be within 5% of a given value or less (e.g., within 3%, 2%, or 1% of the given value).
[0082] An amplitude of the ultrasound transmit signal 802 can be approximately the same across the tones 902-1 to 902-M. In this manner, power is evenly distributed across each tone 902. The quantity of tones 902 (e.g., M) can be determined based on an output power of the speaker 708. Increasing the quantity of tones 902 can increase a likelihood that the hearable 102 can support voice activity7 detection 112 across various conditions including user wear and a physical structure of the user 106’s ear canal 118. However, an amplitude of the ultrasound transmit signal 802 can be limited across these tones 902 based on the output power of the speaker 708. Thus, the quantity of tones 902 can be optimized based on an amount of output power that is available for audioplethysmography 110. [0083] At 1006, the calibration procedure selects one or more tones 904-1 to 904-N to be used for a measurement procedure based on one or more modified characteristics of the ultrasound receive signal 804. The process for selecting the tones 904 is further described with respect to FIG. 11. In general, the calibration procedure determines that the selected tones 904 improve a signal-to- noise ratio for audioplethysmography 110 (or more specifically for voice activity detection 1 12). [0084] At 1008, the hearable 102 performs a measurement procedure using the measurement module 720. In accordance with the measurement procedure, the hearable 102 transmits a second ultrasound transmit signal 802 that propagates within at least the portion of the ear canal 118 of the user 106. If the calibration procedure was performed, the second ultrasound transmit signal 802 can have the selected tones 904-1 to 904-M that were determined by the calibration procedure. The selected tones 904 can be transmitted in parallel or in series over a given time interval.
[0085] An amplitude of the second ultrasound transmit signal 802 can be approximately the same across the selected tones 904-1 to 904-N. In this manner, power is evenly distributed across each selected tone. The amplitude of the second ultrasound transmit signal 802 can be higher than the amplitude of the first ultrasound transmit signal 802 because the available output power is distributed across fewer tones. Additionally or alternatively, a duration of each of the selected tones 904 of the second ultrasound transmit signal 802 can be longer than the duration of the tones 902 of the first ultrasound transmit signal 802. The higher amplitude and/or the longer duration can further improve the signal -to-noise ratio performance of the hearable 102 for audioplethysmography 110. By using a few selected tones 904 that were determined to improve signal-to-noise ratio performance, the measurement procedure can achieve a higher accuracy for voice activity detection 112.
[0086] At 1012, the hearable 102 performs audioplethysmography (e.g., voice processing) using the second ultrasound signal (e.g., the second ultrasound receive signal 804). One aspect of performing audioplethysmography 110 can include performing internal -noise-source filtering 202. as further described with respect to FIGs. 12 and 13. The calibration module 722 is further described with respect to FIG. 11.
[0087] FIG. 11 illustrates an example scheme implemented by the calibration module 722. In the depicted configuration, the calibration module 722 implements a frequency selector, which selects one or more tones 904 for the measurement procedure. In the example implementation, the calibration module 722 includes at least one amplitude detector 1102, at least one phase detector 1104, at least one quality detector 1106, and at least one comparator 1 108. The operations of these components are further described below. [0088] During the calibration procedure, the calibration module 722 accepts the pre-processed signal 910 from the pre-processing module 718, as previously described with respect to FIG. 9. The pre-processed signal 910 can include amplitude and/or phase information associated with the multiple tones 902-1 to 902 -M, which were used to transmit the first ultrasound signal described at 1002 in FIG. 10.
[0089] In this example, the calibration module 722 extracts an amplitude 1 1 10 of the pre- processed signal 910 using the amplitude detector 1102 and extracts a phase 1112 of the pre- processed signal 910 using the phase detector 1104. Alternatively, if in-phase and quadrature components of the pre-processed signal 910 are received separately, the amplitude detector 1102 and the phase detector 1104 can respectively measure the amplitude 1110 and phase 1 112 based on the in-phase and quadrature components.
[0090] The quality detector 1106 measures quality metrics 1114-1 to 1114-2M for each of the tones 902-1 to 902-M and for each of the characteristics (e g., amplitude 1110 and phase 1112). In general, the quality metrics 1114 can represent a variety of different metrics, including peak- to-average ratios and/or signal-to-noise ratios. The peak-to-average ratio represents a peak intensity w ithin a frequency range of interest divided by an average intensity7 within this frequency range. A higher quality metric 1114 indicates a higher-quality signal, or more generally, better performance for audioplethysmography 110.
[0091] In one aspect, the comparator 1108 can evaluate the quality metrics 1114-1 to 1114-2M with respect to a threshold 1116. The threshold 1116 can be set, for example, to a particular value. In other cases, the calibration module 722 can dynamically determine the threshold 1116 and update it over time based on the observed quality metrics 1114-1 of 1114-2M. In an example implementation, the comparator 1108 determines the selected tones 904-1 to 904-N for a subsequent measurement procedure based on the frequencies associated with the quality metrics 1114-1 to 1114-M that are greater than or equal to the threshold 1116.
[0092] Additionally or alternatively, the comparator 1108 can evaluate the quality' metrics 1114-1 to 1114-2M with respect to each other. In an example implementation, the comparator 1108 determines one of the selected tones 904 based on a frequency with the highest quality metric 1114 across the amplitude 1110. Also, the comparator 1108 can determine one of the selected tones 904 based on a frequency with the highest quality metric 1114 across the phase 1112. In other implementations, the comparator 1108 can determine a single selected tone 904 based on a frequency having the highest quality metric 11 14 associated with either the amplitude 1 110 or the phase 1112. [0093] In general, the calibration module 722 enables the selected tones 904-1 to 904-N to be dynamically adjusted prior to the measurement procedure based on a current environment, which can account for a wear of the hearable 102 (e.g., a current insertion depth and/or rotation), a physical structure of the user 106’s ear canal 118, and a response characteristic of the hearable 102 (e.g., speaker, microphone, and/or housing). In this manner, the calibration module 722 can improve the signal-to-noise ratio performance of the hearable 102 for the measurement procedure. The calibration module 722 can also determine which tones 904 generate ultrasound receive signals 804 with desired characteristics for voice activity detection 112. In general, the calibration procedure can be performed whether or not the user 106 is speaking.
[0094] In FIGs. 9 to 11, the calibration procedure and the measurement procedure are described as individual procedures that occur at different time intervals. In particular, the calibration procedure occurs before the measurement procedure. This enables the ultrasound transmit signal 802 for the measurement procedure to be transmitted with fewer tones than the ultrasound transmit signal 802 used for the calibration procedure, which can increase signal-to-noise ratio performance for audioplethysmography 110. In some implementations, however, the hearable 102 can have sufficient output power to perform the measurement procedure with the multiple tones 902-1 to 902-M using a single ultrasound transmit signal 802. In this case, aspects of the calibration module can be integrated within the pre-processing module 718 as a frequency selector, which is further described with respect to FIG. 12. This frequency selector can effectively pass the selected tones 904-1 to 904-N for further processing. Aspects of the measurement procedure are further described with respect to FIG. 12.
Internal-Noise-Source Filtering
[0095] FIG. 12 illustrates an example implementation of the pre-processing module 718. In the depicted configuration, the hearable 102 includes the pre-processing module 718, which is coupled to the measurement module 720 and the calibration module 722. The measurement module 720 is also coupled to the microphone 710 (not shown) and the circuitry’ 304 (not shown). [0096] The pre-processing module 718 includes at least one in-phase and quadrature mixer 1202 (I/Q mixer 1202) and at least one filter 1204. The in-phase and quadrature mixer 1202 performs frequency down-conversion. In an example implementation, the in-phase and quadrature mixer 1202 includes at least two mixers, at least one phase shifter, and at least one combiner (e.g., a summation circuit). The filter 1204 attenuates intermodulation products that are generated by the in-phase and quadrature mixer 1202. In an example implementation, the filter 1204 is implemented using a low -pass filter. [0097] The pre-processing module 718 can optionally include at least one frequency selector 1206. The frequency selector 1206 can identify and select one or more tones 904 (or carrier frequencies) that provide a high-quality signal for later processing. The frequency selector 1206 can further pass the selected tones 904 to other processing modules and filter (or attenuate) other tones that are not selected. The frequency selector 1206 can be implemented in a similar manner as the calibration module 722 of FIG. 1 1. For example, the frequency selector 1206, can include the amplitude detector 1102, the phase detector 1104, the qualify detector 1106, and the comparator 1108.
[0098] During an operation, the in-phase and quadrature mixer 1202 uses the phase shifter and the two mixers to generate in-phase and quadrature components associated with the digital receive signal 908. In particular, the in-phase and quadrature mixer 1202 mixes the digital receive signal 908 with a first version of the digital transmit signal 906 that has a zero-degree phase shift to generate the in-phase component. Additionally, the in-phase and quadrature mixer 1202 mixes the digital receive signal 908 with a second version of the digital transmit signal 906 that has a 180-degree phase shift to generate the quadrature signal. This mixing operation downconverts the digital receive signal 908 from acoustic frequencies to baseband frequencies. Using the combiner, the in-phase and quadrature mixer 1202 combines the in-phase and quadrature components of the digital receive signal 908 to generate a down-converted signal 1208. Use of the in-phase and quadrature mixer 1202 can further improve the signal-to-noise ratio of the down-converted signal 1208 compared to other mixing techniques.
[0099] In this example, the down-converted signal 1208 represents a combination of the in-phase and quadrature components of the mixed-down digital receive signal 908. In alternative implementations, the in-phase and quadrature mixer 1202 doesn’t include the combiner and passes the in-phase and quadrature components separately to the filter 1204. In this manner, the in-phase and quadrature components individually propagate through the filter 1204.
[0100] The filter 1204 generates a filtered signal 1210 based on the down-converted signal 1208. In particular, the filter 1204 filters the down-converted signal 1208 to attenuate spurious or undesired frequencies (e g., intermodulation products), some of which can be associated with an operation of the in-phase and quadrature mixer 1202. In this example, the filtered signal 1210 represents a combination of the in-phase and quadrature components of the down-converted signal 1208. Alternatively, the filtered signal 1210 can represent separate or distinct in-phase and quadrature components, which are individually passed to the frequency selector 1206, the calibration module 722, or the measurement module 720. [0101] During the measurement procedure, the pre-processing module 718 can optionally apply the frequency selector 1206. The frequency selector 1206 passes tones that meet a quality threshold level of performance for audioplethysmography 110. For example, the frequency selector 1206 passes tones 904 having an amplitude 1110 and/or phase 1112 with a quality metric 1114 that is greater than or equal to a threshold 1116. The resulting signal outputted by the frequency selector 1206 is represented by signal 1212. In some implementations, this signal 1212 is passed to the measurement module 720 as the pre-processed signal 910. In other implementations in which the frequency selector 1206 is not implemented, the filtered signal 1210 can be passed to the measurement module 720 and/or the calibration module 722 as the pre- processed signal 910.
[0102] In general, the measurement module 720 can generate the voice data 912 based on the pre- processed signal 910. In this case, the measurement module 720 can analyze the changes in the amplitude 1110 and/or phase 1112 of the voice component 404 of the pre-processed signal 910 for voice processing. This processing technique can be utilized in implementations of the hearable 102 that have minimal (if any) interference betw een the over-the-air signal 402 and the ultrasound signal 410, or in situations in which audioplethysmography 110 is performed in a relatively quiet environment.
[0103] To handle noisy environments, however, the measurement module 720 can generate the voice data 912 based on the pre-processed signal 910 and the received audible signal 504. More specifically, the measurement module 720 can utilize the received audible signal 504 as a reference to attenuate the modulation component 510 within the pre-processed signal 910 and improve the signal-to-noise-ratio associated with the voice component 404.
[0104] To handle situations in which an internal noise source 312 operates during audioplethysmography 110, the measurement module 720 can also apply intemal-noise-source filtering 202. With intemal-noise-source filtering 202, the measurement module 720 utilizes the noise signal 308 as a reference to attenuate the internal noise component 310 within the pre- processed signal 910 and/or within the received audible signal 504. An example implementation of the measurement module 720, which can perform intemal-noise-source filtering 202, is further described with respect to FIG. 13.
[0105] FIG. 13 illustrates an example implementation of the measurement module 720 for performing intemal-noise-source filtering 202 and voice processing. In the depicted configuration, the measurement module 720 includes at least one multi-stage filter 1302 and at least one voice processor 1304 (or speech processor). The multi-stage filter 1302 is coupled between the voice processor 1304 and the pre-processing module 718. Although not explicitly shown, the voice processor 1304 can be coupled to other components of the hearable 102, such as the communication interface 704.
[0106] The voice processor 1304 can perform aspects of voice activity detection 112, speech recognition 114, and/or conversation detection 116. In example implementations, the voice processor 1304 can be implemented using a machine-learned model or another module that performs signal and/or data processing. In general, the voice processor 1304 can analyze changes in the amplitude 1110 and/or phase 1112 of the voice component 404 of the pre-processed signal 910 for voice processing.
[0107] To enable the voice processor 1304 to detect and/or utilize the voice component 404, the multi-stage filter 1302 filters the pre-processed signal 910 to improve the signal-to-noise ratio associated with the voice component 404. In this example, the multi-stage filter 1302 includes at least one intemal-noise-source filter stage 1306 and at least one extemal-noise-source filter stage 1308. Other implementations are also possible in which the multi-stage filter 1302 includes the intemal-noise-source filter stage 1306 and does not include the extemal-noise-source filter stage 1308. This implementation is possible in situations in which audioplethysmography is performed in a relatively quiet environment and/or for implementations in which a design of the hearable 102 minimizes the impact of the modulation component 510.
[0108] The intemal-noise-source filter stage 1306 performs intemal-noise-source filtering 202 to attenuate the internal noise component 310 within the pre-processing signal 910 and/or the received audible signal 504. The intemal-noise-source filter stage 1306 can be implemented using at least two adaptive filters 1310. An example implementation of the intemal-noise-source filter stage 1306 is further described with respect to FIG. 14. Other implementations are also possible in which the intemal-noise-source filter stage 1306 is implemented using a machine-learned model.
[0109] The extemal-noise-source filter stage 1306 uses the received audible signal 504 to attenuate the modulation component 510 within the pre-processed signal 910. The extemal-noise- source filter stage 1308 can be implemented using at least one filter module 1312. An example implementation of the extemal-noise-source filter stage 1308 is further described with respect to FIG. 15. Other implementations are also possible in which the extemal-noise-source filter stage 1308 is implemented using a machine-learned model.
[0110] During an operation, the measurement module 720 accepts the pre-processed signal 910 from the pre-processing module 718. The pre-processed signal 910 can include the internal noise component 310 and the voice component 404. In some implementations, the pre-processed signal 910 also includes the modulation component 510. The internal noise component 310 and the modulation component 510 can make it challenging for the pre-processed signal 910 to be directly used for voice processing. In some cases, the internal noise component 310 can have a larger impact on a noise level of the pre-processed signal 910 relative to the modulation component 510 due to the noise signal 308 being amplified within the ear canal 118.
[0111] The measurement module 720 can also accept the received audible signal 504 from the microphone 710. The received audible signal 504 can include the internal noise component 310, the voice component 404, and the bone-conduction component 512. In noisy environments, the received audible signal 504 can also include the noise component 406.
[0112] The intemal-noise-source filter stage 1306 operates on the pre-processed signal 910 and optionally the received audible signal 504 to attenuate the internal noise component 310. The external -noise-source filter stage 1308 operates on outputs of the internal -noise-source-filter stage 1308 to attenuate the modulation component 510 within the pre-processed signal 910. An output of the extemal-noise-source filter stage 1308 is represented by a denoised signal 1314. The denoised signal 1314 represents a filtered version of the pre-processed signal 910 that includes the voice component 404.
[0113] The voice processor 1304 analyzes the denoised signal 1314 to generate the voice data 912. In particular, the voice processor 1304 can generate the voice data 912 associated with voice activity detection 112, speech recognition 114. and/or conversation detection 116. For voice activity detection 112, the voice data 912 can indicate whether or not the voice component 404 is detected. In some examples, the voice processor 1304 can perform a signal-to-noise ratio detection process, which determines whether or not an amplitude of an input signal exceeds a detection threshold. If the amplitude exceeds the detection threshold, the voice processor 1304 generates the voice data 912 to indicate that the vocalization is detected. Otherwise, if the amplitude does not exceed the detection threshold, the voice processor 1304 generates the voice data 912 to indicate that a vocalization is not detected. Voice activity detection 112, speech recognition 114, and/or conversation detection 116 can be utilized in a variety of different ways to control an operation of the hearable 102 and/or the computing device 104. An example implementation of the multi-stage filter 1302 is further described with respect to FIGs. 14 and 15. [0114] FIG. 14 illustrates an example implementation of the intemal-noise-source filter stage 1306. In the depicted configuration, the intemal-noise-source filter stage 1306 includes two adaptive filters 1310-1 and 1310-2. The adaptive filters 1310-1 and 1310-2 can apply various adaptive filtering techniques, including techniques based on least mean squares (LMS) or recursive least squares (RLS). The adaptive filter 1310-1 performs adaptive filtering using the noise signal 308 as a noise reference to attenuate the internal noise component 310 within the pre- processed signal 910. Likewise, the adaptive filter 1310-2 performs adaptive filtering using the noise signal 308 as a noise reference to attenuate the internal noise component 310 within the received audible signal 504. In this example, the noise signal 308 is not correlated with the desired voice component 404 within the pre-processed signal 910 and the received audible signal 504. As such, the noise signal 308 can be used as a reference signal to attenuate the internal noise component 310 within the pre-processed signal 910 and the received audible signal 504 using adaptive filtering techniques.
[0115] The adaptive filters 1310-1 and 1310-2 respectively generate a denoised pre-processed signal 1402 and a denoised audible signal 1404. The denoised pre-processed signal 1402 includes the voice component 404 and the modulation component 510. The denoised audible signal 1404 signal 1404 includes the voice component 404, the noise component 406, and the bone-conduction component 512. The denoised audible signal 1404 can also be referred to as a denoised signal. The tenn “denoised audible signal 1404” is used to represent a denoised version of the received audible signal 504 and does not necessarily mean that the denoised audible signal 1404 can be heard.
[0116] Generally speaking, the denoised pre-processed signal 1402 and the denoised audible signal 1404 do not include the internal noise component 310 or the internal noise component 310 is substantially attenuated relative to the corresponding input signal. To attenuate the modulation component 510 within the denoised pre-processed signal 1402, the filtering process continues with the extemal-noise-source filter stage 1306, which is further described with respect to FIG. 15.
[0117] FIG. 15 illustrates an example implementation of the extemal-noise-source filter stage 1306. In the depicted configuration, the extemal-noise-source filter stage 1306 includes at least one filter module 1502 and at least one vocalization enhancer 1504. The filter module 1502 can be implemented using at least one adaptive filter 1506 or at least one blind-source separator 1508.
[0118] The adaptive filter 1506 performs adaptive filtering using the received audible signal 504 as a noise reference to separate the voice component 404 of the pre-processed signal 910 from the modulation component 510 (or from the modulated noise component 406). Example adaptive filtering techniques utilized by the adaptive filter 1506 can include techniques based on least mean squares (LMS) or recursive least squares (RLS). The blind-source separator 1508 performs blindsource separation (BSS) using the received audible signal 504 to separate the voice component 404 of the pre-processed signal 910 from the modulation component 510. To perform adaptive filtering or blind-source separation (BSS), the pre-processed signal 910 represents a primary reference and the received audible signal 504 represents a secondary or a noise reference. In general, the adaptive filter 1506 and the blind-source separator 1508 can utilize the received audible signal 504 to significantly attenuate the modulation component 510 (or the modulated noise component 406) within the pre-processed signal 910.
[0119] The vocalization enhancer 1504 enhances (e.g., amplifies relative to a noise level) the voice component 404 within an output signal provided by the filter module 1502. In this way, the vocalization enhancer 1504 can increase sensitivity for voice processing. In an example implementation, the vocalization enhancer 1504 is implemented using a wiener filter 1510.
Example Methods
[0120] FIGs. 16 and 17 depict example methods 1600 and 1700 for implementing aspects of voice activity detection 112 using active acoustic sensing. Methods 1600 and 1700 are shown as sets of operations (or acts) performed but not necessarily limited to the order or combinations in which the operations are shown herein. Further, any of one or more of the operations may be repeated, combined, reorganized, or linked to provide a wide array of additional and/or alternate methods. In portions of the following discussion, reference may be made to the environments 100, 200-1, 200-2, 200-3 of FIGs. 1 and 2, and entities detailed in FIGs. 6 and 7, reference to which is made for example only. The techniques are not limited to performance by one entity or multiple entities operating on one device.
[0121] At 1602, an audible signal is transmitted during a first time period. The audible signal is intended to propagate within at least a portion of an ear canal of a user. For example, the circuitry7 304 transmits (or renders) the audible signal during the first time period. The audible signal is intended to propagate within at least a portion of an ear canal 118 of the user 106. The circuitry 304 can include a speaker 314 or 708, the active-noise-cancellation circuitry 316, or the transparency-mode circuitry 318. The audible signal can be used to play audio content 204 for the user 106 as shown in the environment 200-1, provide active noise cancellation 206 as shown in the environment 200-2, or enable the hearable 102 to operate in accordance w ith a transparency mode 214 as shown in the environment 200-3 of FIG. 2.
[0122] At 1604, an ultrasound transmit signal is transmitted during the first time period. The ultrasound transmit signal is intended to propagate within at least a portion of the ear canal of the user. For example, the transducer 706 (or speaker 708) of the hearable 102 transmits the ultrasound transmit signal 802. The ultrasound transmit signal 802 propagates within at least a portion of the ear canal 118 of the user 106, as described with respect to FIGs. 4 and 8. The ultrasound transmit signal 802 can have at least two tones 904 (or frequencies). In some implementations, the multiple tones 904 can be determined based on a calibration procedure, as described at 1002 in FIG. 10. In other implementations, the multiple tones 904 provide frequency diversity and can improve performance for audioplethysmography 110, as described with respect to FIG. 12.
[0123] At 1606, an ultrasound receive signal is received. The ultrasound receive signal represents a version of the ultrasound transmit signal with one or more waveform characteristics modified based on the propagation within the ear canal and based on a vocalization made by the user during the first time period. The ultrasound receive signal comprises an internal noise component caused by interference generated by the rendering of the audible signal.
[0124] For example, the transducer 706 (or the microphone 710) of the hearable 102 receives the ultrasound receive signal 804. The ultrasound receive signal 804 represents a version of the ultrasound transmit signal 802 with one or more waveform characteristics modified based on the propagation within the ear canal 118 and based on a vocalization made by the user 106 during the first time period. Example waveform characteristics include amplitude, phase, and/or frequency. [0125] The ultrasound receive signal 804 includes the internal noise component 310, which is caused by interference generated by the rendering of the audible signal. The internal noise component 310 can represent a portion of the audible signal that is modulated onto or mixed with the ultrasound receive signal 804. In general, the internal noise component 310 represents a portion of the ultrasound receive signal 804 (e.g.. amplitude, phase, and/or frequency) that is modified by the interference associated with the audible signal.
[0126] The hearable 102 that receives the ultrasound receive signal 804 can be a same hearable 102 that transmitted the ultrasound transmit signal 802 (e.g., the hearable 102-1 or 102-2 in FIG. 8), or another hearable 102 that did not transmit the ultrasound transmit signal 802 (e.g., the hearable 102-2 in FIG. 8). In some implementations, a feedback microphone of an activenoise-cancellation circuit 724 can receive the ultrasound receive signal 804.
[0127] At 1608, a denoised signal is generated by filtering the internal noise component within the received ultrasound receive signal based on a version of the audible signal. For example, the measurement module 720 (e.g., the multi-stage filter 1302) generates the denoised signal 1314 by filtering the internal noise component 310 within the ultrasound receive signal 804 based on a version of the audible signal, which is represented by the noise signal 308 as shown in FIG. 13.
[0128] The version of the audible signal can represent a digital version of the audible signal, a version of the audible signal that is downconverted or upconverted to a particular frequency range for denoising (e.g., a baseband frequency), a pre-processed version of the audible signal, or some combination thereof. In general, the term “the version of the audible signal” means there can be some differences between the playing of the audible signal for the user 106 and the providing of the audible signal as the noise signal 308 for the intemal-noise-source filtering process. Explained another way, the version of the audible signal is related to the audible signal, but can be modified to support an operation of the hearable 102. In this case, the operation involves intemal-noise- source filtering.
[0129] At 1610, the vocalization is detected based on the denoised signal. For example, the hearable 102 uses the measurement module 720 (e.g., the voice processor 1304) to detect the vocalization based on the denoised signal 1314. More specifically, the measurement module 720 analyzes the voice component 404 that is present within the denoised signal 1314 to perform voice activity detection 112, speech recognition 114, and/or conversation detection 116. The measurement module 720 can generate voice data 912, which can be used to control the hearable 102 and/or the computing device 104.
[0130] At 1702 in FIG. 17, active acoustic sensing is performed to detect a pressure wave that propagates within an ear canal of a user and is associated with a vocalization of the user. For example, the hearable 102 performs active acoustic sensing to detect a pressure wave that propagates within an ear canal 118 of a user 106 and is associated with a vocalization of the user 106. To perform active acoustic sensing, the hearable 102 transmits and receives an ultrasound signal 410 (e.g., the ultrasound transmit signal 802 and the ultrasound receive signal 804). The received ultrasound signal 410 includes the voice component 404, which enables audioplethysmography 1 10 to be used for voice processing.
[0131] At 1704, an audible signal that interferes with active acoustic sensing is rendered. For example, circuitry 304 renders an audible signal that interferes with the active acoustic sensing. The circuitry 304 can include a speaker 314 or 708, active-noise-cancellation circuitry 316, and/or transparency-mode circuitry 318. The audible signal can be an audible signal 320 that includes audio content 204, the anti-noise signal 324, and/or the transparency-mode signal 326.
[0132] At 1706, intemal-noise-source filtering is performed to attenuate the interference caused by the rendering of the audible signal. For example, the measurement module 720 performs intemal-noise-source filtering 202 to attenuate the interference (e.g., the internal noise component 310) caused by the rendering of the audible signal. This interference can be attenuated within the pre-processed signal 910 and optionally the received audible signal 504, as shown in FIG. 14.
[0133] At 1708, voice processing is performed based on the active acoustic sensing and the intemal-noise-source filtering. For example, the voice processor 1304 performs voice processing (e.g., voice activity detection 112, speech recognition 114, and/or conversation detection 116) based on the active acoustic sensing and the intemal-noise-source filtering (e.g., based on the denoised pre-processed signal 1402).
[0134] At 1710, a signal that controls an operation of at least one of a hearable or a computing device that is coupled to the hearable is generated. For example, the measurement module 720 generates a voice data 912, which can be used to control an operation of the hearable 102 and/or the computing device 104.
[0135] Throughout the disclosure, the term “version of a signal” is used to indicate that a second signal can be a modified version of a first signal. The “version of the signal” (or the second signal) can represent a digital version of the first signal, an analog version of the first signal, an electrical version of the first signal having a voltage and current, an acoustic version of the first signal having acoustic properties, a downconverted version of the first signal having a lower range of frequencies, an upconverted version of the signal having a higher range of frequencies, a pre- processed version of the signal (e.g., a version whose amplitude, phase, and/or frequency has been modified in some way), a filtered version of the signal, and so forth. In general, the “version of the signal” represents a signal that has been modified using techniques know n in the art to facilitate an operation of the hearable 102.
Example Computing System
[0136] FIG. 18 illustrates various components of an example computing system 1800 that can be implemented as any type of client, server, and/or computing device as described with reference to the previous FIGs. 6 and 7 to implement aspects of active acoustic sensing using a hearable.
[0137] The computing system 1800 includes communication devices 1802 that enable wired and/or wireless communication of device data 1804 (e.g., received data, data that is being received, data scheduled for broadcast, or data packets of the data). The communication devices 1802 or the computing system 1800 can include one or more hearables 102. The device data 1804 or other device content can include configuration settings of the device, media content stored on the device, and/or information associated with a user of the device. Media content stored on the computing system 1800 can include any type of audio, video, and/or image data. The computing system 1800 includes one or more data inputs 1806 via which any type of data, media content, and/or inputs can be received, such as human utterances, user-selectable inputs (explicit or implicit), messages, music, television media content, recorded video content, and any other type of audio, video, and/or image data received from any content and/or data source.
[0138] The computing system 1800 also includes communication interfaces 1808, which can be implemented as any one or more of a serial and/or parallel interface, a wireless interface, any ty pe of network interface, a modem, and as any other type of communication interface. The communication interfaces 1808 provide a connection and/or communication links between the computing system 1800 and a communication network by which other electronic, computing, and communication devices communicate data with the computing system 1800.
[0139] The computing system 1800 includes one or more processors 1810 (e.g., any of microprocessors, controllers, and the like), which process various computer-executable instructions to control the operation of the computing system 1800. Alternatively or in addition, the computing system 1800 can be implemented with any one or combination of hardware, firmware, or fixed logic circuitry that is implemented in connection with processing and control circuits which are generally identified at 1812. Although not shown, the computing system 1800 can include a system bus or data transfer system that couples the various components within the device. A system bus can include any one or combination of different bus structures, such as a memory bus or memory controller, a peripheral bus, a universal serial bus, and/or a processor or local bus that utilizes any of a variety of bus architectures.
[0140] The computing system 1800 also includes a computer-readable medium 1814, such as one or more memory devices that enable persistent and/or non-transitory data storage (i.e., in contrast to mere signal transmission), examples of which include random access memory' (RAM), non-volatile memory (e.g., any one or more of a read-only memory (ROM), flash memory, EPROM, EEPROM, etc ), and a disk storage device. The disk storage device may be implemented as any type of magnetic or optical storage device, such as a hard disk drive, a recordable and/or rewriteable compact disc (CD), any type of a digital versatile disc (DVD), and the like. The computing system 1800 can also include a mass storage medium device (storage medium) 1816.
[0141] The computer-readable medium 1814 provides data storage mechanisms to store the device data 1804, as well as various device applications 1818 and any other types of information and/or data related to operational aspects of the computing system 1800. For example, an operating system 1820 can be maintained as a computer application with the computer-readable medium 1814 and executed on the processors 1810. The device applications 1818 may include a device manager, such as any form of a control application, software application, signal-processing and control module, code that is native to a particular device, a hardware abstraction layer for a particular device, and so on.
[0142] The device applications 1818 also include any system components, engines, or managers to implement intemal-noise-source filtering 202 for active acoustic sensing. In this example, the device applications 1818 include the pre-processing module 718, the measurement module 720, and optionally the calibration module 722. Although not explicitly shown, the device applications 1818 can also include the intemal-noise-source filter stage 1306, the application 606, the voice user interface 608, and/or the voice authenticator 610.
[0143] Throughout this disclosure, examples are described where a computing system 1800 (e.g., the hearable 102, the computing device 104, a client device, a server device, a computer, or another type of computing system) may analyze information (e.g., various audible and/or ultrasound signals) associated with a user, for example, a vocalization. Further to the descriptions above, a user 106 may be provided with controls allowing the user 106 to make an election as to both if and when systems, programs, and/or features described herein may enable collection of information (e.g., information about a user’s social network, social actions, social activities, profession, a user’s preferences, a user’s current location), and if the user 106 is sent content or communications from a server. The computing system 1800 can be configured to only use the information after the computing system 1800 receives explicit permission from the user 106 to use the data. For example, in situations where the hearable 102 analyzes signals for voice activity detection 112, speech recognition 114, and/or conversation detection 116, individual users 106 may be provided with an opportunity to provide input to control whether programs or features of the computing system 1800 can collect and make use of the data. Further, individual users 106 may have constant control over what programs can or cannot do with the information.
[0144] In addition, information collected may be pre-treated in one or more ways before it is transferred, stored, or otherwise used, so that personally-identifiable information is removed. For example, before the computing system 1800 shares data with another device, a user 106’s identity may be treated so that no personally identifiable information can be determined for the user 106. Thus, the user 106 may have control over whether information is collected about the user 106 and the user 106’s device, and how such information, if collected, may be used by the computing system 1800 and/or a remote computing system.
Conclusion
[0145] Although techniques using, and apparatuses including, performing intemal-noise-source filtering for active acoustic sensing have been described in language specific to features and/or methods, it is to be understood that the subject of the appended examples is not necessarily limited to the specific features or methods described. Rather, the specific features and methods are disclosed as example implementations of intemal-noise-source filtering for active acoustic sensing. [0146] Some examples are provided below.
[0147] Example 1: A method comprising: transmitting, during a first time period, an audible signal that propagates within at least a portion of an ear canal of a user; transmitting, during the first time period, an ultrasound transmit signal that propagates within at least a portion of the ear canal of the user; receiving, during the first time period, an ultrasound receive signal, the ultrasound receive signal representing a version of the ultrasound transmit signal with one or more characteristics modified based on the propagation within the ear canal and based on a vocalization made by the user during the first time period, the received ultrasound receive signal comprising an internal noise component caused by interference generated by the rendering of the audible signal; generating a denoised signal by filtering the internal noise component within the received ultrasound receive signal based on a version of the audible signal; and detecting the vocalization based on the denoised signal.
[0148] Example 2: The method of example 1, further comprising: controlling an operation of a hearable and/or an operation of a computing device that is coupled to the hearable based on the detected vocalization.
[0149] Example 3: The method of example 1 or 2, wherein: the detecting the vocalization comprises performing voice processing using the denoised signal; and the method further comprises controlling, based on the voice processing, an operation of a hearable and/or an operation of a computing device that is coupled to the hearable.
[0150] Example 4: The method of example 3. wherein the controlling of the hearable comprises at least one of the following: pausing or resuming the rendering of audio content; adjusting a volume of the audio content being rendered; enabling or disabling active noise cancellation; or enabling or disabling a transparency mode. [0151] Example 5: The method of example 3 or 4, wherein the performing of the voice processing comprises at least one of the following: performing voice activity detection; performing speech recognition; or performing conversation detection.
[0152] Example 6: The method of any previous example, wherein the audible signal comprises audio content requested by the user.
[0153] Example 7: The method of example 6, wherein the audio content comprises music.
[0154] Example 8: The method of any previous example, further comprising: performing active noise cancellation using the audible signal, wherein the audible signal comprises an anti-noise signal.
[0155] Example 9: The method of any previous example, further comprising: operating a hearable in accordance with a transparency mode, wherein the audible signal comprises a transparency-mode signal that provides sound from an external environment to the ear canal of the user.
[0156] Example 10: The method of any previous example, wherein the generating of the denoised signal comprises filtering a version of the received ultrasound receive signal based on a version of the audible signal using an adaptive filter.
[0157] Example 11 : The method of example 10, further comprising: receiving, during the first time period, an over-the-air signal comprising audible frequencies, the audible frequencies comprising the vocalization made by the user during the first time period, wherein: the received ultrasound receive signal comprises a modulation component caused by interference generated by the receiving of the over-the-air signal; and the generating of the denoised signal further comprises attenuating the modulation component within the received ultrasound receive signal using the received over-the-air signal. [0158] Example 12: The method of example 11, wherein: the over-the-air signal comprises the internal noise component; and the generating of the denoised signal comprises: generating a denoised version of the received ultrasound receive signal by filtering the version of the received ultrasound receive signal based on the version of the audible signal using the adaptive filter; generating a denoised audible signal by filtering a version of the over-the-air signal based on the version of the audible signal using another adaptive filter; and generating the denoised signal by filtering the denoised version of the received ultrasound receive signal based on the denoised audible signal.
[0159] Example 13: The method of any one of examples 11 or 12, wherein the receiving of the ultrasound receive signal and the receiving of the over-the-air signal comprises receiving the ultrasound receive signal and the over-the-air signal using a same microphone.
[0160] Example 14: The method of any previous example, wherein the vocalization comprises speech, humming, whistling, or singing.
[0161] Example 15: The method of any previous example, wherein the transmitting of the ultrasound transmit signal comprising transmitting the ultrasound transmit signal with at least two tones.
[0162] Example 16: A computer-readable storage medium comprising instructions that, responsive to execution by a processor, cause a hearable to perform any one of the methods of examples 1 to 15.
[0163] Example 17: A device comprising: at least one transducer; and at least one processor, the device configured to perform, using the at least one transducer and the at least one processor, any one of the methods of examples 1 to 15. [0164] Example 18: The device of example 17, further comprising: a speaker; and an active-noise-cancellation circuit comprising a feedback microphone, wherein the at least one transducer comprises the speaker and the feedback microphone.
[0165] Example 19: The device of example 17, wherein: the at least one transducer comprises a speaker and a microphone; the speaker is configured to be positioned proximate to a first ear of a user; and the microphone is configured to be positioned proximate to a second ear of the user.
[0166] Example 20: The device of any one of examples 17 to 19, wherein the device comprises: at least one earbud; or headphones.

Claims

CLAIMS What is claimed is:
1. A method comprising: transmitting, during a first time period, an audible signal that propagates within at least a portion of an ear canal of a user; transmitting, during the first time period, an ultrasound transmit signal that propagates within at least a portion of the ear canal of the user; receiving, during the first time period, an ultrasound receive signal, the ultrasound receive signal representing a version of the ultrasound transmit signal with one or more characteristics modified based on the propagation within the ear canal and based on a vocalization made by the user during the first time period, the received ultrasound receive signal comprising an internal noise component caused by interference generated by the rendering of the audible signal; generating a denoised signal by filtering the internal noise component within the received ultrasound receive signal based on a version of the audible signal; and detecting the vocalization based on the denoised signal.
2. The method of claim 1, further comprising: controlling an operation of a hearable and/or an operation of a computing device that is coupled to the hearable based on the detected vocalization.
3. The method of claim 1 or 2, wherein: the detecting the vocalization comprises performing voice processing using the denoised signal; and the method further comprises controlling, based on the voice processing, an operation of a hearable and/or an operation of a computing device that is coupled to the hearable.
4. The method of claim 3, wherein the controlling of the hearable comprises at least one of the following: pausing or resuming the rendering of audio content; adjusting a volume of the audio content being rendered; enabling or disabling active noise cancellation; or enabling or disabling a transparency mode.
5. The method of claim 3 or 4, wherein the performing of the voice processing comprises at least one of the following: performing voice activity detection; performing speech recognition; or performing conversation detection.
6. The method of any previous claim, wherein the audible signal comprises audio content requested by the user.
7. The method of claim 6, wherein the audio content comprises music.
8. The method of any previous claim, further comprising: performing active noise cancellation using the audible signal, wherein the audible signal comprises an anti-noise signal.
9. The method of any previous claim, further comprising: operating a hearable in accordance with a transparency mode, wherein the audible signal comprises a transparency-mode signal that provides sound from an external environment to the ear canal of the user.
10. The method of any previous claim, wherein the generating of the denoised signal comprises filtering a version of the received ultrasound receive signal based on a version of the audible signal using an adaptive filter.
11. The method of claim 10, further comprising: receiving, during the first time period, an over-the-air signal comprising audible frequencies, the audible frequencies comprising the vocalization made by the user during the first time period, wherein: the received ultrasound receive signal comprises a modulation component caused by interference generated by the receiving of the over-the-air signal; and the generating of the denoised signal further comprises attenuating the modulation component within the received ultrasound receive signal using the received over-the-air signal.
12. The method of claim 11, wherein: the over-the-air signal comprises the internal noise component; and the generating of the denoised signal comprises: generating a denoised version of the received ultrasound receive signal by filtering the version of the received ultrasound receive signal based on the version of the audible signal using the adaptive filter; generating a denoised audible signal by filtering a version of the over-the-air signal based on the version of the audible signal using another adaptive filter; and generating the denoised signal by filtering the denoised version of the received ultrasound receive signal based on the denoised audible signal.
13. The method of any one of claims 11 or 12, wherein the receiving of the ultrasound receive signal and the receiving of the over-the-air signal comprises receiving the ultrasound receive signal and the over-the-air signal using a same microphone.
14. The method of any previous claim, wherein the vocalization comprises speech, humming, whistling, or singing.
15. The method of any previous claim, wherein the transmitting of the ultrasound transmit signal comprising transmitting the ultrasound transmit signal with at least two tones.
16. A computer-readable storage medium comprising instructions that, responsive to execution by a processor, cause a hearable to perform any one of the methods of claims 1 to 15.
17. A device comprising: at least one transducer; and at least one processor, the device configured to perform, using the at least one transducer and the at least one processor, any one of the methods of claims 1 to 15.
18. The device of claim 17, further comprising: a speaker; and an active-noise-cancellation circuit comprising a feedback microphone, wherein the at least one transducer comprises the speaker and the feedback microphone.
19. The device of claim 17, wherein: the at least one transducer comprises a speaker and a microphone; the speaker is configured to be positioned proximate to a first ear of a user; and the microphone is configured to be positioned proximate to a second ear of the user.
20. The device of any one of claims 17 to 19, wherein the device comprises: at least one earbud; or headphones.
EP23848366.3A 2023-12-29 2023-12-29 Internal-noise-source filtering for active acoustic sensing Active EP4599601B1 (en)

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
PCT/US2023/086468 WO2025144418A1 (en) 2023-12-29 2023-12-29 Internal-noise-source filtering for active acoustic sensing

Publications (2)

Publication Number Publication Date
EP4599601A1 true EP4599601A1 (en) 2025-08-13
EP4599601B1 EP4599601B1 (en) 2026-02-04

Family

ID=89833826

Family Applications (1)

Application Number Title Priority Date Filing Date
EP23848366.3A Active EP4599601B1 (en) 2023-12-29 2023-12-29 Internal-noise-source filtering for active acoustic sensing

Country Status (4)

Country Link
EP (1) EP4599601B1 (en)
JP (1) JP7854106B2 (en)
CN (1) CN120569982A (en)
WO (1) WO2025144418A1 (en)

Families Citing this family (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US12562185B2 (en) * 2023-03-16 2026-02-24 Google Llc Voice activity detection using active acoustic sensing
CN121034334B (en) * 2025-09-05 2026-03-24 长沙幻音科技有限公司 Resolution improving method and system based on audio frequency spectrum analysis

Family Cites Families (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2005048572A2 (en) * 2003-11-11 2005-05-26 Matech, Inc. Two-way communications device having a single transducer
JP4568905B2 (en) * 2004-11-12 2010-10-27 株式会社ケンウッド Microphone device and speech detection device
EP4536070A1 (en) * 2022-06-10 2025-04-16 Google LLC Active acoustic sensing

Also Published As

Publication number Publication date
EP4599601B1 (en) 2026-02-04
CN120569982A (en) 2025-08-29
JP2026504759A (en) 2026-02-10
WO2025144418A1 (en) 2025-07-03
JP7854106B2 (en) 2026-04-30

Similar Documents

Publication Publication Date Title
US12562185B2 (en) Voice activity detection using active acoustic sensing
US11856366B2 (en) Methods and apparatuses for driving audio and ultrasonic signals from the same transducer
EP4599601B1 (en) Internal-noise-source filtering for active acoustic sensing
KR20120034085A (en) Earphone arrangement and method of operation therefor
US20250318800A1 (en) Active Acoustic Sensing
US10878796B2 (en) Mobile platform based active noise cancellation (ANC)
CN116491131A (en) Active Self-Speech Naturalization Using Bone Conduction Sensors
JP5526060B2 (en) Hearing aid adjustment device
US20250082300A1 (en) Respiration Rate Sensing
WO2023240240A1 (en) Audioplethysmography calibration
US11696065B2 (en) Adaptive active noise cancellation based on movement
WO2023283285A1 (en) Wearable audio device with enhanced voice pick-up
WO2024191490A1 (en) Voice activity detection using active acoustic sensing
EP4434450B1 (en) Detecting heart rate variability using a hearable
JP6861303B2 (en) How to operate the hearing aid system and the hearing aid system
US20260030330A1 (en) Authentication Using Active Acoustic Sensing
US20250048020A1 (en) Own voice audio processing for hearing loss
WO2025193636A1 (en) Gating speech processing using active acoustic sensing
WO2025244651A1 (en) Howling prevention
US20260082162A1 (en) Method of operating a hearing device and hearing device

Legal Events

Date Code Title Description
STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: UNKNOWN

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE

PUAI Public reference made under article 153(3) epc to a published international application that has entered the european phase

Free format text: ORIGINAL CODE: 0009012

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE

17P Request for examination filed

Effective date: 20241121

AK Designated contracting states

Kind code of ref document: A1

Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR

GRAP Despatch of communication of intention to grant a patent

Free format text: ORIGINAL CODE: EPIDOSNIGR1

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: GRANT OF PATENT IS INTENDED

RAP3 Party data changed (applicant data changed or rights of an application transferred)

Owner name: GOOGLE LLC

DAV Request for validation of the european patent (deleted)
DAX Request for extension of the european patent (deleted)
INTG Intention to grant announced

Effective date: 20250923

RIN1 Information on inventor provided before grant (corrected)

Inventor name: FAN, XIAORAN

Inventor name: THOR-MUNDSSON, TRAUSTI

GRAS Grant fee paid

Free format text: ORIGINAL CODE: EPIDOSNIGR3

GRAA (expected) grant

Free format text: ORIGINAL CODE: 0009210

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: THE PATENT HAS BEEN GRANTED

P01 Opt-out of the competence of the unified patent court (upc) registered

Free format text: CASE NUMBER: UPC_APP_0018125_4599601/2025

Effective date: 20251218

AK Designated contracting states

Kind code of ref document: B1

Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR

REG Reference to a national code

Ref country code: CH

Ref legal event code: F10

Free format text: ST27 STATUS EVENT CODE: U-0-0-F10-F00 (AS PROVIDED BY THE NATIONAL OFFICE)

Effective date: 20260204

Ref country code: GB

Ref legal event code: FG4D

REG Reference to a national code

Ref country code: DE

Ref legal event code: R096

Ref document number: 602023011714

Country of ref document: DE

REG Reference to a national code

Ref country code: IE

Ref legal event code: FG4D

REG Reference to a national code

Ref country code: NL

Ref legal event code: FP