WO2020240169A1 - Detection of speech - Google Patents
Detection of speech Download PDFInfo
- Publication number
- WO2020240169A1 WO2020240169A1 PCT/GB2020/051270 GB2020051270W WO2020240169A1 WO 2020240169 A1 WO2020240169 A1 WO 2020240169A1 GB 2020051270 W GB2020051270 W GB 2020051270W WO 2020240169 A1 WO2020240169 A1 WO 2020240169A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- signal
- speech
- component
- articulation rate
- speech articulation
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L17/00—Speaker identification or verification techniques
- G10L17/06—Decision making techniques; Pattern matching strategies
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F1/00—Details not covered by groups G06F3/00 - G06F13/00 and G06F21/00
- G06F1/16—Constructional details or arrangements
- G06F1/1613—Constructional details or arrangements for portable computers
- G06F1/163—Wearable computers, e.g. on a belt
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/06—Creation of reference templates; Training of speech recognition systems, e.g. adaptation to the characteristics of the speaker's voice
- G10L15/065—Adaptation
- G10L15/07—Adaptation to the speaker
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/28—Constructional details of speech recognition systems
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L25/00—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
- G10L25/48—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use
- G10L25/51—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use for comparison or discrimination
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L25/00—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
- G10L25/78—Detection of presence or absence of voice signals
Definitions
- This relates to the detection of speech, and in particular to detecting when a speaker is a person who is using a device, for example wearing a wearable accessory such as an earphone.
- Wearable accessories such as earphones, smart glasses and smart watches, are common.
- a device such as a smartphone that has speech recognition functionality
- speech that was spoken by the person wearing the accessory can be supplied to the speech recognition functionality so that any spoken commands can be acted upon, while speech that was not spoken by the person wearing the accessory can be ignored in some circumstances.
- a method of own voice detection for a user of a device comprising:
- determining that the speech has not been generated by the user of the device if a difference between the component of the first signal at the speech articulation rate and the component of the second signal at the speech articulation rate exceeds a threshold value.
- a system for own voice detection for a user of a device, the system comprising:
- At least one filter for filtering the first signal to obtain a component of the first signal at a speech articulation rate and for filtering the second signal to obtain a component of the second signal at the speech articulation rate;
- a comparator for comparing the component of the first signal at the speech articulation rate and the component of the second signal at the speech articulation rate
- a processor for determining that the speech has not been generated by the user of the device, if a difference between the component of the first signal at the speech articulation rate and the component of the second signal at the speech articulation rate exceeds a threshold value.
- a method of detecting a spoof attack on a speaker recognition system comprising:
- a speaker recognition system comprising:
- At least one filter for filtering the first signal to obtain a component of the first signal at a speech articulation rate, and for filtering the second signal to obtain a component of the second signal at the speech articulation rate;
- a comparator for comparing the component of the first signal at the speech articulation rate and the component of the second signal at the speech articulation rate; a processor for determining that the speech has not been generated by a user of the device, if a difference between the component of the first signal at the speech articulation rate and the component of the second signal at the speech articulation rate exceeds a threshold value;
- a speaker recognition block for performing speaker recognition on the first signal representing speech, if the difference between the component of the first signal at the articulation rate and the component of the second signal at the articulation rate does not exceed the threshold value.
- Figure 1 is a schematic view of an electronic device and an associated accessory
- Figure 2 is a further schematic diagram of an electronic device and an accessory
- Figure 3 is a flow chart, illustrating a method
- Figure 4 illustrates a wearable device
- Figure 5 is a block diagram, illustrating a part of a system as described herein;
- Figure 6 is a block diagram, illustrating a part of the system of Figure 5;
- Figure 7 illustrates a stage in the method of Figure 3
- Figure 8 further illustrates a stage in the method of Figure 3;
- Figure 9 is a block diagram, illustrating a form of system as described herein;
- Figure 10 illustrates a stage in the use of the system of Figure 9
- Figure 11 illustrates a stage in the use of the system of Figure 9
- Figure 12 is a block diagram, illustrating an alternative form of system as described herein.
- Figure 13 is a block diagram, illustrating an alternative form of system as described herein.
- Speaker recognition refers to a technique that provides information about the identity of a person speaking. For example, speaker recognition may determine tine identity of a speaker, from amongst a group of previously registered individuals, or may provide information indicating whether a speaker is or is not a particular individual, for the purposes of identification or authentication. Speech recognition refers to a technique for determining the content and/or the meaning of what is spoken, rather than recognising the person speaking.
- FIG. 1 shows a device in accordance with one aspect of the invention.
- the device may be any suitable type of device, such as a mobile computing device, for example a laptop or tablet computer, a games console, a remote control device, a home automation controller or a domestic appliance including a domestic temperature or lighting control system, a toy, a machine such as a robot, an audio player, a video player, or the like, but in this illustrative example the device is a mobile telephone, and specifically a smartphone 10, having a microphone 12 for detecting sounds.
- the smartphone 10 may, by suitable software, be used as the control interface for controlling any other further device or system.
- Figure 1 also shows an accessory, which in this case is a wireless earphone 30, which in this example takes the form of an in-ear earphone or an earbud.
- the earphone 30 may be one of a pair of earphones or a part of a headset, or may be used on its own.
- a wireless earphone 30 is shown, but an earphone with a wired connection to the device may equally be used.
- the accessory may be any suitable device that can be worn by a person and used in conjunction with the device.
- the accessory may be a smart watch, or a pair of smart glasses.
- Figure 2 is a schematic diagram, illustrating the form of the smartphone 10 and the wireless earphone 30. Specifically, Figure 2 shows various interconnected components of the smartphone 10 and the wireless earphone 30. It will be appreciated that the smartphone 10 and the wireless earphone 30 will in practice contain many other components, but the following description is sufficient for an understanding of the present invention. In addition, it will be appreciated that similar components to those shown in Figure 2 may be included in any suitable device and any suitable wearable accessory.
- Figure 2 shows that the smartphone 10 includes the microphone 12 mentioned above.
- Figure 2 also shows a memory 14, which may in practice be provided as a single component or as multiple components.
- the memory 14 is provided for storing data and program instructions.
- FIG. 2 also shows a processor 16, which again may in practice be provided as a single component or as multiple components.
- a processor 16 may be an applications processor of the smartphone 10.
- the memory 14 may act as a tangible computer-readable medium, storing code, for causing the processor 16 to perform methods as described below.
- Figure 2 also shows a transceiver 18, which is provided for allowing the smartphone 10 to communicate with external networks.
- the transceiver 18 may include circuitry for establishing an internet connection either over a WiFi local area network or over a cellular network.
- the transceiver 18 allows the smartphone 10 to communicate with the wireless earphone 30, for example using Bluetooth or another short-range wireless communications protocol.
- FIG 2 also shows audio processing circuitry 20, for performing operations on the audio signals detected by the microphone 12 as required.
- the audio processing circuitry 20 may filter the audio signals or perform other signal processing operations.
- Figure 2 shows that the wireless earphone 30 includes a transceiver 32, which allows the wireless earphone 30 to communicate with the smartphone 10, for example using Bluetooth or another short-range wireless communications protocol.
- Figure 2 also shows that the wireless earphone 30 includes a first sensor 34 and a second sensor 36, which will be described in more detail below. Signals that are generated by the first sensor 34 and the second sensor 36, in response to external stimuli, are transmitted to the smartphone 10 by means of the transceiver 32.
- signals are generated by sensors on an accessory, and these signals are transmitted to a host device, where they are processed.
- the signals are processed on the accessory itself.
- the smartphone 10 is provided with speaker recognition functionality, and with control functionality.
- the smartphone 10 is able to perform various functions in response to spoken commands from an enrolled user.
- the speaker recognition functionality is able to distinguish between spoken commands from an enrolled user, and the same commands when spoken by a different person.
- certain embodiments of the invention relate to operation of a smartphone or another portable electronic device with some sort of voice operability, for example a tablet or laptop computer, a games console, a home control system, a home entertainment system, an in-vehicle entertainment system, a domestic appliance, or the like, in which the speaker recognition functionality is performed in the device that is intended to carry out the spoken command.
- Certain other embodiments relate to systems in which the speaker recognition functionality is performed on a smartphone or other device, which then transmits the commands to a separate device if the speaker recognition functionality is able to confirm that the speaker was an enrolled user.
- the spoken commands are transmitted using the transceiver 18 to a remote speech recognition system, which determines the meaning of the spoken commands.
- the speech recognition system may be located on one or more remote server in a cloud computing environment. Signals based on the meaning of the spoken commands are then returned to the smartphone 10 or other local device.
- a spoken command When a spoken command is received, it is often required to perform a process of speaker verification, in order to confirm that the speaker is an enrolled user of the system. It is known to perform speaker verification by first performing a process of enrolling a user to obtain a model of the user’s speech. Then, when it is desired to determine whether a particular test input is the speech of that user, a first score is obtained by comparing the test input against the model of the user’s speech. In addition, a process of score normalization may be performed. For example, the test input may also be compared against a plurality of models of speech obtained from a plurality of other speakers. These comparisons give a plurality of cohort scores, and statistics can be obtained describing the plurality of cohort scores. The first score can then be normalized using the statistics to obtain a normalized score, and the normalised score can be used for speaker verification.
- the speech signal may be sent for speaker recognition and/or speech recognition. If it is determined that the detected speech was not spoken by the person wearing the accessory 30, a decision may be taken that the speech signal should not be sent for speaker recognition and/or speech recognition.
- Figure 3 is a flow chart illustrating an example of a method in accordance with the present disclosure, and specifically a method for own voice detection for a user wearing a wearable device, that is, a method for detecting whether the person wearing the wearable device is the person speaking. Essentially the same method can be used for detecting whether a person holding a handheld device such as a mobile phone is the person speaking.
- Figure 4 illustrates in more detail the form of the wearable device, in one embodiment. Specifically, Figure 4 shows the earphone 30 being worn in the ear canal 70 of a wearer.
- the first sensor 34 shown in Figure 2 takes the form of an out-of-ear microphone 72, that is, a microphone that detects acoustic signals in the air
- the second sensor 36 shown in Figure 2 takes the form of a bone-conduction sensor 74.
- This may be an in-ear microphone that is able to detect acoustic signals in the wearer’s ear canal, that result from the wearer's speech and that are transmitted through the bones of the wearer’s head, and that may also be able to detect vibrations of the wearer's ear canal itself.
- the bone conduction sensor 74 may be an accelerometer, positioned such that it is in contact with the wearer's ear canal, and can detect contact vibrations that result from the wearer’s speech and that are transmitted through the bones and/or soft tissue of the wearer’s head.
- the first sensor may take the form of an externally directed microphone that picks up the wearer’s speech (and other sounds) by means of conduction through the air
- the second sensor may take the form of an accelerometer that is positioned to be in contact with the wearer’s head, and can detect contact vibrations that result from the wearer’s speech and that are transmitted through the bones and/or soft tissue of the wearer’s head.
- the first sensor may take the form of an externally directed microphone that picks up the wearer’s speech (and other sounds) by means of conduction through the air
- the second sensor may take the form of an accelerometer that is positioned to be in contact with the wearer’s wrist, and can detect contact vibrations that result from the wearer’s speech and that are transmitted through the wearer’s bones and/or soft tissue.
- the first sensor may be the microphone 12 that picks up the wearer’s speech (and other sounds) by means of conduction through the air
- the second sensor may take the form of an accelerometer that is positioned within the handheld device (and therefore not visible in Figure 1), that can detect contact vibrations that result from the user’s speech and that are transmitted through the wearer’s bones and/or soft tissue.
- the accelerometer can detect contact vibrations that result from the wearer's speech and that are transmitted through the bones and/or soft tissue of the user's head, while, when the handheld device is not being pressed against the user’s head, the accelerometer can detect contact vibrations that result from the wearer’s speech and that are transmitted through the bones and/or soft tissue of the user’s arm and hand.
- the method then comprises detecting a first signal representing air-conducted speech using the first microphone 72 of the wearable device 30. That is, when the wearer speaks, the sound leaves their mouth and travels through the air, and can be detected by the microphone 72.
- the method also comprises detecting a second signal representing bone-conducted speech using the bone-conduction sensor 74 of the wearable device 30. That is, when the wearer speaks, the vibrations are conducted through the bones of their head (and/or may be conducted to at least some extent through the surrounding soft tissue), and can be detected by the bone-conduction sensor 74.
- the process of own voice detection can be achieved by comparing the signals generated by the microphone 72 and the bone-conduction sensor 74.
- a typical bone-conduction sensor 74 is based on an accelerometer, and accelerometers typically run at a low sample rate, for example in the region of 100Hz - 1 kHz, the signal generated by the bone-conduction sensor 74 has a restricted frequency range.
- a typical bone-conduction sensor is prone to picking up contact noise (for example resulting from the wearer’s head turning, or from contact with other objects). Further, it has been found that in general voiced speech is transmitted more efficiently via bone conduction than unvoiced speech.
- Figure 5 therefore shows filter circuitry for filtering the signals generated by the microphone 72 and the bone-conduction sensor 74 to make them more useful for the purposes of own voice detection.
- Figure 5 shows the signal from the first sensor 72 (i.e. the first signal mentioned at step 50 of Figure 3) being received at a first input 90, and the signal from the second sensor 72 (i.e. the second signal mentioned at step 52 of Figure 3) being received at a second input 92.
- the second signal will be expected to contain significant signal content only during periods when the wearer’s speech contains voiced speech. Therefore, the first signal, received at the first input 90, is passed to a voiced speech detection block 94. This generates a flag when it is determined that the first signal represents voiced speech.
- Voiced speech can be identified, for example, by: using a deep neural network (DNN), trained against a golden reference, for example using Praat software; performing an autocorrelation with unit delay on the speech signal (because voiced speech has a higher autocorrelation for non-zero lags); performing a linear predictive coding (LPC) analysis (because the initial reflection coefficient is a good indicator of voiced speech); looking at the zero-crossing rate of the speech signal (because unvoiced speech has a higher zero-crossing rate); looking at the short term energy of the signal (which tends to be higher for voiced speech); tracking the first formant frequency F0 (because unvoiced speech does not contain the first format frequency); examining the error in a linear predictive coding (LPC) analysis (because the LPC prediction error is lower for voiced speech); using automatic speech recognition to identify the words being spoken and hence the division of the speech into voiced and unvoiced speech; or fusing any or all of the above.
- DNN deep neural network
- Praat software performing an autocorre
- step 56 of the method of Figure 3 the second signal is filtered to obtain a component of the second signal at a speech articulation rate. Therefore, the second signal, received at the second input 92, is passed to a second articulation rate filter 98.
- Figure 6 is a schematic diagram, illustrating in more detail the form of the first articulation rate filter 96 and the second articulation rate filter 98.
- the respective input signal is passed to a low pass filter 110, with a cutoff frequency that may for example be in the region of 1kHz.
- the low-pass filtered signal is passed to an envelope detector 112 for detecting the envelope of the filtered signal.
- the resulting envelope signal may optionally be passed to a decimator 114, and then to a band-pass filter 116, which is tuned to allow signals at typical articulation rates to pass.
- the band-pass filter 116 may have a pass band between 5-15Hz or between 5-1 OHz.
- the articulation rate filters 96, 98 detect the power modulating at frequencies that correspond to typical articulation rates for speech, where the articulation rate is the rate at which the speaker is speaking, which may for example be measured as the rate at which the speaker generates distinct speech sounds, or phones.
- step 58 of the method of Figure 3 the component of the first signal at the speech articulation rate and the component of the second signal at the speech articulation rate are compared.
- the outputs of the first articulation rate filter 96 and the second articulation rate filter 98 are passed to a comparison and decision block 100.
- step 60 of the method of Figure 3 it is then determined that the speech has not been generated by the user wearing the wearable device, if a difference between the component of the first signal at the speech articulation rate and the component of the second signal at the speech articulation rate exceeds a threshold value. Conversely, if the difference between the component of the first signal at the speech articulation rate and the component of the second signal at the speech articulation rate does not exceed the threshold value, it may be determined that the speech has been generated by the user wearing the wearable device.
- the comparison may be performed only when the flag is generated, indicating that the first signal represents voiced speech. Either the entire filtered first signal may be passed to the comparison and decision block 100, with the comparison being performed only when the flag is generated, indicating that the first signal represents voiced speech, or unvoiced speech may be rejected and only those segments of the filtered first signal representing voiced speech may be passed to the comparison and decision block 100.
- the processing of the signals detected by the first and second sensors is therefore such that, when the wearer of the wearable device is the person speaking, it would be expected that the processed version of the first signal would be similar to the processed version of the second signal.
- the first sensor would stili be able to detect the air- conducted speech, but the second sensor would not be able to detect any bone- conducted speech, and so it would be expected that the processed version of the first signal would be quite different from the processed version of the second signal.
- Figures 7 and 8 are examples of this, showing the magnitudes of the processed versions of the first and second signal in these two cases.
- the signals are divided into frames, having a duration of 20ms (for example), and the magnitude during each frame period is plotted over time.
- Figure 7 shows a situation in which the wearer of the wearable device is the person speaking, and so the processed version of the first signal 130 is similar to the processed version of the second signal 132.
- Figure 8 shows a situation in which the wearer of the wearable device is not the person speaking, and so, while the processed version of the first signal 140 contains components resulting from the air-conducted speech, this is quite different from the processed version of the second signal 142, because this does not contain any component resulting from bone-conducted speech.
- Figure 9 is a schematic illustration, showing a first form of the comparison and decision block 100 from Figure 5, in which the processed version of the first signal, i.e. a component of the first signal at the speech articulation rate, is generated by the first articulation rate filter 96 and passed to a first input 120 of the comparison and decision block 100.
- the processed version of the second signal i.e. a component of the second signal at the speech articulation rate, is generated by the second articulation rate filter 98 and passed to a second input 122 of the comparison and decision block 100.
- the signal at the first input 120 is then passed to a block 124, in which an empirical cumulative distribution function (ECDF) is formed.
- the signal at the second input 122 is then passed to a block 126, in which an empirical cumulative distribution function (ECDF) is formed.
- the ECDF is calculated on a frame-by-frame basis, using the magnitude of the signal during the frames.
- the ECDF then indicates, for each possible signal magnitude, the proportion of the frames in which the actual signal magnitude is below that level.
- Figure 10 shows the two ECDFs calculated in a situation similar to Figure 7, in which the two signals are generally similar. It can thus also be seen that the two ECDFs, namely the ECDF 142 formed from the first signal and the ECDF 144 formed from the second signal, are generally similar.
- One measure of similarity between two ECDFs 142, 144 is to measure the maximum vertical distance between them, in this case d1.
- several measures of the vertical distance between the two ECDFs can be made, and then summed or averaged, in order to arrive at a suitable measure.
- Figure 11 shows the two ECDFs calculated in a situation similar to Figure 8, in which the two signals are substantially different. It can thus also be seen that the two ECDFs, namely the ECDF 152 formed from the first signal and the ECDF 154 formed from the second signal, are substantially different. Specifically, because the second signal does not contain any component resulting from bone-conducted speech, it is generally at a much lower level than the first signal, and so the form of the ECDF shows that the magnitude of the second signal is generally low.
- the maximum vertical distance between the ECDFs 152, 154 is d2.
- the step of comparing the components of the first and second signals at the speech articulation rate may comprise forming respective first and second distribution functions from those components, and calculating a value of a statistical distance between the second distribution function and the first distribution function.
- the value of the statistical distance between the second distribution function and the first distribution function may be calculated as:
- F1 is the first distribution function
- the value of the statistical distance between the second distribution function and the first distribution function may be calculated as:
- F1 is the first distribution function
- the value of the statistical distance between the second distribution function and the first distribution function may be calculated as:
- F1 is the first distribution function
- the step of comparing the components may use a machine learning system that has been trained to distinguish between the components that are generated from a wearer’s own speech and the speech of a non-wearer.
- the distribution functions are cumulative distribution functions, other distribution functions, such as probability distribution functions may be used, with appropriate methods for comparing these functions.
- the methods for comparing may include using machine learning systems as described above.
- the ECDFs are passed to a block 128, which calculates the statistical distance d between the ECDFs, and passes this to a block 130, where this is compared with a threshold q.
- step 60 of the method of Figure 3 it is then determined that tiie speech has not been generated by the user wearing the wearable device, if the statistical distance between the ECDFs generated from the component of the first signal at the speech articulation rate and the component of the second signal at the speech articulation rate exceeds the threshold value q. Conversely, if the statistical distance between the ECDFs generated from the component of the first signal at the speech articulation rate and the component of the second signal at the speech articulation rate does not exceed the threshold value q, it may be determined that the speech has been generated by the user wearing the wearable device.
- Figure 12 is a schematic illustration, showing a second form of the comparison and decision block 100 from Figure 5, in which the processed version of the first signal, i.e. a component of the first signal at the speech articulation rate, is generated by the first articulation rate filter 96 and passed to a first input 160 of the comparison and decision block 100.
- the processed version of the second signal i.e. a component of the second signal at the speech articulation rate, is generated by the second articulation rate filter 98 and passed to a second input 162 of the comparison and decision block 100.
- the signal at the first input 160 is subtracted from the signal at the second input 162 in a subtracter 164, and the difference D is passed to a comparison block 166.
- the value of the difference D that is calculated in each frame, is compared with a threshold value. If the difference exceeds the threshold value in any frame, then it may be determined that the speech has not been generated by the user wearing the wearable device.
- the values of the difference D, calculated in multiple frames are analysed statistically. For example, the mean value of D, calculated over a block of, say, 20 frames, or the running mean, may be compared with the threshold value. As another example, the median value of D, calculated over a block of frames, may be compared with the threshold value. In any of these cases, if the parameter calculated from the separate difference values exceeds the threshold value, then it may be determined that the speech (or at least the relevant part of the speech from which the difference values were calculated) was not generated by the user wearing the wearable device.
- Figure 13 is a block schematic diagram of a system using the method of own speech detection described previously.
- Figure 13 shows a first sensor 200 and a second sensor 202, which may be provided on a wearable device.
- the first sensor 200 may take the form of a microphone that detects acoustic signals
- the second sensor 202 may take the form of an accelerometer for detecting signals transmitted by bone-conduction (including transmission through the wearer's soft tissues).
- the signals from the first sensor 200 and a second sensor 202 are passed to a wear detection block 204, which compares the signals generated by the sensors, and determines whether the wearable device is being worn at that time.
- the wear detection block 204 may take the form of an in-ear detection block.
- the signals detected by the sensor 200 and 202 are substantially the same if the earphone is out of the user’s ear, but are significantly different if the earphone is being worn. Therefore, comparing the signals allows a determination as to whether the wearable device is being worn.
- sensors 200, 202 Other systems are available for detecting whether a wearable device is being worn, and these may or may not use either of the sensors 200, 202.
- optical sensors, conductivity sensors, or proximity sensors may be used.
- the wear detection block 204 may be configured to perform“liveness detection”, that is, to determine whether the wearable device is being worn by a live person at that time.
- the signal generated by the second sensor 202 may be analysed to detect evidence of a wearer’s pulse, to confirm that a wearable device such as an earphone or watch is being worn by a person, rather than being placed in or on an inanimate object.
- the detection block 204 may be configured for determining whether the device is being held by the user (as opposed to, for example, being used while lying on a table or other surface).
- the detection block may receive signals from the sensor 202, or from one or more separate sensor, that can be used to determine whether the device is being held by the user and/or being pressed against the user's head.
- optical sensors, conductivity sensors, or proximity sensors may be used.
- a signal from the detection block 204 is used to close switches 206, 208, and the signals from the sensors 200, 202 are passed to the inputs 90, 92 of the circuitry shown in Figure 5.
- the comparison and decision block 100 generates an output signal indicating whether the detected speech was spoken by the person wearing the wearable device.
- the signal from the first sensor 200 is also passed to a voice keyword detection block 220.
- the voice keyword detection block 220 detects a particular“wake phrase” that is used by a user of the device to wake it from a low power standby mode, and put it into a mode in which speech recognition is possible.
- the signal from the first sensor 200 is sent to a speaker recognition block 222. If the output signal from the comparison and decision block 100 indicates that the detected speech was not spoken by the person wearing the wearable device, this is considered to be a spoof, and so it is not desirable to perform speech recognition on the detected speech. However, if the output signal from the comparison and decision block 100 indicates that the detected speech was spoken by the person wearing the wearable device, a speaker recognition process may be performed by the speaker recognition block 222 on the signal from the first sensor 200. In general terms, the process of speaker recognition extracts features from the speech signal, and compares them with features of a model generated by enrolling a known user into the speaker recognition system. If the comparison finds that the features of the speech are sufficiently similar to the model, then it is determined that the speaker is the enrolled user, to a sufficiently high degree of probability.
- the process of speaker recognition may be omitted and replaced by an ear biometric process, in which the wearer of the earphone is identified. For example, this may be done by examining signals generated by an in- ear microphone provided on the earphone (which microphone may also act as the second sensor 202), and comparing features of the signals with an acoustic model of an enrolled user's ear. If it can be determined with sufficient confidence for the relevant application both that the earphone is being worn by the enrolled user, and also that the speech is the speech of the person wearing the earphone, then this acts as a form of speaker recognition.
- a signal is sent to a speech recognition block 224, which also receives the signal from the first sensor 200.
- the speech recognition block 224 then performs a process of speech recognition on the received signal. If it detects that the speech contains a command, for example, it may send an output signal to a further application to act on that command.
- the availability of a bone-conducted signal can be used for the purposes of own speech detection. Further, the result of the own speech detection can be used to improve the reliability of speaker recognition and speech recognition systems.
- processor control code for example on a nonvolatile carrier medium such as a disk, CD- or DVD-ROM, programmed memory such as read only memory (Firmware), or on a data carrier such as an optical or electrical signal carrier.
- a nonvolatile carrier medium such as a disk, CD- or DVD-ROM
- programmed memory such as read only memory (Firmware)
- a data carrier such as an optical or electrical signal carrier.
- DSP Digital Signal Processor
- ASIC Application Specific
- the code may comprise conventional program code or microcode or, for example code for setting up or controlling an ASIC or FPGA.
- the code may also comprise code for dynamically configuring re-configurable apparatus such as re-programmable logic gate arrays.
- the code may comprise code for a hardware description language such as Verilog TM or VHDL (Very high speed integrated circuit Hardware Description
- the code may be distributed between a plurality of coupled components in communication with one another.
- the embodiments may also be implemented using code running on a field- (re)programmable analogue array or similar device in order to configure analogue hardware.
- module shall be used to refer to a functional unit or block which may be implemented at least partly by dedicated hardware components such as custom defined circuitry and/or at least partly be implemented by one or more software processors or appropriate code running on a suitable general purpose processor or the like.
- a module may itself comprise other modules or functional units.
- a module may be provided by multiple components or sub-modules which need not be co-located and could be provided on different integrated circuits and/or running on different processors.
- Embodiments may be implemented in a host device, especially a portable and/or battery powered host device such as a mobile computing device for example a laptop or tablet computer, a games console, a remote control device, a home automation controller or a domestic appliance including a domestic temperature or lighting control system, a toy, a machine such as a robot, an audio player, a video player, or a mobile telephone for example a smartphone.
- a host device especially a portable and/or battery powered host device such as a mobile computing device for example a laptop or tablet computer, a games console, a remote control device, a home automation controller or a domestic appliance including a domestic temperature or lighting control system, a toy, a machine such as a robot, an audio player, a video player, or a mobile telephone for example a smartphone.
- a portable and/or battery powered host device such as a mobile computing device for example a laptop or tablet computer, a games console, a remote control device, a home automation controller or a domestic appliance including
Landscapes
- Engineering & Computer Science (AREA)
- Human Computer Interaction (AREA)
- Physics & Mathematics (AREA)
- Multimedia (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Acoustics & Sound (AREA)
- Computational Linguistics (AREA)
- Computer Hardware Design (AREA)
- Theoretical Computer Science (AREA)
- Signal Processing (AREA)
- Business, Economics & Management (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Game Theory and Decision Science (AREA)
- Artificial Intelligence (AREA)
- General Engineering & Computer Science (AREA)
- General Physics & Mathematics (AREA)
- Telephone Function (AREA)
Abstract
Description
Claims
Priority Applications (3)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202080031842.7A CN113767431B (en) | 2019-05-30 | 2020-05-26 | Method and system for speech detection |
| GB2114745.9A GB2596752B (en) | 2019-05-30 | 2020-05-26 | Detection of speech |
| KR1020217042234A KR20220015427A (en) | 2019-05-30 | 2020-05-26 | detection of voice |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US201962854388P | 2019-05-30 | 2019-05-30 | |
| US62/854,388 | 2019-05-30 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2020240169A1 true WO2020240169A1 (en) | 2020-12-03 |
Family
ID=70978283
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/GB2020/051270 Ceased WO2020240169A1 (en) | 2019-05-30 | 2020-05-26 | Detection of speech |
Country Status (5)
| Country | Link |
|---|---|
| US (2) | US11488583B2 (en) |
| KR (1) | KR20220015427A (en) |
| CN (1) | CN113767431B (en) |
| GB (1) | GB2596752B (en) |
| WO (1) | WO2020240169A1 (en) |
Families Citing this family (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN113113050A (en) * | 2021-05-10 | 2021-07-13 | 紫光展锐(重庆)科技有限公司 | Voice activity detection method, electronic equipment and device |
| EP4131256A1 (en) * | 2021-08-06 | 2023-02-08 | STMicroelectronics S.r.l. | Voice recognition system and method using accelerometers for sensing bone conduction |
| US12100420B2 (en) * | 2022-02-15 | 2024-09-24 | Google Llc | Speech detection using multiple acoustic sensors |
| CN117476042A (en) | 2022-07-29 | 2024-01-30 | 北京三星通信技术研究有限公司 | Method executed by electronic device, electronic device and storage medium |
Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20190005964A1 (en) * | 2017-06-28 | 2019-01-03 | Cirrus Logic International Semiconductor Ltd. | Detection of replay attack |
| US20190043518A1 (en) * | 2016-02-25 | 2019-02-07 | Dolby Laboratories Licensing Corporation | Capture and extraction of own voice signal |
| US20190075406A1 (en) * | 2016-11-24 | 2019-03-07 | Oticon A/S | Hearing device comprising an own voice detector |
Family Cites Families (22)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| DE69527731T2 (en) * | 1994-05-18 | 2003-04-03 | Nippon Telegraph & Telephone Co., Tokio/Tokyo | Transceiver with an acoustic transducer of the earpiece type |
| AU2571900A (en) * | 1999-02-16 | 2000-09-04 | Yugen Kaisha Gm&M | Speech converting device and method |
| US7695441B2 (en) * | 2002-05-23 | 2010-04-13 | Tympany, Llc | Automated diagnostic hearing test |
| JP3963850B2 (en) * | 2003-03-11 | 2007-08-22 | 富士通株式会社 | Voice segment detection device |
| CA2473195C (en) * | 2003-07-29 | 2014-02-04 | Microsoft Corporation | Head mounted multi-sensory audio input system |
| US7564980B2 (en) * | 2005-04-21 | 2009-07-21 | Sensimetrics Corporation | System and method for immersive simulation of hearing loss and auditory prostheses |
| WO2008086085A2 (en) * | 2007-01-03 | 2008-07-17 | Biosecurity Technologies, Inc. | Ultrasonic and multimodality assisted hearing |
| US20090005964A1 (en) * | 2007-06-28 | 2009-01-01 | Apple Inc. | Intelligent Route Guidance |
| US10112029B2 (en) * | 2009-06-19 | 2018-10-30 | Integrated Listening Systems, LLC | Bone conduction apparatus and multi-sensory brain integration method |
| EP2458586A1 (en) * | 2010-11-24 | 2012-05-30 | Koninklijke Philips Electronics N.V. | System and method for producing an audio signal |
| US9336780B2 (en) * | 2011-06-20 | 2016-05-10 | Agnitio, S.L. | Identification of a local speaker |
| US9837078B2 (en) * | 2012-11-09 | 2017-12-05 | Mattersight Corporation | Methods and apparatus for identifying fraudulent callers |
| US9264824B2 (en) * | 2013-07-31 | 2016-02-16 | Starkey Laboratories, Inc. | Integration of hearing aids with smart glasses to improve intelligibility in noise |
| US9271064B2 (en) * | 2013-11-13 | 2016-02-23 | Personics Holdings, Llc | Method and system for contact sensing using coherence analysis |
| EP3274993B1 (en) * | 2015-04-23 | 2019-06-12 | Huawei Technologies Co. Ltd. | An audio signal processing apparatus for processing an input earpiece audio signal upon the basis of a microphone audio signal |
| US10362415B2 (en) * | 2016-04-29 | 2019-07-23 | Regents Of The University Of Minnesota | Ultrasonic hearing system and related methods |
| GB2552723A (en) * | 2016-08-03 | 2018-02-07 | Cirrus Logic Int Semiconductor Ltd | Speaker recognition |
| US10535364B1 (en) * | 2016-09-08 | 2020-01-14 | Amazon Technologies, Inc. | Voice activity detection using air conduction and bone conduction microphones |
| US10878818B2 (en) * | 2017-09-05 | 2020-12-29 | Massachusetts Institute Of Technology | Methods and apparatus for silent speech interface |
| GB201719734D0 (en) * | 2017-10-30 | 2018-01-10 | Cirrus Logic Int Semiconductor Ltd | Speaker identification |
| US10455324B2 (en) * | 2018-01-12 | 2019-10-22 | Intel Corporation | Apparatus and methods for bone conduction context detection |
| GB2576960B (en) * | 2018-09-07 | 2021-09-22 | Cirrus Logic Int Semiconductor Ltd | Speaker recognition |
-
2020
- 2020-05-18 US US16/876,373 patent/US11488583B2/en active Active
- 2020-05-26 KR KR1020217042234A patent/KR20220015427A/en not_active Withdrawn
- 2020-05-26 WO PCT/GB2020/051270 patent/WO2020240169A1/en not_active Ceased
- 2020-05-26 CN CN202080031842.7A patent/CN113767431B/en active Active
- 2020-05-26 GB GB2114745.9A patent/GB2596752B/en active Active
-
2022
- 2022-09-13 US US17/943,745 patent/US11842725B2/en active Active
Patent Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20190043518A1 (en) * | 2016-02-25 | 2019-02-07 | Dolby Laboratories Licensing Corporation | Capture and extraction of own voice signal |
| US20190075406A1 (en) * | 2016-11-24 | 2019-03-07 | Oticon A/S | Hearing device comprising an own voice detector |
| US20190005964A1 (en) * | 2017-06-28 | 2019-01-03 | Cirrus Logic International Semiconductor Ltd. | Detection of replay attack |
Also Published As
| Publication number | Publication date |
|---|---|
| CN113767431B (en) | 2025-06-13 |
| KR20220015427A (en) | 2022-02-08 |
| CN113767431A (en) | 2021-12-07 |
| GB2596752B (en) | 2023-12-06 |
| US11488583B2 (en) | 2022-11-01 |
| US20200380955A1 (en) | 2020-12-03 |
| GB2596752A (en) | 2022-01-05 |
| US20230005470A1 (en) | 2023-01-05 |
| US11842725B2 (en) | 2023-12-12 |
| GB202114745D0 (en) | 2021-12-01 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US11842725B2 (en) | Detection of speech | |
| US10433075B2 (en) | Low latency audio enhancement | |
| US12417771B2 (en) | Hearing device or system comprising a user identification unit | |
| US11475899B2 (en) | Speaker identification | |
| US9560456B2 (en) | Hearing aid and method of detecting vibration | |
| US11900730B2 (en) | Biometric identification | |
| CN110268470A (en) | Audio device filter modification | |
| US11437021B2 (en) | Processing audio signals | |
| GB2608710A (en) | Speaker identification | |
| CN109346075A (en) | Method and system for recognizing user's voice through human body vibration to control electronic equipment | |
| WO2019008362A1 (en) | Blocked microphone detection | |
| CN115132212A (en) | Voice control method and device | |
| US11894000B2 (en) | Authenticating received speech | |
| CN113132885B (en) | Method for judging wearing state of earphone based on energy difference of double microphones | |
| CN113039601B (en) | A voice control method, device, chip, earphone and system | |
| CN115996349A (en) | Hearing devices including feedback control systems | |
| CN118942491B (en) | Data processing method, electronic device, storage medium, and computer program product | |
| CN111201570A (en) | Analyze speech signals | |
| US20220310057A1 (en) | Methods and apparatus for obtaining biometric data |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 20730689 Country of ref document: EP Kind code of ref document: A1 |
|
| ENP | Entry into the national phase |
Ref document number: 202114745 Country of ref document: GB Kind code of ref document: A Free format text: PCT FILING DATE = 20200526 |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 2114745.9 Country of ref document: GB |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| ENP | Entry into the national phase |
Ref document number: 20217042234 Country of ref document: KR Kind code of ref document: A |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 20730689 Country of ref document: EP Kind code of ref document: A1 |
|
| WWG | Wipo information: grant in national office |
Ref document number: 2114745.9 Country of ref document: GB |
|
| WWG | Wipo information: grant in national office |
Ref document number: 202080031842.7 Country of ref document: CN |


