WO2021234342A1 - Authenticating received speech - Google Patents

Authenticating received speech Download PDF

Info

Publication number
WO2021234342A1
WO2021234342A1 PCT/GB2021/050908 GB2021050908W WO2021234342A1 WO 2021234342 A1 WO2021234342 A1 WO 2021234342A1 GB 2021050908 W GB2021050908 W GB 2021050908W WO 2021234342 A1 WO2021234342 A1 WO 2021234342A1
Authority
WO
WIPO (PCT)
Prior art keywords
correlation
speech
signal
microphone
transducer
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/GB2021/050908
Other languages
French (fr)
Inventor
John Paul Lesso
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Cirrus Logic International Semiconductor Ltd
Original Assignee
Cirrus Logic International Semiconductor Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Cirrus Logic International Semiconductor Ltd filed Critical Cirrus Logic International Semiconductor Ltd
Priority to GB2215008.0A priority Critical patent/GB2608568B/en
Priority to CN202180030771.3A priority patent/CN115461812A/en
Publication of WO2021234342A1 publication Critical patent/WO2021234342A1/en
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L17/00Speaker identification or verification techniques
    • G10L17/22Interactive procedures; Man-machine interfaces
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L17/00Speaker identification or verification techniques
    • G10L17/06Decision making techniques; Pattern matching strategies
    • G10L17/10Multimodal systems, i.e. based on the integration of multiple recognition engines or fusion of expert systems
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L17/00Speaker identification or verification techniques
    • G10L17/06Decision making techniques; Pattern matching strategies
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L17/00Speaker identification or verification techniques
    • G10L17/26Recognition of special voice characteristics, e.g. for use in lie detectors; Recognition of animal voices
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L25/00Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
    • G10L25/03Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the type of extracted parameters
    • G10L25/06Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the type of extracted parameters the extracted parameters being correlation coefficients

Definitions

  • Background Speech recognition systems are known, allowing a user to control a device or system using spoken commands. It is common to use speaker recognition systems in conjunction with speech recognition systems.
  • a speaker recognition system can be used to verify the identity of a person who is speaking, and this can be used to control the operation of the speech recognition system.
  • a spoken command may relate to the personal tastes of the speaker.
  • the spoken command may be “Play my favourite music”, in which case it is necessary to know the identity of the speaker before it is possible to determine which music should be played.
  • a spoken command may relate to a financial transaction.
  • the spoken command may be an instruction that involves transferring money to a specific recipient. In that case, before acting on the spoken command, it is necessary to have a high degree of confidence that the command was spoken by the presumed speaker.
  • Speaker recognition systems often use a voice biometric, where the received speech is compared with a model generated when a person enrols with the system.
  • Many voice-activated devices require the user to speak a predetermined trigger phrase, in order to wake the system from a low-power mode, in order that speech recognition can be performed. It is possible to perform speaker recognition on the speech signal that corresponds to the predetermined trigger phrase, and this speaker recognition process is relatively reliable, because the user will typically have been required to utter the predetermined trigger phrase during the enrolment process. This is referred to as a text-dependent speaker recognition process. Therefore, when the user speaks the predetermined trigger phrase to wake the system, it can be expected that the speech will closely resemble the speech during the enrolment.
  • a method of authenticating a speech signal received by a device comprising first and second transducers, wherein the first transducer comprises a microphone.
  • the method comprises: performing a first voice biometric process on speech contained in a first part of a signal received by the microphone, in order to determine whether the speech is the speech of an enrolled user; determining a first correlation between said first part of the signal received by the microphone and a corresponding part of the signal received by the second transducer; determining a second correlation between said second part of the signal received by the microphone and the corresponding part of the signal received by the second transducer; and determining whether the first correlation and the second correlation satisfy a predetermined condition; and, if it is determined that the speech contained in the first part of the received signal is the speech of an enrolled user and that the first correlation and the second correlation satisfy the predetermined condition, authenticating the received speech signal.
  • a system for authenticating a speech signal received by a device comprising first and second transducers, wherein the first transducer comprises a microphone.
  • the system comprises: at least one input for receiving signals generated by the microphone and by the second transducer; and a processor configured for performing a method comprising: performing a first voice biometric process on speech contained in a first part of the signal generated by the microphone, in order to determine whether the speech is the speech of an enrolled user; determining a first correlation between said first part of the signal generated by the microphone and a corresponding part of the signal generated by the second transducer; determining a second correlation between said second part of the signal generated by the microphone and the corresponding part of the signal generated by the second transducer; and determining whether the first correlation and the second correlation satisfy a predetermined condition; and, if it is determined that the speech contained in the first part of the received signal is the speech of an enrolled user and that the first correlation and the second correlation satisfy the predetermined condition, authenticating the
  • a computer program product comprising machine readable code containing instructions for causing an audio processing circuit to perform a method according to the first aspect.
  • Figure 1 illustrates an example of a device being worn by a user
  • Figure 2 is a schematic diagram, illustrating the form of a host device
  • Figure 3 illustrates in more detail a part of the device of Figure 1;
  • Figure 4 illustrates an example of an attack on a voice-activated device
  • Figure 5 illustrates another example of an attack on a voice-activated device
  • Figure 6 is a flow chart illustrating a method in accordance with the present disclosure
  • Figure 7 is a block diagram illustrating a system for performing the method of Figure 6;
  • Figure 8 is a timing diagram illustrating operation of the system of Figure 7;
  • Figure 9 illustrates operation of a part of the system of Figure 7;
  • Figure 10 illustrates operation of a part of the system of Figure 7
  • Figure 11 illustrates operation of a part of the system of Figure 7;
  • Figure 12 illustrates signals generated in the system of Figure 7, in one example
  • Figure 13 illustrates operation of a part of the system of Figure 7;
  • Figure 14 illustrates operation of a part of the system of Figure 7.
  • Figure 15 illustrates operation of a part of the system of Figure 7.
  • the methods described herein may be implemented in a wide range of devices and systems. However, for ease of explanation of one embodiment, an illustrative example will be described, in which the implementation occurs in a host device, which is used with a wearable accessory. However, in other embodiments, the implementation may occur in a wearable device such as a headset.
  • Figure 1 illustrates an example of a device being worn by a user.
  • Figure 1 illustrates a person wearing an earphone. More specifically, Figure 1 shows a person 10, wearing one wireless earbud 12, 14 in each ear 16, 18. Although this shows a person wearing two earbuds, the method is applicable when only one earbud is being worn.
  • Figure 1 shows a person wearing wireless earbuds
  • the method is applicable to any wired or wireless earbuds or earphones, for example in-ear earphones, supra-aural earphones, or supra-concha earphones.
  • the method is applicable to any wearable device, such as smart glasses.
  • Figure 2 is a schematic diagram, illustrating the form of a host device 20.
  • the host device 20 may for example take the form of a smartphone, a laptop or tablet computer, a smart speaker, a games console, a home control system, a home entertainment system, an in-vehicle entertainment system, a domestic appliance, or any other suitable device.
  • Figure 2 shows various interconnected components of the host device 20.
  • the host device 20 will in practice contain many other components, but the following description is sufficient for an understanding of embodiments of the present disclosure.
  • Figure 2 shows a transceiver 22, which is provided for allowing the host device 20 to communicate with other devices.
  • the transceiver 22 may include circuitry for communicating over a short-range wireless link with an accessory, such as the accessory 10 shown in Figure 1.
  • the transceiver 22 may include circuitry for establishing an internet connection either over a WiFi local area network or over a cellular network.
  • Figure 2 also shows a memory 24, which may in practice be provided as a single component or as multiple components.
  • the memory 24 is provided for storing data and program instructions.
  • FIG. 2 also shows a processor 26, which again may in practice be provided as a single component or as multiple components.
  • a processor 26 may be an applications processor when the host device 20 is a smartphone.
  • Figure 2 also shows audio processing circuitry 28, for performing operations on received audio signals as required.
  • the audio processing circuitry 28 may filter the audio signals or perform other signal processing operations.
  • the host device 20 is provided with voice biometric functionality, and with control functionality.
  • the device 20 is able to perform various functions in response to spoken commands from an enrolled user.
  • the biometric functionality is able to distinguish between spoken commands from the enrolled user, and the same commands when spoken by a different person.
  • certain embodiments of the present disclosure relate to operation of a smartphone or another portable electronic host device with some sort of voice operability, in which the voice biometric functionality is performed in the host device that is intended to carry out the spoken command.
  • Certain other embodiments relate to systems in which the voice biometric functionality is performed on a smartphone or other host device, which then transmits the commands to a separate device if the voice biometric functionality is able to confirm that the speaker was the enrolled user.
  • the spoken commands are transmitted using the transceiver 22 to a remote speech recognition system, which determines the meaning of the spoken commands.
  • the speech recognition system may be located on one or more remote server in a cloud computing environment. Signals based on the meaning of the spoken commands are then returned to the host device 20 or other local device.
  • a first part of the voice biometric functionality is performed on the host device 20 or other device that is located close to the user. Then, as described in more detail below, a signal may be transmitted using the transceiver 22 to a remote system, which performs a second part of the voice biometric functionality.
  • Figure 3 illustrates in more detail a part of the device of Figure 1.
  • Figure 3 illustrates an example where the accessory device is an earphone, which is being worn. More specifically, Figure 3 shows an earbud 30 at the entrance to a wearer’s ear canal 32.
  • the earphone comprises a first transducer and a second transducer. While a person is wearing the earphone, a first transducer is located on an outward facing part of the earphone and a second transducer is located on a part of the earphone facing into the person’s ear canal.
  • the first transducer comprises a microphone 34, located such that it can detect ambient sound in the vicinity of the earbud 30.
  • the earbud 30 also comprises a second microphone 36, located such that it can detect sound in the wearer’s ear canal 32.
  • the earbud 30 also comprises an accelerometer 38, located on the earbud 30 such that it can detect vibrations in the surface of the wearer’s ear canal 32 resulting from the transmission of sound through the wearer’s head.
  • the second transducer mentioned above, can be the second microphone 36, or can be the accelerometer 38.
  • the accessory device may be any suitable wearable device, for example smart glasses, which are provided with a microphone for detecting sound that has travelled through the air, and are also provided with a second transducer such as an accelerometer that is mounted in a position that is in contact with the wearer’s head when the glasses are being worn, such that the accelerometer can detect vibrations in resulting from the transmission of sound through the wearer’s head.
  • a wearable device for example smart glasses, which are provided with a microphone for detecting sound that has travelled through the air, and are also provided with a second transducer such as an accelerometer that is mounted in a position that is in contact with the wearer’s head when the glasses are being worn, such that the accelerometer can detect vibrations in resulting from the transmission of sound through the wearer’s head.
  • embodiments described herein obtain information about the sound conduction path, through the wearer’s head, by comparing the signals detected by the first transducer and the second transducer. More specifically, embodiments described herein obtain information about the sound conduction path, through the wearer’s head, by comparing the signals detected by the first transducer and the second transducer at times when the wearer is speaking.
  • the processing of the signals generated by the external microphone 34, and by the one or more internal transducer 36, 38, may be performed in circuitry provided within the earbud 30 itself. However, in embodiments described herein, the signals generated by the external microphone 34 and by the one or more internal transducer 36, 38 may be transmitted by a suitable wired or wireless connection to the host device 20, where the processing of the signals, as described in more detail below, takes place.
  • Figure 4 illustrates an example of an attack on a voice-activated device.
  • Figure 4 illustrates the operation of a voice-activated device, which, in order to reduce its power consumption, is normally in a low-power sleep mode, and requires an enrolled user to speak a predetermined trigger phrase, in order to wake the system from the low-power mode, in order that speech recognition can be performed.
  • the user speaks the predetermined trigger phrase, which in this case is “Hi phone”, that is used to activate the speech recognition functionality.
  • Figure 5 illustrates another example of an attack on a voice-activated device.
  • Figure 5 again illustrates the operation of a voice-activated device, which, requires the enrolled user to speak the predetermined trigger phrase, in order to wake the system from the low-power mode.
  • the user speaks the predetermined trigger phrase, which in this case is “Hi phone”, that is used to activate the speech recognition functionality.
  • the user speaks a command, namely, in this illustrative example “order me a pizza”. If the speaker recognition system is able to recognise that the person speaking the words “order me a pizza” is the enrolled user, then the system will act on that command, and fulfil the wishes of the enrolled user.
  • the speaker recognition system is able to recognise when a first part of a received signal has been spoken by an enrolled user, but a subsequent part of the received signal has been spoken by a different person.
  • the method disclosed herein proceeds from the recognition that, when the speaker is wearing a wearable accessory, there exists a mechanism for determining whether the speech following the trigger phrase was spoken by the same person as the trigger phrase. If the voice biometric process can be used to determine that the trigger phrase was spoken by the enrolled user, then this additional information can be used for confirming whether the speech following the trigger phrase was spoken by the enrolled user.
  • Figure 6 is a flow chart illustrating a method in accordance with the present disclosure.
  • Figure 6 shows a method of authenticating a speech signal received by a device comprising first and second transducers, where the first transducer comprises a microphone.
  • a first voice biometric process is performed on speech contained in a first part of a signal received by the microphone, in order to determine whether the speech is the speech of an enrolled user.
  • the first voice biometric process may be a text-dependent voice biometric process.
  • step 63 a first correlation is determined between the first part of the signal received by the microphone and a corresponding part of the signal received by the second transducer.
  • step 64 a second correlation is determined between a second part of the signal received by the microphone and the corresponding part of the signal received by the second transducer.
  • step 65 it is determined whether the first correlation and the second correlation satisfy a predetermined condition.
  • step 66 if it is determined that the speech contained in the first part of the received signal is the speech of an enrolled user and that the first correlation and the second correlation satisfy the predetermined condition, the received speech signal is authenticated.
  • the first voice biometric process may be a text- dependent voice biometric process.
  • the first biometric process may be a text-independent voice biometric process.
  • a good text-independent biometric process typically has a high power consumption, so one possibility enabled by this method is to run the first biometric process for a first part of the received signal that lasts a relatively short period of time (for example of the order of 1 second), and then disable the first biometric process, relying on the correlations described above to confirm that the same person was speaking and to authenticate the entire received speech signal.
  • Figure 7 is a block diagram illustrating a system for performing the method of Figure 6.
  • a first input signal S AC is received from a first transducer, in the form of a microphone 70, which is located such that it can detect ambient sound in the vicinity of the wearable accessory device.
  • a second input signal SB C is received from a second transducer 72, which is located such that it can detect vibrations caused by the transmission of sound through the wearer’s body.
  • the second transducer may take the form of a microphone located such that it can detect sound in the wearer’s ear canal 32, or may take the form of an accelerometer, located in the user’s ear canal or elsewhere such that it can detect vibrations in the surface of the wearer’s ear canal resulting from the transmission of sound through the wearer’s head.
  • the second transducer may take the form of an accelerometer, held in position against the user’s head such that it can detect vibrations resulting from the transmission of sound through the wearer’s head.
  • the signal S AC received from the first transducer 70, and the signal SB C received from the second transducer 72 are passed to a buffer 74, where they can be stored for a short period of time for further processing as required.
  • the bones and soft tissue of a person’s head are able to transmit the sounds of voiced speech to a reasonable extent, but are not able to transmit the sounds of unvoiced speech to any significant extent.
  • the signal S AC received from the first transducer 70 and the signal SB C received from the second transducer 72 during periods of voiced speech, but not during periods of unvoiced speech.
  • the received signals S AC and SB C are therefore passed to an acoustic class detection block 76, which detects the acoustic class of received speech, and in particular distinguishes between voiced and unvoiced speech.
  • the presence of voiced speech may be detected by examining a pitch period (F0), for example by consideration of the cepstrum or the Harmonic Product Spectrum (HPS).
  • F0 pitch period
  • HPS Harmonic Product Spectrum
  • the acoustic class of the received speech can be determined satisfactorily from the signal S AC that is received from the first transducer 70, and therefore it is not essential that the signal SB C received from the second transducer 72 should be passed to the acoustic class detection block 76.
  • the acoustic class detection block 76 particularly when the signal-to-noise ratio is low, it is useful for the acoustic class detection block 76 to be able to use the signal SB C received from the second transducer 72 in addition to, or as an alternative to, the signal S AC that is received from the first transducer 70.
  • the signal S AC received from the microphone 70 is also passed to a voice trigger detection block 78, which identifies when the speech in the received signal represents the predetermined trigger phrase.
  • the part of the signal S AC received from the microphone 70 that represents the predetermined trigger phrase is also passed to a voice biometric block 80, which performs a biometric process on the received signal.
  • the voice biometric block 80 may extract features from the received signal, and compare the extracted features with a model of the speech of the enrolled use that was generated during an enrolment process and stored in a database 82.
  • the voice biometric block 80 is intended to perform a speaker recognition process on the predetermined trigger phrase, it can take the form of a text-dependent speaker recognition process.
  • the voice biometric block 80 may perform a text-independent speaker recognition process on the predetermined trigger phrase, or on any part of the received signal, as desired.
  • a control signal is sent to the buffer 74, and the stored signals S AC and SB C are sent to a correlation block 84.
  • signals starting from a point in time before the voice trigger detection block 78 determines that the predetermined trigger phrase has been spoken are sent to the correlation block 84.
  • the point in time is selected to be early enough that the signals that are sent to the correlation block 84 include the signals that correspond to the predetermined trigger phrase itself.
  • the signals that are sent to the correlation block 84 may start from a point in time that is 1 second before the point at which the voice trigger detection block 78 determines that the predetermined trigger phrase has been spoken.
  • the output of the voice biometric block 80 and the output of the correlation block 84 are then combined to provide the final authentication output.
  • the voice biometric block 80 provides an output that indicates whether the predetermined trigger phrase was spoken by the enrolled user.
  • the correlation block 84 provides an output that indicates whether the person speaking during a first part of the received signal is the same person that continues speaking for the entire duration of the received signal.
  • the correlation block can provide a suitable output.
  • Figure 7 shows the output of the voice biometric block 80 being provided to the correlation block 84, and the correlation block 84 providing a combined output.
  • the voice biometric block 80 and the correlation block 84 may provide separate outputs, which may then be combined.
  • the signals SAC and SBC that are provided to the correlation block 84 are expected to be relatively well correlated, provided that it is the person wearing the accessory that is speaking, and provided that the speech is voiced speech.
  • the output of the acoustic class detection block 76 is therefore used as a control input to the correlation block 84.
  • the correlation block 84 examines the correlation between the signals SAC and SBC during periods when they represent voiced speech.
  • the correlation block 84 examines the correlation between the signals SAC and SBC by forming a prediction or estimate of SBC, referred to here as SBC*, from SAC, and then determining whether the actual signal SBC matches the estimate SBC*.
  • SBC* a prediction or estimate of SBC
  • the basis for forming this estimate is shown in Figure 3, where it can be seen that the signal SAC results from the application of the first transfer function TAIR to the originally generated sound S, while the signal SBC results from the application of the second transfer function TBONE to the sound S.
  • Figure 8 is a block diagram, illustrating the process for forming the estimate S BC * from SAC-
  • the received signal SAC is passed to a first block 90, which determines whether, at that specific time, the signal SAC represents voiced speech.
  • This determination is made on the basis of the control signal C received from the acoustic class detection block 76.
  • the signal SAC represents voiced speech
  • it is passed to a filter 92, which multiplies the signal SAC by an estimate of the transfer function TBONE (described with reference to Figure 3), in order to arrive at the estimate SBC*.
  • the filter 92 may be a fixed filter, or may be an adaptive filter.
  • TBONE may be acceptably approximated by a fixed low-order lowpass filter with an 800Hz cut-off frequency. However, in that case, there will be an unknown gain shift between the signals SAC and SBC.
  • An adaptive filter can be used to determine the gain G that is needed to compensate for this.
  • FIG. 9 illustrates the operation of this part of the system.
  • the signal SBC is applied to an adaptive gain block 100, which multiplies the signal by a gain value G.
  • the multiplied signal is applied to one input of a subtractor 102.
  • the estimate SBC* of SBC is applied to a second input of the subtractor 102.
  • the output of the subtractor 102 is an error signal e, which is used to control the gain value G applied by the adaptive gain block 100, in such a way that the value of e is minimised.
  • the resulting final gain value G can be applied to the signal SBC at any convenient point in the system.
  • the filter 92 shown in Figure 8 may alternatively be an adaptive filter.
  • Figure 10 illustrates the mechanism for determining the required form of the filter 92 in this case.
  • the signal SAC is applied to an adaptive filter 110, which multiplies the signal by a filter function TBONE.
  • the multiplied signal is applied to one input of a subtractor 112.
  • the second signal SBC is applied to a second input of the subtractor 112.
  • the output of the subtractor 112 is an error signal e, which is used to control the filter function TBONE applied by the adaptive filter 110, in such a way that the value of e is minimised.
  • the system therefore performs a Least Mean Squares (LMS) method of adaptation.
  • LMS Least Mean Squares
  • the adaptation of the filter function should take place slowly enough that the effect of noise on the signal SAC is averaged out, and hence the filter function of the block 110 becomes equal to the transfer function that needs to be applied to the signal SAC, to make it equal to the signal SBC, i.e. the transfer function TBONE in the equation above.
  • the resulting filter function TBONE when the system has settled is the form of the filter that can be used as the adaptive filter 92 in Figure 8.
  • Figure 11 illustrates signals generated in the system of Figure 7, in one example.
  • Figure 11 shows the signal SAC, the signal SBC* that is derived from SAC as an estimate of SBC, and the signal SBC, in one example, as functions of time.
  • the next step in determining the correlation between the signals is to extract the energy of the signals.
  • Figure 12 illustrates this next step. Specifically, Figure 12 shows the signal SAC being applied to a block 120, which determines an estimate SBC* of the signal SBC, as described with reference to Figure 8.
  • the estimate SBC* is then applied to a first energy calculation block 122, which calculates the energy EBc*of the estimate SBC*.
  • the signal SBC is applied to a second energy calculation block 124, which calculates the energy EBC of the signal SBC.
  • the energy calculation blocks 122, 124 can for example operate by squaring the signals and then low-pass filtering them, or by applying a Teager Kaiser Operator, but other possibilities exist.
  • the outputs of the two energy calculation blocks 122, 124 are passed to a comparison block 126, which compares them, and determines if they are sufficiently similar to meet a similarity threshold.
  • the comparison block 126 may determine the Pearson Correlation Coefficient, the Cosine similarity, the Euclidian distance, or any other statistical distance metric, e.g. Bhattacharya or Mahalanobis, or any other similarity metric, between the outputs of the two energy calculation blocks 122, 124.
  • Figure 13 illustrates the operation of this part of the system of Figure 7.
  • Figure 13(a) shows with line 130 the output EBC* of the first energy calculation block 122, and shows with line 132 the output EBC of the second energy calculation block 124, in the case where the person speaking is the person who is wearing the wearable accessory, and where a fixed filter is used as the filter 92 in Figure 8.
  • Figure 13(a) also shows the amplitude 134 of the difference between the outputs 130 and 132.
  • Figure 13(b) shows with line 140 the output EBC* of the first energy calculation block 122, and shows with line 140 the output E B c of the second energy calculation block 124, in the case where the person speaking is the person who is wearing the wearable accessory, and where an adaptive filter is used as the filter 92 in Figure 8.
  • Figure 13(b) also shows the amplitude 144 of the difference between the outputs 140 and 142.
  • the signal SAC can be used to form a good estimate SBC* of the signal SBC, and hence can be used as a reliable indicator that the person speaking is the person who is wearing the wearable accessory.
  • Figure 13(c) shows with line 150 the output E B c*of the first energy calculation block 122, and shows with line 152 the output EBC of the second energy calculation block 124, in the case where the person speaking is not the person who is wearing the wearable accessory.
  • Figure 13(c) also shows the amplitude 154 of the difference between the outputs 150 and 152.
  • the energy EBC of the signal detected by the second transducer i.e. the transducer located in the wearer’s ear canal, is very small, which is itself a good indication that the person wearing the wearable accessory is not speaking.
  • the correlation process described above is used for determining a first correlation between a first part of the signal SAC received by the microphone and a corresponding part of the signal SBC received by the second transducer, where the first part of the signal may correspond to the predetermined trigger phrase.
  • the degree of correlation between them may be represented by a first correlation value.
  • the correlation process is also used for determining a second correlation between a second part of the signal SAC received by the microphone and the corresponding part of the signal SBC received by the second transducer, where the second part of the signal may correspond to the period following the trigger phrase.
  • the degree of correlation between them may be represented by a second correlation value.
  • the predetermined condition may for example relate to a specific relationship between the first correlation value and the second correlation value.
  • the predetermined condition may be that the first and second correlation values are sufficiently similar that it can be assumed that the person speaking was the same during the first and second parts of the signal.
  • the predetermined condition may be that the first and second correlation values are above a respective threshold.
  • the two thresholds used in this determination may be the same or may be different.
  • the degree of correlation may be high, because it may be possible to set a useful threshold value with a high degree of confidence, whereas it is more difficult to set the threshold value for the second part of the signal, because the speech content is unknown and also of unknown length, and the degree of correlation may be lower.
  • the thresholds may be calculated using Decision Cost Function (DCF) methodology or Neyman-Pearson methodology. As mentioned above, the thresholds may be different, but as an example both may be set such that a correlation factor should exceed 0.8.
  • DCF Decision Cost Function
  • the first and second correlation values are obtained by examining the energies of the respective signals, during the relevant time periods.
  • the first and second correlation values are obtained by calculating the Pearson correlation coefficient between the relevant part of the signal SAC and the corresponding part of the signal SBC.
  • Figure 14 illustrates the operation of the system of Figure 7. Specifically, Figure 14 is a timing diagram illustrating the operation of the system of Figure 7.
  • the top line 160 of Figure 14 shows the words being spoken and being detected by the microphone. Specifically, as shown at 162, Figure 14 shows the words “Hi phone, order me a pizza” being spoken by a first person, who is the wearer of the wearable accessory. In addition, at 164, another person then speaks, and says “with extra anchovies”.
  • the line 166 illustrates the output of the voice trigger detection block 78 in Figure 7.
  • the voice trigger detection block 78 generates an output at time t1, shortly after the predetermined trigger phrase has been spoken.
  • the subsequent processing of the received signals only begins when it has been determined that the predetermined trigger phrase has been spoken, and the stored signals are retrieved from the buffer 74.
  • the following steps will be described as if they are performed directly on the received signals, rather than after a short time delay.
  • the line 168 illustrates the output of the voice biometric block 80 in Figure 7.
  • the voice biometric block 80 also generates an output shortly after the predetermined trigger phrase has been spoken. In this illustrated example, it is assumed that this is a positive output, indicating that the speaker was the enrolled user.
  • the correlation block 84 when the predetermined trigger phrase has been spoken, and the stored signals are retrieved from the buffer, the correlation block 84 is activated.
  • the line 170 illustrates the output of the correlation block 84.
  • the correlation block 84 produces an output that is the result of comparing two correlation values, one obtained from the predetermined trigger phrase, and one obtained from the subsequent speech.
  • the correlation block 84 is able to produce an output. In this case, it produces a positive output, indicating that the speaker was the same person as spoke the predetermined trigger phrase.
  • the correlation block 84 recognises that the person speaking now is not the same person that was speaking before.
  • Figure 15 illustrates the operation of the correlation block 84 during this process.
  • Figure 15 shows with line 180 the output EB C * of the first energy calculation block 122, and shows with line 182 the output EB C of the second energy calculation block 124.
  • Figure 15 also shows the amplitude 184 of the difference between the outputs 180 and 182.
  • the output EB C * of the first energy calculation block 122 is a good estimate of the output EB C of the second energy calculation block 124, and the error signal 184 has a very small amplitude.
  • the output EB C of the second energy calculation block 124 itself has a very small amplitude, and hence the output EB C * of the first energy calculation block 122 is not a good estimate of the output EB C of the second energy calculation block 124, and the error signal 184 has a large amplitude.
  • the correlation block 84 is able to produce an output that confirms that the initial speaker was the enrolled user, but that the words “with extra anchovies” were not spoken by the enrolled user.
  • processor control code for example on a non volatile carrier medium such as a disk, CD- or DVD-ROM, programmed memory such as read only memory (Firmware), or on a data carrier such as an optical or electrical signal carrier.
  • a non volatile carrier medium such as a disk, CD- or DVD-ROM
  • programmed memory such as read only memory (Firmware)
  • a data carrier such as an optical or electrical signal carrier.
  • DSP Digital Signal Processor
  • ASIC Application Specific Integrated Circuit
  • FPGA Field Programmable Gate Array
  • the code may comprise conventional program code or microcode or, for example code for setting up or controlling an ASIC or FPGA.
  • the code may also comprise code for dynamically configuring re-configurable apparatus such as re-programmable logic gate arrays.
  • the code may comprise code for a hardware description language such as Verilog TM or VHDL (Very high speed integrated circuit Hardware Description Language).
  • Verilog TM or VHDL Very high speed integrated circuit Hardware Description Language
  • the code may be distributed between a plurality of coupled components in communication with one another.
  • the embodiments may also be implemented using code running on a field- (re)programmable analogue array or similar device in order to configure analogue hardware.
  • module shall be used to refer to a functional unit or block which may be implemented at least partly by dedicated hardware components such as custom defined circuitry and/or at least partly be implemented by one or more software processors or appropriate code running on a suitable general purpose processor or the like.
  • a module may itself comprise other modules or functional units.
  • a module may be provided by multiple components or sub-modules which need not be co-located and could be provided on different integrated circuits and/or running on different processors.
  • Embodiments may be implemented in a host device, especially a portable and/or battery powered host device such as a mobile computing device for example a laptop or tablet computer, a games console, a remote control device, a home automation controller or a domestic appliance including a domestic temperature or lighting control system, a toy, a machine such as a robot, an audio player, a video player, or a mobile telephone for example a smartphone.
  • a host device especially a portable and/or battery powered host device such as a mobile computing device for example a laptop or tablet computer, a games console, a remote control device, a home automation controller or a domestic appliance including a domestic temperature or lighting control system, a toy, a machine such as a robot, an audio player, a video player, or a mobile telephone for example a smartphone.

Landscapes

  • Engineering & Computer Science (AREA)
  • Multimedia (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Human Computer Interaction (AREA)
  • Physics & Mathematics (AREA)
  • Acoustics & Sound (AREA)
  • Health & Medical Sciences (AREA)
  • Business, Economics & Management (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Game Theory and Decision Science (AREA)
  • Measurement Of The Respiration, Hearing Ability, Form, And Blood Characteristics Of Living Organisms (AREA)
  • User Interface Of Digital Computer (AREA)
  • Circuit For Audible Band Transducer (AREA)

Abstract

A speech signal is received by a device comprising first and second transducers, and the first transducer comprises a microphone. A method comprises performing a first voice biometric process on speech contained in a first part of a signal received by the microphone, in order to determine whether the speech is the speech of an enrolled user. A first correlation is determined, between said first part of the signal received by the microphone and a corresponding part of the signal received by the second transducer. A second correlation is determined, between said second part of the signal received by the microphone and the corresponding part of the signal received by the second transducer. It is then determined whether the first correlation and the second correlation satisfy a predetermined condition. If it is determined that the speech contained in the first part of the received signal is the speech of an enrolled user and that the first correlation and the second correlation satisfy the predetermined condition, the received speech signal is authenticated.

Description

AUTHENTICATING RECEIVED SPEECH
Technical Field Embodiments described herein relate to methods and devices for authenticating a received speech signal.
Background Speech recognition systems are known, allowing a user to control a device or system using spoken commands. It is common to use speaker recognition systems in conjunction with speech recognition systems. A speaker recognition system can be used to verify the identity of a person who is speaking, and this can be used to control the operation of the speech recognition system.
As an illustration of this, a spoken command may relate to the personal tastes of the speaker. For example, the spoken command may be “Play my favourite music”, in which case it is necessary to know the identity of the speaker before it is possible to determine which music should be played.
As another illustration, a spoken command may relate to a financial transaction. For example, the spoken command may be an instruction that involves transferring money to a specific recipient. In that case, before acting on the spoken command, it is necessary to have a high degree of confidence that the command was spoken by the presumed speaker.
Speaker recognition systems often use a voice biometric, where the received speech is compared with a model generated when a person enrols with the system. Many voice-activated devices require the user to speak a predetermined trigger phrase, in order to wake the system from a low-power mode, in order that speech recognition can be performed. It is possible to perform speaker recognition on the speech signal that corresponds to the predetermined trigger phrase, and this speaker recognition process is relatively reliable, because the user will typically have been required to utter the predetermined trigger phrase during the enrolment process. This is referred to as a text-dependent speaker recognition process. Therefore, when the user speaks the predetermined trigger phrase to wake the system, it can be expected that the speech will closely resemble the speech during the enrolment.
However, it is more difficult to perform speaker recognition on the command that follows the predetermined trigger phrase, because the user will typically be speaking a phrase that will not have been uttered during the enrolment process. This is referred to as a text-independent speaker recognition process.
Summary
According to a first aspect of the invention, there is provided a method of authenticating a speech signal received by a device comprising first and second transducers, wherein the first transducer comprises a microphone. The method comprises: performing a first voice biometric process on speech contained in a first part of a signal received by the microphone, in order to determine whether the speech is the speech of an enrolled user; determining a first correlation between said first part of the signal received by the microphone and a corresponding part of the signal received by the second transducer; determining a second correlation between said second part of the signal received by the microphone and the corresponding part of the signal received by the second transducer; and determining whether the first correlation and the second correlation satisfy a predetermined condition; and, if it is determined that the speech contained in the first part of the received signal is the speech of an enrolled user and that the first correlation and the second correlation satisfy the predetermined condition, authenticating the received speech signal.
According to a second aspect of the invention, there is provided a system for authenticating a speech signal received by a device comprising first and second transducers, wherein the first transducer comprises a microphone. The system comprises: at least one input for receiving signals generated by the microphone and by the second transducer; and a processor configured for performing a method comprising: performing a first voice biometric process on speech contained in a first part of the signal generated by the microphone, in order to determine whether the speech is the speech of an enrolled user; determining a first correlation between said first part of the signal generated by the microphone and a corresponding part of the signal generated by the second transducer; determining a second correlation between said second part of the signal generated by the microphone and the corresponding part of the signal generated by the second transducer; and determining whether the first correlation and the second correlation satisfy a predetermined condition; and, if it is determined that the speech contained in the first part of the received signal is the speech of an enrolled user and that the first correlation and the second correlation satisfy the predetermined condition, authenticating the received speech signal.
According to a third aspect of the invention, there is provided a computer program product, comprising machine readable code containing instructions for causing an audio processing circuit to perform a method according to the first aspect.
Brief Description of Drawings
For a better understanding of the present invention, and to show how it may be put into effect, reference will now be made to the accompanying drawings, in which:-
Figure 1 illustrates an example of a device being worn by a user;
Figure 2 is a schematic diagram, illustrating the form of a host device;
Figure 3 illustrates in more detail a part of the device of Figure 1;
Figure 4 illustrates an example of an attack on a voice-activated device;
Figure 5 illustrates another example of an attack on a voice-activated device;
Figure 6 is a flow chart illustrating a method in accordance with the present disclosure;
Figure 7 is a block diagram illustrating a system for performing the method of Figure 6;
Figure 8 is a timing diagram illustrating operation of the system of Figure 7;
Figure 9 illustrates operation of a part of the system of Figure 7;
Figure 10 illustrates operation of a part of the system of Figure 7; Figure 11 illustrates operation of a part of the system of Figure 7;
Figure 12 illustrates signals generated in the system of Figure 7, in one example;
Figure 13 illustrates operation of a part of the system of Figure 7;
Figure 14 illustrates operation of a part of the system of Figure 7; and
Figure 15 illustrates operation of a part of the system of Figure 7.
Detailed Description of Embodiments
The description below sets forth example embodiments according to this disclosure. Further example embodiments and implementations will be apparent to those having ordinary skill in the art. Further, those having ordinary skill in the art will recognize that various equivalent techniques may be applied in lieu of, or in conjunction with, the embodiments discussed below, and all such equivalents should be deemed as being encompassed by the present disclosure.
The methods described herein may be implemented in a wide range of devices and systems. However, for ease of explanation of one embodiment, an illustrative example will be described, in which the implementation occurs in a host device, which is used with a wearable accessory. However, in other embodiments, the implementation may occur in a wearable device such as a headset.
Figure 1 illustrates an example of a device being worn by a user.
Specifically, Figure 1 illustrates a person wearing an earphone. More specifically, Figure 1 shows a person 10, wearing one wireless earbud 12, 14 in each ear 16, 18. Although this shows a person wearing two earbuds, the method is applicable when only one earbud is being worn.
In addition, although Figure 1 shows a person wearing wireless earbuds, the method is applicable to any wired or wireless earbuds or earphones, for example in-ear earphones, supra-aural earphones, or supra-concha earphones. Moreover, the method is applicable to any wearable device, such as smart glasses.
Figure 2 is a schematic diagram, illustrating the form of a host device 20.
The host device 20 may for example take the form of a smartphone, a laptop or tablet computer, a smart speaker, a games console, a home control system, a home entertainment system, an in-vehicle entertainment system, a domestic appliance, or any other suitable device.
Specifically, Figure 2 shows various interconnected components of the host device 20.
It will be appreciated that the host device 20 will in practice contain many other components, but the following description is sufficient for an understanding of embodiments of the present disclosure.
Thus, Figure 2 shows a transceiver 22, which is provided for allowing the host device 20 to communicate with other devices. Specifically, the transceiver 22 may include circuitry for communicating over a short-range wireless link with an accessory, such as the accessory 10 shown in Figure 1. In addition, the transceiver 22 may include circuitry for establishing an internet connection either over a WiFi local area network or over a cellular network.
Figure 2 also shows a memory 24, which may in practice be provided as a single component or as multiple components. The memory 24 is provided for storing data and program instructions.
Figure 2 also shows a processor 26, which again may in practice be provided as a single component or as multiple components. For example, one component of the processor 26 may be an applications processor when the host device 20 is a smartphone.
Figure 2 also shows audio processing circuitry 28, for performing operations on received audio signals as required. For example, the audio processing circuitry 28 may filter the audio signals or perform other signal processing operations. In this embodiment, the host device 20 is provided with voice biometric functionality, and with control functionality. Thus, the device 20 is able to perform various functions in response to spoken commands from an enrolled user. The biometric functionality is able to distinguish between spoken commands from the enrolled user, and the same commands when spoken by a different person. Thus, certain embodiments of the present disclosure relate to operation of a smartphone or another portable electronic host device with some sort of voice operability, in which the voice biometric functionality is performed in the host device that is intended to carry out the spoken command. Certain other embodiments relate to systems in which the voice biometric functionality is performed on a smartphone or other host device, which then transmits the commands to a separate device if the voice biometric functionality is able to confirm that the speaker was the enrolled user.
In some embodiments, while voice biometric functionality is performed on the host device 20 or other device that is located close to the user, the spoken commands are transmitted using the transceiver 22 to a remote speech recognition system, which determines the meaning of the spoken commands. For example, the speech recognition system may be located on one or more remote server in a cloud computing environment. Signals based on the meaning of the spoken commands are then returned to the host device 20 or other local device.
In other embodiments, a first part of the voice biometric functionality is performed on the host device 20 or other device that is located close to the user. Then, as described in more detail below, a signal may be transmitted using the transceiver 22 to a remote system, which performs a second part of the voice biometric functionality.
Figure 3 illustrates in more detail a part of the device of Figure 1.
Specifically, Figure 3 illustrates an example where the accessory device is an earphone, which is being worn. More specifically, Figure 3 shows an earbud 30 at the entrance to a wearer’s ear canal 32.
In general terms, the earphone comprises a first transducer and a second transducer. While a person is wearing the earphone, a first transducer is located on an outward facing part of the earphone and a second transducer is located on a part of the earphone facing into the person’s ear canal.
In the embodiment shown in Figure 3, the first transducer comprises a microphone 34, located such that it can detect ambient sound in the vicinity of the earbud 30.
In the embodiment shown in Figure 3, the earbud 30 also comprises a second microphone 36, located such that it can detect sound in the wearer’s ear canal 32. The earbud 30 also comprises an accelerometer 38, located on the earbud 30 such that it can detect vibrations in the surface of the wearer’s ear canal 32 resulting from the transmission of sound through the wearer’s head. The second transducer, mentioned above, can be the second microphone 36, or can be the accelerometer 38.
As mentioned above, the accessory device may be any suitable wearable device, for example smart glasses, which are provided with a microphone for detecting sound that has travelled through the air, and are also provided with a second transducer such as an accelerometer that is mounted in a position that is in contact with the wearer’s head when the glasses are being worn, such that the accelerometer can detect vibrations in resulting from the transmission of sound through the wearer’s head.
In particular, embodiments described herein obtain information about the sound conduction path, through the wearer’s head, by comparing the signals detected by the first transducer and the second transducer. More specifically, embodiments described herein obtain information about the sound conduction path, through the wearer’s head, by comparing the signals detected by the first transducer and the second transducer at times when the wearer is speaking.
Thus, as shown in Figure 3, when the wearer is speaking and generating a sound S, this is modified by a first transfer function TAIR through the air before it is detected by the external microphone 34, and it is modified by a second transfer function TBONE through the bone and soft tissue of the wearer’s head before it is detected by the internal transducer 36 or 38.
The processing of the signals generated by the external microphone 34, and by the one or more internal transducer 36, 38, may be performed in circuitry provided within the earbud 30 itself. However, in embodiments described herein, the signals generated by the external microphone 34 and by the one or more internal transducer 36, 38 may be transmitted by a suitable wired or wireless connection to the host device 20, where the processing of the signals, as described in more detail below, takes place.
Figure 4 illustrates an example of an attack on a voice-activated device.
Figure 4 illustrates the operation of a voice-activated device, which, in order to reduce its power consumption, is normally in a low-power sleep mode, and requires an enrolled user to speak a predetermined trigger phrase, in order to wake the system from the low-power mode, in order that speech recognition can be performed.
Thus, as shown at 42, the user speaks the predetermined trigger phrase, which in this case is “Hi phone”, that is used to activate the speech recognition functionality.
However, before the user can speak a command, another person speaks, as shown at 44, and says something that might be interpreted as a command, namely, in this illustrative example “Order me a pizza”. If the speaker recognition system is unable to recognise that the person speaking the words “Order me a pizza” is not the enrolled user, then the system might act on that command, which might be against the wishes of the enrolled user.
This is referred to as a competitive command.
Figure 5 illustrates another example of an attack on a voice-activated device.
Figure 5 again illustrates the operation of a voice-activated device, which, requires the enrolled user to speak the predetermined trigger phrase, in order to wake the system from the low-power mode.
Thus, as shown at 52, the user speaks the predetermined trigger phrase, which in this case is “Hi phone”, that is used to activate the speech recognition functionality. In addition, the user speaks a command, namely, in this illustrative example “order me a pizza”. If the speaker recognition system is able to recognise that the person speaking the words “order me a pizza” is the enrolled user, then the system will act on that command, and fulfil the wishes of the enrolled user.
However, as shown at 54, another person then speaks, and says something that might be interpreted as a part of the same command, namely, in this illustrative example “with extra anchovies”.
If the speaker recognition system is unable to recognise that the person speaking the words “with extra anchovies” is not the enrolled user, i.e. is not the person who spoke the words “Order me a pizza”, then the system might act on the entire command,
“Order me a pizza with extra anchovies”, which might be against the wishes of the enrolled user.
This is referred to as a tailgating attack on the system.
As illustrated with reference to Figures 4 and 5, therefore, it is advantageous if the speaker recognition system is able to recognise when a first part of a received signal has been spoken by an enrolled user, but a subsequent part of the received signal has been spoken by a different person.
It is possible to perform speaker recognition on the speech signal that corresponds to the predetermined trigger phrase in a relatively reliable way, because the user will typically have been required to utter the predetermined trigger phrase during the enrolment process. This is referred to as a text-dependent speaker recognition process. Therefore, when the user speaks the predetermined trigger phrase to wake the system, it can be expected that the speech will closely resemble the speech during the enrolment.
However, the text-independent speaker recognition process that is required to confirm that a command has been spoken by the enrolled user is more difficult, at least in the sense of being more computationally intensive, and typically is less reliable.
Thus, the method disclosed herein proceeds from the recognition that, when the speaker is wearing a wearable accessory, there exists a mechanism for determining whether the speech following the trigger phrase was spoken by the same person as the trigger phrase. If the voice biometric process can be used to determine that the trigger phrase was spoken by the enrolled user, then this additional information can be used for confirming whether the speech following the trigger phrase was spoken by the enrolled user.
Figure 6 is a flow chart illustrating a method in accordance with the present disclosure.
Specifically, Figure 6 shows a method of authenticating a speech signal received by a device comprising first and second transducers, where the first transducer comprises a microphone.
In step 62, a first voice biometric process is performed on speech contained in a first part of a signal received by the microphone, in order to determine whether the speech is the speech of an enrolled user. For example, the first voice biometric process may be a text-dependent voice biometric process.
In step 63, a first correlation is determined between the first part of the signal received by the microphone and a corresponding part of the signal received by the second transducer.
In step 64, a second correlation is determined between a second part of the signal received by the microphone and the corresponding part of the signal received by the second transducer.
In step 65, it is determined whether the first correlation and the second correlation satisfy a predetermined condition.
In step 66, if it is determined that the speech contained in the first part of the received signal is the speech of an enrolled user and that the first correlation and the second correlation satisfy the predetermined condition, the received speech signal is authenticated.
As mentioned above, in one example, the first voice biometric process may be a text- dependent voice biometric process. However, in another example, the first biometric process may be a text-independent voice biometric process. A good text-independent biometric process typically has a high power consumption, so one possibility enabled by this method is to run the first biometric process for a first part of the received signal that lasts a relatively short period of time (for example of the order of 1 second), and then disable the first biometric process, relying on the correlations described above to confirm that the same person was speaking and to authenticate the entire received speech signal.
Figure 7 is a block diagram illustrating a system for performing the method of Figure 6.
As described with reference to Figure 3, a first input signal SAC is received from a first transducer, in the form of a microphone 70, which is located such that it can detect ambient sound in the vicinity of the wearable accessory device.
A second input signal SBC is received from a second transducer 72, which is located such that it can detect vibrations caused by the transmission of sound through the wearer’s body. As previously described, when the wearable accessory is a headphone, the second transducer may take the form of a microphone located such that it can detect sound in the wearer’s ear canal 32, or may take the form of an accelerometer, located in the user’s ear canal or elsewhere such that it can detect vibrations in the surface of the wearer’s ear canal resulting from the transmission of sound through the wearer’s head. When the wearable accessory comprises smart glasses, the second transducer may take the form of an accelerometer, held in position against the user’s head such that it can detect vibrations resulting from the transmission of sound through the wearer’s head.
The signal SAC received from the first transducer 70, and the signal SBC received from the second transducer 72 are passed to a buffer 74, where they can be stored for a short period of time for further processing as required.
In general terms, the bones and soft tissue of a person’s head are able to transmit the sounds of voiced speech to a reasonable extent, but are not able to transmit the sounds of unvoiced speech to any significant extent. Thus, when a person wearing a wearable accessory as described herein is speaking, there is usually a good correlation between the signal SAC received from the first transducer 70, and the signal SBC received from the second transducer 72 during periods of voiced speech, but not during periods of unvoiced speech. The received signals SAC and SBC are therefore passed to an acoustic class detection block 76, which detects the acoustic class of received speech, and in particular distinguishes between voiced and unvoiced speech. As is known, the presence of voiced speech may be detected by examining a pitch period (F0), for example by consideration of the cepstrum or the Harmonic Product Spectrum (HPS). In most situations, the acoustic class of the received speech can be determined satisfactorily from the signal SAC that is received from the first transducer 70, and therefore it is not essential that the signal SBC received from the second transducer 72 should be passed to the acoustic class detection block 76. However, particularly when the signal-to-noise ratio is low, it is useful for the acoustic class detection block 76 to be able to use the signal SBC received from the second transducer 72 in addition to, or as an alternative to, the signal SAC that is received from the first transducer 70.
The signal SAC received from the microphone 70 is also passed to a voice trigger detection block 78, which identifies when the speech in the received signal represents the predetermined trigger phrase.
The part of the signal SAC received from the microphone 70 that represents the predetermined trigger phrase is also passed to a voice biometric block 80, which performs a biometric process on the received signal. For example, the voice biometric block 80 may extract features from the received signal, and compare the extracted features with a model of the speech of the enrolled use that was generated during an enrolment process and stored in a database 82. For example, because the voice biometric block 80 is intended to perform a speaker recognition process on the predetermined trigger phrase, it can take the form of a text-dependent speaker recognition process. However, in other embodiments the voice biometric block 80 may perform a text-independent speaker recognition process on the predetermined trigger phrase, or on any part of the received signal, as desired.
When the voice trigger detection block 78 determines that the speech in the received signal represents the predetermined trigger phrase, a control signal is sent to the buffer 74, and the stored signals SAC and SBC are sent to a correlation block 84. Specifically, signals starting from a point in time before the voice trigger detection block 78 determines that the predetermined trigger phrase has been spoken are sent to the correlation block 84. The point in time is selected to be early enough that the signals that are sent to the correlation block 84 include the signals that correspond to the predetermined trigger phrase itself. For example, the signals that are sent to the correlation block 84 may start from a point in time that is 1 second before the point at which the voice trigger detection block 78 determines that the predetermined trigger phrase has been spoken.
The operation of the correlation block 84 will be described in more detail below.
The output of the voice biometric block 80 and the output of the correlation block 84 are then combined to provide the final authentication output. The voice biometric block 80 provides an output that indicates whether the predetermined trigger phrase was spoken by the enrolled user. The correlation block 84 provides an output that indicates whether the person speaking during a first part of the received signal is the same person that continues speaking for the entire duration of the received signal.
If both of these conditions are met, it can be assumed that the enrolled user was speaking for the entire duration of the received signal, and the correlation block can provide a suitable output.
Figure 7 shows the output of the voice biometric block 80 being provided to the correlation block 84, and the correlation block 84 providing a combined output. In other embodiments, the voice biometric block 80 and the correlation block 84 may provide separate outputs, which may then be combined.
As discussed above, the signals SAC and SBC that are provided to the correlation block 84 are expected to be relatively well correlated, provided that it is the person wearing the accessory that is speaking, and provided that the speech is voiced speech.
The output of the acoustic class detection block 76 is therefore used as a control input to the correlation block 84. The correlation block 84 examines the correlation between the signals SAC and SBC during periods when they represent voiced speech.
In this example embodiment, the correlation block 84 examines the correlation between the signals SAC and SBC by forming a prediction or estimate of SBC, referred to here as SBC*, from SAC, and then determining whether the actual signal SBC matches the estimate SBC*. The basis for forming this estimate is shown in Figure 3, where it can be seen that the signal SAC results from the application of the first transfer function TAIR to the originally generated sound S, while the signal SBC results from the application of the second transfer function TBONE to the sound S.
Thus, SAC = S. TAJR and SBC = S. TB0NE, and so: c_ ¾C/ _ ¾C/
' TAIR ' TBONE
Therefore:
SBC = T. SAC , where:
Figure imgf000016_0001
Since TAIR can effectively be ignored, it is reasonable to assume that:
SBC = TBONE- SAC-
Figure 8 is a block diagram, illustrating the process for forming the estimate SBC* from SAC-
As shown in Figure 8, the received signal SAC is passed to a first block 90, which determines whether, at that specific time, the signal SAC represents voiced speech.
This determination is made on the basis of the control signal C received from the acoustic class detection block 76.
During periods when the signal SAC represents voiced speech, it is passed to a filter 92, which multiplies the signal SAC by an estimate of the transfer function TBONE (described with reference to Figure 3), in order to arrive at the estimate SBC*.
In Figure 8, the filter 92 may be a fixed filter, or may be an adaptive filter. For example, TBONE may be acceptably approximated by a fixed low-order lowpass filter with an 800Hz cut-off frequency. However, in that case, there will be an unknown gain shift between the signals SAC and SBC. An adaptive filter can be used to determine the gain G that is needed to compensate for this.
Figure 9 illustrates the operation of this part of the system.
The signal SBC is applied to an adaptive gain block 100, which multiplies the signal by a gain value G. The multiplied signal is applied to one input of a subtractor 102.
The estimate SBC* of SBC is applied to a second input of the subtractor 102.
Thus, the output of the subtractor 102 is an error signal e, which is used to control the gain value G applied by the adaptive gain block 100, in such a way that the value of e is minimised.
The resulting final gain value G can be applied to the signal SBC at any convenient point in the system.
As mentioned above, the filter 92 shown in Figure 8 may alternatively be an adaptive filter.
Figure 10 illustrates the mechanism for determining the required form of the filter 92 in this case.
The signal SAC is applied to an adaptive filter 110, which multiplies the signal by a filter function TBONE. The multiplied signal is applied to one input of a subtractor 112.
The second signal SBC is applied to a second input of the subtractor 112.
Thus, the output of the subtractor 112 is an error signal e, which is used to control the filter function TBONE applied by the adaptive filter 110, in such a way that the value of e is minimised. The system therefore performs a Least Mean Squares (LMS) method of adaptation. The adaptation of the filter function should take place slowly enough that the effect of noise on the signal SAC is averaged out, and hence the filter function of the block 110 becomes equal to the transfer function that needs to be applied to the signal SAC, to make it equal to the signal SBC, i.e. the transfer function TBONE in the equation above.
Thus, the resulting filter function TBONE when the system has settled is the form of the filter that can be used as the adaptive filter 92 in Figure 8. Figure 11 illustrates signals generated in the system of Figure 7, in one example.
Specifically, Figure 11 shows the signal SAC, the signal SBC* that is derived from SAC as an estimate of SBC, and the signal SBC, in one example, as functions of time. The next step in determining the correlation between the signals is to extract the energy of the signals.
Figure 12 illustrates this next step. Specifically, Figure 12 shows the signal SAC being applied to a block 120, which determines an estimate SBC* of the signal SBC, as described with reference to Figure 8.
The estimate SBC* is then applied to a first energy calculation block 122, which calculates the energy EBc*of the estimate SBC*.
At the same time, the signal SBC is applied to a second energy calculation block 124, which calculates the energy EBC of the signal SBC.
The energy calculation blocks 122, 124 can for example operate by squaring the signals and then low-pass filtering them, or by applying a Teager Kaiser Operator, but other possibilities exist.
The outputs of the two energy calculation blocks 122, 124 are passed to a comparison block 126, which compares them, and determines if they are sufficiently similar to meet a similarity threshold. For example, the comparison block 126 may determine the Pearson Correlation Coefficient, the Cosine similarity, the Euclidian distance, or any other statistical distance metric, e.g. Bhattacharya or Mahalanobis, or any other similarity metric, between the outputs of the two energy calculation blocks 122, 124.
Figure 13 illustrates the operation of this part of the system of Figure 7.
Specifically, Figure 13(a) shows with line 130 the output EBC* of the first energy calculation block 122, and shows with line 132 the output EBC of the second energy calculation block 124, in the case where the person speaking is the person who is wearing the wearable accessory, and where a fixed filter is used as the filter 92 in Figure 8.
Figure 13(a) also shows the amplitude 134 of the difference between the outputs 130 and 132.
Figure 13(b) shows with line 140 the output EBC* of the first energy calculation block 122, and shows with line 140 the output EBc of the second energy calculation block 124, in the case where the person speaking is the person who is wearing the wearable accessory, and where an adaptive filter is used as the filter 92 in Figure 8.
Figure 13(b) also shows the amplitude 144 of the difference between the outputs 140 and 142.
In both Figure 13(a) and Figure 13(b), it can be seen that the outputs of the two energy calculation blocks are very similar, and hence that the amplitudes 134, 144 of the difference signals are very small.
This implies that the signal SAC can be used to form a good estimate SBC* of the signal SBC, and hence can be used as a reliable indicator that the person speaking is the person who is wearing the wearable accessory.
Figure 13(c) shows with line 150 the output EBc*of the first energy calculation block 122, and shows with line 152 the output EBC of the second energy calculation block 124, in the case where the person speaking is not the person who is wearing the wearable accessory. Figure 13(c) also shows the amplitude 154 of the difference between the outputs 150 and 152.
In Figure 13(c), it can be seen that the outputs of the two energy calculation blocks are very different, and hence that the amplitude 154 of the difference signal is quite large.
This implies that the signal SAC cannot be used to form a good estimate SBC* of the signal SBC, and hence this can be used as a reliable indicator that the person speaking is not the person who is wearing the wearable accessory.
In fact, in this illustrated case, the energy EBC of the signal detected by the second transducer, i.e. the transducer located in the wearer’s ear canal, is very small, which is itself a good indication that the person wearing the wearable accessory is not speaking.
Thus, as described with reference to Figure 6, the correlation process described above is used for determining a first correlation between a first part of the signal SAC received by the microphone and a corresponding part of the signal SBC received by the second transducer, where the first part of the signal may correspond to the predetermined trigger phrase. The degree of correlation between them may be represented by a first correlation value. The correlation process is also used for determining a second correlation between a second part of the signal SAC received by the microphone and the corresponding part of the signal SBC received by the second transducer, where the second part of the signal may correspond to the period following the trigger phrase.
The degree of correlation between them may be represented by a second correlation value.
It is then determined whether the first correlation and the second correlation satisfy a predetermined condition. The predetermined condition may for example relate to a specific relationship between the first correlation value and the second correlation value. For example, the predetermined condition may be that the first and second correlation values are sufficiently similar that it can be assumed that the person speaking was the same during the first and second parts of the signal.
In other embodiments, the predetermined condition may be that the first and second correlation values are above a respective threshold. The two thresholds used in this determination may be the same or may be different. When the first part of the signal represents a trigger phrase, i.e. the speech content of the first part of the signal is known, the degree of correlation may be high, because it may be possible to set a useful threshold value with a high degree of confidence, whereas it is more difficult to set the threshold value for the second part of the signal, because the speech content is unknown and also of unknown length, and the degree of correlation may be lower.
The thresholds may be calculated using Decision Cost Function (DCF) methodology or Neyman-Pearson methodology. As mentioned above, the thresholds may be different, but as an example both may be set such that a correlation factor should exceed 0.8.
In the embodiment disclosed above, the first and second correlation values are obtained by examining the energies of the respective signals, during the relevant time periods. In an alternative embodiment, the first and second correlation values are obtained by calculating the Pearson correlation coefficient between the relevant part of the signal SAC and the corresponding part of the signal SBC.
Figure 14 illustrates the operation of the system of Figure 7. Specifically, Figure 14 is a timing diagram illustrating the operation of the system of Figure 7.
The top line 160 of Figure 14 shows the words being spoken and being detected by the microphone. Specifically, as shown at 162, Figure 14 shows the words “Hi phone, order me a pizza” being spoken by a first person, who is the wearer of the wearable accessory. In addition, at 164, another person then speaks, and says “with extra anchovies”.
The line 166 illustrates the output of the voice trigger detection block 78 in Figure 7. Thus, the voice trigger detection block 78 generates an output at time t1, shortly after the predetermined trigger phrase has been spoken.
In fact, the subsequent processing of the received signals only begins when it has been determined that the predetermined trigger phrase has been spoken, and the stored signals are retrieved from the buffer 74. However, for ease of reference, the following steps will be described as if they are performed directly on the received signals, rather than after a short time delay.
The line 168 illustrates the output of the voice biometric block 80 in Figure 7. Thus, the voice biometric block 80 also generates an output shortly after the predetermined trigger phrase has been spoken. In this illustrated example, it is assumed that this is a positive output, indicating that the speaker was the enrolled user.
Also, when the predetermined trigger phrase has been spoken, and the stored signals are retrieved from the buffer, the correlation block 84 is activated. The line 170 illustrates the output of the correlation block 84. As described above, the correlation block 84 produces an output that is the result of comparing two correlation values, one obtained from the predetermined trigger phrase, and one obtained from the subsequent speech.
Thus, at time t2, shortly after the predetermined trigger phrase has been completed, and the words “order me a pizza” have started to be spoken, the correlation block 84 is able to produce an output. In this case, it produces a positive output, indicating that the speaker was the same person as spoke the predetermined trigger phrase.
However, at time t3, shortly after the original speaker has finished speaking, and the other person has started to speak the words “with extra anchovies”, the correlation block 84 recognises that the person speaking now is not the same person that was speaking before.
Figure 15 illustrates the operation of the correlation block 84 during this process.
Specifically, Figure 15 shows with line 180 the output EBC* of the first energy calculation block 122, and shows with line 182 the output EBC of the second energy calculation block 124.
Figure 15 also shows the amplitude 184 of the difference between the outputs 180 and 182. Thus, it can be seen that, up until a time of about 3000 samples in Figure 15, the output EBC* of the first energy calculation block 122 is a good estimate of the output EBC of the second energy calculation block 124, and the error signal 184 has a very small amplitude.
However, for subsequent times, the output EBC of the second energy calculation block 124 itself has a very small amplitude, and hence the output EBC* of the first energy calculation block 122 is not a good estimate of the output EBC of the second energy calculation block 124, and the error signal 184 has a large amplitude.
Thus, the correlation block 84 is able to produce an output that confirms that the initial speaker was the enrolled user, but that the words “with extra anchovies” were not spoken by the enrolled user.
This means that, in effect, speaker recognition can be performed, but this is achieved in a reliable way without requiring intensive processing.
The skilled person will recognise that some aspects of the above-described apparatus and methods may be embodied as processor control code, for example on a non volatile carrier medium such as a disk, CD- or DVD-ROM, programmed memory such as read only memory (Firmware), or on a data carrier such as an optical or electrical signal carrier. For many applications embodiments of the invention will be implemented on a DSP (Digital Signal Processor), ASIC (Application Specific Integrated Circuit) or FPGA (Field Programmable Gate Array). Thus the code may comprise conventional program code or microcode or, for example code for setting up or controlling an ASIC or FPGA. The code may also comprise code for dynamically configuring re-configurable apparatus such as re-programmable logic gate arrays. Similarly the code may comprise code for a hardware description language such as Verilog TM or VHDL (Very high speed integrated circuit Hardware Description Language). As the skilled person will appreciate, the code may be distributed between a plurality of coupled components in communication with one another. Where appropriate, the embodiments may also be implemented using code running on a field- (re)programmable analogue array or similar device in order to configure analogue hardware.
Note that as used herein the term module shall be used to refer to a functional unit or block which may be implemented at least partly by dedicated hardware components such as custom defined circuitry and/or at least partly be implemented by one or more software processors or appropriate code running on a suitable general purpose processor or the like. A module may itself comprise other modules or functional units. A module may be provided by multiple components or sub-modules which need not be co-located and could be provided on different integrated circuits and/or running on different processors.
Embodiments may be implemented in a host device, especially a portable and/or battery powered host device such as a mobile computing device for example a laptop or tablet computer, a games console, a remote control device, a home automation controller or a domestic appliance including a domestic temperature or lighting control system, a toy, a machine such as a robot, an audio player, a video player, or a mobile telephone for example a smartphone. It should be noted that the above-mentioned embodiments illustrate rather than limit the invention, and that those skilled in the art will be able to design many alternative embodiments without departing from the scope of the appended claims. The word “comprising” does not exclude the presence of elements or steps other than those listed in a claim, “a” or “an” does not exclude a plurality, and a single feature or other unit may fulfil the functions of several units recited in the claims. Any reference numerals or labels in the claims shall not be construed so as to limit their scope.

Claims

1. A method of authenticating a speech signal received by a device comprising first and second transducers, wherein the first transducer comprises a microphone, the method comprising: performing a first voice biometric process on speech contained in a first part of a signal received by the microphone, in order to determine whether the speech is the speech of an enrolled user; determining a first correlation between said first part of the signal received by the microphone and a corresponding part of the signal received by the second transducer; determining a second correlation between said second part of the signal received by the microphone and the corresponding part of the signal received by the second transducer; and determining whether the first correlation and the second correlation satisfy a predetermined condition; and if it is determined that the speech contained in the first part of the received signal is the speech of an enrolled user and that the first correlation and the second correlation satisfy the predetermined condition, authenticating the received speech signal.
2. The method of claim 1, wherein the second transducer is mechanically coupled to a person wearing the device.
3. The method of claim 2, wherein the second transducer comprises a microphone, positioned to detect sound in an ear canal of the person wearing the device.
4. The method of claim 2, wherein the second transducer comprises an accelerometer, positioned to detect vibrations caused by speech of the person wearing the device.
5. The method of one of claims 1 to 4, wherein the device comprises a headset.
6. The method of one of claims 1 to 4, wherein the device comprises a pair of smart glasses.
7. The method of one of claims 1 to 6, wherein the steps of determining the first correlation and the second correlation comprise: identifying segments of the respective part of the signal received by the microphone and the corresponding part of the signal received by the second transducer in at least one acoustic class; applying a filter to at least one of said segments; and determining a degree of correlation between, the segments of the respective part of the signal received by the microphone and the corresponding part of the signal received by the second transducer, after applying said filter.
8. The method of claim 7, wherein the step of determining a degree of correlation between the segments of the respective part of the signal received by the microphone and the corresponding part of the signal received by the second transducer comprises: calculating energies in a plurality of frames of said segments; and determining whether a difference between the calculated energies is below a threshold level.
9. The method of claim 7, wherein the step of determining a degree of correlation between the segments of the respective part of the signal received by the microphone and the corresponding part of the signal received by the second transducer comprises calculating a correlation coefficient between them.
10. The method of claim 7 or 8, wherein the at least one acoustic class comprises voiced speech.
11. The method of claim 7, 8, 9 or 10, wherein the filter comprises a fixed filter.
12. The method of claim 11 , wherein the filter further comprises an adaptive gain.
13. The method of any of claims 7 to 12, wherein the filter is an adaptive filter, and wherein a filter characteristic of the adaptive filter is determined.
14. The method of any of claims 1 to 13, wherein the first biometric process is a text- dependent biometric process.
15. The method of any of claims 1 to 14, wherein determining whether the first correlation and the second correlation satisfy a predetermined condition comprises determining whether the first correlation and the second correlation have a predetermined relationship.
16. The method of claim 15, wherein determining whether the first correlation and the second correlation satisfy a predetermined condition comprises determining whether the first correlation and the second correlation are sufficiently similar.
17. The method of any of claims 1 to 14, wherein determining whether the first correlation and the second correlation satisfy a predetermined condition comprises determining whether the first correlation and the second correlation both exceed respective threshold values.
18. The method of claim 17, wherein the respective threshold values are different.
19. A system for authenticating a speech signal received by a device comprising first and second transducers, wherein the first transducer comprises a microphone, the system comprising: at least one input for receiving signals generated by the microphone and by the second transducer; and a processor configured for performing a method comprising: performing a first voice biometric process on speech contained in a first part of the signal generated by the microphone, in order to determine whether the speech is the speech of an enrolled user; determining a first correlation between said first part of the signal generated by the microphone and a corresponding part of the signal generated by the second transducer; determining a second correlation between said second part of the signal generated by the microphone and the corresponding part of the signal generated by the second transducer; and determining whether the first correlation and the second correlation satisfy a predetermined condition; and if it is determined that the speech contained in the first part of the received signal is the speech of an enrolled user and that the first correlation and the second correlation satisfy the predetermined condition, authenticating the received speech signal.
20. A computer program product, comprising non-transitory machine readable code containing instructions for causing an audio processing circuit to perform a method according to any of claims 1 to 18.
PCT/GB2021/050908 2020-05-21 2021-04-16 Authenticating received speech Ceased WO2021234342A1 (en)

Priority Applications (2)

Application Number Priority Date Filing Date Title
GB2215008.0A GB2608568B (en) 2020-05-21 2021-04-16 Authenticating received speech
CN202180030771.3A CN115461812A (en) 2020-05-21 2021-04-16 Authenticating received speech

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
US16/880,066 US11341974B2 (en) 2020-05-21 2020-05-21 Authenticating received speech
US16/880,066 2020-05-21

Publications (1)

Publication Number Publication Date
WO2021234342A1 true WO2021234342A1 (en) 2021-11-25

Family

ID=75660068

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/GB2021/050908 Ceased WO2021234342A1 (en) 2020-05-21 2021-04-16 Authenticating received speech

Country Status (4)

Country Link
US (2) US11341974B2 (en)
CN (1) CN115461812A (en)
GB (1) GB2608568B (en)
WO (1) WO2021234342A1 (en)

Families Citing this family (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US11341974B2 (en) 2020-05-21 2022-05-24 Cirrus Logic, Inc. Authenticating received speech
US20240211563A1 (en) * 2022-01-25 2024-06-27 Meta Platforms Technologies, Llc User authentication using combination of vocalization and skin vibration
KR20250098183A (en) * 2023-12-22 2025-07-01 현대자동차주식회사 Apparatus and Method for Recognizing wake-up word

Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20180324518A1 (en) * 2017-05-04 2018-11-08 Apple Inc. Automatic speech recognition triggering system
US20190295554A1 (en) * 2018-03-21 2019-09-26 Cirrus Logic International Semiconductor Ltd. Biometric processes
US10580411B2 (en) * 2017-09-25 2020-03-03 Cirrus Logic, Inc. Talker change detection
US20200075028A1 (en) * 2018-09-05 2020-03-05 Cirrus Logic International Semiconductor Ltd. Speaker recognition and speaker change detection

Family Cites Families (13)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US5548647A (en) * 1987-04-03 1996-08-20 Texas Instruments Incorporated Fixed text speaker verification method and apparatus
DE102008058883B4 (en) * 2008-11-26 2023-07-27 Lumenvox Corporation Method and arrangement for controlling user access
US20110096915A1 (en) * 2009-10-23 2011-04-28 Broadcom Corporation Audio spatialization for conference calls with multiple and moving talkers
US20190147852A1 (en) * 2015-07-26 2019-05-16 Vocalzoom Systems Ltd. Signal processing and source separation
US9838646B2 (en) * 2015-09-24 2017-12-05 Cisco Technology, Inc. Attenuation of loudspeaker in microphone array
US11494473B2 (en) * 2017-05-19 2022-11-08 Plantronics, Inc. Headset for acoustic authentication of a user
GB201804843D0 (en) * 2017-11-14 2018-05-09 Cirrus Logic Int Semiconductor Ltd Detection of replay attack
GB201803570D0 (en) * 2017-10-13 2018-04-18 Cirrus Logic Int Semiconductor Ltd Detection of replay attack
GB2608710B (en) 2018-01-23 2023-05-17 Cirrus Logic Int Semiconductor Ltd Speaker identification
US10692490B2 (en) * 2018-07-31 2020-06-23 Cirrus Logic, Inc. Detection of replay attack
US11900730B2 (en) 2019-12-18 2024-02-13 Cirrus Logic Inc. Biometric identification
US11341974B2 (en) * 2020-05-21 2022-05-24 Cirrus Logic, Inc. Authenticating received speech
WO2022225912A1 (en) * 2021-04-21 2022-10-27 Hourglass Medical Llc Methods for voice blanking muscle movement controlled systems

Patent Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20180324518A1 (en) * 2017-05-04 2018-11-08 Apple Inc. Automatic speech recognition triggering system
US10580411B2 (en) * 2017-09-25 2020-03-03 Cirrus Logic, Inc. Talker change detection
US20190295554A1 (en) * 2018-03-21 2019-09-26 Cirrus Logic International Semiconductor Ltd. Biometric processes
US20200075028A1 (en) * 2018-09-05 2020-03-05 Cirrus Logic International Semiconductor Ltd. Speaker recognition and speaker change detection

Also Published As

Publication number Publication date
GB2608568B (en) 2025-08-13
US11894000B2 (en) 2024-02-06
CN115461812A (en) 2022-12-09
US11341974B2 (en) 2022-05-24
GB2608568A (en) 2023-01-04
GB202215008D0 (en) 2022-11-23
US20210366492A1 (en) 2021-11-25
US20220238121A1 (en) 2022-07-28

Similar Documents

Publication Publication Date Title
US11694695B2 (en) Speaker identification
US12135774B2 (en) Methods, apparatus and systems for biometric processes
US11475899B2 (en) Speaker identification
US20190228778A1 (en) Speaker identification
US20210165866A1 (en) Methods, apparatus and systems for authentication
CN111903112B (en) Ear proximity detection
US12288553B2 (en) Detection of replay attack
GB2584495A (en) Methods, apparatus and systems for authentication
US11894000B2 (en) Authenticating received speech
GB2609093A (en) Speaker identification
US11900730B2 (en) Biometric identification
US11437021B2 (en) Processing audio signals
US11842725B2 (en) Detection of speech
US10896682B1 (en) Speaker recognition based on an inside microphone of a headphone
US11024318B2 (en) Speaker verification
US11710475B2 (en) Methods and apparatus for obtaining biometric data

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 21721164

Country of ref document: EP

Kind code of ref document: A1

ENP Entry into the national phase

Ref document number: 202215008

Country of ref document: GB

Kind code of ref document: A

Free format text: PCT FILING DATE = 20210416

WWE Wipo information: entry into national phase

Ref document number: 2215008.0

Country of ref document: GB

NENP Non-entry into the national phase

Ref country code: DE

WWP Wipo information: published in national office

Ref document number: 2215008.0

Country of ref document: GB

122 Ep: pct application non-entry in european phase

Ref document number: 21721164

Country of ref document: EP

Kind code of ref document: A1

WWG Wipo information: grant in national office

Ref document number: 2215008.0

Country of ref document: GB