WO2016102954A1 - Microphone unit comprising integrated speech analysis - Google Patents

Microphone unit comprising integrated speech analysis Download PDF

Info

Publication number
WO2016102954A1
WO2016102954A1 PCT/GB2015/054122 GB2015054122W WO2016102954A1 WO 2016102954 A1 WO2016102954 A1 WO 2016102954A1 GB 2015054122 W GB2015054122 W GB 2015054122W WO 2016102954 A1 WO2016102954 A1 WO 2016102954A1
Authority
WO
WIPO (PCT)
Prior art keywords
microphone unit
digital
data
speech
microphone
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/GB2015/054122
Other languages
French (fr)
Inventor
John Paul Lesso
John Laurence Melanson
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Cirrus Logic International Semiconductor Ltd
Original Assignee
Cirrus Logic International Semiconductor Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Cirrus Logic International Semiconductor Ltd filed Critical Cirrus Logic International Semiconductor Ltd
Priority to CN201580076624.4A priority Critical patent/CN107251573B/en
Priority to US15/538,619 priority patent/US10297258B2/en
Priority to GB1711576.7A priority patent/GB2551916B/en
Priority to CN202010877951.2A priority patent/CN111933158B/en
Publication of WO2016102954A1 publication Critical patent/WO2016102954A1/en
Anticipated expiration legal-status Critical
Priority to US16/380,106 priority patent/US20190259400A1/en
Ceased legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L19/00Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
    • G10L19/02Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis using spectral analysis, e.g. transform vocoders or subband vocoders
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L19/00Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
    • G10L19/04Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis using predictive techniques
    • G10L19/16Vocoder architecture
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L21/00Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
    • G10L21/02Speech enhancement, e.g. noise reduction or echo cancellation
    • G10L21/0208Noise filtering
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L25/00Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
    • G10L25/03Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the type of extracted parameters
    • G10L25/24Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the type of extracted parameters the extracted parameters being the cepstrum
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L25/00Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
    • G10L25/78Detection of presence or absence of voice signals
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04RLOUDSPEAKERS, MICROPHONES, GRAMOPHONE PICK-UPS OR LIKE ACOUSTIC ELECTROMECHANICAL TRANSDUCERS; ELECTRIC HEARING AIDS; PUBLIC ADDRESS SYSTEMS
    • H04R1/00Details of transducers, loudspeakers or microphones
    • H04R1/02Casings; Cabinets ; Supports therefor; Mountings therein
    • H04R1/04Structural association of microphone with electric circuitry therefor
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04RLOUDSPEAKERS, MICROPHONES, GRAMOPHONE PICK-UPS OR LIKE ACOUSTIC ELECTROMECHANICAL TRANSDUCERS; ELECTRIC HEARING AIDS; PUBLIC ADDRESS SYSTEMS
    • H04R19/00Electrostatic transducers
    • H04R19/04Microphones
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04RLOUDSPEAKERS, MICROPHONES, GRAMOPHONE PICK-UPS OR LIKE ACOUSTIC ELECTROMECHANICAL TRANSDUCERS; ELECTRIC HEARING AIDS; PUBLIC ADDRESS SYSTEMS
    • H04R3/00Circuits for transducers
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04RLOUDSPEAKERS, MICROPHONES, GRAMOPHONE PICK-UPS OR LIKE ACOUSTIC ELECTROMECHANICAL TRANSDUCERS; ELECTRIC HEARING AIDS; PUBLIC ADDRESS SYSTEMS
    • H04R3/00Circuits for transducers
    • H04R3/005Circuits for transducers for combining the signals of two or more microphones
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L19/00Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04RLOUDSPEAKERS, MICROPHONES, GRAMOPHONE PICK-UPS OR LIKE ACOUSTIC ELECTROMECHANICAL TRANSDUCERS; ELECTRIC HEARING AIDS; PUBLIC ADDRESS SYSTEMS
    • H04R2201/00Details of transducers, loudspeakers or microphones covered by H04R1/00 but not provided for in any of its subgroups
    • H04R2201/003Mems transducers or their use
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04RLOUDSPEAKERS, MICROPHONES, GRAMOPHONE PICK-UPS OR LIKE ACOUSTIC ELECTROMECHANICAL TRANSDUCERS; ELECTRIC HEARING AIDS; PUBLIC ADDRESS SYSTEMS
    • H04R2499/00Aspects covered by H04R or H04S not otherwise provided for in their subgroups
    • H04R2499/10General applications
    • H04R2499/11Transducers incorporated or for use in hand-held devices, e.g. mobile phones, PDA's, camera's

Definitions

  • This disclosure relates to reducing the data bit rate on the interface of digital microphones, for example to minimise power consumption in Always-On Voice modes, while still passing enough information to allow downstream keyword detection or speech recognition functions.
  • Audio functionality is becoming increasingly prevalent in portable devices. Such functionality is present not only in devices such as phones that are reliant audio technology, but also in other wearable equipment or devices that may be controlled by voice, for instance voice-responsive toys such as listening-talking teddy bears. Such devices, including phones, will spend little of their time actually transmitting speech, yet one or possibly more microphones may be permanently enabled listening out for some voice command. Even a wearable accessory may be continuously on, awaiting a voice command, and will have little space for a battery, or may rely on some solar or mechanical energy harvesting, and so has severe power consumption requirements in a continuous standby mode as well as in a low-duty-cycie operating mode.
  • Microphone transducer and amplifier technology has improved, but generally a microphone package needs to drive its output signal some distance.
  • the digital microphone signal may have some distance to go from the microphone to a centralised smart codec chip or such, along a ribbon cable, or flex, or even across a densely populated printed circuit board. Even worse are applications where the microphone may be in a headset or earbuds or some acoustically desirable position on the user's clothing, distant from the handset or the main module of a distributed device.
  • a centralised smart codec chip or such, along a ribbon cable, or flex, or even across a densely populated printed circuit board.
  • the microphone may be in a headset or earbuds or some acoustically desirable position on the user's clothing, distant from the handset or the main module of a distributed device.
  • there may be sophisticated signal processing to be performed for example speaker recognition during voice-triggered wake-up, so solutions such as grossly degrading the resolution of the ADC therein may lead to unacceptable downstream processing results.
  • FIG. 1 illustrates a conventional digital microphone 10 communicating with a smart codec 22 in a host device 20, for example a phone
  • Figure 2 illustrates the operating waveforms in a conventional digital microphone interface
  • a host device 20 transmits a clock CLK, typically at a frequency such as 3MHz, to the microphone 10, which uses this to clock an ADC 12 and to clock out from digital buffer interface Dout 14 a 1-bit oversampled delta-sigma stream DAT representing the acoustic signal input Px to the microphone transducer 16 providing the ADC input.
  • Power is consumed in the system by the host 20 transmitting this clock signal CLK, and in particular by the microphone in sending a data stream DAT with an average 1.5MHz transition rate.
  • Power may be reduced by operating at a lower clock rate, say 788kHz, but this greatly increase the in-band quantisation noise or conversely restricts the usable bandwidth for a particular noise level. Even this only reduces the power by a factor of 4, so the power consumption is still significant, particularly in larger form factor devices or long cable runs.
  • Transmitting a delta-sigma stream is notably less efficient in terms of data bit rate and transition rate than transmitting a serial multi-bit pulse-code-moduiated stream, but the latter generally requires an additional clock wire to transmit clocks to mark the start of each multi-bit word.
  • a microphone unit comprising:
  • a transducer for generating an electrical audio signal from a received acoustic signal
  • a speech coder for obtaining compressed speech data from the audio signal; and a digital output, for supplying digital signals representing said compressed speech data.
  • the microphone unit comprises a packaged microphone, for example a MEMS microphone, with on-chip or co-packaged integrated speech coder circuitry.
  • This circuitry transmits data out of this package on a PCB trace or possibly headset cable to downstream circuitry that may perform more complex functions such as speech recognition, the transmitted data representing speech information coded in a speech-compressed format at a low bit rate to reduce the power consumed in physically transmitting the data.
  • uncompressed data can be regarded as a numeric representation of samples in an uniformly sampled system, where the in-band signal is an
  • compressed data is typically derived from uncompressed data in such a way that the digital stream no longer directly represents the above, and has a lower bit rate.
  • Speech coding is an application of data compression of digital audio signals containing speech.
  • Speech coding uses speech-specific parameter estimation using audio signal processing techniques to model the speech signal, and may be combined with generic data compression algorithms to represent the resulting modelled parameters in a compact bitstream.
  • compressed speech data may be [usually digital] data representing an audio signal in terms of speech-specific parameters calculated from the signal. For example, this may be the signal energy in a set of non-uniformiy spaced frequency bins, or may use sub-band coding via say ADPCM of each sub-band.
  • Data compression techniques may then be applied to these time-varying parameters, for example recoding scaiars or vectors according to some codebook.
  • embodiments of the invention may use any speech compression standard, for example one using MDCT, MDCT-Hybrid subband, CELP, ACELP, Two- Stage Noise Feedback Coding (TSNFC), VSELP, RPE-LTP, LPC, Transform coding, or MLT, with suitable examples being AAC, AC-3, ALAC, ALS, AMBE, A r, A R-WB, AMR-WB+, apt-X, ATRAC, BroadVoice, CELT, Codec2, Enhanced AC-3, FLAG, any of the group of G.7xx standards, GSM-FR, iLBC, iSAC, Monkey's Audio, MP2, MPS, Musepack, Nellymoser Asao, Opus, Shorten, SILK, Siren 7, Speex, SVOPC, TTA, TwinVQ, Vorbis, WavPack, or Windows Media Audio.
  • any speech compression standard for example one using MDCT, MDCT-Hybrid subband, CELP, ACELP, Two- Stage
  • Figure 1 illustrates an audio processing system
  • Figure 2 illustrates signals in the audio processing system of Figure 1.
  • Figure 3 illustrates a system, comprising a host device and an accessory.
  • Figure 4 illustrates an audio processing system.
  • Figure 5 illustrates a part of a microphone unit.
  • Figure 6 illustrates a part of a microphone unit.
  • Figure 7 illustrates a part of a microphone unit.
  • Figure 8 illustrates a compressive speech coder.
  • Figure 9 illustrates an audio processing system.
  • Figure 0 illustrates an audio processing system.
  • Figure 1 1 illustrates a part of a microphone unit in the audio processing system of Figure 0.
  • Figure 3 shows an audio system, as just one example of a system using the methods described herein.
  • Figure 3 shows a device 30, which in this example takes the form of a smartphone or tablet computer.
  • the methods described herein may be used with any device, but are described herein with reference to a particular example in which the device is a portable communications device.
  • the host device 30 has audio processing capability.
  • Figure 3 shows an audio input 32, near which there is located a microphone, within the body of the device 30 and therefore not visible in Figure 3. In other devices, there may be multiple microphones.
  • Figure 3 also shows an accessory device 34, which in this example takes the form of a pair of earphones, but which may be any device, in particular any audio accessory device.
  • the pair of earphones has two earpieces 36, 38, each of which includes a speaker for reproducing sound in response to audio signals transferred from the host device 30.
  • Each of the earpieces 36, 38 also includes at least one microphone, for example for detecting ambient noise in the vicinity of the wearer. Signals representing the ambient sound are then transferred from the earphones to the host device 30. The host device may then perform various functions.
  • the host device may perform a noise cancellation function using an algorithm and generate anti-noise signals that it transfers to the earphones for playback.
  • the effect of playing back the anti-noise signals is that the level of ambient noise heard by the wearer is reduced, and the wanted sounds (music, speech, or the like) that are also being transferred from the host device 30 are therefore more audible.
  • the accessory device 34 in this example also includes a microphone 40 that is positioned near to the user's mouth when wearing the earphones. The microphone 40 is suitable for detecting the users speech.
  • the accessory device 34 may be connected to the host device 30 by means of a cable 42.
  • the cable 42 is detachable from at least one of the portable communications device and the audio accessory, in some embodiments, the cable 42 is permanently attached to the accessory device 34, and may be provided with a jack 44, to allow mechanical and electrical connection to or disconnection from the host device via a socket 46 provided on the host device.
  • the cable may be in any suitable format.
  • the host device 30 includes circuitry for receiving signals from the microphone or microphones within the body of the device 30 and/or from the microphones in the earpieces 36, 38, and/or the microphone 40.
  • the circuitry may for example include a codec 52, audio DSP, or other processing circuitry, which in turn may be connected to circuitry within the host device 30 such as an applications processor, and/or may be connected to a remote processor.
  • the processing circuitry may be able to perform a speech processing function, such as recognising the presence of a trigger phrase in a speech input received by one or more of the microphones, identifying a speaker of the speech input, and/or recognising the content of a spoken command in order to be able to control the host device or another connected device on the basis of the user's spoken command.
  • a speech processing function such as recognising the presence of a trigger phrase in a speech input received by one or more of the microphones, identifying a speaker of the speech input, and/or recognising the content of a spoken command in order to be able to control the host device or another connected device on the basis of the user's spoken command.
  • Figure 4 shows an embodiment with a microphone unit 50 having a digital
  • the microphone unit 50 comprises a transducer 54, an ana!ogue-to-information converter (AiC) 58 and a digital output driver 58.
  • AuC ana!ogue-to-information converter
  • the analogue-to-information converter 56 may take several forms. It is well known that brute force digitisation of an audio signal is grossly inefficient in terms of the useful information conveyed, or usually required, as interpreted by say the human ear and brain or some machine equivalent. The basic concept is to extract features of the audio signal that may be particularly useful when interpreted downstream, as illustrated by the data stream Fx in Figure 4. A digital interface 58 then transmits a data stream FDAT carrying this coded speech signal to the codec 52.
  • a dock recognition block 60 in the codec 52 recovers some clock from the incoming data, and then a feature processing block 62 operates on the received feature information to perform functions such as voice activity detect or speaker recognition, delivering appropriate flags VDet to downstream processing circuitry or to control or configure some further or subsequent processing of its own.
  • the codec 52 may comprise a clock generation circuit 66, or may receive a system clock from elsewhere in the host device.
  • the AIC 56 is asynchronous or self-timed in operation, so does not require a dock, and the data transmission may then also be asynchronous, as may at least early stages of processing of the feature data received by the codec. It may comprise an asynchronous ADC, for instance an Asynchronous Delta-Sigma Modulator (ADSM) followed by other analogue asynchronous circuitry or self-timed logic circuitry for digital signal processing.
  • ADSM Asynchronous Delta-Sigma Modulator
  • the microphone may generate its own clock if required by the chosen AIC circuit structure or FDAT data format.
  • the microphone unit may receive at least a low-frequency clock from the codec or elsewhere such as the system real time clock for use to synchronise or tune its internal dock generator using say locked-ioop techniques.
  • the feature data to be transmitted may typically be a frame produced at nominally say 30Hz or 10Hz, and the design of any speech processing function, say speech recognition, may have to accommodate a wide range of pitches and spoken word rate.
  • the dock in the voice recognition mode does not need an accurate or low-jitter sampling clock, so an on-chip uncalibrated low-power clock 84 may be more than adequate.
  • the data may be transmitted as a frame or vector of data at some relatively high bit rate, such that there is a transitionless interval before the each next frame.
  • a microphone unit comprises a transducer and a feature extraction block
  • the transducer may comprise a MEMS microphone, with the MEMS microphone and the feature extraction block being provided in a single integrated circuit.
  • the microphone unit may comprise a packaged microphone, for example a MEMS microphone, with on-chip or co-packaged integrated speech coder circuitry or feature extraction block.
  • This speech coder circuitry may transmit data out of the package on a PCB trace or possibly a cable such as a headset cable to downstream circuitry that may perform more complex functions such as speech recognition, the transmitted data representing speech information coded in a speech-compressed format at a low bit rate to reduce the power consumed in physically transmitting the data.
  • Figure 5 illustrates one embodiment of an AIC 56, in which the analog input signal is presented to an ADC 70, for example a 1-bit deita-sigma ADC, clocked by a sample clock CKM of nominally 768kHz.
  • the de!fa-sigma data stream Dx is then passed to a Decimator, Window block, and a Framer 72 for decimating the data to a sample rate of say 16ks/s, suitable windowing and then framing for presentation to an FFT block 74 for deriving a set of Fourier coefficients representing the power (or magnitude) of the signal in each one of a set of equally spaced frequency bins.
  • This spectral information is then passed through a mel-frequency filter bank 76 to provide estimates of the signal energy in each of a set of non-equally-spaced frequency bands. This set of energy estimates itself may be used for output.
  • each of these energy estimates is passed though a log block 78 to compand the estimate, and then though a Discrete Cosine Transform block 80 to provide cepstra! coefficients, known as Mel-frequency Cepstral Components (MFCC).
  • MFCC Mel-frequency Cepstral Components
  • the output cepstral coefficients comprise 15 channels of 12-bit words at a frame period of 30ms, thus reducing the data rate from the original 1-bit delta-sigma rate of 3Mbs/s or 786kb/s to 6kb/s.
  • Figure 6 illustrates another embodiment of an AIC 56, with some extra functional blocks in the signal path compared with Figure 5, In some other embodiments not ail of these blocks may be present.
  • the analog input signal from the transducer element 90 is presented to an ADC 92, for example a 1-bit delta-sigma ADC, clocked by a sample clock CKM of nominally 768kHz generated by a local dock generator 94, which may be synchronised to a system 32kHz Real-Time Clock for instance, or which may be independent.
  • ADC 92 for example a 1-bit delta-sigma ADC, clocked by a sample clock CKM of nominally 768kHz generated by a local dock generator 94, which may be synchronised to a system 32kHz Real-Time Clock for instance, or which may be independent.
  • the delta-sigma data stream Dx is then decimated in a decimator 96 to a sample rate of, say 16ks/s. It may then be passed to a pre-emphasis block 98 comprising a high-pass filter, to spectrally equalise the speech signal, most likely dominated by low-frequency components. This step may also be advantageous in reducing the effect of low- frequency background noise, for instance wind noise or mechanical acoustic background noise. There may also be a frequency-dependent noise reduction block at this point, as discussed below, to reduce noise in the frequency bands where it is most apparent.
  • the signal may then be passed to a windowing block 100, which may apply say a Hamming window or possibly some other windowing function to extract short-duration frames, say of time duration 10ms to 50ms, over each of which the speech may be considered stationary.
  • the windowing block extracts a stream of short-duration frames by sliding the Hamming window along the speech signal by say half the frame length, or say sliding a 25ms window by 10ms, thus providing a frame of windowed data at a frame rate of 100 frames per second.
  • An FFT 102 block then performs a Fast Fourier Transform (FFT) on the set of windowed samples of each frame, providing a set of Fourier coefficients representing the power (or magnitude) of the signal in each one of a set of equally spaced frequency bins.
  • FFT Fast Fourier Transform
  • Each of these frame-by-frame sets of signal spectral components is then processed by a Mel-filter-bank 104, which maps and combines these linear-spaced spectral components onto frequency bins distributed to correspond more closely to the nonlinear frequency sensitivity of the human ear, with a greater density of bins at low frequencies than at high frequencies. For instance, there may be 23 such bins, each with a triangular band-pass response, with the lowest frequency channel centred at 125Hz and spanning 125kHz while the highest frequency channel in centred at 3657Hz and spans 656Hz. In some embodiments, other numbers of channels or other nonlinear frequency scales such as the Bark scale may be employed.
  • a log block 108 then applies a log scaling to the energy reported from each mei- frequency bin. This helps reduce the sensitivity to very loud or very quiet sounds, in a similar way to the non-linear amplitude sensitivity of human hearing.
  • the logarithmically compressed bin energies are then passed as a set of samples to a Discrete Cosine Transform block DCT 108 which applies a Discrete Cosine Transform on each set of logarithmically compressed bin energies. This serves to separate the slowly varying spectral envelope (or vocal tract) information from the faster varying speech excitation. The former is more useful in speech recognition, so the higher coefficients may be discarded.
  • the higher order (3) coefficients may be generated in parallel with the lower ones.
  • the DCT block 108 may also provide further output data.
  • one component output may be the sum of ail the log energies from each channel, though this may also be derived by a parallel total energy estimator EST 1 10 fed from un-pre-emphasised data.
  • There may also be a dynamic coefficient generator which may generate further coefficients based on the first-order or second-order frame-to-frame differences of the coefficients.
  • An equaliser (EQ) block 112 may adaptively equalise the various components relative to a flat spectrum, for instance using an L S algorithm.
  • the data rate may be further reduced by a Data Compressor (DC) block 1 14, possibly exploiting redundancy or correlation between the coefficients expected due to the nature of speech signals.
  • DC Data Compressor
  • split vector quantisation to compress the MFCC vectors.
  • feature vectors of dimension 14 say may be split into pairs of sub vectors, each quantised to 5 or 6 bits say with a respective codebook at a frame period of 10ms. This may reduce the data rate to 4.4kb/s or lower, say 1.5kb/s if a 30ms frame period is used.
  • the data compressor may employ other standard data compression techniques.
  • the outgoing data stream may be considered compressed speech data in that the output data has been compressed from the input signal in a manner particularly suitable for speech and for communication of the parameters of a speech waveform that convey information rather than general purpose techniques of signal digitisation and of compressing arbitrary data streams.
  • this data now needs to be physically transmitted to the codec or other downstream circuitry, in the case of an accessory connected to a host device by means of a cable (such as the headset 34 containing multiple microphones being connected to an audio device 30, as shown in Figure 1)
  • the output data may be transmitted simply using two wires, one carrying the data (for example 180 bits every 30ms, in the example of Figure 5), and the second carrying a sync pulse or edge every 30ms.
  • the extra power of this low clock rate dock line is negligible compared to the already low power consumption of the data line.
  • a two-wire link may be used between a microphone in the body of a device such as a mobile phone and a codec or similar on a circuit board inside the phone.
  • Standard data formats such as SoundwireTM or SiimbusTM may be used, or standard three-wire interfaces such as I2S.
  • a one-wire serial interface may be employed, transmitting the data in a recurring predefined sequence of frames, in which a unique sync pattern may be sent at the start of every frame of words and recovered by simple and low power data and clock recovery circuitry in the destination device.
  • the clock is preferably a low-power clock inside the microphone, whose precise frequency and jitter is unimportant since the feature data is nowhere near as clock critical as full-resolution PCM.
  • Nibbles of data may be sent using a pulse-length modulated (PLM) one-wire or two- wire format such as disclosed in published US Patent Application
  • PLM pulse-length modulated
  • the data may be sent with a sequence of pulses with a fixed leading edge, with the length of each pulse denoting the binary number.
  • the fixed leading edge makes clock recovery simple.
  • Some slots in the outgoing data stream structure may be reserved for identification or control functions.
  • PLM outgoing data stream structure
  • the speech coding to reduce the data rate and thus the average transition rate on the physical bus may thus greatly reduce the power consumption of the system. This power saving may be offset somewhat by the power consumed by the speech coding itself, but this processing may have otherwise had to be performed somewhere in the system in order to provide the keyword detection or speaker recognition or more general speech recognition function in any case. Also with decreasing transistor size the power required to perform a given digital computation task is falling rapidly with time.
  • FCC Mel-frequency Cepstral Component
  • the generation method may be modified, for instances by raising the !og-me!-amp!itudes (generated by the block 78 in the embodiment shown in Figure 5 or the block 106 in the embodiment shown in Figure 6) to a suitable power (around 2 or 3) before taking the DCT (in the block 80 in the embodiment shown in Figure 5 or the block 108 in the embodiment shown in Figure 6), which reduces the influence of low-energy components.
  • the parameters of the feature extraction may be modified according to a detected or estimated signal-to-noise ratio or other signal- or noise- related parameter associated with the input signal. For instance the number and centre-frequency of the cepstral frequency bins over which the me!-frequency energy is extracted may be modified.
  • the cepstral coding block may comprise or be preceded by a noise reduction block for instance directly after a decimation block 72 or 96 or after a pre-emphasis block 98 that may already have removed some low frequency noise, or operating on the windowed framed data produced by block 100.
  • This noise reduction block may be enabled when necessary by a noise detection block.
  • the noise detection block may be analog and monitor the input signal Ax, or it may be digital and operate on the ADC output Dx.
  • the noise detection block may flag when the level or spectrum or other characteristic of the received signal implies a high noise level or when the ratio of peak or average signal to noise falls below a threshold.
  • the noise reduction circuitry may act to filter the signal to suppress frequency bins where the noise, as monitored in time periods when there appears to be no voice as monitored by a Voice Activity Detector, is likely to exceed the signal at times where there is a signal.
  • a Wiener Filter set up may be used to suppress noise on a frame-by-frame basis.
  • the Wiener filter coefficients may be updated on a frame-by- frame basis and coefficient smoothed via a Mel-frequency filter bank followed by an Inverse Discrete Cosine Transform before application to the actual signal.
  • the Wiener noise reduction may comprise two stages. Each stage may incorporate some dynamic noise enhancement feature where the level of noise reduction performed is dependent on an estimated signal-to-noise ratio or other signal- or noise-related parameter or feature of the signal.
  • Various signal coding techniques where the output data transmitted is derived from the signal energies associated with each of a filter bank with non-uniformly spaced centre frequencies as described above, particularly cepstrai feature extraction using MFCC coding, are compatible with many known downstream voice recognition or speaker recognition algorithms.
  • the MFCC data may actually be forwarded from the codec, for example in an ETSI-standard MFCC form, for further signal processing either within the host device or transmitted to remote servers for processing "in the cloud".
  • the microphone may be required to deliver a more traditional output signal digitising the instantaneous input audio signal in say a 16-bit format at say 16ks/s or 48ks/s.
  • Figure 7 illustrates a microphone unit 130 which may operate in a plurality of modes, with various degrees and methods of signal coding or compression. Thus, Figure 7 shows several different functional blocks. In some other embodiments, only a subset of these blocks is present.
  • ADC 134 for example a 1-bit delta-sigma ADC
  • Dx the resulting deita-sigma data stream Dx is then passed to one or more functional blocks, as described below,
  • the ADC may be clocked by a sample clock CKIV1, which may be generated by a local dock generator 136, or which may be received on a clock input 138 according to the operating mode.
  • the microphone unit may operate in a first, low power , mode in which it uses an internally generated clock and provides compressed speech data and a second, higher power, mode in which it receives an external clock and provides uncompressed data.
  • the operating mode may be controlled by the downstream controi processor via signals received on a control input terminal 140. These inputs may be separate or may be provided by making the digital output line bi-directional, in some embodiments the operating mode may be determined autonomously by circuitry in the microphone unit.
  • a control block 142 receives the control inputs and determines which of the functional blocks are to be active.
  • Figure 7 shows that the data stream Dx may be passed to a PDM formatting block 144, which allows the digitised time-domain output of the microphone to be output directly as a PDM stream.
  • the output of the PDM formatting block 144 is passed to a multiplexer 146, operating under the controi of the control block 142, and the multiplexer output is passed to a driver 148 for generating the digital output DAT.
  • Figure 7 also shows the data stream Dx being passed to a feature extraction block 150, for example for obtaining values based on using non-linear-spaced frequency bins, for instance MFCC values.
  • Figure 7 also shows the data stream Dx being passed to a compressive sampling block 152, for example for deriving a sparse representation of the incoming signal.
  • Figure 7 also shows the data stream Dx being passed to a lossy compression block 154, for example for performing Adaptive differential pulse-code modulation (ADPCM) or a similar form of coding.
  • ADPCM Adaptive differential pulse-code modulation
  • Figure 7 also shows the data stream Dx being passed to a decimator 156.
  • the data stream Dx is also passed to a lossless coding block to provide a suitable output data stream.
  • Figure 7 shows the outputs of the compressive sampling block 152, lossy compression block 154 and decimator 156 being connected to respective data buffer memory blocks 158, 160, 162. These allow the higher quality data generated by these blocks to be stored. Then, if analysis of a lower power data stream suggests that there is a need, power can be expended in transmitting the higher quality data for some further processing or examination that requires such higher quality data.
  • analysis of a lower power data stream might suggest that the audio signal contains a trigger phrase being spoken by a recognized user of the device in a particular time period, in that case, it is possible to read from one of the buffer memory blocks the higher qualify data relating to the same period of time, and to perform further analysis on that data, for example to confirm whether the trigger phrase was in fact spoken, or whether the trigger phrase was spoken by the recognized user, or performing more detailed keyword detection before awakening a greater part of a downstream system.
  • the higher quality data can be used for downstream operations requiring better data, for example downstream speech recognition.
  • Figure 7 also shows the outputs of the feature extraction block 150, compressive audio processing block 152, and lossy compression block 154 being output through respective pulse length modulation (PLM) encoding blocks 164, 166, 168 and through the multiplexer 146, operating under the control of the control block 142, and the driver 148.
  • PLM pulse length modulation
  • Figure 7 also shows the output of the decimator 156 being output through a pulse code modulation (PCM) encoding block 170 and through the multiplexer 146, operating under the control of the control block 142, and the driver 148.
  • PCM pulse code modulation
  • the physical form of the transmitted output may differ according to what operating mode is selected. For instance high-data-rate modes may be transmitted using low-guerage-differential signalling for noise immunity, and the data scrambled to reduce emissions.
  • the signal in low-data rate modes, the signal may be low- bandwidth and not so susceptible to noise and transmission line reflections and suchlike, and is preferably unterminated to save the power consumption associated with driving the termination resistance. In the lower power modes the signal swing, i.e. digital driver supply voltage, may be reduced.
  • circuitry may also be altered according to signal mode. For instance in low data rate modes the speed requirements of the DSP operations may be modest, and the circuitry may thus be operated at a lower logic supply voltage or divided master clock frequency than when performing more complex operations in conjunction with higher rate coding.
  • the AIC or feature-extraction based schemes above may provide particularly efficient methods of encoding and transmitting the essential information in the audio signal, there may be a requirement for the microphone unit to be able to operate also so as to provide a more conventional data format, say for processing by local circuitry or onward transmission for processing in the cloud, where such processing might not understand the more sophisticated signal representation, or where, for example, the current use case is for recording music in high quality.
  • the microphone unit is operable in a first mode in which feature extraction and/or data compression is performed, and in a second mode, in which (for example) a clock is supplied from the codec, and the unit operates in a manner similar to that shown in Figure 1.
  • the digital microphone unit may thus be capable of operating in at least two modes - ADC (analog-digital conversion) or AlC (analog-information conversion).
  • ADC mode the PCM data from the ADC is transmitted, in AlC mode data extracted from the ADC output is coded, particularly for speech.
  • the microphone unit is operable in one mode to perform lossy low-bit-rate PCM coding.
  • the unit may contain a lossy codec such as an ADPCM coder, with a sample rate that in some embodiments may be selectable, for example between 8ks/s - 24ks/s.
  • the microphone unit has a coding block for performing ⁇ -iaw and/or A-iaw coding, or coding to some other telephony standard.
  • the microphone unit in some embodiments has coding blocks for DCT, MDCT-Hybrid subband, CELP, ACELP, Two-Stage Noise Feedback Coding (TSNFC), VSELP, RPE- LTP, LPC, Transform coding, or LT coding.
  • the microphone unit is operable in a mode in which it outputs compressive sampled PCM data, or any scheme that exploits signal sparseness.
  • Figure 8 illustrates an embodiment of a compressive speech coder that may be used in any of the embodiments described or illustrated herein.
  • the output of an ADC 190 is passed through a decimator 192 to provide (for example) 12 bit data at 16ks/s or
  • This data is sampled at an average sample rate of only say 48Hz or 1 kHz but with a sampling time randomised by a suitable random number generator or random pulse generator 194.
  • the sampling circuit samples the input signal at a sample rate less than the input signal bandwidth, and the sampling instants are caused to be distributed randomly in time.
  • Figure 9 shows a system using such a compressive speech coder.
  • the microphone unit 200 including a compressive ADC 202, is connected to supply very low data rate data to a codec 204.
  • downstream circuitry 206 may either perform a partial reconstruction
  • the sparse representation of a signal is matched to a linear combination of a few atoms from a pre-defined dictionary, which atoms may be obtained a priori by using machine learning techniques to learn an over- complete dictionary of primary signals (atoms) directly from data, so that the most relevant properties of the signals may be efficiently captured.
  • Sparse extraction may have some benefits in performance of feature extraction in the presence of noise.
  • the noise is not recognised as comprising any atom component, so does not appear in the encoded data.
  • Such ignorance of input noise may thus avoid unnecessary activation of downstream circuitry and avoid the power consumption increasing in noisy environments relative to quiet environments.
  • Figure 10 shows an embodiment in which a microphone unit 210 is connected to supply very low data rate data to a codec 212, and in which, to further reduce the power consumption, some, if not all, of the feature extraction is performed using Analogue Signal Processing (ASP).
  • ASP Analogue Signal Processing
  • This a signal from a microphone transducer is passed to an analog signal processor 214, and then to one or more anaiog-to-digital converter 216, and then to an optional digital signal processor 218 in the microphone unit 210.
  • Feature recognition 220 is then performed in the codec 212.
  • FIG. 11 shows in more detail the processing inside an embodiment of the
  • microphone unit 210 in which a large part of the signal processing is performed by analogue rather than digital circuitry.
  • an input signal is passed through a plurality of band pass filters (three being shown in Figure 1 purely by way of illustration) 240, 242, 246.
  • the band pass filters are constant Q and equally spaced in mel frequency.
  • the outputs are passed to log function blocks 248, 250, 252, which may be achieved using standard analogue design techniques based for instance on applying the input signal via a voltage-to-current converted signal into an l-V two-port circuit with a logarithmic current-to-voltage conversion such as a semiconductor diode.
  • the outputs are passed to a plurality of parallel ADCs 252, 254, 256.
  • the ADCs may comprise voltage-controlled oscillators, whose frequency is used as a representation of their respective input signals. These are simple and low power, and their linearity is not important in this application. These simple ADCs may have significantly reduced power and area even in total compared to the main ADC. The state of the art for similar circuit blocks in say the field of artificial cochieas is below 20 microwatts.
  • the microphone, ADC, and the speech coding circuitry may advantageously be located close together, to reduce high-data- rate signal paths of the digital data before data-rate reduction. All three components may be packaged together. At least two of these three components may be co- integrated on an integrated circuit.
  • the microphone may be a MEMS transducer, which may be capacitive, piezo-electric or piezo resistive, and co-integrated with at least the ADC.
  • embodiments of the above-described apparatus and methods may be, at least partly, implemented using programmable components rather than dedicated hardwired components.
  • embodiments of the apparatus and methods may be, at least partly embodied as processor control code, for example on a non transitory carrier medium such as a disk, CD- or DVD-ROM, programmed memory such as read only memory (Firmware), or on a data carrier such as an optical or electrical signal carrier.
  • embodiments of the invention may be implemented, at least partly, by a DSP (Digital Signal Processor), ASIC (Application Specific Integrated Circuit) or FPGA (Field Programmable Gate Array),
  • the code may comprise conventional program code or microcode or, for example code for setting up or controlling an ASIC or FPGA,
  • the code may also comprise code for dynamically configuring re-configurable apparatus such as reprogrammable logic gate arrays.
  • the code may comprise code for a hardware description language such as VeriiogTM or VHDL (Very high speed integrated circuit Hardware Description Language), As the skilled person will appreciate, the code may be distributed between a plurality of coupled components in communication with one another.
  • the embodiments may also be implemented using code running on a fieid-(re-)programmable analogue array or similar device in order to configure analogue hardware. It should be understood—especially by those having ordinary skill in the art with the benefit of this disclosure—that that the various operations described herein, particularly in connection with the figures, may be implemented by other circuitry or other hardware components. The order in which each operation of a given method is performed may be changed, and various elements of the systems illustrated herein may be added, reordered, combined, omitted, modified, etc. It is intended that this disclosure embrace all such modifications and changes and, accordingly, the above description should be regarded in an illustrative rather than a restrictive sense.

Landscapes

  • Engineering & Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Acoustics & Sound (AREA)
  • Signal Processing (AREA)
  • Health & Medical Sciences (AREA)
  • Human Computer Interaction (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Computational Linguistics (AREA)
  • Multimedia (AREA)
  • Quality & Reliability (AREA)
  • General Health & Medical Sciences (AREA)
  • Otolaryngology (AREA)
  • Spectroscopy & Molecular Physics (AREA)
  • Compression, Expansion, Code Conversion, And Decoders (AREA)
  • Soundproofing, Sound Blocking, And Sound Damping (AREA)

Abstract

A microphone unit has a transducer, for generating an electrical audio signal from a received acoustic signal; a speech coder, for obtaining compressed speech data from the audio signal; and a digital output, for supplying digital signals representing said compressed speech data. The speech coder may be a lossy speech coder, and may contain a bank of filters with centre frequencies that are non-uniformly spaced, for example mel frequencies.

Description

MICROPHONE UNIT COMPRISING INTEGRATED SPEECH ANALYSIS
This disclosure relates to reducing the data bit rate on the interface of digital microphones, for example to minimise power consumption in Always-On Voice modes, while still passing enough information to allow downstream keyword detection or speech recognition functions.
BACKGROUND
Audio functionality is becoming increasingly prevalent in portable devices. Such functionality is present not only in devices such as phones that are reliant audio technology, but also in other wearable equipment or devices that may be controlled by voice, for instance voice-responsive toys such as listening-talking teddy bears. Such devices, including phones, will spend little of their time actually transmitting speech, yet one or possibly more microphones may be permanently enabled listening out for some voice command. Even a wearable accessory may be continuously on, awaiting a voice command, and will have little space for a battery, or may rely on some solar or mechanical energy harvesting, and so has severe power consumption requirements in a continuous standby mode as well as in a low-duty-cycie operating mode.
Microphone transducer and amplifier technology has improved, but generally a microphone package needs to drive its output signal some distance. Digital
transmission offers advantages including noise immunity, but the usual formats for transmission of digital data from microphones are not particularly efficient in terms of signal line activity and the consequent power consumed in charging parasitic capacitances though a supply voltage at every logic level transition.
In a portable device such as a phone or tablet, containing one or more digital microphone, the digital microphone signal may have some distance to go from the microphone to a centralised smart codec chip or such, along a ribbon cable, or flex, or even across a densely populated printed circuit board. Even worse are applications where the microphone may be in a headset or earbuds or some acoustically desirable position on the user's clothing, distant from the handset or the main module of a distributed device. However, even when largely otherwise inactive, there may be sophisticated signal processing to be performed, for example speaker recognition during voice-triggered wake-up, so solutions such as grossly degrading the resolution of the ADC therein may lead to unacceptable downstream processing results.
There is thus a requirement to reduce the power consumed in sending digital microphone data across a wired digital transmission link, while still conveying enough useful information in the transmitted signal to allow downstream function such as speech recognition.
Figure 1 illustrates a conventional digital microphone 10 communicating with a smart codec 22 in a host device 20, for example a phone, and Figure 2 illustrates the operating waveforms in a conventional digital microphone interface. A host device 20 transmits a clock CLK, typically at a frequency such as 3MHz, to the microphone 10, which uses this to clock an ADC 12 and to clock out from digital buffer interface Dout 14 a 1-bit oversampled delta-sigma stream DAT representing the acoustic signal input Px to the microphone transducer 16 providing the ADC input. Power is consumed in the system by the host 20 transmitting this clock signal CLK, and in particular by the microphone in sending a data stream DAT with an average 1.5MHz transition rate.
Power may be reduced by operating at a lower clock rate, say 788kHz, but this greatly increase the in-band quantisation noise or conversely restricts the usable bandwidth for a particular noise level. Even this only reduces the power by a factor of 4, so the power consumption is still significant, particularly in larger form factor devices or long cable runs.
Transmitting a delta-sigma stream is notably less efficient in terms of data bit rate and transition rate than transmitting a serial multi-bit pulse-code-moduiated stream, but the latter generally requires an additional clock wire to transmit clocks to mark the start of each multi-bit word.
Secondly we note that an unfortunate side effect of reducing the delta-sigma sample clock rate may be to limit the bandwidth usable in terms of background quantisation noise to say 8kHz rather than say 20kHz. This may increase the word error rate (WER) for Voice Key Word detection (VKD). This may in turn lead to a higher incidence of false positives and the system may spend more time in its awake mode thus significantly affecting the average complete system power consumption.
Additionally there is also a prevalent requirement for functions requiring even more accurate input audio data streams, such as speaker identification, as part of a voice- triggered wake-up function. It is known that using a wider bandwidth for the speaker identification captures more speech signal components and thus relaxes the need for high signal~to-noise (SNR) (e.g. relaxes the need for low acoustic background noise, or carefully optimised microphone placement) to get high enough accuracy for biometric purposes. Even in a high SNR environment a relatively wide signal bandwidth may improve speaker verification accuracy. This is at odds with the concept of reducing the frequency of the digital microphone clock to reduce the power consumption.
SUMMARY
According to a first aspect of the invention, there is provided a microphone unit, comprising:
a transducer, for generating an electrical audio signal from a received acoustic signal;
a speech coder, for obtaining compressed speech data from the audio signal; and a digital output, for supplying digital signals representing said compressed speech data.
In embodiments of the invention, the microphone unit comprises a packaged microphone, for example a MEMS microphone, with on-chip or co-packaged integrated speech coder circuitry. This circuitry transmits data out of this package on a PCB trace or possibly headset cable to downstream circuitry that may perform more complex functions such as speech recognition, the transmitted data representing speech information coded in a speech-compressed format at a low bit rate to reduce the power consumed in physically transmitting the data.
In this disclosure, uncompressed data can be regarded as a numeric representation of samples in an uniformly sampled system, where the in-band signal is an
approximation, in the audio band, of the audio input waveform, whereas compressed data is typically derived from uncompressed data in such a way that the digital stream no longer directly represents the above, and has a lower bit rate.
Speech coding is an application of data compression of digital audio signals containing speech. Speech coding uses speech-specific parameter estimation using audio signal processing techniques to model the speech signal, and may be combined with generic data compression algorithms to represent the resulting modelled parameters in a compact bitstream. Thus, compressed speech data may be [usually digital] data representing an audio signal in terms of speech-specific parameters calculated from the signal. For example, this may be the signal energy in a set of non-uniformiy spaced frequency bins, or may use sub-band coding via say ADPCM of each sub-band. Data compression techniques may then be applied to these time-varying parameters, for example recoding scaiars or vectors according to some codebook.
As examples, embodiments of the invention may use any speech compression standard, for example one using MDCT, MDCT-Hybrid subband, CELP, ACELP, Two- Stage Noise Feedback Coding (TSNFC), VSELP, RPE-LTP, LPC, Transform coding, or MLT, with suitable examples being AAC, AC-3, ALAC, ALS, AMBE, A r, A R-WB, AMR-WB+, apt-X, ATRAC, BroadVoice, CELT, Codec2, Enhanced AC-3, FLAG, any of the group of G.7xx standards, GSM-FR, iLBC, iSAC, Monkey's Audio, MP2, MPS, Musepack, Nellymoser Asao, Opus, Shorten, SILK, Siren 7, Speex, SVOPC, TTA, TwinVQ, Vorbis, WavPack, or Windows Media Audio.
Figure 1 illustrates an audio processing system.
Figure 2 illustrates signals in the audio processing system of Figure 1. Figure 3 illustrates a system, comprising a host device and an accessory.
Figure 4 illustrates an audio processing system. Figure 5 illustrates a part of a microphone unit. Figure 6 illustrates a part of a microphone unit.
Figure 7 illustrates a part of a microphone unit. Figure 8 illustrates a compressive speech coder. Figure 9 illustrates an audio processing system. Figure 0 illustrates an audio processing system.
Figure 1 1 illustrates a part of a microphone unit in the audio processing system of Figure 0.
DETAILED DESCRIPTION
Figure 3 shows an audio system, as just one example of a system using the methods described herein.
Specifically, Figure 3 shows a device 30, which in this example takes the form of a smartphone or tablet computer. The methods described herein may be used with any device, but are described herein with reference to a particular example in which the device is a portable communications device. Thus, in this example, the host device 30 has audio processing capability.
Figure 3 shows an audio input 32, near which there is located a microphone, within the body of the device 30 and therefore not visible in Figure 3. In other devices, there may be multiple microphones. Figure 3 also shows an accessory device 34, which in this example takes the form of a pair of earphones, but which may be any device, in particular any audio accessory device. In this example, the pair of earphones has two earpieces 36, 38, each of which includes a speaker for reproducing sound in response to audio signals transferred from the host device 30. Each of the earpieces 36, 38 also includes at least one microphone, for example for detecting ambient noise in the vicinity of the wearer. Signals representing the ambient sound are then transferred from the earphones to the host device 30. The host device may then perform various functions. For example, the host device may perform a noise cancellation function using an algorithm and generate anti-noise signals that it transfers to the earphones for playback. The effect of playing back the anti-noise signals is that the level of ambient noise heard by the wearer is reduced, and the wanted sounds (music, speech, or the like) that are also being transferred from the host device 30 are therefore more audible. The accessory device 34 in this example also includes a microphone 40 that is positioned near to the user's mouth when wearing the earphones. The microphone 40 is suitable for detecting the users speech. The accessory device 34 may be connected to the host device 30 by means of a cable 42. The cable 42 is detachable from at least one of the portable communications device and the audio accessory, in some embodiments, the cable 42 is permanently attached to the accessory device 34, and may be provided with a jack 44, to allow mechanical and electrical connection to or disconnection from the host device via a socket 46 provided on the host device. The cable may be in any suitable format. The host device 30 includes circuitry for receiving signals from the microphone or microphones within the body of the device 30 and/or from the microphones in the earpieces 36, 38, and/or the microphone 40. The circuitry may for example include a codec 52, audio DSP, or other processing circuitry, which in turn may be connected to circuitry within the host device 30 such as an applications processor, and/or may be connected to a remote processor.
For example, the processing circuitry may be able to perform a speech processing function, such as recognising the presence of a trigger phrase in a speech input received by one or more of the microphones, identifying a speaker of the speech input, and/or recognising the content of a spoken command in order to be able to control the host device or another connected device on the basis of the user's spoken command.
Figure 4 shows an embodiment with a microphone unit 50 having a digital
transmission format and method, for communication to a downstream smart codec 52, audio DSP, or other processing circuitry. The microphone unit 50 comprises a transducer 54, an ana!ogue-to-information converter (AiC) 58 and a digital output driver 58.
The analogue-to-information converter 56, or speech coder, or feature extraction block, may take several forms. It is well known that brute force digitisation of an audio signal is grossly inefficient in terms of the useful information conveyed, or usually required, as interpreted by say the human ear and brain or some machine equivalent. The basic concept is to extract features of the audio signal that may be particularly useful when interpreted downstream, as illustrated by the data stream Fx in Figure 4. A digital interface 58 then transmits a data stream FDAT carrying this coded speech signal to the codec 52. in one embodiment, a dock recognition block 60 in the codec 52 recovers some clock from the incoming data, and then a feature processing block 62 operates on the received feature information to perform functions such as voice activity detect or speaker recognition, delivering appropriate flags VDet to downstream processing circuitry or to control or configure some further or subsequent processing of its own. The codec 52 may comprise a clock generation circuit 66, or may receive a system clock from elsewhere in the host device.
Preferably the AIC 56 is asynchronous or self-timed in operation, so does not require a dock, and the data transmission may then also be asynchronous, as may at least early stages of processing of the feature data received by the codec. It may comprise an asynchronous ADC, for instance an Asynchronous Delta-Sigma Modulator (ADSM) followed by other analogue asynchronous circuitry or self-timed logic circuitry for digital signal processing.
However the microphone may generate its own clock if required by the chosen AIC circuit structure or FDAT data format.
In some embodiments the microphone unit may receive at least a low-frequency clock from the codec or elsewhere such as the system real time clock for use to synchronise or tune its internal dock generator using say locked-ioop techniques. However, as will be discussed below, the feature data to be transmitted may typically be a frame produced at nominally say 30Hz or 10Hz, and the design of any speech processing function, say speech recognition, may have to accommodate a wide range of pitches and spoken word rate. Thus in contrast to the use case where music has to be recorded with accurate pitch and where any jitter may lead to unmusical intermoduiation, the dock in the voice recognition mode does not need an accurate or low-jitter sampling clock, so an on-chip uncalibrated low-power clock 84 may be more than adequate. In some embodiments, the data may be transmitted as a frame or vector of data at some relatively high bit rate, such that there is a transitionless interval before the each next frame.
In ail of the embodiments described herein, in which a microphone unit comprises a transducer and a feature extraction block, the transducer may comprise a MEMS microphone, with the MEMS microphone and the feature extraction block being provided in a single integrated circuit.
The microphone unit may comprise a packaged microphone, for example a MEMS microphone, with on-chip or co-packaged integrated speech coder circuitry or feature extraction block.
This speech coder circuitry, or feature extraction block, may transmit data out of the package on a PCB trace or possibly a cable such as a headset cable to downstream circuitry that may perform more complex functions such as speech recognition, the transmitted data representing speech information coded in a speech-compressed format at a low bit rate to reduce the power consumed in physically transmitting the data. Figure 5 illustrates one embodiment of an AIC 56, in which the analog input signal is presented to an ADC 70, for example a 1-bit deita-sigma ADC, clocked by a sample clock CKM of nominally 768kHz. The de!fa-sigma data stream Dx is then passed to a Decimator, Window block, and a Framer 72 for decimating the data to a sample rate of say 16ks/s, suitable windowing and then framing for presentation to an FFT block 74 for deriving a set of Fourier coefficients representing the power (or magnitude) of the signal in each one of a set of equally spaced frequency bins. This spectral information is then passed through a mel-frequency filter bank 76 to provide estimates of the signal energy in each of a set of non-equally-spaced frequency bands. This set of energy estimates itself may be used for output. Alternatively, each of these energy estimates is passed though a log block 78 to compand the estimate, and then though a Discrete Cosine Transform block 80 to provide cepstra! coefficients, known as Mel-frequency Cepstral Components (MFCC).
In one example the output cepstral coefficients comprise 15 channels of 12-bit words at a frame period of 30ms, thus reducing the data rate from the original 1-bit delta-sigma rate of 3Mbs/s or 786kb/s to 6kb/s.
Figure 6 illustrates another embodiment of an AIC 56, with some extra functional blocks in the signal path compared with Figure 5, In some other embodiments not ail of these blocks may be present.
The analog input signal from the transducer element 90 is presented to an ADC 92, for example a 1-bit delta-sigma ADC, clocked by a sample clock CKM of nominally 768kHz generated by a local dock generator 94, which may be synchronised to a system 32kHz Real-Time Clock for instance, or which may be independent.
The delta-sigma data stream Dx is then decimated in a decimator 96 to a sample rate of, say 16ks/s. It may then be passed to a pre-emphasis block 98 comprising a high-pass filter, to spectrally equalise the speech signal, most likely dominated by low-frequency components. This step may also be advantageous in reducing the effect of low- frequency background noise, for instance wind noise or mechanical acoustic background noise. There may also be a frequency-dependent noise reduction block at this point, as discussed below, to reduce noise in the frequency bands where it is most apparent.
The signal may then be passed to a windowing block 100, which may apply say a Hamming window or possibly some other windowing function to extract short-duration frames, say of time duration 10ms to 50ms, over each of which the speech may be considered stationary. The windowing block extracts a stream of short-duration frames by sliding the Hamming window along the speech signal by say half the frame length, or say sliding a 25ms window by 10ms, thus providing a frame of windowed data at a frame rate of 100 frames per second. An FFT 102 block then performs a Fast Fourier Transform (FFT) on the set of windowed samples of each frame, providing a set of Fourier coefficients representing the power (or magnitude) of the signal in each one of a set of equally spaced frequency bins.
Each of these frame-by-frame sets of signal spectral components is then processed by a Mel-filter-bank 104, which maps and combines these linear-spaced spectral components onto frequency bins distributed to correspond more closely to the nonlinear frequency sensitivity of the human ear, with a greater density of bins at low frequencies than at high frequencies. For instance, there may be 23 such bins, each with a triangular band-pass response, with the lowest frequency channel centred at 125Hz and spanning 125kHz while the highest frequency channel in centred at 3657Hz and spans 656Hz. In some embodiments, other numbers of channels or other nonlinear frequency scales such as the Bark scale may be employed.
A log block 108 then applies a log scaling to the energy reported from each mei- frequency bin. This helps reduce the sensitivity to very loud or very quiet sounds, in a similar way to the non-linear amplitude sensitivity of human hearing. The logarithmically compressed bin energies are then passed as a set of samples to a Discrete Cosine Transform block DCT 108 which applies a Discrete Cosine Transform on each set of logarithmically compressed bin energies. This serves to separate the slowly varying spectral envelope (or vocal tract) information from the faster varying speech excitation. The former is more useful in speech recognition, so the higher coefficients may be discarded. However is some embodiments these may be preserved, or possibly combined by weighted addition to provide at least some measure of energy for higher frequencies to aid in distinguishing sibilants or providing more clues for speaker identification, in some embodiments the higher order (3) coefficients may be generated in parallel with the lower ones.
The DCT block 108 may also provide further output data. For instance one component output may be the sum of ail the log energies from each channel, though this may also be derived by a parallel total energy estimator EST 1 10 fed from un-pre-emphasised data. There may also be a dynamic coefficient generator which may generate further coefficients based on the first-order or second-order frame-to-frame differences of the coefficients. An equaliser (EQ) block 112 may adaptively equalise the various components relative to a flat spectrum, for instance using an L S algorithm.
Before transmission, the data rate may be further reduced by a Data Compressor (DC) block 1 14, possibly exploiting redundancy or correlation between the coefficients expected due to the nature of speech signals. For example split vector quantisation to compress the MFCC vectors. In one example feature vectors of dimension 14 say may be split into pairs of sub vectors, each quantised to 5 or 6 bits say with a respective codebook at a frame period of 10ms. This may reduce the data rate to 4.4kb/s or lower, say 1.5kb/s if a 30ms frame period is used.
Additionally or alternatively, the data compressor may employ other standard data compression techniques.
Thus the data rate necessary to carry useful information concerning the speech content of the acoustic input signal has been reduced below that necessary for simple multi-bit or oversampled time-domain representations of the actual waveform by employing compression techniques at least in part reliant on known general characteristics of speech waveforms and of the human perception of speech, for instance in the use of non-iineariy-spaced filter banks and logarithmic compression, or the separation of vocal tract information from excitation information referred to above. The outgoing data stream may be considered compressed speech data in that the output data has been compressed from the input signal in a manner particularly suitable for speech and for communication of the parameters of a speech waveform that convey information rather than general purpose techniques of signal digitisation and of compressing arbitrary data streams.
Having generated the compressed speech data, this data now needs to be physically transmitted to the codec or other downstream circuitry, in the case of an accessory connected to a host device by means of a cable (such as the headset 34 containing multiple microphones being connected to an audio device 30, as shown in Figure 1), the output data may be transmitted simply using two wires, one carrying the data (for example 180 bits every 30ms, in the example of Figure 5), and the second carrying a sync pulse or edge every 30ms. The extra power of this low clock rate dock line is negligible compared to the already low power consumption of the data line. Similarly a two-wire link may be used between a microphone in the body of a device such as a mobile phone and a codec or similar on a circuit board inside the phone.
Standard data formats such as Soundwire™ or Siimbus™ may be used, or standard three-wire interfaces such as I2S. Alternatively, , a one-wire serial interface may be employed, transmitting the data in a recurring predefined sequence of frames, in which a unique sync pattern may be sent at the start of every frame of words and recovered by simple and low power data and clock recovery circuitry in the destination device. The clock is preferably a low-power clock inside the microphone, whose precise frequency and jitter is unimportant since the feature data is nowhere near as clock critical as full-resolution PCM.
Nibbles of data may be sent using a pulse-length modulated (PLM) one-wire or two- wire format such as disclosed in published US Patent Application
(US2013/0197920(A1)). The data may be sent with a sequence of pulses with a fixed leading edge, with the length of each pulse denoting the binary number. The fixed leading edge makes clock recovery simple.
Some slots in the outgoing data stream structure (PLM or non-PLM) may be reserved for identification or control functions. In this application, with continuous streams of data, occasional data bit errors may not have a serious impact. However, in some applications it may be desirable to protect at least the control data with some error- detection and/or correction scheme, e.g. based on Cyclic Redundancy Check bits embedded in the stream. The speech coding to reduce the data rate and thus the average transition rate on the physical bus may thus greatly reduce the power consumption of the system. This power saving may be offset somewhat by the power consumed by the speech coding itself, but this processing may have otherwise had to be performed somewhere in the system in order to provide the keyword detection or speaker recognition or more general speech recognition function in any case. Also with decreasing transistor size the power required to perform a given digital computation task is falling rapidly with time.
It is known that Mel-frequency Cepstral Component ( FCC) values are not very robust in the presence of additive noise. This may lead to false positives from a downstream voice keyword detector, which may lead to this block frequently triggering futile power- up of following circuitry, with a significant effect on average system power
consumption.
In some embodiments the generation method may be modified, for instances by raising the !og-me!-amp!itudes (generated by the block 78 in the embodiment shown in Figure 5 or the block 106 in the embodiment shown in Figure 6) to a suitable power (around 2 or 3) before taking the DCT (in the block 80 in the embodiment shown in Figure 5 or the block 108 in the embodiment shown in Figure 6), which reduces the influence of low-energy components.
In some embodiments the parameters of the feature extraction may be modified according to a detected or estimated signal-to-noise ratio or other signal- or noise- related parameter associated with the input signal. For instance the number and centre-frequency of the cepstral frequency bins over which the me!-frequency energy is extracted may be modified.
In some embodiments, the cepstral coding block may comprise or be preceded by a noise reduction block for instance directly after a decimation block 72 or 96 or after a pre-emphasis block 98 that may already have removed some low frequency noise, or operating on the windowed framed data produced by block 100. This noise reduction block may be enabled when necessary by a noise detection block. The noise detection block may be analog and monitor the input signal Ax, or it may be digital and operate on the ADC output Dx. The noise detection block may flag when the level or spectrum or other characteristic of the received signal implies a high noise level or when the ratio of peak or average signal to noise falls below a threshold.
The noise reduction circuitry may act to filter the signal to suppress frequency bins where the noise, as monitored in time periods when there appears to be no voice as monitored by a Voice Activity Detector, is likely to exceed the signal at times where there is a signal. For instance a Wiener Filter set up may be used to suppress noise on a frame-by-frame basis. The Wiener filter coefficients may be updated on a frame-by- frame basis and coefficient smoothed via a Mel-frequency filter bank followed by an Inverse Discrete Cosine Transform before application to the actual signal.
In some embodiments the Wiener noise reduction may comprise two stages. Each stage may incorporate some dynamic noise enhancement feature where the level of noise reduction performed is dependent on an estimated signal-to-noise ratio or other signal- or noise-related parameter or feature of the signal. Various signal coding techniques where the output data transmitted is derived from the signal energies associated with each of a filter bank with non-uniformly spaced centre frequencies as described above, particularly cepstrai feature extraction using MFCC coding, are compatible with many known downstream voice recognition or speaker recognition algorithms. In some cases the MFCC data may actually be forwarded from the codec, for example in an ETSI-standard MFCC form, for further signal processing either within the host device or transmitted to remote servers for processing "in the cloud". This latter may reduce the data bandwidth required for transmission, and may be used to preserve speech quality in poor transmission conditions. However in some embodiments the microphone may be required to deliver a more traditional output signal digitising the instantaneous input audio signal in say a 16-bit format at say 16ks/s or 48ks/s.
There may also be other applications in which some other format of signal is required. Traditionally this processing and re-formatting of the signal might take place within a phone applications processor or a smart codec with DSP capability. However given the presence of DSP circuitry in the microphone unit, necessary to reduce digital transmission power in stand-by or "Aiways-On" modes, this DSP circuitry may be usable to perform other speech coding methods in other use cases. As semiconductor manufacturing processes evolve with ever-decreasing feature sizes, and as the cost of each of these processes decreases over time with maturity, it becomes more feasible to actually integrate this functionality in the microphone unit itself, leaving any more powerful processing power elsewhere in the system freer to perform higher-level tasks. Or indeed in some end applications, the requirement for other signal-processing DSP may be removed, allowing perhaps some simpler non-DSP controller processor to be used. Figure 7 illustrates a microphone unit 130 which may operate in a plurality of modes, with various degrees and methods of signal coding or compression. Thus, Figure 7 shows several different functional blocks. In some other embodiments, only a subset of these blocks is present.
The analog input signal from a transducer element 132 is presented to an ADC 134, for example a 1-bit delta-sigma ADC, and the resulting deita-sigma data stream Dx is then passed to one or more functional blocks, as described below,
The ADC may be clocked by a sample clock CKIV1, which may be generated by a local dock generator 136, or which may be received on a clock input 138 according to the operating mode. The microphone unit may operate in a first, low power , mode in which it uses an internally generated clock and provides compressed speech data and a second, higher power, mode in which it receives an external clock and provides uncompressed data.
The operating mode may be controlled by the downstream controi processor via signals received on a control input terminal 140. These inputs may be separate or may be provided by making the digital output line bi-directional, in some embodiments the operating mode may be determined autonomously by circuitry in the microphone unit. A control block 142 receives the control inputs and determines which of the functional blocks are to be active.
Thus, Figure 7 shows that the data stream Dx may be passed to a PDM formatting block 144, which allows the digitised time-domain output of the microphone to be output directly as a PDM stream. The output of the PDM formatting block 144 is passed to a multiplexer 146, operating under the controi of the control block 142, and the multiplexer output is passed to a driver 148 for generating the digital output DAT.
Figure 7 also shows the data stream Dx being passed to a feature extraction block 150, for example for obtaining values based on using non-linear-spaced frequency bins, for instance MFCC values. Figure 7 also shows the data stream Dx being passed to a compressive sampling block 152, for example for deriving a sparse representation of the incoming signal.
Figure 7 also shows the data stream Dx being passed to a lossy compression block 154, for example for performing Adaptive differential pulse-code modulation (ADPCM) or a similar form of coding.
Figure 7 also shows the data stream Dx being passed to a decimator 156.
In some embodiments, the data stream Dx is also passed to a lossless coding block to provide a suitable output data stream.
Figure 7 shows the outputs of the compressive sampling block 152, lossy compression block 154 and decimator 156 being connected to respective data buffer memory blocks 158, 160, 162. These allow the higher quality data generated by these blocks to be stored. Then, if analysis of a lower power data stream suggests that there is a need, power can be expended in transmitting the higher quality data for some further processing or examination that requires such higher quality data.
For example, analysis of a lower power data stream might suggest that the audio signal contains a trigger phrase being spoken by a recognized user of the device in a particular time period, in that case, it is possible to read from one of the buffer memory blocks the higher qualify data relating to the same period of time, and to perform further analysis on that data, for example to confirm whether the trigger phrase was in fact spoken, or whether the trigger phrase was spoken by the recognized user, or performing more detailed keyword detection before awakening a greater part of a downstream system. Thus, the higher quality data can be used for downstream operations requiring better data, for example downstream speech recognition.
Figure 7 also shows the outputs of the feature extraction block 150, compressive audio processing block 152, and lossy compression block 154 being output through respective pulse length modulation (PLM) encoding blocks 164, 166, 168 and through the multiplexer 146, operating under the control of the control block 142, and the driver 148. Figure 7 also shows the output of the decimator 156 being output through a pulse code modulation (PCM) encoding block 170 and through the multiplexer 146, operating under the control of the control block 142, and the driver 148.
The physical form of the transmitted output may differ according to what operating mode is selected. For instance high-data-rate modes may be transmitted using low- voitage-differential signalling for noise immunity, and the data scrambled to reduce emissions. On the other hand, in low-data rate modes, the signal may be low- bandwidth and not so susceptible to noise and transmission line reflections and suchlike, and is preferably unterminated to save the power consumption associated with driving the termination resistance. In the lower power modes the signal swing, i.e. digital driver supply voltage, may be reduced.
Other operating parameters of the circuit may also be altered according to signal mode. For instance in low data rate modes the speed requirements of the DSP operations may be modest, and the circuitry may thus be operated at a lower logic supply voltage or divided master clock frequency than when performing more complex operations in conjunction with higher rate coding.
Although the AIC or feature-extraction based schemes above may provide particularly efficient methods of encoding and transmitting the essential information in the audio signal, there may be a requirement for the microphone unit to be able to operate also so as to provide a more conventional data format, say for processing by local circuitry or onward transmission for processing in the cloud, where such processing might not understand the more sophisticated signal representation, or where, for example, the current use case is for recording music in high quality.
In this case, it is advantageous for the initial conversion in ADC to be high quality, requiring a high quality low-jitter clock, and preferably synchronous with codec DSP main clock to avoid issues with sample-rate conversion to be synchronous to the codec master clock and/or the reference sample rate of the standard output digital PCM format. Thus, the microphone unit is operable in a first mode in which feature extraction and/or data compression is performed, and in a second mode, in which (for example) a clock is supplied from the codec, and the unit operates in a manner similar to that shown in Figure 1. The digital microphone unit may thus be capable of operating in at least two modes - ADC (analog-digital conversion) or AlC (analog-information conversion). In ADC mode the PCM data from the ADC is transmitted, in AlC mode data extracted from the ADC output is coded, particularly for speech.
In further embodiments, the microphone unit is operable in one mode to perform lossy low-bit-rate PCM coding. For example, the unit may contain a lossy codec such as an ADPCM coder, with a sample rate that in some embodiments may be selectable, for example between 8ks/s - 24ks/s.
In some embodiments, the microphone unit has a coding block for performing μ-iaw and/or A-iaw coding, or coding to some other telephony standard. For example, the microphone unit in some embodiments has coding blocks for DCT, MDCT-Hybrid subband, CELP, ACELP, Two-Stage Noise Feedback Coding (TSNFC), VSELP, RPE- LTP, LPC, Transform coding, or LT coding.
In other embodiments, the microphone unit is operable in a mode in which it outputs compressive sampled PCM data, or any scheme that exploits signal sparseness.
Figure 8 illustrates an embodiment of a compressive speech coder that may be used in any of the embodiments described or illustrated herein. The output of an ADC 190 is passed through a decimator 192 to provide (for example) 12 bit data at 16ks/s or
48ks/s. This data is sampled at an average sample rate of only say 48Hz or 1 kHz but with a sampling time randomised by a suitable random number generator or random pulse generator 194. Thus, the sampling circuit samples the input signal at a sample rate less than the input signal bandwidth, and the sampling instants are caused to be distributed randomly in time.
Figure 9 shows a system using such a compressive speech coder. Thus, the microphone unit 200, including a compressive ADC 202, is connected to supply very low data rate data to a codec 204. With the aid of prior knowledge of the signal statistics, downstream circuitry 206 may either perform a partial reconstruction
(computationally cheap) to do sparse feature extraction in a low power mode, or a full reconstruction (computationally more expensive) to get Nyquist type voice for onwards transmission. Note there are known post-processing algorithm blocks, such as the block 208, for performing "sparse recognition" that are compatible with such
compressive sampling formats. In such algorithms the sparse representation of a signal is matched to a linear combination of a few atoms from a pre-defined dictionary, which atoms may be obtained a priori by using machine learning techniques to learn an over- complete dictionary of primary signals (atoms) directly from data, so that the most relevant properties of the signals may be efficiently captured.
Sparse extraction may have some benefits in performance of feature extraction in the presence of noise. The noise is not recognised as comprising any atom component, so does not appear in the encoded data. Such ignorance of input noise may thus avoid unnecessary activation of downstream circuitry and avoid the power consumption increasing in noisy environments relative to quiet environments.
Figure 10 shows an embodiment in which a microphone unit 210 is connected to supply very low data rate data to a codec 212, and in which, to further reduce the power consumption, some, if not all, of the feature extraction is performed using Analogue Signal Processing (ASP). This a signal from a microphone transducer is passed to an analog signal processor 214, and then to one or more anaiog-to-digital converter 216, and then to an optional digital signal processor 218 in the microphone unit 210. Feature recognition 220 is then performed in the codec 212.
Figure 11 shows in more detail the processing inside an embodiment of the
microphone unit 210, in which a large part of the signal processing is performed by analogue rather than digital circuitry. Thus, an input signal is passed through a plurality of band pass filters (three being shown in Figure 1 purely by way of illustration) 240, 242, 246. The band pass filters are constant Q and equally spaced in mel frequency. The outputs are passed to log function blocks 248, 250, 252, which may be achieved using standard analogue design techniques based for instance on applying the input signal via a voltage-to-current converted signal into an l-V two-port circuit with a logarithmic current-to-voltage conversion such as a semiconductor diode. The outputs are passed to a plurality of parallel ADCs 252, 254, 256. The ADCs may comprise voltage-controlled oscillators, whose frequency is used as a representation of their respective input signals. These are simple and low power, and their linearity is not important in this application. These simple ADCs may have significantly reduced power and area even in total compared to the main ADC. The state of the art for similar circuit blocks in say the field of artificial cochieas is below 20 microwatts.
In all of the embodiments described herein, the microphone, ADC, and the speech coding circuitry may advantageously be located close together, to reduce high-data- rate signal paths of the digital data before data-rate reduction. All three components may be packaged together. At least two of these three components may be co- integrated on an integrated circuit.
The microphone may be a MEMS transducer, which may be capacitive, piezo-electric or piezo resistive, and co-integrated with at least the ADC.
The skilled person will recognise that various embodiments of the above-described apparatus and methods may be, at least partly, implemented using programmable components rather than dedicated hardwired components. Thus embodiments of the apparatus and methods may be, at least partly embodied as processor control code, for example on a non transitory carrier medium such as a disk, CD- or DVD-ROM, programmed memory such as read only memory (Firmware), or on a data carrier such as an optical or electrical signal carrier. In some applications, embodiments of the invention may be implemented, at least partly, by a DSP (Digital Signal Processor), ASIC (Application Specific Integrated Circuit) or FPGA (Field Programmable Gate Array), Thus the code may comprise conventional program code or microcode or, for example code for setting up or controlling an ASIC or FPGA, The code may also comprise code for dynamically configuring re-configurable apparatus such as reprogrammable logic gate arrays. Similarly the code may comprise code for a hardware description language such as Veriiog™ or VHDL (Very high speed integrated circuit Hardware Description Language), As the skilled person will appreciate, the code may be distributed between a plurality of coupled components in communication with one another. Where appropriate, the embodiments may also be implemented using code running on a fieid-(re-)programmable analogue array or similar device in order to configure analogue hardware. It should be understood— especially by those having ordinary skill in the art with the benefit of this disclosure— that that the various operations described herein, particularly in connection with the figures, may be implemented by other circuitry or other hardware components. The order in which each operation of a given method is performed may be changed, and various elements of the systems illustrated herein may be added, reordered, combined, omitted, modified, etc. It is intended that this disclosure embrace all such modifications and changes and, accordingly, the above description should be regarded in an illustrative rather than a restrictive sense. Similarly, although this disclosure makes reference to specific embodiments, certain modifications and changes can be made to those embodiments without departing from the scope and coverage of this disclosure. Moreover, any benefits, advantages, or solutions to problems that are described herein with regard to specific embodiments are not intended to be construed as a critical, required, or essential feature or element.
Further embodiments likewise, with the benefit of this disclosure, will be apparent to those having ordinary skill in the art, and such embodiments should be deemed as being encompassed herein.

Claims

1 . A microphone unit, comprising:
a transducer, for generating an electrical audio signal from a received acoustic signal;
a speech coder, for obtaining compressed speech data from the audio signal; and
a digital output, for supplying digital signals representing said compressed speech data.
2. A microphone unit as claimed in claim 1 , wherein the speech coder contains a bank of filters with non-uniformly spaced centre frequencies.
3. A microphone unit as claimed in claim 2, wherein the centre frequencies are mei frequencies.
4. A microphone unit as claimed in claim 2 or 3, wherein outputs of the bank of filters are coupled to a log weighting block and a discrete cosine transform block to provide cepstral coefficients.
5. A microphone unit as claimed in one of claims 1 to 4, operable in a mode in which the digital output supplies uncompressed speech data.
6. A microphone unit as claimed in claim 1 , comprising a compressive sampling coder comprising a sampling circuit which samples the input signal at a sample rate less than the input signal bandwidth, wherein the sampling instants are caused to be distributed randomly in time.
7. A microphone unit as claimed in one of claims 1 to 6, wherein the speech coder is a lossy speech coder.
8. A microphone unit as claimed in claim 7, wherein the lossy speech coder uses at least one coding technique selected from ADPCM, MDCT, MDCT-Hybrid subband, CELP, ACELP, Two-Stage Noise Feedback Coding (TSNFC), VSELP, RPE-LTP, LPC, Transform coding, and MLT.
9. A microphone unit as claimed in claimed in one of claims 1 to 8, wherein the speech coder and the digital output are provided on a single integrated circuit.
10. A microphone unit as claimed in claim 9, wherein the transducer is provided on said integrated circuit.
1 1 . A microphone unit as claimed in claim 1 0, wherein the transducer comprises a MEMS microphone.
12. A microphone unit as claimed in one of claims 1 to 1 1 , wherein the transducer is for generating an analog audio signal,
further comprising an analog-to-digital converter, for generating a digital audio signal from the analog audio signal, wherein the speech coder is connected to the analog-to-digital converter, for obtaining speech feature values from the digital audio output signal.
13. A microphone unit as claimed in one of claims 1 to 12, further comprising data compression circuitry, for receiving said cepsfrai feature values from the cepstral feature extraction block, and for generating reduced bit rate signals for supply to said digital output.
14. A microphone unit as claimed in claim 1 3,
wherein said reduced bit rate signals contain data at 20k bits/second or lower.
15. A microphone unit as claimed in one of claims 1 to 14, wherein the digital output is configurable in a first mode, for supplying said digital signals representing said compressed speech data, and in a second mode, for supplying digital signals representing time samples of said audio signal.
16. A microphone unit as claimed in claim 1 5, wherein the digital signals representing time samples of said audio signal are supplied at a higher data rate than the digital signals representing said compressed speech data.
17, A microphone unit as claimed in claim 1 6, wherein the digital signals representing time samples of said audio signal are PCM signals supplied at a sample rate of 16kHz or higher.
18. A microphone unit as claimed in claim 16, wherein the digital signals representing time samples of said audio signal are PDM signals at a sample rate of 0.5MHz or higher.
19. A microphone unit as claimed in one of claims 15 to 18, wherein the digital output is configurable between the first mode and the second mode based on a command received from a separate device to which the digital output is connected.
20. A microphone unit as claimed in one of claims 15 to 18, wherein the digital output is configurable between the first mode and the second mode based on a command generated in the microphone unit.
21 . A microphone unit as claimed in claim 20, comprising a voice activity detector, for generating said command.
22. A microphone unit as claimed in claim 21 , wherein said command causes the digital output to enter the first mode in response to detecting voice activity in the audio signal.
23. A microphone unit as claimed in one of claims 1 to 22, further comprising noise cancelling circuitry, for reducing the effects of ambient noise in the output digital signals.
24. A microphone unit as claimed in claim 23, wherein the noise cancelling circuitry comprises a Wiener filter.
25. A microphone unit as claimed in claim 23, wherein the noise cancelling circuitry comprises an adaptive Wiener filter.
26. A microphone unit, comprising:
a transducer, for generating an audio signal;
a bit rate compression block, for obtaining bit-rate compressed data from the audio signal; and
a digital output, for supplying digital signals representing said bit-rate
compressed data.
27. A microphone unit as claimed in claim 26, wherein the transducer comprises a MEMS microphone.
28. A microphone unit, comprising:
a transducer, for generating an audio signal;
a speech feature extraction block, for obtaining speech feature values from the audio signal: and
a digital output, for supplying digital signals representing said speech feature values.
29. A microphone unit as claimed in claim 27, wherein the speech feature extraction block comprises means for obtaining speech feature values from the audio signal by means of a filter bank for providing estimates of the signal energy in each of a set of non-equally-spaced frequency bands.
30. A microphone unit as claimed in claim 29, wherein the non-equaily-spaced frequency bands are mel frequencies.
31 . A microphone unit as claimed in one of claims 28 to 30, wherein the speech feature extraction block comprises means for obtaining cepstral features.
32. A microphone unit as claimed in one of claims 28 to 31 , wherein the speech feature extraction block and the digital output are provided on a single integrated circuit.
33. A microphone unit as claimed in claim 32, wherein the transducer is provided on said integrated circuit.
34. A microphone unit as claimed in claim 33, wherein the transducer comprises a MEMS microphone.
35. A microphone unit as claimed in one of claims 28 to 34,
wherein the transducer is for generating an analog audio signal,
further comprising an analog-to-digital converter, for generating a digital audio signal from the analog audio signal, wherein the speech feature extraction block is connected to the analog-to-digital converter, for obtaining speech feature values from the digital audio output signal.
36. A microphone unit as claimed in claim 35 when dependent on one of claims 32 to 34, wherein the analog-to-digital converter is provided on said integrated circuit,
37. A microphone unit as claimed in one of claims 28 to 36, further comprising data compression circuitry, for receiving said speech feature values from the speech feature extraction block, and for generating reduced bit rate signals for supply to said digital output.
38. A microphone unit as claimed in claim 37, wherein the data compression circuitry operates using a predetermined code book,
39. A microphone unit as claimed in claim 37, wherein the data compression circuitry performs multi-vector coding.
40, A microphone unit as claimed in one of claims 28 to 39,
wherein said digital signals contain data at 20k bits/second or lower.
41 . A microphone unit as claimed in one of claims 28 to 40, wherein the digital output is configurable in a first mode, for supplying said digital signals representing said speech feature values, and in a second mode, for supplying digital signals representing time samples of said audio signal.
42, A microphone unit as claimed in one of claims 28 to 40, wherein the digital output is configurable in a first mode, for supplying said digital signals representing a data compressed version of said speech feature values, and in a second mode, for supplying digital signals representing time samples of said audio signal.
43. A microphone unit as claimed in claim 42, wherein said digital signals
representing the data compressed version of said speech feature values comprise ADPCM compressed data.
44, A microphone unit as claimed in one of claims 41 to 43, w erein the digital signals representing time samples of said audio signal are supplied at a higher data rate than the digital signals representing said speech feature values.
45. A microphone unit as claimed in claim 44, wherein the digital signals representing time samples of said audio signal are PCM signals supplied at a sample rate of 16kHz or higher.
46. A microphone unit as claimed in claim 44 or 45, wherein the digital signals representing time samples of said audio signal are PDM signals at a sample rate of 0.5MHz or higher.
47. A microphone unit as claimed in one of claims 42 to 46, wherein the digital output is configurable between the first mode and the second mode based on a command received from a separate device to which the digital output is connected.
48. A microphone unit as claimed in one of claims 42 to 46, wherein the digital output is configurable between the first mode and the second mode based on a command generated in the microphone unit.
49. A microphone unit as claimed in claim 48, comprising a voice activity detector, for generating said command,
50. A microphone unit as claimed in claim 49, wherein said command causes the digital output to enter the first mode in response to detecting voice activity in the audio signal.
51 . A microphone unit as claimed in one of claims 28 to 50, further comprising noise cancelling circuitry, for reducing the effects of ambient noise in the output digital signals,
52. A microphone unit as claimed in claim 51 , wherein the noise cancelling circuitry comprises a Wiener filter.
53. A microphone unit as claimed in claim 52, wherein the noise cancelling circuitry comprises an adaptive Wiener filter.
54. A microphone unit as claimed in one of claims 28 to 53, wherein the speech feature extraction block comprises a plurality of analog filters providing signals in different respective frequency bands, and means for extracting features from said signals,
55. A microphone unit as claimed in claim 54, wherein the speech feature extraction block comprises a plurality of analog-to-digital converters connected to the outputs of the analog filters.
56. A microphone unit, comprising:
a transducer, for generating an audio output signal;
a feature extraction block, for deriving speech features from the audio output signal; and
a digital output, for driving digital data representing said speech features over a digital data line.
57. A microphone unit as claimed in claim 56, wherein the transducer comprises a MEMS microphone, and the MEMS microphone and the feature extraction block are provided in a single integrated circuit,
58. A microphone unit, comprising:
a transducer, for generating an audio output signal;
a feature extraction block, for deriving speech features from the audio output signal; and
a digital output, having a first mode for driving digital data representing said audio output over a digital data line at a first data rate, and having a second mode for driving digital data representing said derived features over a digital data line at a second data rate,
59. A microphone unit as claimed in claim 58, wherein the transducer comprises a MEMS microphone, and the MEMS microphone and the feature extraction block are provided in a single integrated circuit.
60. A microphone unit, comprising:
a transducer, for generating an audio output signal; a digitiser, for generating a stream of digitised values representing said audio output signal;
a feature extraction block, for deriving speech features from the audio output signal; and
a digital output, having a first mode for driving said stream of digitised values over a digital data line at a first data rate, and having a second mode for driving digital data representing said derived features over a digital data line at a second data rate.
61 . A microphone unit as claimed in claim 60, wherein the transducer comprises a MEMS microphone, and the MEMS microphone and the feature extraction block are provided in a single integrated circuit,
62. A digital microphone capable of operating in two modes:
a first ADC mode where the output is a digital version of the analogue input signal; and
a second AIC mode where the output is coded using the fact that input signal has a sparse representation in a given domain.
63. A portable device, having a microphone unit as claimed in any of claims 1 to 61 connectabie thereto, the portable device comprising voice processing circuitry for receiving said digital signals representing said speech feature values.
64. A portable device as claimed in claim 63, wherein the voice processing circuitry is adapted for performing voice processing on the received digital signals.
65. A portable device as claimed in claim 64, wherein the voice processing circuitry comprises speech recognition circuitry for recognizing a speaker.
66. A portable device as claimed in claim 64 or 65, wherein the voice processing circuitry comprises speech recognition circuitry for recognizing spoken words in an input.
67. A portable device as claimed in claim 63, wherein the voice processing circuitry is adapted for transmitting data contained in said received digital signals to a remote processor.
68, A portable device as claimed in claim 67, wherein the voice processing circuitry is adapted for transmitting said data to the remote processor in an ETSI-standard MFCC form.
69. A portable device as claimed in one of ciaims 63 to 68, wherein the microphone unit as claimed in one of claims 1 to 61 is connectabie to the portable device by a wired connection.
70, A portable device as claimed in one of ciaims 63 to 69, wherein the microphone unit as claimed in one of claims 1 to 61 is provided in a headset.
71 . A portable device as claimed in one of claims 63 to 70, wherein the portable device comprises a portable communications device, a computing device, a wearable device, a games device, or a noise cancelling device such as a NC headset.
72. An audio processing system, comprising:
a microphone unit and an audio processing unit, wherein the microphone unit and the audio processing unit are connectabie by means of a wired connection,
wherein the microphone unit comprises:
a transducer, for generating an audio output signal;
a feature extraction block, for deriving speech features from the audio output signal; and
a digital output, for driving digital data representing said speech features over the wired connection.
73, An audio processing system as claimed in claim 72, wherein the transducer comprises a MEMS microphone, and the MEMS microphone and the feature extraction block are provided in a single integrated circuit.
PCT/GB2015/054122 2014-12-23 2015-12-22 Microphone unit comprising integrated speech analysis Ceased WO2016102954A1 (en)

Priority Applications (5)

Application Number Priority Date Filing Date Title
CN201580076624.4A CN107251573B (en) 2014-12-23 2015-12-22 Includes microphone unit with integrated speech analysis
US15/538,619 US10297258B2 (en) 2014-12-23 2015-12-22 Microphone unit comprising integrated speech analysis
GB1711576.7A GB2551916B (en) 2014-12-23 2015-12-22 Microphone unit comprising integrated speech analysis
CN202010877951.2A CN111933158B (en) 2014-12-23 2015-12-22 Microphone unit including integrated speech analysis
US16/380,106 US20190259400A1 (en) 2014-12-23 2019-04-10 Microphone unit comprising integrated speech analysis

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
US201462096424P 2014-12-23 2014-12-23
US62/096,424 2014-12-23

Related Child Applications (2)

Application Number Title Priority Date Filing Date
US15/538,619 A-371-Of-International US10297258B2 (en) 2014-12-23 2015-12-22 Microphone unit comprising integrated speech analysis
US16/380,106 Continuation US20190259400A1 (en) 2014-12-23 2019-04-10 Microphone unit comprising integrated speech analysis

Publications (1)

Publication Number Publication Date
WO2016102954A1 true WO2016102954A1 (en) 2016-06-30

Family

ID=53677602

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/GB2015/054122 Ceased WO2016102954A1 (en) 2014-12-23 2015-12-22 Microphone unit comprising integrated speech analysis

Country Status (4)

Country Link
US (2) US10297258B2 (en)
CN (2) CN107251573B (en)
GB (3) GB201509483D0 (en)
WO (1) WO2016102954A1 (en)

Families Citing this family (27)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
GB201509483D0 (en) * 2014-12-23 2015-07-15 Cirrus Logic Internat Uk Ltd Feature extraction
GB2578386B (en) 2017-06-27 2021-12-01 Cirrus Logic Int Semiconductor Ltd Detection of replay attack
GB2563953A (en) 2017-06-28 2019-01-02 Cirrus Logic Int Semiconductor Ltd Detection of replay attack
GB201713697D0 (en) 2017-06-28 2017-10-11 Cirrus Logic Int Semiconductor Ltd Magnetic detection of replay attack
GB201801527D0 (en) 2017-07-07 2018-03-14 Cirrus Logic Int Semiconductor Ltd Method, apparatus and systems for biometric processes
GB201801532D0 (en) 2017-07-07 2018-03-14 Cirrus Logic Int Semiconductor Ltd Methods, apparatus and systems for audio playback
GB201801664D0 (en) 2017-10-13 2018-03-21 Cirrus Logic Int Semiconductor Ltd Detection of liveness
GB2567503A (en) 2017-10-13 2019-04-17 Cirrus Logic Int Semiconductor Ltd Analysing speech signals
GB201804843D0 (en) 2017-11-14 2018-05-09 Cirrus Logic Int Semiconductor Ltd Detection of replay attack
GB201801659D0 (en) * 2017-11-14 2018-03-21 Cirrus Logic Int Semiconductor Ltd Detection of loudspeaker playback
EP3714452B1 (en) * 2017-11-23 2023-02-15 Harman International Industries, Incorporated Method and system for speech enhancement
US11264037B2 (en) 2018-01-23 2022-03-01 Cirrus Logic, Inc. Speaker identification
US11735189B2 (en) 2018-01-23 2023-08-22 Cirrus Logic, Inc. Speaker identification
US11475899B2 (en) 2018-01-23 2022-10-18 Cirrus Logic, Inc. Speaker identification
DE102018204687B3 (en) 2018-03-27 2019-06-13 Infineon Technologies Ag MEMS microphone module
US10692490B2 (en) 2018-07-31 2020-06-23 Cirrus Logic, Inc. Detection of replay attack
US10915614B2 (en) 2018-08-31 2021-02-09 Cirrus Logic, Inc. Biometric authentication
US11637546B2 (en) * 2018-12-14 2023-04-25 Synaptics Incorporated Pulse density modulation systems and methods
CN110191397B (en) * 2019-06-28 2021-10-15 歌尔科技有限公司 Noise reduction method and Bluetooth headset
KR102740717B1 (en) * 2019-08-30 2024-12-11 엘지전자 주식회사 Artificial sound source separation method and device of thereof
US12325627B2 (en) 2020-01-27 2025-06-10 Infineon Technologies Ag Configurable microphone using internal clock changing
US12302064B2 (en) 2020-01-27 2025-05-13 Infineon Technologies Ag Configurable microphone using internal clock changing
US11582560B2 (en) * 2020-11-30 2023-02-14 Infineon Technologies Ag Digital microphone with low data rate interface
CN114187914A (en) * 2021-12-17 2022-03-15 广东电网有限责任公司 Voice recognition method and system
CN116645972B (en) * 2023-02-03 2025-11-28 电子科技大学 Tooth pitch suppression method based on sparse decomposition
US12563348B2 (en) * 2024-04-02 2026-02-24 Infineon Technologies Ag Idle tone mitigation using clock jitter
CN118155608B (en) * 2024-05-11 2024-07-19 米烁网络科技(广州)有限公司 Miniature microphone voice recognition system for multi-noise environment

Citations (7)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US5612974A (en) * 1994-11-01 1997-03-18 Motorola Inc. Convolutional encoder for use on an integrated circuit that performs multiple communication tasks
EP1154408A2 (en) * 2000-05-10 2001-11-14 Kabushiki Kaisha Toshiba Multimode speech coding and noise reduction
US20060034472A1 (en) * 2004-08-11 2006-02-16 Seyfollah Bazarjani Integrated audio codec with silicon audio transducer
WO2010030889A1 (en) * 2008-09-11 2010-03-18 Personics Holdings Inc. Method and system for sound monitoring over a network
US20130197920A1 (en) 2011-12-14 2013-08-01 Wolfson Microelectronics Plc Data transfer
US20140257813A1 (en) * 2013-03-08 2014-09-11 Analog Devices A/S Microphone circuit assembly and system with speech recognition
WO2014189931A1 (en) * 2013-05-23 2014-11-27 Knowles Electronics, Llc Vad detection microphone and method of operating the same

Family Cites Families (24)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US5737485A (en) * 1995-03-07 1998-04-07 Rutgers The State University Of New Jersey Method and apparatus including microphone arrays and neural networks for speech/speaker recognition systems
US5966688A (en) * 1997-10-28 1999-10-12 Hughes Electronics Corporation Speech mode based multi-stage vector quantizer
US20020193989A1 (en) * 1999-05-21 2002-12-19 Michael Geilhufe Method and apparatus for identifying voice controlled devices
US6397186B1 (en) * 1999-12-22 2002-05-28 Ambush Interactive, Inc. Hands-free, voice-operated remote control transmitter
US7076260B1 (en) * 2000-03-21 2006-07-11 Agere Systems Inc. Unbalanced coding for cordless telephony
US20040156520A1 (en) * 2002-04-10 2004-08-12 Poulsen Jens Kristian Miniature digital transducer with reduced number of external terminals
US7099821B2 (en) * 2003-09-12 2006-08-29 Softmax, Inc. Separation of target acoustic signals in a multi-transducer arrangement
US20060206320A1 (en) * 2005-03-14 2006-09-14 Li Qi P Apparatus and method for noise reduction and speech enhancement with microphones and loudspeakers
US20080013747A1 (en) 2006-06-30 2008-01-17 Bao Tran Digital stethoscope and monitoring instrument
WO2009059279A1 (en) * 2007-11-01 2009-05-07 University Of Maryland Compressive sensing system and method for bearing estimation of sparse sources in the angle domain
US8099289B2 (en) * 2008-02-13 2012-01-17 Sensory, Inc. Voice interface and search for electronic devices including bluetooth headsets and remote systems
US8085941B2 (en) * 2008-05-02 2011-12-27 Dolby Laboratories Licensing Corporation System and method for dynamic sound delivery
US8171322B2 (en) * 2008-06-06 2012-05-01 Apple Inc. Portable electronic devices with power management capabilities
KR20110134127A (en) * 2010-06-08 2011-12-14 삼성전자주식회사 Audio data decoding apparatus and method
CN102074245B (en) * 2011-01-05 2012-10-10 瑞声声学科技(深圳)有限公司 Dual-microphone-based speech enhancement device and speech enhancement method
CN102074246B (en) * 2011-01-05 2012-12-19 瑞声声学科技(深圳)有限公司 Dual-microphone based speech enhancement device and method
JP6136218B2 (en) * 2012-12-03 2017-05-31 富士通株式会社 Sound processing apparatus, method, and program
WO2014132168A1 (en) * 2013-02-22 2014-09-04 Marvell World Trade Ltd. Multi-slot multi-point audio interface
CN104247280A (en) * 2013-02-27 2014-12-24 视听公司 Voice-controlled communication connections
US10020008B2 (en) * 2013-05-23 2018-07-10 Knowles Electronics, Llc Microphone and corresponding digital interface
US9111548B2 (en) * 2013-05-23 2015-08-18 Knowles Electronics, Llc Synchronization of buffered data in multiple microphones
US20150350772A1 (en) * 2014-06-02 2015-12-03 Invensense, Inc. Smart sensor for always-on operation
US9549273B2 (en) * 2014-08-28 2017-01-17 Qualcomm Incorporated Selective enabling of a component by a microphone circuit
GB201509483D0 (en) * 2014-12-23 2015-07-15 Cirrus Logic Internat Uk Ltd Feature extraction

Patent Citations (7)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US5612974A (en) * 1994-11-01 1997-03-18 Motorola Inc. Convolutional encoder for use on an integrated circuit that performs multiple communication tasks
EP1154408A2 (en) * 2000-05-10 2001-11-14 Kabushiki Kaisha Toshiba Multimode speech coding and noise reduction
US20060034472A1 (en) * 2004-08-11 2006-02-16 Seyfollah Bazarjani Integrated audio codec with silicon audio transducer
WO2010030889A1 (en) * 2008-09-11 2010-03-18 Personics Holdings Inc. Method and system for sound monitoring over a network
US20130197920A1 (en) 2011-12-14 2013-08-01 Wolfson Microelectronics Plc Data transfer
US20140257813A1 (en) * 2013-03-08 2014-09-11 Analog Devices A/S Microphone circuit assembly and system with speech recognition
WO2014189931A1 (en) * 2013-05-23 2014-11-27 Knowles Electronics, Llc Vad detection microphone and method of operating the same

Non-Patent Citations (1)

* Cited by examiner, † Cited by third party
Title
JASON LASKA ET AL: "Random Sampling for Analog-to-Information Conversion of Wideband Signals", DESIGN, APPLICATIONS, INTEGRATION AND SOFTWARE, 2006 IEEE DALLAS/CAS W ORKSHOP ON, IEEE, PI, 1 October 2006 (2006-10-01), pages 119 - 122, XP031052636, ISBN: 978-1-4244-0669-2 *

Also Published As

Publication number Publication date
GB2535002A (en) 2016-08-10
US20180005636A1 (en) 2018-01-04
CN107251573B (en) 2020-09-25
CN107251573A (en) 2017-10-13
GB2551916A (en) 2018-01-03
GB2551916B (en) 2021-07-07
GB201522666D0 (en) 2016-02-03
US20190259400A1 (en) 2019-08-22
GB201509483D0 (en) 2015-07-15
US10297258B2 (en) 2019-05-21
CN111933158B (en) 2024-09-10
CN111933158A (en) 2020-11-13
GB201711576D0 (en) 2017-08-30

Similar Documents

Publication Publication Date Title
US10297258B2 (en) Microphone unit comprising integrated speech analysis
US10824391B2 (en) Audio user interface apparatus and method
CN110244833B (en) Microphone assembly
US10381021B2 (en) Robust feature extraction using differential zero-crossing counts
US9111548B2 (en) Synchronization of buffered data in multiple microphones
US9412373B2 (en) Adaptive environmental context sample and update for comparing speech recognition
US9775113B2 (en) Voice wakeup detecting device with digital microphone and associated method
US9460720B2 (en) Powering-up AFE and microcontroller after comparing analog and truncated sounds
US9785706B2 (en) Acoustic sound signature detection based on sparse features
US9177546B2 (en) Cloud based adaptive learning for distributed sensors
CN102027536B (en) Adaptively filtering a microphone signal responsive to vibration sensed in a user's face while speaking
US9542933B2 (en) Microphone circuit assembly and system with speech recognition
US20170154620A1 (en) Microphone assembly comprising a phoneme recognizer
CN109346075A (en) Method and system for recognizing user's voice through human body vibration to control electronic equipment
CN104216677A (en) Low-power voice gate for device wake-up
CN102165699A (en) Method and apparatus for signal processing using transform-domain log-companding
CN106104686B (en) Methods in Microphone, Microphone-Component, Microphone-Device
CN103295571A (en) Control of audio commands using temporal and/or spectral compression
JP5027127B2 (en) Improvement of speech intelligibility of mobile communication devices by controlling the operation of vibrator according to background noise
US9978394B1 (en) Noise suppressor
WO2025096392A1 (en) Whispered and other low signal-to-noise voice recognition systems and methods

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 15816854

Country of ref document: EP

Kind code of ref document: A1

WWE Wipo information: entry into national phase

Ref document number: 15538619

Country of ref document: US

NENP Non-entry into the national phase

Ref country code: DE

ENP Entry into the national phase

Ref document number: 201711576

Country of ref document: GB

Kind code of ref document: A

Free format text: PCT FILING DATE = 20151222

122 Ep: pct application non-entry in european phase

Ref document number: 15816854

Country of ref document: EP

Kind code of ref document: A1