WO2021184549A1 - 单耳耳机、智能电子设备、方法和计算机可读介质 - Google Patents

单耳耳机、智能电子设备、方法和计算机可读介质 Download PDF

Info

Publication number
WO2021184549A1
WO2021184549A1 PCT/CN2020/093161 CN2020093161W WO2021184549A1 WO 2021184549 A1 WO2021184549 A1 WO 2021184549A1 CN 2020093161 W CN2020093161 W CN 2020093161W WO 2021184549 A1 WO2021184549 A1 WO 2021184549A1
Authority
WO
WIPO (PCT)
Prior art keywords
microphone
user
ear
ear microphone
mouth
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2020/093161
Other languages
English (en)
French (fr)
Inventor
喻纯
史元春
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Tsinghua University
Original Assignee
Tsinghua University
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Tsinghua University filed Critical Tsinghua University
Publication of WO2021184549A1 publication Critical patent/WO2021184549A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04RLOUDSPEAKERS, MICROPHONES, GRAMOPHONE PICK-UPS OR LIKE ACOUSTIC ELECTROMECHANICAL TRANSDUCERS; ELECTRIC HEARING AIDS; PUBLIC ADDRESS SYSTEMS
    • H04R1/00Details of transducers, loudspeakers or microphones
    • H04R1/10Earpieces; Attachments therefor ; Earphones; Monophonic headphones
    • H04R1/1016Earpieces of the intra-aural type
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L25/00Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
    • G10L25/48Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use
    • G10L25/51Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use for comparison or discrimination
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04RLOUDSPEAKERS, MICROPHONES, GRAMOPHONE PICK-UPS OR LIKE ACOUSTIC ELECTROMECHANICAL TRANSDUCERS; ELECTRIC HEARING AIDS; PUBLIC ADDRESS SYSTEMS
    • H04R1/00Details of transducers, loudspeakers or microphones
    • H04R1/08Mouthpieces; Microphones; Attachments therefor
    • H04R1/083Special constructions of mouthpieces
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04RLOUDSPEAKERS, MICROPHONES, GRAMOPHONE PICK-UPS OR LIKE ACOUSTIC ELECTROMECHANICAL TRANSDUCERS; ELECTRIC HEARING AIDS; PUBLIC ADDRESS SYSTEMS
    • H04R1/00Details of transducers, loudspeakers or microphones
    • H04R1/10Earpieces; Attachments therefor ; Earphones; Monophonic headphones
    • H04R1/1041Mechanical or electronic switches, or control elements
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04RLOUDSPEAKERS, MICROPHONES, GRAMOPHONE PICK-UPS OR LIKE ACOUSTIC ELECTROMECHANICAL TRANSDUCERS; ELECTRIC HEARING AIDS; PUBLIC ADDRESS SYSTEMS
    • H04R1/00Details of transducers, loudspeakers or microphones
    • H04R1/10Earpieces; Attachments therefor ; Earphones; Monophonic headphones
    • H04R1/1091Details not provided for in groups H04R1/1008 - H04R1/1083
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04RLOUDSPEAKERS, MICROPHONES, GRAMOPHONE PICK-UPS OR LIKE ACOUSTIC ELECTROMECHANICAL TRANSDUCERS; ELECTRIC HEARING AIDS; PUBLIC ADDRESS SYSTEMS
    • H04R2201/00Details of transducers, loudspeakers or microphones covered by H04R1/00 but not provided for in any of its subgroups
    • H04R2201/10Details of earpieces, attachments therefor, earphones or monophonic headphones covered by H04R1/10 but not provided for in any of its subgroups

Definitions

  • the present invention generally relates to the field of voice input, and more specifically, to smart electronic devices and voice input triggering methods.
  • voice input After pressing (or holding down) a certain (or some) physical buttons of the mobile device, voice input is activated.
  • the device needs to have a screen; the trigger element occupies the screen content; the limitation of the software UI may lead to a cumbersome trigger method; it is easy to trigger by mistake.
  • the device activates voice input after detecting the corresponding wake-up word.
  • a specific word such as product nickname
  • a single-ear headset having an in-ear microphone and an out-of-ear microphone, and a circuit board with a memory and a processor, and computer executable instructions are stored on the memory.
  • the execution instruction When the execution instruction is executed by the processor, it can perform the following operations: receive the signals collected by the in-ear microphone and the out-of-ear microphone; analyze the signals collected by the in-ear microphone and the out-of-ear microphone, and identify whether the user is making a mouth-covering gesture .
  • the headset is also equipped with a speech detection module for detecting the speech of the user wearing the headset.
  • a speech detection module for detecting the speech of the user wearing the headset.
  • the in-ear microphone and the out-of-ear microphone on the headset are in a closed state, and the speech detection module detects whether the user wearing the headset is speaking, and after recognizing that the user starts to speak, turns on the in-ear microphone and the out-of-ear microphone on the headset , To collect and identify sound signals.
  • the “analyzing the signals collected by the in-ear microphone and the out-of-ear microphone to identify whether the user is making a mouth-covering gesture” includes: responding to the two channels of sounds collected from the in-ear microphone and the out-of-ear microphone The signal is enhanced by human voice signal, the energy amplitudes of the two enhanced signals are calculated separately, the energy amplitude ratio of the two signals is calculated, and the user's voice signal collected by the microphone outside the ear is recognized as being transmitted from the user’s mouth through the air to the outside of the ear Whether the path between the microphones is blocked, and based on this, it is judged whether the user is making a mouth-covering gesture to speak.
  • the microphone outside the ear is an air conduction microphone.
  • the microphone in the ear is an air conduction microphone or a bone conduction microphone.
  • the analyzing the sound signals collected by the in-ear microphone and the out-of-ear microphone, and identifying whether the user is making a mouth-covering gesture includes: calculating the energy amplitude of the user's sound signal received by the in-ear and out-of-ear microphones on the headset Value ratio: When the ratio of the energy amplitude of the user's voice signal received by the in-ear microphone and the out-of-ear microphone exceeds the preset threshold, it is determined that the user is making a mouth-covering gesture.
  • the earphone is operable to wirelessly connect with the smart electronic device, wherein when the earphone recognizes that the user is making a mouth-covering gesture, it transmits a signal indicating the recognition result to the smart electronic device for control
  • the program execution on the smart electronic device includes triggering the corresponding control instruction.
  • the method further includes processing the signals of the in-ear microphone and the out-of-ear microphone to detect whether the user removes the mouth-covering gesture; in response to detecting the user removing the mouth-covering gesture, sending a signal to the smart electronic device to end the interaction process.
  • an electronic device characterized in that it is operable to wirelessly connect with a single earphone below, or is integrated with the single earphone, the single earphone has two microphones, an in-ear microphone and An out-of-ear microphone.
  • the electronic device has a memory and a central processing unit.
  • the memory stores computer-executable instructions. When the computer-executable instructions are executed by the central processing unit, the following operations can be performed: receiving the sound collected by the in-ear microphone and the out-of-ear microphone Signal, analyze the sound signals collected by the microphone in the ear and the microphone outside the ear, and identify whether the user is making a mouth-covering gesture.
  • the electronic device also has a speech detection module for detecting the speech of a user wearing a headset, where before analyzing the sound signals collected by the in-ear microphone and the out-of-ear microphone, and identifying whether the user is making a mouth-covering gesture, before speaking, The in-ear microphone and the out-of-ear microphone on the headset are in a closed state, and the speech detection module detects whether the user wearing the headset is speaking, and after recognizing that the user starts to speak, turns on the in-ear microphone and the out-of-ear microphone on the headset , To collect and identify sound signals.
  • the "analyzing the signals collected by the in-ear microphone and the out-of-ear microphone to identify whether the user is making mouth-covering gestures" includes: making vocal signals on the two channels of sound signals collected from the in-ear microphone and the out-of-ear microphone Enhancement: Calculate the energy amplitudes of the two enhanced signals, calculate the ratio of the energy amplitudes of the two signals, and identify the user's voice signal collected by the external ear microphone between the user’s mouth through the air and the external microphone Whether the path is blocked, and based on this, judge whether the user is making a mouth-covering gesture and vocalize.
  • the microphone outside the ear is an air conduction microphone.
  • the microphone in the ear is an air conduction microphone or a bone conduction microphone.
  • the analyzing the sound signals collected by the in-ear microphone and the out-of-ear microphone, and identifying whether the user is vocalizing in a mouth-covering gesture includes: calculating the energy amplitude of the user's sound signal received by the in-ear microphone on the earphone and the external ear Value ratio: When the ratio of the energy amplitude of the user's voice signal received by the in-ear microphone and the out-of-ear microphone exceeds the preset threshold, it is determined that the user is making a mouth-covering gesture.
  • the operations that can be performed when the computer-executable instructions are executed by the central processing unit further include: in response to recognizing that the user is making a mouth-covering gesture, using a signal indicating the recognition result as a user interactive input control instruction , Control the program execution on the intelligent electronic device, including triggering the corresponding control instruction.
  • the executed control instruction is to trigger other input modes other than the mouth-covering gesture, that is, to process information input by other input modes.
  • the other input methods include one of voice input, non-mouth covering gesture input, line of sight input, blinking input, head movement input, or a combination thereof.
  • the executed control instruction further includes: processing the signal to detect whether the user removes the mouth-covering gesture; in response to detecting that the user removes the mouth-covering gesture, the smart electronic device ends the interaction process.
  • the executed control instruction further includes: providing feedback including any one of visual and auditory senses, prompting the user that the smart electronic device has triggered other input methods.
  • the executed control instruction further includes: the smart electronic device processes the voice input made by the user while keeping the mouth-covering gesture.
  • the smart electronic device is a smart wearable device among a mobile phone, a watch, a smart ring, and a wrist watch.
  • the smart electronic device is a head-mounted smart display device equipped with the in-ear microphone and the out-of-ear microphone.
  • a voice interactive wake-up method for a smart electronic device includes: receiving data collected by the in-ear microphone and the out-of-ear microphone. Sound signal; Analyze the sound signals collected by the microphone in the ear and the microphone outside the ear to identify whether the user is making a mouth-covering gesture; in response to recognizing that the user is making a mouth-covering gesture, the smart device triggers voice input processing , Analyze and make corresponding content output; after responding to the user's mouth-covering gesture, in the case of the user interacting with the smart device, process the sound signals collected by the in-ear microphone and the out-of-ear microphone to determine that the user removes the mouth-covering gesture; After it is determined that the user removes the mouth-covering gesture, the interaction process ends.
  • the content output form includes one or a combination of voice and image.
  • a computer-readable medium having computer-executable instructions stored thereon, and the computer-executable instructions can execute the voice interactive wake-up method as described above when the computer-executable instructions are executed by a computer.
  • the present invention uses two microphones inside the same earphone—in-ear microphone and out-of-ear microphone—to identify whether the user is making a mouth-covering gesture, and then trigger the voice input, which can accurately recognize the cover
  • the voice input under the mouth gesture can trigger the voice input very conveniently and accurately.
  • the use efficiency is higher. It can be used with one hand. No need to switch between different user interfaces/applications, no need to hold down a button, just lift your hand to your mouth to use it.
  • the radio quality is high.
  • the voice input signal collected by the in-ear microphone and the out-of-ear microphone of the headset is clear, and is less affected by environmental sounds.
  • High privacy and sociality Determine whether to trigger the voice input application based on the inherent characteristics of the sound captured by the in-ear microphone and the out-of-ear microphone of the same headset configuration. There is no need for traditional physical button triggering, interface element triggering, wake-up word detection, and the interaction is more natural.
  • Fig. 1 schematically shows the following scenario.
  • the user wears a single-ear headset, and the single-ear headset is equipped with both an in-ear microphone and an out-of-ear microphone, and the user makes a mouth-covering gesture and speaks in a low voice at the same time.
  • This situation may occur, for example, in a conference room, when the user does not want to influence others but still needs to speak in a low or silent voice.
  • Figure 2 schematically shows the change in the energy of the mouth-covering action when the user’s sound is propagated in the air, so that the sound entering the microphone outside the earphone becomes smaller; in contrast, the microphone inside the earphone receives the sound through the ear canal and The sound transmitted by the head is not affected by the mouth covering movement.
  • Figure 3 schematically shows the different sources of the user's speaking voice received by the in-ear microphone, where the user's speaking voice received by the in-ear microphone is from the throat or mouth, sounds from the ear canal, or from the head Sound conducted by muscles and bones.
  • Fig. 4 shows an overall flow chart of using a single ear headset equipped with an in-ear microphone and an out-of-ear microphone to recognize whether a user is making a mouth-covering gesture and vocalize according to an embodiment of the present invention.
  • the inventive concept of the present invention will be introduced first.
  • the main change is the path of the user’s voice reaching the microphone outside the ear, which has relatively little influence on the propagation path of the microphone in the ear.
  • the conduction path of the out-of-ear microphone receiving the user's voice is different. Therefore, the ratio of the energy amplitude of the user's voice signal received by the in-ear microphone and the out-of-ear microphone on the earphone can be used to determine whether the user is speaking while covering the mouth.
  • the voice input can be triggered at the initial moment when it is recognized that the user is making a voice while covering the mouth.
  • Fig. 1 schematically shows the following scenario.
  • the user wears a single-ear headset, and the single-ear headset is equipped with both an in-ear microphone and an out-of-ear microphone, and the user makes a mouth-covering gesture and speaks in a low voice at the same time.
  • This situation may occur, for example, in a conference room, when the user does not want to influence others but still needs to speak in a low voice.
  • the direction of the microphone in the ear is toward the ear to collect the sound in the ear; the direction of the microphone outside the ear is outward to collect the sound in the environment, including conduction through the external air Voice of users.
  • Figure 2 schematically shows the change in the energy of the mouth-covering action when the user’s sound is propagated in the air, so that the sound entering the microphone outside the earphone becomes smaller; in contrast, the microphone inside the earphone receives the sound through the ear canal and The sound transmitted by the head is not affected by the mouth covering movement.
  • Figure 3 schematically shows the different sources of the user's speaking voice received by the in-ear microphone, where the user's speaking voice received by the in-ear microphone is emitted by the throat or mouth, and the sound transmitted through the ear canal or through the head Sound conducted by muscles and bones.
  • Fig. 4 shows an overall flow chart of using a single ear headset equipped with an in-ear microphone and an out-of-ear microphone to recognize whether a user is making a mouth-covering gesture and vocalize according to an embodiment of the present invention.
  • the method is preferably executed on a monaural headset.
  • the monaural headset has a circuit board, and the circuit board has a memory and a processor.
  • the memory stores computer-executable instructions. When the computer-executable instructions are executed by the processor, The method can be performed.
  • the method can also be executed on a smart electronic device that cooperates with a monaural headset, such as a smart phone.
  • a monaural headset such as a smart phone.
  • the in-ear microphone and the out-of-ear microphone of the monaural headset need to be connected.
  • the collected two signals are sent to the smart electronic device.
  • step S401 the signals collected by the in-ear microphone and the out-of-ear microphone are received.
  • step S402 the signals collected by the in-ear microphone and the out-of-ear microphone are analyzed to identify whether the user is making a mouth-covering gesture or not.
  • the microphone outside the ear may be an air conduction microphone
  • the microphone inside the ear may be an air conduction microphone or a bone conduction microphone.
  • analyzing the sound signals collected by the in-ear microphone and the out-of-ear microphone, and identifying whether the user is making a mouth-covering gesture includes: calculating the energy amplitude of the user's sound signal received by the in-ear microphone and the out-of-ear microphone on the headset Value ratio: When the ratio of the energy amplitude of the user's voice signal received by the in-ear microphone and the out-of-ear microphone exceeds the preset threshold, it is determined that the user is making a mouth-covering gesture.
  • analyzing the signals collected by the in-ear microphone and the out-of-ear microphone, and identifying whether the user is making a mouth-covering gesture may include: vocalizing the two channels of sound signals collected from the in-ear microphone and the out-of-ear microphone Signal enhancement; respectively calculate the energy amplitude of the two enhanced signals, calculate the ratio of the energy amplitude of the two signals, and identify that the user's voice signal collected by the external ear microphone is transmitted from the user’s mouth through the air to the external microphone Whether the path of is blocked, and based on this, judge whether the user is making a mouth-covering gesture to speak.
  • the headset is also equipped with a speech detection module for detecting the speech of the user wearing the headset, in which the sound signals collected by the in-ear microphone and the out-of-ear microphone are analyzed to identify whether the user is making a mouth-covering gesture before making a voice action.
  • the in-ear microphone and the out-of-ear microphone on the headset are in a closed state, and the speech detection module detects whether the user wearing the headset is speaking, and after recognizing that the user starts to speak, turns on the in-ear microphone and the out-of-ear microphone on the headset
  • the microphone collects and recognizes sound signals.
  • the earphone is operable to connect wirelessly with the smart electronic device, wherein when the earphone recognizes that the user is making a mouth-covering gesture, it transmits a signal indicating the recognition result to the smart electronic device for Control the execution of the program on the smart electronic device, including triggering the corresponding control instruction.
  • the operation performed by the headset further includes processing the signals of the in-ear microphone and the out-of-the-ear microphone to detect whether the user removes the mouth-covering gesture; in response to detecting that the user removes the mouth-covering gesture, sending a signal to the smart electronic device to end the interaction process.
  • an electronic device that is operable to connect wirelessly with a single earphone below, or is integrated with the single earphone, the single earphone has two microphones, an in-ear microphone and an out-of-ear microphone
  • the electronic device has a memory and a central processing unit.
  • the memory stores computer-executable instructions. When the computer-executable instructions are executed by the central processing unit, the following operations can be performed: receiving the sound signals collected by the in-ear microphone and the out-of-ear microphone, and analyzing The sound signals collected by the in-ear microphone and the out-of-ear microphone identify whether the user is making mouth-covering gestures.
  • the electronic device may also have a speech detection module for detecting the speech of a user wearing a headset.
  • a speech detection module for detecting the speech of a user wearing a headset. Before analyzing the sound signals collected by the in-ear microphone and the out-of-ear microphone, and identifying whether the user is making a mouth-covering gesture, the headset The in-ear microphone and the out-of-ear microphone on the headset are in a closed state, the speech detection module detects whether the user wearing the headset is speaking, and after recognizing that the user starts to speak, turns on the in-ear microphone and the out-of-ear microphone on the headset to perform sound Signal acquisition and identification.
  • the “analyzing the signals collected by the in-ear microphone and the out-of-ear microphone to identify whether the user is making a mouth-covering gesture” includes: performing human voice on the two channels of sound signals collected from the in-ear microphone and the out-of-ear microphone Signal enhancement; respectively calculate the energy amplitude of the two enhanced signals, calculate the ratio of the energy amplitude of the two signals, and identify that the user's voice signal collected by the external ear microphone is transmitted from the user’s mouth through the air to the external microphone Whether the path of is blocked, and based on this, judge whether the user is making a mouth-covering gesture to speak.
  • the microphone outside the ear is an air conduction microphone
  • the microphone inside the ear is an air conduction microphone or a bone conduction microphone.
  • the analyzing the sound signals collected by the in-ear microphone and the out-of-ear microphone and identifying whether the user is making a mouth-covering gesture includes:
  • the operations that can be performed when the computer-executable instructions are executed by the central processing unit further include: in response to recognizing that the user is making a mouth-covering gesture, using a signal indicating the recognition result as an instruction for user interactive input control, Control the program execution on the smart electronic device, including triggering corresponding control instructions or triggering other input methods.
  • the executed control instruction is to trigger other input methods other than the mouth-covering gesture, that is, to process information input by other input methods.
  • the other input methods include one of voice input, non-mouth covering gesture input, line of sight input, blinking input, head movement input, or a combination thereof.
  • the smart electronic device also processes the in-ear microphone signal and the out-of-ear microphone signal to detect whether the user removes the mouth-covering gesture; in response to detecting that the user removes the mouth-covering gesture, the smart electronic device ends the interaction process.
  • any feedback including visual and auditory senses, prompting the user that the smart electronic device has triggered other input methods.
  • the smart electronic device is, for example, a smart wearable device among a mobile phone, a watch, a smart ring, and a wrist watch.
  • the smart electronic device is a head-mounted smart display device equipped with the in-ear microphone and the out-of-ear microphone.
  • a voice interactive wake-up method for a smart electronic device includes: receiving sound signals collected by the in-ear microphone and the out-of-ear microphone, and analyzing The sound signals collected by the in-ear microphone and the out-of-ear microphone identify whether the user is making a mouth-covering gesture; in response to determining that the user keeps the mouth-covering gesture with his hand, according to the type and intelligence of the mouth-covering gesture.
  • the interactive content currently applied by the device analyzes the user’s interactive intention; according to the parsed interactive intention, the smart device will receive the user’s input, analyze and output the corresponding content; after responding to the user’s mouth-covering gesture,
  • the sound signals collected by the in-ear microphone and the out-of-ear microphone are processed to determine the user to remove the mouth-covering gesture; in response to determining that the user removes the mouth-covering gesture, the interaction process is ended.
  • the content output form includes one or a combination of voice and image.
  • a computer-readable medium having computer-executable instructions stored thereon, and the computer-executable instructions can execute the above-mentioned voice interactive wake-up method when the computer-executable instructions are executed by a computer.
  • the present invention uses two microphones inside the same earphone—in-ear microphone and out-of-ear microphone—to identify whether the user is making a mouth-covering gesture, and then trigger the voice input, which can accurately recognize the cover
  • the voice input under the mouth gesture can trigger the voice input very conveniently and accurately.
  • the use efficiency is higher. It can be used with one hand. No need to switch between different user interfaces/applications, no need to hold down a button, just lift your hand to your mouth to use it.
  • the radio quality is high.
  • the voice input signal collected by the in-ear microphone and the out-of-ear microphone of the headset is clear, and is less affected by environmental sounds.
  • High privacy and sociality Determine whether to trigger the voice input application based on the inherent characteristics of the sound captured by the in-ear microphone and the out-of-ear microphone of the same headset configuration. There is no need for traditional physical button triggering, interface element triggering, wake-up word detection, and the interaction is more natural.

Landscapes

  • Engineering & Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Acoustics & Sound (AREA)
  • Signal Processing (AREA)
  • Computational Linguistics (AREA)
  • Health & Medical Sciences (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Human Computer Interaction (AREA)
  • Multimedia (AREA)
  • User Interface Of Digital Computer (AREA)
  • Telephone Function (AREA)

Abstract

一种能够识别用户捂嘴手势下发声的单耳耳机、智能电子便携设备和语音交互唤醒方法,该单耳耳机具有耳内麦克风和耳外麦克风,以及具有一块电路板,电路板上具有存储器和处理器,存储器上存储有计算机可执行指令,计算机可执行指令被处理器执行时能够执行如下操作:接收所述耳内麦克风和耳外麦克风采集的信号;分析耳内麦克风和耳外麦克风采集的信号,识别用户是否在做捂嘴手势的状态下发声。所述识别结果可以触发语音输入。本发明能够准确地识别出捂嘴手势下的语音输入;另外在由耳机自身电路板对信号进行接受和处理的情况下,不需要额外解决数据传输和信号的时间同步问题,节省电能,且保证高识别精度;使用效率更高、收音质量高、隐私性与社会性高。

Description

单耳耳机、智能电子设备、方法和计算机可读介质
本申请要求于2020年3月19日提交中国专利局、申请号为202010198596.6、名称为“单耳耳机、智能电子设备、方法和计算机可读介质”的中国专利申请的优先权,其全部内容通过引用结合在本申请中。
技术领域
本发明总的来说涉及语音输入领域,且更为具体地,涉及智能电子设备、语音输入触发方法。
背景技术
随着计算机技术的发展,语音识别算法日益成熟,语音输入因其在交互方式上的高自然性与有效性而正变得越来越重要。用户可以通过语音与移动设备(手机、手表等)进行交互,完成指令输入、信息查询、语音聊天等多种任务。
而在何时触发语音输入这一点上,现有的解决方案都有一些缺陷:
1.物理按键触发
按下(或按住)移动设备的某个(或某些)物理按键后,激活语音输入。
该方案的缺点是:需要物理按键;容易误触发;需要用户按键。
2.界面元素触发
点击(或按住)移动设备的屏幕上的界面元素(如图标),激活语音输入。
该方案的缺点是:需要设备具备屏幕;触发元素占用屏幕内容;受限于软件UI限制,可能导致触发方式繁琐;容易误触发。
3.唤醒词(语音)检测
以某个特定词语(如产品昵称)为唤醒词,设备检测到对应的唤醒词后激活语音输入。
该方案的缺点是:隐私性和社会性较差;交互效率较低。
发明内容
针对上述问题,本申请人先前提交了几份专利申请,在如下四个方面上提出了多项新的技术方案:1、基于人类说话时风噪声特征的语音输入触发,具体地,通过识别人说话时候的语音和风噪声音来直接启动语音输入并将接收的声音信号作为语音输入处理;2、基于多个麦克风接收的声音信号的差别的语音输入触发;3、基于低声说话方式识别的语音输入触发;4、基于麦克风的声音信号的距离判断的语音输入触发,相关专利申请公开案号为CN110262767A、CN110223711A、CN110428806A、CN110111776A、CN110097875A、CN110164440A,本文将这几篇专利文献全文并入,作为本公开的内容。
根据本发明的一个方面,提供了一种单耳耳机,具有耳内麦克风和耳外麦克风,以及具有一块电路板,电路板上具有存储器和处理器,存储器上存储有计算机可执行指令,计算机可执行指令被处理器执行时能够执行如下操作:接收所述耳内麦克风和耳外麦克风采集的信号;分析耳内麦克风和耳外麦克风采集的信号,识别用户是否在做捂嘴手势的状态下发声。
可选地,耳机还具备用于检测佩戴耳机的用户说话的说话检测模块,其中在分析耳内麦克风和耳外麦克风采集的声音信号,识别用户是否在做捂嘴手势的状态下发声动作之前,所述耳机上的耳内麦克风和耳外麦克风处于关闭状态,所述说话检测模块检测佩戴耳机的用户是否在说话,以及在识别到用户开始说话之后,打开耳机上的耳内麦克风和耳外麦克风,进行声音信号采集并识别。
可选地,所述“分析耳内麦克风和耳外麦克风采集的信号,识别用户是否在做捂嘴手势的状态下发声”,包括:对从耳内麦克风和耳外麦克风采集到的两路声音信号做人声信号增强,分别计算两路增强后信号的能量幅值,计算所述两路信号的能量幅值比值,识别耳外麦克风采集的用户声音信号在从用户口腔发出通过空气传到耳外麦克风之间的路径上有没有被遮挡,并基于此判断用户是否在做捂嘴手势的状态下发声。
可选地,所述耳外的麦克风是空气传导麦克风。
可选地,所述耳内的麦克风为空气传导麦克风或骨传导麦克风。
可选地,所述分析耳内麦克风和耳外麦克风采集的声音信号,识别用户是否在做捂嘴手势的状态下发声包括:计算耳机上耳内和耳外麦克风接收到的用户声音信号能量幅值比值;在耳内麦克风和耳外麦克风接收到的用户声音信号能量幅值比值超过预设阈值时,判断用户是在做捂嘴手势的状态下发声。
可选地,所述耳机可操作来与智能电子设备无线连接,其中当耳机识别出用户是在做捂嘴手势的状态下发声时,将指示识别结果的信号传递给智能电子设备,用于控制智能电子设备上的程序执行,包括触发相应的控制指令。
可选地,还包括处理所述耳内麦克风和耳外麦克风信号以检测用户是否去除捂嘴手势;响应于检测到用户去除捂嘴手势,发送信号给智能电子设备结束所述交互过程。
根据本发明的另一方面,提供了一种电子设备,特征在于:可操作来与下面的单个耳机无线连接,或者集成有所述单个耳机,所述单个耳机具有两个麦克风,耳内麦克风和耳外麦克风,电子设备具有存储器和中央处理器,存储器上存储有计算机可执行指令,计算机可执行指令被中央处理器执行时能够执行如下操作:接收所述耳内麦克风和耳外麦克风采集的声音信号,分析耳内麦克风和耳外麦克风采集的声音信号,识别用户是否在做捂嘴手势的状态下发声。
可选地,电子设备还具备用于检测佩戴耳机的用户说话的说话检测模块,其中在分析耳内麦克风和耳外麦克风采集的声音信号,识别用户是否在做捂嘴手势的状态下发声之前,所述耳机上的耳内麦克风和耳外麦克风处于关闭状态,所述说话检测模块检测佩戴耳机的用户是否在说话,以及在识别到用户开始说话之后,打开耳机上的耳内麦克风和耳外麦克风,进行声音信号采集并识别。
可选地,所述“分析耳内麦克风和耳外麦克风采集的信号,识别用户是否在做捂嘴手势”,包括:对从耳内麦克风和耳外麦克风采集到的两路声 音信号做人声信号增强;分别计算两路增强后信号的能量幅值,计算所述两路信号的能量幅值比值,识别耳外麦克风采集的用户声音信号在从用户口腔发出通过空气传到耳外麦克风之间的路径上有没有被遮挡,并基于此判断用户是否在做捂嘴手势的状态下发声。
可选地,所述耳外的麦克风是空气传导麦克风。
可选地,所述耳内的麦克风为空气传导麦克风或骨传导麦克风。
可选地,所述分析耳内麦克风和耳外麦克风采集的声音信号,识别用户是否在做捂嘴手势的状态下发声包括:计算耳机上耳内麦克风和耳外接收到的用户声音信号能量幅值比值;在耳内麦克风和耳外麦克风接收到的用户声音信号能量幅值比值超过预设阈值时,判断用户是在做捂嘴手势的状态下发声。
可选地,计算机可执行指令被中央处理器执行时能够执行的操作还包括:响应于识别出用户是在做出捂嘴手势的状态下,将指示识别结果的信号作为用户交互输入控制的指示,控制智能电子设备上的程序执行,包括触发相应的控制指令。
可选地,执行的控制指令为触发除捂嘴手势外的其它输入方式,即处理其它输入方式输入的信息。
可选地,所述其他输入方式包括语音输入、非捂嘴手势输入、视线输入、眨眼输入、头动输入之一或者其组合。
可选地,执行的控制指令还包括:处理所述信号以检测用户是否去除捂嘴手势;响应于检测到用户去除捂嘴手势,智能电子设备结束所述交互过程。
可选地,执行的控制指令还包括:提供包括视觉、听觉任一项反馈,提示用户智能电子设备已经触发其他输入方式。
可选地,执行的控制指令还包括:智能电子设备对用户在保持捂嘴手势同时进行的语音输入进行处理。
可选地,所述智能电子设备为手机、手表、智能戒指、腕表中的一种智能穿戴设备。
可选地,所述智能电子设备为头戴式智能显示设备,装备有所述耳内 麦克风和耳外麦克风。
根据本发明的另一方面,提供了一种如上所述的智能电子设备的语音交互唤醒方法,所述智能电子设备执行的语音交互唤醒方法包括:接收所述耳内麦克风和耳外麦克风采集的声音信号;分析耳内麦克风和耳外麦克风采集的声音信号,识别用户是否在做捂嘴手势的状态下发声;响应于识别到用户在做捂嘴手势的状态下发声,智能设备触发语音输入处理,分析及做出相应的内容输出;响应用户捂嘴手势后,在用户与智能设备交互情况下,处理所述耳内麦克风和耳外麦克风采集的声音信号,以确定用户去除捂嘴手势;响应于确定用户去除捂嘴手势,结束所述交互过程。
可选地,所述内容输出形式包括语音、图像中一种或其组合。
根据本发明另一方面,提供了一种计算机可读介质,其上存储有计算机可执行指令,计算机可执行指令被计算机执行时能够执行如上所述的语音交互唤醒方法。
本发明的技术方案具有至少下述优势中的一个或多个:
1.本发明利用同一个耳机内部的两个麦克风——耳内麦克风和耳外麦克风——来识别用户是否在做捂嘴手势的状态下发声,进而触发语音输入,这样能够准确地识别出捂嘴手势下的语音输入,能够非常便利准确地触发语音输入。
2.在由耳机自身电路板对耳机上的耳内麦克风和耳外麦克风的两路信号进行接受和处理的情况下,不需要额外解决数据传输和信号的时间同步问题,会节省电能,且保证高识别精度,
3.使用效率更高。单手即可使用。无需在不同的用户界面/应用之间切换,也不需按住某个按键,直接抬起手到嘴边就能使用。
4.收音质量高。耳机的耳内麦克风和耳外麦克风收取的语音输入信号清晰,受环境音的影响较小。
5.高隐私性与社会性。基于同一耳机配置的耳内麦克风和耳外麦克风捕捉的声音内在特征,来确定是否触发语音输入应用,其中无需传统的物理按键触发、界面元素触发、唤醒词检测,交互更加自然。
6.做出捂嘴手势,用户进行语音输入对他人的干扰较小,同时具有较好的隐私保护,降低用户语音输入时的心理负担。
附图说明
从下面结合附图对本发明实施例的详细描述中,本发明的上述和/或其它目的、特征和优势将变得更加清楚并更容易理解。其中:
图1示意性地示出了如下情境,用户佩戴单耳耳机,单耳耳机上同时配置有耳内麦克风和耳外麦克风,以及用户做出捂嘴手势并同时低声说话。这种情况可能发生在例如一种会议室中,用户不想影响他人但仍需要低声或无声说话的时候。
图2示意了示出了捂嘴动作对于用户发出的声音在空气中传播时能量的改变,让进入到耳机外麦克风的声音变小;相比而言,耳机内部的麦克风接收到通过耳道和头部传播的声音,不受捂嘴动作的影响。
图3示意性地示出了耳内麦克风所接收的用户说话声音的不同来源,其中耳内麦克风所接收到的用户说话声音是喉咙或口腔发出、通过耳道传出的声音或者通过头部的肌肉、骨骼传导的声音。
图4示出了根据本发明实施例的利用配备有耳内麦克风和耳外麦克风的单耳耳机来识别用户是否在做捂嘴手势的状态下发声的总体流程图。
具体实施方式
为了使本领域技术人员更好地理解本发明,下面结合附图和具体实施方式对本发明作进一步详细说明。
为便于理解,在详细介绍之前,首先介绍下本发明的发明构思。在用户佩戴单耳耳机的情况下,当用户做捂嘴动作时,主要改变的是用户声音达到耳外麦克风的路径,对耳内麦克风接受人声的传播路径影响相对较小,耳内麦克风和耳外麦克风接受用户说话声音的传导路径不同,因此,可以通过耳机上耳内麦克风和耳外麦克风接收到的用户声音信号能量幅值比值来判断用户是否在做捂嘴动作的状态下发声。进而,可以在识别到用户在做捂嘴动作的状态下发声的初始时刻,触发语音输入。
图1示意性地示出了如下情境,用户佩戴单耳耳机,单耳耳机上同时配置有耳内麦克风和耳外麦克风,以及用户做出捂嘴手势并同时低声说话。这种情况可能发生在例如一种会议室中,用户不想影响他人但仍需要低声说话的时候。如图1所示,用户佩戴该耳机时,耳内麦克风的收音方向朝着耳朵内,收集耳朵内的声音;耳外麦克风的收音方向向外,采集环境中的声音,也包括通过外部空气传导的用户说话声音。
图2示意了示出了捂嘴动作对于用户发出的声音在空气中传播时能量的改变,让进入到耳机外麦克风的声音变小;相比而言,耳机内部的麦克风接收到通过耳道和头部传播的声音,不受捂嘴动作的影响。
图3示意性地示出了耳内麦克风所接收的用户说话声音的不同来源,其中耳内麦克风所接收到的用户说话声音是喉咙或口腔发出,通过耳道传出的声音或者通过头部的肌肉、骨骼传导的声音。
图4示出了根据本发明实施例的利用配备有耳内麦克风和耳外麦克风的单耳耳机来识别用户是否在做捂嘴手势的状态下发声的总体流程图。
所述方法优选是在单耳耳机上执行的,此时单耳耳机具有一块电路板,电路板上具有存储器和处理器,存储器上存储有计算机可执行指令,计算机可执行指令被处理器执行时能够执行所述方法。
不过所述方法也可以在与单耳耳机协作的智能电子设备上执行,例如在智能手机上执行,此时在方法执行之前,需要将所述单耳耳机的所述耳内麦克风和耳外麦克风采集的这两路信号发送到智能电子设备上。
如图4所示,在步骤S401中,接收所述耳内麦克风和耳外麦克风采集的信号。
在步骤S402中,分析耳内麦克风和耳外麦克风采集的信号,识别用户是否在做捂嘴手势的状态下发声。
在一个示例中,耳外的麦克风可以是空气传导麦克风,耳内的麦克风为空气传导麦克风或骨传导麦克风。
在一个示例中,分析耳内麦克风和耳外麦克风采集的声音信号,识别用户是否在做捂嘴手势的状态下发声包括:计算耳机上耳内麦克风和耳外麦克风接收到的用户声音信号能量幅值比值;在耳内麦克风和耳外麦克风 接收到的用户声音信号能量幅值比值超过预设阈值时,判断用户是在做捂嘴手势的状态下发声。
在一个示例中,分析耳内麦克风和耳外麦克风采集的信号,识别用户是否在做捂嘴手势的状态下发声可以包括:对从耳内麦克风和耳外麦克风采集到的两路声音信号做人声信号增强;分别计算两路增强后信号的能量幅值,计算所述两路信号的能量幅值比值,识别耳外麦克风采集的用户声音信号在从用户口腔发出通过空气传到耳外麦克风之间的路径上有没有被遮挡,并基于此判断用户是否在做捂嘴手势的状态下发声。
在一个示例中,耳机还具备用于检测佩戴耳机的用户说话的说话检测模块,其中在分析耳内麦克风和耳外麦克风采集的声音信号,识别用户是否在做捂嘴手势的状态下发声动作之前,所述耳机上的耳内麦克风和耳外麦克风处于关闭状态,所述说话检测模块检测佩戴耳机的用户是否在说话,以及在识别到用户开始说话之后,打开耳机上的耳内麦克风和耳外麦克风,进行声音信号采集并识别。
在一个示例中,所述耳机可操作来与智能电子设备无线连接,其中当耳机识别出用户是在做捂嘴手势的状态下发声时,将指示识别结果的信号传递给智能电子设备,用于控制智能电子设备上的程序执行,包括触发相应的控制指令。
在一个示例中,耳机执行操作还包括处理所述耳内麦克风和耳外麦克风信号以检测用户是否去除捂嘴手势;响应于检测到用户去除捂嘴手势,发送信号给智能电子设备结束所述交互过程。
根据本发明另一实施例,提供了一种电子设备,可操作来与下面的单个耳机无线连接,或者集成有所述单个耳机,所述单个耳机具有两个麦克风,耳内麦克风和耳外麦克风,电子设备具有存储器和中央处理器,存储器上存储有计算机可执行指令,计算机可执行指令被中央处理器执行时能够执行如下操作:接收所述耳内麦克风和耳外麦克风采集的声音信号,分析耳内麦克风和耳外麦克风采集的声音信号,识别用户是否在做捂嘴手势。
电子设备还可以具备用于检测佩戴耳机的用户说话的说话检测模块,其中在分析耳内麦克风和耳外麦克风采集的声音信号,识别用户是否在做 捂嘴手势的状态下发声之前,所述耳机上的耳内麦克风和耳外麦克风处于关闭状态,所述说话检测模块检测佩戴耳机的用户是否在说话,以及在识别到用户开始说话之后,打开耳机上的耳内麦克风和耳外麦克风,进行声音信号采集并识别。
在一个示例中,所述“分析耳内麦克风和耳外麦克风采集的信号,识别用户是否在做捂嘴手势”,包括:对从耳内麦克风和耳外麦克风采集到的两路声音信号做人声信号增强;分别计算两路增强后信号的能量幅值,计算所述两路信号的能量幅值比值,识别耳外麦克风采集的用户声音信号在从用户口腔发出通过空气传到耳外麦克风之间的路径上有没有被遮挡,并基于此判断用户是否在做捂嘴手势的状态下发声。
例如,所述耳外的麦克风是空气传导麦克风,所述耳内的麦克风为空气传导麦克风或骨传导麦克风。
作为示例,所述分析耳内麦克风和耳外麦克风采集的声音信号,识别用户是否在做捂嘴手势的状态下发声包括:
计算耳机上耳内麦克风和耳外麦克风接收到的用户声音信号能量幅值比值;
在耳内麦克风和耳外麦克风接收到的用户声音信号能量幅值比值超过预设阈值时,判断用户是在做捂嘴手势的状态下发声。
作为示例,计算机可执行指令被中央处理器执行时能够执行的操作还包括:响应于识别出用户是在做出捂嘴手势的状态下,将指示识别结果的信号作为用户交互输入控制的指示,控制智能电子设备上的程序执行,包括触发相应的控制指令或者触发其他输入方式。
作为示例,执行的控制指令为触发除捂嘴手势外的其它输入方式,即处理其它输入方式输入的信息。
作为示例,所述其他输入方式包括语音输入、非捂嘴手势输入、视线输入、眨眼输入、头动输入之一或者其组合。
智能电子设备还处理所述耳内麦克风信号和耳外麦克风信号以检测用户是否去除捂嘴手势;响应于检测到用户去除捂嘴手势,智能电子设备结束所述交互过程。
作为示例,提供包括视觉、听觉任一项反馈,提示用户智能电子设备已经触发其他输入方式。
所述智能电子设备例如为手机、手表、智能戒指、腕表中的一种智能穿戴设备。
例如,所述智能电子设备为头戴式智能显示设备,装备有所述耳内麦克风和耳外麦克风。
根据本发明另一实施例,提供了一种智能电子设备的语音交互唤醒方法,所述智能电子设备执行的语音交互唤醒方法包括:接收所述耳内麦克风和耳外麦克风采集的声音信号,分析耳内麦克风和耳外麦克风采集的声音信号,识别用户是否在做捂嘴手势的状态下发声;响应于确定用户将手放在嘴边持续保持捂嘴手势,根据所做捂嘴手势类别、智能设备当前应用的交互内容,对于用户的交互意图进行解析;根据解析得到的交互意图,智能设备将对于用户的输入信息进行接收,分析及做出相应的内容输出;响应用户捂嘴手势后,在用户与智能设备交互情况下,处理所述耳内麦克风和耳外麦克风采集的声音信号,信号以确定用户去除捂嘴手势;响应于确定用户去除捂嘴手势,结束所述交互过程。
作为示例,内容输出形式包括语音、图像中一种或其组合。
根据本发明的另一方面,提供了一种计算机可读介质,其上存储有计算机可执行指令,计算机可执行指令被计算机执行时能够执行上述语音交互唤醒方法。
本发明各个实施例的方案可以提供下述一种或几种优势:
1.本发明利用同一个耳机内部的两个麦克风——耳内麦克风和耳外麦克风——来识别用户是否在做捂嘴手势的状态下发声,进而触发语音输入,这样能够准确地识别出捂嘴手势下的语音输入,能够非常便利准确地触发语音输入。
2.在由耳机自身电路板对耳机上的耳内麦克风和耳外麦克风的两路信号进行接受和处理的情况下,不需要额外解决数据传输和信号的时间同步问题,会节省电能,且保证高识别精度,
3.使用效率更高。单手即可使用。无需在不同的用户界面/应用之间切换,也不需按住某个按键,直接抬起手到嘴边就能使用。
4.收音质量高。耳机的耳内麦克风和耳外麦克风收取的语音输入信号清晰,受环境音的影响较小。
5.高隐私性与社会性。基于同一耳机配置的耳内麦克风和耳外麦克风捕捉的声音内在特征,来确定是否触发语音输入应用,其中无需传统的物理按键触发、界面元素触发、唤醒词检测,交互更加自然。
6.做出捂嘴手势,用户进行语音输入对他人的干扰较小,同时具有较好的隐私保护,降低用户语音输入时的心理负担。
以上已经描述了本发明的各实施例,上述说明是示例性的,并非穷尽性的,并且也不限于所披露的各实施例。在不偏离所说明的各实施例的范围和精神的情况下,对于本技术领域的普通技术人员来说许多修改和变更都是显而易见的。因此,本发明的保护范围应该以权利要求的保护范围为准。

Claims (25)

  1. 一种单耳耳机,其特征在于,具有耳内麦克风和耳外麦克风,以及具有一块电路板,电路板上具有存储器和处理器,存储器上存储有计算机可执行指令,计算机可执行指令被处理器执行时能够执行如下操作:
    接收所述耳内麦克风和耳外麦克风采集的信号;
    分析耳内麦克风和耳外麦克风采集的信号,识别用户是否在做捂嘴手势的状态下发声。
  2. 根据权利要求1的耳机,其特征在于,还具备用于检测佩戴耳机的用户说话的说话检测模块,其中
    在分析耳内麦克风和耳外麦克风采集的声音信号,识别用户是否在做捂嘴手势的状态下发声动作之前,所述耳机上的耳内麦克风和耳外麦克风处于关闭状态,
    所述说话检测模块检测佩戴耳机的用户是否在说话,以及在识别到用户开始说话之后,打开耳机上的耳内麦克风和耳外麦克风,进行声音信号采集并识别。
  3. 根据权利要求1的耳机,其特征在于,所述“分析耳内麦克风和耳外麦克风采集的信号,识别用户是否在做捂嘴手势的状态下发声”,包括:
    对从耳内麦克风和耳外麦克风采集到的两路声音信号做人声信号增强
    分别计算两路增强后信号的能量幅值,计算所述两路信号的能量幅值比值,识别耳外麦克风采集的用户声音信号在从用户口腔发出通过空气传到耳外麦克风之间的路径上有没有被遮挡,并基于此判断用户是否在做捂嘴手势的状态下发声。
  4. 根据权利要求1的耳机,其特征在于,所述耳外的麦克风是空气传导麦克风。
  5. 根据权利要求1的耳机,其特征在于,所述耳内的麦克风为空气传导麦克风或骨传导麦克风。
  6. 根据权利要求1的耳机,其特征在于,所述分析耳内麦克风和耳外麦克风采集的声音信号,识别用户是否在做捂嘴手势的状态下发声包括:
    计算耳机上的耳内麦克风和耳外麦克风接收到的用户声音信号能量幅 值比值;
    在耳内麦克风和耳外麦克风接收到的用户声音信号能量幅值比值超过预设阈值时,判断用户是在做捂嘴手势的状态下发声。
  7. 根据权利要求1的耳机,其特征在于,所述耳机可操作来与智能电子设备无线连接,其中
    当耳机识别出用户是在做捂嘴手势的状态下发声时,将指示识别结果的信号传递给智能电子设备,用于控制智能电子设备上的程序执行,包括触发相应的控制指令。
  8. 根据权利要求7的耳机,其特征在于,还包括处理所述耳内麦克风和耳外麦克风信号以检测用户是否去除捂嘴手势;
    响应于检测到用户去除捂嘴手势,发送信号给智能电子设备结束所述交互过程。
  9. 一种智能电子设备,特征在于:
    可操作来与下面的单个耳机无线连接,或者集成有所述单个耳机,所述单个耳机具有两个麦克风,耳内麦克风和耳外麦克风,
    智能电子设备具有存储器和中央处理器,存储器上存储有计算机可执行指令,计算机可执行指令被中央处理器执行时能够执行如下操作:
    接收所述耳内麦克风和耳外麦克风采集的声音信号,
    分析耳内麦克风和耳外麦克风采集的声音信号,识别用户是否在做捂嘴手势的状态下发声。
  10. 根据权利要求9的智能电子设备,其特征在于,还具备用于检测佩戴耳机的用户说话的说话检测模块,其中
    在分析耳内麦克风和耳外麦克风采集的声音信号,识别用户是否在做捂嘴手势的状态下发声之前,所述耳机上的耳内麦克风和耳外麦克风处于关闭状态,
    所述说话检测模块检测佩戴耳机的用户是否在说话,以及在识别到用户开始说话之后,打开耳机上的耳内麦克风和耳外麦克风,进行声音信号采集并识别。
  11. 根据权利要求9的智能电子设备,其特征在于,所述“分析耳内 麦克风和耳外麦克风采集的信号,识别用户是否在做捂嘴手势”,包括:
    对从耳内麦克风和耳外麦克风采集到的两路声音信号做人声信号增强
    分别计算两路增强后信号的能量幅值,计算所述两路信号的能量幅值比值,识别耳外麦克风采集的用户声音信号在从用户口腔发出通过空气传到耳外麦克风之间的路径上有没有被遮挡,并基于此判断用户是否在做捂嘴手势的状态下发声。
  12. 根据权利要求9的智能电子设备,其特征在于,所述耳外的麦克风是空气传导麦克风。
  13. 根据权利要求9的智能电子设备,其特征在于,所述耳内的麦克风为空气传导麦克风或骨传导麦克风。
  14. 根据权利要求9的智能电子设备,其特征在于,所述分析耳内麦克风和耳外麦克风采集的声音信号,识别用户是否在做捂嘴手势的状态下发声包括:
    计算耳机上耳内麦克风和耳外麦克风接收到的用户声音信号能量幅值比值;
    在耳内麦克风和耳外麦克风接收到的用户声音信号能量幅值比值超过预设阈值时,判断用户是在做捂嘴手势的状态下发声。
  15. 根据权利要求9的智能电子设备,其特征在于,计算机可执行指令被中央处理器执行时能够执行的操作还包括:
    响应于识别出用户是在做出捂嘴手势的状态下,将指示识别结果的信号作为用户交互输入控制的指示,控制智能电子设备上的程序执行,包括触发相应的控制指令。
  16. 根据权利要求15的智能电子设备,其特征在于,执行的控制指令为触发除捂嘴手势外的其它输入方式,即处理其它输入方式输入的信息。
  17. 根据权利要求16的智能电子设备,其特征在于,所述其他输入方式包括语音输入、非捂嘴手势输入、视线输入、眨眼输入、头动输入之一或者其组合。
  18. 根据权利要求15的智能电子设备,其特征在于,处理所述信号以检测用户是否去除捂嘴手势;
    响应于检测到用户去除捂嘴手势,智能电子设备结束所述交互过程。
  19. 根据权利要求15所述的智能电子设备,其特征在于,提供包括视觉、听觉任一项反馈,提示用户智能电子设备已经触发其他输入方式。
  20. 根据权利要求15的智能电子设备,其特征在于,智能电子设备对用户在保持捂嘴手势同时进行的语音输入进行处理。
  21. 根据权利要求9的智能电子设备,其特征在于,所述智能电子设备为手机、手表、智能戒指、腕表中的一种智能穿戴设备。
  22. 根据权利要求9的智能电子设备,其特征在于,所述智能电子设备为头戴式智能显示设备,装备有所述耳内麦克风和耳外麦克风。
  23. 一种如权利要求1到22任一项所述的智能电子设备的语音交互唤醒方法,其特征在于,所述智能电子设备执行的语音交互唤醒方法包括:
    接收所述耳内麦克风和耳外麦克风采集的声音信号,
    分析耳内麦克风和耳外麦克风采集的声音信号,识别用户是否在做捂嘴手势的状态下发声;
    响应于识别到用户在做捂嘴手势的状态下发声,智能设备触发语音输入处理,分析及做出相应的内容输出;
    响应用户捂嘴手势后,在用户与智能设备交互情况下,处理所述耳内麦克风和耳外麦克风采集的声音信号,以确定用户去除捂嘴手势;
    响应于确定用户去除捂嘴手势,结束所述交互过程。
  24. 根据权力要求23的语音交互唤醒方法,其特征在于,所述内容输出形式包括语音、图像中一种或其组合。
  25. 一种计算机可读介质,其特征在于,其上存储有计算机可执行指令,计算机可执行指令被计算机执行时能够执行权利要求23-24任一项所述的语音交互唤醒方法。
PCT/CN2020/093161 2020-03-19 2020-05-29 单耳耳机、智能电子设备、方法和计算机可读介质 Ceased WO2021184549A1 (zh)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN202010198596.6A CN111432303B (zh) 2020-03-19 2020-03-19 单耳耳机、智能电子设备、方法和计算机可读介质
CN202010198596.6 2020-03-19

Publications (1)

Publication Number Publication Date
WO2021184549A1 true WO2021184549A1 (zh) 2021-09-23

Family

ID=71555389

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2020/093161 Ceased WO2021184549A1 (zh) 2020-03-19 2020-05-29 单耳耳机、智能电子设备、方法和计算机可读介质

Country Status (2)

Country Link
CN (1) CN111432303B (zh)
WO (1) WO2021184549A1 (zh)

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN114120603A (zh) * 2021-11-26 2022-03-01 歌尔科技有限公司 语音控制方法、耳机和存储介质

Families Citing this family (8)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN110164440B (zh) * 2019-06-03 2022-08-09 交互未来(北京)科技有限公司 基于捂嘴动作识别的语音交互唤醒电子设备、方法和介质
CN112055278B (zh) * 2020-08-17 2022-03-08 大象声科(深圳)科技有限公司 融合入耳麦克风和耳外麦克风的深度学习降噪设备
CN112133313A (zh) * 2020-10-21 2020-12-25 交互未来(北京)科技有限公司 基于单耳机语音对话过程捂嘴手势的识别方法
CN112259124B (zh) * 2020-10-21 2021-06-15 交互未来(北京)科技有限公司 基于音频频域特征的对话过程捂嘴手势识别方法
CN115132212A (zh) * 2021-03-24 2022-09-30 华为技术有限公司 一种语音控制方法和装置
CN113473299B (zh) * 2021-07-22 2025-12-12 立讯电子科技(昆山)有限公司 控制方法和穿戴式装置
CN113825063B (zh) * 2021-11-24 2022-03-15 珠海深圳清华大学研究院创新中心 耳机的语音识别启动方法及耳机的语音识别方法
CN114143651A (zh) * 2021-11-26 2022-03-04 思必驰科技股份有限公司 用于骨传导耳机的语音唤醒方法和装置

Citations (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20170186446A1 (en) * 2015-12-24 2017-06-29 Michal Wosk Mouth proximity detection
CN108305637A (zh) * 2018-01-23 2018-07-20 广东欧珀移动通信有限公司 耳机语音处理方法、终端设备及存储介质
CN108882087A (zh) * 2018-06-12 2018-11-23 歌尔科技有限公司 一种智能语音检测方法、无线耳机、tws耳机及终端
CN109949810A (zh) * 2019-03-28 2019-06-28 华为技术有限公司 一种语音唤醒方法、装置、设备及介质
CN110164440A (zh) * 2019-06-03 2019-08-23 清华大学 基于捂嘴动作识别的语音交互唤醒电子设备、方法和介质
CN110837353A (zh) * 2018-08-17 2020-02-25 宏达国际电子股份有限公司 补偿耳内音频信号的方法、电子装置及记录介质

Family Cites Families (7)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
KR102128562B1 (ko) * 2008-11-10 2020-06-30 구글 엘엘씨 멀티센서 음성 검출
CN205283527U (zh) * 2015-12-22 2016-06-01 深圳市中安瑞科通信有限公司 半双工无线机搭配蓝牙的送受话系统
US10477328B2 (en) * 2016-08-01 2019-11-12 Qualcomm Incorporated Audio-based device control
EP3611612A1 (en) * 2018-08-14 2020-02-19 Nokia Technologies Oy Determining a user input
CN110265036A (zh) * 2019-06-06 2019-09-20 湖南国声声学科技股份有限公司 语音唤醒方法、系统、电子设备及计算机可读存储介质
CN110121129B (zh) * 2019-06-20 2021-04-20 歌尔股份有限公司 耳机的麦克风阵列降噪方法、装置、耳机及tws耳机
CN110445931A (zh) * 2019-08-01 2019-11-12 花豹科技有限公司 语音识别开启方法及电子设备

Patent Citations (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20170186446A1 (en) * 2015-12-24 2017-06-29 Michal Wosk Mouth proximity detection
CN108305637A (zh) * 2018-01-23 2018-07-20 广东欧珀移动通信有限公司 耳机语音处理方法、终端设备及存储介质
CN108882087A (zh) * 2018-06-12 2018-11-23 歌尔科技有限公司 一种智能语音检测方法、无线耳机、tws耳机及终端
CN110837353A (zh) * 2018-08-17 2020-02-25 宏达国际电子股份有限公司 补偿耳内音频信号的方法、电子装置及记录介质
CN109949810A (zh) * 2019-03-28 2019-06-28 华为技术有限公司 一种语音唤醒方法、装置、设备及介质
CN110164440A (zh) * 2019-06-03 2019-08-23 清华大学 基于捂嘴动作识别的语音交互唤醒电子设备、方法和介质

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN114120603A (zh) * 2021-11-26 2022-03-01 歌尔科技有限公司 语音控制方法、耳机和存储介质
CN114120603B (zh) * 2021-11-26 2023-08-08 歌尔科技有限公司 语音控制方法、耳机和存储介质

Also Published As

Publication number Publication date
CN111432303B (zh) 2023-01-10
CN111432303A (zh) 2020-07-17

Similar Documents

Publication Publication Date Title
CN111432303B (zh) 单耳耳机、智能电子设备、方法和计算机可读介质
US12112756B2 (en) Voice interaction wakeup electronic device, method and medium based on mouth-covering action recognition
CN108710615B (zh) 翻译方法及相关设备
CN110097875B (zh) 基于麦克风信号的语音交互唤醒电子设备、方法和介质
CN110428806B (zh) 基于麦克风信号的语音交互唤醒电子设备、方法和介质
CN110223711B (zh) 基于麦克风信号的语音交互唤醒电子设备、方法和介质
JP5998861B2 (ja) 情報処理装置、情報処理方法及びプログラム
CN105988768B (zh) 智能设备控制方法、信号获取方法及相关设备
CN106714023A (zh) 一种基于骨传导耳机的语音唤醒方法、系统及骨传导耳机
KR101598400B1 (ko) 이어셋 및 그 제어 방법
WO2020244257A1 (zh) 语音唤醒方法、系统、电子设备及计算机可读存储介质
CN106685459B (zh) 一种可穿戴设备操作的控制方法及可穿戴设备
WO2022199405A1 (zh) 一种语音控制方法和装置
CN109067965B (zh) 翻译方法、翻译装置、可穿戴装置及存储介质
CN106713569B (zh) 一种可穿戴设备的操作控制方法及可穿戴设备
CN111105796A (zh) 无线耳机控制装置及控制方法、语音控制设置方法和系统
JP2009178783A (ja) コミュニケーションロボット及びその制御方法
US20250232787A1 (en) Voice control method and apparatus chip, earphones, and system
CN106686231A (zh) 一种可穿戴设备的消息播放方法及可穿戴设备
CN110111776A (zh) 基于麦克风信号的语音交互唤醒电子设备、方法和介质
CN106774915A (zh) 一种可穿戴设备通信消息的收发控制方法及可穿戴设备
US20170024380A1 (en) System and method for the translation of sign languages into synthetic voices
CN112259124B (zh) 基于音频频域特征的对话过程捂嘴手势识别方法
CN112133313A (zh) 基于单耳机语音对话过程捂嘴手势的识别方法
CN205582480U (zh) 一种智能声控系统

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 20925451

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 20925451

Country of ref document: EP

Kind code of ref document: A1