WO2020244401A1 - 基于靠近嘴部检测的语音输入唤醒装置、方法和介质 - Google Patents
基于靠近嘴部检测的语音输入唤醒装置、方法和介质 Download PDFInfo
- Publication number
- WO2020244401A1 WO2020244401A1 PCT/CN2020/092066 CN2020092066W WO2020244401A1 WO 2020244401 A1 WO2020244401 A1 WO 2020244401A1 CN 2020092066 W CN2020092066 W CN 2020092066W WO 2020244401 A1 WO2020244401 A1 WO 2020244401A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- user
- electronic device
- mouth
- smart electronic
- voice input
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F3/00—Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
- G06F3/16—Sound input; Sound output
- G06F3/167—Audio in a user interface, e.g. using voice commands for navigating, audio feedback
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04M—TELEPHONIC COMMUNICATION
- H04M1/00—Substation equipment, e.g. for use by subscribers
- H04M1/72—Mobile telephones; Cordless telephones, i.e. devices for establishing wireless links to base stations without route selection
- H04M1/724—User interfaces specially adapted for cordless or mobile telephones
- H04M1/72448—User interfaces specially adapted for cordless or mobile telephones with means for adapting the functionality of the device according to specific conditions
Definitions
- the present invention generally relates to the field of voice input, and more specifically, to smart electronic devices and voice input triggering methods.
- voice input After pressing (or holding down) a certain (or certain) physical buttons of the mobile device, voice input is activated.
- the device needs to have a screen; the trigger element occupies the screen content; the limitation of the software UI leads to a cumbersome trigger method; and it is easy to trigger by mistake.
- the device activates voice input after detecting the corresponding wake-up word.
- a specific word such as product nickname
- the voice input is automatically activated.
- an intelligent electronic device including a sensor system, capable of capturing a signal from which it can be determined that the intelligent electronic device is close to the user's mouth.
- the intelligent electronic device includes a memory and a processor, and a computer is stored in the memory. Executable instructions, when the computer-executable instructions are executed by the processor, are operable to process the signal to determine whether the smart electronic device is close to the user's mouth, and to activate voice input in response to determining that it is close to the user's mouth.
- the sensor system can also capture a signal from which it can determine that the smart electronic device touches the face part near the mouth, and the smart electronic device processes the signal to determine whether to touch the face part near the user's mouth, and responds To determine that the user touches the face part near the user's mouth, the smart electronic device is determined to be close to the user's mouth, and the voice input is activated.
- the smart electronic device determines that it is close to the user's mouth.
- the recognized specific posture includes one or more of the following: the smart electronic device is close to the user's mouth, but does not touch the user's face, and the approach distance is 0 to 3 cm; the smart electronic device is close to the user's mouth, but Do not touch the user’s face, and the approach distance is 3-10 cm; the smart electronic device is close to the user’s mouth, touches the user’s face, and the part of the face near the touched mouth is the nose; the smart electronic device is close to the user’s mouth, Touching the user’s face, the part of the face near the mouth touched is the part between the upper lip and the nose; the smart electronic device is close to the user’s mouth, touching the user’s face, near the mouth touched by the smart electronic device
- the face part of is the chin; the smart electronic device approaches the user's mouth and touches the user's face, and the face part near the mouth touched by the smart electronic device is the cheek.
- the smart electronic device also recognizes the posture of the smart electronic device approaching the user's mouth through the sensor signal; in response to the recognized specific posture, the smart electronic device processes the voice input in a specific manner.
- the smart electronic device in response to determining that the smart electronic device is close to the user's mouth, provides at least one of image, sound, or tactile feedback to prompt the user to perform voice input.
- the signal is processed to determine whether the user's mouth leaves the smart electronic device; in response to determining that the user's mouth leaves the smart electronic device, the voice input is ended.
- the smart electronic device processing the signal to determine whether the smart electronic device is close to the user's mouth includes: calculating the probability of the smart electronic device being close to the user's mouth; comparing the probability with a predetermined probability threshold, when the probability is greater than or equal to When the predetermined probability threshold is set, it is determined that the smart electronic device is close to the user's mouth.
- activating the voice input includes: judging whether to activate the voice input in combination with the intelligent electronic portable device itself, the user and the environment.
- the judging whether to activate the voice input according to the situation of the smart electronic portable device itself, the user and the environment includes: determining whether the user is a specific authorized user by the voiceprint, and in the case of determining that the user is a specific authorized user, activating Voice input.
- the sensor system includes a common camera.
- the sensor system includes an infrared camera.
- the sensor system includes a depth camera.
- the sensor system package is close to the light sensor.
- the sensor system includes a distance sensor.
- the sensor system includes a wide-angle camera.
- the sensor system includes a capacitive sensing sensor.
- the sensor system includes a motion sensor.
- the sensor system includes a camera
- the smart electronic device analyzes the image signal collected by the camera, detects whether there are features of the face near the mouth at a close range in the image, and recognizes whether the smart electronic device is close to the mouth.
- the sensor system further includes a distance sensor, which detects the distance between the smart electronic device and the user's face through the distance sensor signal.
- the sensor system further includes a proximity light sensor, which recognizes whether the smart electronic device is close to the user's face through the proximity light sensor signal.
- the sensor system further includes a capacitive sensing sensor, which recognizes whether the smart electronic device touches the user's face through the capacitive sensing sensor signal.
- the smart electronic device recognizes the touched user's face part through the capacitive sensor signal.
- the sensor system on the smart phone includes an accelerometer, a gyroscope, and a proximity light sensor; the proximity light sensor recognizes that the front of the smart electronic device is blocked, and recognizes the user's action of putting the mobile phone to his mouth based on the previous accelerometer and gyroscope signals , Instead of placing it to the ear.
- a voice input trigger method for a smart electronic device includes a sensor system capable of capturing a signal from which it can be determined that the smart electronic device is close to the user's mouth.
- the voice input triggers includes: processing the signal to determine whether the smart electronic device is close to the user's mouth; in response to determining that the smart electronic device is close to the user's mouth, activating voice input.
- the sensor system can also capture a signal from which it can determine that the smart electronic device touches the face part near the mouth, and the smart electronic device processes the signal to determine whether to touch the face part near the mouth of the user , In response to determining that it touches the face part near the user's mouth, it is determined that the smart electronic device is close to the user's mouth to activate the voice input.
- the smart electronic device determines that it is close to the user's mouth.
- the recognized specific posture includes one or more of the following: the smart electronic device is close to the user's mouth, but does not touch the user's face, and the approach distance is 0 to 3 cm; the smart electronic device is close to the user's mouth, but Do not touch the user’s face, and the approach distance is 3-10 cm; the smart electronic device is close to the user’s mouth, touches the user’s face, and the part of the face near the touched mouth is the nose; the smart electronic device is close to the user’s mouth, Touching the user’s face, the part of the face near the mouth touched is the part between the upper lip and the nose; the smart electronic device is close to the user’s mouth, touching the user’s face, near the mouth touched by the smart electronic device
- the face part of is the chin; the smart electronic device approaches the user's mouth and touches the user's face, and the face part near the mouth touched by the smart electronic device is the cheek.
- the voice input triggering method further includes: recognizing the posture of the smart electronic device approaching the user's mouth through the sensor signal; in response to the recognized specific posture, the smart electronic device processes the voice input in a specific manner.
- the voice input trigger method further includes: in response to determining that it is close to the user's mouth, the smart electronic device provides at least one of image, sound or tactile feedback to prompt the user to perform voice input.
- the voice input triggering method further includes processing the signal to determine whether the user's mouth leaves the smart electronic device after activating the voice input; in response to determining that the user's mouth leaves the smart electronic device, ending the voice input.
- the smart electronic device processing the signal to determine whether the smart electronic device is close to the user's mouth includes: calculating the probability of the smart electronic device being close to the user's mouth; comparing the probability with a predetermined probability threshold, when the probability is greater than or equal to When the predetermined probability threshold is set, it is determined that the smart electronic device is close to the user's mouth.
- activating the voice input includes: judging whether to activate the voice input in combination with the intelligent electronic portable device itself, the user and the environment.
- the judging whether to activate the voice input according to the situation of the smart electronic portable device itself, the user and the environment includes: determining whether the user is a specific authorized user by the voiceprint, and in the case of determining that the user is a specific authorized user, activating Voice input.
- the mobile devices here include, but are not limited to, mobile phones, watches, and smaller smart wearable devices such as smart rings and watches.
- a computer-readable storage medium having computer-readable instructions stored thereon, and the instructions, when executed by a computer, are operable to perform any of the methods described above.
- the use efficiency is higher. It can be used with one hand. No need to switch between different user interfaces/applications, no need to hold down a button, just lift your hand to your mouth to use it.
- the radio quality is high.
- the device's voice recorder is at the user's mouth, and the voice input signal received is clear, and it is less affected by environmental sounds.
- Fig. 1 is a schematic flowchart of a voice input interaction method according to an embodiment of the present invention.
- Fig. 2 is a schematic front view of an upper mouth covering posture in a trigger posture according to an embodiment of the present invention.
- Fig. 3 is a schematic side view of an upper mouth-covering posture in a trigger posture according to an embodiment of the present invention.
- Fig. 4 is a schematic diagram of a nose-touching gesture in a triggering gesture according to an embodiment of the present invention.
- Fig. 5 is a schematic diagram of a no-nosing posture in a trigger posture according to an embodiment of the present invention.
- the proximity of the electronic device to the user’s mouth means that the distance between the electronic device and the user’s mouth is within a predetermined distance threshold or based on the probability that the probability of the electronic device being close to the user’s mouth is greater than the predetermined probability threshold, including the electronic device touching the face near the mouth The contact condition of the part. Determining that the probability that the electronic device is close to the user's mouth is greater than the predetermined probability threshold includes both explicit calculation of the probability and implicit judgment, such as autonomous learning through a deep neural network to determine whether the electronic device is close to the user's mouth.
- an intelligent electronic device including a sensor system, capable of capturing a signal from which it can be determined that the intelligent electronic device is close to the user's mouth.
- the intelligent electronic device includes a memory and a processor, and a computer is stored in the memory. Executable instructions, when the computer-executable instructions are executed by the processor, are operable to process the signal to determine whether the smart electronic device is close to the user's mouth, and to activate voice input in response to determining that it is close to the user's mouth.
- the smart electronic device here may be a smart phone, a smart watch, a smart ring, and so on.
- mobile phones are mainly used as examples of smart electronic devices.
- Fig. 1 is a schematic flowchart of a voice input interaction method according to an embodiment of the present invention.
- the user moves the smart electronic device to his mouth to enable voice input.
- Figures 2 to 5 show several cases where the user moves the smart electronic device to his mouth to trigger voice input.
- 2 and 3 are schematic diagrams of the front and side of the upper mouth-covering posture in the trigger posture, respectively.
- the user moves the upper end of the phone between the nose and the lips, that is, near the person, covering the mouth.
- the upper end of the mobile phone can be placed on top of a person, or 1-10 cm away from the face.
- 4 and 5 are schematic diagrams of the nose-touching posture and the no-nosing posture in the triggering posture, respectively.
- the above description of the triggering postures is exemplary, not exhaustive, and is not limited to the disclosed postures.
- step S102 the smart electronic device receives the signal sensed by its own sensor, processes the signal, and detects that it is moved to the user's mouth.
- step S103 the smart electronic device processes the signal detected by the aforementioned sensor to determine whether the smart electronic device is close to the user's mouth.
- the smart electronic device When the user moves the smart electronic device to the mouth, the smart electronic device detects and recognizes whether it has been moved to the user's mouth through its various sensors.
- some types of sensors are taken as examples for description, in which it is determined that the user is moved to the user's mouth and it is interpreted that the user needs to trigger a voice input.
- the first example sensor system includes a proximity sensor and a camera
- Proximity sensor is a general term for sensors that replace contact detection methods such as limit switches and detect objects without touching the detection object.
- Proximity sensors include inductive, capacitive, ultrasonic, photoelectric, magnetic, and other types of sensors. .
- the camera image is triggered to collect the camera image, and the camera image is used to determine whether the smart device is near the mouth.
- the second example sensor system includes accelerometer and camera
- the accelerometer is used to detect the state of the smart device from moving to still, and the camera image is triggered to collect the camera image, and judge whether the smart device is in the mouth by whether facial features, including nose, mouth, etc. appear in the camera image.
- the third example is the case where the sensor system on a smartphone includes an accelerometer, a gyroscope, and a proximity sensor
- the proximity light sensor recognizes that the front of the smart device is blocked, and based on the previous accelerometer and gyroscope signals, it recognizes the user's movement of placing the mobile phone to the mouth instead of the movement to the ear.
- the movement state of the smartphone is accelerated first and then stopped.
- This mode can be detected by the accelerometer; in the final stage of the movement, the movement direction of the smartphone is almost vertical
- this mode can be detected by the direction of acceleration; the rotation and orientation changes during the overall movement of the smartphone can be detected by the gyroscope.
- the fourth example sensor system includes a camera case
- the image captured by the front camera detects the specific features of the user's face, such as the eyes, mouth, skin, and features of other objects, such as glasses, taken at a very close distance, it is judged that the user is located near the user's mouth.
- the fifth example sensor system includes capacitive touch screen case
- the smart electronic device When the user uses the nose-touch gesture as shown in FIG. 4 to trigger a voice input, the smart electronic device will record the capacitive image signal between the nose and the center of the screen, and infer that it is located next to the user's mouth.
- the sixth example is that the smart electronic device detects the posture of the user using the device through the sensor system.
- the capacitive screen detects the capacitive image of the user's nose and the one that is not detected corresponds to two different postures ( Figure 4 and Figure 5).
- the device responds and processes the user’s voice information differently. For example, when not touching the nose, the device understands and processes the user’s voice information according to natural language; while when touching the nose, the device uniformly follows Send a voice message to understand and execute.
- the seventh example sensor system includes a distance sensor, and detects the distance between the smart electronic device and the user's face through ToF (time of flight) distance sensor signals.
- ToF time of flight
- the smart electronic device determines that it is close to the user's mouth.
- the smart electronic device also recognizes the posture of the smart electronic device approaching the user's mouth through the sensor signal; in response to the recognized specific posture, the smart electronic device processes the voice input in a specific manner.
- the recognized specific gesture includes one or more of the following:
- the smart electronic device is close to the user's mouth without touching it, and the proximity distance is 0 to 3 cm;
- the smart electronic device is close to the user's mouth without touching it, and the proximity distance is 3-10 cm;
- the part of the face near the mouth touched by the smart electronic device is the nose
- the part of the face near the mouth touched by the smart electronic device is the part between the upper lip and the nose;
- the part of the face near the mouth touched by the smart electronic device is the chin
- the part of the face near the mouth touched by the smart electronic device is the cheek.
- the smart electronic device in response to determining that the smart electronic device is close to the user's mouth, the smart electronic device provides at least one of image, sound, or tactile feedback to prompt the user to perform voice input.
- the signal is processed to determine whether the user's mouth leaves the smart electronic device, and in response to determining that the user's mouth leaves the smart electronic device, the voice input is ended.
- processing the signal by the smart electronic device to determine whether the smart electronic device is close to the user's mouth includes: calculating the probability of the smart electronic device being close to the user's mouth; comparing the probability with a predetermined probability threshold, when the probability When it is greater than or equal to the predetermined probability threshold, it is determined that the smart electronic device is close to the user's mouth.
- activating the voice input includes: determining whether to activate the voice input in combination with the smart electronic portable device itself, the user and the environment.
- determining whether to activate voice input may include: determining whether the user is a specific authorized user by voiceprint, and in the case of determining that the user is a specific authorized user, Activate voice input.
- the sensor system includes one or more of the following: a normal camera; an infrared camera; a depth camera; a proximity light sensor; a distance sensor; a wide-angle camera; a capacitive sensing sensor; a motion sensor.
- step S104 in response to determining that it is close to the user's mouth, the smart electronic device directly activates the voice input.
- the smart electronic device detects that it moves to the user's mouth, that is, when the user needs to use voice input, the smart electronic device activates the voice input, for example, turns on the microphone to record the user's voice information.
- the smart electronic device can make a feedback output to help the user confirm that the voice input can be started.
- the feedback output here is to notify the user that the voice input application has been started and is in the mode of recording and parsing and understanding the voice, rather than requesting input instructions from the user.
- the feedback output includes, but is not limited to, vibration, voice, image and other prompt methods: when the feedback mode is vibration, the user can obtain feedback that the voice input has been activated by feeling the vibration of the smart electronic device in the hand; when the feedback mode is voice, the device passes A short beep or natural voice "please input voice" prompts the user; when the feedback mode is an image, the device's screen will greatly change the screen color, so that the user can observe it even at a very close distance. After receiving the corresponding feedback, the user performs voice input to the smart electronic device by speaking.
- the smart electronic device records the user's voice content, according to the different tasks and contexts, combined with natural language processing technology to understand the user's voice input and complete the corresponding tasks.
- the method of detecting whether the mouth is removed is similar to the method of detecting the proximity of the mouth.
- a mobile phone is used as an example of a smart electronic device, but the smart electronic device is not limited to this, and may also be a wearable smart electronic watch, a smart bracelet, a smart ring, and so on.
- the previous article describes when the smart electronic device is close to the mouth. Taking the smart electronic device as a portable mobile phone as an example, it is explained that the mobile phone is moved to the mouth. However, this is an example. The smart electronic device can also be kept still. The mouth is close to the smart electronic device. For example, when the user is driving, assuming that the smart electronic device is fixed on the steering wheel, the user can actively make the mouth close to the smart electronic device.
- the use efficiency is higher. Users can use it with one hand, no need to switch between different user interfaces/applications, and no need to hold down a button. For example, in the case of a mobile phone, it can be used directly by raising the hand to the mouth.
- the radio quality is high.
- the device's voice recorder is at the user's mouth, and the voice input signal received is clear, and it is less affected by environmental sounds.
Landscapes
- Engineering & Computer Science (AREA)
- Human Computer Interaction (AREA)
- Theoretical Computer Science (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Multimedia (AREA)
- Health & Medical Sciences (AREA)
- Signal Processing (AREA)
- General Health & Medical Sciences (AREA)
- Physics & Mathematics (AREA)
- General Engineering & Computer Science (AREA)
- General Physics & Mathematics (AREA)
- Computer Networks & Wireless Communication (AREA)
- Telephone Function (AREA)
- User Interface Of Digital Computer (AREA)
Abstract
一种基于靠近嘴部检测的语音输入唤醒装置、方法和介质。该方法应用于具有传感器系统的智能电子设备。当智能电子设备位于用户的嘴边时,自动激活语音输入。移动设备通过传感器系统捕捉信号,判断智能电子设备是否接近用户嘴部。响应于确定自身位于用户嘴边,激活语音输入。在用户进行语音输入的情况下,传感器系统检测用户嘴部离开智能电子设备的信号,以结束语音输入应用。这适于用户在智能电子设备上进行语音输入,提高了语音输入的收音质量、效率与隐私性,并使得交互更加自然。
Description
本申请要求于2019年06月03日提交中国专利局、申请号为201910476243.5、发明名称为“基于靠近嘴部检测的语音输入唤醒装置、方法和介质”的中国专利申请的优先权,其全部内容通过引用结合在本申请中。
本发明总的来说涉及语音输入领域,且更为具体地,涉及智能电子设备、语音输入触发方法。
随着计算机技术的发展,语音识别算法日益成熟,语音输入因其在交互方式上的高自然性与有效性而正变得越来越重要。用户可以通过语音与移动设备(手机、手表等)进行交互,完成指令输入、信息查询、语音聊天等多种任务。
而在何时触发语音输入这一点上,现有的解决方案都有一些缺陷:
1.物理按键触发
按下(或按住)移动设备的某个(或某些)物理按键后,激活语音输入。
该方案的缺点是:需要物理按键;容易误触发;需要用户按键。
2.界面元素触发
点击(或按住)移动设备的屏幕上的界面元素(如图标),激活语音输入。
该方案的缺点是:需要设备具备屏幕;触发元素占用屏幕内容;受限于软件UI限制,导致触发方式繁琐;容易误触发。
3.唤醒词(语音)检测
以某个特定词语(如产品昵称)为唤醒词,设备检测到对应的唤醒词后激活语音输入。
该方案的缺点是:隐私性和社会性较差;交互效率较低;检测的准确率与语音信号相关,容易在日常对话中误触发。
发明内容
鉴于上述情况,提出了本发明:
当用户将移动设备移动到自己嘴部附近时,自动激活语音输入。
根据本发明的一个方面,提供了一种智能电子设备,包括传感器系统,能够捕捉到从其能判定智能电子设备接近用户嘴部的信号,智能电子设备包括存储器和处理器,存储器上存储有计算机可执行指令,所述计算机可执行指令被处理器执行时可操作来:处理所述信号以确定智能电子设备是否接近用户嘴部,响应于确定自身接近用户嘴部,激活语音输入。
优选的,传感器系统还能够捕捉到从其能判定智能电子设备碰触嘴部附近的脸部部位的信号,智能电子设备处理所述信号以确定是否碰触用户嘴部附近的脸部部位,响应于确定自身碰触用户嘴部附近的脸部部位,确定智能电子设备接近用户嘴部,激活语音输入。
优选的,在确定智能电子设备与用户脸部的距离处于0~10厘米范围内时,智能电子设备确定自身接近用户嘴部。
优选的,所识别的特定姿态包括下面的一种或者多种:智能电子设备接近用户嘴部,但不碰触用户脸部,接近距离为0~3厘米;智能电子设备接近用户嘴部,但不碰触用户脸部,接近距离为3~10厘米;智能电子设备接近用户嘴部,碰触用户脸部,碰触的嘴部附近的脸部部位为鼻子;智能电子设备接近用户嘴部,碰触用户脸部,碰触的嘴部附近的脸部部位为上嘴唇和鼻子之间的部位;智能电子设备接近用户嘴部,碰触用户脸部,智能电子设备所碰触的嘴部附近的脸部部位为下巴;智能电子设备接近用户嘴部,碰触用户脸部,智能电子设备所碰触的嘴部附近的脸部部位为面颊。
优选的,智能电子设备还通过所述传感器信号,识别智能电子设备接近用户嘴部的姿态;响应于识别的特定姿态,智能电子设备以特定的方式 处理语音输入。
优选的,智能电子设备响应于确定自身接近用户嘴部,智能电子设备提供图像、声音或者触觉反馈中的至少一种,提示用户进行语音输入。
优选的,在激活语音输入之后,处理所述信号以确定用户嘴部是否离开智能电子设备;响应于确定用户嘴部离开智能电子设备,结束语音输入。
优选的,智能电子设备处理所述信号以确定智能电子设备是否接近用户嘴部包括:计算智能电子设备接近用户嘴部的概率;将所述概率与预定概率阈值相比较,当所述概率大于等于预定概率阈值时,确定智能电子设备接近用户嘴部。
优选的,所述响应于确定自身接近用户嘴部,激活语音输入包括:结合智能电子便携设备自身、用户与环境的情况,判断是否激活语音输入。
优选的,所述结合智能电子便携设备自身、用户与环境的情况,判断是否激活语音输入包括:由声纹判定用户是否是特定的授权用户,在判定用户是特定的授权用户的情况下,激活语音输入。
优选的,传感器系统包括普通摄像头。
优选的,传感器系统包括红外摄像头。
优选的,传感器系统包括深度摄像头。
优选的,传感器系统包接近光传感器。
优选的,传感器系统包括距离传感器。
优选的,传感器系统包括广角摄像头。
优选的,传感器系统包括电容感应传感器。
优选的,传感器系统包括运动传感器。
优选的,所述传感器系统包括摄像头,智能电子设备分析摄像头采集的图像信号,检测图像中是否存在近距离拍摄嘴部附近的脸部部位特征,识别智能电子设备是否接近嘴部。
优选的,传感器系统还包括距离传感器,通过距离传感器信号检测智能电子设备与用户脸部的距离。
优选的,传感器系统还包括接近光传感器,通过接近光传感器信号识别智能电子设备是否接近用户脸部。
优选的,传感器系统还包括电容感应传感器,通过电容感应传感器信号识别智能电子设备是否碰触用户脸部。
优选的,其中智能电子设备通过电容感应传感器信号识别碰触的用户脸部部位。
优选的,智能手机上传感器系统包括加速度计、陀螺仪和接近光传感器;接近光传感器识别智能电子设备前方被遮挡,基于此前的加速度计和陀螺仪的信号识别用户将手机放到嘴边的动作,而非放置到耳朵边的动作。
根据本发明的另一方面,提供了一种智能电子设备的语音输入触发方法,智能电子设备包括传感器系统,能够捕捉到从其能判定智能电子设备接近用户嘴部的信号,所述语音输入触发方法包括:处理所述信号以确定智能电子设备是否接近用户嘴部;响应于确定自身接近用户嘴部,激活语音输入。
优选的,所述传感器系统还能够捕捉到从其能判定智能电子设备碰触嘴部附近的脸部部位的信号,智能电子设备处理所述信号以确定是否碰触用户嘴部附近的脸部部位,响应于确定自身碰触用户嘴部附近的脸部部位,确定智能电子设备接近用户嘴部,激活语音输入。
优选的,在确定智能电子设备与用户脸部的距离处于0~10厘米范围内时,智能电子设备确定自身接近用户嘴部。
优选的,所识别的特定姿态包括下面的一种或者多种:智能电子设备接近用户嘴部,但不碰触用户脸部,接近距离为0~3厘米;智能电子设备接近用户嘴部,但不碰触用户脸部,接近距离为3~10厘米;智能电子设备接近用户嘴部,碰触用户脸部,碰触的嘴部附近的脸部部位为鼻子;智能电子设备接近用户嘴部,碰触用户脸部,碰触的嘴部附近的脸部部位为上嘴唇和鼻子之间的部位;智能电子设备接近用户嘴部,碰触用户脸部,智能电子设备所碰触的嘴部附近的脸部部位为下巴;智能电子设备接近用户嘴部,碰触用户脸部,智能电子设备所碰触的嘴部附近的脸部部位为面颊。
优选的,语音输入触发方法还包括:通过所述传感器信号,识别智能电子设备接近用户嘴部的姿态;响应于识别的特定姿态,智能电子设备以特定的方式处理语音输入。
优选的,语音输入触发方法还包括:响应于确定自身接近用户嘴部,智能电子设备提供图像、声音或者触觉反馈中的至少一种,提示用户进行语音输入。
优选的,语音输入触发方法还包括在激活语音输入之后,处理所述信号以确定用户嘴部是否离开智能电子设备;响应于确定用户嘴部离开智能电子设备,结束语音输入。
优选的,智能电子设备处理所述信号以确定智能电子设备是否接近用户嘴部包括:计算智能电子设备接近用户嘴部的概率;将所述概率与预定概率阈值相比较,当所述概率大于等于预定概率阈值时,确定智能电子设备接近用户嘴部。
优选的,所述响应于确定自身接近用户嘴部,激活语音输入包括:结合智能电子便携设备自身、用户与环境的情况,判断是否激活语音输入。
优选的,所述结合智能电子便携设备自身、用户与环境的情况,判断是否激活语音输入包括:由声纹判定用户是否是特定的授权用户,在判定用户是特定的授权用户的情况下,激活语音输入。
此处的移动设备包括但不限于手机、手表,以及智能戒指、腕表等更小型的智能穿戴设备。
根据本发明的另一方面,提供了一种计算机可读存储介质,其上存储有计算机可读指令,所述指令当被计算机执行时,可操作来执行前面任一项所述的方法。
根据本发明实施例的智能电子设备和语音触发方法具有如下优势中的一个或多个:
1.交互更加自然。将设备放在嘴前即触发语音输入,符合用户习惯与认知。
2.使用效率更高。单手即可使用。无需在不同的用户界面/应用之间切换,也不需按住某个按键,直接抬起手到嘴边就能使用。
3.收音质量高。设备的录音机在用户嘴边,收取的语音输入信号清晰,受环境音的影响较小。
4.高隐私性与社会性。设备在嘴前,则用户只需发出相对较小的声音 即可完成高质量的语音输入,对他人的干扰较小,同时具有较好的隐私保护。
从下面结合附图对本发明实施例的详细描述中,本发明的上述和/或其它目的、特征和优势将变得更加清楚并更容易理解。其中:
图1是根据本发明实施例的语音输入交互方法的示意性流程图。
图2是根据本发明实施例的触发姿势中的上端遮嘴姿势的正面示意图。
图3是根据本发明实施例的触发姿势中的上端遮嘴姿势的侧面示意图。
图4是根据本发明实施例的触发姿势中的碰触鼻子姿势的示意图。
图5是根据本发明实施例的触发姿势中的不碰鼻子姿势的示意图。
为了使本领域技术人员更好地理解本发明,下面结合附图和具体实施方式对本发明作进一步详细说明。
在本文中电子设备接近用户嘴部指的是电子设备与用户嘴部距离在预定距离阈值或者基于概率判定电子设备与用户嘴部接近的概率大于预定概率阈值,包括电子设备接触嘴部附近的脸部部位的触碰情况。判定电子设备与用户嘴部接近的概率大于预定概率阈值包括显式地计算概率,也包括隐式地判断,例如通过深度神经网络来进行自主学习、判定电子设备是否接近用户嘴部。
根据本发明一个实施例,提供了一种智能电子设备,包括传感器系统,能够捕捉到从其能判定智能电子设备接近用户嘴部的信号,智能电子设备包括存储器和处理器,存储器上存储有计算机可执行指令,所述计算机可执行指令被处理器执行时可操作来:处理所述信号以确定智能电子设备是否接近用户嘴部,响应于确定自身接近用户嘴部,激活语音输入。
作为例子而非作为限制,这里的智能电子设备可以是智能手机、智能 手表、智能指环等等。
下文中,主要以手机作为智能电子设备的例子。
图1是根据本发明实施例的语音输入交互方法的示意性流程图。
如图1所示,S101,用户通过将智能电子设备移动到嘴边,以启用语音输入。
图2至图5显示了几例用户将智能电子设备移动到嘴边以触发语音输入的情况。其中,图2与图3分别是触发姿势中的上端遮嘴姿势的正面与侧面示意图。在这种姿势下,用户将手机的上端移动到鼻子与嘴唇之间,也即人中附近,遮挡嘴部。根据不同用户的使用习惯,手机上端既可以顶在人中上,也可以距离脸部1~10厘米。图4与图5分别是触发姿势中的碰触鼻子姿势与不碰鼻子姿势的示意图。上述对触发姿势的说明是示例性的,并非穷尽性的,并且也不限于所披露的各姿势。
在步骤S102中,智能电子设备接收自身的传感器感测的信号,处理所述信号,检测自身被移动到用户嘴前。
在步骤S103中,智能电子设备处理上述传感器检测的信号,以确定智能电子设备是否接近用户嘴部。
当用户将智能电子设备移动到嘴边时,智能电子设备通过自身的各种传感器,检测和识别自身是否被移动到用户嘴边。下面以某几种传感器为例进行说明,其中判断到自身被移动到用户嘴边被解释为用户需要触发语音输入。
需要说明的是,以下所有实施例示例仅从某单个传感器本身出发,给出该传感器预测的概率值,不过这仅为示例,而非作为限制,实际应用中,识别算法很可能会综合传感器系统中的多种传感器结果,给出最终的识别结果。
第一示例传感器系统包括接近传感器和摄像头的情况
接近传感器是替代限位开关等接触式检测方式,以无需接触检测对象进行检测为目的的传感器的总称,接近传感器例如有感应型、静电容量型、超声波型、光电型、磁力型等种类的传感器。
接近传感器的读数值从不接近变更为接近时,例如状态显示变更为接 近时,触发采集摄像头图像,通过摄像头图像中是否出现脸部特征,包括鼻子、嘴等判断智能设备是否处于嘴边。
第二示例传感器系统包括加速度计和摄像头的情况
利用加速度计检测到智能设备从运动到静止的状态,触发采集摄像头图像,通过摄像头图像中是否出现脸部特征,包括鼻子、嘴等判断智能设备是否处于嘴边。
第三示例智能手机上传感器系统包括加速度计、陀螺仪和接近光传感器的情况
接近光传感器识别智能设备前方被遮挡,基于此前的加速度计和陀螺仪的信号识别用户将手机放到嘴边的动作,而非放置到耳朵边的动作。
具体地,例如当用户持握智能手机移动到嘴部附近时,智能手机的运动状态为先加速后停止,该模式可以通过加速度计检测;在移动的最后阶段,智能手机是运动方向是近乎垂直于手机平面的,该模式可以通过加速度的方向检测;智能手机整体移动过程中会发生转动和朝向的变化,可以通过陀螺仪检测。
第四示例传感器系统包括摄像头情况
通过前置摄像头拍摄的图像,检测到用户脸部特定特征,如极近距离拍摄到的眼睛、嘴部、皮肤,以及其他物体的特征,如眼镜等时,判断自身位于用户嘴边。
第五示例传感器系统包括电容触摸屏情况
在用户使用如图4所示的碰触鼻子姿势触发语音输入时,智能电子设备会记录下鼻子与屏幕中央的电容图像信号,进而推断出自身位于用户嘴边。
第六示例智能电子设备通过传感器系统检测用户使用设备的姿势。比如,电容屏幕检测到用户鼻子的电容图像与没检测到的即对应为两种不同 的姿势(图4和图5)。在不同的姿势下,设备对用户的语音信息进行不同的响应与处理,如在不碰触鼻子时,设备按照自然语言理解并处理用户的语音信息;而在碰触鼻子时,设备则统一按照发送语音消息理解并执行。
第七示例传感器系统包括距离传感器,通过ToF(time of flight)距离传感器信号检测智能电子设备与用户脸部的距离。
在一个示例中,在确定智能电子设备与用户脸部的距离处于0~10厘米范围内时,智能电子设备确定自身接近用户嘴部。
在一个示例中,智能电子设备还通过所述传感器信号,识别智能电子设备接近用户嘴部的姿态;响应于识别到的特定姿态,智能电子设备以特定的方式处理语音输入。
在一个示例中,所识别的特定姿态包括下面的一种或者多种:
智能电子设备接近用户嘴部,但不碰触,接近距离为0~3厘米;
智能电子设备接近用户嘴部,但不碰触,接近距离为3~10厘米;
智能电子设备所碰触的嘴部附近的脸部部位为鼻子;
智能电子设备所碰触的嘴部附近的脸部部位为上嘴唇和鼻子之间的部位;
智能电子设备所碰触的嘴部附近的脸部部位为下巴;
智能电子设备所碰触的嘴部附近的脸部部位为面颊。
在一个示例中,智能电子设备响应于确定自身接近用户嘴部,智能电子设备提供图像、声音或者触觉反馈中的至少一种,提示用户进行语音输入。
在一个示例中,在激活语音输入之后,处理所述信号以确定用户嘴部是否离开智能电子设备,以及响应于确定用户嘴部离开智能电子设备,结束语音输入。
在一个示例中,智能电子设备处理所述信号以确定智能电子设备是否接近用户嘴部包括:计算智能电子设备接近用户嘴部的概率;将所述概率与预定概率阈值相比较,当所述概率大于等于预定概率阈值时,确定智能电子设备接近用户嘴部。
在一个示例中,响应于确定自身接近用户嘴部,激活语音输入包括:结合智能电子便携设备自身、用户与环境的情况,判断是否激活语音输入。
在一个示例中,结合智能电子便携设备自身、用户与环境的情况,判断是否激活语音输入可以包括:由声纹判定用户是否是特定的授权用户,在判定用户是特定的授权用户的情况下,激活语音输入。
作为示例而非作为限制,传感器系统包括下述中的一种或多种:普通摄像头;红外摄像头;深度摄像头;接近光传感器;距离传感器;广角摄像头;电容感应传感器;运动传感器。
在步骤S104中,响应于确定自身接近用户嘴部,智能电子设备直接激活语音输入。在智能电子设备检测到自身移动到用户嘴边,也即用户需要使用语音输入时,智能电子设备激活语音输入,例如开启麦克风录制用户的语音信息。
可选地,智能电子设备可以做出反馈输出,帮助用户确认已可开始语音输入。需要说明的是,这里的反馈输出是向用户通知,语音输入应用已经启动,处于进行录音和解析理解语音的模式下,而并非是向用户来请求输入指令。其中所述反馈输出包括但不限于震动、语音、图像等提示方式:反馈模式是震动时,用户可以通过感受手中智能电子设备的震动获得已开启语音输入的反馈;反馈模式是语音时,设备通过发出短促的提示音或自然语音“请进行语音输入”提示用户;反馈模式是图像时,设备的屏幕会大幅改变屏幕色调,使得用户在极近距离下也能通过余光观察到。用户在接受到相应的反馈后,通过说话的方式对智能电子设备进行语音输入。智能电子设备录制下用户的语音内容,根据任务与上下文的不同,再结合自然语言处理技术理解用户的语音输入并完成相应的任务。
最后,用户将智能电子设备移开嘴边,以结束该次语音输入语音,这种结束语音输入的方式非常自然。检测是否移开嘴边的方法与前述检测接近嘴边的方法相似。
前文中将手机作为智能电子设备的例子,不过智能电子设备并不局限于此,还可以例如为可穿戴式智能电子手表、智能手环、智能戒指等等。
前文中是描述智能电子设备与嘴边接近时,以智能电子设备是便携式 的手机为例,说明将手机移动到嘴边,不过此为示例,也可以智能电子设备保持不动,用户主动移动将嘴部靠近智能电子设备,例如在用户开车的状态下,假设智能电子设备固定于方向盘,用户可以主动使得嘴部接近智能电子设备。
利用本发明实施例的智能电子设备,具有以下优势中的一个或多个:
1.交互更加自然。将设备放在嘴前即触发语音输入,而不需要用户进行额外的按下按钮等确认操作,符合用户习惯与认知。
2.使用效率更高。用户单手即可使用,无需在不同的用户界面/应用之间切换,也不需按住某个按键,例如在手机的情况下直接抬起手到嘴边就能使用。
3.收音质量高。设备的录音机在用户嘴边,收取的语音输入信号清晰,受环境音的影响较小。
4.高隐私性与社会性。设备在嘴前,则用户只需发出相对较小的声音即可完成高质量的语音输入,对他人的干扰较小,同时具有较好的隐私保护。
以上已经描述了本发明的各实施例,上述说明是示例性的,并非穷尽性的,并且也不限于所披露的各实施例。在不偏离所说明的各实施例的范围和精神的情况下,对于本技术领域的普通技术人员来说许多修改和变更都是显而易见的。因此,本发明的保护范围应该以权利要求的保护范围为准。
Claims (35)
- 一种智能电子设备,包括传感器系统,能够捕捉到从其能判定智能电子设备接近用户嘴部的信号,智能电子设备包括存储器和处理器,存储器上存储有计算机可执行指令,所述计算机可执行指令被处理器执行时可操作来:处理所述信号以确定智能电子设备是否接近用户嘴部,响应于确定自身接近用户嘴部,激活语音输入。
- 根据权利要求1的智能电子设备,所述传感器系统还能够捕捉到从其能判定智能电子设备碰触嘴部附近的脸部部位的信号,智能电子设备处理所述信号以确定是否碰触用户嘴部附近的脸部部位,响应于确定自身碰触用户嘴部附近的脸部部位,确定智能电子设备接近用户嘴部,激活语音输入。
- 根据权利要求1的智能电子设备,在确定智能电子设备与用户脸部的距离处于0~10厘米范围内时,智能电子设备确定自身接近用户嘴部。
- 根据权利要求1-2的智能电子设备,所识别的特定姿态包括下面的一种或者多种:智能电子设备接近用户嘴部,但不碰触用户脸部,接近距离为0~3厘米,智能电子设备接近用户嘴部,但不碰触用户脸部,接近距离为3~10厘米,智能电子设备接近用户嘴部,碰触用户脸部,碰触的嘴部附近的脸部部位为鼻子,智能电子设备接近用户嘴部,碰触用户脸部,碰触的嘴部附近的脸部部位为上嘴唇和鼻子之间的部位,智能电子设备接近用户嘴部,碰触用户脸部,智能电子设备所碰触的嘴部附近的脸部部位为下巴,智能电子设备接近用户嘴部,碰触用户脸部,智能电子设备所碰触的嘴部附近的脸部部位为面颊。
- 根据权利要求1-4的智能电子设备,还包括:通过所述传感器信号,识别智能电子设备接近用户嘴部的姿态;响应于识别的特定姿态,智能电子设备以特定的方式处理语音输入。
- 根据权利要求1的智能电子设备,还包括:响应于确定自身接近用户嘴部,智能电子设备提供图像、声音或者触觉反馈中的至少一种,提示用户进行语音输入。
- 根据权利要求1的智能电子设备,在激活语音输入之后,处理所述信号以确定用户嘴部是否离开智能电子设备;响应于确定用户嘴部离开智能电子设备,结束语音输入。
- 根据权利要求1的智能电子设备,智能电子设备处理所述信号以确定智能电子设备是否接近用户嘴部包括:计算智能电子设备接近用户嘴部的概率;将所述概率与预定概率阈值相比较,当所述概率大于等于预定概率阈值时,确定智能电子设备接近用户嘴部。
- 根据权利要求1的智能电子设备,所述响应于确定自身接近用户嘴部,激活语音输入包括:结合智能电子便携设备自身、用户与环境的情况,判断是否激活语音输入。
- 根据权利要求9的智能电子设备,所述结合智能电子便携设备自身、用户与环境的情况,判断是否激活语音输入包括:由声纹判定用户是否是特定的授权用户,在判定用户是特定的授权用户的情况下,激活语音输入。
- 根据权利要求1的智能电子设备,传感器系统包括普通摄像头。
- 根据权利要求1的智能电子设备,传感器系统包括红外摄像头。
- 根据权利要求1的智能电子设备,传感器系统包括深度摄像头。
- 根据权利要求1的智能电子设备,传感器系统包接近光传感器。
- 根据权利要求1的智能电子设备,传感器系统包括距离传感器。
- 根据权利要求1的智能电子设备,传感器系统包括广角摄像头。
- 根据权利要求1的智能电子设备,传感器系统包括电容感应传感器。
- 根据权利要求1的智能电子设备,传感器系统包括运动传感器。
- 根据权利要求1的智能电子设备,所述传感器系统包括摄像头,智能电子设备分析摄像头采集的图像信号,检测图像中是否存在近距离拍摄嘴部附近的脸部部位特征,识别智能电子设备是否接近嘴部。
- 根据权利要求19的智能电子设备,传感器系统还包括距离传感器,通过距离传感器信号检测智能电子设备与用户脸部的距离。
- 根据权利要求19的智能电子设备,传感器系统还包括接近光传感器,通过接近光传感器信号识别智能电子设备是否接近用户脸部。
- 根据权利要求1的智能电子设备,传感器系统还包括电容感应传感器,通过电容感应传感器信号识别智能电子设备是否碰触用户脸部。
- 根据权利要求22的智能电子设备,其中智能电子设备通过电容感应传感器信号识别碰触的用户脸部部位。
- 根据权利要求1的智能电子设备,智能手机上传感器系统包括加速度计、陀螺仪和接近光传感器;接近光传感器识别智能电子设备前方被遮挡,基于此前的加速度计和陀螺仪的信号识别用户将手机放到嘴边的动作,而非放置到耳朵边的动作。
- 一种智能电子设备的语音输入触发方法,智能电子设备包括传感器系统,能够捕捉到从其能判定智能电子设备接近用户嘴部的信号,所述语音输入触发方法包括:处理所述信号以确定智能电子设备是否接近用户嘴部,响应于确定自身接近用户嘴部,激活语音输入。
- 根据权利要求25的语音输入触发方法,所述传感器系统还能够捕捉到从其能判定智能电子设备碰触嘴部附近的脸部部位的信号,智能电子设备处理所述信号以确定是否碰触用户嘴部附近的脸部部位,响应于确定自身碰触用户嘴部附近的脸部部位,确定智能电子设备接近用户嘴部,激活语音输入。
- 根据权利要求25的语音输入触发方法,在确定智能电子设备与用户脸部的距离处于0~10厘米范围内时,智能电子设备确定自身接近用户嘴部。
- 根据权利要求25-26的语音输入触发方法,所识别的特定姿态包括下面的一种或者多种:智能电子设备接近用户嘴部,但不碰触用户脸部,接近距离为0~3厘米,智能电子设备接近用户嘴部,但不碰触用户脸部,接近距离为3~10厘米,智能电子设备接近用户嘴部,碰触用户脸部,碰触的嘴部附近的脸部部位为鼻子,智能电子设备接近用户嘴部,碰触用户脸部,碰触的嘴部附近的脸部部位为上嘴唇和鼻子之间的部位,智能电子设备接近用户嘴部,碰触用户脸部,智能电子设备所碰触的嘴部附近的脸部部位为下巴,智能电子设备接近用户嘴部,碰触用户脸部,智能电子设备所碰触的嘴部附近的脸部部位为面颊。
- 根据权利要求25-28的语音输入触发方法,还包括:通过所述传感器信号,识别智能电子设备接近用户嘴部的姿态;响应于识别的特定姿态,智能电子设备以特定的方式处理语音输入。
- 根据权利要求25的语音输入触发方法,还包括:响应于确定自身接近用户嘴部,智能电子设备提供图像、声音或者触觉反馈中的至少一种,提示用户进行语音输入。
- 根据权利要求25的语音输入触发方法,还包括在激活语音输入之后,处理所述信号以确定用户嘴部是否离开智能电子设备;响应于确定用户嘴部离开智能电子设备,结束语音输入。
- 根据权利要求25的语音输入触发方法,智能电子设备处理所述信号以确定智能电子设备是否接近用户嘴部包括:计算智能电子设备接近用户嘴部的概率;将所述概率与预定概率阈值相比较,当所述概率大于等于预定概率阈值时,确定智能电子设备接近用户嘴部。
- 根据权利要求25的语音输入触发方法,所述响应于确定自身接近用户嘴部,激活语音输入包括:结合智能电子便携设备自身、用户与环境的情况,判断是否激活语音输入。
- 根据权利要求33的语音输入触发方法,所述结合智能电子便携设备自身、用户与环境的情况,判断是否激活语音输入包括:由声纹判定用户是否是特定的授权用户,在判定用户是特定的授权用户的情况下,激活语音输入。
- 一种计算机可读存储介质,其上存储有计算机可读指令,所述指令当被计算机执行时,可操作来执行权利要求25-34任一项所述的方法。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN201910476243.5A CN110262767B (zh) | 2019-06-03 | 2019-06-03 | 基于靠近嘴部检测的语音输入唤醒装置、方法和介质 |
| CN201910476243.5 | 2019-06-03 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2020244401A1 true WO2020244401A1 (zh) | 2020-12-10 |
Family
ID=67916429
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2020/092066 Ceased WO2020244401A1 (zh) | 2019-06-03 | 2020-05-25 | 基于靠近嘴部检测的语音输入唤醒装置、方法和介质 |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN110262767B (zh) |
| WO (1) | WO2020244401A1 (zh) |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2024258424A1 (en) * | 2023-06-15 | 2024-12-19 | Google Llc | Selectively invoking an automated assistant according to a result of shape detection at a capacitive array |
Families Citing this family (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN110262767B (zh) * | 2019-06-03 | 2022-03-11 | 交互未来(北京)科技有限公司 | 基于靠近嘴部检测的语音输入唤醒装置、方法和介质 |
| CN114283798A (zh) * | 2021-07-15 | 2022-04-05 | 海信视像科技股份有限公司 | 手持设备的收音方法及手持设备 |
| CN117746849B (zh) * | 2022-09-14 | 2025-10-03 | 荣耀终端股份有限公司 | 一种语音交互方法、装置及终端 |
Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN105869639A (zh) * | 2016-03-21 | 2016-08-17 | 广东小天才科技有限公司 | 一种语音识别的方法及系统 |
| CN105991827A (zh) * | 2015-02-11 | 2016-10-05 | 中兴通讯股份有限公司 | 呼叫处理方法及装置 |
| EP3139246A1 (en) * | 2014-05-16 | 2017-03-08 | ZTE Corporation | Control method and apparatus, electronic device, and computer storage medium |
| CN110262767A (zh) * | 2019-06-03 | 2019-09-20 | 清华大学 | 基于靠近嘴部检测的语音输入唤醒装置、方法和介质 |
Family Cites Families (12)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN103516861B (zh) * | 2012-06-15 | 2015-05-27 | 国基电子(上海)有限公司 | 手持设备及其自动接听电话及锁定触摸屏的控制方法 |
| JP6393021B2 (ja) * | 2012-08-28 | 2018-09-19 | 京セラ株式会社 | 電子機器、制御方法、及び制御プログラム |
| CN104007809B (zh) * | 2013-02-27 | 2017-09-01 | 联想(北京)有限公司 | 一种控制方法及电子设备 |
| KR102129786B1 (ko) * | 2013-04-03 | 2020-07-03 | 엘지전자 주식회사 | 단말기 및 이의 제어방법 |
| EP2801974A3 (en) * | 2013-05-09 | 2015-02-18 | DSP Group Ltd. | Low power activation of a voice activated device |
| US10228904B2 (en) * | 2014-11-12 | 2019-03-12 | Lenovo (Singapore) Pte. Ltd. | Gaze triggered voice recognition incorporating device velocity |
| CN104657105B (zh) * | 2015-01-30 | 2016-10-26 | 腾讯科技(深圳)有限公司 | 一种开启终端的语音输入功能的方法和装置 |
| CN104978165A (zh) * | 2015-06-23 | 2015-10-14 | 上海卓易科技股份有限公司 | 一种语音信息的处理方法、系统及电子设备 |
| CN106598268B (zh) * | 2016-11-10 | 2019-01-11 | 清华大学 | 文本输入方法和电子设备 |
| CN106933483A (zh) * | 2017-02-28 | 2017-07-07 | 清华大学 | 一种可以感知用户感受的触摸交互方式 |
| CN108510986A (zh) * | 2018-03-07 | 2018-09-07 | 北京墨丘科技有限公司 | 语音交互方法、装置、电子设备及计算机可读存储介质 |
| CN109584879B (zh) * | 2018-11-23 | 2021-07-06 | 华为技术有限公司 | 一种语音控制方法及电子设备 |
-
2019
- 2019-06-03 CN CN201910476243.5A patent/CN110262767B/zh active Active
-
2020
- 2020-05-25 WO PCT/CN2020/092066 patent/WO2020244401A1/zh not_active Ceased
Patent Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| EP3139246A1 (en) * | 2014-05-16 | 2017-03-08 | ZTE Corporation | Control method and apparatus, electronic device, and computer storage medium |
| CN105991827A (zh) * | 2015-02-11 | 2016-10-05 | 中兴通讯股份有限公司 | 呼叫处理方法及装置 |
| CN105869639A (zh) * | 2016-03-21 | 2016-08-17 | 广东小天才科技有限公司 | 一种语音识别的方法及系统 |
| CN110262767A (zh) * | 2019-06-03 | 2019-09-20 | 清华大学 | 基于靠近嘴部检测的语音输入唤醒装置、方法和介质 |
Cited By (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2024258424A1 (en) * | 2023-06-15 | 2024-12-19 | Google Llc | Selectively invoking an automated assistant according to a result of shape detection at a capacitive array |
| US12430152B2 (en) | 2023-06-15 | 2025-09-30 | Google Llc | Selectively invoking an automated assistant according to a result of shape detection at a capacitive array |
Also Published As
| Publication number | Publication date |
|---|---|
| CN110262767A (zh) | 2019-09-20 |
| CN110262767B (zh) | 2022-03-11 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US12112756B2 (en) | Voice interaction wakeup electronic device, method and medium based on mouth-covering action recognition | |
| KR102181588B1 (ko) | 동작-음성의 다중 모드 명령에 기반한 최적 제어 방법 및 이를 적용한 전자 장치 | |
| EP3387628B1 (en) | Apparatus, system, and methods for interfacing with a user and/or external apparatus by stationary state detection | |
| JP5998861B2 (ja) | 情報処理装置、情報処理方法及びプログラム | |
| CN109712621B (zh) | 一种语音交互控制方法及终端 | |
| CN110262767B (zh) | 基于靠近嘴部检测的语音输入唤醒装置、方法和介质 | |
| US20130211843A1 (en) | Engagement-dependent gesture recognition | |
| US20150123919A1 (en) | Information input apparatus, information input method, and computer program | |
| US20150077381A1 (en) | Method and apparatus for controlling display of region in mobile device | |
| US20200286484A1 (en) | Methods and systems for speech detection | |
| KR20190022109A (ko) | 음성 인식 서비스를 활성화하는 방법 및 이를 구현한 전자 장치 | |
| JP2015532559A5 (zh) | ||
| JPH0981309A (ja) | 入力装置 | |
| CN112634895A (zh) | 语音交互免唤醒方法和装置 | |
| EP4004908B1 (en) | Activating speech recognition | |
| CN111580656B (zh) | 可穿戴设备及其控制方法、装置 | |
| CN108549802A (zh) | 一种基于人脸识别的解锁方法、装置以及移动终端 | |
| CN105869639A (zh) | 一种语音识别的方法及系统 | |
| CN105843400A (zh) | 一种体感交互方法及装置、穿戴设备 | |
| WO2022083486A1 (zh) | 用于识别手势的方法、电子电路、电子设备和介质 | |
| CN108133708B (zh) | 一种语音助手的控制方法、装置及移动终端 | |
| WO2017165023A1 (en) | Under-wrist mounted gesturing | |
| US20140297257A1 (en) | Motion sensor-based portable automatic interpretation apparatus and control method thereof | |
| CN110519517A (zh) | 临摹引导方法、电子设备及计算机可读存储介质 | |
| KR20140122498A (ko) | 사용자 인터페이스 인식 장치 및 방법 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 20818603 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 20818603 Country of ref document: EP Kind code of ref document: A1 |