WO2026007709A1 - 语音唤醒方法、装置、电子设备以及存储介质 - Google Patents
语音唤醒方法、装置、电子设备以及存储介质Info
- Publication number
- WO2026007709A1 WO2026007709A1 PCT/CN2025/102042 CN2025102042W WO2026007709A1 WO 2026007709 A1 WO2026007709 A1 WO 2026007709A1 CN 2025102042 W CN2025102042 W CN 2025102042W WO 2026007709 A1 WO2026007709 A1 WO 2026007709A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- model
- feature extraction
- speech feature
- extraction sub
- speech
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/02—Feature extraction for speech recognition; Selection of recognition unit
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/06—Creation of reference templates; Training of speech recognition systems, e.g. adaptation to the characteristics of the speaker's voice
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/08—Speech classification or search
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/08—Speech classification or search
- G10L15/16—Speech classification or search using artificial neural networks
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/20—Speech recognition techniques specially adapted for robustness in adverse environments, e.g. in noise, of stress induced speech
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/22—Procedures used during a speech recognition process, e.g. man-machine dialogue
Definitions
- This disclosure relates to a voice wake-up method, apparatus, electronic device, and storage medium.
- Wake-up methods include touch wake-up (such as the lock screen button), timed wake-up (such as an alarm clock), and passive wake-up (such as a phone call).
- touch wake-up such as the lock screen button
- timed wake-up such as an alarm clock
- passive wake-up such as a phone call
- voice wake-up technology is increasingly being adopted to wake devices via voice commands, switching them from sleep to active state.
- existing voice wake-up solutions struggle to effectively address wake-up issues in high-noise environments. Specifically, in noisy environments such as subways, airports, and shopping malls, it is difficult to detect wake-up keywords, resulting in a low wake-up rate.
- This disclosure provides a voice wake-up method, apparatus, electronic device, and storage medium to solve the problem of low wake-up rate caused by difficulty in detecting wake-up keywords in high-noise environments.
- embodiments of this disclosure provide a voice wake-up method, the method comprising:
- the voice signal to be processed acquired by the target device is determined, wherein the voice signal to be processed supports carrying preset wake-up keywords for waking up the target device;
- a reference speech feature extraction sub-model associated with the target device is determined, and speech feature information to be processed is extracted from the speech signal to be processed through the reference speech feature extraction sub-model.
- the reference speech feature extraction sub-model is obtained by adjusting a candidate speech feature extraction sub-model that has been pre-trained by self-supervised training, and the reference speech feature extraction sub-model has inherited the anti-noise interference performance of the candidate speech feature extraction sub-model that has been pre-trained by self-supervised training when performing speech feature extraction.
- the target device or the target application is woken up based on the voice feature information to be processed.
- embodiments of this disclosure also provide a voice wake-up device, the device comprising:
- the first determining module is used to determine the voice signal to be processed acquired by the target device, wherein the voice signal to be processed supports carrying a preset wake-up keyword for waking up the target device or a target application associated with the target device;
- the second determining module is used to determine the reference speech feature extraction sub-model associated with the target device, and extract speech feature information to be processed from the speech signal to be processed through the reference speech feature extraction sub-model.
- the reference speech feature extraction sub-model is obtained by adjusting the candidate speech feature extraction sub-model based on the pre-self-supervised training, and the reference speech feature extraction sub-model has inherited the anti-noise interference performance of the pre-self-supervised training candidate speech feature extraction sub-model when performing speech feature extraction.
- the control module is used to perform wake-up control on the target device or the target application based on the voice feature information to be processed.
- embodiments of this disclosure also provide an electronic device, the electronic device comprising:
- One or more processors are One or more processors;
- Storage device for storing one or more programs.
- the one or more processors When the one or more programs are executed by the one or more processors, the one or more processors implement the voice wake-up method as provided in any embodiment of this disclosure.
- embodiments of this disclosure also provide a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the voice wake-up method as provided in any embodiment of this disclosure.
- Figure 1 is a schematic flowchart of a voice wake-up method provided in an embodiment of this disclosure
- Figure 2 is a schematic diagram of a voice wake-up device provided in an embodiment of this disclosure.
- Figure 3 is a schematic diagram of the structure of an electronic device that implements a voice wake-up method according to an embodiment of this disclosure.
- a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information.
- This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
- sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format.
- the pop-up window can also include a selection control allowing the user to choose "agree” or “disagree” to provide personal information to the electronic device.
- FIG. 1 is a flowchart illustrating a voice wake-up method provided in an embodiment of this disclosure. This embodiment is applicable to situations where a device is woken up by voice to switch from a sleep state to a working state.
- the voice wake-up method can be executed by a voice wake-up device, which can be implemented in software and/or hardware and is generally integrated into any electronic device with network communication capabilities, such as a mobile terminal, a PC, or a server.
- the voice wake-up method of this disclosure embodiment may include the following process:
- S110 Determine the voice signal to be processed acquired by the target device, wherein the voice signal to be processed supports carrying preset wake-up keywords for waking up the target device or a target application associated with the target device.
- the target device can be a device that switches from sleep mode to working mode by a specific voice command containing a preset wake-up keyword.
- the target device When in standby, sleep, or inactive mode, the target device needs to receive a specific voice command containing a preset wake-up keyword to start or activate its main functions.
- a smart speaker is in a low-power standby mode when not woken up. Only when the user says a preset wake-up word will it begin to respond to and execute subsequent voice commands from the user, such as playing music or checking the weather.
- some smart home appliances such as smart TVs or smart air conditioners, can also be set to a voice-activated mode. Only after receiving a correct voice wake-up command containing a preset wake-up keyword will it switch from sleep mode to working mode, ready to receive further operation commands.
- the target device can be a mobile phone or other terminal device, which may have a target application (such as a voice assistant) installed that can be launched and controlled via voice commands.
- This target application can be activated by a wake word received by the mobile phone or other terminal device.
- the target device can be headphones connected to a mobile phone or other smart device, which may have a target application (such as a voice assistant) installed that can be launched via voice commands. The user can activate the voice assistant by inputting a wake word through the headphones, input various voice commands through the headphones, and receive voice responses from the voice assistant through the headphones.
- the voice signal to be processed can include sound information acquired within a preset distance range around the target device's location.
- This sound information contains various voice content, such as speech from a sound source near the target device, a mixture of noise and speech in the target device's environment, or preset keywords related to waking up the target device or its associated application. Therefore, the voice signal to be processed supports carrying preset wake-up keywords for waking up the target device.
- These preset wake-up keywords can be specific voice commands or phrases used to trigger the device to switch from sleep or standby mode to working mode, preparing to receive and process subsequent voice commands or perform related operations.
- the target device or its associated application receives the preset wake-up keyword via voice, it can switch from sleep mode to working mode.
- the speech signal to be processed can encompass sound-related information collected within a preset distance range around the target device's location in a reference noise environment.
- the reference noise environment specifically refers to the environment where the sound source is located, where the noise intensity is relatively high, such as a noise loudness greater than a preset noise intensity.
- speech is often significantly interfered with, resulting in a complex and noisy overall acoustic environment. This inevitably makes the key detection of speech signals in noisy environments more difficult, thus negatively impacting the success rate of voice wake-up and reducing the voice wake-up rate.
- determining the voice signal to be processed acquired by the target device includes the following steps A1-A2:
- Step A1 Use the microphone configured on the target device to acquire the voice signal within a preset distance range of the target device in real time.
- Step A2 The real-time acquired speech signal is preprocessed to obtain the speech signal to be processed.
- the preprocessing includes at least one of noise removal and filtering.
- the microphone configured on the target device can be a device capable of converting sound from the surrounding environment of the target device's location into electrical signals.
- the microphone on the target device has specific sensitivity and receiving range, and can be used to collect sound information within a preset distance range around the target device's location.
- at least one preprocessing method such as noise removal and filtering, can be applied to the real-time acquired speech signal to obtain the speech signal to be processed.
- the preset distance range can be a pre-defined distance area; for example, if the preset distance range is 3 meters, the microphone will collect sound within a 3-meter radius centered on the target device.
- the microphone configured on the target device will continuously and instantly convert the sound information generated within the preset distance range of the target device's location into electrical signals and transmit them to the target device or the target application associated with the target device for subsequent analysis, processing, or execution of corresponding operations, such as waking up the device. This enables timely acquisition of voice information within a specific distance around the target device, providing data support for various voice-related functions of the device.
- S120 Determine the reference speech feature extraction sub-model associated with the target device, and extract the speech feature information to be processed from the speech signal to be processed through the reference speech feature extraction sub-model.
- the reference speech feature extraction sub-model is obtained by adjusting the candidate speech feature extraction sub-model based on the pre-self-supervised training, and the reference speech feature extraction sub-model has inherited the anti-noise interference performance of the pre-self-supervised training candidate speech feature extraction sub-model when performing speech feature extraction.
- a reference speech feature extraction sub-model can be a model used to extract speech feature information from a speech signal to be processed.
- the reference speech feature extraction sub-model is obtained by adjusting a pre-trained, self-supervised candidate speech feature extraction sub-model, ensuring that the adjustment process is achieved through adjustments to the candidate sub-model.
- the reference speech feature extraction sub-model automatically captures speech features from the speech signal, gradually inheriting the noise resistance performance of the pre-trained, self-supervised candidate sub-model during speech feature extraction.
- the speech feature information extracted by the reference speech feature extraction sub-model includes, but is not limited to, the speech's spectral characteristics, temporal features, and acoustic features.
- a specific reference speech feature extraction sub-model associated with the target device is obtained.
- This reference speech feature extraction sub-model is then used to process the speech signal, thereby extracting relevant speech feature information.
- the reference speech feature extraction sub-model is not generated out of thin air, but is obtained by adjusting a pre-trained candidate speech feature extraction sub-model through self-supervised training.
- the candidate speech feature extraction sub-model is adjusted based on a pre-trained self-supervised candidate speech feature extraction sub-model.
- the candidate speech feature extraction sub-model By adjusting the candidate speech feature extraction sub-model, it can be better adapted to the environment and noise conditions of the target device.
- the parameters of the candidate speech feature extraction sub-model can be fine-tuned according to the type and intensity of noise to improve its anti-interference ability in noisy environments. This ensures that the reference speech feature extraction sub-model inherits the ability of the pre-trained self-supervised candidate speech feature extraction sub-model to resist noise interference when extracting speech features.
- the candidate speech feature extraction sub-model can extract speech features relatively accurately even in noisy conditions.
- this reference speech feature extraction sub-model inherits the ability of the pre-self-supervised candidate speech feature extraction sub-model to resist noise interference in speech feature extraction. This means that even if the speech signal to be processed is in a noisy environment, the reference speech feature extraction sub-model can still extract useful speech features relatively accurately, reducing the adverse effects of noise on feature extraction, thereby providing a more reliable and effective foundation for subsequent speech processing tasks (such as speech recognition, speech understanding, etc.).
- the process of determining the reference speech feature extraction sub-model in this embodiment includes the following steps B1-B2:
- Step B1 Determine the candidate speech feature extraction sub-model corresponding to the target device.
- the candidate speech feature extraction sub-model is obtained by training and adjusting the speech feature extraction sub-model on the candidate speech dataset using a self-supervised learning method based on mask prediction.
- the amount of data in the candidate speech dataset is greater than the preset amount of data.
- the candidate speech feature extraction sub-model has stable robustness to noise.
- a large amount of speech data can be collected, including speech signals in various noisy environments and sound information within a preset distance range around the target device's location.
- a speech dataset of millions of hours can be used for training to improve the model's generalization ability and robustness to noise.
- a self-supervised learning model based on masked prediction (such as the pre-training method of HuberT) learns the intrinsic representation of speech by randomly masking a portion of frames in the input speech data and then predicting the features of the masked frames. This is then fine-tuned on a large-scale speech dataset, using optimization algorithms such as stochastic gradient descent.
- the noise resistance performance of the adjusted reference speech feature extraction sub-model is evaluated to ensure it can effectively extract speech features in noisy environments.
- a candidate speech feature extraction model that is robust to noise is obtained and can be used for practical speech processing tasks such as speech recognition, speech enhancement, and voice wake-up.
- these candidate speech feature extraction sub-models are trained by a self-supervised learning method based on mask prediction on a large number of candidate speech datasets.
- the amount of data in the candidate speech dataset needs to be greater than the preset amount of data.
- the model data of the candidate speech feature extraction sub-model is relatively large, which can ensure that the model can learn enough speech features, so that the candidate speech feature extraction sub-model has stable robustness to noise, so as to have the performance of resisting noise interference to a certain extent.
- Step B2 Based on the pre-trained self-supervised candidate speech feature extraction sub-model, perform model compression to obtain the reference speech feature extraction sub-model associated with the target device, so that the reference speech feature extraction sub-model requires less computational resources to extract speech features than the candidate speech feature extraction sub-model.
- model compression After determining the candidate speech feature extraction sub-models, model compression is required.
- the purpose of model compression is to reduce the computational resource requirements of the candidate speech feature extraction sub-models while maintaining their performance.
- Model compression yields a reference speech feature extraction sub-model associated with the target device. This reference sub-model not only inherits the noise immunity performance of the candidate speech feature extraction sub-models but also requires fewer computational resources than the candidate models.
- model compression based on a pre-self-supervised trained candidate speech feature extraction sub-model may include: using model compression methods such as pruning, quantization, low-rank decomposition, and knowledge distillation to compress the model based on the pre-self-supervised trained candidate speech feature extraction sub-model.
- Parameter quantization involves converting model weights from high precision (e.g., 32-bit floating-point numbers) to low precision (e.g., 8-bit integers) to reduce storage space and computational cost. Pruning reduces the number of parameters by removing unimportant connections or neurons. Knowledge distillation transfers knowledge from a complex, large model (teacher model) to a smaller model (student model), enabling the student model to achieve near-teacher model performance on a smaller scale. Low-rank decomposition reduces the number of parameters by decomposing the model's weight matrix into a product of low-rank matrices.
- high precision e.g., 32-bit floating-point numbers
- 8-bit integers e.g. 8-bit integers
- the reference speech feature extraction sub-model due to its meticulous compression processing, can run extremely efficiently on the target device. It accurately extracts speech feature information while significantly reducing computational resource consumption. Compared to the uncompressed original model, its memory requirements and computational load are significantly reduced. This allows for smooth and rapid execution of speech feature extraction tasks on resource-constrained target devices, such as mobile terminals or embedded devices with relatively weak processing power, without compromising the accuracy and reliability of the extracted speech feature information.
- the reference speech feature extraction sub-model supports speech feature extraction under reference noise, which is the ambient noise in which the target device or target application is woken up by voice, and the noise level of the reference noise is greater than the preset noise level.
- the reference speech feature extraction sub-model boasts superior performance in extracting speech features under reference noise conditions.
- Reference noise refers to the ambient noise within a preset distance range surrounding the target device or its associated application when receiving a voice wake-up command, and the noise level exceeds a pre-set noise threshold.
- the fact that the reference speech feature extraction sub-model can perform speech feature extraction under reference noise conditions demonstrates its strong ability to handle severe noise situations. Even in environments with high noise levels, it can still effectively extract valuable speech features. This provides a solid and reliable foundation for subsequent speech processing and analysis, effectively ensuring the smoothness and accuracy of the entire speech processing workflow.
- model compression is performed based on a pre-trained self-supervised candidate speech feature extraction sub-model to obtain a reference speech feature extraction sub-model associated with the target device, including the following steps C1-C2:
- Step C1 Determine the second speech feature extraction sub-model corresponding to the first speech feature extraction sub-model.
- the first speech feature extraction sub-model is a candidate speech feature extraction sub-model that has been pre-trained in a self-supervised manner.
- the second speech feature extraction sub-model is a speech feature extraction sub-model used for speech feature extraction and whose model size is smaller than that of the first speech feature extraction sub-model.
- Step C2 Guide the second speech feature extraction sub-model to learn in the direction of reducing the reference loss during the speech feature extraction training task through the first speech feature extraction sub-model, so as to transfer the noise interference resistance performance of the first speech feature extraction sub-model during speech feature extraction to the second speech feature extraction sub-model.
- the reference loss includes the prediction loss of the second speech feature extraction sub-model in the speech feature extraction training task and the output difference loss between the first speech feature extraction sub-model and the second speech feature extraction sub-model.
- the first speech feature extraction sub-model is a candidate speech feature extraction sub-model obtained in advance through self-supervised training. A second speech feature extraction sub-model with a smaller model size is then found. The purpose of this is to transfer the noise immunity performance of the first speech feature extraction sub-model to the second speech feature extraction sub-model during subsequent training, while reducing computational resource consumption.
- the reference loss includes the prediction loss of the second speech feature extraction sub-model in the speech feature extraction training task and the output difference loss between the first and second speech feature extraction sub-models.
- the first speech feature extraction sub-model is used to guide the training of the second speech feature extraction sub-model, learning its noise-resistant performance during speech feature extraction.
- the second speech feature extraction sub-model can learn how to extract speech features resilient to noise interference, thereby improving speech recognition accuracy in noisy environments.
- the robustness of the larger first speech feature extraction sub-model in speech feature extraction is transferred to the smaller second speech feature extraction sub-model.
- it has lower computational costs and smaller model size, making it easier to deploy and apply on resource-constrained devices.
- the output of the first sub-model (e.g., probability distribution, feature representation) is typically used as a soft objective to guide the training of the second sub-model.
- real labels can be incorporated as a hard objective to achieve a more effective transfer of robustness during speech feature extraction.
- the second speech feature extraction sub-model may approach or even surpass the first speech feature extraction sub-model in terms of accuracy, especially when the amount of data is relatively small or the computing resources are limited.
- a first speech feature extraction sub-model and a second speech feature extraction sub-model are selected: First, a first speech feature extraction sub-model with excellent performance, large scale, and complexity is determined, along with a smaller first speech feature extraction sub-model with lower computational resource requirements. Next, the first speech feature extraction sub-model is fully trained using a large amount of training data to achieve good performance.
- a reference loss in knowledge distillation is defined, including the prediction loss of the second speech feature extraction sub-model when extracting speech features from the speech signal, and the difference loss between the output of the second speech feature extraction sub-model and the output of the first speech feature extraction sub-model (e.g., the difference loss can be based on the difference in probability distributions).
- the second speech feature extraction sub-model is guided by the first speech feature extraction sub-model in the speech feature extraction training task. By adjusting the parameters of the second speech feature extraction sub-model, it is made to simultaneously minimize the prediction loss and the difference loss between the outputs of the two models. Based on the training results, the second speech feature extraction sub-model is fine-tuned, for example, by adjusting the learning rate, the number of training epochs, and hyperparameters, to further improve its performance.
- the second speech feature extraction sub-model is continuously guided by the first speech feature extraction sub-model, and through continuous trial and adjustment, the most suitable parameter settings are found. This allows the second speech feature extraction sub-model to effectively inherit the anti-interference ability of the first speech feature extraction sub-model during speech feature extraction, and achieve a better balance between performance and resource utilization.
- the scale of the speech feature extraction sub-model can be described in detail using at least one of the following dimensions: training data scale, number of model parameters, model computational cost, model memory usage, and model complexity.
- the scale of training data can include the number of samples, feature dimensions, and total data volume (in bytes). For example, if the training data contains billions or even tens of billions of samples, and each sample has thousands of features, then the data scale is considered large.
- the number of model parameters can include the number of learnable parameters in the statistical model. When the number of parameters reaches millions, tens of millions, or even billions, it is generally considered a large-scale model.
- Model computational cost can be measured by the number of floating-point operations required to calculate the model during training or inference. If the computational cost is measured in billions, tens of billions, or even higher orders of magnitude, it indicates a large-scale model.
- Model memory usage can be considered in terms of the memory space required by the model during training or inference.
- Model complexity can be measured by the number of layers, the number of neurons, and the complexity of the network structure. Complex network structures are usually associated with large-scale models.
- the voice feature information to be processed includes not only acoustic features such as prosody and timbre, but also acoustic features used to describe preset wake-up keywords. By identifying the voice feature information to be processed, it is determined whether preset wake-up keywords exist in the voice signal to be processed. Then, it is possible to control the wake-up of the target device or target application by identifying whether preset wake-up keywords can be identified from the voice signal to be processed.
- the reference speech feature extraction sub-model inherits the noise resistance of the pre-self-supervised candidate speech feature extraction sub-model, extracting more accurate and reliable speech features from the speech signal. Even in noisy environments, it effectively captures key speech features, enabling target devices or applications to accurately recognize wake-up commands in various complex noise environments, significantly improving the success rate and stability of device wake-up. Accurately extracted speech features reduce false wake-ups, preventing unnecessary energy consumption and system burden caused by incorrectly identifying environmental noise or irrelevant speech as wake-up commands, thus improving device energy efficiency and operational efficiency.
- wake-up control of the target device or target application based on the voice feature information to be processed includes the following steps D1-D2:
- Step D1 Input the speech feature information to be processed into the reference speech feature detection sub-model associated with the target device.
- the output of the reference speech feature extraction sub-model is connected in series with the input of the reference speech feature detection sub-model.
- the reference speech feature detection sub-model can support the recognition of the speech features output by the reference speech feature extraction sub-model and detect the possibility of the presence of preset wake-up keywords in the speech signal based on the speech features.
- Step D2 Based on the possibility that the speech signal to be processed output by the reference speech feature detection sub-model has preset wake-up keywords, wake-up control is performed on the target device or target application.
- a speech feature detection submodel can be chained after a reference speech feature extraction submodel and trained on a wake-up speech dataset using a classification loss.
- the specific steps are as follows: Collect and organize a training dataset containing preset wake-up keywords.
- a suitable neural network architecture for speech processing can be selected as the speech feature extraction submodel, such as a recurrent neural network (RNN), a long short-term memory network (LSTM), or a convolutional neural network (CNN), and a classification layer can be added after the Hubert model.
- a speech feature detection submodel is chained after the reference speech feature extraction submodel, for example, using a model suitable for classification tasks.
- the model is trained using the prepared dataset, and the parameters are continuously adjusted through backpropagation to minimize the loss function, resulting in the reference speech feature detection submodel.
- the reference speech feature detection sub-model is used to detect the presence of preset wake-up keywords in a speech signal. Its core function is to analyze and process the input speech features to determine whether they contain information related to the wake-up keywords.
- the reference speech feature detection sub-model typically employs a series of algorithms and techniques to perform feature extraction, pattern matching, and classification on the input speech features. By inputting the speech feature information to be processed into the reference speech feature detection sub-model, it can analyze and process this information and output a result representing the probability of the presence of preset wake-up keywords in the speech signal. This result can be used to implement wake-up control for target devices or applications; for example, when the probability exceeds a certain threshold, the wake-up operation of the target device or application is triggered.
- wake-up control is performed on the target device or target application, including the following steps E1-E2:
- Step E1 If the probability of a preset wake-up keyword in the voice signal to be processed is greater than the preset probability threshold, then the target device or target application will be switched from sleep state to working state.
- Step E2 If the probability of a preset wake-up keyword in the voice signal to be processed is not greater than a preset probability threshold, then the target device or target application will not be switched from sleep state to working state.
- the system can accurately identify truly valid wake-up commands, avoid false wake-ups, and improve wake-up accuracy. Switching is only initiated when the probability exceeds the threshold, ensuring that only a second voice signal with a high probability of containing the preset wake-up keyword can wake the device, reducing unnecessary state switching due to misjudgment. Switching is not initiated when the probability is below the threshold; the target device or application will continue to wait for a valid wake-up command, preventing the device from being mistakenly woken from sleep mode due to uncertainty or incorrect judgment, thus saving device resources and energy.
- Determining whether to switch device states based on the probability of the preset wake-up keyword in the voice signal allows for the rational allocation of device resources and energy consumption. This approach ensures timely response to valid user wake-up commands while avoiding the inconvenience of false wake-ups, making the voice wake-up function convenient and comfortable for users.
- the technical solution of this disclosure determines the voice signal to be processed corresponding to the target device, carrying preset wake-up keywords, providing a data basis for subsequent wake-up operations on the target device or target application. Furthermore, by using a reference voice feature extraction sub-model adjusted based on a pre-trained self-supervised candidate voice feature extraction sub-model, feature information is extracted from the voice signal to be processed. Since the reference voice feature extraction sub-model inherits the noise resistance performance of the candidate voice feature extraction sub-model during voice feature extraction, it can effectively extract key features from complex voice signals and reduce noise interference. Therefore, this technical solution significantly improves the ability to detect wake-up keywords in high-noise environments through a feature extraction sub-model with excellent noise resistance, thereby greatly improving the wake-up rate. This allows the target device or target application to be woken up more accurately and reliably in various noisy environments, greatly improving the user experience in complex noise scenarios and enhancing the applicability and practicality of the device.
- FIG. 2 is a schematic diagram of a voice wake-up device provided in an embodiment of this disclosure. This embodiment is applicable to situations where a device is woken up by voice to switch from a sleep state to a working state.
- the voice wake-up device can be implemented in the form of software and/or hardware, and is generally integrated into any electronic device with network communication function, such as a mobile terminal, PC, or server.
- the voice wake-up device of this embodiment may include the following:
- the first determining module 210 is used to determine the voice signal to be processed acquired by the target device, wherein the voice signal to be processed supports carrying a preset wake-up keyword for waking up the target device or a target application associated with the target device.
- the second determining module 220 is used to determine the reference speech feature extraction sub-model associated with the target device, and extract speech feature information to be processed from the speech signal to be processed through the reference speech feature extraction sub-model.
- the reference speech feature extraction sub-model is obtained by adjusting the candidate speech feature extraction sub-model based on the pre-self-supervised training, and the reference speech feature extraction sub-model has inherited the anti-noise interference performance of the pre-self-supervised training candidate speech feature extraction sub-model when performing speech feature extraction.
- the control module 230 is used to perform wake-up control on the target device or the target application based on the voice feature information to be processed.
- determining the voice signal to be processed acquired by the target device includes:
- the microphone configured on the target device is used to acquire the voice signal within a preset distance range of the target device in real time;
- the real-time acquired speech signal is preprocessed to obtain the speech signal to be processed.
- the preprocessing includes noise removal and filtering.
- the process of determining the reference speech feature extraction sub-model includes:
- a candidate speech feature extraction sub-model corresponding to the target device is determined.
- the candidate speech feature extraction sub-model is obtained by training and adjusting the speech feature extraction sub-model on the candidate speech dataset using a self-supervised learning method based on mask prediction.
- the data volume of the candidate speech dataset is greater than a preset data volume, and the candidate speech feature extraction sub-model has stable robustness to noise.
- model compression is performed to obtain a reference speech feature extraction sub-model associated with the target device, so that the reference speech feature extraction sub-model requires less computational resources to perform speech feature extraction than the candidate speech feature extraction sub-model.
- model compression is performed on the pre-self-supervised trained candidate speech feature extraction sub-model to obtain the reference speech feature extraction sub-model associated with the target device, including:
- the first speech feature extraction sub-model is a candidate speech feature extraction sub-model that has been pre-trained by self-supervised training.
- the second speech feature extraction sub-model is a speech feature extraction sub-model that is used to perform speech feature extraction and has a smaller model size than the first speech feature extraction sub-model.
- the first speech feature extraction sub-model guides the second speech feature extraction sub-model to learn in the direction of reducing the reference loss during the speech feature extraction training task, so as to transfer the noise interference resistance performance of the first speech feature extraction sub-model during speech feature extraction to the second speech feature extraction sub-model.
- the reference loss includes the prediction loss of the second speech feature extraction sub-model in the speech feature extraction training task and the output difference loss between the first speech feature extraction sub-model and the second speech feature extraction sub-model.
- the reference speech feature extraction sub-model supports speech feature extraction under reference noise, wherein the reference noise is the ambient noise in which the target device or the target application is woken up by voice, and the noise level of the reference noise is greater than a preset noise level.
- wake-up control of the target device or the target application based on the voice feature information to be processed includes:
- the speech feature information to be processed is input into the reference speech feature detection sub-model associated with the target device.
- the output of the reference speech feature extraction sub-model is connected in series with the input of the reference speech feature detection sub-model.
- the reference speech feature detection sub-model can support the recognition of the speech features of the output of the reference speech feature extraction sub-model and detect the possibility of the existence of preset wake-up keywords in the speech signal based on the speech features.
- wake-up control is performed on the target device or the target application.
- wake-up control is performed on the target device or the target application based on the possibility that the speech signal to be processed output by the reference speech feature detection sub-model contains a preset wake-up keyword, including:
- the target device or the target application will be switched from sleep state to working state.
- the target device or the target application will not be switched from sleep state to working state.
- the technical solution of this disclosure determines the voice signal to be processed corresponding to the target device, carrying preset wake-up keywords, providing a data basis for subsequent wake-up operations on the target device or target application. Furthermore, by using a reference voice feature extraction sub-model adjusted based on a pre-trained self-supervised candidate voice feature extraction sub-model, feature information is extracted from the voice signal to be processed. Since the reference voice feature extraction sub-model inherits the noise resistance performance of the candidate voice feature extraction sub-model during voice feature extraction, it can effectively extract key features from complex voice signals and reduce noise interference. Therefore, this technical solution significantly improves the ability to detect wake-up keywords in high-noise environments through a feature extraction sub-model with excellent noise resistance, thereby greatly improving the wake-up rate. This allows the target device or target application to be woken up more accurately and reliably in various noisy environments, greatly improving the user experience in complex noise scenarios and enhancing the applicability and practicality of the device.
- the voice wake-up device provided in this disclosure can execute the voice wake-up method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects of the execution method.
- FIG. 3 is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure.
- the terminal device in the embodiments of this disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers.
- the electronic device shown in Figure 3 is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this disclosure.
- the electronic device 300 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303.
- the RAM 303 also stores various programs and data required for the operation of the electronic device 300.
- the processing unit 301, ROM 302, and RAM 303 are interconnected via a bus 304.
- An edit/output (I/O) interface 305 is also connected to the bus 304.
- I/O interface 305 I/O interface 305
- input devices 306 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.
- output devices 307 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.
- storage devices 308 including, for example, magnetic tapes, hard disks, etc.
- communication devices 309 allows electronic device 300 to communicate wirelessly or wiredly with other devices to exchange data.
- Figure 3 shows electronic device 300 with various devices, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
- embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts.
- the computer program can be downloaded and installed from a network via communication device 309, or installed from storage device 308, or installed from ROM 302.
- processing device 301 When the computer program is executed by processing device 301, it performs the functions defined in the methods of embodiments of this disclosure.
- the electronic device provided in this embodiment and the voice wake-up method provided in the above embodiments belong to the same inventive concept.
- Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.
- This disclosure provides a computer storage medium storing a computer program that, when executed by a processor, implements the voice wake-up method provided in the above embodiments.
- the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof.
- a computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof.
- a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device.
- a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof.
- a computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.
- the program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
- clients and servers may communicate using any currently known or future-developed network protocol, such as HTTP (Hypertext Transfer Protocol), and may interconnect with digital data communication (e.g., communication networks) of any form or medium.
- network protocols such as HTTP (Hypertext Transfer Protocol)
- communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and end-to-end networks (e.g., ad-hoc end-to-end networks), as well as any currently known or future-developed networks.
- LANs local area networks
- WANs wide area networks
- the Internet e.g., the Internet of Things
- end-to-end networks e.g., ad-hoc end-to-end networks
- the aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
- the aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: determine a speech signal to be processed acquired by a target device, wherein the speech signal to be processed supports carrying a preset wake-up keyword for waking up the target device or a target application associated with the target device; determine a reference speech feature extraction sub-model associated with the target device, and extract speech feature information to be processed from the speech signal to be processed through the reference speech feature extraction sub-model, wherein the reference speech feature extraction sub-model is obtained by adjusting a candidate speech feature extraction sub-model trained in advance under self-supervised training, and the reference speech feature extraction sub-model has inherited the noise interference resistance performance of the candidate speech feature extraction sub-model trained in advance under self-supervised training; and perform wake-up control on the target device or the target application based on the speech feature information to be processed.
- Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages.
- the program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server.
- the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
- LAN local area network
- WAN wide area network
- each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function.
- the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved.
- each block in the block diagrams and/or flowcharts, and combinations of blocks in the block diagrams and/or flowcharts can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
- the units described in the embodiments of this disclosure can be implemented in software or in hardware.
- the name of a unit does not necessarily limit the unit itself; for example, the first acquisition unit can also be described as "a unit that acquires at least two Internet Protocol addresses".
- exemplary types of hardware logic components include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
- FPGAs Field Programmable Gate Arrays
- ASICs Application-Specific Integrated Circuits
- ASSPs Application Standard Products
- SoCs System-on-Chip
- CPLDs Complex Programmable Logic Devices
- a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device.
- a machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium.
- a machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing.
- machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
- RAM random access memory
- ROM read-only memory
- EPROM or flash memory erasable programmable read-only memory
- CD-ROM compact disk read-only memory
- magnetic storage devices or any suitable combination of the foregoing.
Landscapes
- Engineering & Computer Science (AREA)
- Physics & Mathematics (AREA)
- Computational Linguistics (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Human Computer Interaction (AREA)
- Acoustics & Sound (AREA)
- Multimedia (AREA)
- Artificial Intelligence (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Evolutionary Computation (AREA)
- Telephone Function (AREA)
- Measurement Of Mechanical Vibrations Or Ultrasonic Waves (AREA)
Abstract
本公开实施例提供了一种语音唤醒方法、装置、电子设备以及存储介质。该语音唤醒方法包括:确定目标设备获取的待处理语音信号;确定目标设备关联的参考语音特征提取子模型,并通过参考语音特征提取子模型从待处理语音信号中提取待处理语音特征信息,参考语音特征提取子模型是基于预先自监督训练的候选语音特征提取子模型进行调整得到,且参考语音特征提取子模型已承继预先自监督训练的候选语音特征提取子模型在进行语音特征提取时的抗噪声干扰性能;基于待处理语音特征信息对目标设备或目标应用进行唤醒控制。本公开技术方案提升了在高噪声环境下检测唤醒关键词的能力,大大提高了唤醒率。
Description
本申请要求于2024年7月2日递交的中国专利申请第202410883020.1号的优先权,在此全文引用上述中国专利申请公开的内容以作为本申请的一部分。
本公开实施例涉及一种语音唤醒方法、装置、电子设备以及存储介质。
通常情况下,设备需先被唤醒,由休眠状态转变为工作状态,才能够正常处理指令,比如唤醒方式包括触摸唤醒(如锁屏键)、定时唤醒(如闹钟)、被动唤醒(如电话)等。为了更加便捷的进行设备唤醒,逐步采用语音唤醒技术来通过语音方式对设备唤醒,以将设备从休眠状态切换到工作状态。但是,相关语音唤醒方案在高噪声环境下的唤醒问题难以有效解决,具体表现为当处于较大的外部噪声环境下时,例如地铁、机场、商场等,难以检测到唤醒关键词,从而唤醒率偏低。
本公开提供一种语音唤醒方法、装置、电子设备以及存储介质,以解决在高噪声环境下较难检测到唤醒关键词导致唤醒率较低的问题。
第一方面,本公开实施例提供了一种语音唤醒方法,所述方法包括:
确定目标设备获取的待处理语音信号,所述待处理语音信号中支持携带用于对所述目标设备进行唤醒的预设唤醒关键词;
确定目标设备关联的参考语音特征提取子模型,并通过所述参考语音特征提取子模型从所述待处理语音信号中提取待处理语音特征信息,所述参考语音特征提取子模型是基于预先自监督训练的候选语音特征提取子模型进行调整得到,且所述参考语音特征提取子模型已承继所述预先自监督训练的候选语音特征提取子模型在进行语音特征提取时的抗噪声干扰性能;
基于所述待处理语音特征信息对所述目标设备或所述目标应用进行唤醒控制。
第二方面,本公开实施例还提供了一种语音唤醒装置,所述装置包括:
第一确定模块,用于确定目标设备获取的待处理语音信号,所述待处理语音信号中支持携带用于对所述目标设备或与目标设备相关联的目标应用进行唤醒的预设唤醒关键词;
第二确定模块,用于确定目标设备关联的参考语音特征提取子模型,并通过所述参考语音特征提取子模型从所述待处理语音信号中提取待处理语音特征信息,所述参考语音特征提取子模型是基于预先自监督训练的候选语音特征提取子模型进行调整得到,且所述参考语音特征提取子模型已承继所述预先自监督训练的候选语音特征提取子模型在进行语音特征提取时的抗噪声干扰性能;
控制模块,用于基于所述待处理语音特征信息对所述目标设备或所述目标应用进行唤醒控制。
第三方面,本公开实施例还提供了一种电子设备,所述电子设备包括:
一个或多个处理器;
存储装置,用于存储一个或多个程序,
当所述一个或多个程序被所述一个或多个处理器执行,使得所述一个或多个处理器实现如本公开任意实施例提供的语音唤醒方法。
第四方面,本公开实施例还提供了一种包含计算机可执行指令的存储介质,所述计算机可执行指令在由计算机处理器执行时用于执行如本公开任意实施例提供的语音唤醒方法。
应当理解,本部分所描述的内容并非旨在标识本公开的实施例的关键或重要特征,也不用于限制本公开的范围。本公开的其它特征将通过以下的说明书而变得容易理解。
结合附图并参考以下具体实施方式,本公开各实施例的上述和其他特征、优点及方面将变得更加明显。贯穿附图中,相同或相似的附图标记表示相同或相似的元素。应当理解附图是示意性的,原件和元素不一定按照比例绘制。
图1是本公开实施例所提供的一种语音唤醒方法流程示意图;
图2是本公开实施例所提供的一种语音唤醒装置结构示意图;以及
图3是本公开实施例所提供的一种实现语音唤醒方法的电子设备的结构示意图。
下面将参照附图更详细地描述本公开的实施例。虽然附图中显示了本公开的某些实施例,然而应当理解的是,本公开可以通过各种形式来实现,而且不应该被解释为限于这里阐述的实施例,相反提供这些实施例是为了更加透彻和完整地理解本公开。应当理解的是,本公开的附图及实施例仅用于示例性作用,并非用于限制本公开的保护范围。
应当理解,本公开的方法实施方式中记载的各个步骤可以按照不同的顺序执行,和/或并行执行。此外,方法实施方式可以包括附加的步骤和/或省略执行示出的步骤。本公开的范围在此方面不受限制。
本文使用的术语“包括”及其变形是开放性包括,即“包括但不限于”。术语“基于”是“至少部分地基于”。术语“一个实施例”表示“至少一个实施例”;术语“另一实施例”表示“至少一个另外的实施例”;术语“一些实施例”表示“至少一些实施例”。其他术语的相关定义将在下文描述中给出。
需要注意,本公开中提及的“第一”、“第二”等概念仅用于对不同的装置、模块或单元进行区分,并非用于限定这些装置、模块或单元所执行的功能的顺序或者相互依存关系。
需要注意,本公开中提及的“一个”、“多个”的修饰是示意性而非限制性的,本领域技术人员应当理解,除非在上下文另有明确指出,否则应该理解为“一个或多个”。
本公开实施方式中的多个装置之间所交互的消息或者信息的名称仅用于说明性的目的,而并不是用于对这些消息或信息的范围进行限制。
可以理解的是,在使用本公开各实施例公开的技术方案之前,均应当依据相关法律法规通过恰当的方式对本公开所涉及个人信息的类型、使用范围、使用场景等告知用户并获得用户的授权。
例如,在响应于接收到用户的主动请求时,向用户发送提示信息,以明确地提示用户,其请求执行的操作将需要获取和使用到用户的个人信息。从而,使得用户可以根据提示信息来自主地选择是否向执行本公开技术方案的操作的电子设备、应用程序、服务器或存储介质等软件或硬件提供个人信息。
作为一种可选的但非限定性的实现方式,响应于接收到用户的主动请求,向用户发送提示信息的方式例如可以是弹窗的方式,弹窗中可以以文字的方式呈现提示信息。此外,弹窗中还可以承载供用户选择“同意”或者“不同意”向电子设备提供个人信息的选择控件。
可以理解的是,上述通知和获取用户授权过程仅是示意性的,不对本公开的实现方式构成限定,其它满足相关法律法规的方式也可应用于本公开的实现方式中。
可以理解的是,本技术方案所涉及的数据(包括但不限于数据本身、数据的获取或使用)应当遵循相应法律法规及相关规定的要求。
图1为本公开实施例所提供的一种语音唤醒方法的流程示意图,本公开实施例适用于通过语音方式对设备进行唤醒让设备从休眠状态切换到工作状态的情况,该语音唤醒方法可以由语音唤醒装置来执行,该语音唤醒装置可以通过软件和/或硬件的形式进行实现,并一般集成在任何具有网络通信功能的电子设备上,该电子设备可以是移动终端、PC端或服务器等。
如图1所示,本公开实施例的语音唤醒方法可包括以下过程:
S110、确定目标设备获取的待处理语音信号,待处理语音信号中支持携带用于对目标设备或与目标设备相关联的目标应用进行唤醒的预设唤醒关键词。
目标设备可以是指通过包括有预设唤醒关键词的特定语音指令进行语音唤醒实现从休眠状态切换到工作状态的设备,目标设备在处于待机、休眠或未激活状态时,需要通过接收包括有预设唤醒关键词的特定语音指令来启动或激活其主要功能。例如,智能音箱在未被唤醒时处于低功耗的待机状态,只有当用户说出预设的唤醒词后,才会开始响应并执行用户后续的语音指令,如播放音乐、查询天气等。又如,一些智能家电,如智能电视或智能空调,也可以设置为需要语音唤醒的模式,只有接收到正确的包括预设唤醒关键词的语音唤醒指令之后,才会从休眠状态切换到工作状态,准备接收进一步操作指令。
又例如,目标设备可以手机等终端设备,手机等终端设备上可以安装有能够通过语音指令启动、控制的目标应用程序(例如语音助手),该目标应用程序可以通过手机等终端设备接收到的唤醒词进行唤醒。又例如,目标设备可以是连接到手机等智能设备上的耳机,手机等智能设备上可以安装有能够通过语音指令来启动的目标应用程序(例如语音助手),用户可以通过耳机向语音助手输入唤醒词来唤醒语音助手,通过耳机向语音助手输入各种语音指令,并可以通过耳机获得语音助手提供的语音响应。
待处理语音信号可以包括在目标设备所处位置的周边预设距离范围内所获取到的声音信息,待处理语音信号对应的声音信息中包含了各种语音内容,可以是由靠近目标设备的发声源发出的话语、目标设备所处位置环境中的噪音与语音的混合,或者是与对目标设备或与目标设备相关联的目标应用进行唤醒相关的预设关键词等,为此待处理语音信号中支持携带用于对目标设备进行唤醒的预设唤醒关键词。其中,预设唤醒关键词可以是一种特定的语音指令或短语,用于触发设备从休眠或待机状态转换为工作状态,并准备接收和处理后续的语音命令或进行相关操作,当目标设备或与目标设备相关联的目标应用通过语音方式接收到预设唤醒关键词后可以从休眠状态切换到工作状态。
可选地,待处理的语音信号能够涵盖在参考噪声环境下,在目标设备所处位置的周边预设距离范围内所采集获取到的声音相关信息。参考噪声环境具体可以是指发出声音的位置所处环境中,其噪声强度处于相对较高水平的声音环境,比如噪声响度大于预设噪声强度。在这样一种噪声强度颇高的环境当中,说话的声音往往会受到明显的干扰,整体的声学环境呈现出较为繁杂喧闹的状态,这必然会致使对噪声环境中语音信号的关键检测面临更大的难度,进而对语音唤醒的成功率产生不利影响,降低语音唤醒的比率。
作为一种可选的但非限定性的实现方式,确定目标设备获取的待处理语音信号,包括以下步骤A1-A2:
步骤A1、采用目标设备上配置的拾音器对目标设备的预设距离范围内语音信号进行实时获取。
步骤A2、将实时获取的语音信号进行预处理后得到待处理语音信号,预处理包括去除噪声与滤波处理中的至少一项。
目标设备上配置的拾音器可以是一种能够将目标设备所处位置的周围环境中声音转换为电信号的设备,目标设备配置的拾音器具有特定的灵敏度和接收范围,可以利用安装在目标设备上的拾音器采集获取目标设备所处位置的周边预设距离范围内的声音信息。进而,可以对实时采集获取的语音信号进行去除噪声与滤波处理等至少一项预处理方式得到待处理语音信号。
预设距离范围可以是预先设定的一个距离区域;例如,如果预设距离范围是3米,那么拾音器会采集以目标设备为中心,半径3米范围内的声音。目标设备上配置的拾音器会不间断地、即时地将目标设备所在位置的预设距离范围内产生的声音信息转换为电信号并传递给目标设备或与目标设备相关联的目标应用,以便进行后续的分析、处理或执行相应的操作,比如唤醒设备等,实现及时获取目标设备周围特定距离内的语音信息,为设备的各种语音相关功能提供数据支持。
S120、确定目标设备关联的参考语音特征提取子模型,并通过参考语音特征提取子模型从待处理语音信号中提取待处理语音特征信息,参考语音特征提取子模型是基于预先自监督训练的候选语音特征提取子模型进行调整得到,且参考语音特征提取子模型已承继预先自监督训练的候选语音特征提取子模型在进行语音特征提取时的抗噪声干扰性能。
参考语音特征提取子模型可以是一种用于从待处理语音信号中提取语音特征信息的模型,参考语音特征提取子模型是基于预先自监督训练的候选语音特征提取子模型进行调整得到,且保证在调整形成参考语音特征提取子模型的过程中,通过候选语音特征提取子模型进行调整实现,参考语音特征提取子模型自动捕捉语音信号中的语音特征,逐步承继预先自监督训练的候选语音特征提取子模型在进行语音特征提取时的抗噪声干扰性能。其中,参考语音特征提取子模型提取的语音特征信息包括但不限于语音的频谱、时域特征、声学特征等。
针对目标设备获取的待处理语音信号,获取与目标设备相关联的一个特定的参考语音特征提取子模型,利用这个参考语音特征提取子模型来处理待处理语音信号,从而从中提取出有关的待处理语音特征信息。参考语音特征提取子模型并非凭空产生,而是在预先通过自监督训练好的候选语音特征提取子模型的基础上进行调整而得到的。
并且,根据预先自监督训练的候选语音特征提取子模型进行调整得到,通过调整候选语音特征提取子模型,可以使其更好地适应目标设备所处的环境和噪声情况。例如,可以根据噪声的类型和强度,对候选语音特征提取子模型的参数进行微调,以提高其在噪声环境下的抗干扰能力,实现参考语音特征提取子模型继承了预先自监督训练的候选语音特征提取子模型在提取语音特征时抵抗噪声干扰的能力,也就是说候选语音特征提取子模型能够在有噪声的情况下较为准确地提取语音特征。
之所以选择这个参考语音特征提取子模型,是因为参考语音特征提取子模型继承了预先自监督训练的候选语音特征提取子模型在语音特征提取方面抵抗噪声干扰的能力,这意味着即使待处理的语音信号处于有噪声的环境中,该参考语音特征提取子模型也能够相对准确地提取出有用的语音特征,减少噪声对特征提取的不良影响,从而为后续的语音处理任务(如语音识别、语音理解等)提供更可靠和有效的基础。
作为一种可选的但非限定性的实现方式,本公开实施例中参考语音特征提取子模型的确定过程,包括以下步骤B1-B2:
步骤B1、确定目标设备对应的候选语音特征提取子模型,候选语音特征提取子模型是在候选语音数据集上基于掩码预测的自监督学习方式进行语音特征提取子模型的训练调整得到,候选语音数据集的数据量大于预设数据量,候选语音特征提取子模型对噪音具备稳定地鲁棒性。
可选地,收集大量的语音数据,包括在各种噪声环境下的语音信号,以及目标设备所处位置的周边预设距离范围内的声音信息,比如采用百万小时级别的语音数据集进行训练,以提高模型的泛化能力和对噪音的鲁棒性。接着,基于掩码预测的自监督学习模型(比如HuBERT的预训练方式),通过在输入语音数据中随机掩码一部分帧,然后预测被掩码的帧的特征,从而学习到语音的内在表示,并在大规模语音数据集上进行微调训练,比如可以采用随机梯度下降等优化算法进行训练微调,对调整后的参考语音特征提取子模型进行抗噪声干扰性能评估,以确保其在噪声环境下能够有效地提取语音特征。经过多次调整和优化,得到一个对噪音较鲁棒的候选语音特征提取模型,可以用于实际的语音处理任务,如语音识别、语音增强、语音唤醒等。
对于目标设备对应的候选语音特征提取子模型而言,这些候选语音特征提取子模型是通过在大量的候选语音数据集上进行基于掩码预测的自监督学习方式训练得到的,候选语音数据集的数据量需要大于预设数据量,候选语音特征提取子模型的模型数据量比较大,能确保模型能够学习到足够的语音特征,使得候选语音特征提取子模型对噪音具备稳定的鲁棒性,以便在一定程度上具备抵抗噪声干扰的性能。
步骤B2、基于预先自监督训练的候选语音特征提取子模型进行模型压缩,得到目标设备关联的参考语音特征提取子模型,以使参考语音特征提取子模型在进行语音特征提取是所需计算资源少于候选语音特征提取子模型进行语音特征提取是所需计算资源。
在确定候选语音特征提取子模型后,需要对候选语音特征提取子模型进行模型压缩,通过模型压缩的目的是减少候选语音特征提取子模型的计算资源需求,同时保持模型的性能。通过模型压缩可以得到目标设备关联的参考语音特征提取子模型,该参考语音特征提取子模型不仅承继了候选语音特征提取子模型在进行语音特征提取时的抗噪声干扰性能,同时在进行语音特征提取时所需的计算资源也会少于候选语音特征提取子模型。
可选地,基于预先自监督训练的候选语音特征提取子模型进行模型压缩可以包括:使用剪枝、量化、低秩分解、知识蒸馏等模型压缩方式,来基于预先自监督训练的候选语音特征提取子模型进行模型压缩。
其中,参数量化可以是将模型的权重参数从高精度(如32位浮点数)量化为低精度(如8位整数),以减少参数的存储空间和计算量。剪枝可以是通过删除模型中不重要的连接或神经元,减少模型的参数数量。知识蒸馏可以是将复杂大型模型(教师模型)的知识转移到较小的模型(学生模型)中,使学生模型能够在较小规模下达到接近教师模型的性能。低秩分解可以是将模型的权重矩阵分解为低秩矩阵的乘积,降低参数数量。
采用上述方案,参考语音特征提取子模型由于经过了精心的压缩处理,能够在目标设备上以极其高效的方式运行,并且在显著降低了对计算资源的占用的情况下,依然能够精准无误地提取出准确的语音特征信息。与未压缩的原始模型相比,其所需的内存空间大幅缩减,计算量也显著减少,从而使得在资源有限的目标设备上,例如处理能力相对较弱的移动终端或者嵌入式设备中,能够流畅而迅速地执行语音特征提取任务,并且在减少计算资源消耗的同时,丝毫不影响提取语音特征信息的准确性和可靠性。
作为一种可选的但非限定性的实现方式,参考语音特征提取子模型支持在参考噪声下进行语音特征提取,参考噪声为目标设备或目标应用被进行语音唤醒时所处环境噪声,且参考噪声的噪声大小大于预设噪声大小。
参考语音特征提取子模型拥有在参考噪声环境下开展语音特征提取的卓越性能,参考噪声可以指目标设备或与目标设备相关联的目标应用在接受语音唤醒指令时所处的周边预设距离范围内的环境噪声,并且参考噪声的噪声量级已然超越了预先设定的噪声大小阈值。参考语音特征提取子模型支持在参考噪声下进行语音特征提取充分表明,参考语音特征提取子模型具备应对颇为恶劣的噪声状况的强大实力,即便置身于噪声强度颇高的环境之中,依然能够卓有成效地提取出极具价值的语音特征。由此,为后续的语音处理与分析等相关工作构筑了坚实且可靠的基础,有力地保障了整个语音处理流程的顺畅与精准。
作为一种可选的但非限定性的实现方式,基于预先自监督训练的候选语音特征提取子模型进行模型压缩,得到目标设备关联的参考语音特征提取子模型,包括以下步骤C1-C2:
步骤C1、确定第一语音特征提取子模型对应的第二语音特征提取子模型,第一语音特征提取子模型为预先自监督训练的候选语音特征提取子模型,第二语音特征提取子模型为用于进行语音特征提取且模型规模量小于第一语音特征提取子模型的模型规模量的语音特征提取子模型。
步骤C2、通过第一语音特征提取子模型引导第二语音特征提取子模型在进行语音特征提取训练任务中向降低参考损失的方向进行学习,以将第一语音特征提取子模型在进行语音特征提取时的抗噪声干扰性能传递给第二语音特征提取子模型,参考损失包括第二语音特征提取子模型在语音特征提取训练任务中预测损失以及第一语音特征提取子模型与第二语音特征提取子模型之间的输出差异损失。
第一语音特征提取子模型是预先通过自监督训练得到的候选语音特征提取子模型,找到一个模型规模量小于第一语音特征提取子模型的第二语音特征提取子模型。这样做的目的是为了在后续的训练中,将第一语音特征提取子模型的抗噪声干扰性能传递给第二语音特征提取子模型,同时减少计算资源的消耗。
参考损失包括第二语音特征提取子模型在语音特征提取训练任务中的预测损失以及第一语音特征提取子模型与第二语音特征提取子模型之间的输出差异损失。使用第一语音特征提取子模型来引导第二语音特征提取子模型的训练学习第一语音特征提取子模型在语音特征提取时的抗噪声性能,通过最小化参考损失第二语音特征提取子模型可以学习到如何具有抗噪声干扰性能的提取语音特征,从而提高在噪声环境下的语音识别准确率。
将规模较大的第一语音特征提取子模型在语音特征提取时的抗干扰能力转移到规模较小的第二语音特征提取子模型,实现让规模较小、计算资源需求较低的第二语音特征提取子模型能够学习到第一语音特征提取子模型所蕴含的知识和模式,从而在性能上接近甚至达到第一语音特征提取子模型的水平,同时具有更低的计算成本和更小的模型尺寸,便于在资源受限的设备上部署和应用。
在将规模较大的第一语音特征提取子模型在语音特征提取时的抗干扰能力转移到规模较小的第二语音特征提取子模型的过程中,通常会使用第一语音特征提取子模型的输出(如概率分布、特征表示等)作为软目标来指导第二语音特征提取子模型的训练,同时也可以结合真实的标签作为硬目标,以实现更有效的在语音特征提取时的抗干扰能力传递。通过不断调整第二语音特征提取子模型的参数,使其能够拟合第一语音特征提取子模型的输出和真实数据的分布,从而实现知识的迁移和模型的压缩优化。
采用上述过程,如果第一语音特征提取子模型具有丰富而准确的知识,并且蒸馏过程能够有效地将这些知识传递给第二语音特征提取子模型,同时第二语音特征提取子模型具有足够的学习能力来吸收这些知识,那么第二语音特征提取子模型有可能在准确度上接近甚至超过第一语音特征提取子模型,尤其是在数据量相对较少或者计算资源有限的情况下。
示例性地,选择第一语音特征提取子模型和第二语音特征提取子模型:首先确定一个语音特征提取性能优秀、规模较大且复杂的第一语音特征提取子模型,以及一个规模较小、计算资源需求较低的第一语音特征提取子模型。接着,使用大量的训练数据对第一语音特征提取子模型进行充分训练,使其达到较好的性能。定义知识蒸馏中的参考损失,包括第二语音特征提取子模型对语音信号进行语音特征提取时的预测损失,以及第二语音特征提取子模型的输出与第一语音特征提取子模型的输出之间的差异损失(比如差异损失可以基于概率分布的差异)。通过第一语音特征提取子模型引导第二语音特征提取子模型在进行语音特征提取训练任务,通过调整第二语音特征提取子模型的参数,使其同时最小化的预测损失和与两个模型输出的差异损失。根据训练效果,对第二语音特征提取子模型进行微调,例如调整学习率、训练轮数、超参数等,以进一步提高第二语音特征提取子模型的性能。
采用上述方案,通过第一语音特征提取子模型对第二语音特征提取子模型不断进行引导,不断尝试和调整,以找到最适合的参数设置,从而使第二语音特征提取子模型能够有效地从第一语音特征提取子模型中继承在语音特征提取时的抗干扰能力,并在性能和资源利用之间达到较好的平衡。
可选地,语音特征提取子模型的模型规模量可以采用以下至少一个维度指标进行详细描述:训练数据规模、模型参数数量、模型计算量、模型的内存占用量以及模型复杂度等。
其中,训练数据规模可以包括训练数据的样本数量、特征维度、数据总量(以字节为单位)等。例如,如果训练数据包含数十亿甚至数百亿个样本,且每个样本具有数千个特征,那么认为数据规模较大。模型参数数量可以包括统计模型中可学习的参数个数,参数数量达到数百万、数千万甚至数十亿时,通常被认为是大规模模型。模型计算量可以为计算模型在训练或推理过程中所需的浮点运算次数。如果计算量以数十亿、数百亿甚至更高的数量级来衡量,说明模型具有较大规模。模型内存占用可以考虑模型在训练或推理时所需的内存空间。当模型需要大量的内存来存储参数、中间结果和数据时,表明其规模较大。模型复杂度可以为模型的层数、神经元数量、网络结构的复杂程度等。复杂的网络结构通常与大规模模型相关。
S130、基于待处理语音特征信息对目标设备或目标应用进行唤醒控制。
待处理语音特征信息中不仅包括如韵律、音色等声学特征,还包括用于描述预设唤醒关键词的声学特征,通过识别待处理语音特征信息确定待处理语音信号中是否存在预设唤醒关键词,进而可以是否从待处理语音信号中识别到预设唤醒关键词来实现对目标设备或目标应用进行唤醒控制。
参考语音特征提取子模型承继了预先自监督训练的候选语音特征提取子模型的抗噪声干扰性能,从待处理语音信号中提取的待处理语音特征信息更加准确和可靠,即便在嘈杂的环境中也能有效地捕捉到关键的语音特征,这使得目标设备或目标应用在面对各种复杂的噪声环境时,能够准确地识别出唤醒指令,大大提高了设备唤醒的成功率和稳定性。准确提取的待处理语音特征信息能够减少误唤醒的情况发生,避免了设备因错误地将环境噪声或无关语音识别为唤醒指令而造成不必要的能源消耗和系统负担,提高了设备的能效和运行效率。
作为一种可选的但非限定性的实现方式,基于待处理语音特征信息对目标设备或目标应用进行唤醒控制,包括以下步骤D1-D2:
步骤D1、将待处理语音特征信息输入到目标设备关联的参考语音特征检测子模型,参考语音特征提取子模型的输出串联连接参考语音特征检测子模型的输入,参考语音特征检测子模型能支持识别参考语音特征提取子模型的输出的语音特征并基于语音特征检测语音信号中存在预设唤醒关键词的可能性。
步骤D2、根据参考语音特征检测子模型输出的待处理语音信号存在预设唤醒关键词的可能性,对目标设备或目标应用进行唤醒控制。
在参考语音特征提取子模型后面串联一个语音特征检测子模型,可以利用分类损失在唤醒的语音数据集上进行训练。具体步骤如下:收集并整理包含预设唤醒关键词的训练数据集。接着,可以选择适合语音处理的神经网络架构作为语音特征提取子模型,例如循环神经网络(RNN)、长短时记忆网络(LSTM)或卷积神经网络(CNN)等,并在Hubert模型后面添加分类层。在参考语音特征提取子模型后面串联一个语音特征检测子模型,比如使用适合分类任务模型,使用准备好的数据集对模型进行训练,通过反向传播算法不断调整模型的参数,以最小化损失函数,得到参考语音特征检测子模型。
参考语音特征检测子模型可以用于检测语音信号中是否存在预设唤醒关键词的模型,其核心作用在于对输入的语音特征予以分析及处理,从而判定其中是否涵盖与唤醒关键词相关的信息。参考语音特征检测子模型通常会运用一系列的算法与技术,针对输入的语音特征展开特征提取、模式匹配以及分类等操作。通过将待处理语音特征信息输入至参考语音特征检测子模型里,参考语音特征检测子模型能够对这些待处理语音特征信息进行剖析和处置,并输出一个用以表征待处理语音信号中存在预设唤醒关键词可能性的结果。此结果能够用于对目标设备或目标应用实施唤醒控制,譬如当可能性超出一定阈值时,便触发目标设备或目标应用的唤醒操作。
作为一种可选的但非限定性的实现方式,根据参考语音特征检测子模型输出的待处理语音信号存在预设唤醒关键词的可能性,对目标设备或目标应用进行唤醒控制,包括以下步骤E1-E2:
步骤E1、若待处理语音信号中存在预设唤醒关键词的可能性大于预设概率阈值,则启动将目标设备或目标应用从休眠状态切换到工作状态。
步骤E2、若待处理语音信号中存在预设唤醒关键词的可能性不大于预设概率阈值,则不将目标设备或目标应用从休眠状态切换到工作状态。
通过设置预设概率阈值来判断是否启动设备的状态切换,能够准确地识别出真正有效的唤醒指令,避免误唤醒,提高了唤醒的精准性。当可能性大于阈值时才启动切换,确保只有高度可能包含预设唤醒关键词的第二语音信号能唤醒设备,减少因误判导致的不必要的状态切换。当可能性不大于阈值时不启动切换,目标设备或目标应用将继续等待有效的唤醒指令输入,防止因不确定或错误的判断而错误地将设备从休眠状态唤醒,节省设备资源和能源。根据语音信号中存在预设唤醒关键词的可能性来决定是否切换设备状态,能够合理地分配设备的资源和能耗。通过上述方案既能及时响应用户的有效唤醒指令,又能避免因误唤醒带来的困扰,使用户在使用语音唤醒功能时感到便捷和舒适。
本公开实施例的技术方案,确定目标设备对应的携带预设唤醒关键词的待处理语音信号,为后续对目标设备或目标应用的唤醒操作提供数据基础;并且通过使用基于预先自监督训练的候选语音特征提取子模型调整得到的参考语音特征提取子模型,从待处理语音信号中提取特征信息,由于参考语音特征提取子模型承继了候选语音特征提取子模型在语音特征提取时的抗噪声干扰性能,能够有效地从复杂的语音信号中提取出关键特征,降低噪声的干扰,因此本技术方案通过具有优秀抗噪性能的特征提取子模型,显著提升了在高噪声环境下检测唤醒关键词的能力,从而大大提高了唤醒率,使得目标设备或目标应用在各种嘈杂的环境中都能更准确、可靠地被唤醒,极大地改善了用户在复杂噪声场景下的使用体验,增强了设备的适用性和实用性。
图2为本公开实施例所提供的一种语音唤醒装置的结构示意图,本公开实施例适用于通过语音方式对设备进行唤醒让设备从休眠状态切换到工作状态的情况,该语音唤醒装置可以通过软件和/或硬件的形式进行实现,并一般集成在任何具有网络通信功能的电子设备上,该电子设备可以是移动终端、PC端或服务器等。
如图2所示,本公开实施例的语音唤醒装置可包括以下:
第一确定模块210,用于确定目标设备获取的待处理语音信号,所述待处理语音信号中支持携带用于对所述目标设备或与目标设备相关联的目标应用进行唤醒的预设唤醒关键词;
第二确定模块220,用于确定目标设备关联的参考语音特征提取子模型,并通过所述参考语音特征提取子模型从所述待处理语音信号中提取待处理语音特征信息,所述参考语音特征提取子模型是基于预先自监督训练的候选语音特征提取子模型进行调整得到,且所述参考语音特征提取子模型已承继所述预先自监督训练的候选语音特征提取子模型在进行语音特征提取时的抗噪声干扰性能;
控制模块230,用于基于所述待处理语音特征信息对所述目标设备或所述目标应用进行唤醒控制。
在上述实施例的基础上,可选地,确定目标设备获取的待处理语音信号,包括:
采用目标设备上配置的拾音器对所述目标设备的预设距离范围内语音信号进行实时获取;
将实时获取的语音信号进行预处理后得到所述待处理语音信号,所述预处理包括去除噪声与滤波处理。
在上述实施例的基础上,可选地,所述参考语音特征提取子模型的确定过程,包括:
确定目标设备对应的候选语音特征提取子模型,所述候选语音特征提取子模型是在候选语音数据集上基于掩码预测的自监督学习方式进行语音特征提取子模型的训练调整得到,所述候选语音数据集的数据量大于预设数据量,所述候选语音特征提取子模型对噪音具备稳定地鲁棒性;基于所述预先自监督训练的候选语音特征提取子模型进行模型压缩,得到所述目标设备关联的参考语音特征提取子模型,以使所述参考语音特征提取子模型在进行语音特征提取是所需计算资源少于所述候选语音特征提取子模型进行语音特征提取是所需计算资源。
在上述实施例的基础上,可选地,基于所述预先自监督训练的候选语音特征提取子模型进行模型压缩,得到所述目标设备关联的参考语音特征提取子模型,包括:
确定第一语音特征提取子模型对应的第二语音特征提取子模型,所述第一语音特征提取子模型为预先自监督训练的候选语音特征提取子模型,所述第二语音特征提取子模型为用于进行语音特征提取且模型规模量小于所述第一语音特征提取子模型的模型规模量的语音特征提取子模型;
通过所述第一语音特征提取子模型引导所述第二语音特征提取子模型在进行语音特征提取训练任务中向降低参考损失的方向进行学习,以将所述第一语音特征提取子模型在进行语音特征提取时的抗噪声干扰性能传递给所述第二语音特征提取子模型,所述参考损失包括所述第二语音特征提取子模型在语音特征提取训练任务中预测损失以及所述第一语音特征提取子模型与所述第二语音特征提取子模型之间的输出差异损失。
在上述实施例的基础上,可选地,所述参考语音特征提取子模型支持在参考噪声下进行语音特征提取,所述参考噪声为所述目标设备或所述目标应用被进行语音唤醒时所处环境噪声,且所述参考噪声的噪声大小大于预设噪声大小。
在上述实施例的基础上,可选地,基于所述待处理语音特征信息对所述目标设备或所述目标应用进行唤醒控制,包括:
将所述待处理语音特征信息输入到目标设备关联的参考语音特征检测子模型,所述参考语音特征提取子模型的输出串联连接所述参考语音特征检测子模型的输入,所述参考语音特征检测子模型能支持识别所述参考语音特征提取子模型的输出的语音特征并基于语音特征检测语音信号中存在预设唤醒关键词的可能性;
根据所述参考语音特征检测子模型输出的所述待处理语音信号存在预设唤醒关键词的可能性,对所述目标设备或所述目标应用进行唤醒控制。
在上述实施例的基础上,可选地,根据所述参考语音特征检测子模型输出的所述待处理语音信号存在预设唤醒关键词的可能性,对所述目标设备或所述目标应用进行唤醒控制,包括:
若所述待处理语音信号中存在预设唤醒关键词的可能性大于预设概率阈值,则启动将所述目标设备或所述目标应用从休眠状态切换到工作状态;
若所述待处理语音信号中存在预设唤醒关键词的可能性不大于预设概率阈值,则不将所述目标设备或所述目标应用从休眠状态切换到工作状态。
本公开实施例的技术方案,确定目标设备对应的携带预设唤醒关键词的待处理语音信号,为后续对目标设备或目标应用的唤醒操作提供数据基础;并且通过使用基于预先自监督训练的候选语音特征提取子模型调整得到的参考语音特征提取子模型,从待处理语音信号中提取特征信息,由于参考语音特征提取子模型承继了候选语音特征提取子模型在语音特征提取时的抗噪声干扰性能,能够有效地从复杂的语音信号中提取出关键特征,降低噪声的干扰,因此本技术方案通过具有优秀抗噪性能的特征提取子模型,显著提升了在高噪声环境下检测唤醒关键词的能力,从而大大提高了唤醒率,使得目标设备或目标应用在各种嘈杂的环境中都能更准确、可靠地被唤醒,极大地改善了用户在复杂噪声场景下的使用体验,增强了设备的适用性和实用性。
本公开实施例所提供的语音唤醒装置可执行本公开任意实施例所提供的语音唤醒方法,具备执行方法相应的功能模块和有益效果。
值得注意的是,上述装置所包括的各个单元和模块只是按照功能逻辑进行划分的,但并不局限于上述的划分,只要能够实现相应的功能即可;另外,各功能单元的具体名称也只是为了便于相互区分,并不用于限制本公开实施例的保护范围。
图3为本公开实施例所提供的一种电子设备的结构示意图。下面参考图3,其示出了适于用来实现本公开实施例的电子设备(例如图3中的终端设备或服务器)300的结构示意图。本公开实施例中的终端设备可以包括但不限于诸如移动电话、笔记本电脑、数字广播接收器、PDA(个人数字助理)、PAD(平板电脑)、PMP(便携式多媒体播放器)、车载终端(例如车载导航终端)等等的移动终端以及诸如数字TV、台式计算机等等的固定终端。图3示出的电子设备仅仅是一个示例,不应对本公开实施例的功能和使用范围带来任何限制。
如图3所示,电子设备300可以包括处理装置(例如中央处理器、图形处理器等)301,其可以根据存储在只读存储器(ROM)302中的程序或者从存储装置308加载到随机访问存储器(RAM)303中的程序而执行各种适当的动作和处理。在RAM 303中,还存储有电子设备300操作所需的各种程序和数据。处理装置301、ROM 302以及RAM 303通过总线304彼此相连。编辑/输出(I/O)接口305也连接至总线304。
通常,以下装置可以连接至I/O接口305:包括例如触摸屏、触摸板、键盘、鼠标、摄像头、麦克风、加速度计、陀螺仪等的输入装置306;包括例如液晶显示器(LCD)、扬声器、振动器等的输出装置307;包括例如磁带、硬盘等的存储装置308;以及通信装置309。通信装置309可以允许电子设备300与其他设备进行无线或有线通信以交换数据。虽然图3示出了具有各种装置的电子设备300,但是应理解的是,并不要求实施或具备所有示出的装置。可以替代地实施或具备更多或更少的装置。
特别地,根据本公开的实施例,上文参考流程图描述的过程可以被实现为计算机软件程序。例如,本公开的实施例包括一种计算机程序产品,其包括承载在非暂态计算机可读介质上的计算机程序,该计算机程序包含用于执行流程图所示的方法的程序代码。在这样的实施例中,该计算机程序可以通过通信装置309从网络上被下载和安装,或者从存储装置308被安装,或者从ROM 302被安装。在该计算机程序被处理装置301执行时,执行本公开实施例的方法中限定的上述功能。
本公开实施方式中的多个装置之间所交互的消息或者信息的名称仅用于说明性的目的,而并不是用于对这些消息或信息的范围进行限制。
本公开实施例提供的电子设备与上述实施例提供的语音唤醒方法属于同一发明构思,未在本实施例中详尽描述的技术细节可参见上述实施例,并且本实施例与上述实施例具有相同的有益效果。
本公开实施例提供了一种计算机存储介质,其上存储有计算机程序,该程序被处理器执行时实现上述实施例所提供的语音唤醒方法。
需要说明的是,本公开上述的计算机可读介质可以是计算机可读信号介质或者计算机可读存储介质或者是上述两者的任意组合。计算机可读存储介质例如可以是——但不限于——电、磁、光、电磁、红外线、或半导体的系统、装置或器件,或者任意以上的组合。计算机可读存储介质的更具体的例子可以包括但不限于:具有一个或多个导线的电连接、便携式计算机磁盘、硬盘、随机访问存储器(RAM)、只读存储器(ROM)、可擦式可编程只读存储器(EPROM或闪存)、光纤、便携式紧凑磁盘只读存储器(CD-ROM)、光存储器件、磁存储器件、或者上述的任意合适的组合。在本公开中,计算机可读存储介质可以是任何包含或存储程序的有形介质,该程序可以被指令执行系统、装置或者器件使用或者与其结合使用。而在本公开中,计算机可读信号介质可以包括在基带中或者作为载波一部分传播的数据信号,其中承载了计算机可读的程序代码。这种传播的数据信号可以采用多种形式,包括但不限于电磁信号、光信号或上述的任意合适的组合。计算机可读信号介质还可以是计算机可读存储介质以外的任何计算机可读介质,该计算机可读信号介质可以发送、传播或者传输用于由指令执行系统、装置或者器件使用或者与其结合使用的程序。计算机可读介质上包含的程序代码可以用任何适当的介质传输,包括但不限于:电线、光缆、RF(射频)等等,或者上述的任意合适的组合。
在一些实施方式中,客户端、服务器可以利用诸如HTTP(HyperText Transfer Protocol,超文本传输协议)之类的任何当前已知或未来研发的网络协议进行通信,并且可以与任意形式或介质的数字数据通信(例如,通信网络)互连。通信网络的示例包括局域网(“LAN”),广域网(“WAN”),网际网(例如,互联网)以及端对端网络(例如,ad hoc端对端网络),以及任何当前已知或未来研发的网络。
上述计算机可读介质可以是上述电子设备中所包含的;也可以是单独存在,而未装配入该电子设备中。
上述计算机可读介质承载有一个或者多个程序,当上述一个或者多个程序被该电子设备执行时,使得该电子设备:确定目标设备获取的待处理语音信号,所述待处理语音信号中支持携带用于对所述目标设备或与目标设备相关联的目标应用进行唤醒的预设唤醒关键词;确定目标设备关联的参考语音特征提取子模型,并通过所述参考语音特征提取子模型从所述待处理语音信号中提取待处理语音特征信息,所述参考语音特征提取子模型是基于预先自监督训练的候选语音特征提取子模型进行调整得到,且所述参考语音特征提取子模型已承继所述预先自监督训练的候选语音特征提取子模型在进行语音特征提取时的抗噪声干扰性能;基于所述待处理语音特征信息对所述目标设备或所述目标应用进行唤醒控制。
可以以一种或多种程序设计语言或其组合来编写用于执行本公开的操作的计算机程序代码,上述程序设计语言包括但不限于面向对象的程序设计语言—诸如Java、Smalltalk、C++,还包括常规的过程式程序设计语言—诸如“C”语言或类似的程序设计语言。程序代码可以完全地在用户计算机上执行、部分地在用户计算机上执行、作为一个独立的软件包执行、部分在用户计算机上部分在远程计算机上执行、或者完全在远程计算机或服务器上执行。在涉及远程计算机的情形中,远程计算机可以通过任意种类的网络——包括局域网(LAN)或广域网(WAN)—连接到用户计算机,或者,可以连接到外部计算机(例如利用因特网服务提供商来通过因特网连接)。
附图中的流程图和框图,图示了按照本公开各种实施例的系统、方法和计算机程序产品的可能实现的体系架构、功能和操作。在这点上,流程图或框图中的每个方框可以代表一个模块、程序段、或代码的一部分,该模块、程序段、或代码的一部分包含一个或多个用于实现规定的逻辑功能的可执行指令。也应当注意,在有些作为替换的实现中,方框中所标注的功能也可以以不同于附图中所标注的顺序发生。例如,两个接连地表示的方框实际上可以基本并行地执行,它们有时也可以按相反的顺序执行,这依所涉及的功能而定。也要注意的是,框图和/或流程图中的每个方框、以及框图和/或流程图中的方框的组合,可以用执行规定的功能或操作的专用的基于硬件的系统来实现,或者可以用专用硬件与计算机指令的组合来实现。
描述于本公开实施例中所涉及到的单元可以通过软件的方式实现,也可以通过硬件的方式来实现。其中,单元的名称在某种情况下并不构成对该单元本身的限定,例如,第一获取单元还可以被描述为“获取至少两个网际协议地址的单元”。
本文中以上描述的功能可以至少部分地由一个或多个硬件逻辑部件来执行。例如,非限制性地,可以使用的示范类型的硬件逻辑部件包括:现场可编程门阵列(FPGA)、专用集成电路(ASIC)、专用标准产品(ASSP)、片上系统(SOC)、复杂可编程逻辑设备(CPLD)等等。
在本公开的上下文中,机器可读介质可以是有形的介质,其可以包含或存储以供指令执行系统、装置或设备使用或与指令执行系统、装置或设备结合地使用的程序。机器可读介质可以是机器可读信号介质或机器可读储存介质。机器可读介质可以包括但不限于电子的、磁性的、光学的、电磁的、红外的、或半导体系统、装置或设备,或者上述内容的任何合适组合。机器可读存储介质的更具体示例会包括基于一个或多个线的电气连接、便携式计算机盘、硬盘、随机存取存储器(RAM)、只读存储器(ROM)、可擦除可编程只读存储器(EPROM或快闪存储器)、光纤、便捷式紧凑盘只读存储器(CD-ROM)、光学储存设备、磁储存设备、或上述内容的任何合适组合。
以上描述仅为本公开的较佳实施例以及对所运用技术原理的说明。本领域技术人员应当理解,本公开中所涉及的公开范围,并不限于上述技术特征的特定组合而成的技术方案,同时也应涵盖在不脱离上述公开构思的情况下,由上述技术特征或其等同特征进行任意组合而形成的其它技术方案。例如上述特征与本公开中公开的(但不限于)具有类似功能的技术特征进行互相替换而形成的技术方案。
此外,虽然采用特定次序描绘了各操作,但是这不应当理解为要求这些操作以所示出的特定次序或以顺序次序执行来执行。在一定环境下,多任务和并行处理可能是有利的。同样地,虽然在上面论述中包含了若干具体实现细节,但是这些不应当被解释为对本公开的范围的限制。在单独的实施例的上下文中描述的某些特征还可以组合地实现在单个实施例中。相反地,在单个实施例的上下文中描述的各种特征也可以单独地或以任何合适的子组合的方式实现在多个实施例中。
尽管已经采用特定于结构特征和/或方法逻辑动作的语言描述了本主题,但是应当理解所附权利要求书中所限定的主题未必局限于上面描述的特定特征或动作。相反,上面所描述的特定特征和动作仅仅是实现权利要求书的示例形式。
Claims (10)
- 一种语音唤醒方法,包括:确定目标设备获取的待处理语音信号,所述待处理语音信号中支持携带用于对所述目标设备或与目标设备相关联的目标应用进行唤醒的预设唤醒关键词;确定目标设备关联的参考语音特征提取子模型,并通过所述参考语音特征提取子模型从所述待处理语音信号中提取待处理语音特征信息,所述参考语音特征提取子模型是基于预先自监督训练的候选语音特征提取子模型进行调整得到,且所述参考语音特征提取子模型已承继所述预先自监督训练的候选语音特征提取子模型在进行语音特征提取时的抗噪声干扰性能;基于所述待处理语音特征信息对所述目标设备或所述目标应用进行唤醒控制。
- 根据权利要求1所述的方法,其中,所述确定目标设备获取的待处理语音信号,包括:采用目标设备上配置的拾音器对所述目标设备的预设距离范围内语音信号进行实时获取;将实时获取的语音信号进行预处理后得到所述待处理语音信号,所述预处理包括去除噪声与滤波处理中的至少一项。
- 根据权利要求1或2所述的方法,其中,所述参考语音特征提取子模型的确定过程,包括:确定目标设备对应的候选语音特征提取子模型,所述候选语音特征提取子模型是在候选语音数据集上基于掩码预测的自监督学习方式进行语音特征提取子模型的训练调整得到,所述候选语音数据集的数据量大于预设数据量,所述候选语音特征提取子模型对噪音具备稳定地鲁棒性;基于所述预先自监督训练的候选语音特征提取子模型进行模型压缩,得到所述目标设备关联的参考语音特征提取子模型,以使所述参考语音特征提取子模型在进行语音特征提取是所需计算资源少于所述候选语音特征提取子模型进行语音特征提取是所需计算资源。
- 根据权利要求3所述的方法,其中,所述基于所述预先自监督训练的候选语音特征提取子模型进行模型压缩,得到所述目标设备关联的参考语音特征提取子模型,包括:确定第一语音特征提取子模型对应的第二语音特征提取子模型,所述第一语音特征提取子模型为预先自监督训练的候选语音特征提取子模型,所述第二语音特征提取子模型为用于进行语音特征提取且模型规模量小于所述第一语音特征提取子模型的模型规模量的语音特征提取子模型;通过所述第一语音特征提取子模型引导所述第二语音特征提取子模型在进行语音特征提取训练任务中向降低参考损失的方向进行学习,以将所述第一语音特征提取子模型在进行语音特征提取时的抗噪声干扰性能传递给所述第二语音特征提取子模型,所述参考损失包括所述第二语音特征提取子模型在语音特征提取训练任务中预测损失以及所述第一语音特征提取子模型与所述第二语音特征提取子模型之间的输出差异损失。
- 根据权利要求3或4所述的方法,其中,所述参考语音特征提取子模型支持在参考噪声下进行语音特征提取,所述参考噪声为所述目标设备或所述目标应用被进行语音唤醒时所处环境噪声,且所述参考噪声的噪声大小大于预设噪声大小。
- 根据权利要求1至5中任一项所述的方法,其中,所述基于所述待处理语音特征信息对所述目标设备或所述目标应用进行唤醒控制,包括:将所述待处理语音特征信息输入到目标设备关联的参考语音特征检测子模型,所述参考语音特征提取子模型的输出串联连接所述参考语音特征检测子模型的输入,所述参考语音特征检测子模型能支持识别所述参考语音特征提取子模型的输出的语音特征并基于语音特征检测语音信号中存在预设唤醒关键词的可能性;根据所述参考语音特征检测子模型输出的所述待处理语音信号存在预设唤醒关键词的可能性,对所述目标设备或所述目标应用进行唤醒控制。
- 根据权利要求6所述的方法,其中,所述根据所述参考语音特征检测子模型输出的所述待处理语音信号存在预设唤醒关键词的可能性,对所述目标设备或所述目标应用进行唤醒控制,包括:若所述待处理语音信号中存在预设唤醒关键词的可能性大于预设概率阈值,则启动将所述目标设备或所述目标应用从休眠状态切换到工作状态;若所述待处理语音信号中存在预设唤醒关键词的可能性不大于预设概率阈值,则不将所述目标设备或所述目标应用从休眠状态切换到工作状态。
- 一种语音唤醒装置,包括:第一确定模块,被配置为确定目标设备获取的待处理语音信号,所述待处理语音信号中支持携带用于对所述目标设备或与目标设备相关联的目标应用进行唤醒的预设唤醒关键词;第二确定模块,被配置为确定目标设备关联的参考语音特征提取子模型,并通过所述参考语音特征提取子模型从所述待处理语音信号中提取待处理语音特征信息,所述参考语音特征提取子模型是基于预先自监督训练的候选语音特征提取子模型进行调整得到,且所述参考语音特征提取子模型已承继所述预先自监督训练的候选语音特征提取子模型在进行语音特征提取时的抗噪声干扰性能;控制模块,被配置为基于所述待处理语音特征信息对所述目标设备或所述目标应用进行唤醒控制。
- 一种电子设备,包括:一个或多个处理器;存储装置,被配置为存储一个或多个程序,其中,当所述一个或多个程序被所述一个或多个处理器执行时,使得所述一个或多个处理器实现如权利要求1-7中任一项所述的语音唤醒方法。
- 一种包含计算机可执行指令的存储介质,其中,所述计算机可执行指令在由计算机处理器执行时用于执行如权利要求1-7中任一项所述的语音唤醒方法。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202410883020.1A CN121281506A (zh) | 2024-07-02 | 2024-07-02 | 语音唤醒方法、装置、电子设备以及存储介质 |
| CN202410883020.1 | 2024-07-02 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2026007709A1 true WO2026007709A1 (zh) | 2026-01-08 |
Family
ID=98232743
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2025/102042 Pending WO2026007709A1 (zh) | 2024-07-02 | 2025-06-19 | 语音唤醒方法、装置、电子设备以及存储介质 |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN121281506A (zh) |
| WO (1) | WO2026007709A1 (zh) |
Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN110634468A (zh) * | 2019-09-11 | 2019-12-31 | 中国联合网络通信集团有限公司 | 语音唤醒方法、装置、设备及计算机可读存储介质 |
| CN114038457A (zh) * | 2021-11-04 | 2022-02-11 | 北京房江湖科技有限公司 | 用于语音唤醒的方法、电子设备、存储介质和程序 |
| CN115223551A (zh) * | 2021-03-30 | 2022-10-21 | 暗物智能科技(广州)有限公司 | 一种基于语音相似性匹配的语音唤醒方法及系统 |
| US20220366898A1 (en) * | 2021-05-14 | 2022-11-17 | Microsoft Technology Licensing, Llc | Unified speech representation learning |
| CN116343781A (zh) * | 2023-03-23 | 2023-06-27 | 京东科技信息技术有限公司 | 语音识别模型的训练方法及装置、存储介质及电子设备 |
-
2024
- 2024-07-02 CN CN202410883020.1A patent/CN121281506A/zh active Pending
-
2025
- 2025-06-19 WO PCT/CN2025/102042 patent/WO2026007709A1/zh active Pending
Patent Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN110634468A (zh) * | 2019-09-11 | 2019-12-31 | 中国联合网络通信集团有限公司 | 语音唤醒方法、装置、设备及计算机可读存储介质 |
| CN115223551A (zh) * | 2021-03-30 | 2022-10-21 | 暗物智能科技(广州)有限公司 | 一种基于语音相似性匹配的语音唤醒方法及系统 |
| US20220366898A1 (en) * | 2021-05-14 | 2022-11-17 | Microsoft Technology Licensing, Llc | Unified speech representation learning |
| CN114038457A (zh) * | 2021-11-04 | 2022-02-11 | 北京房江湖科技有限公司 | 用于语音唤醒的方法、电子设备、存储介质和程序 |
| CN116343781A (zh) * | 2023-03-23 | 2023-06-27 | 京东科技信息技术有限公司 | 语音识别模型的训练方法及装置、存储介质及电子设备 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN121281506A (zh) | 2026-01-06 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US11244672B2 (en) | Speech recognition method and apparatus, and storage medium | |
| CN110288979B (zh) | 一种语音识别方法及装置 | |
| CN112735418B (zh) | 一种语音交互的处理方法、装置、终端及存储介质 | |
| CN107481718B (zh) | 语音识别方法、装置、存储介质及电子设备 | |
| WO2021093449A1 (zh) | 基于人工智能的唤醒词检测方法、装置、设备及介质 | |
| CN110570840B (zh) | 一种基于人工智能的智能设备唤醒方法和装置 | |
| US11393490B2 (en) | Method, apparatus, device and computer-readable storage medium for voice interaction | |
| CN111326146A (zh) | 语音唤醒模板的获取方法、装置、电子设备及计算机可读存储介质 | |
| CN111522592A (zh) | 一种基于人工智能的智能终端唤醒方法和装置 | |
| WO2025064046A1 (en) | Low power always-on listening artificial intelligence (ai) system | |
| CN114360510A (zh) | 一种语音识别方法和相关装置 | |
| WO2026007709A1 (zh) | 语音唤醒方法、装置、电子设备以及存储介质 | |
| WO2021146661A2 (en) | Systems and methods for generating wake signals from known users | |
| CN113761952B (zh) | 一种文本翻译方法和相关装置 | |
| CN115995014A (zh) | 一种喇叭单体的检测方法、音频检测的方法以及相关装置 | |
| CN117012184A (zh) | 一种语音识别方法和相关装置 | |
| US20260088019A1 (en) | Method for training wake-up word detection model, wake-up word detection method, and non-transient computer-readable storage medium | |
| CN117012202B (zh) | 语音通道识别方法、装置、存储介质及电子设备 | |
| HK40084314A (zh) | 一种喇叭单体的检测方法、音频检测的方法以及相关装置 | |
| Kadam et al. | Proximity-Gated Wake-Word Spotting on Low-Power Embedded Hardware Using Quantized CNNs | |
| Delgado | Design of Embedded Systems for Real-Time Voice Recognition | |
| WO2026007785A1 (zh) | 语音唤醒方法、装置、电子设备以及存储介质 | |
| HK40042005A (zh) | 一种语音交互的处理方法、装置、终端及存储介质 | |
| HK40027344A (zh) | 一种基於人工智能的智能终端唤醒方法和装置 | |
| HK40021088B (zh) | 一种基於人工智能的智能设备唤醒方法和装置 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 25832062 Country of ref document: EP Kind code of ref document: A1 |