EP4413563A1 - Procédé d'analyse d'un signal sonore bruité pour la reconnaissance de mots clé de commande et d'un locuteur du signal sonore bruité analysé - Google Patents
Procédé d'analyse d'un signal sonore bruité pour la reconnaissance de mots clé de commande et d'un locuteur du signal sonore bruité analyséInfo
- Publication number
- EP4413563A1 EP4413563A1 EP22793176.3A EP22793176A EP4413563A1 EP 4413563 A1 EP4413563 A1 EP 4413563A1 EP 22793176 A EP22793176 A EP 22793176A EP 4413563 A1 EP4413563 A1 EP 4413563A1
- Authority
- EP
- European Patent Office
- Prior art keywords
- sound signal
- sound
- group
- speaker
- noisy
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/08—Speech classification or search
- G10L15/16—Speech classification or search using artificial neural networks
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L25/00—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
- G10L25/78—Detection of presence or absence of voice signals
- G10L25/84—Detection of presence or absence of voice signals for discriminating voice from noise
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/08—Learning methods
- G06N3/09—Supervised learning
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/06—Creation of reference templates; Training of speech recognition systems, e.g. adaptation to the characteristics of the speaker's voice
- G10L15/063—Training
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/20—Speech recognition techniques specially adapted for robustness in adverse environments, e.g. in noise, of stress induced speech
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L17/00—Speaker identification or verification techniques
- G10L17/04—Training, enrolment or model building
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L17/00—Speaker identification or verification techniques
- G10L17/18—Artificial neural networks; Connectionist approaches
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/08—Speech classification or search
- G10L2015/088—Word spotting
Definitions
- TITLE Process for the analysis of a noisy sound signal for the recognition of command keywords and of a speaker of the noisy sound signal analyzed
- the technical field of the invention is that of the analysis of sound signals and in particular that of the analysis of noisy sound signals for the recognition of command keywords and their speaker.
- the present invention relates to a method for analyzing a noisy sound signal and in particular a method for analyzing a noisy sound signal for the recognition of at least one group of command keywords and a speaker of the noisy sound signal.
- the present invention also relates to a system for implementing the method according to the invention.
- These connected speakers are capable of analyzing sound signals to identify and recognize predefined command keywords present in the analyzed sound signal and send the corresponding commands via a wireless link to the home automation objects concerned.
- the identification and recognition of command keywords is generally carried out as soon as a particular keyword, known as the activation keyword, has been detected to avoid triggering commands inadvertently.
- the connected speaker does not always succeed in recognizing the command keywords present in a sound signal, in particular when the speaker has a particular accent or uses a language not represented in the training database.
- the invention offers a solution to the problems mentioned above, by making it possible to recognize each command keyword present in a sound signal regardless of the linguistic specificities of its speaker.
- a first aspect of the invention relates to a method for analyzing a noisy sound signal for the recognition of at least one group of command keywords and of a speaker of the analyzed noisy sound signal, the sound signal noise to be analyzed being recorded by at least one microphone and the method comprising the following steps:
- the surrounding noise being a noise generated by the sound environment of the speaker
- Supervised training of an artificial neural network on the training data base to obtain a trained artificial neural network capable of providing from a sound signature obtained from a noisy sound signal, a speaker prediction and at least one command keyword group prediction;
- the artificial neural network is trained to be able both to recognize each command keyword present in the analyzed sound signal, and to identify the speaker of the analyzed sound signal.
- the training is carried out on a training database comprising, for each speaker to be recognized, a plurality of sound signatures obtained from non-noisy sound signals recorded by the speaker himself, thus presenting the specificities of the speaker such as their language or accent, and the command keywords the speaker wants to use.
- the recognition of the speaker by the artificial neural network is therefore facilitated without the need for a phoneme translation step or a language understanding step, and each speaker can personalize the command keywords used.
- the training database takes into account the noise near the microphone, which improves the performance of the artificial neural network on the sound signals recorded by the microphone presenting a similar noise.
- the training is carried out on a single training database allowing on the one hand the learning of the command keywords by the artificial neural network, and on the other hand the identification of characteristics biometrics allowing the speaker to be recognized by the artificial neural network.
- the quantity of data necessary for training the artificial neural network is therefore much lower than what is necessary in the state of the art where these two tasks are carried out separately on two distinct training databases.
- the method according to the first aspect of the invention may have one or more additional characteristics among the following, considered individually or according to all technically possible combinations.
- the trained artificial neural network is also capable of providing, from a sound signature, an activation binary prediction relating to the detection or not of at least one group of words activation keys, each sound signature of the training database being further associated with an activation binary, the step of using the trained artificial neural network further making it possible to obtain a prediction of binary d activation.
- the trained artificial neural network is also capable of providing, from a sound signature, a prediction of a termination bit relating to the detection or not of at least one group of termination keywords, each sound signature of the training database being further associated with a termination binary, the step of using the trained artificial neural network making it possible to further obtain a termination bit prediction.
- the trained artificial neural network is also capable of providing, from a sound signature, at least one binding binary prediction relating to the detection or not of at least one group of linking keywords, each sound signature of the training database being further associated with at least one linking bit and, if the value of the linking bit corresponds to the detection of at at least one group of binding keywords, to at least a second group of control keywords, the step of using the trained artificial neural network further obtaining a binding bit prediction and at least one prediction second group of command keywords.
- the performance of the artificial neural network for the recognition of command keywords is increased in the case where the analyzed sound signal has at least a first group of command keywords and a second group of command keywords , since the range of the analyzed sound signal comprising the first group of control keywords is delimited by the group of activation keywords on the one hand and the group of linking keywords on the other hand.
- At least one non-noisy sound signal recorded during the step of forming the training database is pronounced by a moving speaker.
- the artificial neural network has better performance for the recognition of command keywords on the sound signals uttered by mobile speakers, without having to multiply the number of microphones required, thanks to the spatialization of the data.
- the training database is updated on request, at regular intervals, or automatically after detection of a modification of the sound environment of the microphone.
- the training database is updated to adapt to the noise near the microphone, which may vary.
- the supervised training step of the artificial neural network is performed as soon as the training database is updated.
- a second aspect of the invention relates to a system for implementing the method according to the invention comprising: at least one microphone configured to record noisy or non-noisy sound signals, and the surrounding noise; at least one local computer configured to: calculate sound signatures from noisy sound signals obtained via at least one microphone; using the artificial neural network trained on calculated sound signatures; at least one main computer configured to: constitute the training database from sound signatures calculated by the local computer; supervised training of the artificial neural network on the constituted training database.
- the system according to the invention further comprises at least one storage device configured to store each noiseless sound signal recorded.
- the method according to the invention can be carried out offline, that is to say locally.
- the system according to the invention comprises a plurality of independent or coupled microphones.
- the quality of the recorded sound signals is better, in particular the errors due to echoes are reduced.
- the system according to the invention comprises a local computer per microphone.
- the central computer carries out the training of the artificial neural network which requires significant computing resources, and communicates the trained artificial neural network to each local computer which processes the sound signals recorded by the corresponding microphone.
- the local computer and the central computer correspond to a single computer.
- a third aspect of the invention relates to a computer program product comprising instructions which, when the program is executed on a computer, lead the latter to implement the steps of the method according to the invention.
- a fourth aspect of the invention relates to a computer-readable recording medium comprising instructions which, when executed by a computer, lead the latter to implement the steps of the method according to the invention.
- Figure 1 is a block diagram illustrating the sequence of steps of a method according to the invention.
- Figure 2 shows a schematic representation of a first embodiment of a system according to the invention.
- Figure 3 shows a schematic representation of a second embodiment of the system according to the invention.
- Figure 4 shows a schematic representation of a third embodiment of the system according to the invention.
- a first aspect of the invention relates to a method for analyzing a sound signal making it possible both to recognize each group of command keywords present in the analyzed sound signal and to identify the speaker of the analyzed sound signal. .
- the analyzed sound signal is recorded by at least one microphone and noisy, that is to say it includes a useful non-noisy sound signal pronounced by the speaker and noise generated by the sound environment of the speaker, otherwise known as environmental noise, for example a signal generated by a television or a vacuum cleaner.
- environmental noise for example a signal generated by a television or a vacuum cleaner.
- Surrounding noise is continuously changing and can be both stationary, for example generated by ventilation, and unsteady, for example generated by a computer keyboard.
- the invention was tested by considering the sound files corresponding to the following environments: interior of a vehicle, traffic noise, vacuum cleaner, drill, keyboard, musical instruments, singing, white noise, etc.
- non-noisy sound signal means a sound signal whose signal-to-noise ratio is strictly greater than 15 dB.
- noise signal means a sound signal whose signal-to-noise ratio is less than 15 dB.
- microphone designates both a single microphone and a network of microphones comprising a plurality of microphones located in the same place and aimed at improving the quality of the recorded sound signals.
- keyword group means an intent sentence or "intent sentence” in English.
- group of keywords does not need to have any meaning, nor to be in an existing language.
- group of command keywords means a group of words making it possible to trigger a command for a connected electronic device.
- the group of command keywords “lower the sound” makes it possible to trigger a command from a connected speaker broadcasting music so that the speaker lowers the volume
- the group of command keywords “turn off the light” allows you to trigger a command from a connected lamp illuminating a room so that the lamp turns off.
- a group of command keywords comprises at least one word.
- the number of commands that can be triggered is limited and depends in particular on the number of connected electronic devices.
- the commands that can be triggered are chosen by a user.
- Each command is associated with at least one group of command keywords making it possible to trigger the command.
- the command to turn off an air conditioner can be associated with both the "turn off the air conditioner” command keyword group and the "turn off the air conditioner” command keyword group.
- the speaker of the analyzed sound signal is identified from among a group of speakers comprising a finite number of speakers.
- FIG. 1 is a block diagram illustrating the sequence of steps of the method 100 according to the invention.
- a first step 101 of the method 100 according to the invention consists in building a training database.
- the first step 101 comprises a first sub-step 1011 consisting, for each speaker of the group of speakers, in recording at least one non-noisy sound signal pronounced by the speaker.
- each non-noisy sound signal can be recorded by the microphone that recorded the analyzed noisy sound signal or by another microphone.
- Each non-noisy sound signal is for example uttered by the speaker when he is moving, that is to say that the non-noisy sound signal is uttered in different distinct positions.
- a second sub-step 1012 consists for the microphone having recorded the analyzed noisy sound signal, in recording the surrounding noise.
- a third sub-step 1013 consists in adding the noise recorded in the second sub-step 1012 to each non-noisy sound signal recorded in the first sub-step 101 to obtain a noisy sound signal.
- a fourth sub-step 1014 consists in calculating a sound signature for each noisy sound signal obtained in the third sub-step 1013.
- a fifth sub-step 1015 consists, for each sound signature calculated in the fourth step 1014, in associating the calculated sound signature: with the speaker who uttered the non-noisy sound signal on the basis of which the sound signature was calculated; at least one group of control keywords present in the non-noisy sound signal.
- the information associated with each non-noisy sound signal during the fifth sub-step 1015 is for example provided by the speaker during a configuration phase.
- the training database constituted then comprises each sound signature calculated in the fourth sub-step 1014 associated with the speaker and with the group of command keywords associated with the sound signature during the fifth sub-step 1015.
- the training database constituted in the first step 101 can be updated on request, at regular intervals, or automatically after detection of a modification of the sound environment of the microphone.
- the microphone To detect a change in the sound environment of the microphone, the microphone records for example the surrounding noise permanently, on request or at regular intervals and it is considered for example that there is a change in the sound environment of the microphone if a difference of at least 3 dB is observed between two recordings of the surrounding noise by the microphone.
- a second step 102 of the method 100 according to the invention consists in training in a supervised manner an artificial neural network on the training database constituted in the first step 101 .
- the artificial neural network can be any artificial neural network capable of performing multi-label classification or "multi-label classification" in English.
- Supervised training otherwise called supervised learning, makes it possible to train an artificial neural network for a predefined task, by updating its parameters so as to minimize a cost function corresponding to the error between the data of output provided by the artificial neural network and the real output datum, i.e. what the artificial neural network should output to fulfill the predefined task on a certain input datum.
- a training database therefore comprises input data, each associated with a real output data.
- the training database comprises a plurality of sound signatures, each sound signature of the plurality of sound signatures being obtained from a noisy sound signal and associated with: a speaker of the noisy sound signal corresponding to the signature sound; at least one group of control keywords identified in the noisy sound signal corresponding to the sound signature.
- the input data are the sound signatures and the real output data are the speaker and the command keyword group(s).
- the supervised training of the artificial neural network therefore consists in updating the parameters so as to minimize a cost function taking into account the error between the speaker prediction provided by the artificial neural network from a sound signature from the training database and the speaker associated with the sound signature in the training database, and the error between the control keyword group prediction provided by the artificial neural network to from the sound signature and the command keyword group associated with the sound signature in the training database.
- the cost function is for example the binary cross-entropy function.
- Each sound signature of the training database is, for example, of the cepstral coefficient type of frequency Mel of the corresponding noisy sound signal, of the i-vector type obtained from the corresponding noisy sound signal or of the x-vector type obtained from the corresponding noisy sound signal.
- the second step 102 is for example carried out as soon as the training database is updated.
- a third step 103 of the method 100 according to the invention consists in calculating a sound signature from the analyzed sound signal.
- the sound signature calculated in the third step 103 is of the same type as the sound signatures of the training database.
- a fourth step 104 of the method 100 according to the invention consists in using the artificial neural network trained in the second step 102 on the sound signature calculated in the third step 103.
- the artificial neural network then provides a speaker prediction, and at least one command keyword group prediction.
- the speaker prediction corresponds to a speaker among the group of speakers or to a parameter indicating that the speaker is not known.
- the command keyword group prediction corresponds to a group of command keywords encountered during supervised training or to a parameter indicating that the command keyword group is not known or does not exist.
- the artificial neural network therefore performs a multi-label classification giving the identity of the speaker or detecting an unknown speaker via a first group of labels and giving the group of control keywords possibly detected for the speaker detected via a second label group.
- the analyzed sound signal can also include a group of activation keywords preceding the group or groups of command keywords.
- a group of activation keywords comprises at least one word.
- a group of activation keywords is for example “hello” or "please”.
- a sound signal allowing the triggering of a command causing the stopping of an air conditioner therefore includes, for example, the useful sound signal "hello stop air conditioning >>, "hello” being the activation keyword group and "stop air conditioning” being the command keyword group.
- each sound signature of the training database is also associated with an activation bit relating to the detection or not of at least one group of activation keywords in the noisy sound signal. corresponding to the sound signature, that is to say worth 1 if at least one group of activation keywords is present and 0 otherwise, and the artificial neural network also provides, at the fourth step 104, a binary prediction activation.
- the speaker Alternatively to the use of a group of activation keywords, before the recording of a sound signal, the speaker must for example wait for a certain duration, for example of the order of one second, before speaking the command keyword group(s).
- the analyzed sound signal can also include a group of termination keywords following the group or groups of command keywords.
- a group of termination keywords comprises at least one word.
- a group of termination keywords is for example “end” or “thank you”.
- a sound signal allowing the triggering of a command causing the stopping of an air conditioner therefore includes, for example, the useful sound signal “hello, stop the air conditioning thank you", where "hello” is the enable keyword group, "stop air conditioning” is the command keyword group, and "thank you” is the termination keyword group.
- each sound signature of the training database is also associated with a termination bit relating to the detection or not of at least one group of termination keywords in the noisy sound signal corresponding to the sound signature, and the artificial neural network also provides in the fourth step 104, a termination bit prediction.
- the analyzed sound signal can also comprise a group of linking keywords situated between two groups of command keywords.
- a group of linking keywords comprises at least one word.
- a group of linking keywords is for example “and” or “then”.
- a sound signal allowing the triggering of a command causing the stopping of an air conditioner and the extinction of the light therefore comprises for example the helpful beep "hello turn off the aircon then turn off the light", "hello” being the activation keyword group, "stop the aircon” being the first control keyword group, "then >> being the linking keyword group, and "turn off the light” being the second command keyword group.
- each sound signature of the training database is also associated with at least one link bit relating to the detection or not of at least one group of link keywords in the noisy sound signal.
- the artificial neural network also provides to the fourth step 104, a link bit prediction and a second control keyword group prediction.
- the training database includes sound signatures obtained from noisy sound signals uttered by moving speakers, an average absolute error of 9% is obtained for the prediction of command keyword groups.
- a second aspect of the invention relates to a system 200 allowing the implementation of the method 100 according to the invention.
- FIG. 2 shows a schematic representation of a first embodiment of the system 200 according to the invention.
- FIG. 3 shows a schematic representation of a second embodiment of the system 200 according to the invention.
- FIG. 4 shows a schematic representation of a third embodiment of the system 200 according to the invention.
- the system 200 comprises: at least one microphone 201 configured to record noisy sound signals, non-noisy sound signals and surrounding noise; at least one local computer 202-1 configured to: calculate sound signatures from noisy sound signals obtained via at least one microphone 201; using the artificial neural network trained on calculated sound signatures; at least one central computer 202-2 configured to: constitute the training database from calculated sound signatures; supervised training of the artificial neural network on the constituted training database; the local computer 202-1 possibly being confused with the central computer 202-2.
- the system 200 comprises for example a plurality of independent or coupled microphones 201.
- the system 200 comprises for example four microphones 201, which makes it possible to cover 360°.
- the system 200 comprises at least one microphone 201 physically connected to a single computer 202 playing both the role of local computer 202-1 and the role of central computer 202-2.
- the system 200 comprises a single microphone 201 physically connected to a computer 202.
- the system 200 comprises at least one microphone 201 connected via a wired or wireless link to a single computer 200 playing both the role of local computer 202-1 and the role of central computer 202-2.
- the system 200 comprises two microphones 201 connected via a wireless link to a computer 202.
- the system 200 according to the third embodiment comprises at least one microphone 201, each microphone 201 being connected physically or via a wired or wireless link to a local computer 202-1 and each local computer 202-1 being connected via a wired or wireless link to a central computer 202-2.
- the system 200 includes two microphones 201 each physically connected to a local computer 202-1 and each local computer 202-1 being connected via a wireless link to a central computer 202-2.
- the system 200 according to the invention can also comprise a storage device 203, for example a memory.
- the storage device 203 stores for example each non-noisy sound signal recorded during the first sub-step 101 1 or each non-noisy sound signal recorded during the first sub-step 101 1 by a given microphone 201.
- the system 200 according to the invention is for example a communication gateway and more particularly a connected enclosure.
- the proposed approach is more generic than commercial tools.
- the proposed approach makes it possible to achieve the noise or speech detection (VAD), command detection (CMD), sound environment identification (ASC) and speaker identification (SPEAKER ID); while the proposed commercial solutions only allow noise or speech detection (VAD) and command detection (CMD) or speaker identification (SPEAKER ID).
- the proposed approach therefore adapts effectively and quickly to the conditions of use in which it is implemented.
- the approach guarantees the robustness of speaker recognition (“voice” columns) and command word recognition (“command” column ).
- voice voice
- command command
- the learning time of the model is systematically less than 1 s to achieve a high success rate, whereas with tools known from the state of the art this time is greater than 1 h.
- the proposed approach requires a small amount of memory compared to the known methods of the state of the art, which generally require more than 1 GB of RAM for training the model and disk space to store the training database.
Landscapes
- Engineering & Computer Science (AREA)
- Physics & Mathematics (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Human Computer Interaction (AREA)
- Acoustics & Sound (AREA)
- Multimedia (AREA)
- Computational Linguistics (AREA)
- Artificial Intelligence (AREA)
- Evolutionary Computation (AREA)
- Theoretical Computer Science (AREA)
- Signal Processing (AREA)
- Biophysics (AREA)
- Data Mining & Analysis (AREA)
- Biomedical Technology (AREA)
- General Health & Medical Sciences (AREA)
- Molecular Biology (AREA)
- Computing Systems (AREA)
- General Engineering & Computer Science (AREA)
- General Physics & Mathematics (AREA)
- Mathematical Physics (AREA)
- Software Systems (AREA)
- Life Sciences & Earth Sciences (AREA)
- Measurement Of Mechanical Vibrations Or Ultrasonic Waves (AREA)
- Soundproofing, Sound Blocking, And Sound Damping (AREA)
Abstract
Description
Claims
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| FR2110510A FR3127839B1 (fr) | 2021-10-05 | 2021-10-05 | Procédé d’analyse d’un signal sonore bruité pour la reconnaissance de mots clé de commande et d’un locuteur du signal sonore bruité analysé |
| PCT/EP2022/077461 WO2023057384A1 (fr) | 2021-10-05 | 2022-10-03 | Procédé d'analyse d'un signal sonore bruité pour la reconnaissance de mots clé de commande et d'un locuteur du signal sonore bruité analysé |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP4413563A1 true EP4413563A1 (fr) | 2024-08-14 |
Family
ID=78483395
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP22793176.3A Pending EP4413563A1 (fr) | 2021-10-05 | 2022-10-03 | Procédé d'analyse d'un signal sonore bruité pour la reconnaissance de mots clé de commande et d'un locuteur du signal sonore bruité analysé |
Country Status (4)
| Country | Link |
|---|---|
| US (1) | US20240296859A1 (fr) |
| EP (1) | EP4413563A1 (fr) |
| FR (1) | FR3127839B1 (fr) |
| WO (1) | WO2023057384A1 (fr) |
Families Citing this family (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2022253003A1 (fr) * | 2021-05-31 | 2022-12-08 | 华为技术有限公司 | Procédé d'amélioration de la parole et dispositif associé |
| US12597434B2 (en) * | 2021-11-09 | 2026-04-07 | Dolby Laboratories Licensing Corporation | Control of speech preservation in speech enhancement |
| US12586598B2 (en) * | 2023-06-05 | 2026-03-24 | Infineon Technologies Americas Corp. | Audio distortion removal based on a set of reference audio samples |
Family Cites Families (13)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US9432611B1 (en) * | 2011-09-29 | 2016-08-30 | Rockwell Collins, Inc. | Voice radio tuning |
| KR102420518B1 (ko) * | 2015-09-09 | 2022-07-13 | 삼성전자주식회사 | 자연어 처리 시스템, 자연어 처리 장치, 자연어 처리 방법 및 컴퓨터 판독가능 기록매체 |
| US11693622B1 (en) * | 2015-09-28 | 2023-07-04 | Amazon Technologies, Inc. | Context configurable keywords |
| US11373672B2 (en) * | 2016-06-14 | 2022-06-28 | The Trustees Of Columbia University In The City Of New York | Systems and methods for speech separation and neural decoding of attentional selection in multi-speaker environments |
| EP4235645B1 (fr) * | 2016-07-06 | 2025-01-29 | DRNC Holdings, Inc. | Système et procédé pour personnaliser des interfaces vocales pour maison intelligente à l'aide de profils vocaux personnalisés |
| US9990926B1 (en) * | 2017-03-13 | 2018-06-05 | Intel Corporation | Passive enrollment method for speaker identification systems |
| US11775891B2 (en) * | 2017-08-03 | 2023-10-03 | Telepathy Labs, Inc. | Omnichannel, intelligent, proactive virtual agent |
| US10757148B2 (en) * | 2018-03-02 | 2020-08-25 | Ricoh Company, Ltd. | Conducting electronic meetings over computer networks using interactive whiteboard appliances and mobile devices |
| TWI719385B (zh) * | 2019-01-11 | 2021-02-21 | 緯創資通股份有限公司 | 電子裝置及其語音指令辨識方法 |
| US11276397B2 (en) * | 2019-03-01 | 2022-03-15 | DSP Concepts, Inc. | Narrowband direction of arrival for full band beamformer |
| WO2021062705A1 (fr) * | 2019-09-30 | 2021-04-08 | 大象声科(深圳)科技有限公司 | Procédé de détection en temps réel de mot-clé de discours à canal monophonique résistant |
| US11798550B2 (en) * | 2020-03-26 | 2023-10-24 | Snap Inc. | Speech-based selection of augmented reality content |
| US11881219B2 (en) * | 2020-09-28 | 2024-01-23 | Hill-Rom Services, Inc. | Voice control in a healthcare facility |
-
2021
- 2021-10-05 FR FR2110510A patent/FR3127839B1/fr active Active
-
2022
- 2022-10-03 WO PCT/EP2022/077461 patent/WO2023057384A1/fr not_active Ceased
- 2022-10-03 EP EP22793176.3A patent/EP4413563A1/fr active Pending
- 2022-10-03 US US18/697,907 patent/US20240296859A1/en active Pending
Also Published As
| Publication number | Publication date |
|---|---|
| WO2023057384A1 (fr) | 2023-04-13 |
| FR3127839A1 (fr) | 2023-04-07 |
| FR3127839B1 (fr) | 2024-04-12 |
| US20240296859A1 (en) | 2024-09-05 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| EP4413563A1 (fr) | Procédé d'analyse d'un signal sonore bruité pour la reconnaissance de mots clé de commande et d'un locuteur du signal sonore bruité analysé | |
| KR102509464B1 (ko) | 발언 분류기 | |
| KR102374519B1 (ko) | 문맥상의 핫워드들 | |
| EP3234945B1 (fr) | Affectation d'applications dans des systèmes basés sur la parole | |
| US8635243B2 (en) | Sending a communications header with voice recording to send metadata for use in speech recognition, formatting, and search mobile search application | |
| US8949266B2 (en) | Multiple web-based content category searching in mobile search application | |
| CN110097870B (zh) | 语音处理方法、装置、设备和存储介质 | |
| CN111460111A (zh) | 评估自动对话服务的重新训练推荐 | |
| US20170148429A1 (en) | Keyword detector and keyword detection method | |
| US20110054899A1 (en) | Command and control utilizing content information in a mobile voice-to-speech application | |
| US20110054896A1 (en) | Sending a communications header with voice recording to send metadata for use in speech recognition and formatting in mobile dictation application | |
| US20110054895A1 (en) | Utilizing user transmitted text to improve language model in mobile dictation application | |
| US20110054900A1 (en) | Hybrid command and control between resident and remote speech recognition facilities in a mobile voice-to-speech application | |
| US20110060587A1 (en) | Command and control utilizing ancillary information in a mobile voice-to-speech application | |
| US20110054894A1 (en) | Speech recognition through the collection of contact information in mobile dictation application | |
| US20110054898A1 (en) | Multiple web-based content search user interface in mobile search application | |
| US20110054897A1 (en) | Transmitting signal quality information in mobile dictation application | |
| FR2743238A1 (fr) | Dispositif de telecommunication reagissant a des ordres vocaux et procede d'utilisation de celui-ci | |
| CN113889091A (zh) | 语音识别方法、装置、计算机可读存储介质及电子设备 | |
| CN108877779B (zh) | 用于检测语音尾点的方法和装置 | |
| EP3627510B1 (fr) | Filtrage d'un signal sonore acquis par un systeme de reconnaissance vocale | |
| CN116013370A (zh) | 一种语音情感分类及合成方法、系统、装置及存储介质 | |
| CN119559941A (zh) | 基于意图识别的智能打断语音机器人对话的方法及装置 | |
| US20250182773A1 (en) | Methods and apparatuses for speech enhancement | |
| CN113241061B (zh) | 语音识别结果的处理方法、装置、电子设备和存储介质 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: UNKNOWN |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE |
|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE |
|
| 17P | Request for examination filed |
Effective date: 20240412 |
|
| AK | Designated contracting states |
Kind code of ref document: A1 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR |
|
| DAV | Request for validation of the european patent (deleted) | ||
| DAX | Request for extension of the european patent (deleted) | ||
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: EXAMINATION IS IN PROGRESS |
|
| 17Q | First examination report despatched |
Effective date: 20250908 |