WO2018014469A1 - 语音识别处理方法和装置 - Google Patents
语音识别处理方法和装置 Download PDFInfo
- Publication number
- WO2018014469A1 WO2018014469A1 PCT/CN2016/105080 CN2016105080W WO2018014469A1 WO 2018014469 A1 WO2018014469 A1 WO 2018014469A1 CN 2016105080 W CN2016105080 W CN 2016105080W WO 2018014469 A1 WO2018014469 A1 WO 2018014469A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- acoustic model
- mandarin acoustic
- mandarin
- province
- model
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/06—Creation of reference templates; Training of speech recognition systems, e.g. adaptation to the characteristics of the speaker's voice
- G10L15/063—Training
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/06—Creation of reference templates; Training of speech recognition systems, e.g. adaptation to the characteristics of the speaker's voice
- G10L15/065—Adaptation
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/06—Creation of reference templates; Training of speech recognition systems, e.g. adaptation to the characteristics of the speaker's voice
- G10L15/065—Adaptation
- G10L15/07—Adaptation to the speaker
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/08—Speech classification or search
- G10L15/16—Speech classification or search using artificial neural networks
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/01—Assessment or evaluation of speech recognition systems
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/06—Creation of reference templates; Training of speech recognition systems, e.g. adaptation to the characteristics of the speaker's voice
- G10L15/063—Training
- G10L2015/0631—Creating reference templates; Clustering
Definitions
- the present invention relates to the field of speech recognition technologies, and in particular, to a speech recognition processing method and apparatus.
- the performance of speech recognition is one of the key factors affecting the practical application of speech recognition products.
- Acoustic model as the main component of speech recognition, plays a key role in the performance of speech recognition.
- how to comprehensively utilize various information to enhance the performance and promotion ability of acoustic models has important theoretical research and practical application value for the speech recognition industry.
- the user's Mandarin pronunciation may have a certain degree of dialect accent.
- “h” and “f” are often not found, and in the Mandarin speech recognition product.
- the Mandarin acoustic models are for national users and do not take into account the accent differences in the user's Mandarin.
- the object of the present invention is to solve at least one of the above technical problems to some extent.
- a first object of the present invention is to provide a speech recognition processing method which establishes a Mandarin acoustic model with a dialect accent based on the accent difference of users in different regions, thereby improving the performance of speech recognition.
- a second object of the present invention is to provide a speech recognition processing apparatus.
- a third object of the invention is to propose an apparatus.
- a fourth object of the present invention is to provide a non-volatile computer storage medium.
- the first aspect of the present invention provides a voice recognition processing method, including the following steps:
- training is performed on a preset processing model to generate a general Mandarin acoustic model
- adaptive training is performed on the common Mandarin acoustic model to generate a Mandarin acoustic model with dialect accents corresponding to each province.
- the speech recognition processing method of the embodiment of the present invention tests and evaluates the common Mandarin acoustic model and the Mandarin acoustic model with dialect accent according to the voice test data of each province, and the recognition performance of the Mandarin acoustic model with dialect accent Above the general Mandarin acoustic model, a Mandarin acoustic model with a dialect accent is deployed.
- the Putonghua acoustic model with dialect accent whose recognition performance is higher than that of the general Mandarin acoustic model is deployed online, which ensures the practicability of the speech recognition processing method.
- the speech recognition processing method of the embodiment of the present invention further has the following additional technical features:
- the voice sample data of all regions in the country is trained on a preset processing model to generate a general Mandarin acoustic model, including:
- the training is performed on the preset deep neural network model, and the model structure based on the deep long and short time memory unit and the acoustic model connecting the time series classification framework are generated.
- the performing adaptive training on the universal Mandarin acoustic model includes:
- the method further includes:
- the recognition performance of the Mandarin acoustic model with a dialect accent is higher than the general Mandarin acoustic model, the Mandarin acoustic model with a dialect accent is deployed online.
- the method further includes:
- the speech information is input to the universal Mandarin acoustic model for speech recognition.
- the second aspect of the present invention provides a voice recognition processing apparatus, including: a first generation module, configured to perform training on a preset processing model according to voice sample data of all regions in the country to generate a universal Mandarin acoustic model;
- a second generating module configured to perform adaptive training on the common Mandarin acoustic model according to the voice sample data of each province, and generate a Mandarin acoustic model with a dialect accent corresponding to each province.
- the speech recognition processing device of the embodiment of the present invention tests and evaluates the general-purpose Mandarin acoustic model and the Mandarin-acoustic acoustic model with dialect accent according to the voice test data of each province, and the recognition performance of the Mandarin acoustic model with dialect accent Above the general Mandarin acoustic model, a Mandarin acoustic model with a dialect accent is deployed.
- the Putonghua acoustic model with dialect accent whose recognition performance is higher than that of the general Mandarin acoustic model is deployed online, which ensures the practicability of the speech recognition processing method.
- the speech recognition processing apparatus of the embodiment of the present invention further has the following additional technical features:
- the first generating module is configured to:
- the training is performed on the preset deep neural network model, and the model structure based on the deep long and short time memory unit and the acoustic model connecting the time series classification framework are generated.
- the second generation module performs adaptive training on the universal Mandarin acoustic model, respectively, including:
- the device further includes:
- An evaluation module configured to perform test evaluation on the universal Mandarin acoustic model and the Mandarin acoustic model with dialect accent according to voice test data of each province;
- a deployment module configured to deploy the Mandarin acoustic model with a dialect accent on the line when the recognition performance of the Mandarin acoustic model with a dialect accent is higher than the common Mandarin acoustic model.
- the device further includes:
- a receiving module configured to receive voice information that is sent by a user and carries network address information
- a determining module configured to determine, according to the network address information, province information where the user is located;
- a judging module configured to determine whether a Mandarin acoustic model with a dialect accent corresponding to the province information is deployed
- a processing module configured to input the voice information into a Mandarin acoustic model with a dialect accent corresponding to the province information for voice recognition when a Mandarin acoustic model with a dialect accent corresponding to the province information is deployed ;
- the processing module is further configured to input the voice information into the common Mandarin acoustic model for voice recognition when a Mandarin acoustic model with a dialect accent corresponding to the province information is not deployed.
- a third aspect of the present invention provides an apparatus, including: one or more processors; a memory; one or more programs, the one or more programs being stored in the memory, When executed by the one or more processors, performing the following steps: performing training on a preset processing model according to voice sample data of all regions of the country to generate a general-purpose Mandarin acoustic model;
- adaptive training is performed on the common Mandarin acoustic model to generate a Mandarin acoustic model with dialect accents corresponding to each province.
- a fourth aspect of the present invention provides a nonvolatile computer storage medium storing one or more programs when the one or more programs are executed by one device. And causing the device to perform the following steps: performing training on a preset processing model according to voice sample data of all regions of the country to generate a common Mandarin acoustic model;
- adaptive training is performed on the common Mandarin acoustic model to generate a Mandarin acoustic model with dialect accents corresponding to each province.
- FIG. 1 is a flow chart of a speech recognition processing method in accordance with one embodiment of the present invention.
- FIG. 2 is a flow chart showing generation of a Mandarin-like acoustic model with an accent in accordance with one embodiment of the present invention
- FIG. 3 is a flowchart of a voice recognition processing method according to another embodiment of the present invention.
- FIG. 4 is a flowchart of a voice recognition processing method according to still another embodiment of the present invention.
- FIG. 5 is a schematic structural diagram of a voice recognition processing apparatus according to an embodiment of the present invention.
- FIG. 6 is a schematic structural diagram of a voice recognition processing apparatus according to another embodiment of the present invention.
- FIG. 7 is a schematic structural diagram of a speech recognition processing apparatus according to still another embodiment of the present invention.
- FIG. 1 is a flowchart of a voice recognition processing method according to an embodiment of the present invention. As shown in FIG. 1, the method includes:
- S110 Perform training on a preset processing model according to the voice sample data of all regions of the country to generate a common Mandarin acoustic model.
- a processing model for generating a common-sound acoustic model is preset, for example, a depth neural network model is preset, and then voice sample data of all regions in the country is collected, and the voice sample data is input into a preset processing model.
- the processing model extracts the speech features in the speech sample data, maps the speech features to the language basic unit, generates a general-purpose Mandarin acoustic model, and based on the universal Mandarin acoustic model, the speech of the national user can be recognized.
- S120 Perform adaptive training on the common Mandarin acoustic model according to the voice sample data of each province, and generate a Mandarin acoustic model with a dialect accent corresponding to each province.
- the user's Mandarin pronunciation may have a certain degree of dialect accent.
- the pronunciations of "c” and “ch” are the same.
- “c” and “ch” have obvious boundaries, which makes it impossible to accurately identify the user's voice data.
- the speech recognition processing method of the embodiment of the present invention performs training based on the original general-purpose Mandarin acoustic model, and optimizes the general-purpose Mandarin acoustic model based on the pronunciation features of the dialect accents of different provinces, for each different dialect.
- the accent establishes a corresponding Mandarin acoustic model with a dialect accent, so that the voice data input by the user can be accurately identified by the Mandarin acoustic model with different dialect accents.
- voice sample data of provinces in the country are collected as adaptive data, wherein the voice sample data collected by each province may be relatively small, such as may be several hundred hours of voice.
- adaptive training is performed on the common Mandarin acoustic model, and adaptive training is provided for each province to obtain the corresponding Mandarin acoustic model.
- the above adaptive training refers to: in the process of processing and analyzing the voice sample data of the collected provinces in the country, automatically adjusting the processing parameters, boundary conditions or constraints of the acoustic model of the Mandarin according to the data characteristics of the voice sample data.
- the general Mandarin model is optimized to a Mandarin acoustic model that is compatible with the statistical distribution features and structural features of the speech sample data of each province.
- the collected voice sample data of the above five provinces may be separately input.
- the general Mandarin acoustic model adaptive training is performed on the common Mandarin acoustic model according to the voice sample data of each province, and a Mandarin acoustic model with a Henan accent corresponding to the above five provinces is generated, with a Hebei accent.
- the speech recognition processing method of the embodiment of the present invention performs training on a preset processing model according to the voice sample data of all regions in the country, and generates a general-purpose Mandarin acoustic model, and according to the voice sample data of each province, respectively Adaptive training is performed on the general Mandarin acoustic model to generate a Mandarin acoustic model with dialect accents corresponding to each province.
- a Mandarin acoustic model with dialect accent is established. Type, improving the performance of speech recognition.
- the performance of the generated acoustic model with the dialect accent can be verified. Therefore, only the Mandarin acoustic model with dialect accent, which is improved in performance compared to the ordinary acoustic model, is deployed online.
- FIG. 3 is a flowchart of a voice recognition processing method according to another embodiment of the present invention. As shown in FIG. 3, the method includes:
- a depth neural network model may be preset, the input of the deep neural network model being a single-frame or multi-frame stitched speech acoustic feature, and the output is a context-dependent vowel unit, ie, based on input acoustic characteristics Classification of context-dependent vocoding units to generate correlated acoustic models.
- the voice sample data of all regions of the country is input into the deep neural network model for training, and based on the acoustic characteristics of the input voice sample data, the classification processing of the context-dependent vocoding unit is performed to generate a memory unit based on the deep long and short time.
- S320 Perform adaptive training on the common Mandarin acoustic model according to the voice sample data of each province, and generate a Mandarin acoustic model with a dialect accent corresponding to each province.
- adaptive training can be performed on the common Mandarin acoustic model by using multiple adaptive training methods:
- an adaptive training method of tuning the basic model with a small learning rate can be employed, and adaptive training is performed on the general Mandarin acoustic model.
- the above two adaptive updating methods can be updated by the standard cross entropy criterion and the error backpropagation method, and the regularized objective function can be expressed as:
- ⁇ denotes a regular term coefficient
- o t denotes a feature of the t-th frame sample
- q t denotes a mark corresponding to the t-th frame sample
- W denotes a model parameter
- W 0 denotes a current model parameter.
- the probability expression of the target is a linear interpolation of the distribution under the true mark of the adaptive model and the adaptive data.
- an adaptive training method that only tunes part of the model parameters can be used to perform adaptive training on the general-purpose Mandarin acoustic model.
- the deep bottleneck layer is added by the singular value decomposition method to perform adaptive updating of fewer parameters, thereby reducing the amount of model parameters that the adaptive model needs to update.
- an adaptive training method that introduces new features can be employed to perform adaptive training on a common Mandarin acoustic model.
- the adaptive training method in this example takes into account the particularity of the dialect accent, introduces the more classic ivector and speaker coding methods in voiceprint and adaptive training, and features various types of complex information through each dialect speech.
- Vector extraction which is added to the input features for adaptive training.
- the corresponding ivector vector is extracted for the speech data of each packet for decoding.
- M is the mean supervector of all training corpora
- m is the mean supervector of the target speech accumulated to the current packet data
- T is the load matrix
- w is the ivector to be obtained.
- each frame feature After obtaining the ivector in the current corpus data, each frame feature will be stitched onto the ivector feature to form a new feature and then retrain the acoustic model.
- the model parameter weights of the ivector feature part are updated, while the original model parameters are kept unchanged, so as to ensure that the model does not over-fitting, and at the same time, the updated model and the original model do not change too much.
- the promotion effect of the generated Mandarin acoustic model with dialect accent is guaranteed.
- the generated Mandarin acoustic model with dialect accents corresponding to each province is not too biased toward the general Mandarin acoustic model, and in practical applications, the performance of the Mandarin acoustic model with dialect accent may not be higher than General Mandarin acoustic model.
- the speech test data of the Henan accent is input to the common Mandarin acoustic model and the Mandarin acoustic model with the river south accent, and the accuracy of the speech recognition based on the general Mandarin acoustic model and the Mandarin acoustic model with the river south accent is Performance is tested and evaluated.
- the recognition performance of the Mandarin acoustic model with dialect accent is higher than that of the general Mandarin acoustic model, it indicates that the Mandarin acoustic model with dialect accent can more accurately identify the accent with the accent acoustic model. Mandarin, thus deploying a Mandarin acoustic model with a dialect accent.
- the speech recognition processing method of the embodiment of the present invention tests and evaluates the common Mandarin acoustic model and the Mandarin acoustic model with dialect accent according to the voice test data of each province, if the Mandarin has a dialect accent.
- the acoustic model's recognition performance is higher than that of the general Mandarin acoustic model, and the Mandarin acoustic model with dialect accent is deployed online.
- the Putonghua acoustic model with dialect accent whose recognition performance is higher than that of the general Mandarin acoustic model is deployed online, which ensures the practicability of the speech recognition processing method.
- the dialect accent to which the user belongs can be determined in various ways to input the user according to the Mandarin acoustic model corresponding to the dialect accent. Voice information is identified.
- the personal information of the user may be obtained, and the dialect accent of the user is determined according to the province to which the citizen belongs, so as to identify the voice information input by the user according to the acoustic model of the Mandarin corresponding to the dialect accent.
- the network address information to which the user sends the voice recognition request may be obtained, and the province to which the network address information belongs may be determined to obtain the dialect accent of the user, so that the user may input according to the acoustic model of the Mandarin corresponding to the dialect accent.
- the voice information is identified.
- FIG. 4 is a flowchart of a voice recognition processing method according to still another embodiment of the present invention. As shown in FIG. 4, after step S340 shown in FIG. 3, the method includes:
- S410 Receive voice information that is sent by a user and carries network address information.
- the voice information carrying the network address information sent by the user may be received, and the province in which the network address information is located may be queried according to the network address information.
- the province information to which the network information belongs may be determined according to the IP address in the network address information.
- the voice information is input to a Mandarin acoustic model with a dialect accent corresponding to the province information for voice recognition.
- the speech information is input to the Mandarin acoustic model with dialect accent corresponding to the provincial information for speech recognition.
- the voice recognition processing method of the embodiment of the present invention determines the province information of the user according to the voice information of the network address information sent by the user, and deploys the Mandarin with the dialect accent corresponding to the province information.
- the speech acoustic information of the user is recognized using the Mandarin acoustic model with a dialect accent. Thereby, the performance of speech recognition is improved.
- FIG. 5 is a schematic structural diagram of a speech recognition processing apparatus according to an embodiment of the present invention. As shown in FIG. 5, the apparatus includes: a first generation module. 10 and a second generation module 20.
- the first generation module 10 is configured to perform training on a preset processing model according to the voice sample data of all regions in the country to generate a common Mandarin acoustic model.
- a processing model for generating a common-sound acoustic model is preset, for example, a depth neural network model is preset, and then voice sample data of all regions in the country is collected, and the voice sample data is input into a preset processing model.
- the first generation module 10 extracts the voice features in the voice sample data by processing the model, maps the voice features to the language basic unit, and generates a general-purpose Mandarin acoustic model, and the voice recognition of the national users can be realized based on the universal Mandarin acoustic model.
- the second generating module 20 is configured to perform adaptive training on the common Mandarin acoustic model according to the voice sample data of each province, and generate a Mandarin acoustic model with a dialect accent corresponding to each province.
- voice sample data of provinces in the country are collected as adaptive data, wherein the voice sample data collected by each province may be relatively small, such as may be several hundred hours of voice.
- the second generation module 20 performs adaptive training on the common Mandarin acoustic model based on the voice sample data collected by each province, and performs adaptive training for each province to obtain a corresponding Mandarin acoustic model.
- voice recognition processing method embodiment is also applicable to the voice of the embodiment.
- the identification processing device is similar in implementation principle and will not be described here.
- the speech recognition processing device of the embodiment of the present invention performs training on a preset processing model according to the voice sample data of all regions in the country, generates a general-purpose Mandarin acoustic model, and respectively according to the voice sample data of each province. Adaptive training is performed on the general Mandarin acoustic model to generate a Mandarin acoustic model with dialect accents corresponding to each province. Therefore, based on the accent difference of users in different regions, a common-sound acoustic model with dialect accent is established, which improves the performance of speech recognition.
- the performance of the generated acoustic model with the dialect accent can be verified. Therefore, only the Mandarin acoustic model with dialect accent, which is improved in performance compared to the ordinary acoustic model, is deployed online.
- FIG. 6 is a schematic structural diagram of a voice recognition processing apparatus according to another embodiment of the present invention. As shown in FIG. 6, the apparatus further includes an evaluation module 30 and a deployment module 40, as shown in FIG.
- the evaluation module 30 is configured to perform a test evaluation on the common Mandarin acoustic model and the Mandarin acoustic model with dialect accents according to the voice test data of each province.
- the deployment module 40 is configured to deploy a Mandarin acoustic model with a dialect accent when the recognition performance of the Mandarin acoustic model with a dialect accent is higher than that of the general Mandarin acoustic model.
- the first generation module 10 also inputs the voice sample data of all regions of the country into the deep neural network model for training, and classifies the context-dependent vowel unit based on the acoustic features of the input voice sample data.
- the training process is performed to generate a model structure based on deep long and short time memory cells, and an acoustic model connected to the time series classification frame.
- the second generation module 20 can adopt an adaptive training mode of tuning the basic model with a small learning rate, an adaptive training mode of only tuning part of the model parameters, and an adaptive training method of introducing new features, etc. on the common Mandarin acoustic model. Adaptive training is performed to generate a Mandarin acoustic model with a dialect accent.
- the evaluation module 30 needs to test and evaluate the common Mandarin acoustic model and the Mandarin acoustic model with dialect accents according to the voice test data of each province.
- the deployment module 40 then deploys a Mandarin acoustic model with a dialect accent.
- voice recognition processing method embodiment is also applicable to the voice recognition processing apparatus of the embodiment, and the implementation principle thereof is similar, and details are not described herein again.
- the speech recognition processing apparatus performs test evaluation on the common Mandarin acoustic model and the Mandarin acoustic model with dialect accent according to the voice test data of each province, if dialect is used.
- the recognition performance of the accented Mandarin acoustic model is higher than that of the general Mandarin acoustic model, and the Mandarin acoustic model with the dialect accent is deployed online.
- the Putonghua acoustic model with dialect accent whose recognition performance is higher than that of the general Mandarin acoustic model is deployed online, which ensures the practicability of the speech recognition processing method.
- the dialect accent to which the user belongs can be determined in various ways to input the user according to the Mandarin acoustic model corresponding to the dialect accent. Voice information is identified.
- FIG. 7 is a schematic structural diagram of a voice recognition processing apparatus according to still another embodiment of the present invention. As shown in FIG. 7, the apparatus further includes: a receiving module 50, a determining module 60, and a determining module. 70 and processing module 80.
- the receiving module 50 is configured to receive voice information that is sent by the user and carries network address information.
- the determining module 60 is configured to determine the province information where the user is located according to the network address information.
- the receiving module 50 can receive the voice information that is sent by the user and carry the network address information, and the determining module 60 can determine the province in which the network is located according to the network address information.
- the identifier can be determined according to the IP address in the network address information.
- Verification information etc.
- the determining module 70 is configured to determine whether a Mandarin acoustic model with a dialect accent corresponding to the province information is deployed.
- the processing module 80 is configured to input the voice information into the Mandarin acoustic model with the dialect accent corresponding to the province information for voice recognition when the Mandarin acoustic model with the dialect accent corresponding to the province information is deployed.
- the processing module 80 is further configured to input the voice information into the common Mandarin acoustic model for voice recognition when the Mandarin acoustic model with the dialect accent corresponding to the province information is not deployed.
- the determining module 70 may determine whether a Mandarin acoustic model with a dialect accent corresponding to the province information is deployed, and if deployed, indicating that the voice recognition performance is higher than the Mandarin acoustic model.
- a Mandarin acoustic model with a dialect accent corresponding to the province information and thus the processing module 80 inputs the voice information into a Mandarin acoustic model with a dialect accent corresponding to the province information for speech recognition.
- the processing module 80 If not deployed, it indicates that there is no Mandarin acoustic model with dialect accent corresponding to the provincial information corresponding to the Mandarin speech model, and thus the processing module 80 inputs the speech information into the general Mandarin acoustic model for speech recognition.
- voice recognition processing method embodiment is also applicable to the voice recognition processing apparatus of the embodiment, and the implementation principle thereof is similar, and details are not described herein again.
- the voice recognition processing apparatus determines the province information in which the user is located according to the voice information carried by the user carrying the network address information, and deploys the Mandarin with the dialect accent corresponding to the province information.
- the speech acoustic information of the user is recognized using the Mandarin acoustic model with a dialect accent. Thereby, the performance of speech recognition is improved.
Landscapes
- Engineering & Computer Science (AREA)
- Physics & Mathematics (AREA)
- Computational Linguistics (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Human Computer Interaction (AREA)
- Acoustics & Sound (AREA)
- Multimedia (AREA)
- Artificial Intelligence (AREA)
- Evolutionary Computation (AREA)
- Machine Translation (AREA)
- Telephonic Communication Services (AREA)
- Soundproofing, Sound Blocking, And Sound Damping (AREA)
Abstract
一种语音识别处理方法和装置,其中,方法包括:根据全国所有地区的语音样本数据在预设的处理模型上进行训练,生成通用普通话声学模型(S110);根据各省份的语音样本数据,分别在通用普通话声学模型上进行自适应训练,生成与各省份对应的带有方言口音的普通话声学模型(S120)。由此,基于不同地区的用户的口音差异,建立带有方言口音的普通话声学模型,提高了语音识别的性能。
Description
相关申请的交叉引用
本申请要求百度在线网络技术(北京)有限公司于2016年7月22日提交的、发明名称为“语音识别处理方法和装置”的、中国专利申请号“201610585160.6”的优先权。
本发明涉及语音识别技术领域,尤其涉及一种语音识别处理方法和装置。
语音识别的性能是影响语音识别产品实用化的关键因素之一,声学模型作为语音识别的主要组成部分,对语音识别性能的好坏起到了关键的作用。在声学模型的训练中,如何综合利用各种信息提升声学模型的表现和推广能力,对于语音识别产业具有重要的理论研究和实际应用的价值。
通常情况下,用户的普通话发音可能会带有一定程度的方言口音,比如带有湖南口音的用户的普通话发音中,则常会出现“h”“f”不分的情况,而普通话语音识别产品中的普通话声学模型都是面向全国用户的,没有考虑到用户普通话中的口音差异。
发明内容
本发明的目的旨在至少在一定程度上解决上述的技术问题之一。
为此,本发明的第一个目的在于提出一种语音识别处理方法,该方法基于不同地区的用户的口音差异,建立带有方言口音的普通话声学模型,提高了语音识别的性能。
本发明的第二个目的在于提出一种语音识别处理装置。
本发明的第三个目的在于提出一种设备。
本发明的第四个目的在于提出一种非易失性计算机存储介质。
为了实现上述目的,本发明第一方面实施例提出了一种语音识别处理方法,包括以下步骤:
根据全国所有地区的语音样本数据在预设的处理模型上进行训练,生成通用普通话声学模型;
根据各省份的语音样本数据,分别在所述通用普通话声学模型上进行自适应训练,生成与各省份对应的带有方言口音的普通话声学模型。
本发明实施例的语音识别处理方法,根据各省份的语音测试数据,分别对通用普通话声学模型,以及带有方言口音的普通话声学模型进行测试评估,如果带有方言口音的普通话声学模型的识别性能高于通用普通话声学模型,则将带有方言口音的普通话声学模型部署上线。由此,将识别性能高于通用普通话声学模型的带有方言口音的普通话声学模型部署上线,保证了语音识别处理方法的实用性。
另外,本发明实施例的语音识别处理方法还具有如下附加的技术特征:
在本发明的一个实施例中,所述根据全国所有地区的语音样本数据在预设的处理模型上进行训练,生成通用普通话声学模型,包括:
根据全国所有地区的语音样本数据在预设的深度神经网络模型上进行训练,生成基于深层长短时记忆单元的模型结构,以及连接时序分类框架的声学模型。
在本发明的一个实施例中,所述分别在所述通用普通话声学模型上进行自适应训练,包括:
较小学习率调优基本模型的自适应训练方式;或者,
只调优部分模型参数的自适应训练方式;或者,
引入新特征的自适应训练方式。
在本发明的一个实施例中,在所述生成与各省份对应的带有方言口音的普通话声学模型之后,还包括:
根据各省份的语音测试数据,分别对所述通用普通话声学模型,以及所述带有方言口音的普通话声学模型进行测试评估;
如果所述带有方言口音的普通话声学模型的识别性能高于所述通用普通话声学模型,则将所述带有方言口音的普通话声学模型部署上线。
在本发明的一个实施例中,在所述将所述带有方言口音的普通话声学模型部署上线之后,还包括:
接收用户发送的携带网络地址信息的语音信息;
根据所述网络地址信息确定所述用户所在的省份信息;
判断是否部署有与所述省份信息对应的带有方言口音的普通话声学模型;
如果部署,则将所述语音信息输入到与所述省份信息对应的带有方言口音的普通话声学模型进行语音识别;
如果没有部署,则将所述语音信息输入到所述通用普通话声学模型进行语音识别。
为了实现上述目的,本发明第二方面实施例提出了一种语音识别处理装置,包括:第一生成模块,用于根据全国所有地区的语音样本数据在预设的处理模型上进行训练,生成通用普通话声学模型;
第二生成模块,用于根据各省份的语音样本数据,分别在所述通用普通话声学模型上进行自适应训练,生成与各省份对应的带有方言口音的普通话声学模型。
本发明实施例的语音识别处理装置,根据各省份的语音测试数据,分别对通用普通话声学模型,以及带有方言口音的普通话声学模型进行测试评估,如果带有方言口音的普通话声学模型的识别性能高于通用普通话声学模型,则将带有方言口音的普通话声学模型部署上线。由此,将识别性能高于通用普通话声学模型的带有方言口音的普通话声学模型部署上线,保证了语音识别处理方法的实用性。
另外,本发明实施例的语音识别处理装置,还具有如下附加的技术特征:
在本发明的一个实施例中,所述第一生成模块用于:
根据全国所有地区的语音样本数据在预设的深度神经网络模型上进行训练,生成基于深层长短时记忆单元的模型结构,以及连接时序分类框架的声学模型。
在本发明的一个实施例中,所述第二生成模块分别在所述通用普通话声学模型上进行自适应训练,包括:
较小学习率调优基本模型的自适应训练方式;或者,
只调优部分模型参数的自适应训练方式;或者,
引入新特征的自适应训练方式。
在本发明的一个实施例中,所述装置还包括:
评估模块,用于根据各省份的语音测试数据,分别对所述通用普通话声学模型,以及所述带有方言口音的普通话声学模型进行测试评估;
部署模块,用于在所述带有方言口音的普通话声学模型的识别性能高于所述通用普通话声学模型时,将所述带有方言口音的普通话声学模型部署上线。
在本发明的一个实施例中,所述装置还包括:
接收模块,用于接收用户发送的携带网络地址信息的语音信息;
确定模块,用于根据所述网络地址信息确定所述用户所在的省份信息;
判断模块,用于判断是否部署有与所述省份信息对应的带有方言口音的普通话声学模型;
处理模块,用于在部署有与所述省份信息对应的带有方言口音的普通话声学模型时,将所述语音信息输入到与所述省份信息对应的带有方言口音的普通话声学模型进行语音识别;
所述处理模块还用于在没有部署与所述省份信息对应的带有方言口音的普通话声学模型时,将所述语音信息输入到所述通用普通话声学模型进行语音识别。
本发明附加的方面和优点将在下面的描述中部分给出,部分将从下面的描述中变得明
显,或通过本发明的实践了解到。
为了实现上述实施例,本发明第三方面实施例提供了一种设备,包括:一个或者多个处理器;存储器;一个或者多个程序,所述一个或者多个程序存储在所述存储器中,当被所述一个或者多个处理器执行时,执行以下步骤:根据全国所有地区的语音样本数据在预设的处理模型上进行训练,生成通用普通话声学模型;
根据各省份的语音样本数据,分别在所述通用普通话声学模型上进行自适应训练,生成与各省份对应的带有方言口音的普通话声学模型。
为了实现上述目的,本发明第四方面实施例提供了一种非易失性计算机存储介质,所述计算机存储介质存储有一个或者多个程序,当所述一个或者多个程序被一个设备执行时,使得所述设备执行以下步骤:根据全国所有地区的语音样本数据在预设的处理模型上进行训练,生成通用普通话声学模型;
根据各省份的语音样本数据,分别在所述通用普通话声学模型上进行自适应训练,生成与各省份对应的带有方言口音的普通话声学模型。
本发明的上述和/或附加的方面和优点从结合下面附图对实施例的描述中将变得明显和容易理解,其中:
图1是根据本发明一个实施例的语音识别处理方法的流程图;
图2是根据本发明一个实施例的生成带有口音的普通话声学模型的生成流程图;
图3是根据本发明另一个实施例的语音识别处理方法的流程图;
图4是根据本发明又一个实施例的语音识别处理方法的流程图;
图5是根据本发明一个实施例的语音识别处理装置的结构示意图;
图6是根据本发明另一个实施例的语音识别处理装置的结构示意图;以及
图7是根据本发明又一个实施例的语音识别处理装置的结构示意图。
下面详细描述本发明的实施例,所述实施例的示例在附图中示出,其中自始至终相同或类似的标号表示相同或类似的元件或具有相同或类似功能的元件。下面通过参考附图描述的实施例是示例性的,旨在用于解释本发明,而不能理解为对本发明的限制。
下面参考附图描述本发明实施例的语音识别处理方法和装置。
图1是根据本发明一个实施例的语音识别处理方法的流程图,如图1所示,该方法包括:
S110,根据全国所有地区的语音样本数据在预设的处理模型上进行训练,生成通用普通话声学模型。
具体地,预设一训练生成普通话声学模型的处理模型,比如预设一深度神经网络模型等,进而采集全国所有地区的语音样本数据,将该语音样本数据输入预设的处理模型。
进而,处理模型提取语音样本数据中的语音特征,将语音特征映射到语言基本单元,生成通用普通话声学模型,基于该通用普通话声学模型可实现对全国用户的语音的识别。
S120,根据各省份的语音样本数据,分别在通用普通话声学模型上进行自适应训练,生成与各省份对应的带有方言口音的普通话声学模型。
应当理解的是,在实际应用时,用户的普通话发音可能会带有一定程度的方言口音,例如,在带有四川口音的普通话发音中,其“c”和“ch”的发音是相同的,而普通话声学模型中“c”和“ch”具有明显的区分界线,导致不能对用户的语音数据进行准确地识别。
为了解决上述问题,本发明实施例的语音识别处理方法,在原有的通用普通话声学模型的基础上进行训练,基于不同省份的方言口音的发音特征,优化通用普通话声学模型,对每个不同的方言口音建立对应的带有方言口音的普通话声学模型,从而可以通过带有不同的方言口音的普通话声学模型,对用户输入的语音数据进行准确的识别。
具体地,在实际应用中,采集全国各省份的语音样本数据作为自适应数据,其中,每个省份所采集的语音样本数据,可能数量相对较少,比如可能为几百个小时的语音数量级,进而基于每个省份所采集的语音样本数据,分别在通用普通话声学模型上进行自适应训练,为各个省份进行自适应训练得到对应的普通话声学模型。
其中,上述自适应训练是指:在对采集的全国各省份的语音样本数据进行处理和分析过程中,根据语音样本数据的数据特征,自动调整普通话声学模型的处理参数、边界条件或约束条件等,使得通用普通话模型优化为与各省份的语音样本数据的统计分布特征、结构特征相适应的普通话声学模型。
举例而言,如图2所示,在生成带有广东、河北、河南、广西、四川五个省份的口音的普通话声学模型时,可将采集到的以上五个省份的语音样本数据,分别输入到通用普通话声学模型中,进而根据各省份的语音样本数据,分别在通用普通话声学模型上进行自适应训练,生成与以上五个省份对应的带有河南口音的普通话声学模型、带有河北口音的普通话声学模型等。
综上所述,本发明实施例的语音识别处理方法,根据全国所有地区的语音样本数据在预设的处理模型上进行训练,生成通用普通话声学模型,并根据各省份的语音样本数据,分别在通用普通话声学模型上进行自适应训练,生成与各省份对应的带有方言口音的普通话声学模型。由此,基于不同地区的用户的口音差异,建立带有方言口音的普通话声学模
型,提高了语音识别的性能。
基于以上实施例,为了进一步保证语音识别处理方法的实用性,在生成与各省份对应的带有方言口音的普通话声学模型之后,还可对生成的带有方言口音的声学模型进行性能的验证,从而只对相较于普通声学模型性能得到提升的,带有方言口音的普通话声学模型部署上线。
图3是根据本发明另一个实施例的语音识别处理方法的流程图,如图3所示,该方法包括:
S310,根据全国所有地区的语音样本数据在预设的深度神经网络模型上进行训练,生成基于深层长短时记忆单元的模型结构,以及连接时序分类框架的声学模型。
在本发明的一个实施例中,可预先设置深度神经网络模型,该深度神经网络模型的输入为单帧或多帧拼接的语音声学特征,输出为上下文相关的声韵母单元,即基于输入声学特征对上下文相关的声韵母单元的分类,以生成相关声学模型。
具体而言,将全国所有地区的语音样本数据输入该深度神经网络模型进行训练,基于输入语音样本数据的声学特征,对上下文相关的声韵母单元的分类等训练处理,生成基于深层长短时记忆单元的模型结构,以及连接时序分类框架的声学模型。
S320,根据各省份的语音样本数据,分别在通用普通话声学模型上进行自适应训练,生成与各省份对应的带有方言口音的普通话声学模型。
需要说明的是,根据具体应用场景的不同,可采用多种自适应训练方式在通用普通话声学模型上进行自适应训练:
第一种示例,可采用较小学习率调优基本模型的自适应训练方式,在通用普通话声学模型上进行自适应训练。
在本示例中,在对通用普通话声学模型调优时,利用带有口音的语音样本数据在通用普通话声学模型上采用较低的学习率进行微调。
而由于目前通用普通话声学模型的线上模型参数过大,一般小数据量学习容易造成模型过拟合,推广性不强,因此在进行自适应训练时,可采用L2范数正则化以及KL散度正则化的自适应更新方法,进行自适应训练。
其中,α表示正则项系数,ot表示第t帧样本的特征,qt表示第t帧样本对应的标记,W表示模型参数,W0表示当前模型参数。在KL散步正则下,目标的概率表达式是需要更新模型的分布和自适应数据的真实标记下的分布的线性插值。
第二种示例,可采用只调优部分模型参数的自适应训练方式,在通用普通话声学模型上进行自适应训练。
在本示例中,保持住大部分的模型参数与原有的通用模型一致,只对输出层或者隐层的偏置进行调整。并且由于更新的参数规模不大,一般不容易过拟合。
在具体实施过程中,可采用只更新输出层的参数,以及采用奇异值分解的方法加入深瓶颈层以进行较少参数的自适应更新,从而减少自适应模型需要更新的模型参数量。
第三种示例,可采用引入新特征的自适应训练方式,在通用普通话声学模型上进行自适应训练。
本示例中的自适应训练方式考虑到方言口音的特殊性,引入在声纹和自适应训练中较为经典的ivector和说话人编码的方式,通过对每一个方言语音进行包含各类复杂信息的特征矢量提取,将其加入到输入特征进行中自适应训练。
其中,在ivector的提取中,通过采用实时的ivector提取方法,在实际解码中,对每一个包的语音数据提取出相应的ivector矢量进行解码。具体而言,可使用公式M=m+Tw提取ivector。
其中M是所有训练语料的均值超矢量,m是目标语音的积累到当前包数据的均值超矢量,T是载荷矩阵,w则是需要得到的ivector。
在得到当前语料数据中的ivector之后,每一帧特征将拼接上该ivector特征,形成新的特征进而重新训练声学模型。在训练过程中,只更新ivector特征部分的模型参数权重,而保持原有的模型参数不变,以保证模型不会过拟合,同时保证更新后的模型与原有模型不会变化太多,保证生成的带有方言口音的普通话声学模型的推广效果。
S330,根据各省份的语音测试数据,分别对通用普通话声学模型,以及带有方言口音的普通话声学模型进行测试评估。
具体地,生成的与各省份对应的带有方言口音的普通话声学模型,并不过于偏向通用普通话声学模型,且在实际应用时,有可能带有方言口音的普通话声学模型的性能并不高于通用普通话声学模型。
因此,为了保证部署上线的声学模型的性能得到提升,需要根据各省份的语音测试数据,分别对通用普通话声学模型,以及带有方言口音的普通话声学模型进行测试评估。
比如,分别向通用普通话声学模型和带有河南方言口音的普通话声学模型,输入河南口音的语音测试数据,根据通用普通话声学模型和带有河南方言口音的普通话声学模型语音识别的准确率,对其性能进行测试评估。
S340,如果带有方言口音的普通话声学模型的识别性能高于通用普通话声学模型,则将带有方言口音的普通话声学模型部署上线。
具体地,如果带有方言口音的普通话声学模型的识别性能高于通用普通话声学模型,则表明该带有方言口音的普通话声学模型,相较于通用普通话声学模型能够更加准确地识别带有口音的普通话,因而将带有方言口音的普通话声学模型部署上线。
综上所述,本发明实施例的语音识别处理方法,根据各省份的语音测试数据,分别对通用普通话声学模型,以及带有方言口音的普通话声学模型进行测试评估,如果带有方言口音的普通话声学模型的识别性能高于通用普通话声学模型,则将带有方言口音的普通话声学模型部署上线。由此,将识别性能高于通用普通话声学模型的带有方言口音的普通话声学模型部署上线,保证了语音识别处理方法的实用性。
基于以上描述,在实际应用中,将带有方言口音的普通话声学模型部署上线之后,可采用多种方式确定用户所属的方言口音,以根据与该方言口音对应的普通话声学模型,对用户输入的语音信息进行识别。
第一种示例,可以获取用户的个人信息,根据个人信息中的籍贯所属省份,确定用户所属方言口音,以便根据与该方言口音对应的普通话声学模型,对用户输入的语音信息进行识别。
第二种示例,可以获取用户发出语音识别请求所属的网络地址信息,确定该网络地址信息所属的省份,以获取用户所属方言口音,从而可根据与该方言口音对应的普通话声学模型,对用户输入的语音信息进行识别。
为了更加清楚的说明,如何确定用户所属的方言口音,以根据与该方言口音对应的普通话声学模型,对用户输入的语音信息进行识别,下面结合附图4,基于以上第二种示例的具体实施过程,进行举例说明:
图4是根据本发明又一个实施例的语音识别处理方法的流程图,如图4所示,在如图3所示的步骤S340后,该方法包括:
S410,接收用户发送的携带网络地址信息的语音信息。
S420,根据网络地址信息确定用户所在的省份信息。
具体地,可接收用户发送的携带网络地址信息的语音信息,进而可根据该网络地址信息查询确定其所在的省份,比如,可根据网络地址信息中的IP地址确定其所属的省份信息等。
S430,判断是否部署有与省份信息对应的带有方言口音的普通话声学模型。
S440,如果部署,则将语音信息输入到与省份信息对应的带有方言口音的普通话声学模型进行语音识别。
S450,如果没有部署,则将语音信息输入到通用普通话声学模型进行语音识别。
具体地,在确定用户所在的省份信息后,可判断是否部署有与省份信息对应的带有方言口音的普通话声学模型,如果部署,则表明存在语音识别性能高于普通话声学模型的,与省份信息对应的带有方言口音的普通话声学模型,因而将语音信息输入到与省份信息对应的带有方言口音的普通话声学模型进行语音识别。
如果没有部署,则表明没有语音识别性能高于普通话声学模型的,与省份信息对应的带有方言口音的普通话声学模型,因而将语音信息输入到通用普通话声学模型进行语音识别。
综上所述,本发明实施例的语音识别处理方法,根据用户发送的携带网络地址信息的语音信息,确定用户所在的省份信息,并在部署有与该省份信息对应的带有方言口音的普通话声学模型时,使用该带有方言口音的普通话声学模型识别用户的语音信息。由此,提高了语音识别的性能。
为了实现上述实施例,本发明还提出了一种语音识别处理装置,图5是根据本发明一个实施例的语音识别处理装置的结构示意图,如图5所示,该装置包括:第一生成模块10和第二生成模块20。
其中,第一生成模块10,用于根据全国所有地区的语音样本数据在预设的处理模型上进行训练,生成通用普通话声学模型。
具体地,预设一训练生成普通话声学模型的处理模型,比如预设一深度神经网络模型等,进而采集全国所有地区的语音样本数据,将该语音样本数据输入预设的处理模型。
进而,第一生成模块10通过处理模型提取语音样本数据中的语音特征,将语音特征映射到语言基本单元,生成通用普通话声学模型,基于该通用普通话声学模型可实现对全国用户的语音的识别。
第二生成模块20,用于根据各省份的语音样本数据,分别在通用普通话声学模型上进行自适应训练,生成与各省份对应的带有方言口音的普通话声学模型。
具体地,在实际应用中,采集全国各省份的语音样本数据作为自适应数据,其中,每个省份所采集的语音样本数据,可能数量相对较少,比如可能为几百个小时的语音数量级,进而第二生成模块20基于每个省份所采集的语音样本数据,分别在通用普通话声学模型上进行自适应训练,为各个省份进行自适应训练得到对应的普通话声学模型。
需要说明的是,前述对语音识别处理方法实施例的解释说明也适用于该实施例的语音
识别处理装置,其实现原理类似,此处不再赘述。
综上所述,本发明实施例的语音识别处理装置,根据全国所有地区的语音样本数据在预设的处理模型上进行训练,生成通用普通话声学模型,并根据各省份的语音样本数据,分别在通用普通话声学模型上进行自适应训练,生成与各省份对应的带有方言口音的普通话声学模型。由此,基于不同地区的用户的口音差异,建立带有方言口音的普通话声学模型,提高了语音识别的性能。
基于以上实施例,为了进一步保证语音识别处理方法的实用性,在生成与各省份对应的带有方言口音的普通话声学模型之后,还可对生成的带有方言口音的声学模型进行性能的验证,从而只对相较于普通声学模型性能得到提升的,带有方言口音的普通话声学模型部署上线。
图6是根据本发明另一个实施例的语音识别处理装置的结构示意图,如图6所示,在如图5所示的基础上,该装置还包括:评估模块30和部署模块40。
其中,评估模块30,用于根据各省份的语音测试数据,分别对通用普通话声学模型,以及带有方言口音的普通话声学模型进行测试评估。
部署模块40,用于在带有方言口音的普通话声学模型的识别性能高于通用普通话声学模型时,将带有方言口音的普通话声学模型部署上线。
在本发明的一个实施例中,第一生成模块10还将全国所有地区的语音样本数据输入该深度神经网络模型进行训练,基于输入语音样本数据的声学特征,对上下文相关的声韵母单元的分类等训练处理,生成基于深层长短时记忆单元的模型结构,以及连接时序分类框架的声学模型。
进而,第二生成模块20可采用较小学习率调优基本模型的自适应训练方式、只调优部分模型参数的自适应训练方式、引入新特征的自适应训练方式等在通用普通话声学模型上进行自适应训练,以生成带有方言口音的普通话声学模型。
为了保证部署上线的声学模型的性能得到提升,评估模块30需要根据各省份的语音测试数据,分别对通用普通话声学模型,以及带有方言口音的普通话声学模型进行测试评估。
进一步地,如果带有方言口音的普通话声学模型的识别性能高于通用普通话声学模型,则表明该带有方言口音的普通话声学模型,相较于通用普通话声学模型能够更加准确地识别带有口音的普通话,因而部署模块40将带有方言口音的普通话声学模型部署上线。
需要说明的是,前述对语音识别处理方法实施例的解释说明也适用于该实施例的语音识别处理装置,其实现原理类似,此处不再赘述。
综上所述,本发明实施例的语音识别处理装置,根据各省份的语音测试数据,分别对通用普通话声学模型,以及带有方言口音的普通话声学模型进行测试评估,如果带有方言
口音的普通话声学模型的识别性能高于通用普通话声学模型,则将带有方言口音的普通话声学模型部署上线。由此,将识别性能高于通用普通话声学模型的带有方言口音的普通话声学模型部署上线,保证了语音识别处理方法的实用性。
基于以上描述,在实际应用中,将带有方言口音的普通话声学模型部署上线之后,可采用多种方式确定用户所属的方言口音,以根据与该方言口音对应的普通话声学模型,对用户输入的语音信息进行识别。
图7是根据本发明又一个实施例的语音识别处理装置的结构示意图,如图7所示,在如图6所示的基础上,该装置还包括:接收模块50、确定模块60、判断模块70和处理模块80。
其中,接收模块50,用于接收用户发送的携带网络地址信息的语音信息。
确定模块60,用于根据网络地址信息确定用户所在的省份信息。
具体地,接收模块50可接收用户发送的携带网络地址信息的语音信息,进而确定模块60可根据该网络地址信息查询确定其所在的省份,比如,可根据网络地址信息中的IP地址确定其所属的省份信息等。
判断模块70,用于判断是否部署有与省份信息对应的带有方言口音的普通话声学模型。
处理模块80,用于在部署有与省份信息对应的带有方言口音的普通话声学模型时,将语音信息输入到与省份信息对应的带有方言口音的普通话声学模型进行语音识别。
处理模块80还用于在没有部署与省份信息对应的带有方言口音的普通话声学模型时,将语音信息输入到通用普通话声学模型进行语音识别。
具体地,在确定用户所在的省份信息后,判断模块70可判断是否部署有与省份信息对应的带有方言口音的普通话声学模型,如果部署,则表明存在语音识别性能高于普通话声学模型的,与省份信息对应的带有方言口音的普通话声学模型,因而处理模块80将语音信息输入到与省份信息对应的带有方言口音的普通话声学模型进行语音识别。
如果没有部署,则表明没有语音识别性能高于普通话声学模型的,与省份信息对应的带有方言口音的普通话声学模型,因而处理模块80将语音信息输入到通用普通话声学模型进行语音识别。
需要说明的是,前述对语音识别处理方法实施例的解释说明也适用于该实施例的语音识别处理装置,其实现原理类似,此处不再赘述。
综上所述,本发明实施例的语音识别处理装置,根据用户发送的携带网络地址信息的语音信息,确定用户所在的省份信息,并在部署有与该省份信息对应的带有方言口音的普通话声学模型时,使用该带有方言口音的普通话声学模型识别用户的语音信息。由此,提高了语音识别的性能。
在本说明书的描述中,参考术语“一个实施例”、“一些实施例”、“示例”、“具体示例”、或“一些示例”等的描述意指结合该实施例或示例描述的具体特征、结构、材料或者特点包含于本发明的至少一个实施例或示例中。在本说明书中,对上述术语的示意性表述不必须针对的是相同的实施例或示例。而且,描述的具体特征、结构、材料或者特点可以在任一个或多个实施例或示例中以合适的方式结合。此外,在不相互矛盾的情况下,本领域的技术人员可以将本说明书中描述的不同实施例或示例以及不同实施例或示例的特征进行结合和组合。
尽管上面已经示出和描述了本发明的实施例,可以理解的是,上述实施例是示例性的,不能理解为对本发明的限制,本领域的普通技术人员在本发明的范围内可以对上述实施例进行变化、修改、替换和变型。
Claims (12)
- 一种语音识别处理方法,其特征在于,包括以下步骤:根据全国所有地区的语音样本数据在预设的处理模型上进行训练,生成通用普通话声学模型;根据各省份的语音样本数据,分别在所述通用普通话声学模型上进行自适应训练,生成与各省份对应的带有方言口音的普通话声学模型。
- 如权利要求1所述的方法,其特征在于,所述根据全国所有地区的语音样本数据在预设的处理模型上进行训练,生成通用普通话声学模型,包括:根据全国所有地区的语音样本数据在预设的深度神经网络模型上进行训练,生成基于深层长短时记忆单元的模型结构,以及连接时序分类框架的声学模型。
- 如权利要求1-2任一所述的方法,其特征在于,所述分别在所述通用普通话声学模型上进行自适应训练,包括:较小学习率调优基本模型的自适应训练方式;或者,只调优部分模型参数的自适应训练方式;或者,引入新特征的自适应训练方式。
- 如权利要求1-3任一所述的方法,其特征在于,在所述生成与各省份对应的带有方言口音的普通话声学模型之后,还包括:根据各省份的语音测试数据,分别对所述通用普通话声学模型,以及所述带有方言口音的普通话声学模型进行测试评估;如果所述带有方言口音的普通话声学模型的识别性能高于所述通用普通话声学模型,则将所述带有方言口音的普通话声学模型部署上线。
- 如权利要求1-4任一所述的方法,其特征在于,在所述将所述带有方言口音的普通话声学模型部署上线之后,还包括:接收用户发送的携带网络地址信息的语音信息;根据所述网络地址信息确定所述用户所在的省份信息;判断是否部署有与所述省份信息对应的带有方言口音的普通话声学模型;如果部署,则将所述语音信息输入到与所述省份信息对应的带有方言口音的普通话声学模型进行语音识别;如果没有部署,则将所述语音信息输入到所述通用普通话声学模型进行语音识别。
- 一种语音识别处理装置,其特征在于,包括:第一生成模块,用于根据全国所有地区的语音样本数据在预设的处理模型上进行训练, 生成通用普通话声学模型;第二生成模块,用于根据各省份的语音样本数据,分别在所述通用普通话声学模型上进行自适应训练,生成与各省份对应的带有方言口音的普通话声学模型。
- 如权利要求6所述的装置,其特征在于,所述第一生成模块用于:根据全国所有地区的语音样本数据在预设的深度神经网络模型上进行训练,生成基于深层长短时记忆单元的模型结构,以及连接时序分类框架的声学模型。
- 如权利要求6或7所述的装置,其特征在于,所述第二生成模块分别在所述通用普通话声学模型上进行自适应训练,包括:较小学习率调优基本模型的自适应训练方式;或者,只调优部分模型参数的自适应训练方式;或者,引入新特征的自适应训练方式。
- 如权利要求6-8任一所述的方法,其特征在于,还包括:评估模块,用于根据各省份的语音测试数据,分别对所述通用普通话声学模型,以及所述带有方言口音的普通话声学模型进行测试评估;部署模块,用于在所述带有方言口音的普通话声学模型的识别性能高于所述通用普通话声学模型时,将所述带有方言口音的普通话声学模型部署上线。
- 如权利要求6-9任一所述的装置,其特征在于,还包括:接收模块,用于接收用户发送的携带网络地址信息的语音信息;确定模块,用于根据所述网络地址信息确定所述用户所在的省份信息;判断模块,用于判断是否部署有与所述省份信息对应的带有方言口音的普通话声学模型;处理模块,用于在部署有与所述省份信息对应的带有方言口音的普通话声学模型时,将所述语音信息输入到与所述省份信息对应的带有方言口音的普通话声学模型进行语音识别;所述处理模块还用于在没有部署与所述省份信息对应的带有方言口音的普通话声学模型时,将所述语音信息输入到所述通用普通话声学模型进行语音识别。
- 一种设备,其特征在于,包括:一个或者多个处理器;存储器;一个或者多个程序,所述一个或者多个程序存储在所述存储器中,当被所述一个或者多个处理器执行时,执行以下步骤:根据全国所有地区的语音样本数据在预设的处理模型上进行训练,生成通用普通 话声学模型;根据各省份的语音样本数据,分别在所述通用普通话声学模型上进行自适应训练,生成与各省份对应的带有方言口音的普通话声学模型。
- 一种非易失性计算机存储介质,其特征在于,所述计算机存储介质存储有一个或者多个程序,当所述一个或者多个程序被一个设备执行时,使得所述设备执行以下步骤:根据全国所有地区的语音样本数据在预设的处理模型上进行训练,生成通用普通话声学模型;根据各省份的语音样本数据,分别在所述通用普通话声学模型上进行自适应训练,生成与各省份对应的带有方言口音的普通话声学模型。
Priority Applications (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP2019502659A JP6774551B2 (ja) | 2016-07-22 | 2016-11-08 | 音声認識処理方法及び装置 |
| US16/318,809 US11138967B2 (en) | 2016-07-22 | 2016-11-08 | Voice recognition processing method, device and computer storage medium |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN201610585160.6A CN106251859B (zh) | 2016-07-22 | 2016-07-22 | 语音识别处理方法和装置 |
| CN201610585160.6 | 2016-07-22 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2018014469A1 true WO2018014469A1 (zh) | 2018-01-25 |
Family
ID=57604542
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2016/105080 Ceased WO2018014469A1 (zh) | 2016-07-22 | 2016-11-08 | 语音识别处理方法和装置 |
Country Status (4)
| Country | Link |
|---|---|
| US (1) | US11138967B2 (zh) |
| JP (1) | JP6774551B2 (zh) |
| CN (1) | CN106251859B (zh) |
| WO (1) | WO2018014469A1 (zh) |
Cited By (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN110223674A (zh) * | 2019-04-19 | 2019-09-10 | 平安科技(深圳)有限公司 | 语音语料训练方法、装置、计算机设备和存储介质 |
| CN112382266A (zh) * | 2020-10-30 | 2021-02-19 | 北京有竹居网络技术有限公司 | 一种语音合成方法、装置、电子设备及存储介质 |
| CN113889122A (zh) * | 2021-09-29 | 2022-01-04 | 马上消费金融股份有限公司 | 声纹识别模型训练方法、声纹识别方法及相关设备 |
| WO2026060321A1 (en) | 2024-09-13 | 2026-03-19 | Olema Pharmaceuticals, Inc. | Kat6 degrader compounds and method thereof |
Families Citing this family (50)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN108281137A (zh) * | 2017-01-03 | 2018-07-13 | 中国科学院声学研究所 | 一种全音素框架下的通用语音唤醒识别方法及系统 |
| CN108269568B (zh) * | 2017-01-03 | 2021-07-30 | 中国科学院声学研究所 | 一种基于ctc的声学模型训练方法 |
| CN106887226A (zh) * | 2017-04-07 | 2017-06-23 | 天津中科先进技术研究院有限公司 | 一种基于人工智能识别的语音识别算法 |
| US10446136B2 (en) * | 2017-05-11 | 2019-10-15 | Ants Technology (Hk) Limited | Accent invariant speech recognition |
| CN107481717B (zh) * | 2017-08-01 | 2021-03-19 | 百度在线网络技术(北京)有限公司 | 一种声学模型训练方法及系统 |
| CN107909715A (zh) * | 2017-09-29 | 2018-04-13 | 嘉兴川森智能科技有限公司 | 自动售货机中的语音识别系统及方法 |
| CN108039168B (zh) * | 2017-12-12 | 2020-09-11 | 科大讯飞股份有限公司 | 声学模型优化方法及装置 |
| CN108417203A (zh) * | 2018-01-31 | 2018-08-17 | 广东聚晨知识产权代理有限公司 | 一种人体语音识别传输方法及系统 |
| CN108735199B (zh) * | 2018-04-17 | 2021-05-28 | 北京声智科技有限公司 | 一种声学模型的自适应训练方法及系统 |
| CN108670128A (zh) * | 2018-05-21 | 2018-10-19 | 深圳市沃特沃德股份有限公司 | 语音控制扫地机器人的方法和扫地机器人 |
| CN110600032A (zh) * | 2018-05-23 | 2019-12-20 | 北京语智科技有限公司 | 一种语音识别方法及装置 |
| CN118737132A (zh) * | 2018-07-13 | 2024-10-01 | 谷歌有限责任公司 | 端到端流关键词检出 |
| CN108877784B (zh) * | 2018-09-05 | 2022-12-06 | 河海大学 | 一种基于口音识别的鲁棒语音识别方法 |
| CN109243461B (zh) * | 2018-09-21 | 2020-04-14 | 百度在线网络技术(北京)有限公司 | 语音识别方法、装置、设备及存储介质 |
| CN110941188A (zh) * | 2018-09-25 | 2020-03-31 | 珠海格力电器股份有限公司 | 智能家居控制方法及装置 |
| CN111063338B (zh) * | 2018-09-29 | 2023-09-19 | 阿里巴巴集团控股有限公司 | 音频信号识别方法、装置、设备、系统和存储介质 |
| CN109291049B (zh) * | 2018-09-30 | 2021-03-05 | 北京木业邦科技有限公司 | 数据处理方法、装置及控制设备 |
| CN111107380B (zh) * | 2018-10-10 | 2023-08-15 | 北京默契破冰科技有限公司 | 一种用于管理音频数据的方法、设备和计算机存储介质 |
| CN111031329B (zh) * | 2018-10-10 | 2023-08-15 | 北京默契破冰科技有限公司 | 一种用于管理音频数据的方法、设备和计算机存储介质 |
| KR102718582B1 (ko) * | 2018-10-19 | 2024-10-17 | 삼성전자주식회사 | 음성을 인식하는 장치 및 방법, 음성 인식 모델을 트레이닝하는 장치 및 방법 |
| CN109346059B (zh) * | 2018-12-20 | 2022-05-03 | 广东小天才科技有限公司 | 一种方言语音的识别方法及电子设备 |
| CN109979439B (zh) * | 2019-03-22 | 2021-01-29 | 泰康保险集团股份有限公司 | 基于区块链的语音识别方法、装置、介质及电子设备 |
| CN109887497B (zh) * | 2019-04-12 | 2021-01-29 | 北京百度网讯科技有限公司 | 语音识别的建模方法、装置及设备 |
| CN110033760B (zh) * | 2019-04-15 | 2021-01-29 | 北京百度网讯科技有限公司 | 语音识别的建模方法、装置及设备 |
| CN111354349A (zh) * | 2019-04-16 | 2020-06-30 | 深圳市鸿合创新信息技术有限责任公司 | 一种语音识别方法及装置、电子设备 |
| CN110047467B (zh) * | 2019-05-08 | 2021-09-03 | 广州小鹏汽车科技有限公司 | 语音识别方法、装置、存储介质及控制终端 |
| CN112116909A (zh) * | 2019-06-20 | 2020-12-22 | 杭州海康威视数字技术股份有限公司 | 语音识别方法、装置及系统 |
| CN112133290A (zh) * | 2019-06-25 | 2020-12-25 | 南京航空航天大学 | 一种针对民航陆空通话领域的基于迁移学习的语音识别方法 |
| CN110349571B (zh) * | 2019-08-23 | 2021-09-07 | 北京声智科技有限公司 | 一种基于连接时序分类的训练方法及相关装置 |
| CN110570837B (zh) * | 2019-08-28 | 2022-03-11 | 卓尔智联(武汉)研究院有限公司 | 一种语音交互方法、装置及存储介质 |
| CN110534116B (zh) * | 2019-08-29 | 2022-06-03 | 北京安云世纪科技有限公司 | 应用于智能设备的语音识别模型设置方法及装置 |
| CN110930995B (zh) * | 2019-11-26 | 2022-02-11 | 中国南方电网有限责任公司 | 一种应用于电力行业的语音识别模型 |
| CN110956954B (zh) * | 2019-11-29 | 2020-12-11 | 百度在线网络技术(北京)有限公司 | 一种语音识别模型训练方法、装置以及电子设备 |
| CN111477234A (zh) * | 2020-03-05 | 2020-07-31 | 厦门快商通科技股份有限公司 | 一种声纹数据注册方法和装置以及设备 |
| CN111599349B (zh) * | 2020-04-01 | 2023-04-18 | 云知声智能科技股份有限公司 | 一种训练语言模型的方法及系统 |
| CN111816165A (zh) * | 2020-07-07 | 2020-10-23 | 北京声智科技有限公司 | 语音识别方法、装置及电子设备 |
| CN112331182B (zh) * | 2020-10-26 | 2024-07-09 | 平安科技(深圳)有限公司 | 语音数据生成方法、装置、计算机设备及存储介质 |
| CN112614485A (zh) * | 2020-12-30 | 2021-04-06 | 竹间智能科技(上海)有限公司 | 识别模型构建方法、语音识别方法、电子设备及存储介质 |
| CN112802455B (zh) * | 2020-12-31 | 2023-04-11 | 北京捷通华声科技股份有限公司 | 语音识别方法及装置 |
| CN113593525B (zh) * | 2021-01-26 | 2024-08-06 | 腾讯科技(深圳)有限公司 | 口音分类模型训练和口音分类方法、装置和存储介质 |
| CN113223542B (zh) * | 2021-04-26 | 2024-04-12 | 北京搜狗科技发展有限公司 | 音频的转换方法、装置、存储介质及电子设备 |
| CN113345451B (zh) * | 2021-04-26 | 2023-08-22 | 北京搜狗科技发展有限公司 | 一种变声方法、装置及电子设备 |
| CN113192491B (zh) * | 2021-04-28 | 2024-05-03 | 平安科技(深圳)有限公司 | 声学模型生成方法、装置、计算机设备及存储介质 |
| CN113593534B (zh) * | 2021-05-28 | 2023-07-14 | 思必驰科技股份有限公司 | 针对多口音语音识别的方法和装置 |
| CN114254649A (zh) * | 2021-12-15 | 2022-03-29 | 科大讯飞股份有限公司 | 一种语言模型的训练方法、装置、存储介质及设备 |
| CN114596845A (zh) * | 2022-04-13 | 2022-06-07 | 马上消费金融股份有限公司 | 语音识别模型的训练方法、语音识别方法及装置 |
| CN115497453A (zh) * | 2022-08-31 | 2022-12-20 | 海尔优家智能科技(北京)有限公司 | 识别模型的评估方法及装置、存储介质及电子装置 |
| CN115496916B (zh) * | 2022-09-30 | 2023-08-22 | 北京百度网讯科技有限公司 | 图像识别模型的训练方法、图像识别方法以及相关装置 |
| CN116935832A (zh) * | 2023-04-24 | 2023-10-24 | 广西壮族自治区通信产业服务有限公司技术服务分公司 | 一种基于强化学习的轻量化多方言识别方法和系统 |
| CN119049507B (zh) * | 2024-07-26 | 2025-11-18 | 浙江大学 | 一种采用声学单位的汉语方言口音矫正方法和系统 |
Citations (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN103310788A (zh) * | 2013-05-23 | 2013-09-18 | 北京云知声信息技术有限公司 | 一种语音信息识别方法及系统 |
| CN104036774A (zh) * | 2014-06-20 | 2014-09-10 | 国家计算机网络与信息安全管理中心 | 藏语方言识别方法及系统 |
| US20140288928A1 (en) * | 2013-03-25 | 2014-09-25 | Gerald Bradley PENN | System and method for applying a convolutional neural network to speech recognition |
| CN105229725A (zh) * | 2013-03-11 | 2016-01-06 | 微软技术许可有限责任公司 | 多语言深神经网络 |
| CN105336323A (zh) * | 2015-10-14 | 2016-02-17 | 清华大学 | 维语语音识别方法和装置 |
| CN105632501A (zh) * | 2015-12-30 | 2016-06-01 | 中国科学院自动化研究所 | 一种基于深度学习技术的自动口音分类方法及装置 |
Family Cites Families (16)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPH071435B2 (ja) * | 1993-03-16 | 1995-01-11 | 株式会社エイ・ティ・アール自動翻訳電話研究所 | 音響モデル適応方式 |
| US20080147404A1 (en) * | 2000-05-15 | 2008-06-19 | Nusuara Technologies Sdn Bhd | System and methods for accent classification and adaptation |
| US8892443B2 (en) * | 2009-12-15 | 2014-11-18 | At&T Intellectual Property I, L.P. | System and method for combining geographic metadata in automatic speech recognition language and acoustic models |
| US8468012B2 (en) * | 2010-05-26 | 2013-06-18 | Google Inc. | Acoustic model adaptation using geographic information |
| US9966064B2 (en) * | 2012-07-18 | 2018-05-08 | International Business Machines Corporation | Dialect-specific acoustic language modeling and speech recognition |
| US9153231B1 (en) * | 2013-03-15 | 2015-10-06 | Amazon Technologies, Inc. | Adaptive neural network speech recognition models |
| JP5777178B2 (ja) | 2013-11-27 | 2015-09-09 | 国立研究開発法人情報通信研究機構 | 統計的音響モデルの適応方法、統計的音響モデルの適応に適した音響モデルの学習方法、ディープ・ニューラル・ネットワークを構築するためのパラメータを記憶した記憶媒体、及び統計的音響モデルの適応を行なうためのコンピュータプログラム |
| CN103680493A (zh) * | 2013-12-19 | 2014-03-26 | 百度在线网络技术(北京)有限公司 | 区分地域性口音的语音数据识别方法和装置 |
| US10319374B2 (en) * | 2015-11-25 | 2019-06-11 | Baidu USA, LLC | Deployed end-to-end speech recognition |
| US10229672B1 (en) * | 2015-12-31 | 2019-03-12 | Google Llc | Training acoustic models using connectionist temporal classification |
| US20180018973A1 (en) * | 2016-07-15 | 2018-01-18 | Google Inc. | Speaker verification |
| US10431206B2 (en) * | 2016-08-22 | 2019-10-01 | Google Llc | Multi-accent speech recognition |
| US10629192B1 (en) * | 2018-01-09 | 2020-04-21 | Electronic Arts Inc. | Intelligent personalized speech recognition |
| US11170761B2 (en) * | 2018-12-04 | 2021-11-09 | Sorenson Ip Holdings, Llc | Training of speech recognition systems |
| US10388272B1 (en) * | 2018-12-04 | 2019-08-20 | Sorenson Ip Holdings, Llc | Training speech recognition systems using word sequences |
| US10839788B2 (en) * | 2018-12-13 | 2020-11-17 | i2x GmbH | Systems and methods for selecting accent and dialect based on context |
-
2016
- 2016-07-22 CN CN201610585160.6A patent/CN106251859B/zh active Active
- 2016-11-08 US US16/318,809 patent/US11138967B2/en not_active Expired - Fee Related
- 2016-11-08 JP JP2019502659A patent/JP6774551B2/ja not_active Expired - Fee Related
- 2016-11-08 WO PCT/CN2016/105080 patent/WO2018014469A1/zh not_active Ceased
Patent Citations (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN105229725A (zh) * | 2013-03-11 | 2016-01-06 | 微软技术许可有限责任公司 | 多语言深神经网络 |
| US20140288928A1 (en) * | 2013-03-25 | 2014-09-25 | Gerald Bradley PENN | System and method for applying a convolutional neural network to speech recognition |
| CN103310788A (zh) * | 2013-05-23 | 2013-09-18 | 北京云知声信息技术有限公司 | 一种语音信息识别方法及系统 |
| CN104036774A (zh) * | 2014-06-20 | 2014-09-10 | 国家计算机网络与信息安全管理中心 | 藏语方言识别方法及系统 |
| CN105336323A (zh) * | 2015-10-14 | 2016-02-17 | 清华大学 | 维语语音识别方法和装置 |
| CN105632501A (zh) * | 2015-12-30 | 2016-06-01 | 中国科学院自动化研究所 | 一种基于深度学习技术的自动口音分类方法及装置 |
Cited By (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN110223674A (zh) * | 2019-04-19 | 2019-09-10 | 平安科技(深圳)有限公司 | 语音语料训练方法、装置、计算机设备和存储介质 |
| CN110223674B (zh) * | 2019-04-19 | 2023-05-26 | 平安科技(深圳)有限公司 | 语音语料训练方法、装置、计算机设备和存储介质 |
| CN112382266A (zh) * | 2020-10-30 | 2021-02-19 | 北京有竹居网络技术有限公司 | 一种语音合成方法、装置、电子设备及存储介质 |
| CN113889122A (zh) * | 2021-09-29 | 2022-01-04 | 马上消费金融股份有限公司 | 声纹识别模型训练方法、声纹识别方法及相关设备 |
| WO2026060321A1 (en) | 2024-09-13 | 2026-03-19 | Olema Pharmaceuticals, Inc. | Kat6 degrader compounds and method thereof |
Also Published As
| Publication number | Publication date |
|---|---|
| CN106251859A (zh) | 2016-12-21 |
| US11138967B2 (en) | 2021-10-05 |
| CN106251859B (zh) | 2019-05-31 |
| US20190189112A1 (en) | 2019-06-20 |
| JP2019527852A (ja) | 2019-10-03 |
| JP6774551B2 (ja) | 2020-10-28 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| CN106251859B (zh) | 语音识别处理方法和装置 | |
| US11393492B2 (en) | Voice activity detection method, method for establishing voice activity detection model, computer device, and storage medium | |
| CN109147758B (zh) | 一种说话人声音转换方法及装置 | |
| CN108428447B (zh) | 一种语音意图识别方法及装置 | |
| CN110517664B (zh) | 多方言识别方法、装置、设备及可读存储介质 | |
| CN108305641B (zh) | 情感信息的确定方法和装置 | |
| CN109523616B (zh) | 一种面部动画生成方法、装置、设备及可读存储介质 | |
| CN109767778B (zh) | 一种融合Bi-LSTM和WaveNet的语音转换方法 | |
| CN111048064A (zh) | 基于单说话人语音合成数据集的声音克隆方法及装置 | |
| CN110349597B (zh) | 一种语音检测方法及装置 | |
| WO2017076211A1 (zh) | 基于语音的角色分离方法及装置 | |
| CN108319666A (zh) | 一种基于多模态舆情分析的供电服务评估方法 | |
| WO2013020329A1 (zh) | 参数语音合成方法和系统 | |
| CN106847259B (zh) | 一种音频关键词模板的筛选和优化方法 | |
| CN110211594B (zh) | 一种基于孪生网络模型和knn算法的说话人识别方法 | |
| CN102280106A (zh) | 用于移动通信终端的语音网络搜索方法及其装置 | |
| CN114882868A (zh) | 语音合成、情绪迁移、交互方法、存储介质、程序产品 | |
| US20020026309A1 (en) | Speech processing system | |
| CN103559289B (zh) | 语种无关的关键词检索方法及系统 | |
| CN110751941B (zh) | 语音合成模型的生成方法、装置、设备及存储介质 | |
| CN112331207A (zh) | 服务内容监控方法、装置、电子设备和存储介质 | |
| CN115101058A (zh) | 一种语音数据处理方法、装置、存储介质及设备 | |
| CN111833842B (zh) | 合成音模板发现方法、装置以及设备 | |
| CN117475989A (zh) | 一种少量数据的自动训练的音色克隆方法 | |
| CN108831486B (zh) | 基于dnn与gmm模型的说话人识别方法 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 16909396 Country of ref document: EP Kind code of ref document: A1 |
|
| ENP | Entry into the national phase |
Ref document number: 2019502659 Country of ref document: JP Kind code of ref document: A |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 16909396 Country of ref document: EP Kind code of ref document: A1 |