WO2017113739A1 - 一种语音识别方法及装置 - Google Patents
一种语音识别方法及装置 Download PDFInfo
- Publication number
- WO2017113739A1 WO2017113739A1 PCT/CN2016/089579 CN2016089579W WO2017113739A1 WO 2017113739 A1 WO2017113739 A1 WO 2017113739A1 CN 2016089579 W CN2016089579 W CN 2016089579W WO 2017113739 A1 WO2017113739 A1 WO 2017113739A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- gauss
- cluster
- soft
- gaussian
- clustering
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/08—Speech classification or search
- G10L15/10—Speech classification or search using distance or distortion measures between unknown speech and reference templates
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/08—Speech classification or search
- G10L15/14—Speech classification or search using statistical models, e.g. Hidden Markov Models [HMMs]
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/06—Creation of reference templates; Training of speech recognition systems, e.g. adaptation to the characteristics of the speaker's voice
- G10L15/063—Training
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/08—Speech classification or search
- G10L15/14—Speech classification or search using statistical models, e.g. Hidden Markov Models [HMMs]
- G10L15/142—Hidden Markov Models [HMMs]
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/08—Speech classification or search
- G10L15/18—Speech classification or search using natural language modelling
- G10L15/183—Speech classification or search using natural language modelling using context dependencies, e.g. language models
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L25/00—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
- G10L25/27—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the analysis technique
- G10L25/39—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the analysis technique using genetic algorithms
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/06—Creation of reference templates; Training of speech recognition systems, e.g. adaptation to the characteristics of the speaker's voice
- G10L15/063—Training
- G10L2015/0631—Creating reference templates; Clustering
Definitions
- the patent application relates to voice technology, and in particular to a voice recognition method and device.
- System size is controllable: The number of Gauss in the Gaussian mixture model is easy to control during training.
- the system speed can be controlled: the use of dynamic Gaussian selection technology can greatly reduce the computing time
- Gaussian selection is to cluster all Gauss in the speech recognition system as a member Gauss in the model training stage to form cluster Gauss; in the identification, first evaluate each using acoustic features. Clustering Gauss, the Gaussian member of the Gaussian cluster with high likelihood is selected for further evaluation. The other member Gauss was discarded.
- Traditional Gaussian selection techniques have the following disadvantages:
- Hard clustering is used in clustering, that is, a member Gauss belongs to only one cluster Gauss. Clustering accuracy is low.
- the Gaussian selection at the time of recognition cannot be dynamically updated, resulting in too many members of Gaussian remaining in the calculation, and the recognition speed is slow.
- the speed and accuracy of the assessment can reduce the number of Gaussians in an acoustic model that need to be evaluated in the speech recognition process, which is more accurate and efficient than the conventional Gaussian selection, thereby improving the acoustic model.
- an embodiment of the present invention provides a voice recognition method, including the following steps:
- the soft clustering calculation is performed according to the N Gauss obtained through the model training, and M soft cluster Gaussians are obtained;
- the speech is converted into a feature vector, and the top L soft cluster Gauss with the highest score is calculated according to the feature vector, wherein L is smaller than the M;
- the Gaussian of each member in the L soft clusters is used as the Gaussian in the acoustic model in the speech recognition process, and the acoustic model likelihood is calculated.
- Embodiments of the present invention also provide a voice recognition apparatus, including:
- the soft clustering obtaining module is configured to perform soft clustering calculation according to N Gauss obtained through model training, and obtain M soft clustering Gauss;
- a vector conversion module configured to convert a voice into a feature vector when performing voice recognition
- a selection module configured to calculate a top L soft cluster Gauss with the highest score according to the feature vector, and select each member of the first L soft cluster Gauss as a selected Gauss; the L is smaller than the M ;
- a calculation module configured to calculate Gaussian of the selection module as a Gaussian in the acoustic model in the speech recognition process, and perform acoustic model likelihood calculation.
- the Gaussian score calculation amount of each member in the GMM can be reduced from about 70% of the entire calculation time to 20%, thereby improving the acoustic model likelihood evaluation speed and accuracy, and is particularly suitable.
- Local speech recognition, wake-up, and speech endpoint detection detecting the starting point of speech).
- the cluster Gauss is re-estimated to obtain M soft cluster Gauss.
- each member Gauss can belong to multiple cluster Gauss, which is improved.
- the minimum clustering cost of each cluster Gauss is calculated when the K-means algorithm is used to re-estimate the cluster Gaussian;
- the mean and variance of each cluster Gauss are calculated, and the re-estimated cluster Gauss is obtained;
- the re-estimated cluster Gauss is used as M soft cluster Gauss.
- the value of L is a minimum value that satisfies the following conditions:
- the Y represents the feature vector
- ⁇ is a compression index for the Gaussian "posterior" probability
- G i represents the i-th cluster Gauss
- Y) represents the i-th cluster Gaussian Test "probability.
- the number of Gaussians to be evaluated in the acoustic model in the recognition process is small, and the acoustic model likelihood evaluation speed is improved.
- the step of calculating the top L soft cluster Gauss scores according to the feature vector includes the following substeps:
- the Y represents the feature vector
- ⁇ m represents the mean of the mth soft cluster Gauss
- ⁇ m represents the variance of the mth soft cluster Gauss.
- FIG. 1 is a schematic diagram of a speech recognition system in accordance with some embodiments of the present invention.
- FIG. 3 is a flowchart of a voice recognition method according to a first embodiment
- FIG. 4 is a schematic diagram of dynamic Gaussian selection according to a first embodiment
- an HMM+GMM-based recognition system reads a segment of speech by frame, and the system changes each frame of speech signal into a feature vector.
- the system combines the per-frame feature vector to evaluate the likelihood of each Gaussian in the acoustic model, and assumes the combination of multiple words.
- the combination of these words uses the language model for likelihood evaluation, the sum of acoustic likelihood and language likelihood. The highest word combination is output as the recognition result.
- a first embodiment of the present invention relates to a speech recognition method.
- it is necessary to perform soft clustering calculation based on N Gauss obtained by model training in advance, and obtain M soft cluster Gaussians.
- the number of Gauss members to be calculated is controlled by dynamic Gaussian selection.
- the calculation flow of the soft clustering is as shown in FIG. 2.
- N Gaussians are obtained through model training, such as 1000 Gaussians.
- step 202 N Gaussians are assigned to cluster Gauss by a predetermined weight.
- step 203 the clustering Gauss is re-estimated according to the update weights of the Gaussians of each Gauss pair, and M soft cluster Gaussians are obtained.
- the Gaussian mixture model is used in speech recognition to describe the probability distribution of each state of the Hidden Markov Model (HMM), each state using several Gaussians to express its own probability distribution.
- a Gaussian distribution has its own mean ⁇ and variance ⁇ .
- Gauss In order to effectively use Gaussian selection in the recognition system, Gauss must be shared between states. This acoustic model of shared Gaussian is called a semi-continuous Markov model. In the case of using the same number of Gaussian, semi-continuous Gaussian improves the description ability of the model, thereby increasing the recognition rate.
- N in the local identification system, N generally takes 1000
- Gauss the distance criterion between Gauss must be clarified before clustering.
- weighted symmetric KL divergence (WSKLD) is employed as the distance criterion.
- the SKLD of a distance between Gaussian m and Gauss n is:
- ⁇ n is the mean of Gauss n
- ⁇ m is the mean of Gaussian m
- I is the unit matrix.
- WSKLD is:
- N strm is the number of subspaces of the Gaussian model.
- K-means algorithm For the calculation of soft clustering, in the specific implementation, any of the following algorithms can be used: K-means algorithm, C-means algorithm, self-organizing graph algorithm.
- the K-means algorithm is taken as an example to illustrate the following:
- g(i,n) represents the update weight of the nth Gaussian to the ith cluster Gauss
- ⁇ is the preset cluster softness parameter
- WSKLD represents the weighted symmetric KL dispersion as the Gaussian distance criterion degree.
- the mean of clustering Gauss the variance and the weight of each member Gaussian to update each cluster Gaussian:
- the first step is to get the best update weight:
- the second step is to obtain the best mean and variance of the cluster Gauss based on the optimal weight.
- the method for updating the clustering Gauss mean is as follows:
- an auxiliary matrix Z can be constructed.
- the covariance matrix is limited to a diagonal matrix.
- the imposed condition causes the cluster to not converge, but does not affect the clustering accuracy, so that the re-estimated cluster Gauss is obtained as M soft cluster Gauss.
- the recognition system calculates the minimum clustering cost of each cluster Gauss, and then deducts each minimum clustering cost, thereby obtaining the update weight of each member Gaussian for each cluster Gaussian. Then, according to the update weight, the mean and variance of each cluster Gauss are calculated, and the re-estimated cluster Gauss is obtained as M soft cluster Gauss.
- step 301 the recognition system reads a segment of speech by frame, for example, each frame is 10 milliseconds in length.
- step 302 the recognition system changes each frame of the speech signal into a feature vector, and the resulting feature vector is used to evaluate the soft clustering Gauss.
- step 303 the top L soft cluster Gausss with the highest score (where L is less than M) are calculated from the feature vectors.
- ⁇ m represents the mean of the mth soft cluster Gauss
- ⁇ m represents the variance of the mth soft cluster Gaussian
- the value of L is a minimum value that satisfies the following conditions:
- Y represents the feature vector
- ⁇ is a compression index for the Gaussian "posterior" probability
- G i represents the i-th cluster Gauss
- Y) represents the "a posteriori" of the i-th cluster Gaussian Probability.
- a member Gauss is selected and calculated depends on the member Gaussian and cluster Gauss maps and the cluster Gaussian selection list. As shown in FIG. 4, "1" in the clustering Gaussian selection table indicates that the corresponding cluster Gauss is selected at the current time in the recognition process. In the "cluster-member Gauss map", the Gauss corresponding to the selected cluster Gauss is queried for calculation. The likelihood of the unselected member Gauss is replaced by a small value.
- step 305 it is determined whether there are still unread speech frames. If the result of the determination is yes, indicating that there is a speech frame that needs to be recognized, then returning to step 301, the next speech frame is read to continue the recognition. Otherwise, the voice recognition has been completed, and the process ends.
- step 306 the recognition result is output.
- the result of the speech recognition in this step is the sum of the acoustic likelihood and the language likelihood.
- the hard Gaussian clustering means that each member function belongs to only one cluster Gauss, and the cluster only uses the mean as a vector.
- Soft precision clustering is a method described in some embodiments of the invention.
- a system that does not use Gaussian clustering is used as a baseline. It can be seen that the hard Gaussian cluster is inferior in accuracy to the method of some embodiments of the present invention. The speed of the two is quite the same.
- the baseline system is inferior in speed and accuracy to the system of some embodiments of the present invention.
- the K-Means method is used to perform soft clustering on Gauss in the system training phase (that is, one member Gauss may belong to multiple cluster Gauss), and the number of clusters Gradually increase, and the way each increase reflects the law of model distribution.
- dynamic Gaussian selection is used to control the number of Gauss members that need to be calculated. Thereby the acoustic model likelihood evaluation speed and accuracy are improved. More accurate and efficient than traditional Gaussian selection.
- a second embodiment of the present invention relates to a speech recognition method.
- the second embodiment is substantially the same as the first embodiment, and the main difference is that in the first embodiment, Gaussian soft clustering is performed using a precise K-means algorithm in the system training phase.
- the C-means algorithm is used to perform soft clustering on Gauss in the system training phase.
- the specific implementation of the soft clustering calculation by using the C-means algorithm is basically the same as the K-means algorithm, and will not be described in detail in this embodiment.
- a third embodiment of the present invention relates to a speech recognition method.
- the third embodiment is substantially the same as the first embodiment, and the main difference is that in the first embodiment, Gaussian soft clustering is performed using a precise K-means algorithm in the system training phase.
- the self-organizing graph algorithm is used to perform soft clustering on Gauss in the system training phase.
- the specific implementation of the soft clustering calculation by using the self-organizing graph algorithm is only slightly different in step 203, and the self-organizing graph algorithm is a well-known technology of the existing clustering algorithm, and will not be described in detail in this embodiment.
- a fourth embodiment of the present invention relates to a voice recognition apparatus, as shown in FIG. 5, including:
- the soft clustering obtaining module is configured to perform soft clustering calculation according to N Gauss obtained through model training, and obtain M soft clustering Gauss;
- a vector conversion module configured to convert a voice into a feature vector when performing voice recognition
- a selection module configured to calculate the top L soft cluster Gauss with the highest score according to the feature vector, and select Gauss of each of the first L soft cluster Gauss as the selected Gauss, where L is less than M;
- the calculation module is configured to calculate the acoustic model likelihood by using Gauss selected by the selection module as Gauss in the acoustic model in the speech recognition process.
- the soft cluster acquisition module includes:
- a weight allocation module configured to allocate N Gaussians according to a preset weight to the cluster Gauss
- the re-estimation module is configured to re-estimate the cluster Gauss according to the update weight of each Gaussian of each Gauss pair, and obtain M soft cluster Gauss.
- the present embodiment is a system embodiment corresponding to the first embodiment, and the present embodiment can be implemented in cooperation with the first embodiment.
- the related technical details mentioned in the first embodiment are still effective in the present embodiment, and are not described herein again in order to reduce repetition. Accordingly, the related art details mentioned in the present embodiment can also be applied to the first embodiment.
- each module involved in this embodiment is a logic module.
- a logical unit may be a physical unit, a part of a physical unit, or multiple physical entities. A combination of units is implemented.
- the present embodiment does not introduce a unit that is not closely related to solving the technical problem proposed by the present invention, but this does not mean that there are no other units in the present embodiment.
Landscapes
- Engineering & Computer Science (AREA)
- Physics & Mathematics (AREA)
- Multimedia (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Human Computer Interaction (AREA)
- Computational Linguistics (AREA)
- Acoustics & Sound (AREA)
- Artificial Intelligence (AREA)
- Probability & Statistics with Applications (AREA)
- Life Sciences & Earth Sciences (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Bioinformatics & Computational Biology (AREA)
- Evolutionary Biology (AREA)
- Signal Processing (AREA)
- Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
Abstract
一种语音识别方法及装置,预先根据通过模型训练得到的N个高斯进行软性聚类计算,得到M个软聚类高斯(203);在进行语音识别时,将语音转换得到特征向量,并根据该特征向量计算得分最高的前L个软聚类高斯,其中L小于M(303);将L个软聚类高斯内的各成员高斯,作为语音识别过程中声学模型里需要参与计算的高斯,进行声学模型似然度的计算(304)。在语音识别的时候采用动态高斯选择的方式,减少识别过程中声学模型里需要评估的高斯个数,提高了声学模型似然度评估的速度和准确性。
Description
交叉引用
本申请要求于2015年12月30日提交中国专利局、申请号为201511027242.0的中国专利申请的优先权,其全部内容通过引用结合在本申请中。
本专利申请涉及语音技术,特别涉及一种语音识别方法及装置。
发明人在实现本发明的过程中发现,随着语音识别技术的发展,近年来语音识别技术的准确率随着深度学习的推广取得了巨大的进步,特别是在基于云的服务中。现有的语音识别服务多数在云端实现,语音需要上传至服务器,服务器对上传的语音进行声学评估,从而给出识别结果。为了提高识别率,服务器大多采用深度学习的方法对语音进行评估。但深度学习需要耗费巨大的计算资源,在本地或者嵌入式设备中不适用。而且在很多不能联网的使用场景下,只能依赖本地语音识别技术。由于本地计算和存储资源有限,隐马尔科夫模型(HMM)和高斯混合模型(GMM)仍然是不可或缺的技术选择。这种技术框架具有以下优点:
1、系统尺寸可控:高斯混合模型中的高斯数量易于在训练时控制。
2、系统速度可控:使用动态高斯选择技术可以大幅度降低运算时间
所谓高斯选择即在模型训练阶段,把语音识别系统中所有的高斯作为成员高斯进行聚类,形成聚类高斯;在识别的时候首先利用声学特征评估每个
聚类高斯,那些似然度高的聚类高斯所对应的成员高斯被选中进行进一步的评估。而其他成员高斯被丢弃。传统的高斯选择技术有以下缺点:
1、在聚类的时候采用硬聚类,即一个成员高斯只属于一个聚类高斯。聚类精确度较低。
2、聚类时直接把成员高斯的均值和方差作为聚类的输入,在训练聚类高斯的时候直接把均值和方差做简单的算术平均,聚类精度极低。
3、聚类的时候,没有有效的迭代方法,致使聚类收敛于局部最优。
4、识别时的高斯选择不能做到动态更新,导致过多的成员高斯保留在计算中,识别速度慢。
发明内容
本发明部分实施例的目的在于提供一种语音识别方法及装置,使得语音识别过程中可以减少声学模型里需要评估的高斯个数,比传统的高斯选择更加准确和高效,从而提高了声学模型似然度评估的速度和准确性。
为解决上述技术问题,本发明的实施方式提供了一种语音识别方法,包含以下步骤:
预先根据通过模型训练得到的N个高斯,进行软性聚类计算,得到M个软聚类高斯;
在进行语音识别时,将语音转换得到特征向量,并根据所述特征向量计算得分最高的前L个软聚类高斯,其中L小于所述M;
将L个软聚类高斯内的各成员高斯,作为语音识别过程中声学模型里需要参与计算的高斯,进行声学模型似然度的计算。
本发明的实施方式还提供了一种语音识别装置,包含:
软性聚类获取模块,用于根据通过模型训练得到的N个高斯,进行软性聚类计算,得到M个软聚类高斯;
向量转换模块,用于在进行语音识别时,将语音转换得到特征向量;
选择模块,用于根据所述特征向量计算得分最高的前L个软聚类高斯,并将所述前L个软聚类高斯的各成员高斯,作为选择的高斯;所述L小于所述M;
计算模块,用于将所述选择模块选择的高斯,作为语音识别过程中声学模型里需要参与计算的高斯,进行声学模型似然度的计算。
本发明实施方式相对于现有技术而言,通过对模型训练得到的N个高斯进行软性聚类,得到M个软聚类高斯,再根据特征向量对M个软聚类高斯进行计算得到分数最高的前L个软聚类高斯,然后将L个软聚类高斯内的各成员高斯进行声学模型似然度的计算,得到识别输出结果。通过软性聚类可以使一个成员高斯属于多个聚类高斯,提高了聚类的精确度,而且在识别的时候采用动态高斯选择的方式,减少了识别过程中声学模型里需要评估的高斯个数,使得在本地识别过程中,可将GMM中每个成员高斯的得分计算量从整个计算时间的70%左右降低到20%,从而提高了声学模型似然度评估速度和准确率,尤其适用于本地语音识别,唤醒,和语音端点检测(检测语音的起始点)。
在一个实施例中,根据通过模型训练得到的N个高斯,进行软性聚类计算的步骤中,包含以下子步骤:
将N个高斯按预设权重分配给聚类高斯;
根据各高斯对所属的各聚类高斯的更新权重,重新估计聚类高斯,得到M个软聚类高斯。
通过软性聚类计算,使得每个成员高斯可以属于多个聚类高斯,提高了
模型的描述能力,从而提高识别率。
在一个实施例中,在采用K均值算法重新估计聚类高斯时,计算各聚类高斯的最小聚类代价;
对最小聚类代价求导,获取每个成员高斯对每个聚类高斯的更新权重;
根据获取到的每个成员高斯对每个聚类高斯的更新权重,计算各聚类高斯的均值和方差,得到重新估计的聚类高斯;
将该重新估计的聚类高斯,作为M个软聚类高斯。
通过计算各聚类高斯的最小聚类代价使得聚类高斯的划分达到平方误差最小。采用精确的K均值(K-Means)方法对高斯进行软性聚类(即一个成员高斯可属于多个聚类高斯),聚类个数逐步增加,并且每次增加的方式反映了模型分布的规律,一方面保证了同一聚类内各成员高斯的相似度,另一方面可使得类与类之间的区别明显,从而提高了聚类的精度。
在一个实施例中,所述L的取值为满足下列条件的最小值:
其中,p(Gi|Y)≥p(Gi+1|Y)
所述Y表示所述特征向量,α是一个对高斯“后验”概率的压缩指数,Gi表示第i个聚类高斯,p(Gi|Y)表示第i个聚类高斯的“后验”概率。
将根据上述公式计算得出的最小值作为L的取值,可以使识别过程中声学模型里需要评估的高斯个数较少,提高了声学模型似然度评估速度。
在一个实施例中,根据特征向量计算出得分最高的前L个软聚类高斯的步骤中,包含以下子步骤:
根据以下公式,获取各软聚类高斯的得分:
所述Y表示所述特征向量,μm表示第m个软聚类高斯的均值,Σm表示第m个软聚类高斯的方差。
图1是根据本发明部分实施方式的语音识别系统示意图;
图2是根据第一实施方式中软性聚类的计算流程图;
图3是根据第一实施方式的语音识别方法流程图;
图4是根据第一实施方式的动态高斯选择示意图;
图5是根据第四实施方式的语音识别装置结构示意图。
为使本发明部分实施例的目的、技术方案和优点更加清楚,下面将结合附图对本发明的各实施方式进行详细的阐述。然而,本领域的普通技术人员可以理解,在本发明各实施方式中,为了使读者更好地理解本申请而提出了许多技术细节。但是,即使没有这些技术细节和基于以下各实施方式的种种变化和修改,也可以实现本申请各权利要求所要求保护的技术方案。
语音识别目的是在观察到一段语音信号的情况下,给出可能性最高的文本。如图1所示,一个基于HMM+GMM的识别系统按帧读取一段语音,系统把每帧语音信号变成特征向量。系统结合每帧特征向量评估声学模型中每个高斯的似然度,同时假设多种词的组合,对这些词的组合利用语言模型进行似然度评估,声学似然度和语言似然度总和最高的词组合作为识别结果输出。
本发明的第一实施方式涉及一种语音识别方法。在本实施方式中,需要预先根据通过模型训练得到的N个高斯,进行软性聚类计算,得到M个软聚类高斯。在进行语音识别时,通过用动态高斯选择的方式,控制需要计算的成员高斯个数。在本实施方式中,软性聚类的计算流程如图2所示。
在步骤201中,通过模型训练得到N个高斯,如得到1000个高斯。
在步骤202中,将N个高斯按预设权重分配给聚类高斯。
在步骤203中,根据各高斯对所属的各聚类高斯的更新权重,重新估计聚类高斯,得到M个软聚类高斯。
本领域技术人员可以理解,高斯混合模型在语音识别中用来描述隐马尔科夫模型(HMM)每个状态的概率分布,每个状态使用若干个高斯来表述自己的概率分布。一个高斯分布有自己的均值μ和方差Σ。为了在识别系统中有效使用高斯选择,状态间必须共享高斯。这种共享高斯的声学模型叫做半连续马尔科夫模型。在使用相同数量高斯的情况下,半连续高斯会提高模型的描述能力,从而提高识别率。通过模型训练得到N(在本地识别系统中,N一般取值1000)个高斯,在聚类前必须明确高斯之间的距离判据。在本实施方式中,采用加权对称KL散度(WSKLD)作为距离判据。一个高斯m和高斯n之间的距离的SKLD为:
如果高斯模型分成多个子空间,每个子空间都有自己的权重β,则WSKLD为:
其中Nstrm为高斯模型的子空间个数。
软性聚类的计算,在具体实现时,可以采用以下任意算法:K均值算法、C均值算法、自组织图算法。下面以K均值算法为例,进行具体说明:
该算法可以用下述伪码来描述:
1、把聚类高斯的个数m设为1,使用所有高斯作为成员高斯估计出一个聚类高斯。
2、while m<M(M是聚类高斯的个数的目标值)
2c.For循环τfrom 1 to T
2c-1.For聚类高斯i,i from 1 to m
2c-1-1.For成员高斯n,n from 1 to N,其中N是成员高斯的个数
上述伪码中聚类的目标是让聚类代价Q最小,其中,Q的计算公式如下:
其中,g(i,n)表示第n个高斯对第i个聚类高斯的更新权重;γ为预设的聚类软硬度参数;WSKLD表示作为高斯之间距离判据的加权对称KL散度。
通过迭代可以得到以下参数:聚类高斯的均值,方差和每个成员高斯对更新每个聚类高斯的权重:
在获取上述参数的迭代过程中,第一步是获取最佳的更新权重:
第二步是基于最佳权重获取聚类高斯最佳的均值和方差。更新聚类高斯均值的方法如下:
为了计算聚类高斯的方差,可以构造一个辅助矩阵Z。
基于Z的构造,它有DP个整的的特征值和与之对称的负的DP个特征值,其中DP是均值和方差的维度。此时构造一个2DP-by-DP的矩阵V,它列是DP个Z的正特征值对应的特征向量。把V分成上半部分U和下半部分
W:
则聚类高斯的协方差矩阵估计如下:
均值和协方差矩阵交替迭代几轮后,协方差矩阵被限制为对角阵。这个强加的条件在少数情况下会导致聚类不收敛,但是不影响聚类准确性,从而得到重新估计的聚类高斯,作为M个软聚类高斯。
也就是说,在本实施方式中,识别系统通过计算各聚类高斯的最小聚类代价,再对每个最小聚类代价求导,从而获取每个成员高斯对每个聚类高斯的更新权重,然后根据该更新权重,计算各聚类高斯的均值和方差,得到重新估计的聚类高斯,作为M个软聚类高斯。
在得到M个软聚类高斯后对语音进行识别,具体流程如图3所示:
在步骤301中,识别系统按帧读取一段语音,比如说,每帧长度为10毫秒。
在步骤302中,识别系统把每帧语音信号变成特征向量,得到的特征向量用于对软聚类高斯进行评估。
在步骤303中,根据特征向量计算出得分最高的前L个软聚类高斯(其中L小于M)。
具体地说,如图4所示:在语音识别的过程中,当一阵语音被转换成特征向量Y后,所有的聚类高斯首先利用该向量进行评估,得分最高的前L个聚类高斯被选中放在聚类高斯选择表。根据以下公式,可以获取各软聚类高斯的得分:
其中Y表示所述特征向量,μm表示第m个软聚类高斯的均值,Σm表示第m个软聚类高斯的方差。在得到M个聚类高斯的得分后,取得分最高的前L个聚类高斯,作为选中的聚类高斯。
在本实施方式中,L的取值为满足下列条件的最小值:
其中,Y表示特征向量,α是一个对高斯“后验”概率的压缩指数,Gi表示第i个聚类高斯,p(Gi|Y)表示第i个聚类高斯的“后验”概率。
在步骤304中,将L个软聚类高斯内的各成员高斯,作为语音识别过程中声学模型里需要参与计算的高斯,进行声学模型似然度的计算。
也就是说,一个成员高斯是否被选择并计算取决于成员高斯和聚类高斯映射表和聚类高斯选择列表。如图4中,聚类高斯选择表中“1”表示相应的聚类高斯在识别过程中的当前时刻被选中。在“聚类-成员高斯映射表”中查询被选中的聚类高斯对应的成员高斯,进行计算。未被选中的成员高斯的似然度用一个小值代替。
在步骤305中,判断是否还存在未读取的语音帧。如果判断结果为是,说明还有需要识别的语音帧,则回到步骤301读取下一个语音帧继续进行识别。否则说明语音识别已经全部完成,则结束流程。
在步骤306中,输出识别结果。具体地说,本步骤中的语音识别的结果为声学似然度和语言似然度总和,本步骤与现有技术相同,在此不再赘述。
为了验证本实施方式中的语音识别方法的实用性,在一个测试集上,测试了几种发放的CPU时间和识别率,结果如表1所示:
其中硬高斯聚类是指每个成员函数只属于一个聚类高斯,而且聚类仅仅是把均值当做向量进行。软精确聚类是本发明部分实施例中描述的方法。不使用高斯聚类的系统作为基线。可以看到硬高斯聚类在精确方面比本发明部分实施例的方法要差。二者速度相当。基线系统在速度和精度都比本发明部分实施例的系统要差。
表1
不难发现,本发明的实施方式,在系统训练阶段采用精确的K均值(K-Means)方法对高斯进行软性聚类(即一个成员高斯可属于多个聚类高斯),聚类个数逐步增加,并且每次增加的方式反映了模型分布的规律。在识别的时候采用动态高斯选择的方式,控制需要计算的成员高斯个数。从而提高了声学模型似然度评估速度和准确率。比传统的高斯选择更加准确和高效。
本发明的第二实施方式涉及一种语音识别方法。第二实施方式与第一实施方式大致相同,主要区别之处在于:在第一实施方式中,在系统训练阶段采用精确的K均值(K-Means)算法对高斯进行软性聚类。而在本发明第二实施方式中,在系统训练阶段采用C均值算法对高斯进行软性聚类。由于采用C均值算法进行软性聚类计算的具体实现方式,与K均值算法基本相同,在本实施方式中不再赘述。
本发明的第三实施方式涉及一种语音识别方法。第三实施方式与第一实施方式大致相同,主要区别之处在于:在第一实施方式中,在系统训练阶段采用精确的K均值(K-Means)算法对高斯进行软性聚类。而在本发明第三实施方式中,在系统训练阶段采用自组织图算法对高斯进行软性聚类。由于采用自组织图算法进行软性聚类计算的具体实现方式,仅在步骤203中略有不同,而自组织图算法为现有的聚类算法的公知技术,本实施方式中也不再赘述。
上面各种方法的步骤划分,只是为了描述清楚,实现时可以合并为一个步骤或者对某些步骤进行拆分,分解为多个步骤,只要包含相同的逻辑关系,都在本专利的保护范围内;对算法中或者流程中添加无关紧要的修改或者引入无关紧要的设计,但不改变其算法和流程的核心设计都在该专利的保护范围内。
本发明第四实施方式涉及一种语音识别装置,如图5所示,包含:
软性聚类获取模块,用于根据通过模型训练得到的N个高斯,进行软性聚类计算,得到M个软聚类高斯;
向量转换模块,用于在进行语音识别时,将语音转换得到特征向量;
选择模块,用于根据特征向量计算出得分最高的前L个软聚类高斯,并将前L个软聚类高斯的各成员高斯,作为选择的高斯,其中L小于M;
计算模块,用于将选择模块选择的高斯,作为语音识别过程中声学模型里需要参与计算的高斯,进行声学模型似然度的计算。
其中软性聚类获取模块包含:
权重分配模块,用于将N个高斯按预设权重分配给聚类高斯;
重估计模块,用于根据各高斯对所属的各聚类高斯的更新权重,重新估计聚类高斯,得到M个软聚类高斯。
不难发现,本实施方式为与第一实施方式相对应的系统实施例,本实施方式可与第一实施方式互相配合实施。第一实施方式中提到的相关技术细节在本实施方式中依然有效,为了减少重复,这里不再赘述。相应地,本实施方式中提到的相关技术细节也可应用在第一实施方式中。
值得一提的是,本实施方式中所涉及到的各模块均为逻辑模块,在实际应用中,一个逻辑单元可以是一个物理单元,也可以是一个物理单元的一部分,还可以以多个物理单元的组合实现。此外,为了突出本发明的创新部分,本实施方式中并没有将与解决本发明所提出的技术问题关系不太密切的单元引入,但这并不表明本实施方式中不存在其它的单元。
本领域的普通技术人员可以理解,上述各实施方式是实现本发明的具体实施例,而在实际应用中,可以在形式上和细节上对其作各种改变,而不偏离本发明的精神和范围。
Claims (10)
- 一种语音识别方法,包含以下步骤:预先根据通过模型训练得到的N个高斯,进行软性聚类计算,得到M个软聚类高斯;在进行语音识别时,将语音转换得到特征向量,并根据所述特征向量计算出得分最高的前L个软聚类高斯,所述L小于所述M;将所述L个软聚类高斯内的各成员高斯,作为语音识别过程中声学模型里需要参与计算的高斯,进行声学模型似然度的计算。
- 根据权利要求1所述的语音识别方法,其中,所述根据通过模型训练得到的N个高斯,进行软性聚类计算的步骤中,包含以下子步骤:将所述N个高斯按预设权重分配给聚类高斯;根据各高斯对所属的各聚类高斯的更新权重,重新估计聚类高斯,得到所述M个软聚类高斯。
- 根据权利要求1或2所述的语音识别方法,其中,所述根据通过模型训练得到的N个高斯,进行软性聚类计算的步骤中,采用以下任意算法,进行所述软性聚类的计算:K均值算法、C均值算法、自组织图算法。
- 根据权利要求3所述的语音识别方法,其中,在采用K均值算法重新估计聚类高斯时,计算各聚类高斯的最小聚类代价;对所述最小聚类代价求导,获取每个成员高斯对每个聚类高斯的更新权重;根据获取到的每个成员高斯对每个聚类高斯的更新权重,计算各聚类高 斯的均值和方差,得到所述重新估计的聚类高斯;将所述重新估计的聚类高斯,作为所述M个软聚类高斯。
- 根据权利要求1至7任一项所述的语音识别方法,其中,在所述将语音转换得到特征向量的步骤中,将每个语音帧转换为一个所述特征向量。
- 一种语音识别装置,包含:软性聚类获取模块,用于根据通过模型训练得到的N个高斯,进行软性聚类计算,得到M个软聚类高斯;向量转换模块,用于在进行语音识别时,将语音转换得到特征向量;选择模块,用于根据所述特征向量计算出得分最高的前L个软聚类高斯,并将所述前L个软聚类高斯的各成员高斯,作为选择的高斯;所述L小于所述M;计算模块,用于将所述选择模块选择的高斯,作为语音识别过程中声学模型里需要参与计算的高斯,进行声学模型似然度的计算。
- 根据权利要求9所述的语音识别装置,其中,所述软性聚类获取模块包含:权重分配模块,用于将所述N个高斯按预设权重分配给聚类高斯;重估计模块,用于根据各高斯对所属的各聚类高斯的更新权重,重新估计聚类高斯,得到所述M个软聚类高斯。
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US15/240,119 US20170193987A1 (en) | 2015-12-30 | 2016-08-18 | Speech recognition method and device |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN201511027242.0 | 2015-12-30 | ||
| CN201511027242.0A CN105895089A (zh) | 2015-12-30 | 2015-12-30 | 一种语音识别方法及装置 |
Related Child Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| US15/240,119 Continuation US20170193987A1 (en) | 2015-12-30 | 2016-08-18 | Speech recognition method and device |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2017113739A1 true WO2017113739A1 (zh) | 2017-07-06 |
Family
ID=57002535
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2016/089579 Ceased WO2017113739A1 (zh) | 2015-12-30 | 2016-07-10 | 一种语音识别方法及装置 |
Country Status (3)
| Country | Link |
|---|---|
| US (1) | US20170193987A1 (zh) |
| CN (1) | CN105895089A (zh) |
| WO (1) | WO2017113739A1 (zh) |
Families Citing this family (7)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN110473536B (zh) * | 2019-08-20 | 2021-10-15 | 北京声智科技有限公司 | 一种唤醒方法、装置和智能设备 |
| CN113470416B (zh) * | 2020-03-31 | 2023-02-17 | 上汽通用汽车有限公司 | 利用嵌入式系统实现车位检测的系统、方法和存储介质 |
| CN111640419B (zh) * | 2020-05-26 | 2023-04-07 | 合肥讯飞数码科技有限公司 | 语种识别方法、系统、电子设备及存储介质 |
| CN112037773B (zh) * | 2020-11-05 | 2021-01-29 | 北京淇瑀信息科技有限公司 | 一种n最优口语语义识别方法、装置及电子设备 |
| CN112329746B (zh) * | 2021-01-04 | 2021-04-16 | 中国科学院自动化研究所 | 多模态谎言检测方法、装置、设备 |
| CN116189671B (zh) * | 2023-04-27 | 2023-07-07 | 凌语国际文化艺术传播股份有限公司 | 一种用于语言教学的数据挖掘方法及系统 |
| CN119152847A (zh) * | 2023-06-14 | 2024-12-17 | 华为技术有限公司 | 一种自定义唤醒词的语音唤醒方法及装置 |
Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN1655232A (zh) * | 2004-02-13 | 2005-08-17 | 松下电器产业株式会社 | 上下文相关的汉语语音识别建模方法 |
| CN102486922A (zh) * | 2010-12-03 | 2012-06-06 | 株式会社理光 | 说话人识别方法、装置和系统 |
| US20120330664A1 (en) * | 2011-06-24 | 2012-12-27 | Xin Lei | Method and apparatus for computing gaussian likelihoods |
| US20140214420A1 (en) * | 2013-01-25 | 2014-07-31 | Microsoft Corporation | Feature space transformation for personalization using generalized i-vector clustering |
Family Cites Families (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US9436759B2 (en) * | 2007-12-27 | 2016-09-06 | Nant Holdings Ip, Llc | Robust information extraction from utterances |
| US8583416B2 (en) * | 2007-12-27 | 2013-11-12 | Fluential, Llc | Robust information extraction from utterances |
| EP2189976B1 (en) * | 2008-11-21 | 2012-10-24 | Nuance Communications, Inc. | Method for adapting a codebook for speech recognition |
| US9269368B2 (en) * | 2013-03-15 | 2016-02-23 | Broadcom Corporation | Speaker-identification-assisted uplink speech processing systems and methods |
| US9293140B2 (en) * | 2013-03-15 | 2016-03-22 | Broadcom Corporation | Speaker-identification-assisted speech processing systems and methods |
-
2015
- 2015-12-30 CN CN201511027242.0A patent/CN105895089A/zh active Pending
-
2016
- 2016-07-10 WO PCT/CN2016/089579 patent/WO2017113739A1/zh not_active Ceased
- 2016-08-18 US US15/240,119 patent/US20170193987A1/en not_active Abandoned
Patent Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN1655232A (zh) * | 2004-02-13 | 2005-08-17 | 松下电器产业株式会社 | 上下文相关的汉语语音识别建模方法 |
| CN102486922A (zh) * | 2010-12-03 | 2012-06-06 | 株式会社理光 | 说话人识别方法、装置和系统 |
| US20120330664A1 (en) * | 2011-06-24 | 2012-12-27 | Xin Lei | Method and apparatus for computing gaussian likelihoods |
| US20140214420A1 (en) * | 2013-01-25 | 2014-07-31 | Microsoft Corporation | Feature space transformation for personalization using generalized i-vector clustering |
Also Published As
| Publication number | Publication date |
|---|---|
| US20170193987A1 (en) | 2017-07-06 |
| CN105895089A (zh) | 2016-08-24 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2017113739A1 (zh) | 一种语音识别方法及装置 | |
| US10468032B2 (en) | Method and system of speaker recognition using context aware confidence modeling | |
| US10332507B2 (en) | Method and device for waking up via speech based on artificial intelligence | |
| CN108346428B (zh) | 语音活动检测及其模型建立方法、装置、设备及存储介质 | |
| KR102072235B1 (ko) | 자동 발화속도 분류 방법 및 이를 이용한 음성인식 시스템 | |
| CN107633842B (zh) | 语音识别方法、装置、计算机设备及存储介质 | |
| US20150199960A1 (en) | I-Vector Based Clustering Training Data in Speech Recognition | |
| US11631414B2 (en) | Speech recognition method and speech recognition apparatus | |
| WO2021136029A1 (zh) | 重打分模型训练方法及装置、语音识别方法及装置 | |
| WO2018227781A1 (zh) | 语音识别方法、装置、计算机设备及存储介质 | |
| CN110349597B (zh) | 一种语音检测方法及装置 | |
| Jansen et al. | Towards Unsupervised Training of Speaker Independent Acoustic Models. | |
| CN113555005B (zh) | 模型训练、置信度确定方法及装置、电子设备、存储介质 | |
| JP5752060B2 (ja) | 情報処理装置、大語彙連続音声認識方法及びプログラム | |
| WO2018161763A1 (zh) | 语音数据集训练方法、计算机设备和计算机可读存储介质 | |
| CN104575519A (zh) | 特征提取方法、装置及重音检测的方法、装置 | |
| CN106297769B (zh) | 一种应用于语种识别的鉴别性特征提取方法 | |
| CN105845141A (zh) | 基于信道鲁棒的说话人确认模型及说话人确认方法和装置 | |
| CN114360552A (zh) | 用于说话人识别的网络模型训练方法、装置及存储介质 | |
| CN114023336B (zh) | 模型训练方法、装置、设备以及存储介质 | |
| Lu et al. | Probabilistic linear discriminant analysis for acoustic modeling | |
| US20120109650A1 (en) | Apparatus and method for creating acoustic model | |
| KR20160098910A (ko) | 음성 인식 데이터 베이스 확장 방법 및 장치 | |
| US20180061395A1 (en) | Apparatus and method for training a neural network auxiliary model, speech recognition apparatus and method | |
| CN112309374B (zh) | 服务报告生成方法、装置和计算机设备 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 16880537 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 16880537 Country of ref document: EP Kind code of ref document: A1 |
















