WO2011037587A1 - Downsampling schemes in a hierarchical neural network structure for phoneme recognition - Google Patents
Downsampling schemes in a hierarchical neural network structure for phoneme recognition Download PDFInfo
- Publication number
- WO2011037587A1 WO2011037587A1 PCT/US2009/058563 US2009058563W WO2011037587A1 WO 2011037587 A1 WO2011037587 A1 WO 2011037587A1 US 2009058563 W US2009058563 W US 2009058563W WO 2011037587 A1 WO2011037587 A1 WO 2011037587A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- downsampling
- phoneme
- vectors
- posterior
- posterior vectors
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/02—Feature extraction for speech recognition; Selection of recognition unit
Definitions
- the present invention relates to automatic speech recognition, and more specifically phoneme recognition in a hierarchical neural network.
- a phoneme is the minimal unit of speech sound in a language that can serve to distinguish meaning.
- Phoneme recognition can be applied to improve automatic speech recognition.
- Other applications of phoneme recognition can also be found in speaker recognition, language identification, and keyword spotting.
- phoneme recognition has received much attention in the field of automatic speech recognition.
- HMM hidden Markov model
- MLP Multilayered Perceptron
- Pinto Another phoneme recognition structure was proposed in J. Pinto et al, Exploiting Contextual Information For Improved Phoneme Recognition, in Proc. ICASSP, 2008, pp. 4449-4452, (hereinafter "Pinto", incorporated herein by reference). Pinto suggested estimating phoneme posteriors using a two-layer hierarchical structure. A first MLP estimates intermediate phoneme posteriors based on a temporal window of cepstral features, and then a second MLP estimates final phoneme posteriors based on a temporal window of intermediate posterior features. The final phoneme posteriors are then input to a phonetic decoder for obtaining a final recognized phoneme sequence.
- Embodiments of the present invention are directed to phoneme recognition.
- a sequence of intermediate output posterior vectors is generated from an input sequence of cepstral features using a first layer perceptron.
- the intermediate output posterior vectors are then downsampled to form a reduced input set of intermediate posterior vectors for a second layer perceptron.
- a sequence of final posterior vectors is generated from the reduced input set of intermediate posterior vectors using the second layer perceptron. Then the final posterior vectors are decoded to determine an output recognized phoneme sequence representative of the input sequence of cepstral features.
- intra-phonetic information and/or inter-phonetic information may be used for decoding the final posterior vectors.
- the downsampling may be based on a window downsampling arrangement or a temporal downsampling arrangement, for example, using uniform downsampling or non-uniform downsampling.
- the downsampling may be based on using an intermediate phoneme decoder to determine possible phoneme boundaries.
- the perceptrons may be arrangements within a Hybrid Hidden Markov Model-Multilayer Perceptron (HMM-MLP) phoneme classifier.
- HMM-MLP Hybrid Hidden Markov Model-Multilayer Perceptron
- Embodiments also include an application or device adapted to perform the method according to any of the above.
- the application may be a real time application and/or an embedded application.
- Figure 1 shows a two-layer hierarchical structure for estimating posterior vectors for phoneme recognition.
- Figure 2 compares the standard hierarchical approach for handling intermediate posterior vectors with a window downsampling approach.
- Figure 3 A-B shows sequences of posterior vectors in a temporal context C which stretches over a period of frames.
- Figure 4 compares the standard hierarchical approach for handling intermediate posterior vectors with a uniform downsampling approach.
- Figure 5 illustrates a non-uniform downsampling approach.
- the emission probabilities in an HMM state can be estimated by scaled likelihoods according to the Bayes' rule:
- the scaled likelihoods are applied to a Viterbi decoder for obtaining a recognized phoneme sequence.
- the MLP estimates a posterior probability for each phoneme or phoneme-state.
- the out ut of the MLP is a posterior feature vector:
- O the total number of phonemes or phoneme-states.
- Embodiments of the present invention are directed to an improved neural network for phoneme recognition based on a two-layer hierarchical structure based on a multi-layer perceptron (MLP).
- Intermediate posterior vectors generated by a first layer MLP e.g., cepstral features level
- a second layer MLP e.g., intermediate posteriors level
- Figure 1 shows a two-layer hierarchical structure for estimating posterior vectors according to embodiments of the present invention.
- an MLP estimates intermediate phoneme posterior vectors p t from a context of Ncepstral features x . Redundant information contained in the intermediate posterior vectors is removed by one or more downsampling arrangements, and the second layer MLP, which provides context modeling at the posterior level, based on a window of C intermediate posterior vectors:
- window downsampling redundant information is removed at the input window of the second layer MLP. That allows the number of parameters to be greatly reduced and reduces computational time while maintaining system accuracy requirements.
- temporal downsampling reduces the number of frames at the input of the second layer MLP by downsampling the intermediate posteriors based on a intermediate phonetic decoder. Such downsampling arrangements make it feasible to implement a hierarchical scheme in a real time application or in a embedded system.
- Inter-phonetic information can be modeled by the second layer of the hierarchical MLP, and a window downsampling technique can then be used to make feasible a real time implementation.
- a two-layer MLP hierarchical structure can be used that estimates the posterior vectors.
- the first layer MLP provides context modeling at the feature level by estimating intermediate phoneme posterior vectors p ⁇ from a context of Ncepstral features xf .
- the second layer MLP provides context modeling at the posterior level by estimating phoneme posterior vectors based on a window of M posterior features pf covering a context of C frames.
- M C since all posterior vectors M contained in a context window of C frames were taken.
- Figure 2A shows the standard hierarchical approach where a temporal window of M consecutive intermediate posterior vectors are input to a second layer MLP without downsampling.
- M different windows of cepstral features, shifted from each other by one frame, are used to generate M consecutive intermediate posterior vectors:
- the window of intermediate posterior vectors pf is input to a second layer MLP to estimate a final posterior vector q ⁇ .
- the sequence of intermediate posterior vectors p is generated by a single MLP which is shifted over time.
- the M consecutive posterior vectors can be viewed as if they were derived by M MLPs situated at M consecutive time instances.
- each MLP has been trained based on its corresponding label l t+j for ⁇ , ⁇ ⁇ .
- Context modeling at the posterior level as described in the foregoing is based on inter-phonetic information, which is useful for improving recognition accuracy since it captures how a phoneme is bounded by other phonemes.
- the number of posterior vectors M at the input of the second layer MLP is given by:
- sequences of posterior vectors p are given showing only those relevant components with highest posterior values within a temporal context C. It can be observed that a particular component p dominates during certain sub-intervals of the whole context C since a phoneme stretches over several frames. Therefore, there is similar information contained in each sub-interval. Having all this repeated information may be irrelevant for the task of phoneme classification based on inter-phonetic information.
- the TIMIT corpus was used without the SA dialect sentences, dividing the database into three parts.
- the training data set contained 3346 utterances from 419 speakers
- the cross-validation data set contained 350 utterances from 44 speakers
- the standard test data set contained 1344 utterances from 168 speakers.
- a 39 phoneme set (from Lee and Hon) was used with the difference that closures were merged to the regarding burst.
- Feature vectors were of 39-dimensions with 13 PLPs with delta and double delta coefficients.
- the feature vectors were under global mean and variance normalization, with each feature vector extracted from a 25 msec speech window with a 10 msec shift.
- the MLPs were trained with the Quicknet software tool to implement three layer perceptrons with 1000 hidden units.
- the number of output units corresponded to 39 and 117 for 1 -state and 3 -state modeling respectively with the softmax nonlinearity function at the output.
- a standard back-propagation algorithm with cross-entropy error criteria was used for training the neural network, where the learning rate reduction and stop training criteria were controlled by the frame error rate in cross-validation to avoid overtraining.
- a phoneme insertion penalty was set to give maximum phoneme accuracy in the cross-validation.
- a Viterbi decoder was implemented with a minimum duration of three states per phoneme. It also was assumed that all phonemes and states were equally distributed. No language model was used and silences were discarded for evaluation.
- phoneme accuracy (PA) was used in all experiments as a measure of performance.
- the MLPs estimate phoneme posteriors or state posteriors for 1 -state or 3-state modeling respectively.
- Figure 2B shows how a window downsampling approach can be used to select fewer posterior vectors which are separated by some given number of frames T s .
- the number of MLPs M has been highly reduced according to the equation in the preceding paragraph, where each MLP is separated a number of samples T s . This reduces the number of inputs to the second layer MLP, thereby decreasing thus the number of parameters while still preserving acceptable performance levels.
- the posterior vectors generated by the hierarchical downsampling approach are still covering the same temporal context C.
- the intra-phonetic can be used as a first hierarchical step, then the posterior vectors generated used as the input to a second hierarchical step, given by the inter-phonetic approach.
- the aim such an arrangement is first to better classify a phoneme based on the temporal transition information within the phoneme. Then, temporal transition information among different phonemes is utilized to continue improving phoneme classification.
- FIG. 4A shows the intermediate posterior vectors for each frame sampling period T t according to the standard hierarchical approach as in Pinto.
- the MLP receives a set of intermediate posterior vectors which are sampled based on a impulse train which denotes how often the sampling is performed. This can be viewed as a filtering process where the intermediate posterior vectors that are not sampled are ignored by the MLP.
- Each intermediate posterior vector has a
- each intermediate posterior vector corresponds to a window of C consecutive posterior vectors.
- the uniform downsampling arrangement described above has the disadvantage that important information can be lost when the sampling period is highly increased.
- the intermediate posterior vectors corresponding to short phonemes can be totally ignored after performing the downsampling. For this reason, it may be useful to sample the set of intermediate posterior vectors every time potentially important information appears.
- the sampling points can be estimated by an intermediate Viterbi decoder which takes at its input the set of intermediate posteriors. Then, the sampling points correspond to those points in time where different phonemes have been recognized, generating a non-uniform downsampling arrangement.
- Figure 5 illustrates the idea of a non-uniform downsampling arrangement.
- a set of intermediate posterior vectors generated by the first layer MLP is input to an intermediate Viterbi decoder, which gives at its output an intermediate recognized phoneme sequence together with the phoneme boundaries. Then the recognized phoneme sequence not used further, but just the time boundaries.
- Each segment corresponding to a recognized phoneme is uniformly divided into three sub-segments. Then, the sampling points are indicated by the central frame of all sub-segments.
- the second MLP gave an average speed of 359.63 MCPS (Million Connections Per Second) for forward propagation phase. This MLP was able to process 418.6 frames/sec. On the other hand, the MLP for 3-state modeling gave an average speed of 340.87 MCPS which processed 132.4 frames/sec. Table 5 shows the average number of frames per utterance at the input of the second layer MLP:
- Table 5 Performance of downs ampling arrangements measured in average frames per utterance and computational time.
- Table 6 shows phoneme accuracies of the different downsampling methods:
- Table 6 Phoneme accuracies of downsampling arrangements.
- the second layer MLP estimates phoneme posteriors (1-state) or state posteriors (3-state). Minimum duration constraints of 1-state and 3-state per phoneme are evaluated in the final Viterbi decoder.
- Table 5 shows that a high reduction in the number of frames is achieved— by 59.5% and 66.8% for 1-state and 3- state modeling, respectively— which reduced computational time in the same proportion. Moreover, the computational time required by the intermediate decoder for processing one utterance was neglected compared to the time required by the second layer MLP in processing the same utterance. Table 6 shows that a similar accuracy is obtained, compared to the standard approach. This shows a significant advantage of non-uniform downsampling where a high decrease of computational time is obtained while keeping good performance.
- Embodiments of the invention may be implemented in any conventional computer programming language. For example, some or all of an embodiment may be implemented in a procedural programming language (e.g., "C") or an object oriented programming language (e.g. , "C++", Python). Alternative embodiments of the invention may be implemented as pre-programmed hardware elements, other related components, or as a combination of hardware and software components.
- a procedural programming language e.g., "C”
- object oriented programming language e.g. , "C++”, Python
- Alternative embodiments of the invention may be implemented as pre-programmed hardware elements, other related components, or as a combination of hardware and software components.
- Embodiments can be implemented in whole or in part as a computer program product for use with a computer system.
- Such implementation may include a series of computer instructions fixed either on a tangible medium, such as a computer readable medium (e.g., a diskette, CD-ROM, ROM, or fixed disk) or transmittable to a computer system, via a modem or other interface device, such as a communications adapter connected to a network over a medium.
- the medium may be either a tangible medium (e.g., optical or analog communications lines) or a medium implemented with wireless techniques (e.g., microwave, infrared or other transmission techniques).
- the series of computer instructions embodies all or part of the functionality previously described herein with respect to the system.
- Such computer instructions can be written in a number of programming languages for use with many computer architectures or operating systems. Furthermore, such instructions may be stored in any memory device, such as semiconductor, magnetic, optical or other memory devices, and may be transmitted using any communications technology, such as optical, infrared, microwave, or other transmission technologies. It is expected that such a computer program product may be distributed as a removable medium with accompanying printed or electronic documentation (e.g., shrink wrapped software), preloaded with a computer system (e.g., on system ROM or fixed disk), or distributed from a server or electronic bulletin board over the network (e.g., the Internet or World Wide Web). Of course, some embodiments of the invention may be implemented as a combination of both software (e.g. , a computer program product) and hardware. Still other embodiments of the invention are implemented as entirely hardware, or entirely software (e.g., a computer program product).
Landscapes
- Engineering & Computer Science (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Computational Linguistics (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Human Computer Interaction (AREA)
- Physics & Mathematics (AREA)
- Acoustics & Sound (AREA)
- Multimedia (AREA)
- Machine Translation (AREA)
Abstract
An approach for phoneme recognition is described. A sequence of intermediate output posterior vectors is generated from an input sequence of cepstral features using a first layer perceptron. The intermediate output posterior vectors are then downsampled to form a reduced input set of intermediate posterior vectors for a second layer perceptron. A sequence of final posterior vectors is generated from the reduced input set of intermediate posterior vectors using the second layer perceptron. Then the final posterior vectors are decoded to determine an output recognized phoneme sequence representative of the input sequence of cepstral features.
Description
TITLE
Downsampling Schemes In A Hierarchical Neural Network Structure For Phoneme
Recognition
FIELD OF THE INVENTION
[0001] The present invention relates to automatic speech recognition, and more specifically phoneme recognition in a hierarchical neural network.
BACKGROUND ART
[0002] A phoneme is the minimal unit of speech sound in a language that can serve to distinguish meaning. Phoneme recognition can be applied to improve automatic speech recognition. Other applications of phoneme recognition can also be found in speaker recognition, language identification, and keyword spotting. Thus, phoneme recognition has received much attention in the field of automatic speech recognition.
[0003] One common and successful approach for phoneme recognition uses a hierarchical neural network structure based on a hybrid hidden Markov model (HMM)- Multilayered Perceptron (MLP) arrangement. The MLP outputs are used as HMM state emission probabilities in a Viterbi decoder. This approach has the considerable advantage that the MLP can be trained to discriminatively classify phonemes. The MLP also can easily incorporate a long temporal context without making explicit assumptions. This property is particularly important for phoneme recognition because phoneme
characteristics can be spread over a large temporal context.
[0004] Many different approaches have been proposed to continue to exploit the contextual information of a phoneme. One approach is based on a combination of different specialized classifiers that provides considerable improvements over simple generic classifiers. For instance, in the approach known as TRAPS, long temporal information is divided into frequency bands, and then, several classifiers are independently trained using specific frequency information over a long temporal range. See H. Hermansky and S. Sharma, Temporal Patterns (TRAPS) in ASR of Noisy Speech, in Proc. ICASSP, 1999, vol. 1, pp. 289-292, incorporated herein by reference. Another different technique splits a long
temporal context in time. See D. Vasquez et al, On Expanding Context By Temporal Decomposition For Improving Phoneme Recognition, in SPECOM, 2009, incorporated herein by reference. A combination of these two approaches which splits the context in time and frequency is evaluated in P. Schwarz et al., Hierarchical Structures
Of Neural Networks For Phoneme Recognition, in Proc. ICASSP, 2006, pp. 325-328, incorporated herein by reference.
[0005] Another phoneme recognition structure was proposed in J. Pinto et al, Exploiting Contextual Information For Improved Phoneme Recognition, in Proc. ICASSP, 2008, pp. 4449-4452, (hereinafter "Pinto", incorporated herein by reference). Pinto suggested estimating phoneme posteriors using a two-layer hierarchical structure. A first MLP estimates intermediate phoneme posteriors based on a temporal window of cepstral features, and then a second MLP estimates final phoneme posteriors based on a temporal window of intermediate posterior features. The final phoneme posteriors are then input to a phonetic decoder for obtaining a final recognized phoneme sequence.
[0006] The hierarchical approach described by Pinto significantly increases system accuracy, compared to a non-hierarchical scheme (a single layer). But computational time is greatly increased because the second MLP has to process the same number of speech frames as were processed by the first MLP. In addition, the second MLP has an input window with a large number of consecutive frames, so there are a high number of parameters that must be processed. These factors make it less practical to implement such a hierarchical approach in a real time application or in an embedded system.
SUMMARY OF THE INVENTION
[0007] Embodiments of the present invention are directed to phoneme recognition. A sequence of intermediate output posterior vectors is generated from an input sequence of cepstral features using a first layer perceptron. The intermediate output posterior vectors are then downsampled to form a reduced input set of intermediate posterior vectors for a second layer perceptron. A sequence of final posterior vectors is generated from the reduced input set of intermediate posterior vectors using the second layer perceptron. Then the final posterior vectors are decoded to determine an output recognized phoneme
sequence representative of the input sequence of cepstral features.
[0008] In further embodiments, intra-phonetic information and/or inter-phonetic information may be used for decoding the final posterior vectors. The downsampling may be based on a window downsampling arrangement or a temporal downsampling arrangement, for example, using uniform downsampling or non-uniform downsampling. The downsampling may be based on using an intermediate phoneme decoder to determine possible phoneme boundaries. And the perceptrons may be arrangements within a Hybrid Hidden Markov Model-Multilayer Perceptron (HMM-MLP) phoneme classifier.
[0009] Embodiments also include an application or device adapted to perform the method according to any of the above. For example, the application may be a real time application and/or an embedded application.
BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 shows a two-layer hierarchical structure for estimating posterior vectors for phoneme recognition.
[0011] Figure 2 compares the standard hierarchical approach for handling intermediate posterior vectors with a window downsampling approach.
[0012] Figure 3 A-B shows sequences of posterior vectors in a temporal context C which stretches over a period of frames.
[0013] Figure 4 compares the standard hierarchical approach for handling intermediate posterior vectors with a uniform downsampling approach.
[0014] Figure 5 illustrates a non-uniform downsampling approach.
DETAILED DESCRIPTION OF SPECIFIC EMBODIMENTS
[0015] In phoneme recognition, the goal is to find the most probable corresponding phoneme or phoneme-state sequence S = {s\ , ¾ , ¾} given a sequence of observations
or cepstral vectors Xo = {xi , xt , χκ}, where K > L. In the following discussion, the phoneme recognition approach is discussed based on a Hybrid HMM/MLP where the MLP outputs are used as HMM state emission probabilities. An MLP has the additional advantage of incorporating a long context in the form of a temporal window of Ncepstral features. Therefore, the set of acoustic vectors at the input of the MLP is
X = {x1 N,.., xf , x^} where:
Xi — {X Χί>"·> X
2 2
[0016] The emission probabilities in an HMM state can be estimated by scaled likelihoods according to the Bayes' rule:
where pt' = p(st = i | xf ) refers to the output of the MLP. The prior state probability P(st = i) can be estimated after performing a forced alignment of the training data to its true labels. Since p(x ) is independent of the HMM state, it can be ignored for deriving scaled likelihoods. The scaled likelihoods are applied to a Viterbi decoder for obtaining a recognized phoneme sequence.
[0017] At each time instant t, the MLP estimates a posterior probability for each phoneme or phoneme-state. Thus the out ut of the MLP is a posterior feature vector:
with O being the total number of phonemes or phoneme-states. In this discussion, we refer to 1 -state or 3 -state modeling when the number of output units equals the total number of phonemes or phoneme-states respectively.
Hierarchical Downsampling
[0018] Embodiments of the present invention are directed to an improved neural network for phoneme recognition based on a two-layer hierarchical structure based on a multi-layer perceptron (MLP). Intermediate posterior vectors generated by a first layer MLP (e.g., cepstral features level) are input to a second layer MLP (e.g., intermediate
posteriors level). Figure 1 shows a two-layer hierarchical structure for estimating posterior vectors according to embodiments of the present invention. In the first layer which is defined as context modeling at the feature level, an MLP estimates intermediate phoneme posterior vectors pt from a context of Ncepstral features x . Redundant information contained in the intermediate posterior vectors is removed by one or more downsampling arrangements, and the second layer MLP, which provides context modeling at the posterior level, based on a window of C intermediate posterior vectors:
to estimate final phoneme posterior vectors q = p(st = i | p ) given phoneme posterior trajectories which constitute to some extended language model properties.
[0019] For example, in one approach referred to as window downsampling, redundant information is removed at the input window of the second layer MLP. That allows the number of parameters to be greatly reduced and reduces computational time while maintaining system accuracy requirements. Another approach referred to as temporal downsampling reduces the number of frames at the input of the second layer MLP by downsampling the intermediate posteriors based on a intermediate phonetic decoder. Such downsampling arrangements make it feasible to implement a hierarchical scheme in a real time application or in a embedded system.
Inter-Phonetic Information
[0020] Inter-phonetic information can be modeled by the second layer of the hierarchical MLP, and a window downsampling technique can then be used to make feasible a real time implementation. As explained above, a two-layer MLP hierarchical structure can be used that estimates the posterior vectors. The first layer MLP provides context modeling at the feature level by estimating intermediate phoneme posterior vectors p^ from a context of Ncepstral features xf . The second layer MLP provides context modeling at the posterior level by estimating phoneme posterior vectors based on a window of M posterior features pf covering a context of C frames. In the hierarchical approach, M = C since all posterior vectors M contained in a context window of C frames were taken.
[0021] Figure 2A shows the standard hierarchical approach where a temporal window of M consecutive intermediate posterior vectors are input to a second layer MLP without downsampling. At the feature level, M different windows of cepstral features, shifted from each other by one frame, are used to generate M consecutive intermediate posterior vectors:
covering a context of C frames. Then, at the posterior level, the window of intermediate posterior vectors pf is input to a second layer MLP to estimate a final posterior vector q^.
[0022] In practice, the sequence of intermediate posterior vectors p is generated by a single MLP which is shifted over time. The M consecutive posterior vectors can be viewed as if they were derived by M MLPs situated at M consecutive time instances. In this case, each MLP has been trained based on its corresponding label lt+j for < ,· < . The
2 2
M generated posterior vectors may have inter-phonetic information among them since the labels, which cover a context of C frames, may belong to different phonemes. Then, the final posterior vectors q = p(st = i | p ) are estimated based on some inter-phonetic posterior trajectories which correspond to some extended language model properties.
[0023] Context modeling at the posterior level as described in the foregoing is based on inter-phonetic information, which is useful for improving recognition accuracy since it captures how a phoneme is bounded by other phonemes. The number of posterior vectors M at the input of the second layer MLP is given by:
T.
where C is the total context of frames and Ts is the frame sampling period. As an example, Figure 3A shows a context of C = 31 frames with a period of Ts = 5, giving a number of Μ = Ί intermediate posterior vectors at the input of the second layer MLP. On the other hand, when Ts = 1 , then M = C and the hierarchical approach given in Pinto is obtained.
[0024] In Figure 3 A-B, sequences of posterior vectors p are given showing only those relevant components with highest posterior values within a temporal context C. It can be observed that a particular component p dominates during certain sub-intervals of the whole context C since a phoneme stretches over several frames. Therefore, there is similar information contained in each sub-interval. Having all this repeated information may be irrelevant for the task of phoneme classification based on inter-phonetic information.
Testing Framework
[0025] In one specific set of testing experiments, the TIMIT corpus was used without the SA dialect sentences, dividing the database into three parts. The training data set contained 3346 utterances from 419 speakers, the cross-validation data set contained 350 utterances from 44 speakers, and the standard test data set contained 1344 utterances from 168 speakers. A 39 phoneme set (from Lee and Hon) was used with the difference that closures were merged to the regarding burst. Feature vectors were of 39-dimensions with 13 PLPs with delta and double delta coefficients. The feature vectors were under global mean and variance normalization, with each feature vector extracted from a 25 msec speech window with a 10 msec shift.
[0026] The MLPs were trained with the Quicknet software tool to implement three layer perceptrons with 1000 hidden units. The number of output units corresponded to 39 and 117 for 1 -state and 3 -state modeling respectively with the softmax nonlinearity function at the output. A standard back-propagation algorithm with cross-entropy error criteria was used for training the neural network, where the learning rate reduction and stop training criteria were controlled by the frame error rate in cross-validation to avoid overtraining. In addition, a phoneme insertion penalty was set to give maximum phoneme accuracy in the cross-validation. A Viterbi decoder was implemented with a minimum duration of three states per phoneme. It also was assumed that all phonemes and states were equally distributed. No language model was used and silences were discarded for evaluation. Finally, phoneme accuracy (PA) was used in all experiments as a measure of performance.
[0027] Table 1 shows the results of a set of experiments for a standard HMM/MLP
system representing context modeling at the feature level for 1 -state and 3-state modeling with a window of N = 9 consecutive cepstral vectors.
Table 1-Phoneme recognition accuracies for different context modeling levels. The MLPs estimate phoneme posteriors or state posteriors for 1 -state or 3-state modeling respectively.
In addition, results of the hierarchical approach are also given for M = 21 in the second row of Table 1 where it can be observed that a remarkable increase in performance has been achieved by the hierarchical approach, as in Pinto.
Window Downsampling
[0028] Figure 2B shows how a window downsampling approach can be used to select fewer posterior vectors which are separated by some given number of frames Ts. In this case, the number of MLPs Mhas been highly reduced according to the equation in the preceding paragraph, where each MLP is separated a number of samples Ts. This reduces the number of inputs to the second layer MLP, thereby decreasing thus the number of parameters while still preserving acceptable performance levels. However, the posterior vectors generated by the hierarchical downsampling approach are still covering the same temporal context C.
[0029] Several experiments were performed to test this approach in which the frame sample period Ts was varied while keeping almost constant the temporal context C. Table 2 gives the results for 1 -state and 3-state modeling showing phoneme recognition accuracies:
Table 2- Phoneme recognition accuracies for the hierarchical downsampling approach. Several sampling rates have been tested giving M intermediate posterior vectors at the input of the second layer MLP covering a temporal context of C frames.
As expected, the performance significantly decreases when Ts increases up to 10 frames. On the other hand, it is interesting to observe that the performance remains almost constant for Ts = 3 and Ts = 5. In particular, for Ts = 5, good performance was obtained, but the number of input posterior vectors was greatly reduced from 21 to 5. These results confirm that there is redundant information at the input of the second layer MLP which can be ignored.
Intra-Phonetic Information
[0030] The foregoing shows how system performance can be increased when inter- phonetic information is used. In addition, the transition information within a particular phoneme can be modeled better to reflect the fact that a phoneme behaves differently at its beginning, middle, and end. The MLPs that have been trained as described above can also be used to obtain this intra-phonetic information. The inter-phonetic information in different MLPs trained according to Fig. 2B was obtained with M = 5, with the same label It , which belongs to the label of the middle of the context C. Thus, there is an implicit assumption that a phoneme occupies a total context of C frames.
[0031] Table 3 gives the results when M = 5 different MLPs were used for modeling inter and intra-phonetic information, as shown in Figure 2B:
j System PA j
1 -state 3-state j
j no hierarchy 68,13 71.21 j
hierarchy inter 71.61 74.20 j
j hierarchy intra 70.91 73.65 j
Table 3- Phoneme recognition accuracies in a hierarchical framework under inter or intra phonetic constraints. The number of intermediate posterior vectors at the input of the second layer MLP is M=5, covering a temporal context of C=21 frames.
Indeed, the inter-phonetic arrangement is the same as shown in Table 2 for Ts = 5.
According to Table 3, it can be observed that the inter-phonetic approach out performs the intra-phonetic approach, meaning that inter- information is more useful to better classify a
phoneme. However, both approaches considerably out perform the non-hierarchical approach. And while these approaches are based on different criteria, they can serve to complement each other and further improve performance.
Combining Inter- and Intra-Phonetic Information
[0032] The foregoing discussion describes how considering inter- and intra-phonetic information can be used to considerably improve phoneme recognition performance over a non-hierarchical system. Since these two approaches were generated based on different criteria, we now consider the complementariness of both approaches in order to achieve further improvements. To that end, the intra-phonetic can be used as a first hierarchical step, then the posterior vectors generated used as the input to a second hierarchical step, given by the inter-phonetic approach. The aim such an arrangement is first to better classify a phoneme based on the temporal transition information within the phoneme. Then, temporal transition information among different phonemes is utilized to continue improving phoneme classification.
[0033] In one set of experiments, both inter-phonetic derivations at the second hierarchical step (i.e., the conventional hierarchical approach given in Pinto) and the window downsampling approach above. Table 4 shows results of the combination together with the downsampling parameters utilized in the inter-phonetic step.
Table 4-Phoneme recognition accuracies for a combination of inter- and intra-phonetic information.
The introduction of intra-phonetic information as an intermediate step in a hierarchical approach achieves further improvements. These results verify the assumption that both criteria carry complementary information, useful for improving system performance. On the other hand, results concerning the downsampling technique, using different
sampling periods Ts , show the advantage of removing redundant information while keeping good performance.
Temporal Downsampling.
[0034] Another approach for removing redundant information from the intermediate posterior vectors is based on temporal downsampling. Figure 4A shows the intermediate posterior vectors for each frame sampling period Tt according to the standard hierarchical approach as in Pinto. The MLP receives a set of intermediate posterior vectors which are sampled based on a impulse train which denotes how often the sampling is performed. This can be viewed as a filtering process where the intermediate posterior vectors that are not sampled are ignored by the MLP. Each intermediate posterior vector has a
corresponding final posterior vector since the sampling period is Tt = 1 frame. In Fig. 4A, each intermediate posterior vector corresponds to a window of C consecutive posterior vectors.
[0035] Figure 4B shows a uniform downsampling scheme when Tt = 3. Under this approach, it can be observed that a significant number of frames are reduced when Tt is considerably increased. It can be shown that having just a few number of frames of final posterior vectors decreases system accuracy under constrain of minimum phoneme duration of 3 -states in the Viterbi decoder. This suggests further testing this approach with minimum phoneme duration of 1 -state. It is also important to mention that during training, the true labels together with the training set of intermediate posterior vectors are also downsampled, significantly reducing the training time.
Non-Uniform Downsampling
[0036] The uniform downsampling arrangement described above has the disadvantage that important information can be lost when the sampling period is highly increased. In particular, the intermediate posterior vectors corresponding to short phonemes can be totally ignored after performing the downsampling. For this reason, it may be useful to sample the set of intermediate posterior vectors every time potentially important information appears. The sampling points can be estimated by an intermediate Viterbi decoder which takes at its input the set of intermediate posteriors. Then, the sampling
points correspond to those points in time where different phonemes have been recognized, generating a non-uniform downsampling arrangement. By this means, the loss of possible important information is highly alleviated, while significantly reducing the number of frames at the input of the second layer MLP.
[0037] Figure 5 illustrates the idea of a non-uniform downsampling arrangement. A set of intermediate posterior vectors generated by the first layer MLP is input to an intermediate Viterbi decoder, which gives at its output an intermediate recognized phoneme sequence together with the phoneme boundaries. Then the recognized phoneme sequence not used further, but just the time boundaries. Each segment corresponding to a recognized phoneme is uniformly divided into three sub-segments. Then, the sampling points are indicated by the central frame of all sub-segments.
[0038] Three sampling points per recognized phoneme were used in order to keep consistency with a phoneme model of minimum 3-state duration. In addition, no word insertion penalty was used in the intermediate Viterbi decoder. Moreover, as it is performed for the uniform downsampling scheme, the training set of intermediate posteriors together with the true labels were also downsampled based on the non-uniform sampling points. For training the second layer MLP, a window of C consecutive posterior vectors was used.
[0039] In one set of experiments testing these ideas, we tested various specific MLP arrangements at the posterior level, corresponding to different hierarchical approaches: no downsampling (standard), uniform downsampling and non-uniform downsampling. All the tested MLPs, had the same number of parameters, and therefore, the computational time reduction was only due to a decrease in number of frames at the input of the second layer MLP. As in other experiments, the input of the second layer MLP for
all approaches consisted of a window of C = 21 intermediate posterior vectors.
[0040] For 1-state modeling, the second MLP gave an average speed of 359.63 MCPS (Million Connections Per Second) for forward propagation phase. This MLP was able to process 418.6 frames/sec. On the other hand, the MLP for 3-state modeling
gave an average speed of 340.87 MCPS which processed 132.4 frames/sec. Table 5 shows the average number of frames per utterance at the input of the second layer MLP:
Table 5— Performance of downs ampling arrangements measured in average frames per utterance and computational time.
Table 6 shows phoneme accuracies of the different downsampling methods:
Table 6— Phoneme accuracies of downsampling arrangements. The second layer MLP estimates phoneme posteriors (1-state) or state posteriors (3-state). Minimum duration constraints of 1-state and 3-state per phoneme are evaluated in the final Viterbi decoder.
[0041] For uniform downsampling, three different sampling periods Tt were tested. And for = 1, the standard hierarchical approach was obtained as it is given in Pinto. It can be seen in Table 5 that the computational time was considerably reduced when Tt was increased, since there were fewer frames to process by the second layer MLP. But, the system accuracy dropped significantly. For this approach, it can be observed that it is necessary to use a minimum duration of 1-state in the final decoder in order to keep a reasonably good accuracy.
[0042] One reason of the low accuracy obtained when Tt was increased was caused by the poor classification of short phonemes. When the sampling period is extremely increased, only a few samples corresponding to short phonemes may remain, or they even may totally disappear. This effect was verified by measuring the phoneme recognition of short phonemes when Tt = 5. The shortest phonemes of the TIMIT corpus are given by /dx/ and /dh/ with an average number of frames of 2.9 and 3.5, respectively. And in fact,
those phonemes suffered the highest deterioration after downsampling with a decrease in phoneme recognition from 69% to 50% and from 68% to 59%, for /dx/ and /dh/ respectively.
[0043] For the non-uniform downsampling arrangement, Table 5 shows that a high reduction in the number of frames is achieved— by 59.5% and 66.8% for 1-state and 3- state modeling, respectively— which reduced computational time in the same proportion. Moreover, the computational time required by the intermediate decoder for processing one utterance was neglected compared to the time required by the second layer MLP in processing the same utterance. Table 6 shows that a similar accuracy is obtained, compared to the standard approach. This shows a significant advantage of non-uniform downsampling where a high decrease of computational time is obtained while keeping good performance.
[0044] These results confirm that there is a large amount of redundant information contained in the intermediate posterior vectors, and therefore, it is not necessary for the second layer MLP to again process each individual frame of the entire utterance. Various strategies have been described for downsampling the intermediate posterior vectors. One especially promising approach was obtained by non-uniform downsampling where the sampling frequency is estimated by an intermediate Viterbi decoder. This reduced the computational time by 67% while maintaining system accuracy comparable to the standard hierarchical approach. These results suggest the viability of implementing a hierarchical structure in a real-time application.
[0045] Embodiments of the invention may be implemented in any conventional computer programming language. For example, some or all of an embodiment may be implemented in a procedural programming language (e.g., "C") or an object oriented programming language (e.g. , "C++", Python). Alternative embodiments of the invention may be implemented as pre-programmed hardware elements, other related components, or as a combination of hardware and software components.
[0046] Embodiments can be implemented in whole or in part as a computer program
product for use with a computer system. Such implementation may include a series of computer instructions fixed either on a tangible medium, such as a computer readable medium (e.g., a diskette, CD-ROM, ROM, or fixed disk) or transmittable to a computer system, via a modem or other interface device, such as a communications adapter connected to a network over a medium. The medium may be either a tangible medium (e.g., optical or analog communications lines) or a medium implemented with wireless techniques (e.g., microwave, infrared or other transmission techniques). The series of computer instructions embodies all or part of the functionality previously described herein with respect to the system. Those skilled in the art should appreciate that such computer instructions can be written in a number of programming languages for use with many computer architectures or operating systems. Furthermore, such instructions may be stored in any memory device, such as semiconductor, magnetic, optical or other memory devices, and may be transmitted using any communications technology, such as optical, infrared, microwave, or other transmission technologies. It is expected that such a computer program product may be distributed as a removable medium with accompanying printed or electronic documentation (e.g., shrink wrapped software), preloaded with a computer system (e.g., on system ROM or fixed disk), or distributed from a server or electronic bulletin board over the network (e.g., the Internet or World Wide Web). Of course, some embodiments of the invention may be implemented as a combination of both software (e.g. , a computer program product) and hardware. Still other embodiments of the invention are implemented as entirely hardware, or entirely software (e.g., a computer program product).
[0047] Although various exemplary embodiments of the invention have been disclosed, it should be apparent to those skilled in the art that various changes and modifications can be made which will achieve some of the advantages of the invention without departing from the true scope of the invention.
Claims
1. A method for phoneme recognition comprising:
generating a sequence of intermediate output posterior vectors from an input sequence of cepstral features using a first layer perceptron;
downsampling the intermediate output posterior vectors to form a reduced input set of intermediate posterior vectors for a second layer perceptron.
2. A method according to claim I, further comprising:
generating a sequence of final posterior vectors from the reduced input set of
intermediate posterior vectors using the second layer perceptron.
3. A method according to claim 2, further comprising:
decoding the final posterior vectors to determine an output recognized phoneme sequence representative of the input sequence of cepstral features.
4. A method according to claim 3, wherein decoding the final posterior vectors uses intra-phonetic information.
5. A method according to claim 3, wherein decoding the final posterior vectors uses inter-phonetic information.
6. A method according to claim 1, wherein the downsampling is based on a window downsampling arrangement.
7. A method according to claim 1, wherein the downsampling is based on a temporal downsampling arrangement.
8. A method according to claim 7, wherein the downsampling is based on a uniform downsampling arrangement.
9. A method according to claim 7, wherein the downsampling is based on a non-uniform downsampling arrangement.
10. A method according to claim 9, wherein the downsampling is based on using an intermediate phoneme decoder to determine possible phoneme boundaries.
11. A method according to claim 1, wherein the perceptrons are arrangements within a Hybrid Hidden Markov Model-Multilayer Perceptron (HMM-MLP) phoneme classifier.
12. An application adapted to perform the method according to any of claims 1-1 1.
13. An application according to claim 12, wherein the application is a real time application.
14. An application according to claim 12, wherein the application is an embedded application.
Priority Applications (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US13/497,119 US9595257B2 (en) | 2009-09-28 | 2009-09-28 | Downsampling schemes in a hierarchical neural network structure for phoneme recognition |
| PCT/US2009/058563 WO2011037587A1 (en) | 2009-09-28 | 2009-09-28 | Downsampling schemes in a hierarchical neural network structure for phoneme recognition |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/US2009/058563 WO2011037587A1 (en) | 2009-09-28 | 2009-09-28 | Downsampling schemes in a hierarchical neural network structure for phoneme recognition |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2011037587A1 true WO2011037587A1 (en) | 2011-03-31 |
Family
ID=42224159
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/US2009/058563 Ceased WO2011037587A1 (en) | 2009-09-28 | 2009-09-28 | Downsampling schemes in a hierarchical neural network structure for phoneme recognition |
Country Status (2)
| Country | Link |
|---|---|
| US (1) | US9595257B2 (en) |
| WO (1) | WO2011037587A1 (en) |
Families Citing this family (69)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US9318108B2 (en) | 2010-01-18 | 2016-04-19 | Apple Inc. | Intelligent automated assistant |
| US8977255B2 (en) | 2007-04-03 | 2015-03-10 | Apple Inc. | Method and system for operating a multi-function portable electronic device using voice-activation |
| US8676904B2 (en) | 2008-10-02 | 2014-03-18 | Apple Inc. | Electronic devices with voice command and contextual data processing capabilities |
| US10255566B2 (en) | 2011-06-03 | 2019-04-09 | Apple Inc. | Generating and processing task items that represent tasks to perform |
| US10276170B2 (en) | 2010-01-18 | 2019-04-30 | Apple Inc. | Intelligent automated assistant |
| US8775341B1 (en) | 2010-10-26 | 2014-07-08 | Michael Lamport Commons | Intelligent control with hierarchical stacked neural networks |
| US9015093B1 (en) | 2010-10-26 | 2015-04-21 | Michael Lamport Commons | Intelligent control with hierarchical stacked neural networks |
| US10057736B2 (en) | 2011-06-03 | 2018-08-21 | Apple Inc. | Active transport based notifications |
| US10417037B2 (en) | 2012-05-15 | 2019-09-17 | Apple Inc. | Systems and methods for integrating third party services with a digital assistant |
| US9202464B1 (en) * | 2012-10-18 | 2015-12-01 | Google Inc. | Curriculum learning for speech recognition |
| DE112014000709B4 (en) | 2013-02-07 | 2021-12-30 | Apple Inc. | METHOD AND DEVICE FOR OPERATING A VOICE TRIGGER FOR A DIGITAL ASSISTANT |
| US10652394B2 (en) | 2013-03-14 | 2020-05-12 | Apple Inc. | System and method for processing voicemail |
| US10748529B1 (en) | 2013-03-15 | 2020-08-18 | Apple Inc. | Voice activated device for use with a voice-based digital assistant |
| KR101959188B1 (en) | 2013-06-09 | 2019-07-02 | 애플 인크. | Device, method, and graphical user interface for enabling conversation persistence across two or more instances of a digital assistant |
| US10176167B2 (en) | 2013-06-09 | 2019-01-08 | Apple Inc. | System and method for inferring user intent from speech inputs |
| KR101749009B1 (en) | 2013-08-06 | 2017-06-19 | 애플 인크. | Auto-activating smart responses based on activities from remote devices |
| WO2015184186A1 (en) | 2014-05-30 | 2015-12-03 | Apple Inc. | Multi-command single utterance input method |
| US9715875B2 (en) | 2014-05-30 | 2017-07-25 | Apple Inc. | Reducing the need for manual start/end-pointing and trigger phrases |
| US10170123B2 (en) | 2014-05-30 | 2019-01-01 | Apple Inc. | Intelligent assistant for home automation |
| US9338493B2 (en) | 2014-06-30 | 2016-05-10 | Apple Inc. | Intelligent automated assistant for TV user interactions |
| US9824684B2 (en) * | 2014-11-13 | 2017-11-21 | Microsoft Technology Licensing, Llc | Prediction-based sequence recognition |
| US9886953B2 (en) | 2015-03-08 | 2018-02-06 | Apple Inc. | Virtual assistant activation |
| US10460227B2 (en) | 2015-05-15 | 2019-10-29 | Apple Inc. | Virtual assistant in a communication session |
| US10200824B2 (en) | 2015-05-27 | 2019-02-05 | Apple Inc. | Systems and methods for proactively identifying and surfacing relevant content on a touch-sensitive device |
| US20160378747A1 (en) | 2015-06-29 | 2016-12-29 | Apple Inc. | Virtual assistant for media playback |
| KR102371188B1 (en) * | 2015-06-30 | 2022-03-04 | 삼성전자주식회사 | Apparatus and method for speech recognition, and electronic device |
| US10747498B2 (en) | 2015-09-08 | 2020-08-18 | Apple Inc. | Zero latency digital assistant |
| US10740384B2 (en) | 2015-09-08 | 2020-08-11 | Apple Inc. | Intelligent automated assistant for media search and playback |
| US10671428B2 (en) | 2015-09-08 | 2020-06-02 | Apple Inc. | Distributed personal assistant |
| US10331312B2 (en) | 2015-09-08 | 2019-06-25 | Apple Inc. | Intelligent automated assistant in a media environment |
| US11587559B2 (en) | 2015-09-30 | 2023-02-21 | Apple Inc. | Intelligent device identification |
| US10141010B1 (en) * | 2015-10-01 | 2018-11-27 | Google Llc | Automatic censoring of objectionable song lyrics in audio |
| US10691473B2 (en) | 2015-11-06 | 2020-06-23 | Apple Inc. | Intelligent automated assistant in a messaging environment |
| US10956666B2 (en) | 2015-11-09 | 2021-03-23 | Apple Inc. | Unconventional virtual assistant interactions |
| US10043517B2 (en) * | 2015-12-09 | 2018-08-07 | International Business Machines Corporation | Audio-based event interaction analytics |
| US10223066B2 (en) | 2015-12-23 | 2019-03-05 | Apple Inc. | Proactive assistance based on dialog communication between devices |
| US12223282B2 (en) | 2016-06-09 | 2025-02-11 | Apple Inc. | Intelligent automated assistant in a home environment |
| US10586535B2 (en) | 2016-06-10 | 2020-03-10 | Apple Inc. | Intelligent digital assistant in a multi-tasking environment |
| DK201670540A1 (en) | 2016-06-11 | 2018-01-08 | Apple Inc | Application integration with a digital assistant |
| DK179415B1 (en) | 2016-06-11 | 2018-06-14 | Apple Inc | Intelligent device arbitration and control |
| US10726832B2 (en) | 2017-05-11 | 2020-07-28 | Apple Inc. | Maintaining privacy of personal information |
| DK180048B1 (en) | 2017-05-11 | 2020-02-04 | Apple Inc. | MAINTAINING THE DATA PROTECTION OF PERSONAL INFORMATION |
| DK179496B1 (en) | 2017-05-12 | 2019-01-15 | Apple Inc. | USER-SPECIFIC Acoustic Models |
| DK201770428A1 (en) | 2017-05-12 | 2019-02-18 | Apple Inc. | Low-latency intelligent automated assistant |
| DK179745B1 (en) | 2017-05-12 | 2019-05-01 | Apple Inc. | SYNCHRONIZATION AND TASK DELEGATION OF A DIGITAL ASSISTANT |
| DK201770411A1 (en) | 2017-05-15 | 2018-12-20 | Apple Inc. | MULTI-MODAL INTERFACES |
| DK179560B1 (en) | 2017-05-16 | 2019-02-18 | Apple Inc. | Far-field extension for digital assistant services |
| US20180336892A1 (en) | 2017-05-16 | 2018-11-22 | Apple Inc. | Detecting a trigger of a digital assistant |
| US10303715B2 (en) | 2017-05-16 | 2019-05-28 | Apple Inc. | Intelligent automated assistant for media exploration |
| US10818288B2 (en) | 2018-03-26 | 2020-10-27 | Apple Inc. | Natural assistant interaction |
| US10928918B2 (en) | 2018-05-07 | 2021-02-23 | Apple Inc. | Raise to speak |
| US11145294B2 (en) | 2018-05-07 | 2021-10-12 | Apple Inc. | Intelligent automated assistant for delivering content from user experiences |
| DK179822B1 (en) | 2018-06-01 | 2019-07-12 | Apple Inc. | Voice interaction at a primary device to access call functionality of a companion device |
| DK201870355A1 (en) | 2018-06-01 | 2019-12-16 | Apple Inc. | Virtual assistant operation in multi-device environments |
| DK180639B1 (en) | 2018-06-01 | 2021-11-04 | Apple Inc | DISABILITY OF ATTENTION-ATTENTIVE VIRTUAL ASSISTANT |
| US10892996B2 (en) | 2018-06-01 | 2021-01-12 | Apple Inc. | Variable latency device coordination |
| US11462215B2 (en) | 2018-09-28 | 2022-10-04 | Apple Inc. | Multi-modal inputs for voice commands |
| US11348573B2 (en) | 2019-03-18 | 2022-05-31 | Apple Inc. | Multimodality in digital assistant systems |
| DK201970509A1 (en) | 2019-05-06 | 2021-01-15 | Apple Inc | Spoken notifications |
| US11307752B2 (en) | 2019-05-06 | 2022-04-19 | Apple Inc. | User configurable task triggers |
| US11140099B2 (en) | 2019-05-21 | 2021-10-05 | Apple Inc. | Providing message response suggestions |
| DK180129B1 (en) | 2019-05-31 | 2020-06-02 | Apple Inc. | USER ACTIVITY SHORTCUT SUGGESTIONS |
| DK201970510A1 (en) | 2019-05-31 | 2021-02-11 | Apple Inc | Voice identification in digital assistant systems |
| US11227599B2 (en) | 2019-06-01 | 2022-01-18 | Apple Inc. | Methods and user interfaces for voice-based control of electronic devices |
| US11061543B1 (en) | 2020-05-11 | 2021-07-13 | Apple Inc. | Providing relevant data items based on context |
| US11183193B1 (en) | 2020-05-11 | 2021-11-23 | Apple Inc. | Digital assistant hardware abstraction |
| US11490204B2 (en) | 2020-07-20 | 2022-11-01 | Apple Inc. | Multi-device audio adjustment coordination |
| US11438683B2 (en) | 2020-07-21 | 2022-09-06 | Apple Inc. | User identification using headphones |
| JP7487794B2 (en) * | 2020-11-25 | 2024-05-21 | 日本電信電話株式会社 | Labeling processing method, labeling processing device, and labeling processing program |
Citations (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20040172247A1 (en) * | 2003-02-24 | 2004-09-02 | Samsung Electronics Co., Ltd. | Continuous speech recognition method and system using inter-word phonetic information |
Family Cites Families (16)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US5377302A (en) * | 1992-09-01 | 1994-12-27 | Monowave Corporation L.P. | System for recognizing speech |
| JPH075892A (en) * | 1993-04-29 | 1995-01-10 | Matsushita Electric Ind Co Ltd | Speech recognition method |
| JPH10511472A (en) * | 1994-12-08 | 1998-11-04 | ザ リージェンツ オブ ザ ユニバーシティ オブ カリフォルニア | Method and apparatus for improving speech recognition between speech impaired persons |
| WO2000013170A1 (en) * | 1998-09-01 | 2000-03-09 | Swisscom Ag | Neural network and its use for speech recognition |
| GB2355834A (en) * | 1999-10-29 | 2001-05-02 | Nokia Mobile Phones Ltd | Speech recognition |
| US6513004B1 (en) * | 1999-11-24 | 2003-01-28 | Matsushita Electric Industrial Co., Ltd. | Optimized local feature extraction for automatic speech recognition |
| JP3520022B2 (en) * | 2000-01-14 | 2004-04-19 | 株式会社国際電気通信基礎技術研究所 | Foreign language learning device, foreign language learning method and medium |
| US20030004720A1 (en) * | 2001-01-30 | 2003-01-02 | Harinath Garudadri | System and method for computing and transmitting parameters in a distributed voice recognition system |
| US7454342B2 (en) * | 2003-03-19 | 2008-11-18 | Intel Corporation | Coupled hidden Markov model (CHMM) for continuous audiovisual speech recognition |
| US8229744B2 (en) * | 2003-08-26 | 2012-07-24 | Nuance Communications, Inc. | Class detection scheme and time mediated averaging of class dependent models |
| EP1661124A4 (en) * | 2003-09-05 | 2008-08-13 | Stephen D Grody | Methods and apparatus for providing services using speech recognition |
| US8260611B2 (en) * | 2005-04-01 | 2012-09-04 | Qualcomm Incorporated | Systems, methods, and apparatus for highband excitation generation |
| DE102007033472A1 (en) * | 2007-07-18 | 2009-01-29 | Siemens Ag | Method for speech recognition |
| KR100911429B1 (en) * | 2007-08-22 | 2009-08-11 | 한국전자통신연구원 | Method and apparatus for generating noise adaptive acoustic model for moving environment |
| US8560307B2 (en) * | 2008-01-28 | 2013-10-15 | Qualcomm Incorporated | Systems, methods, and apparatus for context suppression using receivers |
| US8131541B2 (en) * | 2008-04-25 | 2012-03-06 | Cambridge Silicon Radio Limited | Two microphone noise reduction system |
-
2009
- 2009-09-28 WO PCT/US2009/058563 patent/WO2011037587A1/en not_active Ceased
- 2009-09-28 US US13/497,119 patent/US9595257B2/en not_active Expired - Fee Related
Patent Citations (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20040172247A1 (en) * | 2003-02-24 | 2004-09-02 | Samsung Electronics Co., Ltd. | Continuous speech recognition method and system using inter-word phonetic information |
Non-Patent Citations (2)
| Title |
|---|
| HERMANSKY H ET AL: "Tandem connectionist feature extraction for conventional HMM systems", ACOUSTICS, SPEECH, AND SIGNAL PROCESSING, 2000. ICASSP '00. PROCEEDING S. 2000 IEEE INTERNATIONAL CONFERENCE ON 5-9 JUNE 2000, PISCATAWAY, NJ, USA,IEEE, vol. 3, 5 June 2000 (2000-06-05), pages 1635 - 1638, XP010507669, ISBN: 978-0-7803-6293-2 * |
| JOEL PINTO ET AL: "Exploiting contextual information for improved phoneme recognition", ACOUSTICS, SPEECH AND SIGNAL PROCESSING, 2008. ICASSP 2008. IEEE INTERNATIONAL CONFERENCE ON, IEEE, PISCATAWAY, NJ, USA, 31 March 2008 (2008-03-31), pages 4449 - 4452, XP031251585, ISBN: 978-1-4244-1483-3 * |
Also Published As
| Publication number | Publication date |
|---|---|
| US9595257B2 (en) | 2017-03-14 |
| US20120239403A1 (en) | 2012-09-20 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US9595257B2 (en) | Downsampling schemes in a hierarchical neural network structure for phoneme recognition | |
| CN108447490B (en) | Method and device for voiceprint recognition based on memory bottleneck feature | |
| KR102294638B1 (en) | Combined learning method and apparatus using deepening neural network based feature enhancement and modified loss function for speaker recognition robust to noisy environments | |
| Peddinti et al. | A time delay neural network architecture for efficient modeling of long temporal contexts | |
| US8554555B2 (en) | Method for automated training of a plurality of artificial neural networks | |
| US8838446B2 (en) | Method and apparatus of transforming speech feature vectors using an auto-associative neural network | |
| KR101704926B1 (en) | Statistical Model-based Voice Activity Detection with Ensemble of Deep Neural Network Using Acoustic Environment Classification and Voice Activity Detection Method thereof | |
| US20180232632A1 (en) | Efficient connectionist temporal classification for binary classification | |
| JP2000099080A (en) | Speech Recognition Method Using Reliability Scale Evaluation | |
| EP2903003A1 (en) | Online maximum-likelihood mean and variance normalization for speech recognition | |
| Mallidi et al. | Autoencoder based multi-stream combination for noise robust speech recognition. | |
| EP0725383B1 (en) | Pattern adaptation system using tree scheme | |
| KR102429656B1 (en) | A speaker embedding extraction method and system for automatic speech recognition based pooling method for speaker recognition, and recording medium therefor | |
| CN102237082B (en) | Self-adaption method of speech recognition system | |
| Zhao et al. | Stranded Gaussian mixture hidden Markov models for robust speech recognition | |
| Frikha et al. | A comparative survey of ANN and hybrid HMM/ANN architectures for robust speech recognition | |
| Ali et al. | Enhancing Embeddings for Speech Classification in Noisy Conditions. | |
| Li et al. | DNN online adaptation for automatic speech recognition | |
| Räsänen et al. | A noise robust method for pattern discovery in quantized time series: the concept matrix approach. | |
| Abdelaziz | Turbo Decoders for Audio-Visual Continuous Speech Recognition. | |
| Vaásquez et al. | On speeding phoneme recognition in a hierarchical MLP structure | |
| Li | Noise robust speech recognition using deep neural network | |
| Hirota et al. | Experimental evaluation of structure of garbage model generated from in-vocabulary words | |
| KR102935627B1 (en) | Method for real-time signal processing using convolution recurrent neural network | |
| Moons et al. | Resource aware design of a deep convolutional-recurrent neural network for speech recognition through audio-visual sensor fusion |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 09793049 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 13497119 Country of ref document: US |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 09793049 Country of ref document: EP Kind code of ref document: A1 |




