EP1514258B1 - Distribution de frequence de distance vectorielle minimale pour alignement temporel dynamique - Google Patents

Distribution de frequence de distance vectorielle minimale pour alignement temporel dynamique Download PDF

Info

Publication number
EP1514258B1
EP1514258B1 EP03729040A EP03729040A EP1514258B1 EP 1514258 B1 EP1514258 B1 EP 1514258B1 EP 03729040 A EP03729040 A EP 03729040A EP 03729040 A EP03729040 A EP 03729040A EP 1514258 B1 EP1514258 B1 EP 1514258B1
Authority
EP
European Patent Office
Prior art keywords
vectors
template
matching
utterance
distribution
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Expired - Lifetime
Application number
EP03729040A
Other languages
German (de)
English (en)
Other versions
EP1514258A4 (fr
EP1514258A1 (fr
Inventor
Veton Z. Kepuska
Harinath K. Reddy
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
ThinkEngine Networks Inc
Original Assignee
ThinkEngine Networks Inc
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by ThinkEngine Networks Inc filed Critical ThinkEngine Networks Inc
Publication of EP1514258A1 publication Critical patent/EP1514258A1/fr
Publication of EP1514258A4 publication Critical patent/EP1514258A4/fr
Application granted granted Critical
Publication of EP1514258B1 publication Critical patent/EP1514258B1/fr
Anticipated expiration legal-status Critical
Expired - Lifetime legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/08Speech classification or search
    • G10L15/12Speech classification or search using dynamic programming techniques, e.g. dynamic time warping [DTW]
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/08Speech classification or search
    • G10L2015/088Word spotting

Definitions

  • This description relates to dynamic time warping of speech.
  • Speech is a time-dependent process of high variability.
  • One variability is the duration of a spoken word. Multiple utterances of a particular word by a single speaker may have different durations. Even when utterances of the word happen to have the same duration, a particular part of the word will often have different durations among the utterances. Durations also vary between speakers for utterances of a given word or part of a word.
  • Speech processing for example, speech recognition, often involves comparing two instances of a word, such as comparing an uttered word to a model of a word.
  • the durational variations of uttered words and parts of words can be accommodated by a non-linear time warping designed to align speech features of two speech instances that correspond to the same acoustic events before comparing the two speech instances.
  • Dynamic time warping is a dynamic programming technique suitable to match patterns that are time dependent.
  • the result of applying DTW is a measure of similarity of a test pattern (for example, an uttered word) and a reference pattern (e.g., a template or model of a word).
  • a test pattern for example, an uttered word
  • a reference pattern e.g., a template or model of a word.
  • Each test pattern and each reference pattern may be represented as a sequence of vectors.
  • the two speech patterns are aligned in time and DTW measures a global distance between the two sequences of vectors.
  • a time-time matrix 10 illustrates the alignment process.
  • the uttered word is represented by a sequence of feature vectors 12 (also called frames) arrayed along the horizontal axis.
  • the template or model of the word is represented by a sequence of feature vectors 14 (also called frames) arrayed along the vertical axis.
  • the feature vectors are generated at intervals of, for example, 0.01 sec (e.g., 100 feature vectors per second).
  • Each feature vector captures properties of speech typically centered within 20-30 msec. Properties of the speech signal generally do not change significantly within a time duration of the analysis window (i.e., 20-30 msec).
  • the analysis window is shifted by 0.01 sec to capture the properties of the speech signal in each successive time instance.
  • the utterance is SsPEEhH, a noisy version of the template SPEECH.
  • the utterance SsPEEhH will typically be compared to all other templates (i.e., reference patterns or models that correspond to other words) in a repository to find the template that is the best match.
  • the best matching template is deemed to be the one that has the lowest global distance from the utterance, computed along a path 16 that best aligns the utterance with a given template, i.e., produces the lowest global distance of any alignment path between the utterance and the given template.
  • a path we mean a series of associations between frames of the utterance and corresponding frames of the template. The complete universe of possible paths includes every possible set of associations between the frames.
  • a global distance of a path is the sum of local distances for each of the associations of the path.
  • One way to find the path that yields the best match (i.e., lowest global distance) between the utterance and a given template is by evaluating all possible paths in the universe. That approach is time consuming because the number of possible paths is exponential with the length of the utterance.
  • the matching process can be shortened by requiring that (a) a path cannot go backwards in time; (i.e., to the left or down in figure 1) (b) a path must include an association for every frame in the utterance, and (c) local distance scores are combined by adding to give a global distance.
  • DP dynamic programming
  • a point (i, j) the next selected point on the path, comes from one among (i-1, j-1), (i-1, j) or (i, j-1) that has the lowest distance.
  • DTW refers to this application of dynamic programming to speech recognition.
  • DP finds the lowest distance path through the matrix, while minimizing the amount of computation.
  • the DP algorithm operates in a time-synchronous manner by considering each column of the time-time matrix in succession (which is equivalent to processing the utterance frame-by-frame). For a template of length N (corresponding to an N-row matrix), the maximum number of paths being considered at any time is N.
  • a test utterance feature vector j is compared to all reference template features, 1...N, thus generating a vector of corresponding local distances d(1, j), d(2, j), ... d(N, j).
  • DP For basic speech recognition, DP has a small memory requirement.
  • the only storage required by the search (as distinct from storage required for the templates) is an array that holds a single column of the time-time matrix.
  • Equation 1 enforces the rule that the only directions in which a path can move when at ( i, j ) in the time-time matrix is up, right, or diagonally up and right.
  • equation 1 is in a form that could be recursively programmed. However, unless the language is optimized for recursion, this method can be slow even for relatively small pattern sizes. Another method that is both quicker and requires less memory storage uses two nested "for" loops. This method only needs two arrays that hold adjacent columns of the time-time matrix.
  • the algorithm to find the least global distance path is as follows (Note that, in figure 2, which shows a representative set of rows and columns, the cells at (i, j) 22 and (i, 0) have different possible originator cells.
  • the path to (i, 0) 24 can originate only from (i - 1, 0). But the path to any other (i, j) can originate from the three standard locations 26, 28, 29):
  • the algorithm is repeated for each template.
  • the template file that gives the lowest global matching score is picked as the most likely word.
  • the minimum global matching score for a template need be compared only with a relatively limited number of alternative score values representing other minimum global matching scores (that is, even though there may be 1000 templates, many templates will share the same value for minimum global score). All that is needed for a correct recognition is the best matching score to be produced by a corresponding template. The best matching score is simply the one that is relatively the lowest compared to other scores. Thus, we may call this a "relative scoring" approach to word matching.
  • the situation in which two matches share a common score can be resolved in various ways, for example, by a tie breaking rule, by asking the user to confirm, or by picking the one that has lowest maximal local distance. However, the case of ties is irrelevant for recognizing a single "OnWord" and practically never occurs.
  • This algorithm works well for tasks having a relatively small number of possible choices, for example, recognizing one word from among 10-100 possible ones.
  • the average number of alternatives for a given recognition cycle of an utterance is called the perplexity of the recognition.
  • the algorithm is not practical for real-time tasks that have nearly infinite perplexity, for example, correctly detecting and recognizing a specific word/command phrase (for example, a so-called wake-up word, hot word or OnWord) from all other possible words/phrases/sounds. It is impractical to have a corresponding model for every possible word/phrase/sound that is not the word to be recognized. And absolute values of matching scores are not suited to select correct word recognition because of wide variability in the scores.
  • templates against which an utterance are to be matched may be divided between those that are within a vocabulary of interest (called in-vocabulary or INV) and those that are outside a vocabulary of interest (called out-of-vocabulary or OOV). Then a threshold can be set so that an utterance that yields a test score below the threshold is deemed to be INV, and an utterance that has a score greater than the threshold is considered not be a correct word (OOV).
  • this approach can correctly recognize less than 50% of INV words (correct acceptance) and treats 50% of uttered words as OOV (false rejection).
  • OOV false rejection
  • using the same threshold about 5% - 10% of utterances of OOV words would be recognized as INV (false acceptance).
  • US-A-5 710 864 discloses a speech recognition confidence measure.
  • the invention as defined in the independent claims provides a method comprising measuring distances between vectors that represent an utterance and vectors that represent a template, generating a distribution of values associated with the vectors of the template, the values indicating how many times reference template vectors produce a minimum local distance in matching with vectors of the utterance, and making a matching decision based on the measured distances and on the generated distribution.
  • Implementations of the invention may include one or more of the following.
  • the generating of information includes producing a distribution of values associated with the vectors of the template, the values indicating the frequency with which reference template vectors produce a minimum local distance in matching with vectors of the utterance.
  • the matching decision is based on the extent to which the distribution is non-uniform.
  • the matching decision is based on the spikiness of the distribution.
  • the matching decision is based on how well the entire set of vectors representing the template are used in the matching.
  • the measuring of distances includes generating a raw dynamic time warping score and rescoring the score based on the information indicative of how well the vectors of the utterance match the vectors of the template.
  • the rescoring is based on both the spikiness of the distribution and the on how well the entire set of vectors representing the template are used in the matching.
  • one way to improve the relative scoring method is to repeat each matching 102 of the speech features 104 of an utterance (which we also call a test pattern) with a given template 106 (e.g., an on word template) by using the same frames of the template but reversing their order 108 .
  • a test pattern that is INV (as defined earlier) and that has a relatively good matching score in the initial matching is expected to have a significantly worse matching score in the second matching 110 with the reversed-order template.
  • One reason for this difference in scores is the constraint of DTW that prohibits a matching path from progressing backward in time.
  • an OOV (i.e., out-of-vocabulary) test pattern would be expected to have comparable matching scores for both the normal and reverse orders of frames because matching of features may be more or less random regardless of the order.
  • the template is for the word DISCO VE R and the test pattern is for the word t r a ve l
  • the features corresponding to the underlined sounds would match well but not the ones marked in italics.
  • the other features of the words would match poorly.
  • R E V OCSID and t r a v el The global matching score will be more or less about the same as in normal order.
  • a similar effect can be achieved by reversing the order of frames in the test pattern 120 instead of the template (which we also call the reference pattern).
  • Which approach is used may depend on the hardware platform and software architectural solution of the application. One goal is the minimization of the latency of the response of the system. If CPU and memory resources are not an issue, both approaches can be used and all three resulting scores may be used to achieve a more accurate recognition and rejection.
  • Figure 3 shows a scatter plot 30 of average global distance (i.e., global distance divided by number of test vectors) of the reversed reference pattern and average global distance of the pattern in its original order.
  • Each point plots the distance of the same uttered word (token) against an original-order template and a reversed-order template
  • a point on the 45-degree line 32 would have the same distance for matching done in the normal order of frames and for matching done in the reverse order of frames.
  • Diamond points 34 are OOV and square points 36 are INV.
  • the OOV points are clustered largely along the diagonal line 32, which indicates that average global distance for the original order matching and for the reverse order matching are generally uncorrelated with the direction of the frames.
  • the scatter of INV points is largely above the diagonal indicating that the average global distances are correlated with the direction of the frames.
  • the INV test words that correspond to square points below the diagonal line largely depict cases associated with a failure of the voice activity detector to correctly determine the beginning and ending of the test word, thus producing partial segmentation of a word or missing it completely.
  • the remaining INV test words correspond to square points that can be recovered by normalization rescoring described later.
  • the line 38 in figure 3 depicts a discriminating surface that would maximally separate OOV points from INV points (e.g., a minimal classification error rate in Bayesian terms).
  • the reverse order matching step can be used in a variety of ways:
  • This second approach supplements the matching of original features (frames, vectors) by a measure of how well the feature vectors of the test pattern (the uttered word) are matched by the feature vectors of the reference pattern.
  • An INV word ideally would match each reference pattern's vectors only once.
  • An OOV word would typically match only sections of each reference pattern (e.g., in the words DISCO VE R and t r a ve l, the features corresponding to the sounds that are underscored would match well as would the features corresponding to the sounds that are italicized, but the sequence of the matches would be in reversed order).
  • the counts in the counters together represent a distribution of the number of minimum distances as a function of index.
  • the distribution is more or less uniform and flat, because all vectors of the reference pattern form minimum distance matches with the vectors of the test pattern. This situation is typically true for an INV word. If test vectors are from INV words, good matching would occur for each test feature vector (e.g., each sound) distributed across all reference vectors. For an OOV word, by contrast, the distribution would be significantly less uniform and peaky-er, because only some sections of the feature vectors would match, as in the example above (figure 4), while other vectors would match only at random. Thus, the dissimilarity of the distribution of the number of minimum distance matches from a uniform distribution is a good discriminator of INV and OOV words.
  • This method is useful in cases where an OOV test pattern ⁇ T ⁇ yields a small DTW global matching score (incorrectly suggesting a good match) and thus makes recognition invalid (false acceptance). Because the DMDI of this match significantly departs from uniform, it is possible to exploit the non-uniformity as a basis for disregarding the putative match to prevent an invalid recognition.
  • Figure 4 depicts DMDIs for an INV word (shown by squares), an OOV word (diamonds) and a uniform distribution (triangles) for comparison.
  • Figure 4 shows that the DMDI of an OOV word is spikier than the DMDI of an INV word, and that the number of reference vectors that do not become minimal distance vectors is significantly larger for OOV (38 vectors in the example) than for INV (21 vectors).
  • a measure of uniformity of distribution that will account for these effects provides a way to rescore an original DTW score to avoid invalid recognitions of OOV words.
  • the uniformity measure would be used to reduce the DTW raw score for a DMDI that is closer to a uniform distribution and to increase the DTW raw score for a DMDI that departs from a uniform distribution.
  • One example of a uniformity measure would be based on two functions.
  • R represents the number of reference vectors in the reference pattern.
  • a second function is based on a value, ⁇ , that represents a measure of how well the entire set of vectors of a reference pattern are used in matching.
  • is the total number of reference vectors that did not have minimal distance from any test vector.
  • the values of the two functions ⁇ and ⁇ are combined to give one DMDI penalty; e.g, ⁇ * ⁇ that is used to rescore an original DTW raw score of a DMDI.
  • the ratio ⁇ /R determines the rate of the proportion of reference vectors that were not selected as best representing vectors out of the total number of reference vectors. ⁇ is used to control minimal number ⁇ must have before this measure can have any significant effect.
  • the results of matching can be subjected to a first rescoring 122 based on a scatter plot, for example, of the kind shown in figure 3 and a second rescoring 124 based on a DMDI, for example, of the kind shown in figure 4.
  • the rescorings can then be combined 126.
  • Combining of scores may be done in the following way.
  • the dynamic time warping module emits its final raw matching scores 128, 130 for original ordered reference template vectors of the template and reversed ordered reference template vectors (or alternatively reversed ordered test template vectors or both).
  • aver_score extract_h ( L_shl ( L_mult ( raw_score , den ) , 6 ) )
  • aver_revscore extract_h ( L_shl ( L_mult ( raw_revscore , den ) , 6 ) )
  • den is a fixed point representation of 1/T
  • L_mult is a fixed point multiplication operator
  • L_shl is a shift left operator
  • extract_h is a high order byte extraction operation.
  • n_zeros and n_bestmatchs are the parameters that are computed: the number of zero matches and the distribution distortion.
  • Fixed-point implementation of the computation of rescoring parameters respectively denoted by n_zeros and n_bestmatchs may be done as follows: where sub denotes a fixed point subtraction operator, L_mac denotes a multiply -add fixed point operator, and sature denotes a saturation operator ensuring that the content of the variable saturates to the given fixed point precision without overflowing.
  • n_zeros of reverse and normal order matches
  • norm_n_zeros n_zeros + T ⁇ R T
  • norm_reverse_n_zeros reverse_n_zeros + T ⁇ R T
  • a correction term is computed as the difference of norm_n_zeros-norm_reverse_n_zeros.
  • delta_n_zeros norm_n_zeros ⁇ norm_reverse_n_zeros
  • norm_n_zeros has lower values than norm_reverse_n_zeros thus this correction will reduce the score.
  • norm_reverse_n_zeros For OOV words those two parameters have similar values thus not contributing to adjustment of the original score at all.
  • delta_scatter L_shl(L_sub(L_add(L_shl(aver_score,1),mult(aver_score,21845)), aver_revscore), 2);
  • a speech signal 50 is received from a telephone line, a network medium, wirelessly, or in any other possible way, at a speech processor 48.
  • the speech processor may perform one or more of a wide variety of speech processing functions for one or more (including a very large number of) incoming speech signals 50.
  • One of the functions may be speech recognition.
  • the speech processor may be part of a larger system.
  • the speech processor may include a dynamic time warping device that implements the matching elements 102, 110 and that uses the speech signal and speech templates 106, 108, 120 to generate a raw DTW score 56.
  • the rescoring devices modify the raw DTW score to generate rescored DTW values.
  • the rescored DTW values may then be used in generating recognized speech 60.
  • the recognized speech may be used for a variety of purposes.

Landscapes

  • Engineering & Computer Science (AREA)
  • Acoustics & Sound (AREA)
  • Health & Medical Sciences (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Human Computer Interaction (AREA)
  • Physics & Mathematics (AREA)
  • Computational Linguistics (AREA)
  • Multimedia (AREA)
  • Image Analysis (AREA)
  • Design And Manufacture Of Integrated Circuits (AREA)
  • Piezo-Electric Or Mechanical Vibrators, Or Delay Or Filter Circuits (AREA)
  • Machine Translation (AREA)
  • Filters And Equalizers (AREA)
  • Measuring Frequencies, Analyzing Spectra (AREA)
  • Radar Systems Or Details Thereof (AREA)

Claims (9)

  1. Procédé de reconnaissance de la parole, comprenant les étapes de:
    mesure de distances entre des vecteurs qui représentent une émission de son et des vecteurs qui représentent un gabarit;
    génération d'une distribution de valeurs qui sont associées aux vecteurs du gabarit, les valeurs indiquant combien de fois des vecteurs de gabarit de référence produisent une distance locale minimum en correspondance avec des vecteurs de l'émission de son; et
    réalisation d'une décision de correspondance sur la base des distances mesurées et sur la base de la distribution générée.
  2. Procédé selon la revendication 1, dans lequel la décision de correspondance est basée sur l'étendue selon laquelle la distribution est non uniforme.
  3. Procédé selon la revendication 2, dans lequel la décision de correspondance est basée sur la caractéristique de pointe de la distribution.
  4. Procédé selon la revendication 2, dans lequel la décision de correspondance est basée sur le degré de pertinence selon lequel le jeu complet de vecteurs représentant le gabarit est utilisé au niveau de la correspondance.
  5. Procédé selon la revendication 1, dans lequel la mesure de distances inclut la génération d'un score de dérive temporelle dynamique brute et le recalcul du score sur la base de l'information indicative du degré de pertinence selon lequel les vecteurs de l'émission de son correspondent aux vecteurs du gabarit.
  6. Procédé selon la revendication 5, dans lequel le recalcul du score est basé sur à la fois la caractéristique de pointe de la distribution et le degré de pertinence selon lequel le jeu complet de vecteurs représentant le gabarit est utilisé au niveau de la correspondance.
  7. Dispositif de reconnaissance de la parole, comprenant:
    un moyen qui est adapté pour mesurer des distances entre des vecteurs qui représentent une émission de son et des vecteurs qui représentent un gabarit;
    un moyen qui est adapté pour générer une distribution de valeurs qui sont associées aux vecteurs du gabarit, les valeurs indiquant combien de fois des vecteurs de gabarit de référence produisent une distance locale minimum au niveau d'une correspondance avec des vecteurs de l'émission de son; et
    un moyen qui est adapté pour réaliser une décision de correspondance sur la base des distances mesurées et sur la base de la distribution générée.
  8. Appareil comprenant:
    un port d'entrée qui est connecté pour recevoir une parole numérisée; et
    un dispositif de reconnaissance de la parole selon la revendication 7.
  9. Support lisible par ordinateur qui est porteur d'instructions qui, lorsqu'elles sont exécutées, ont pour effet qu'un ordinateur met en oeuvre les étapes qui suivent:
    mesure de distances entre des vecteurs qui représentent une émission de son et des vecteurs qui représentent un gabarit;
    génération d'une distribution de valeurs qui sont associées aux vecteurs du gabarit, les valeurs indiquant combien de fois des vecteurs de gabarit de référence produisent une distance locale minimum au niveau d'une correspondance avec des vecteurs de l'émission de son; et
    réalisation d'une décision de correspondance sur la base des distances mesurées et sur la base de la distribution générée.
EP03729040A 2002-05-21 2003-05-21 Distribution de frequence de distance vectorielle minimale pour alignement temporel dynamique Expired - Lifetime EP1514258B1 (fr)

Applications Claiming Priority (3)

Application Number Priority Date Filing Date Title
US152447 2002-05-21
US10/152,447 US6983246B2 (en) 2002-05-21 2002-05-21 Dynamic time warping using frequency distributed distance measures
PCT/US2003/015900 WO2003100768A1 (fr) 2002-05-21 2003-05-21 Distribution de frequence de distance vectorielle minimale pour la transformation par alignement dynamique

Publications (3)

Publication Number Publication Date
EP1514258A1 EP1514258A1 (fr) 2005-03-16
EP1514258A4 EP1514258A4 (fr) 2005-08-24
EP1514258B1 true EP1514258B1 (fr) 2006-08-16

Family

ID=29548481

Family Applications (1)

Application Number Title Priority Date Filing Date
EP03729040A Expired - Lifetime EP1514258B1 (fr) 2002-05-21 2003-05-21 Distribution de frequence de distance vectorielle minimale pour alignement temporel dynamique

Country Status (7)

Country Link
US (1) US6983246B2 (fr)
EP (1) EP1514258B1 (fr)
AT (1) ATE336777T1 (fr)
AU (1) AU2003233603A1 (fr)
CA (1) CA2486204A1 (fr)
DE (1) DE60307633T2 (fr)
WO (1) WO2003100768A1 (fr)

Families Citing this family (24)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
KR100462472B1 (ko) * 2002-09-11 2004-12-17 학교법인 포항공과대학교 동적 타임 워핑 디바이스와 이를 이용한 음성 인식 장치
EP1723636A1 (fr) * 2004-03-12 2006-11-22 Siemens Aktiengesellschaft Determination de seuils de fiabilite et de rejet avec adaptation a l'utilisateur et au vocabulaire
EP1693830B1 (fr) * 2005-02-21 2017-12-20 Harman Becker Automotive Systems GmbH Système de données à commande vocale
JP5396044B2 (ja) * 2008-08-20 2014-01-22 株式会社コナミデジタルエンタテインメント ゲーム装置、ゲーム装置の制御方法、及びプログラム
US8639506B2 (en) * 2010-03-11 2014-01-28 Telefonica, S.A. Fast partial pattern matching system and method
US8340429B2 (en) 2010-09-18 2012-12-25 Hewlett-Packard Development Company, Lp Searching document images
US8639508B2 (en) * 2011-02-14 2014-01-28 General Motors Llc User-specific confidence thresholds for speech recognition
US9324323B1 (en) 2012-01-13 2016-04-26 Google Inc. Speech recognition using topic-specific language models
US8775177B1 (en) 2012-03-08 2014-07-08 Google Inc. Speech recognition process
EP3077999B1 (fr) * 2013-12-06 2022-02-02 The ADT Security Corporation Application activée par la voix pour dispositifs mobiles
US10796805B2 (en) 2015-10-08 2020-10-06 Cordio Medical Ltd. Assessment of a pulmonary condition by speech analysis
US10089989B2 (en) * 2015-12-07 2018-10-02 Semiconductor Components Industries, Llc Method and apparatus for a low power voice trigger device
US10847177B2 (en) 2018-10-11 2020-11-24 Cordio Medical Ltd. Estimating lung volume by speech analysis
US12512114B2 (en) 2019-03-12 2025-12-30 Cordio Medical Ltd. Analyzing speech using speech models and segmentation based on acoustic features
US12494224B2 (en) 2019-03-12 2025-12-09 Cordio Medical Ltd. Analyzing speech using speech-sample alignment and segmentation based on acoustic features
US11024327B2 (en) 2019-03-12 2021-06-01 Cordio Medical Ltd. Diagnostic techniques based on speech models
US12488805B2 (en) 2019-03-12 2025-12-02 Cordio Medical Ltd. Using optimal articulatory event-types for computer analysis of speech
US11011188B2 (en) 2019-03-12 2021-05-18 Cordio Medical Ltd. Diagnostic techniques based on speech-sample alignment
KR20210137503A (ko) * 2019-03-12 2021-11-17 코디오 메디칼 리미티드 음성 모델에 기반한 진단 기법
US11484211B2 (en) 2020-03-03 2022-11-01 Cordio Medical Ltd. Diagnosis of medical conditions using voice recordings and auscultation
US11417342B2 (en) 2020-06-29 2022-08-16 Cordio Medical Ltd. Synthesizing patient-specific speech models
US12334105B2 (en) 2020-11-23 2025-06-17 Cordio Medical Ltd. Detecting impaired physiological function by speech analysis
US12518774B2 (en) 2023-02-05 2026-01-06 Cordio Medical Ltd. Identifying optimal articulatory event-types for computer analysis of speech
US12555595B2 (en) 2023-05-18 2026-02-17 Cordio Medical Ltd. Converting a sequence of speech records of a human subject into a sequence of indicators of a physiological state of the subject

Family Cites Families (9)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CA1204855A (fr) * 1982-03-23 1986-05-20 Phillip J. Bloom Methode et appareil utilises dans le traitement des signaux
GB2145864B (en) * 1983-09-01 1987-09-03 King Reginald Alfred Voice recognition
JPS62232000A (ja) * 1986-03-25 1987-10-12 インタ−ナシヨナル・ビジネス・マシ−ンズ・コ−ポレ−シヨン 音声認識装置
US4803729A (en) * 1987-04-03 1989-02-07 Dragon Systems, Inc. Speech recognition method
US5159637A (en) * 1988-07-27 1992-10-27 Fujitsu Limited Speech word recognizing apparatus using information indicative of the relative significance of speech features
US5621859A (en) * 1994-01-19 1997-04-15 Bbn Corporation Single tree method for grammar directed, very large vocabulary speech recognizer
US5710864A (en) * 1994-12-29 1998-01-20 Lucent Technologies Inc. Systems, methods and articles of manufacture for improving recognition confidence in hypothesized keywords
US6076054A (en) * 1996-02-29 2000-06-13 Nynex Science & Technology, Inc. Methods and apparatus for generating and using out of vocabulary word models for speaker dependent speech recognition
US6223155B1 (en) * 1998-08-14 2001-04-24 Conexant Systems, Inc. Method of independently creating and using a garbage model for improved rejection in a limited-training speaker-dependent speech recognition system

Also Published As

Publication number Publication date
US6983246B2 (en) 2006-01-03
US20030220790A1 (en) 2003-11-27
EP1514258A4 (fr) 2005-08-24
ATE336777T1 (de) 2006-09-15
EP1514258A1 (fr) 2005-03-16
CA2486204A1 (fr) 2003-12-04
DE60307633D1 (de) 2006-09-28
WO2003100768A1 (fr) 2003-12-04
DE60307633T2 (de) 2007-08-16
AU2003233603A1 (en) 2003-12-12

Similar Documents

Publication Publication Date Title
EP1514258B1 (fr) Distribution de frequence de distance vectorielle minimale pour alignement temporel dynamique
US5684925A (en) Speech representation by feature-based word prototypes comprising phoneme targets having reliable high similarity
US6493667B1 (en) Enhanced likelihood computation using regression in a speech recognition system
US6226612B1 (en) Method of evaluating an utterance in a speech recognition system
EP0501631B1 (fr) Procédé de décorrélation temporelle pour vérification robuste de locuteur
JP4218982B2 (ja) 音声処理
US6260013B1 (en) Speech recognition system employing discriminatively trained models
CN102129860B (zh) 基于无限状态隐马尔可夫模型的与文本相关的说话人识别方法
US6490555B1 (en) Discriminatively trained mixture models in continuous speech recognition
US8374869B2 (en) Utterance verification method and apparatus for isolated word N-best recognition result
EP0504485B1 (fr) Appareil pour coder un label, indépendant du locuteur
US7617103B2 (en) Incrementally regulated discriminative margins in MCE training for speech recognition
US20090119103A1 (en) Speaker recognition system
KR100307623B1 (ko) 엠.에이.피 화자 적응 조건에서 파라미터의 분별적 추정 방법 및 장치 및 이를 각각 포함한 음성 인식 방법 및 장치
EP3513404A1 (fr) Sélection de microphones et segmentation de multiples locuteurs avec reconnaissance vocale automatique (asr) ambiante
US5825977A (en) Word hypothesizer based on reliably detected phoneme similarity regions
US5621849A (en) Voice recognizing method and apparatus
JP2023532844A (ja) 患者固有の音声モデルの合成
US7346497B2 (en) High-order entropy error functions for neural classifiers
US7085717B2 (en) Scoring and re-scoring dynamic time warping of speech
JP2001083986A (ja) 統計モデル作成方法
US20050192806A1 (en) Probability density function compensation method for hidden markov model and speech recognition method and apparatus using the same
Savchenko et al. Optimization of gain in symmetrized itakura-saito discrimination for pronunciation learning
Khasawneh et al. The application of polynomial discriminant function classifiers to isolated Arabic speech recognition
JP3009962B2 (ja) 音声認識装置

Legal Events

Date Code Title Description
PUAI Public reference made under article 153(3) epc to a published international application that has entered the european phase

Free format text: ORIGINAL CODE: 0009012

17P Request for examination filed

Effective date: 20041220

AK Designated contracting states

Kind code of ref document: A1

Designated state(s): AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HU IE IT LI LU MC NL PT RO SE SI SK TR

AX Request for extension of the european patent

Extension state: AL LT LV MK

A4 Supplementary search report drawn up and despatched

Effective date: 20050711

RIC1 Information provided on ipc code assigned before grant

Ipc: 7G 10L 15/12 A

DAX Request for extension of the european patent (deleted)
GRAP Despatch of communication of intention to grant a patent

Free format text: ORIGINAL CODE: EPIDOSNIGR1

RIN1 Information on inventor provided before grant (corrected)

Inventor name: REDDY, HARINATH, K.

Inventor name: KEPUSKA, VETON, Z.

GRAS Grant fee paid

Free format text: ORIGINAL CODE: EPIDOSNIGR3

GRAA (expected) grant

Free format text: ORIGINAL CODE: 0009210

AK Designated contracting states

Kind code of ref document: B1

Designated state(s): AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HU IE IT LI LU MC NL PT RO SE SI SK TR

PG25 Lapsed in a contracting state [announced via postgrant information from national office to epo]

Ref country code: IT

Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT;WARNING: LAPSES OF ITALIAN PATENTS WITH EFFECTIVE DATE BEFORE 2007 MAY HAVE OCCURRED AT ANY TIME BEFORE 2007. THE CORRECT EFFECTIVE DATE MAY BE DIFFERENT FROM THE ONE RECORDED.

Effective date: 20060816

Ref country code: LI

Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT

Effective date: 20060816

Ref country code: CH

Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT

Effective date: 20060816

Ref country code: CZ

Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT

Effective date: 20060816

Ref country code: SK

Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT

Effective date: 20060816

Ref country code: BE

Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT

Effective date: 20060816

Ref country code: SI

Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT

Effective date: 20060816

Ref country code: FI

Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT

Effective date: 20060816

Ref country code: RO

Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT

Effective date: 20060816

Ref country code: AT

Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT

Effective date: 20060816

Ref country code: NL

Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT

Effective date: 20060816

REG Reference to a national code

Ref country code: GB

Ref legal event code: FG4D

REG Reference to a national code

Ref country code: CH

Ref legal event code: EP

REG Reference to a national code

Ref country code: IE

Ref legal event code: FG4D

REF Corresponds to:

Ref document number: 60307633

Country of ref document: DE

Date of ref document: 20060928

Kind code of ref document: P

PG25 Lapsed in a contracting state [announced via postgrant information from national office to epo]

Ref country code: DK

Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT

Effective date: 20061116

Ref country code: SE

Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT

Effective date: 20061116

Ref country code: BG

Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT

Effective date: 20061116

PG25 Lapsed in a contracting state [announced via postgrant information from national office to epo]

Ref country code: ES

Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT

Effective date: 20061127

PG25 Lapsed in a contracting state [announced via postgrant information from national office to epo]

Ref country code: PT

Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT

Effective date: 20070116

NLV1 Nl: lapsed or annulled due to failure to fulfill the requirements of art. 29p and 29m of the patents act
REG Reference to a national code

Ref country code: CH

Ref legal event code: PL

ET Fr: translation filed
PLBE No opposition filed within time limit

Free format text: ORIGINAL CODE: 0009261

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: NO OPPOSITION FILED WITHIN TIME LIMIT

26N No opposition filed

Effective date: 20070518

PG25 Lapsed in a contracting state [announced via postgrant information from national office to epo]

Ref country code: MC

Free format text: LAPSE BECAUSE OF NON-PAYMENT OF DUE FEES

Effective date: 20070531

PG25 Lapsed in a contracting state [announced via postgrant information from national office to epo]

Ref country code: GR

Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT

Effective date: 20061117

PG25 Lapsed in a contracting state [announced via postgrant information from national office to epo]

Ref country code: EE

Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT

Effective date: 20060816

PG25 Lapsed in a contracting state [announced via postgrant information from national office to epo]

Ref country code: CY

Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT

Effective date: 20060816

Ref country code: LU

Free format text: LAPSE BECAUSE OF NON-PAYMENT OF DUE FEES

Effective date: 20070521

PG25 Lapsed in a contracting state [announced via postgrant information from national office to epo]

Ref country code: HU

Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT

Effective date: 20070217

Ref country code: TR

Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT

Effective date: 20060816

PGFP Annual fee paid to national office [announced via postgrant information from national office to epo]

Ref country code: IE

Payment date: 20090825

Year of fee payment: 7

REG Reference to a national code

Ref country code: GB

Ref legal event code: 732E

Free format text: REGISTERED BETWEEN 20101021 AND 20101027

REG Reference to a national code

Ref country code: FR

Ref legal event code: TP

REG Reference to a national code

Ref country code: DE

Ref legal event code: R081

Ref document number: 60307633

Country of ref document: DE

Owner name: CASTELL SOFTWARE LLC, US

Free format text: FORMER OWNER: THINKENGINE NETWORKS INC., MARLBOROUGH, US

Effective date: 20110208

Ref country code: DE

Ref legal event code: R081

Ref document number: 60307633

Country of ref document: DE

Owner name: CASTELL SOFTWARE LLC, WILMINGTON, US

Free format text: FORMER OWNER: THINKENGINE NETWORKS INC., MARLBOROUGH, MASS., US

Effective date: 20110208

PG25 Lapsed in a contracting state [announced via postgrant information from national office to epo]

Ref country code: IE

Free format text: LAPSE BECAUSE OF NON-PAYMENT OF DUE FEES

Effective date: 20100521

REG Reference to a national code

Ref country code: FR

Ref legal event code: PLFP

Year of fee payment: 14

REG Reference to a national code

Ref country code: FR

Ref legal event code: PLFP

Year of fee payment: 15

REG Reference to a national code

Ref country code: FR

Ref legal event code: PLFP

Year of fee payment: 16

PGFP Annual fee paid to national office [announced via postgrant information from national office to epo]

Ref country code: FR

Payment date: 20210420

Year of fee payment: 19

Ref country code: DE

Payment date: 20210413

Year of fee payment: 19

PGFP Annual fee paid to national office [announced via postgrant information from national office to epo]

Ref country code: GB

Payment date: 20210428

Year of fee payment: 19

REG Reference to a national code

Ref country code: DE

Ref legal event code: R119

Ref document number: 60307633

Country of ref document: DE

GBPC Gb: european patent ceased through non-payment of renewal fee

Effective date: 20220521

PG25 Lapsed in a contracting state [announced via postgrant information from national office to epo]

Ref country code: FR

Free format text: LAPSE BECAUSE OF NON-PAYMENT OF DUE FEES

Effective date: 20220531

PG25 Lapsed in a contracting state [announced via postgrant information from national office to epo]

Ref country code: GB

Free format text: LAPSE BECAUSE OF NON-PAYMENT OF DUE FEES

Effective date: 20220521

Ref country code: DE

Free format text: LAPSE BECAUSE OF NON-PAYMENT OF DUE FEES

Effective date: 20221201