EP2036078A1 - Verfahren und vorrichtung zur natürlichsprachlichen erkennung einer sprachäusserung - Google Patents
Verfahren und vorrichtung zur natürlichsprachlichen erkennung einer sprachäusserungInfo
- Publication number
- EP2036078A1 EP2036078A1 EP07764643A EP07764643A EP2036078A1 EP 2036078 A1 EP2036078 A1 EP 2036078A1 EP 07764643 A EP07764643 A EP 07764643A EP 07764643 A EP07764643 A EP 07764643A EP 2036078 A1 EP2036078 A1 EP 2036078A1
- Authority
- EP
- European Patent Office
- Prior art keywords
- recognition
- speech
- grammar
- utterance
- speech signal
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Withdrawn
Links
- 238000000034 method Methods 0.000 title claims abstract description 40
- 230000001755 vocal effect Effects 0.000 title abstract 3
- 238000001514 detection method Methods 0.000 claims description 9
- 238000011156 evaluation Methods 0.000 claims description 2
- 238000004590 computer program Methods 0.000 claims 2
- 238000013461 design Methods 0.000 description 3
- 238000012545 processing Methods 0.000 description 2
- 238000006243 chemical reaction Methods 0.000 description 1
- 238000004891 communication Methods 0.000 description 1
- 230000007423 decrease Effects 0.000 description 1
- 230000001419 dependent effect Effects 0.000 description 1
- 239000000203 mixture Substances 0.000 description 1
- 238000012795 verification Methods 0.000 description 1
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/08—Speech classification or search
- G10L15/18—Speech classification or search using natural language modelling
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/28—Constructional details of speech recognition systems
- G10L15/32—Multiple recognisers used in sequence or in parallel; Score combination systems therefor, e.g. voting systems
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/02—Feature extraction for speech recognition; Selection of recognition unit
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/08—Speech classification or search
- G10L15/18—Speech classification or search using natural language modelling
- G10L15/183—Speech classification or search using natural language modelling using context dependencies, e.g. language models
- G10L15/19—Grammatical context, e.g. disambiguation of the recognition hypotheses based on word sequence rules
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/28—Constructional details of speech recognition systems
Definitions
- the invention relates to a method and a device for natural language recognition of a speech utterance, in particular on the basis of a speech recognition system, which can be executed, for example, on an electronic data processing system.
- Speech recognition systems are intended for use in a variety of applications. For example, speech recognition systems are used in conjunction with office applications to capture texts or in conjunction with technical devices for their control and command input. Speech recognition systems are also used to control information and communication devices, e.g. Radio, mobile phone and navigation systems used. Companies also use voice-dialogue systems for customer advice and information, which are also based on speech recognition systems. On these latter, the patent application is related.
- the object of the invention is therefore to realize a speech recognition method and system with a large scope of recognition with a small amount of grammar.
- the inventive method is based on the detection of a speech utterance of a person and conversion into a speech signal to be processed for a speech recognition device, the analysis of the speech signal in parallel or sequentially in several speech recognition branches of the speech recognition device using multiple grammars, and the successful termination of the recognition process, if the analysis the speech signal in at least one speech recognition branch delivers a positive recognition result.
- a simultaneous analysis of the utterance by two or more independent grammars takes place.
- the speech utterance of a person triggers two or more simultaneous recognition processes, which independently analyze and evaluate the utterance.
- a comparatively small main grammar with a low recognition scope a more comprehensive secondary grammar with an extended scope of recognition is provided here. Both grammars are without common intersection.
- a second embodiment of the invention relates to a grammar cascade.
- different grammars are used in succession, ie sequentially. The moment a grammar provides a recognition result, the cascade is left and the recognition process is terminated. In this method, 100% of all utterances to be recognized are compared to the first grammar. Depending on the performance and design of this grammar, a share of, for example, 20% of unrecognized utterances to a second recognition level is passed on. In the event that a third recognition stage is installed, it can be assumed that a share of, for example, 5% of all incoming utterances reaches this third recognition stage.
- Both methods of recognition are intended to cover a wide range of statements with several "smaller" grammars, which, in combination, nevertheless guarantee a high level of recognition security, which can be done as described above in the form of a simultaneous or a successive recognition procedure.
- Figure 1 shows schematically a first embodiment of the speech recognition system with parallel-working speech recognition branches.
- Figure 2 shows schematically a second embodiment of the speech recognition system with sequentially operating, cascaded speech recognition branches.
- a speech utterance of a person present as speech signal 10 is simultaneously supplied to two speech recognition branches and analyzed by two grammars 12 and 14 (grammar A and grammar B).
- the two grammars 12, 14 have no common intersection, that is, they are based on different sets of rules.
- the parallel processing of the speech signal increases the analysis effort and thus the necessary computer load when using the method on a computer. However, this fact is outweighed by the faster recognition and significantly improved recognition security.
- a comparison 16 of the speech signal with the grammar (A) 12 results in either a positive recognition result (Yes) or a negative recognition result (No).
- a comparison 18 of the speech signal with the grammar (B) 14 results in either a positive recognition result (Yes) or a negative recognition result (No).
- four possible recognition cases result, which can be evaluated by logic 20 with different methods.
- the recognition cases 1 to 3 are unproblematic insofar as they provide clear results: Case 1 forces a non-recognition of the speech signal and thus a rejection, position 24. Cases 2 and 3 only provide a positive result and thus clearly show a recognition of the Voice signal on, position 22.
- Figure 2 shows another preferred embodiment of the invention.
- several grammars 12, 14 and 26 (grammars A, B and C) are sequentially linked together in a cascade. That is, in the grammar cascade, the various grammars 12, 14, and 26 are not addressed simultaneously, but successively.
- the recognition process can be represented as follows: the moment a grammar yields a positive recognition result, the cascade is left and the recognition process ends, position 22.
- the speech signal 10 is first supplied to a first grammar (A) 12 and analyzed there.
- a comparison 16 of the speech signal with the grammar (A) 12 results in either a positive recognition result (Yes) at which the recognition process is successfully completed or a negative recognition result (No) at which the speech signal is analyzed for further analysis of a second grammar (B ) 14 is supplied.
- a comparison 18 of the speech signal 10 with the second grammar (B) 14 results in either a positive recognition result (Yes) in which the recognition process is successfully completed or a negative recognition result (No) in which the speech signal is used for further analysis of a third grammar (C) 26 is supplied.
- a comparison 28 of the speech signal with the third grammar (C) 26 results in either a positive recognition result (Yes), in which the recognition process is successfully terminated, or a negative recognition result (No), in which the speech signal is rejected as unrecognized, Position 24th
Landscapes
- Engineering & Computer Science (AREA)
- Computational Linguistics (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Human Computer Interaction (AREA)
- Physics & Mathematics (AREA)
- Acoustics & Sound (AREA)
- Multimedia (AREA)
- Artificial Intelligence (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Machine Translation (AREA)
Abstract
Description
Claims
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| DE102006029755A DE102006029755A1 (de) | 2006-06-27 | 2006-06-27 | Verfahren und Vorrichtung zur natürlichsprachlichen Erkennung einer Sprachäußerung |
| PCT/EP2007/005224 WO2008000353A1 (de) | 2006-06-27 | 2007-06-14 | Verfahren und vorrichtung zur natürlichsprachlichen erkennung einer sprachäusserung |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP2036078A1 true EP2036078A1 (de) | 2009-03-18 |
Family
ID=38543007
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP07764643A Withdrawn EP2036078A1 (de) | 2006-06-27 | 2007-06-14 | Verfahren und vorrichtung zur natürlichsprachlichen erkennung einer sprachäusserung |
Country Status (9)
| Country | Link |
|---|---|
| US (1) | US9208787B2 (de) |
| EP (1) | EP2036078A1 (de) |
| KR (1) | KR20090033459A (de) |
| CN (1) | CN101484934B (de) |
| BR (1) | BRPI0713987A2 (de) |
| CA (1) | CA2656114C (de) |
| DE (1) | DE102006029755A1 (de) |
| RU (1) | RU2432623C2 (de) |
| WO (1) | WO2008000353A1 (de) |
Families Citing this family (17)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN101558443B (zh) | 2006-12-15 | 2012-01-04 | 三菱电机株式会社 | 声音识别装置 |
| DE102008025532B4 (de) * | 2008-05-28 | 2014-01-09 | Audi Ag | Kommunikationssystem und Verfahren zum Durchführen einer Kommunikation zwischen einem Nutzer und einer Kommunikationseinrichtung |
| DE102010040553A1 (de) * | 2010-09-10 | 2012-03-15 | Siemens Aktiengesellschaft | Spracherkennungsverfahren |
| DE102010049869B4 (de) * | 2010-10-28 | 2023-03-16 | Volkswagen Ag | Verfahren zum Bereitstellen einer Sprachschnittstelle in einem Fahrzeug und Vorrichtung dazu |
| US9431012B2 (en) | 2012-04-30 | 2016-08-30 | 2236008 Ontario Inc. | Post processing of natural language automatic speech recognition |
| US9093076B2 (en) * | 2012-04-30 | 2015-07-28 | 2236008 Ontario Inc. | Multipass ASR controlling multiple applications |
| US9601111B2 (en) * | 2012-11-13 | 2017-03-21 | GM Global Technology Operations LLC | Methods and systems for adapting speech systems |
| EP3232436A3 (de) * | 2012-11-16 | 2017-10-25 | 2236008 Ontario Inc. | Anwendungsdienst-schnittstelle zu asr |
| US9135916B2 (en) * | 2013-02-26 | 2015-09-15 | Honeywell International Inc. | System and method for correcting accent induced speech transmission problems |
| KR101370539B1 (ko) | 2013-03-15 | 2014-03-06 | 포항공과대학교 산학협력단 | 지시 표현 처리에 기반한 대화 처리 방법 및 장치 |
| US10186262B2 (en) * | 2013-07-31 | 2019-01-22 | Microsoft Technology Licensing, Llc | System with multiple simultaneous speech recognizers |
| US10885918B2 (en) | 2013-09-19 | 2021-01-05 | Microsoft Technology Licensing, Llc | Speech recognition using phoneme matching |
| US9698999B2 (en) | 2013-12-02 | 2017-07-04 | Amazon Technologies, Inc. | Natural language control of secondary device |
| US9601108B2 (en) * | 2014-01-17 | 2017-03-21 | Microsoft Technology Licensing, Llc | Incorporating an exogenous large-vocabulary model into rule-based speech recognition |
| US9552817B2 (en) * | 2014-03-19 | 2017-01-24 | Microsoft Technology Licensing, Llc | Incremental utterance decoder combination for efficient and accurate decoding |
| US10749989B2 (en) | 2014-04-01 | 2020-08-18 | Microsoft Technology Licensing Llc | Hybrid client/server architecture for parallel processing |
| CN113259736B (zh) * | 2021-05-08 | 2022-08-09 | 深圳市康意数码科技有限公司 | 一种语音控制电视机的方法及电视机 |
Family Cites Families (29)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US6249761B1 (en) * | 1997-09-30 | 2001-06-19 | At&T Corp. | Assigning and processing states and arcs of a speech recognition model in parallel processors |
| US6499013B1 (en) * | 1998-09-09 | 2002-12-24 | One Voice Technologies, Inc. | Interactive user interface using speech recognition and natural language processing |
| DE19910234A1 (de) * | 1999-03-09 | 2000-09-21 | Philips Corp Intellectual Pty | Verfahren mit mehreren Spracherkennern |
| US6526380B1 (en) * | 1999-03-26 | 2003-02-25 | Koninklijke Philips Electronics N.V. | Speech recognition system having parallel large vocabulary recognition engines |
| US7058573B1 (en) * | 1999-04-20 | 2006-06-06 | Nuance Communications Inc. | Speech recognition system to selectively utilize different speech recognition techniques over multiple speech recognition passes |
| JP4465564B2 (ja) * | 2000-02-28 | 2010-05-19 | ソニー株式会社 | 音声認識装置および音声認識方法、並びに記録媒体 |
| EP1273004A1 (de) * | 2000-04-06 | 2003-01-08 | One Voice Technologies Inc. | System zum verarbeiten eines natürlichen sprach-dialog-systems |
| US6912498B2 (en) * | 2000-05-02 | 2005-06-28 | Scansoft, Inc. | Error correction in speech recognition by correcting text around selected area |
| US7464033B2 (en) * | 2000-07-31 | 2008-12-09 | Texas Instruments Incorporated | Decoding multiple HMM sets using a single sentence grammar |
| JP2002116796A (ja) * | 2000-10-11 | 2002-04-19 | Canon Inc | 音声処理装置、音声処理方法及び記憶媒体 |
| US20020107695A1 (en) * | 2001-02-08 | 2002-08-08 | Roth Daniel L. | Feedback for unrecognized speech |
| US6964020B1 (en) * | 2001-02-22 | 2005-11-08 | Sprint Communications Company L.P. | Method and system for facilitating construction of a canned message in a microbrowser environment |
| BR0207642A (pt) * | 2001-02-28 | 2004-06-01 | Voice Insight | Sistema de consulta de linguagem natural para acessar um sistema de informação |
| US7072837B2 (en) | 2001-03-16 | 2006-07-04 | International Business Machines Corporation | Method for processing initially recognized speech in a speech recognition session |
| FR2832524A1 (fr) * | 2001-11-22 | 2003-05-23 | Cegetel Groupe | Procede de gestion d'un document principal |
| US6898567B2 (en) * | 2001-12-29 | 2005-05-24 | Motorola, Inc. | Method and apparatus for multi-level distributed speech recognition |
| US7177814B2 (en) * | 2002-02-07 | 2007-02-13 | Sap Aktiengesellschaft | Dynamic grammar for voice-enabled applications |
| US7016849B2 (en) * | 2002-03-25 | 2006-03-21 | Sri International | Method and apparatus for providing speech-driven routing between spoken language applications |
| US7184957B2 (en) * | 2002-09-25 | 2007-02-27 | Toyota Infotechnology Center Co., Ltd. | Multiple pass speech recognition method and system |
| US20040158468A1 (en) * | 2003-02-12 | 2004-08-12 | Aurilab, Llc | Speech recognition with soft pruning |
| EP1599867B1 (de) * | 2003-03-01 | 2008-02-13 | Robert E. Coifman | Verbesserung der transkriptionsgenauigkeit von spracherkennungssoftware |
| US7603267B2 (en) * | 2003-05-01 | 2009-10-13 | Microsoft Corporation | Rules-based grammar for slots and statistical model for preterminals in natural language understanding system |
| US7647645B2 (en) * | 2003-07-23 | 2010-01-12 | Omon Ayodele Edeki | System and method for securing computer system against unauthorized access |
| US7852993B2 (en) * | 2003-08-11 | 2010-12-14 | Microsoft Corporation | Speech recognition enhanced caller identification |
| EP1766940A4 (de) * | 2004-06-04 | 2012-04-11 | Systems Ltd Keyless | System zur verbesserung der dateneingabe in einer mobilen und festen umgebung |
| JP4574390B2 (ja) * | 2005-02-22 | 2010-11-04 | キヤノン株式会社 | 音声認識方法 |
| DE102005030967B4 (de) * | 2005-06-30 | 2007-08-09 | Daimlerchrysler Ag | Verfahren und Vorrichtung zur Interaktion mit einem Spracherkennungssystem zur Auswahl von Elementen aus Listen |
| JP2007057844A (ja) * | 2005-08-24 | 2007-03-08 | Fujitsu Ltd | 音声認識システムおよび音声処理システム |
| US8688451B2 (en) * | 2006-05-11 | 2014-04-01 | General Motors Llc | Distinguishing out-of-vocabulary speech from in-vocabulary speech |
-
2006
- 2006-06-27 DE DE102006029755A patent/DE102006029755A1/de not_active Ceased
-
2007
- 2007-06-14 BR BRPI0713987-0A patent/BRPI0713987A2/pt not_active Application Discontinuation
- 2007-06-14 RU RU2009102507/09A patent/RU2432623C2/ru active
- 2007-06-14 WO PCT/EP2007/005224 patent/WO2008000353A1/de not_active Ceased
- 2007-06-14 US US12/306,350 patent/US9208787B2/en not_active Expired - Fee Related
- 2007-06-14 CA CA2656114A patent/CA2656114C/en not_active Expired - Fee Related
- 2007-06-14 EP EP07764643A patent/EP2036078A1/de not_active Withdrawn
- 2007-06-14 CN CN2007800246599A patent/CN101484934B/zh not_active Expired - Fee Related
- 2007-06-14 KR KR1020097001732A patent/KR20090033459A/ko not_active Ceased
Non-Patent Citations (1)
| Title |
|---|
| See references of WO2008000353A1 * |
Also Published As
| Publication number | Publication date |
|---|---|
| CA2656114C (en) | 2016-02-09 |
| KR20090033459A (ko) | 2009-04-03 |
| CN101484934B (zh) | 2013-01-02 |
| DE102006029755A1 (de) | 2008-01-03 |
| US20100114577A1 (en) | 2010-05-06 |
| RU2432623C2 (ru) | 2011-10-27 |
| CN101484934A (zh) | 2009-07-15 |
| CA2656114A1 (en) | 2008-01-03 |
| WO2008000353A1 (de) | 2008-01-03 |
| US9208787B2 (en) | 2015-12-08 |
| RU2009102507A (ru) | 2010-08-10 |
| BRPI0713987A2 (pt) | 2012-11-20 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| EP2036078A1 (de) | Verfahren und vorrichtung zur natürlichsprachlichen erkennung einer sprachäusserung | |
| DE69831991T2 (de) | Verfahren und Vorrichtung zur Sprachdetektion | |
| DE60111329T2 (de) | Anpassung des phonetischen Kontextes zur Verbesserung der Spracherkennung | |
| DE69031284T2 (de) | Verfahren und Einrichtung zur Spracherkennung | |
| DE3337353C2 (de) | Sprachanalysator auf der Grundlage eines verborgenen Markov-Modells | |
| DE3236832C2 (de) | Verfahren und Gerät zur Sprachanalyse | |
| DE69814104T2 (de) | Aufteilung von texten und identifizierung von themen | |
| DE3236834C2 (de) | Verfahren und Gerät zur Sprachanalyse | |
| EP1927980B1 (de) | Verfahren zur Klassifizierung der gesprochenen Sprache in Sprachdialogsystemen | |
| DE69933623T2 (de) | Spracherkennung | |
| EP0987683B1 (de) | Spracherkennungsverfahren mit Konfidenzmassbewertung | |
| EP0994461A2 (de) | Verfahren zur automatischen Erkennung einer buchstabierten sprachlichen Äusserung | |
| DE19942178C1 (de) | Verfahren zum Aufbereiten einer Datenbank für die automatische Sprachverarbeitung | |
| EP1273003B1 (de) | Verfahren und vorrichtung zum bestimmen prosodischer markierungen | |
| EP1340222A1 (de) | Verfahren und system zur multilingualen spracherkennung | |
| EP1466317A1 (de) | Betriebsverfahren eines automatischen spracherkenners zur sprecherunabhängigen spracherkennung von worten aus verschiedenen sprachen und automatischer spracherkenner | |
| DE69519229T2 (de) | Verfahren und vorrichtung zur anpassung eines spracherkenners an dialektische sprachvarianten | |
| EP0836175A2 (de) | Verfahren und Anordnung zum Ableiten wenigstens einer Folge von Wörtern aus einem Sprachsignal | |
| DE102013101871A1 (de) | Wortwahlbasierte Sprachanalyse und Sprachanalyseeinrichtung | |
| EP2034472B1 (de) | Spracherkennungsverfahren und Spracherkennungsvorrichtung | |
| EP1125278B1 (de) | Datenverarbeitungssystem oder kommunikationsendgerät mit einer einrichtung zur erkennung gesprochener sprache und verfahren zur erkennung bestimmter akustischer objekte | |
| DE69331035T2 (de) | Zeichenerkennungssystem | |
| DE60217313T2 (de) | Verfahren zur durchführung der spracherkennung dynamischer äusserungen | |
| DE69621674T2 (de) | Trainingssystem für Referenzmuster und dieses Trainingssystem benutzendes Spracherkennungssystem | |
| EP1457966A1 (de) | Verfahren zum Ermitteln der Verwechslungsgefahr von Vokabulareinträgen bei der phonembasierten Spracherkennung |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| 17P | Request for examination filed |
Effective date: 20090112 |
|
| AK | Designated contracting states |
Kind code of ref document: A1 Designated state(s): AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HU IE IS IT LI LT LU LV MC MT NL PL PT RO SE SI SK TR |
|
| AX | Request for extension of the european patent |
Extension state: AL BA HR MK RS |
|
| RIN1 | Information on inventor provided before grant (corrected) |
Inventor name: MARKEFKA, GUNTBERT Inventor name: LIEDTKE, KLAUS-DIETER Inventor name: HAYN, EKKEHARD |
|
| 17Q | First examination report despatched |
Effective date: 20110926 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE APPLICATION IS DEEMED TO BE WITHDRAWN |
|
| 18D | Application deemed to be withdrawn |
Effective date: 20120207 |
|
| DAX | Request for extension of the european patent (deleted) |