JP4709663B2 - ユーザ適応型の音声認識方法及び音声認識装置 - Google Patents
ユーザ適応型の音声認識方法及び音声認識装置 Download PDFInfo
- Publication number
- JP4709663B2 JP4709663B2 JP2006060671A JP2006060671A JP4709663B2 JP 4709663 B2 JP4709663 B2 JP 4709663B2 JP 2006060671 A JP2006060671 A JP 2006060671A JP 2006060671 A JP2006060671 A JP 2006060671A JP 4709663 B2 JP4709663 B2 JP 4709663B2
- Authority
- JP
- Japan
- Prior art keywords
- reliability
- recognition
- user
- group
- candidate
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Active
Links
- 238000000034 method Methods 0.000 title claims description 29
- 230000003044 adaptive effect Effects 0.000 title claims description 5
- 238000004364 calculation method Methods 0.000 claims description 37
- 238000012790 confirmation Methods 0.000 claims description 24
- 230000004044 response Effects 0.000 claims description 17
- 238000010586 diagram Methods 0.000 description 9
- 230000008569 process Effects 0.000 description 8
- 238000000605 extraction Methods 0.000 description 7
- 230000015572 biosynthetic process Effects 0.000 description 4
- 238000003786 synthesis reaction Methods 0.000 description 4
- 230000006870 function Effects 0.000 description 3
- 238000004422 calculation algorithm Methods 0.000 description 2
- 239000000284 extract Substances 0.000 description 2
- 230000005484 gravity Effects 0.000 description 2
- 230000005236 sound signal Effects 0.000 description 2
- 230000006835 compression Effects 0.000 description 1
- 238000007906 compression Methods 0.000 description 1
- 238000007796 conventional method Methods 0.000 description 1
- 230000000694 effects Effects 0.000 description 1
- 238000003909 pattern recognition Methods 0.000 description 1
- 238000011946 reduction process Methods 0.000 description 1
Images
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/22—Procedures used during a speech recognition process, e.g. man-machine dialogue
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/08—Speech classification or search
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/06—Creation of reference templates; Training of speech recognition systems, e.g. adaptation to the characteristics of the speaker's voice
- G10L15/063—Training
- G10L2015/0631—Creating reference templates; Clustering
Landscapes
- Engineering & Computer Science (AREA)
- Computational Linguistics (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Human Computer Interaction (AREA)
- Physics & Mathematics (AREA)
- Acoustics & Sound (AREA)
- Multimedia (AREA)
- Navigation (AREA)
- Machine Translation (AREA)
Description
120 認識部
130 信頼度計算部
140 閾値設定部
310 判断部
320 分類部
330 閾値計算部
340 保存部
Claims (12)
- ユーザから入力された音声の認識結果による認識候補の信頼度を計算するステップと、
前記計算した認識候補の信頼度が閾値以上であれば、前記認識候補を音声認識結果として出力し、前記計算した認識候補の信頼度が閾値未満であれば、該認識候補をユーザに提供してユーザから該認識候補についての確認応答を受けるステップと、
所定回数以上の音声入力に対してデータが得られたのち、個々の音声入力に対する前記認識候補についてユーザに確認した結果と前記計算した認識候補の信頼度とを利用して、ユーザに適応した新たな閾値を設定するステップと、を含むことを特徴とするユーザ適応型の音声認識方法。 - 前記新たな閾値を設定するステップは、
前記認識候補についてユーザに確認した結果、ユーザが正解であると応答した認識候補を第1グループに分類し、ユーザが不正解であると応答した認識候補を第2グループに分類するステップと、
前記第1グループに分類された認識候補の信頼度が分布する第1信頼度区間と、前記第2グループに分類された認識候補の信頼度が分布する第2信頼度区間とが重畳しない場合、前記第2グループに分類された認識候補の信頼度のうち最も高い信頼度以上であり、前記第1グループに分類された認識候補の信頼度のうち最も低い信頼度以下である範囲内の値を持つように前記新たな閾値を計算するステップと、を含むことを特徴とする請求項1に記載のユーザ適応型の音声認識方法。 - 前記新たな閾値は、前記第1グループに分類された認識候補の信頼度のうち最も低い信頼度と、前記第2グループに分類された認識候補の信頼度のうち最も高い信頼度との平均値であることを特徴とする請求項2に記載のユーザ適応型の音声認識方法。
- 前記第1信頼度区間と前記第2信頼度区間とが重畳する場合、前記第1グループに分類された認識候補の信頼度のうち最も低い信頼度以上であり、前記第2グループに分類された認識候補の信頼度のうち最も高い信頼度以下である範囲内の値を持つように、前記新たな閾値を計算するステップをさらに含むことを特徴とする請求項2に記載のユーザ適応型の音声認識方法。
- 前記新たな閾値は、所定の信頼度範囲以内に含まれ、前記信頼度範囲は、前記第1グループに分類された認識候補のうち前記信頼度範囲の下限値未満の信頼度を持つ認識候補の数と、前記第2グループに分類された認識候補のうち前記信頼度範囲の上限値以上の信頼度を持つ認識候補の数との割合を所定の割合に最も近い割合にする範囲で計算されることを特徴とする請求項4に記載のユーザ適応型の音声認識方法。
- 前記新たな閾値は、前記信頼度範囲の上限値以上の信頼度を持つ認識候補の信頼度のうち最も低い信頼度と、前記信頼度範囲の下限値以下の信頼度を持つ認識候補の信頼度のうち最も高い信頼度との平均値であることを特徴とする請求項5に記載のユーザ適応型の音声認識方法。
- ユーザから入力された音声の認識結果による認識候補の信頼度を計算する信頼度計算部と、
前記計算した認識候補の信頼度が閾値以上であれば、前記認識候補を音声認識結果として出力し、前記計算した認識候補の信頼度が閾値未満であれば、該認識候補をユーザに提供してユーザから該認識候補についての確認応答を受ける制御部と、
所定回数以上の音声入力に対してデータが得られたのち、個々の音声入力に対する前記認識候補についてユーザに確認した結果と前記計算した認識候補の信頼度とを利用して、ユーザに適応した新たな閾値を設定する閾値設定部と、
を備えることを特徴とするユーザ適応型の音声認識装置。 - 前記閾値設定部は、
前記認識候補についてユーザに確認した結果、ユーザが正解であると応答した認識候補を第1グループに分類し、ユーザが不正解であると応答した認識候補を第2グループに分類する分類部と、
前記第1グループに分類された認識候補の信頼度が分布する第1信頼度区間と、前記第2グループに分類された認識候補の信頼度が分布する第2信頼度区間とが重畳しない場合、前記第2グループに分類された認識候補の信頼度のうち最も高い信頼度以上であり、前記第1グループに分類された認識候補の信頼度のうち最も低い信頼度以下である範囲内の値を持つように前記新たな閾値を計算する閾値計算部と、を備えることを特徴とする請求項7に記載のユーザ適応型の音声認識装置。 - 前記新たな閾値は、前記第1グループに分類された認識候補の信頼度のうち最も低い信頼度と、前記第2グループに分類された認識候補の信頼度のうち最も高い信頼度との平均値であることを特徴とする請求項8に記載のユーザ適応型の音声認識装置。
- 前記閾値計算部は、前記第1信頼度区間と前記第2信頼度区間とが重畳する場合、前記第1グループに分類された認識候補の信頼度のうち最も低い信頼度以上であり、前記第2グループに分類された認識候補の信頼度のうち最も高い信頼度以下である範囲内の値を持つように、前記新たな閾値を計算するステップをさらに含むことを特徴とする請求項8に記載のユーザ適応型の音声認識装置。
- 前記新たな閾値は、所定の信頼度範囲以内に含まれ、前記信頼度範囲は、前記第1グループに分類された認識候補のうち前記信頼度範囲の下限値未満の信頼度を持つ認識候補の数と、前記第2グループに分類された認識候補のうち前記信頼度範囲の上限値以上の信頼度を持つ認識候補の数との割合を所定の割合に最も近い割合にする範囲で計算されることを特徴とする請求項10に記載のユーザ適応型の音声認識装置。
- 前記新たな閾値は、前記信頼度範囲の上限値以上の信頼度を持つ認識候補の信頼度のうち最も低い信頼度と、前記信頼度範囲の下限値以下の信頼度を持つ認識候補の信頼度のうち最も高い信頼度との平均値であることを特徴とする請求項11に記載のユーザ適応型の音声認識装置。
Applications Claiming Priority (2)
Application Number | Priority Date | Filing Date | Title |
---|---|---|---|
KR1020050018786A KR100679044B1 (ko) | 2005-03-07 | 2005-03-07 | 사용자 적응형 음성 인식 방법 및 장치 |
KR10-2005-0018786 | 2005-03-07 |
Publications (2)
Publication Number | Publication Date |
---|---|
JP2006251800A JP2006251800A (ja) | 2006-09-21 |
JP4709663B2 true JP4709663B2 (ja) | 2011-06-22 |
Family
ID=36945180
Family Applications (1)
Application Number | Title | Priority Date | Filing Date |
---|---|---|---|
JP2006060671A Active JP4709663B2 (ja) | 2005-03-07 | 2006-03-07 | ユーザ適応型の音声認識方法及び音声認識装置 |
Country Status (3)
Country | Link |
---|---|
US (1) | US7996218B2 (ja) |
JP (1) | JP4709663B2 (ja) |
KR (1) | KR100679044B1 (ja) |
Families Citing this family (188)
Publication number | Priority date | Publication date | Assignee | Title |
---|---|---|---|---|
US8645137B2 (en) | 2000-03-16 | 2014-02-04 | Apple Inc. | Fast, language-independent method for user authentication by voice |
US9530050B1 (en) | 2007-07-11 | 2016-12-27 | Ricoh Co., Ltd. | Document annotation sharing |
US8965145B2 (en) | 2006-07-31 | 2015-02-24 | Ricoh Co., Ltd. | Mixed media reality recognition using multiple specialized indexes |
US8856108B2 (en) * | 2006-07-31 | 2014-10-07 | Ricoh Co., Ltd. | Combining results of image retrieval processes |
US9171202B2 (en) | 2005-08-23 | 2015-10-27 | Ricoh Co., Ltd. | Data organization and access for mixed media document system |
US9405751B2 (en) | 2005-08-23 | 2016-08-02 | Ricoh Co., Ltd. | Database for mixed media document system |
US8825682B2 (en) * | 2006-07-31 | 2014-09-02 | Ricoh Co., Ltd. | Architecture for mixed media reality retrieval of locations and registration of images |
US8868555B2 (en) | 2006-07-31 | 2014-10-21 | Ricoh Co., Ltd. | Computation of a recongnizability score (quality predictor) for image retrieval |
US7812986B2 (en) | 2005-08-23 | 2010-10-12 | Ricoh Co. Ltd. | System and methods for use of voice mail and email in a mixed media environment |
US7702673B2 (en) | 2004-10-01 | 2010-04-20 | Ricoh Co., Ltd. | System and methods for creation and use of a mixed media environment |
US9373029B2 (en) | 2007-07-11 | 2016-06-21 | Ricoh Co., Ltd. | Invisible junction feature recognition for document security or annotation |
US8176054B2 (en) | 2007-07-12 | 2012-05-08 | Ricoh Co. Ltd | Retrieving electronic documents by converting them to synthetic text |
US10192279B1 (en) | 2007-07-11 | 2019-01-29 | Ricoh Co., Ltd. | Indexed document modification sharing with mixed media reality |
US8838591B2 (en) | 2005-08-23 | 2014-09-16 | Ricoh Co., Ltd. | Embedding hot spots in electronic documents |
US8949287B2 (en) | 2005-08-23 | 2015-02-03 | Ricoh Co., Ltd. | Embedding hot spots in imaged documents |
US8156116B2 (en) | 2006-07-31 | 2012-04-10 | Ricoh Co., Ltd | Dynamic presentation of targeted information in a mixed media reality recognition system |
US9384619B2 (en) | 2006-07-31 | 2016-07-05 | Ricoh Co., Ltd. | Searching media content for objects specified using identifiers |
US8732025B2 (en) * | 2005-05-09 | 2014-05-20 | Google Inc. | System and method for enabling image recognition and searching of remote content on display |
US7783135B2 (en) * | 2005-05-09 | 2010-08-24 | Like.Com | System and method for providing objectified image renderings using recognition information from images |
US7760917B2 (en) | 2005-05-09 | 2010-07-20 | Like.Com | Computer-implemented method for performing similarity searches |
US7657126B2 (en) | 2005-05-09 | 2010-02-02 | Like.Com | System and method for search portions of objects in images and features thereof |
US7660468B2 (en) | 2005-05-09 | 2010-02-09 | Like.Com | System and method for enabling image searching using manual enrichment, classification, and/or segmentation |
US7519200B2 (en) | 2005-05-09 | 2009-04-14 | Like.Com | System and method for enabling the use of captured images through recognition |
US7945099B2 (en) | 2005-05-09 | 2011-05-17 | Like.Com | System and method for use of images with recognition analysis |
US20080177640A1 (en) | 2005-05-09 | 2008-07-24 | Salih Burak Gokturk | System and method for using image analysis and search in e-commerce |
US8677377B2 (en) | 2005-09-08 | 2014-03-18 | Apple Inc. | Method and apparatus for building an intelligent automated assistant |
US8571272B2 (en) * | 2006-03-12 | 2013-10-29 | Google Inc. | Techniques for enabling or establishing the use of face recognition algorithms |
US9690979B2 (en) | 2006-03-12 | 2017-06-27 | Google Inc. | Techniques for enabling or establishing the use of face recognition algorithms |
US9020966B2 (en) | 2006-07-31 | 2015-04-28 | Ricoh Co., Ltd. | Client device for interacting with a mixed media reality recognition system |
US8201076B2 (en) | 2006-07-31 | 2012-06-12 | Ricoh Co., Ltd. | Capturing symbolic information from documents upon printing |
US9176984B2 (en) | 2006-07-31 | 2015-11-03 | Ricoh Co., Ltd | Mixed media reality retrieval of differentially-weighted links |
US8489987B2 (en) | 2006-07-31 | 2013-07-16 | Ricoh Co., Ltd. | Monitoring and analyzing creation and usage of visual content using image and hotspot interaction |
US9063952B2 (en) | 2006-07-31 | 2015-06-23 | Ricoh Co., Ltd. | Mixed media reality recognition with image tracking |
US9318108B2 (en) | 2010-01-18 | 2016-04-19 | Apple Inc. | Intelligent automated assistant |
US8977255B2 (en) | 2007-04-03 | 2015-03-10 | Apple Inc. | Method and system for operating a multi-function portable electronic device using voice-activation |
US9423996B2 (en) * | 2007-05-03 | 2016-08-23 | Ian Cummings | Vehicle navigation user interface customization methods |
US8403919B2 (en) * | 2007-06-05 | 2013-03-26 | Alcon Refractivehorizons, Inc. | Nomogram computation and application system and method for refractive laser surgery |
JP4973731B2 (ja) * | 2007-07-09 | 2012-07-11 | 富士通株式会社 | 音声認識装置、音声認識方法、および、音声認識プログラム |
US8416981B2 (en) | 2007-07-29 | 2013-04-09 | Google Inc. | System and method for displaying contextual supplemental content based on image content |
KR100933946B1 (ko) * | 2007-10-29 | 2009-12-28 | 연세대학교 산학협력단 | 음성 분석구간 중첩길이의 가변적 선택을 이용한 특징 벡터추출 방법 및 이를 이용한 화자 인식 시스템 |
US9330720B2 (en) | 2008-01-03 | 2016-05-03 | Apple Inc. | Methods and apparatus for altering audio output signals |
US8996376B2 (en) | 2008-04-05 | 2015-03-31 | Apple Inc. | Intelligent text-to-speech conversion |
US8239203B2 (en) * | 2008-04-15 | 2012-08-07 | Nuance Communications, Inc. | Adaptive confidence thresholds for speech recognition |
US10496753B2 (en) | 2010-01-18 | 2019-12-03 | Apple Inc. | Automatically adapting user interfaces for hands-free interaction |
KR101056511B1 (ko) | 2008-05-28 | 2011-08-11 | (주)파워보이스 | 실시간 호출명령어 인식을 이용한 잡음환경에서의음성구간검출과 연속음성인식 시스템 |
JP2010008601A (ja) * | 2008-06-25 | 2010-01-14 | Fujitsu Ltd | 案内情報表示装置、案内情報表示方法及びプログラム |
AU2009270946A1 (en) * | 2008-07-14 | 2010-01-21 | Google Inc. | System and method for using supplemental content items for search criteria for identifying other content items of interest |
US20100030549A1 (en) | 2008-07-31 | 2010-02-04 | Lee Michael M | Mobile device having human language translation capability with positional feedback |
WO2010067118A1 (en) * | 2008-12-11 | 2010-06-17 | Novauris Technologies Limited | Speech recognition involving a mobile device |
KR101217524B1 (ko) * | 2008-12-22 | 2013-01-18 | 한국전자통신연구원 | 고립어 엔베스트 인식결과를 위한 발화검증 방법 및 장치 |
US20100180127A1 (en) * | 2009-01-14 | 2010-07-15 | Motorola, Inc. | Biometric authentication based upon usage history |
US10088976B2 (en) | 2009-01-15 | 2018-10-02 | Em Acquisition Corp., Inc. | Systems and methods for multiple voice document narration |
US8359202B2 (en) * | 2009-01-15 | 2013-01-22 | K-Nfb Reading Technology, Inc. | Character models for document narration |
US8370151B2 (en) | 2009-01-15 | 2013-02-05 | K-Nfb Reading Technology, Inc. | Systems and methods for multiple voice document narration |
US20100313141A1 (en) * | 2009-06-03 | 2010-12-09 | Tianli Yu | System and Method for Learning User Genres and Styles and for Matching Products to User Preferences |
US10241644B2 (en) | 2011-06-03 | 2019-03-26 | Apple Inc. | Actionable reminder entries |
US9858925B2 (en) | 2009-06-05 | 2018-01-02 | Apple Inc. | Using context information to facilitate processing of commands in a virtual assistant |
US10241752B2 (en) | 2011-09-30 | 2019-03-26 | Apple Inc. | Interface for a virtual digital assistant |
US10706373B2 (en) | 2011-06-03 | 2020-07-07 | Apple Inc. | Performing actions associated with task items that represent tasks to perform |
US9431006B2 (en) | 2009-07-02 | 2016-08-30 | Apple Inc. | Methods and apparatuses for automatic speech recognition |
JP4951035B2 (ja) * | 2009-07-08 | 2012-06-13 | 日本電信電話株式会社 | 音声単位別尤度比モデル作成装置、音声単位別尤度比モデル作成方法、音声認識信頼度算出装置、音声認識信頼度算出方法、プログラム |
TWI421857B (zh) * | 2009-12-29 | 2014-01-01 | Ind Tech Res Inst | 產生詞語確認臨界值的裝置、方法與語音辨識、詞語確認系統 |
US10276170B2 (en) | 2010-01-18 | 2019-04-30 | Apple Inc. | Intelligent automated assistant |
US10679605B2 (en) | 2010-01-18 | 2020-06-09 | Apple Inc. | Hands-free list-reading by intelligent automated assistant |
US10553209B2 (en) | 2010-01-18 | 2020-02-04 | Apple Inc. | Systems and methods for hands-free notification summaries |
US10705794B2 (en) | 2010-01-18 | 2020-07-07 | Apple Inc. | Automatically adapting user interfaces for hands-free interaction |
US8682667B2 (en) | 2010-02-25 | 2014-03-25 | Apple Inc. | User profiling for selecting user specific voice input processing information |
JP5533042B2 (ja) * | 2010-03-04 | 2014-06-25 | 富士通株式会社 | 音声検索装置、音声検索方法、プログラム及び記録媒体 |
TWI459828B (zh) * | 2010-03-08 | 2014-11-01 | Dolby Lab Licensing Corp | 在多頻道音訊中決定語音相關頻道的音量降低比例的方法及系統 |
US8392186B2 (en) | 2010-05-18 | 2013-03-05 | K-Nfb Reading Technology, Inc. | Audio synchronization for document narration with user-selected playback |
US9263034B1 (en) * | 2010-07-13 | 2016-02-16 | Google Inc. | Adapting enhanced acoustic models |
US8639508B2 (en) * | 2011-02-14 | 2014-01-28 | General Motors Llc | User-specific confidence thresholds for speech recognition |
US9262612B2 (en) | 2011-03-21 | 2016-02-16 | Apple Inc. | Device access using voice authentication |
US10057736B2 (en) | 2011-06-03 | 2018-08-21 | Apple Inc. | Active transport based notifications |
US9058331B2 (en) | 2011-07-27 | 2015-06-16 | Ricoh Co., Ltd. | Generating a conversation in a social network based on visual search results |
US8994660B2 (en) | 2011-08-29 | 2015-03-31 | Apple Inc. | Text correction processing |
US10134385B2 (en) | 2012-03-02 | 2018-11-20 | Apple Inc. | Systems and methods for name pronunciation |
US9483461B2 (en) | 2012-03-06 | 2016-11-01 | Apple Inc. | Handling speech synthesis of content for multiple languages |
US9280610B2 (en) | 2012-05-14 | 2016-03-08 | Apple Inc. | Crowd sourcing information to fulfill user requests |
KR20130133629A (ko) | 2012-05-29 | 2013-12-09 | 삼성전자주식회사 | 전자장치에서 음성명령을 실행시키기 위한 장치 및 방법 |
US9721563B2 (en) | 2012-06-08 | 2017-08-01 | Apple Inc. | Name recognition system |
US9495129B2 (en) | 2012-06-29 | 2016-11-15 | Apple Inc. | Device, method, and user interface for voice-activated navigation and browsing of a document |
US9576574B2 (en) | 2012-09-10 | 2017-02-21 | Apple Inc. | Context-sensitive handling of interruptions by intelligent digital assistant |
US9547647B2 (en) | 2012-09-19 | 2017-01-17 | Apple Inc. | Voice-based media searching |
CN103076893B (zh) * | 2012-12-31 | 2016-08-17 | 百度在线网络技术(北京)有限公司 | 一种用于实现语音输入的方法与设备 |
US9368114B2 (en) | 2013-03-14 | 2016-06-14 | Apple Inc. | Context-sensitive handling of interruptions |
WO2014144579A1 (en) | 2013-03-15 | 2014-09-18 | Apple Inc. | System and method for updating an adaptive speech recognition model |
WO2014144949A2 (en) | 2013-03-15 | 2014-09-18 | Apple Inc. | Training an at least partial voice command system |
WO2014197336A1 (en) | 2013-06-07 | 2014-12-11 | Apple Inc. | System and method for detecting errors in interactions with a voice-based digital assistant |
WO2014197334A2 (en) | 2013-06-07 | 2014-12-11 | Apple Inc. | System and method for user-specified pronunciation of words for speech synthesis and recognition |
US9582608B2 (en) | 2013-06-07 | 2017-02-28 | Apple Inc. | Unified ranking with entropy-weighted information for phrase-based semantic auto-completion |
WO2014197335A1 (en) | 2013-06-08 | 2014-12-11 | Apple Inc. | Interpreting and acting upon commands that involve sharing information with remote devices |
KR101772152B1 (ko) | 2013-06-09 | 2017-08-28 | 애플 인크. | 디지털 어시스턴트의 둘 이상의 인스턴스들에 걸친 대화 지속성을 가능하게 하기 위한 디바이스, 방법 및 그래픽 사용자 인터페이스 |
US10176167B2 (en) | 2013-06-09 | 2019-01-08 | Apple Inc. | System and method for inferring user intent from speech inputs |
CN105265005B (zh) | 2013-06-13 | 2019-09-17 | 苹果公司 | 用于由语音命令发起的紧急呼叫的系统和方法 |
US9899021B1 (en) * | 2013-12-20 | 2018-02-20 | Amazon Technologies, Inc. | Stochastic modeling of user interactions with a detection system |
US10540979B2 (en) * | 2014-04-17 | 2020-01-21 | Qualcomm Incorporated | User interface for secure access to a device using speaker verification |
CN104142909B (zh) * | 2014-05-07 | 2016-04-27 | 腾讯科技(深圳)有限公司 | 一种汉字注音方法及装置 |
US9620105B2 (en) | 2014-05-15 | 2017-04-11 | Apple Inc. | Analyzing audio input for efficient speech and music recognition |
US10592095B2 (en) | 2014-05-23 | 2020-03-17 | Apple Inc. | Instantaneous speaking of content on touch devices |
US9502031B2 (en) | 2014-05-27 | 2016-11-22 | Apple Inc. | Method for supporting dynamic grammars in WFST-based ASR |
US9715875B2 (en) | 2014-05-30 | 2017-07-25 | Apple Inc. | Reducing the need for manual start/end-pointing and trigger phrases |
US10078631B2 (en) | 2014-05-30 | 2018-09-18 | Apple Inc. | Entropy-guided text prediction using combined word and character n-gram language models |
US9760559B2 (en) | 2014-05-30 | 2017-09-12 | Apple Inc. | Predictive text input |
US9842101B2 (en) | 2014-05-30 | 2017-12-12 | Apple Inc. | Predictive conversion of language input |
US10289433B2 (en) | 2014-05-30 | 2019-05-14 | Apple Inc. | Domain specific language for encoding assistant dialog |
US10170123B2 (en) | 2014-05-30 | 2019-01-01 | Apple Inc. | Intelligent assistant for home automation |
WO2015184186A1 (en) | 2014-05-30 | 2015-12-03 | Apple Inc. | Multi-command single utterance input method |
US9430463B2 (en) | 2014-05-30 | 2016-08-30 | Apple Inc. | Exemplar-based natural language processing |
US9734193B2 (en) | 2014-05-30 | 2017-08-15 | Apple Inc. | Determining domain salience ranking from ambiguous words in natural speech |
US9633004B2 (en) | 2014-05-30 | 2017-04-25 | Apple Inc. | Better resolution when referencing to concepts |
US9785630B2 (en) | 2014-05-30 | 2017-10-10 | Apple Inc. | Text prediction using combined word N-gram and unigram language models |
US9384738B2 (en) * | 2014-06-24 | 2016-07-05 | Google Inc. | Dynamic threshold for speaker verification |
US9338493B2 (en) | 2014-06-30 | 2016-05-10 | Apple Inc. | Intelligent automated assistant for TV user interactions |
US10659851B2 (en) | 2014-06-30 | 2020-05-19 | Apple Inc. | Real-time digital assistant knowledge updates |
US20160063990A1 (en) * | 2014-08-26 | 2016-03-03 | Honeywell International Inc. | Methods and apparatus for interpreting clipped speech using speech recognition |
KR102357321B1 (ko) | 2014-08-27 | 2022-02-03 | 삼성전자주식회사 | 음성 인식이 가능한 디스플레이 장치 및 방법 |
US10446141B2 (en) | 2014-08-28 | 2019-10-15 | Apple Inc. | Automatic speech recognition based on user feedback |
US9818400B2 (en) | 2014-09-11 | 2017-11-14 | Apple Inc. | Method and apparatus for discovering trending terms in speech requests |
US10789041B2 (en) | 2014-09-12 | 2020-09-29 | Apple Inc. | Dynamic thresholds for always listening speech trigger |
US10074360B2 (en) | 2014-09-30 | 2018-09-11 | Apple Inc. | Providing an indication of the suitability of speech recognition |
US9668121B2 (en) | 2014-09-30 | 2017-05-30 | Apple Inc. | Social reminders |
US9646609B2 (en) | 2014-09-30 | 2017-05-09 | Apple Inc. | Caching apparatus for serving phonetic pronunciations |
US9886432B2 (en) | 2014-09-30 | 2018-02-06 | Apple Inc. | Parsimonious handling of word inflection via categorical stem + suffix N-gram language models |
US10127911B2 (en) | 2014-09-30 | 2018-11-13 | Apple Inc. | Speaker identification and unsupervised speaker adaptation techniques |
US9953644B2 (en) | 2014-12-01 | 2018-04-24 | At&T Intellectual Property I, L.P. | Targeted clarification questions in speech recognition with concept presence score and concept correctness score |
US10552013B2 (en) | 2014-12-02 | 2020-02-04 | Apple Inc. | Data detection |
US9711141B2 (en) | 2014-12-09 | 2017-07-18 | Apple Inc. | Disambiguating heteronyms in speech synthesis |
US9865280B2 (en) | 2015-03-06 | 2018-01-09 | Apple Inc. | Structured dictation using intelligent automated assistants |
US9886953B2 (en) | 2015-03-08 | 2018-02-06 | Apple Inc. | Virtual assistant activation |
US9721566B2 (en) | 2015-03-08 | 2017-08-01 | Apple Inc. | Competing devices responding to voice triggers |
US10567477B2 (en) | 2015-03-08 | 2020-02-18 | Apple Inc. | Virtual assistant continuity |
US9899019B2 (en) | 2015-03-18 | 2018-02-20 | Apple Inc. | Systems and methods for structured stem and suffix language models |
US9842105B2 (en) | 2015-04-16 | 2017-12-12 | Apple Inc. | Parsimonious continuous-space phrase representations for natural language processing |
US10083688B2 (en) | 2015-05-27 | 2018-09-25 | Apple Inc. | Device voice control for selecting a displayed affordance |
US10127220B2 (en) | 2015-06-04 | 2018-11-13 | Apple Inc. | Language identification from short strings |
US10101822B2 (en) | 2015-06-05 | 2018-10-16 | Apple Inc. | Language input correction |
US9578173B2 (en) | 2015-06-05 | 2017-02-21 | Apple Inc. | Virtual assistant aided communication with 3rd party service in a communication session |
US10255907B2 (en) | 2015-06-07 | 2019-04-09 | Apple Inc. | Automatic accent detection using acoustic models |
US10186254B2 (en) | 2015-06-07 | 2019-01-22 | Apple Inc. | Context-based endpoint detection |
US11025565B2 (en) | 2015-06-07 | 2021-06-01 | Apple Inc. | Personalized prediction of responses for instant messaging |
US10671428B2 (en) | 2015-09-08 | 2020-06-02 | Apple Inc. | Distributed personal assistant |
US10747498B2 (en) | 2015-09-08 | 2020-08-18 | Apple Inc. | Zero latency digital assistant |
US9997161B2 (en) * | 2015-09-11 | 2018-06-12 | Microsoft Technology Licensing, Llc | Automatic speech recognition confidence classifier |
US9697820B2 (en) | 2015-09-24 | 2017-07-04 | Apple Inc. | Unit-selection text-to-speech synthesis using concatenation-sensitive neural networks |
US10366158B2 (en) | 2015-09-29 | 2019-07-30 | Apple Inc. | Efficient word encoding for recurrent neural network language models |
US11010550B2 (en) | 2015-09-29 | 2021-05-18 | Apple Inc. | Unified language modeling framework for word prediction, auto-completion and auto-correction |
US11587559B2 (en) | 2015-09-30 | 2023-02-21 | Apple Inc. | Intelligent device identification |
CN106653010B (zh) * | 2015-11-03 | 2020-07-24 | 络达科技股份有限公司 | 电子装置及其透过语音辨识唤醒的方法 |
US10691473B2 (en) | 2015-11-06 | 2020-06-23 | Apple Inc. | Intelligent automated assistant in a messaging environment |
US10706852B2 (en) | 2015-11-13 | 2020-07-07 | Microsoft Technology Licensing, Llc | Confidence features for automated speech recognition arbitration |
US10049668B2 (en) | 2015-12-02 | 2018-08-14 | Apple Inc. | Applying neural network language models to weighted finite state transducers for automatic speech recognition |
CN108369451B (zh) * | 2015-12-18 | 2021-10-29 | 索尼公司 | 信息处理装置、信息处理方法及计算机可读存储介质 |
US10223066B2 (en) | 2015-12-23 | 2019-03-05 | Apple Inc. | Proactive assistance based on dialog communication between devices |
US10446143B2 (en) | 2016-03-14 | 2019-10-15 | Apple Inc. | Identification of voice inputs providing credentials |
US9934775B2 (en) | 2016-05-26 | 2018-04-03 | Apple Inc. | Unit-selection text-to-speech synthesis based on predicted concatenation parameters |
US9972304B2 (en) | 2016-06-03 | 2018-05-15 | Apple Inc. | Privacy preserving distributed evaluation framework for embedded personalized systems |
US10249300B2 (en) | 2016-06-06 | 2019-04-02 | Apple Inc. | Intelligent list reading |
US10049663B2 (en) | 2016-06-08 | 2018-08-14 | Apple, Inc. | Intelligent automated assistant for media exploration |
DK179309B1 (en) | 2016-06-09 | 2018-04-23 | Apple Inc | Intelligent automated assistant in a home environment |
US10490187B2 (en) | 2016-06-10 | 2019-11-26 | Apple Inc. | Digital assistant providing automated status report |
US10586535B2 (en) | 2016-06-10 | 2020-03-10 | Apple Inc. | Intelligent digital assistant in a multi-tasking environment |
US10192552B2 (en) | 2016-06-10 | 2019-01-29 | Apple Inc. | Digital assistant providing whispered speech |
US10067938B2 (en) | 2016-06-10 | 2018-09-04 | Apple Inc. | Multilingual word prediction |
US10509862B2 (en) | 2016-06-10 | 2019-12-17 | Apple Inc. | Dynamic phrase expansion of language input |
DK201670540A1 (en) | 2016-06-11 | 2018-01-08 | Apple Inc | Application integration with a digital assistant |
DK179415B1 (en) | 2016-06-11 | 2018-06-14 | Apple Inc | Intelligent device arbitration and control |
DK179343B1 (en) | 2016-06-11 | 2018-05-14 | Apple Inc | Intelligent task discovery |
DK179049B1 (en) | 2016-06-11 | 2017-09-18 | Apple Inc | Data driven natural language event detection and classification |
US10043516B2 (en) | 2016-09-23 | 2018-08-07 | Apple Inc. | Intelligent automated assistant |
US10169319B2 (en) * | 2016-09-27 | 2019-01-01 | International Business Machines Corporation | System, method and computer program product for improving dialog service quality via user feedback |
US10593346B2 (en) | 2016-12-22 | 2020-03-17 | Apple Inc. | Rank-reduced token representation for automatic speech recognition |
DK201770439A1 (en) | 2017-05-11 | 2018-12-13 | Apple Inc. | Offline personal assistant |
DK179745B1 (en) | 2017-05-12 | 2019-05-01 | Apple Inc. | SYNCHRONIZATION AND TASK DELEGATION OF A DIGITAL ASSISTANT |
DK179496B1 (en) | 2017-05-12 | 2019-01-15 | Apple Inc. | USER-SPECIFIC Acoustic Models |
DK201770432A1 (en) | 2017-05-15 | 2018-12-21 | Apple Inc. | Hierarchical belief states for digital assistants |
DK201770431A1 (en) | 2017-05-15 | 2018-12-20 | Apple Inc. | Optimizing dialogue policy decisions for digital assistants using implicit feedback |
DK179560B1 (en) | 2017-05-16 | 2019-02-18 | Apple Inc. | FAR-FIELD EXTENSION FOR DIGITAL ASSISTANT SERVICES |
KR102348593B1 (ko) * | 2017-10-26 | 2022-01-06 | 삼성에스디에스 주식회사 | 기계 학습 기반의 객체 검출 방법 및 그 장치 |
TWI682385B (zh) * | 2018-03-16 | 2020-01-11 | 緯創資通股份有限公司 | 語音服務控制裝置及其方法 |
JP2021529978A (ja) * | 2018-05-10 | 2021-11-04 | エル ソルー カンパニー, リミテッドLlsollu Co., Ltd. | 人工知能サービス方法及びそのための装置 |
US11087748B2 (en) * | 2018-05-11 | 2021-08-10 | Google Llc | Adaptive interface in a voice-activated network |
KR20200007496A (ko) * | 2018-07-13 | 2020-01-22 | 삼성전자주식회사 | 개인화 ASR(automatic speech recognition) 모델을 생성하는 전자 장치 및 이를 동작하는 방법 |
US11170770B2 (en) * | 2018-08-03 | 2021-11-09 | International Business Machines Corporation | Dynamic adjustment of response thresholds in a dialogue system |
CN110111775B (zh) * | 2019-05-17 | 2021-06-22 | 腾讯科技(深圳)有限公司 | 一种流式语音识别方法、装置、设备及存储介质 |
JP2022544984A (ja) | 2019-08-21 | 2022-10-24 | ドルビー ラボラトリーズ ライセンシング コーポレイション | ヒト話者の埋め込みを会話合成に適合させるためのシステムおよび方法 |
WO2021149923A1 (ko) * | 2020-01-20 | 2021-07-29 | 주식회사 씨오티커넥티드 | 영상 검색 제공 방법 및 장치 |
US11620993B2 (en) * | 2021-06-09 | 2023-04-04 | Merlyn Mind, Inc. | Multimodal intent entity resolver |
Citations (6)
Publication number | Priority date | Publication date | Assignee | Title |
---|---|---|---|---|
JPH0432900A (ja) * | 1990-05-29 | 1992-02-04 | Ricoh Co Ltd | 音声認識装置 |
JPH1185189A (ja) * | 1997-09-10 | 1999-03-30 | Hitachi Ltd | 音声認識装置 |
JP2000122689A (ja) * | 1998-10-20 | 2000-04-28 | Mitsubishi Electric Corp | 話者適応化装置及び音声認識装置 |
JP2000181482A (ja) * | 1998-12-17 | 2000-06-30 | Sony Internatl Europ Gmbh | 音声認識装置及び自動音声認識装置の非教示及び/又はオンライン適応方法 |
JP2001013991A (ja) * | 1999-06-30 | 2001-01-19 | Toshiba Corp | 音声認識支援方法及び音声認識システム |
JP2004325635A (ja) * | 2003-04-23 | 2004-11-18 | Sharp Corp | 音声処理装置、音声処理方法、音声処理プログラム、および、プログラム記録媒体 |
Family Cites Families (8)
Publication number | Priority date | Publication date | Assignee | Title |
---|---|---|---|---|
US5305244B2 (en) * | 1992-04-06 | 1997-09-23 | Computer Products & Services I | Hands-free user-supported portable computer |
US5732187A (en) | 1993-09-27 | 1998-03-24 | Texas Instruments Incorporated | Speaker-dependent speech recognition using speaker independent models |
US5559925A (en) * | 1994-06-24 | 1996-09-24 | Apple Computer, Inc. | Determining the useability of input signals in a data recognition system |
US6567778B1 (en) | 1995-12-21 | 2003-05-20 | Nuance Communications | Natural language speech recognition using slot semantic confidence scores related to their word recognition confidence scores |
KR100277105B1 (ko) | 1998-02-27 | 2001-01-15 | 윤종용 | 음성 인식 데이터 결정 장치 및 방법 |
US7103542B2 (en) * | 2001-12-14 | 2006-09-05 | Ben Franklin Patent Holding Llc | Automatically improving a voice recognition system |
EP1378886A1 (en) * | 2002-07-02 | 2004-01-07 | Ubicall Communications en abrégé "UbiCall" S.A. | Speech recognition device |
US7788103B2 (en) * | 2004-10-18 | 2010-08-31 | Nuance Communications, Inc. | Random confirmation in speech based systems |
-
2005
- 2005-03-07 KR KR1020050018786A patent/KR100679044B1/ko not_active IP Right Cessation
-
2006
- 2006-02-16 US US11/354,942 patent/US7996218B2/en active Active
- 2006-03-07 JP JP2006060671A patent/JP4709663B2/ja active Active
Patent Citations (6)
Publication number | Priority date | Publication date | Assignee | Title |
---|---|---|---|---|
JPH0432900A (ja) * | 1990-05-29 | 1992-02-04 | Ricoh Co Ltd | 音声認識装置 |
JPH1185189A (ja) * | 1997-09-10 | 1999-03-30 | Hitachi Ltd | 音声認識装置 |
JP2000122689A (ja) * | 1998-10-20 | 2000-04-28 | Mitsubishi Electric Corp | 話者適応化装置及び音声認識装置 |
JP2000181482A (ja) * | 1998-12-17 | 2000-06-30 | Sony Internatl Europ Gmbh | 音声認識装置及び自動音声認識装置の非教示及び/又はオンライン適応方法 |
JP2001013991A (ja) * | 1999-06-30 | 2001-01-19 | Toshiba Corp | 音声認識支援方法及び音声認識システム |
JP2004325635A (ja) * | 2003-04-23 | 2004-11-18 | Sharp Corp | 音声処理装置、音声処理方法、音声処理プログラム、および、プログラム記録媒体 |
Also Published As
Publication number | Publication date |
---|---|
KR100679044B1 (ko) | 2007-02-06 |
JP2006251800A (ja) | 2006-09-21 |
KR20060097895A (ko) | 2006-09-18 |
US7996218B2 (en) | 2011-08-09 |
US20060200347A1 (en) | 2006-09-07 |
Similar Documents
Publication | Publication Date | Title |
---|---|---|
JP4709663B2 (ja) | ユーザ適応型の音声認識方法及び音声認識装置 | |
US8280733B2 (en) | Automatic speech recognition learning using categorization and selective incorporation of user-initiated corrections | |
US7974843B2 (en) | Operating method for an automated language recognizer intended for the speaker-independent language recognition of words in different languages and automated language recognizer | |
US7401017B2 (en) | Adaptive multi-pass speech recognition system | |
KR100679042B1 (ko) | 음성인식 방법 및 장치, 이를 이용한 네비게이션 시스템 | |
EP2048655B1 (en) | Context sensitive multi-stage speech recognition | |
EP1269464B1 (en) | Discriminative training of hidden markov models for continuous speech recognition | |
EP1936606A1 (en) | Multi-stage speech recognition | |
US20110196678A1 (en) | Speech recognition apparatus and speech recognition method | |
JP2006038895A (ja) | 音声処理装置および音声処理方法、プログラム、並びに記録媒体 | |
JP2000122691A (ja) | 綴り字読み式音声発話の自動認識方法 | |
JP2008009153A (ja) | 音声対話システム | |
EP1734509A1 (en) | Method and system for speech recognition | |
US20150310853A1 (en) | Systems and methods for speech artifact compensation in speech recognition systems | |
JP2000099087A (ja) | 言語音声モデルを適応させる方法及び音声認識システム | |
CN108806691B (zh) | 语音识别方法及系统 | |
EP1213706B1 (en) | Method for online adaptation of pronunciation dictionaries | |
JPH09179581A (ja) | 音声認識システム | |
JP3444108B2 (ja) | 音声認識装置 | |
JP2011053312A (ja) | 適応化音響モデル生成装置及びプログラム | |
KR100622019B1 (ko) | 음성 인터페이스 시스템 및 방법 | |
KR20060098673A (ko) | 음성 인식 방법 및 장치 | |
JP2003044085A (ja) | コマンド入力機能つきディクテーション装置 | |
JPH08314490A (ja) | ワードスポッティング型音声認識方法と装置 | |
JPH08248975A (ja) | 標準パターン学習装置およびこの装置を使用した音声認識装置 |
Legal Events
Date | Code | Title | Description |
---|---|---|---|
RD02 | Notification of acceptance of power of attorney |
Free format text: JAPANESE INTERMEDIATE CODE: A7422 Effective date: 20061101 |
|
RD04 | Notification of resignation of power of attorney |
Free format text: JAPANESE INTERMEDIATE CODE: A7424 Effective date: 20061114 |
|
A131 | Notification of reasons for refusal |
Free format text: JAPANESE INTERMEDIATE CODE: A131 Effective date: 20090616 |
|
A131 | Notification of reasons for refusal |
Free format text: JAPANESE INTERMEDIATE CODE: A131 Effective date: 20100511 |
|
A601 | Written request for extension of time |
Free format text: JAPANESE INTERMEDIATE CODE: A601 Effective date: 20100810 |
|
A602 | Written permission of extension of time |
Free format text: JAPANESE INTERMEDIATE CODE: A602 Effective date: 20100813 |
|
A601 | Written request for extension of time |
Free format text: JAPANESE INTERMEDIATE CODE: A601 Effective date: 20100910 |
|
A602 | Written permission of extension of time |
Free format text: JAPANESE INTERMEDIATE CODE: A602 Effective date: 20100915 |
|
A521 | Written amendment |
Free format text: JAPANESE INTERMEDIATE CODE: A523 Effective date: 20101008 |
|
A131 | Notification of reasons for refusal |
Free format text: JAPANESE INTERMEDIATE CODE: A131 Effective date: 20101102 |
|
A521 | Written amendment |
Free format text: JAPANESE INTERMEDIATE CODE: A523 Effective date: 20110202 |
|
A01 | Written decision to grant a patent or to grant a registration (utility model) |
Free format text: JAPANESE INTERMEDIATE CODE: A01 Effective date: 20110222 |
|
A61 | First payment of annual fees (during grant procedure) |
Free format text: JAPANESE INTERMEDIATE CODE: A61 Effective date: 20110318 |
|
R250 | Receipt of annual fees |
Free format text: JAPANESE INTERMEDIATE CODE: R250 |
|
R250 | Receipt of annual fees |
Free format text: JAPANESE INTERMEDIATE CODE: R250 |
|
R250 | Receipt of annual fees |
Free format text: JAPANESE INTERMEDIATE CODE: R250 |
|
R250 | Receipt of annual fees |
Free format text: JAPANESE INTERMEDIATE CODE: R250 |
|
R250 | Receipt of annual fees |
Free format text: JAPANESE INTERMEDIATE CODE: R250 |
|
R250 | Receipt of annual fees |
Free format text: JAPANESE INTERMEDIATE CODE: R250 |