DE602007005833D1 - Sprachaktivitätdetektionssystem und verfahren - Google Patents

Sprachaktivitätdetektionssystem und verfahren

Info

Publication number
DE602007005833D1
DE602007005833D1 DE602007005833T DE602007005833T DE602007005833D1 DE 602007005833 D1 DE602007005833 D1 DE 602007005833D1 DE 602007005833 T DE602007005833 T DE 602007005833T DE 602007005833 T DE602007005833 T DE 602007005833T DE 602007005833 D1 DE602007005833 D1 DE 602007005833D1
Authority
DE
Germany
Prior art keywords
classes
frames
discrimination
feature vectors
detection system
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Active
Application number
DE602007005833T
Other languages
English (en)
Inventor
Zica Valsan
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
International Business Machines Corp
Original Assignee
International Business Machines Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by International Business Machines Corp filed Critical International Business Machines Corp
Publication of DE602007005833D1 publication Critical patent/DE602007005833D1/de
Active legal-status Critical Current
Anticipated expiration legal-status Critical

Links

Classifications

    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS OR SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING; SPEECH OR AUDIO CODING OR DECODING
    • G10L25/00Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
    • G10L25/78Detection of presence or absence of voice signals
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS OR SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/02Feature extraction for speech recognition; Selection of recognition unit
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS OR SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING; SPEECH OR AUDIO CODING OR DECODING
    • G10L25/00Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
    • G10L25/03Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the type of extracted parameters
DE602007005833T 2006-11-16 2007-10-26 Sprachaktivitätdetektionssystem und verfahren Active DE602007005833D1 (de)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
EP06124228 2006-11-16
PCT/EP2007/061534 WO2008058842A1 (en) 2006-11-16 2007-10-26 Voice activity detection system and method

Publications (1)

Publication Number Publication Date
DE602007005833D1 true DE602007005833D1 (de) 2010-05-20

Family

ID=38857912

Family Applications (1)

Application Number Title Priority Date Filing Date
DE602007005833T Active DE602007005833D1 (de) 2006-11-16 2007-10-26 Sprachaktivitätdetektionssystem und verfahren

Country Status (9)

Country Link
US (2) US8311813B2 (de)
EP (1) EP2089877B1 (de)
JP (1) JP4568371B2 (de)
KR (1) KR101054704B1 (de)
CN (1) CN101548313B (de)
AT (1) ATE463820T1 (de)
CA (1) CA2663568C (de)
DE (1) DE602007005833D1 (de)
WO (1) WO2008058842A1 (de)

Families Citing this family (74)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
KR20010103333A (ko) * 2000-05-09 2001-11-23 류영선 즉석두부용 분말 제조방법
US8131543B1 (en) * 2008-04-14 2012-03-06 Google Inc. Speech detection
WO2010070839A1 (ja) * 2008-12-17 2010-06-24 日本電気株式会社 音声検出装置、音声検出プログラムおよびパラメータ調整方法
US8554348B2 (en) * 2009-07-20 2013-10-08 Apple Inc. Transient detection using a digital audio workstation
JP5334142B2 (ja) * 2009-07-21 2013-11-06 独立行政法人産業技術総合研究所 混合音信号中の混合比率推定方法及びシステム並びに音素認識方法
CN102044242B (zh) * 2009-10-15 2012-01-25 华为技术有限公司 语音激活检测方法、装置和电子设备
CN102714034B (zh) * 2009-10-15 2014-06-04 华为技术有限公司 信号处理的方法、装置和系统
JP2013508773A (ja) * 2009-10-19 2013-03-07 テレフオンアクチーボラゲット エル エム エリクソン(パブル) 音声エンコーダの方法およびボイス活動検出器
US8626498B2 (en) * 2010-02-24 2014-01-07 Qualcomm Incorporated Voice activity detection based on plural voice activity detectors
EP2561508A1 (de) 2010-04-22 2013-02-27 Qualcomm Incorporated Sprachaktivitätserkennung
US8762144B2 (en) * 2010-07-21 2014-06-24 Samsung Electronics Co., Ltd. Method and apparatus for voice activity detection
CN102446506B (zh) * 2010-10-11 2013-06-05 华为技术有限公司 音频信号的分类识别方法及装置
US8898058B2 (en) 2010-10-25 2014-11-25 Qualcomm Incorporated Systems, methods, and apparatus for voice activity detection
CN102741918B (zh) * 2010-12-24 2014-11-19 华为技术有限公司 用于话音活动检测的方法和设备
PT3493205T (pt) 2010-12-24 2021-02-03 Huawei Tech Co Ltd Método e aparelho para detetar de forma adaptativa uma atividade de voz num sinal de áudio de entrada
CN102097095A (zh) * 2010-12-28 2011-06-15 天津市亚安科技电子有限公司 一种语音端点检测方法及装置
US20130090926A1 (en) * 2011-09-16 2013-04-11 Qualcomm Incorporated Mobile device context information using speech detection
US9235799B2 (en) 2011-11-26 2016-01-12 Microsoft Technology Licensing, Llc Discriminative pretraining of deep neural networks
US8965763B1 (en) * 2012-02-02 2015-02-24 Google Inc. Discriminative language modeling for automatic speech recognition with a weak acoustic model and distributed training
US8543398B1 (en) 2012-02-29 2013-09-24 Google Inc. Training an automatic speech recognition system using compressed word frequencies
US8374865B1 (en) 2012-04-26 2013-02-12 Google Inc. Sampling training data for an automatic speech recognition system based on a benchmark classification distribution
US8805684B1 (en) 2012-05-31 2014-08-12 Google Inc. Distributed speaker adaptation
US8571859B1 (en) 2012-05-31 2013-10-29 Google Inc. Multi-stage speaker adaptation
US8880398B1 (en) 2012-07-13 2014-11-04 Google Inc. Localized speech recognition with offload
US9123333B2 (en) 2012-09-12 2015-09-01 Google Inc. Minimum bayesian risk methods for automatic speech recognition
US10304465B2 (en) 2012-10-30 2019-05-28 Google Technology Holdings LLC Voice control user interface for low power mode
US10373615B2 (en) 2012-10-30 2019-08-06 Google Technology Holdings LLC Voice control user interface during low power mode
US10381001B2 (en) 2012-10-30 2019-08-13 Google Technology Holdings LLC Voice control user interface during low-power mode
US9584642B2 (en) 2013-03-12 2017-02-28 Google Technology Holdings LLC Apparatus with adaptive acoustic echo control for speakerphone mode
US9477925B2 (en) 2012-11-20 2016-10-25 Microsoft Technology Licensing, Llc Deep neural networks training for speech and pattern recognition
US9454958B2 (en) 2013-03-07 2016-09-27 Microsoft Technology Licensing, Llc Exploiting heterogeneous data in deep neural network-based speech recognition systems
US9570087B2 (en) * 2013-03-15 2017-02-14 Broadcom Corporation Single channel suppression of interfering sources
CN107093991B (zh) 2013-03-26 2020-10-09 杜比实验室特许公司 基于目标响度的响度归一化方法和设备
US9466292B1 (en) * 2013-05-03 2016-10-11 Google Inc. Online incremental adaptation of deep neural networks using auxiliary Gaussian mixture models in speech recognition
US9997172B2 (en) * 2013-12-02 2018-06-12 Nuance Communications, Inc. Voice activity detection (VAD) for a coded speech bitstream without decoding
US8768712B1 (en) 2013-12-04 2014-07-01 Google Inc. Initiating actions based on partial hotwords
EP2945303A1 (de) 2014-05-16 2015-11-18 Thomson Licensing Verfahren und Vorrichtung zur Auswahl oder Beseitigung von Audiokomponentenarten
US10650805B2 (en) * 2014-09-11 2020-05-12 Nuance Communications, Inc. Method for scoring in an automatic speech recognition system
US9324320B1 (en) * 2014-10-02 2016-04-26 Microsoft Technology Licensing, Llc Neural network-based speech processing
US9842608B2 (en) 2014-10-03 2017-12-12 Google Inc. Automatic selective gain control of audio data for speech recognition
CN105529038A (zh) * 2014-10-21 2016-04-27 阿里巴巴集团控股有限公司 对用户语音信号进行处理的方法及其系统
US10403269B2 (en) 2015-03-27 2019-09-03 Google Llc Processing audio waveforms
US10515301B2 (en) 2015-04-17 2019-12-24 Microsoft Technology Licensing, Llc Small-footprint deep neural network
US10121471B2 (en) * 2015-06-29 2018-11-06 Amazon Technologies, Inc. Language model speech endpointing
CN104980211B (zh) * 2015-06-29 2017-12-12 北京航天易联科技发展有限公司 一种信号处理方法和装置
US10229700B2 (en) * 2015-09-24 2019-03-12 Google Llc Voice activity detection
US10339921B2 (en) 2015-09-24 2019-07-02 Google Llc Multichannel raw-waveform neural networks
US10347271B2 (en) * 2015-12-04 2019-07-09 Synaptics Incorporated Semi-supervised system for multichannel source enhancement through configurable unsupervised adaptive transformations and supervised deep neural network
US9959887B2 (en) 2016-03-08 2018-05-01 International Business Machines Corporation Multi-pass speech activity detection strategy to improve automatic speech recognition
US10490209B2 (en) * 2016-05-02 2019-11-26 Google Llc Automatic determination of timing windows for speech captions in an audio stream
CN107564512B (zh) * 2016-06-30 2020-12-25 展讯通信(上海)有限公司 语音活动侦测方法及装置
US10475471B2 (en) * 2016-10-11 2019-11-12 Cirrus Logic, Inc. Detection of acoustic impulse events in voice applications using a neural network
US10242696B2 (en) 2016-10-11 2019-03-26 Cirrus Logic, Inc. Detection of acoustic impulse events in voice applications
WO2018118744A1 (en) * 2016-12-19 2018-06-28 Knowles Electronics, Llc Methods and systems for reducing false alarms in keyword detection
CN106782529B (zh) * 2016-12-23 2020-03-10 北京云知声信息技术有限公司 语音识别的唤醒词选择方法及装置
US10810995B2 (en) * 2017-04-27 2020-10-20 Marchex, Inc. Automatic speech recognition (ASR) model training
US10311874B2 (en) 2017-09-01 2019-06-04 4Q Catalyst, LLC Methods and systems for voice-based programming of a voice-controlled device
US10403303B1 (en) * 2017-11-02 2019-09-03 Gopro, Inc. Systems and methods for identifying speech based on cepstral coefficients and support vector machines
CN107808659A (zh) * 2017-12-02 2018-03-16 宫文峰 智能语音信号模式识别系统装置
CN109065027B (zh) * 2018-06-04 2023-05-02 平安科技(深圳)有限公司 语音区分模型训练方法、装置、计算机设备及存储介质
WO2019244298A1 (ja) * 2018-06-21 2019-12-26 日本電気株式会社 属性識別装置、属性識別方法、およびプログラム記録媒体
CN108922556B (zh) * 2018-07-16 2019-08-27 百度在线网络技术(北京)有限公司 声音处理方法、装置及设备
US20200074997A1 (en) * 2018-08-31 2020-03-05 CloudMinds Technology, Inc. Method and system for detecting voice activity in noisy conditions
CN111199733A (zh) * 2018-11-19 2020-05-26 珠海全志科技股份有限公司 多级识别语音唤醒方法及装置、计算机存储介质及设备
CN111524536B (zh) 2019-02-01 2023-09-08 富士通株式会社 信号处理方法和信息处理设备
CN109754823A (zh) * 2019-02-26 2019-05-14 维沃移动通信有限公司 一种语音活动检测方法、移动终端
CN110349597B (zh) * 2019-07-03 2021-06-25 山东师范大学 一种语音检测方法及装置
KR20210044559A (ko) 2019-10-15 2021-04-23 삼성전자주식회사 출력 토큰 결정 방법 및 장치
US11270720B2 (en) * 2019-12-30 2022-03-08 Texas Instruments Incorporated Background noise estimation and voice activity detection system
CN112420022A (zh) * 2020-10-21 2021-02-26 浙江同花顺智能科技有限公司 一种噪声提取方法、装置、设备和存储介质
CN112509598A (zh) * 2020-11-20 2021-03-16 北京小米松果电子有限公司 音频检测方法及装置、存储介质
CN112466056B (zh) * 2020-12-01 2022-04-05 上海旷日网络科技有限公司 一种基于语音识别的自助柜取件系统及方法
KR102318642B1 (ko) * 2021-04-16 2021-10-28 (주)엠제이티 음성 분석 결과를 이용하는 온라인 플랫폼
US20230327892A1 (en) * 2022-04-07 2023-10-12 Bank Of America Corporation System And Method For Managing Exception Request Blocks In A Blockchain Network

Family Cites Families (18)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US4696039A (en) * 1983-10-13 1987-09-22 Texas Instruments Incorporated Speech analysis/synthesis system with silence suppression
US4780906A (en) * 1984-02-17 1988-10-25 Texas Instruments Incorporated Speaker-independent word recognition method and system based upon zero-crossing rate and energy measurement of analog speech signal
DE4422545A1 (de) 1994-06-28 1996-01-04 Sel Alcatel Ag Start-/Endpunkt-Detektion zur Worterkennung
US6314396B1 (en) * 1998-11-06 2001-11-06 International Business Machines Corporation Automatic gain control in a speech recognition system
US6556967B1 (en) * 1999-03-12 2003-04-29 The United States Of America As Represented By The National Security Agency Voice activity detector
US6615170B1 (en) * 2000-03-07 2003-09-02 International Business Machines Corporation Model-based voice activity detection system and method using a log-likelihood ratio and pitch
US6901362B1 (en) * 2000-04-19 2005-05-31 Microsoft Corporation Audio segmentation and classification
JP3721948B2 (ja) 2000-05-30 2005-11-30 株式会社国際電気通信基礎技術研究所 音声始端検出方法、音声認識装置における音声区間検出方法および音声認識装置
US6754626B2 (en) * 2001-03-01 2004-06-22 International Business Machines Corporation Creating a hierarchical tree of language models for a dialog system based on prompt and dialog context
EP1443498B1 (de) * 2003-01-24 2008-03-19 Sony Ericsson Mobile Communications AB Rauschreduzierung und audiovisuelle Sprachaktivitätsdetektion
FI20045315A (fi) * 2004-08-30 2006-03-01 Nokia Corp Ääniaktiivisuuden havaitseminen äänisignaalissa
US20070033042A1 (en) * 2005-08-03 2007-02-08 International Business Machines Corporation Speech detection fusing multi-class acoustic-phonetic, and energy features
US20070036342A1 (en) * 2005-08-05 2007-02-15 Boillot Marc A Method and system for operation of a voice activity detector
CN100573663C (zh) * 2006-04-20 2009-12-23 南京大学 基于语音特征判别的静音检测方法
US20080010065A1 (en) * 2006-06-05 2008-01-10 Harry Bratt Method and apparatus for speaker recognition
US20080300875A1 (en) * 2007-06-04 2008-12-04 Texas Instruments Incorporated Efficient Speech Recognition with Cluster Methods
KR100930584B1 (ko) * 2007-09-19 2009-12-09 한국전자통신연구원 인간 음성의 유성음 특징을 이용한 음성 판별 방법 및 장치
US8131543B1 (en) * 2008-04-14 2012-03-06 Google Inc. Speech detection

Also Published As

Publication number Publication date
US20120330656A1 (en) 2012-12-27
KR101054704B1 (ko) 2011-08-08
KR20090083367A (ko) 2009-08-03
CN101548313B (zh) 2011-07-13
JP4568371B2 (ja) 2010-10-27
EP2089877B1 (de) 2010-04-07
US8554560B2 (en) 2013-10-08
CA2663568A1 (en) 2008-05-22
US20100057453A1 (en) 2010-03-04
CA2663568C (en) 2016-01-05
CN101548313A (zh) 2009-09-30
US8311813B2 (en) 2012-11-13
ATE463820T1 (de) 2010-04-15
WO2008058842A1 (en) 2008-05-22
JP2010510534A (ja) 2010-04-02
EP2089877A1 (de) 2009-08-19

Similar Documents

Publication Publication Date Title
DE602007005833D1 (de) Sprachaktivitätdetektionssystem und verfahren
WO2009100410A3 (en) Method and system for analysis of flow cytometry data using support vector machines
KR20180084576A (ko) 행동-인식 연결 학습 기반 의도 이해 장치, 방법 및 그 방법을 수행하기 위한 기록 매체
WO2009046359A3 (en) Detection and classification of running vehicles based on acoustic signatures
WO2017134416A3 (en) Touchscreen panel signal processing
DE602005008041D1 (de) Verfahren und system zur klassifizierung eines audiosignals
NZ603953A (en) Apparatus and method for analysing events from sensor data by optimization
CN101164105A (zh) 用于减小音频噪声的系统和方法
WO2006099218A3 (en) Methods and systems for evaluating and generating anomaly detectors
TW200745975A (en) System and methods for quantitatively evaluating complexity of computing system configuration
CN104538041A (zh) 异常声音检测方法及系统
MX2021014721A (es) Sistemas y metodos para aprendizaje de maquina de atributos de voz.
WO2012138917A3 (en) Gesture-activated input using audio recognition
GB201222640D0 (en) Method and apparatus for detecting and classifying signals
ATE506890T1 (de) Vorrichtung und verfahren zur vorhersage eines kontrollverlustes über einen muskel
TW200743372A (en) Apparatus and method for reducing temporal noise
WO2010118233A3 (en) Cadence analysis of temporal gait patterns for seismic discrimination
MY182556A (en) Image pattern recognition system and method
GB2493030B (en) Method of sound analysis and associated sound synthesis
CN106356074A (zh) 声音信号处理方法
TW200736951A (en) Identification of input sequences
WO2007121431A3 (en) Classification of composite actions involving interaction with objects
GB2456985A (en) System and method for determing seismic event location
KR101749254B1 (ko) 딥 러닝 기반의 통합 음향 정보 인지 시스템
MY191125A (en) Audio data processing method and terminal

Legal Events

Date Code Title Description
8320 Willingness to grant licences declared (paragraph 23)
8364 No opposition during term of opposition