WO2020164397A1 - 一种语音识别方法及系统 - Google Patents

一种语音识别方法及系统 Download PDF

Info

Publication number
WO2020164397A1
WO2020164397A1 PCT/CN2020/074178 CN2020074178W WO2020164397A1 WO 2020164397 A1 WO2020164397 A1 WO 2020164397A1 CN 2020074178 W CN2020074178 W CN 2020074178W WO 2020164397 A1 WO2020164397 A1 WO 2020164397A1
Authority
WO
WIPO (PCT)
Prior art keywords
different
recognition
doa
doas
signal
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2020/074178
Other languages
English (en)
French (fr)
Inventor
张仕良
雷鸣
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Alibaba Group Holding Ltd
Original Assignee
Alibaba Group Holding Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Alibaba Group Holding Ltd filed Critical Alibaba Group Holding Ltd
Priority to US17/428,015 priority Critical patent/US12315527B2/en
Publication of WO2020164397A1 publication Critical patent/WO2020164397A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L21/00Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
    • G10L21/02Speech enhancement, e.g. noise reduction or echo cancellation
    • G10L21/0208Noise filtering
    • G10L21/0216Noise filtering characterised by the method used for estimating noise
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04RLOUDSPEAKERS, MICROPHONES, GRAMOPHONE PICK-UPS OR LIKE ACOUSTIC ELECTROMECHANICAL TRANSDUCERS; ELECTRIC HEARING AIDS; PUBLIC ADDRESS SYSTEMS
    • H04R3/00Circuits for transducers
    • H04R3/005Circuits for transducers for combining the signals of two or more microphones
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04RLOUDSPEAKERS, MICROPHONES, GRAMOPHONE PICK-UPS OR LIKE ACOUSTIC ELECTROMECHANICAL TRANSDUCERS; ELECTRIC HEARING AIDS; PUBLIC ADDRESS SYSTEMS
    • H04R3/00Circuits for transducers
    • H04R3/04Circuits for transducers for correcting frequency response
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/06Creation of reference templates; Training of speech recognition systems, e.g. adaptation to the characteristics of the speaker's voice
    • G10L15/063Training
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L21/00Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
    • G10L21/02Speech enhancement, e.g. noise reduction or echo cancellation
    • G10L21/0208Noise filtering
    • G10L21/0216Noise filtering characterised by the method used for estimating noise
    • G10L2021/02161Number of inputs available containing the signal or the noise to be suppressed
    • G10L2021/02166Microphone arrays; Beamforming
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L21/00Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
    • G10L21/02Speech enhancement, e.g. noise reduction or echo cancellation
    • G10L21/0208Noise filtering
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04RLOUDSPEAKERS, MICROPHONES, GRAMOPHONE PICK-UPS OR LIKE ACOUSTIC ELECTROMECHANICAL TRANSDUCERS; ELECTRIC HEARING AIDS; PUBLIC ADDRESS SYSTEMS
    • H04R2430/00Signal processing covered by H04R, not provided for in its groups
    • H04R2430/20Processing of the output signals of the acoustic transducers of an array for obtaining a desired directivity characteristic

Definitions

  • This application relates to, but is not limited to, signal processing technology, especially a voice recognition method and system.
  • the far-field speech recognition system mainly includes two components: one is the front-end signal enhancement part, which is used to process the received multi-channel noisy speech signal to obtain an enhanced single-channel speech signal.
  • the front-end signal enhancement part removes certain noise interference by using the correlation between the voice signals of multiple channels to improve the signal-to-noise ratio of the signal; the other is the commonly used back-end speech recognition (ASR) part for the front-end signal.
  • ASR back-end speech recognition
  • This application provides a voice recognition method and system, which can ensure the accuracy of the voice recognition result.
  • the embodiment of the present invention provides a voice recognition method, including:
  • the dividing signal sources according to different directions of arrival DoA includes:
  • the DoA angle includes at least one of the following: 30 degrees, 60 degrees, 90 degrees, 120 degrees, and 150 degrees.
  • performing enhancement processing on signal sources corresponding to different DoAs respectively includes:
  • the signal sources corresponding to different DoAs are respectively subjected to a beamforming method based on delay superposition DAS to obtain the enhanced signal.
  • performing enhancement processing on signal sources corresponding to different DoAs respectively includes:
  • the MVDR beamforming method is respectively performed on the signal sources corresponding to different DoAs to obtain the enhanced signal.
  • the method further includes: dividing the space according to different DoAs; performing speech enhancement processing on the speech signals in different areas to obtain different enhanced signal samples;
  • the sample training obtains the acoustic models corresponding to different DoAs.
  • the inputting the recognition results of different DoA into respective acoustic models, and performing fusion processing on the output results of each acoustic model to obtain the recognition results includes:
  • the recognition results corresponding to different DoAs are input into the respective acoustic models; the output results of each acoustic model are fused to obtain the recognition results.
  • the fusion is achieved through a ROVER-based fusion system.
  • This application also provides a computer-readable storage medium that stores computer-executable instructions, and the computer-executable instructions are used to execute any of the speech recognition methods described above.
  • This application further provides a device for realizing information sharing, including a memory and a processor, wherein the memory stores the following instructions that can be executed by the processor: step.
  • the present application also provides a sound box including a memory and a processor, wherein the memory stores the following instructions that can be executed by the processor: for executing the steps of any one of the above-mentioned voice recognition methods.
  • This application further provides a speech recognition system, including: a preprocessing module, a first processing module, a second processing module, and a recognition module; wherein,
  • the preprocessing module is used to divide the signal source according to different DoA
  • the first processing module is used to perform enhancement processing on signal sources corresponding to different DoAs
  • the second processing module is configured to perform voice recognition on the enhanced processed signals corresponding to different DoAs to obtain recognition results corresponding to different DoAs;
  • the recognition module is used to input the recognition results of different DoA into their respective acoustic models, and perform fusion processing on the output results of each acoustic model to obtain the recognition results.
  • the device further includes: a training module, configured to divide the space according to different DoAs; perform voice enhancement processing on voice signals in different regions to obtain different enhanced signal samples; The acoustic models corresponding to different DoAs are obtained by training using the obtained samples.
  • a training module configured to divide the space according to different DoAs; perform voice enhancement processing on voice signals in different regions to obtain different enhanced signal samples;
  • the acoustic models corresponding to different DoAs are obtained by training using the obtained samples.
  • This application includes: dividing signal sources according to different directions of arrival DoA; performing enhancement processing on signal sources corresponding to different DoAs; performing voice recognition on signals corresponding to different DoAs after enhancement processing, and obtaining signals corresponding to different DoAs
  • Recognition results Input the recognition results of different DoA into their respective acoustic models, and perform fusion processing on the output results of each acoustic model to obtain the recognition results.
  • the space is divided into several regions through a preset DoA angle, thereby dividing the signal source into different spatial regions; further, the signals of different spatial regions are enhanced and identified and then fused to obtain the identification result of the signal source.
  • This application does not need to estimate the true signal source direction at each moment, avoiding the problem of inaccurate recognition caused by estimating the signal-to-noise ratio and signal source direction in a complex environment, thereby ensuring the accuracy of speech recognition results Sex.
  • Figure 1 is a schematic flow diagram of the speech recognition method of this application.
  • FIG. 2 is an example diagram of a beamforming method based on Delay-and-Sum in this application;
  • Figure 3 is a schematic diagram of the composition structure of the speech recognition system of this application.
  • the computing device includes one or more processors (CPUs), input/output interfaces, network interfaces, and memory.
  • processors CPUs
  • input/output interfaces network interfaces
  • memory volatile and non-volatile memory
  • the memory may include non-permanent memory in computer readable media, random access memory (RAM) and/or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). Memory is an example of computer readable media.
  • RAM random access memory
  • ROM read-only memory
  • flash RAM flash memory
  • Computer-readable media includes permanent and non-permanent, removable and non-removable media, and information storage can be realized by any method or technology.
  • the information can be computer-readable instructions, data structures, program modules, or other data.
  • Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, CD-ROM, digital versatile disc (DVD) or other optical storage, Magnetic cassettes, magnetic tape storage or other magnetic storage devices or any other non-transmission media can be used to store information that can be accessed by computing devices.
  • computer-readable media does not include non-transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.
  • Figure 1 is a schematic flowchart of the speech recognition method of this application, as shown in Figure 1, including:
  • Step 100 Divide signal sources according to different directions of arrival (DoA, Direction Of Arrival).
  • the delay of sound reaching different microphones of the microphone array can be calculated by calculating the target sound source, that is, the signal source in step 100, which may be within a certain angle of space, that is, within a certain DOA angle.
  • the target sound source that is, the signal source in step 100
  • the inventor of the present application found that when the DoA cannot be accurately estimated, the space can be divided into different directions, and then it is assumed that the target sound source is in this direction.
  • the space is divided into multiple regions according to at least one preset DoA angle, such as 30 degrees, 60 degrees, 90 degrees, 120 degrees, 150 degrees, etc., so as to assume that the signal source appears in these DoA Within the angle area, that is, the signal source is divided into areas formed by different DoA angles. What needs to be explained here is that the signal source is mobile, so it may be in different DoA angle areas at different times, but it will definitely be in a certain DoA angle area.
  • DoA angle such as 30 degrees, 60 degrees, 90 degrees, 120 degrees, 150 degrees, etc.
  • the spatial Divide into several areas, thus assuming that the signal source appears in these DoA angle areas.
  • the subsequent signal enhancement processing can be performed separately for the signal source in each area.
  • Step 101 Perform enhancement processing on signal sources corresponding to different DoAs.
  • the enhancement processing may include:
  • the signal sources corresponding to different DoA are respectively subjected to a beamforming method based on delay-and-sum (DAS, Delay-and-Sum) to obtain an enhanced signal.
  • DAS delay-and-sum
  • Figure 2 is an example of a Delay-and-Sum beam forming method based on this application.
  • DAS delay-and-sum
  • Figure 2 is an example of a Delay-and-Sum beam forming method based on this application.
  • the enhancement processing may include:
  • MVDR Minimum Variance Distortionless Response
  • Capon is an adaptive spatial wavenumber spectrum estimation algorithm proposed by Capon in 1967.
  • Step 102 Perform voice recognition on the enhanced processed signals corresponding to different DoAs to obtain recognition results corresponding to different DoAs.
  • speech recognition may include, for example, an ASR system.
  • Step 103 Input the recognition results of different DoA into respective acoustic models, and perform fusion processing on the output results of each theological model to obtain the recognition results corresponding to the signal source.
  • this step also includes: partitioning the space according to different DoAs, and then performing speech enhancement processing on the speech signals in different areas to obtain different enhanced signal samples, and using the obtained samples to train to obtain acoustic models corresponding to different DoAs.
  • partitioning the space according to different DoAs and then performing speech enhancement processing on the speech signals in different areas to obtain different enhanced signal samples, and using the obtained samples to train to obtain acoustic models corresponding to different DoAs.
  • this step may include: inputting the recognition results corresponding to different DoAs into the respective trained acoustic models, and then fusing the output results of each acoustic model using, for example, a ROVER-based fusion system to obtain the final signal The recognition result corresponding to the source.
  • the fusion may be implemented by a fusion system based on a recognition result voting error reduction (ROVER, Recognizer Output Voting Error Reduction) method.
  • ROVER Recognizer Output Voting Error Reduction
  • the space is divided into several regions through a preset DoA angle, thereby dividing the signal source into different spatial regions; further, the signals of different spatial regions are enhanced and identified and then fused to obtain the identification result of the signal source.
  • This application does not need to estimate the true signal source direction at each moment, avoiding the problem of inaccurate recognition caused by estimating the signal-to-noise ratio and signal source direction in a complex environment, thereby ensuring the accuracy of speech recognition results Sex.
  • the present application also provides a computer-readable storage medium that stores computer-executable instructions, and the computer-executable instructions are used to execute any of the above-mentioned speech recognition methods.
  • the present application further provides a device for realizing information sharing, including a memory and a processor, wherein the memory stores the steps of any one of the aforementioned voice recognition methods.
  • the present application also provides a sound box including a memory and a processor, wherein the memory stores the following instructions that can be executed by the processor: for executing the steps of any one of the above-mentioned voice recognition methods.
  • Fig. 3 is a schematic diagram of the composition structure of the speech recognition system of this application. As shown in Fig. 3, it at least includes: a preprocessing module, a first processing module, a second processing module, and a recognition module;
  • the preprocessing module is used to divide the signal source according to different DoA
  • the first processing module is used to perform enhancement processing on signal sources corresponding to different DoAs
  • the second processing module is used to perform voice recognition on the enhanced processed signals corresponding to different DoAs to obtain recognition results;
  • the recognition module is used to input the recognition results of different DoA into their respective acoustic models, and perform fusion processing on the output results of each acoustic model to obtain the recognition results.
  • the preprocessing module is specifically used for:
  • the space is divided into multiple regions, so as to assume that the signal source appears in these DoA angle regions, that is, Divide the signal source into areas formed by different DoA angles.
  • the first processing module is specifically configured to:
  • the MVDR beamforming method is performed on the signal sources corresponding to different DoAs to obtain the enhanced signal.
  • MVDR is an adaptive spatial wavenumber spectrum estimation algorithm proposed by Capon in 1967.
  • the second processing module may be an ASR system.
  • the identification module is specifically used for:
  • the speech recognition device of the present application also includes: a training module for dividing the space according to different DoAs; performing speech enhancement processing on the speech signals in different areas to obtain different enhanced signal samples; and training using the obtained samples Obtain the acoustic models corresponding to different DoAs.
  • modules in the speech recognition system of the present application may be individually set in different physical devices, or may be set in multiple physical devices after a reasonable combination, or all may be set in the same physical device.

Landscapes

  • Engineering & Computer Science (AREA)
  • Acoustics & Sound (AREA)
  • Physics & Mathematics (AREA)
  • Health & Medical Sciences (AREA)
  • Signal Processing (AREA)
  • Computational Linguistics (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Human Computer Interaction (AREA)
  • Multimedia (AREA)
  • Otolaryngology (AREA)
  • General Health & Medical Sciences (AREA)
  • Quality & Reliability (AREA)
  • Artificial Intelligence (AREA)
  • Circuit For Audible Band Transducer (AREA)

Abstract

本申请公开了一种语音识别方法及系统,本申请实施例通过预先设置的DoA角度,将空间分成若干个区域,从而将信号源划分至不同的空间区域;进而对不同空间区域的信号进行增强和识别处理后融合得到信号源的识别结果。本申请不需要估计每个时刻真实的信号源方向,避免了在复杂环境下,由于估计信号的信噪比和信号源方向而带来的识别不准确的问题,从而保障了语音识别结果的准确性。

Description

一种语音识别方法及系统
本申请要求2019年02月12日递交的申请号为201910111593.1、发明名称为“一种语音识别方法及系统”的中国专利申请的优先权,其全部内容通过引用结合在本申请中。
技术领域
本申请涉及但不限于信号处理技术,尤指一种语音识别方法及系统。
背景技术
相关技术中,远场语音识别系统主要包括两个组成部分:一是前端信号增强部分,用于对接收到的多通道带噪语音信号进行处理,从而得到增强后的单通道语音信号。前端信号增强部分通过利用多个通道的语音信号之间的关联性,去除一定的噪声干扰,提升信号的信噪比;另一个是后端常用的语音识别(ASR)部分,用于对前端信号增强部分处理后的单通道语音信号输入一个通用的语音识别系统,以得到最终的语音识别结果。
在复杂环境下,很难估计出信号的信噪比和信号源方向,也就是说,相关技术中的远场语音识别技术很难保障语音识别结果的准确性。
发明内容
本申请提供一种语音识别方法及系统,能够保障语音识别结果的准确性。
本发明实施例提供了一种语音识别方法,包括:
按照不同的波达方向DoA对信号源进行划分;
对对应于不同DoA的信号源分别进行增强处理;
对增强处理后的对应不同DoA的信号分别进行语音识别,得到对应不同DoA的识别结果;
将不同DoA的识别结果输入各自的声学模型,并对各声学模型的输出结果进行融合处理,得到识别结果。
在一种示例性实例中,所述按照不同的波达方向DoA对信号源进行划分,包括:
将空间划分为多个区域,从而将信号源划分至不同DoA角度形成的区域。
在一种示例性实例中,所述DoA角度包括以下至少之一:30度、60度、90度、120度、150度。
在一种示例性实例中,所述对对应于不同DoA的信号源分别进行增强处理,包括:
对所述对应于不同DoA的信号源都分别进行基于延时叠加DAS的波束形成方法,得到所述增强处理后的信号。
在一种示例性实例中,所述对对应于不同DoA的信号源分别进行增强处理,包括:
对所述对应于不同DoA的信号源都分别进行MVDR的波束形成方法,得到所述增强处理后的信号。
在一种示例性实例中,所述方法之前还包括:根据不同的所述DoA对空间进行区域划分;对不同区域内的语音信号进行语音增强处理,得到不同的增强信号样本;利用得到的各样本训练得到对应不同DoA的所述声学模型。
在一种示例性实例中,所述将不同DoA的识别结果输入各自的声学模型,并对各声学模型的输出结果进行融合处理,得到识别结果,包括:
将对应不同DoA的所述识别结果输入各自所述声学模型;对各声学模型的输出结果进行融合,得到所述识别结果。
在一种示例性实例中,所述融合通过基于ROVER的融合系统实现。
本申请还提供了一种计算机可读存储介质,存储有计算机可执行指令,所述计算机可执行指令用于执行上述任一项所述的语音识别方法。
本申请又提供了一种用于实现信息分享的装置,包括存储器和处理器,其中,存储器中存储有以下可被处理器执行的指令:用于执行上述任一项所述的语音识别方法的步骤。
本申请还提供了一种音箱,包括存储器和处理器,其中,存储器中存储有以下可被处理器执行的指令:用于执行上述任一项所述的语音识别方法的步骤。
本申请再提供了一种语音识别系统,包括:预处理模块、第一处理模块、第二处理模块、识别模块;其中,
预处理模块,用于按照不同的DoA对信号源进行划分;
第一处理模块,用于对对应于不同DoA的信号源分别进行增强处理;
第二处理模块,用于对增强处理后的对应不同DoA的信号分别进行语音识别,得到对应不同DoA的识别结果;
识别模块,用于将不同DoA的识别结果输入各自的声学模型,并对各声学模型的输出结果进行融合处理,得到识别结果。
在一种示例性实例中,所述装置还包括:训练模块,用于根据不同的所述DoA对空间进行区域划分;对不同区域内的语音信号进行语音增强处理,得到不同的增强信号样 本;利用得到的各样本训练得到对应不同DoA的所述声学模型。
本申请包括:按照不同的波达方向DoA对信号源进行划分;对对应于不同DoA的信号源分别进行增强处理;对增强处理后的对应不同DoA的信号分别进行语音识别,得到对应不同DoA的识别结果;将不同DoA的识别结果输入各自的声学模型,并对各声学模型的输出结果进行融合处理,得到识别结果。本申请实施例通过预先设置的DoA角度,将空间分成若干个区域,从而将信号源划分至不同的空间区域;进而对不同空间区域的信号进行增强和识别处理后融合得到信号源的识别结果。本申请不需要估计每个时刻真实的信号源方向,避免了在复杂环境下,由于估计信号的信噪比和信号源方向而带来的识别不准确的问题,从而保障了语音识别结果的准确性。
本发明的其它特征和优点将在随后的说明书中阐述,并且,部分地从说明书中变得显而易见,或者通过实施本发明而了解。本发明的目的和其他优点可通过在说明书、权利要求书以及附图中所特别指出的结构来实现和获得。
附图说明
附图用来提供对本申请技术方案的进一步理解,并且构成说明书的一部分,与本申请的实施例一起用于解释本申请的技术方案,并不构成对本申请技术方案的限制。
图1为本申请语音识别方法的流程示意图;
图2为本申请一种基于Delay-and-Sum波束形成方法的示例图;
图3为本申请语音识别系统的组成结构示意图。
具体实施方式
为使本申请的目的、技术方案和优点更加清楚明白,下文中将结合附图对本申请的实施例进行详细说明。需要说明的是,在不冲突的情况下,本申请中的实施例及实施例中的特征可以相互任意组合。
在本申请一个典型的配置中,计算设备包括一个或多个处理器(CPU)、输入/输出接口、网络接口和内存。
内存可能包括计算机可读介质中的非永久性存储器,随机存取存储器(RAM)和/或非易失性内存等形式,如只读存储器(ROM)或闪存(flash RAM)。内存是计算机可读介质的示例。
计算机可读介质包括永久性和非永久性、可移动和非可移动媒体可以由任何方法或 技术来实现信息存储。信息可以是计算机可读指令、数据结构、程序的模块或其他数据。计算机的存储介质的例子包括,但不限于相变内存(PRAM)、静态随机存取存储器(SRAM)、动态随机存取存储器(DRAM)、其他类型的随机存取存储器(RAM)、只读存储器(ROM)、电可擦除可编程只读存储器(EEPROM)、快闪记忆体或其他内存技术、只读光盘只读存储器(CD-ROM)、数字多功能光盘(DVD)或其他光学存储、磁盒式磁带,磁带磁盘存储或其他磁性存储设备或任何其他非传输介质,可用于存储可以被计算设备访问的信息。按照本文中的界定,计算机可读介质不包括非暂存电脑可读媒体(transitory media),如调制的数据信号和载波。
在附图的流程图示出的步骤可以在诸如一组计算机可执行指令的计算机系统中执行。并且,虽然在流程图中示出了逻辑顺序,但是在某些情况下,可以以不同于此处的顺序执行所示出或描述的步骤。
图1为本申请语音识别方法的流程示意图,如图1所示,包括:
步骤100:按照不同的波达方向(DoA,Direction Of Arrival)对信号源进行划分。
声音到达麦克风阵列不同麦克风的延迟,通过这个延迟可以计算出目标声源即步骤100中的信号源可能在空间的某个角度内即某个DOA角度内。本申请发明人发现,当不能准确估计DoA时,可以将空间划分成不同的方向,然后假设目标声源在这个方向。
在一种示例性实例中,按照预先设置的至少一个DoA角度,比如30度、60度、90度、120度、150度等,将空间划分为多个区域,从而假设信号源出现在这些DoA角度区域内,也就是说,将信号源划分至不同DoA角度形成的区域。这里需要说明的是,信号源是移动的,所以不同时刻可能处于不同的DoA角度区域内,但是肯定会处于某个DoA角度区域内。
在复杂环境下,很难估计信号的信噪比和信号源方向,因此,本申请实施例中,并不需要估计每个时刻真实的信号源方向,而是通过预先设置的DoA角度,将空间分成若干个区域,从而假设信号源出现在这些DoA角度区域内。通过假设信号源总是会处于其中某一个DoA角度范围内,使得后续可以针对每个区域的信号源分别进行信号增强处理。
步骤101:对对应于不同DoA的信号源分别进行增强处理。
在一种示例性实例中,增强处理可以包括:
对对应于不同DoA的信号源都分别进行基于延时叠加(DAS,Delay-and-Sum)的波束形成方法,得到增强处理后的信号。图2为本申请一种基于Delay-and-Sum波束形 成方法的示例,具体实现可以参见相关技术,这里仅仅是举例说明,并不用于限定本申请的保护范围。
在一种示例性实例中,增强处理可以包括:
对对应于不同DoA的信号源都分别进行MVDR的波束形成方法,得到增强处理后的信号。其中,MVDR(Minimum Variance Distortionless Response)是Capon于1967年提出的一种自适应的空间波数谱估计算法。
步骤102:对增强处理后的对应不同DoA的信号分别进行语音识别,得到对应不同DoA的识别结果。
在一种示例性实例中,语音识别可以包括如ASR系统。
本申请中,由于对对应不同DoA的信号都进行了波束形成,因此,经过语音识别如ASR系统后会得到若干个对应不同DoA的识别结果。
步骤103:将不同DoA的识别结果输入各自的声学模型,并对各神学模型的输出结果进行融合处理,得到信号源对应的识别结果。
本步骤之前还包括:根据不同的DoA对空间进行区域划分,然后对不同区域的语音信号进行语音增强处理,得到不同的增强信号样本,利用得到的各样本训练得到对应不同DoA的声学模型。训练的方法很多,可以采用相关技术来实现,具体实现并不用于限定本申请的保护范围。
在一种示例性实例中,本步骤可以包括:将对应不同DoA的识别结果输入各自训练好的声学模型,然后再将各声学模型的输出结果采用如基于ROVER的融合系统进行融合,得到最终信号源对应的识别结果。
在一种示例性实例中,融合可以通过一种基于识别结果投票的错误降低(ROVER,Recognizer Output Voting Error Reduction)方法的融合系统来实现。
本申请实施例通过预先设置的DoA角度,将空间分成若干个区域,从而将信号源划分至不同的空间区域;进而对不同空间区域的信号进行增强和识别处理后融合得到信号源的识别结果。本申请不需要估计每个时刻真实的信号源方向,避免了在复杂环境下,由于估计信号的信噪比和信号源方向而带来的识别不准确的问题,从而保障了语音识别结果的准确性。
本申请还提供一种计算机可读存储介质,存储有计算机可执行指令,所述计算机可执行指令用于执行上述任一项的语音识别方法。
本申请再提供一种实现信息分享的装置,包括存储器和处理器,其中,存储器中存 储有上述任一项的语音识别方法的步骤。
本申请还提供一种音箱,包括存储器和处理器,其中,存储器中存储有以下可被处理器执行的指令:用于执行上述任一项所述的语音识别方法的步骤。
图3为本申请语音识别系统的组成结构示意图,如图3所示,至少包括:预处理模块、第一处理模块、第二处理模块、识别模块;其中,
预处理模块,用于按照不同的DoA对信号源进行划分;
第一处理模块,用于对对应于不同DoA的信号源分别进行增强处理;
第二处理模块,用于对增强处理后的对应不同DoA的信号分别进行语音识别,得到识别结果;
识别模块,用于将不同DoA的识别结果输入各自的声学模型,并对各声学模型的输出结果进行融合处理,得到识别结果。
在一种示例性实例中,预处理模块具体用于:
按照预先设置的至少一个DoA角度,比如30度、60度、90度、120度、150度等,将空间划分为多个区域,从而假设信号源出现在这些DoA角度区域内,也就是说,将信号源划分至不同DoA角度形成的区域。
在一种示例性实例中,第一处理模块具体用于:
对对应于不同DoA的信号源都分别进行基于DAS的波束形成方法,得到增强后的信号;
或者,对对应于不同DoA的信号源都分别进行MVDR的波束形成方法,得到增强后的信号。其中,MVDR是Capon于1967年提出的一种自适应的空间波数谱估计算法。
在一种示例性实例中,第二处理模块可以是ASR系统。
在一种示例性实例中,识别模块具体用于:
将对应不同DoA的识别结果输入训练好的各自的声学模型,将各声学模型的识别结果采用如基于ROVER的融合系统进行融合,得到信号源对应的识别结果。
本申请语音识别装置还包括:训练模块,用于根据不同的所述DoA对空间进行区域划分;对不同区域内的语音信号进行语音增强处理,得到不同的增强信号样本;利用得到的各样本训练得到对应不同DoA的所述声学模型。
需要说明的是,本申请语音识别系统中的各模块可以单独设置在不同的实体设备中,也可以合理组合后设置在多个实体设备中,还可以是都设置在同一实体设备中。
虽然本申请所揭露的实施方式如上,但所述的内容仅为便于理解本申请而采用的实 施方式,并非用以限定本申请。任何本申请所属领域内的技术人员,在不脱离本申请所揭露的精神和范围的前提下,可以在实施的形式及细节上进行任何的修改与变化,但本申请的专利保护范围,仍须以所附的权利要求书所界定的范围为准。

Claims (13)

  1. 一种语音识别方法,包括:
    按照不同的波达方向DoA对信号源进行划分;
    对对应于不同DoA的信号源分别进行增强处理;
    对增强处理后的对应不同DoA的信号分别进行语音识别,得到对应不同DoA的识别结果;
    将不同DoA的识别结果输入各自的声学模型,并对各声学模型的输出结果进行融合处理,得到识别结果。
  2. 根据权利要求1所述的语音识别方法,其中,所述按照不同的波达方向DoA对信号源进行划分,包括:
    将空间划分为多个区域,从而将信号源划分至不同DoA角度形成的区域。
  3. 根据权利要求2所述的语音识别方法,其中,所述DoA角度包括以下至少之一:30度、60度、90度、120度、150度。
  4. 根据权利要求1所述的语音识别方法,其中,所述对对应于不同DoA的信号源分别进行增强处理,包括:
    对所述对应于不同DoA的信号源都分别进行基于延时叠加DAS的波束形成方法,得到所述增强处理后的信号。
  5. 根据权利要求1所述的语音识别方法,其中,所述对对应于不同DoA的信号源分别进行增强处理,包括:
    对所述对应于不同DoA的信号源都分别进行MVDR的波束形成方法,得到所述增强处理后的信号。
  6. 根据权利要求1所述的语音识别方法,所述方法之前还包括:根据不同的所述DoA对空间进行区域划分;对不同区域内的语音信号进行语音增强处理,得到不同的增强信号样本;利用得到的各样本训练得到对应不同DoA的所述声学模型。
  7. 根据权利要求6所述的语音识别方法,其中,所述将不同DoA的识别结果输入各自的声学模型,并对各声学模型的输出结果进行融合处理,得到识别结果,包括:
    将对应不同DoA的所述识别结果输入各自所述声学模型;对各声学模型的输出结果进行融合,得到所述识别结果。
  8. 根据权利要求6所述的语音识别方法,其中,所述融合通过基于ROVER的融合系统实现。
  9. 一种计算机可读存储介质,存储有计算机可执行指令,所述计算机可执行指令用于执行权利要求1~权利要求8任一项所述的语音识别方法。
  10. 一种用于实现信息分享的装置,包括存储器和处理器,其中,存储器中存储有以下可被处理器执行的指令:用于执行权利要求1~权利要求8任一项所述的语音识别方法的步骤。
  11. 一种音箱,包括存储器和处理器,其中,存储器中存储有以下可被处理器执行的指令:用于执行权利要求1~权利要求8任一项所述的语音识别方法的步骤。
  12. 一种语音识别系统,包括:预处理模块、第一处理模块、第二处理模块、识别模块;其中,
    预处理模块,用于按照不同的DoA对信号源进行划分;
    第一处理模块,用于对对应于不同DoA的信号源分别进行增强处理;
    第二处理模块,用于对增强处理后的对应不同DoA的信号分别进行语音识别,得到对应不同DoA的识别结果;
    识别模块,用于将不同DoA的识别结果输入各自的声学模型,并对各声学模型的输出结果进行融合处理,得到识别结果。
  13. 根据权利要求12所述的语音识别系统,所述系统还包括:训练模块,用于根据不同的所述DoA对空间进行区域划分;对不同区域内的语音信号进行语音增强处理,得到不同的增强信号样本;利用得到的各样本训练得到对应不同DoA的所述声学模型。
PCT/CN2020/074178 2019-02-12 2020-02-03 一种语音识别方法及系统 Ceased WO2020164397A1 (zh)

Priority Applications (1)

Application Number Priority Date Filing Date Title
US17/428,015 US12315527B2 (en) 2019-02-12 2020-02-03 Method and system for speech recognition

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN201910111593.1A CN111627425B (zh) 2019-02-12 2019-02-12 一种语音识别方法及系统
CN201910111593.1 2019-02-12

Publications (1)

Publication Number Publication Date
WO2020164397A1 true WO2020164397A1 (zh) 2020-08-20

Family

ID=72045480

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2020/074178 Ceased WO2020164397A1 (zh) 2019-02-12 2020-02-03 一种语音识别方法及系统

Country Status (3)

Country Link
US (1) US12315527B2 (zh)
CN (1) CN111627425B (zh)
WO (1) WO2020164397A1 (zh)

Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN102271299A (zh) * 2010-06-01 2011-12-07 索尼公司 声音信号处理装置和声音信号处理方法
CN107742522A (zh) * 2017-10-23 2018-02-27 科大讯飞股份有限公司 基于麦克风阵列的目标语音获取方法及装置
CN109272989A (zh) * 2018-08-29 2019-01-25 北京京东尚科信息技术有限公司 语音唤醒方法、装置和计算机可读存储介质

Family Cites Families (61)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CA2069356C (en) 1991-07-17 1997-05-06 Gary Wayne Elko Adjustable filter for differential microphones
JP2780676B2 (ja) 1995-06-23 1998-07-30 日本電気株式会社 音声認識装置及び音声認識方法
EP0856832A1 (fr) 1997-02-03 1998-08-05 Koninklijke Philips Electronics N.V. Procédé de reconnaissance vocale de mots et dispositif dans lequel ledit procédé est mis en application
US6633842B1 (en) 1999-10-22 2003-10-14 Texas Instruments Incorporated Speech recognition front-end feature extraction for noisy speech
US6574597B1 (en) 1998-05-08 2003-06-03 At&T Corp. Fully expanded context-dependent networks for speech recognition
JP4169921B2 (ja) 2000-09-29 2008-10-22 パイオニア株式会社 音声認識システム
US6850887B2 (en) * 2001-02-28 2005-02-01 International Business Machines Corporation Speech recognition in noisy environments
US6678210B2 (en) 2001-08-28 2004-01-13 Rowe-Deines Instruments, Inc. Frequency division beamforming for sonar arrays
US20040024599A1 (en) 2002-07-31 2004-02-05 Intel Corporation Audio search conducted through statistical pattern matching
US20090030552A1 (en) * 2002-12-17 2009-01-29 Japan Science And Technology Agency Robotics visual and auditory system
JP3632099B2 (ja) * 2002-12-17 2005-03-23 独立行政法人科学技術振興機構 ロボット視聴覚システム
EP1691344B1 (en) * 2003-11-12 2009-06-24 HONDA MOTOR CO., Ltd. Speech recognition system
ES2443033T3 (es) * 2005-06-08 2014-02-17 Massachusetts Institute Of Technology Supervisión continua de poblaciones y comportamiento de peces a escala de plataforma continental
JP4234746B2 (ja) 2006-09-25 2009-03-04 株式会社東芝 音響信号処理装置、音響信号処理方法及び音響信号処理プログラム
US8275615B2 (en) * 2007-07-13 2012-09-25 International Business Machines Corporation Model weighting, selection and hypotheses combination for automatic speech recognition and machine translation
US20100217590A1 (en) * 2009-02-24 2010-08-26 Broadcom Corporation Speaker localization system and method
US8548807B2 (en) 2009-06-09 2013-10-01 At&T Intellectual Property I, L.P. System and method for adapting automatic speech recognition pronunciation by acoustic model restructuring
WO2011055410A1 (ja) 2009-11-06 2011-05-12 株式会社 東芝 音声認識装置
NO334170B1 (no) * 2011-05-16 2013-12-30 Radionor Comm As Fremgangsmåte og system for langdistanse, adaptivt, mobilt, stråleformende adhoc-kommunikasjonssystem med integrert posisjonering
US9881616B2 (en) * 2012-06-06 2018-01-30 Qualcomm Incorporated Method and systems having improved speech recognition
US9076450B1 (en) * 2012-09-21 2015-07-07 Amazon Technologies, Inc. Directed audio for speech recognition
US9131041B2 (en) * 2012-10-19 2015-09-08 Blackberry Limited Using an auxiliary device sensor to facilitate disambiguation of detected acoustic environment changes
US9653070B2 (en) 2012-12-31 2017-05-16 Intel Corporation Flexible architecture for acoustic signal processing engine
US10475440B2 (en) * 2013-02-14 2019-11-12 Sony Corporation Voice segment detection for extraction of sound source
US9286897B2 (en) * 2013-09-27 2016-03-15 Amazon Technologies, Inc. Speech recognizer with multi-directional decoding
US20150161999A1 (en) * 2013-12-09 2015-06-11 Ravi Kalluri Media content consumption with individualized acoustic speech recognition
US9443516B2 (en) 2014-01-09 2016-09-13 Honeywell International Inc. Far-field speech recognition systems and methods
US20160034811A1 (en) * 2014-07-31 2016-02-04 Apple Inc. Efficient generation of complementary acoustic models for performing automatic speech recognition system combination
US9299347B1 (en) 2014-10-22 2016-03-29 Google Inc. Speech recognition using associative mapping
KR102351366B1 (ko) * 2015-01-26 2022-01-14 삼성전자주식회사 음성 인식 방법 및 장치
KR101658001B1 (ko) * 2015-03-18 2016-09-21 서강대학교산학협력단 강인한 음성 인식을 위한 실시간 타겟 음성 분리 방법
US9697826B2 (en) * 2015-03-27 2017-07-04 Google Inc. Processing multi-channel audio waveforms
US20190147852A1 (en) * 2015-07-26 2019-05-16 Vocalzoom Systems Ltd. Signal processing and source separation
CN105161092B (zh) * 2015-09-17 2017-03-01 百度在线网络技术(北京)有限公司 一种语音识别方法和装置
US9980055B2 (en) * 2015-10-12 2018-05-22 Oticon A/S Hearing device and a hearing system configured to localize a sound source
US10229672B1 (en) * 2015-12-31 2019-03-12 Google Llc Training acoustic models using connectionist temporal classification
WO2017138934A1 (en) * 2016-02-10 2017-08-17 Nuance Communications, Inc. Techniques for spatially selective wake-up word recognition and related systems and methods
WO2017164954A1 (en) * 2016-03-23 2017-09-28 Google Inc. Adaptive audio enhancement for multichannel speech recognition
US10170134B2 (en) * 2017-02-21 2019-01-01 Intel IP Corporation Method and system of acoustic dereverberation factoring the actual non-ideal acoustic environment
US10499139B2 (en) * 2017-03-20 2019-12-03 Bose Corporation Audio signal processing for noise reduction
CN107030691B (zh) * 2017-03-24 2020-04-14 华为技术有限公司 一种看护机器人的数据处理方法及装置
US10297267B2 (en) * 2017-05-15 2019-05-21 Cirrus Logic, Inc. Dual microphone voice processing for headsets with variable microphone array orientation
CN108877827B (zh) * 2017-05-15 2021-04-20 福州瑞芯微电子股份有限公司 一种语音增强交互方法及系统、存储介质及电子设备
US10943583B1 (en) * 2017-07-20 2021-03-09 Amazon Technologies, Inc. Creation of language models for speech recognition
CN109686378B (zh) * 2017-10-13 2021-06-08 华为技术有限公司 语音处理方法和终端
WO2019104681A1 (zh) * 2017-11-30 2019-06-06 深圳市大疆创新科技有限公司 拍摄方法和装置
EP3729833B1 (en) * 2017-12-20 2025-02-26 Harman International Industries, Incorporated Virtual test environment for active noise management systems
CN110047478B (zh) * 2018-01-16 2021-06-08 中国科学院声学研究所 基于空间特征补偿的多通道语音识别声学建模方法及装置
US10867610B2 (en) * 2018-05-04 2020-12-15 Microsoft Technology Licensing, Llc Computerized intelligent assistant for conferences
US20190341053A1 (en) * 2018-05-06 2019-11-07 Microsoft Technology Licensing, Llc Multi-modal speech attribution among n speakers
CN110364166B (zh) * 2018-06-28 2022-10-28 腾讯科技(深圳)有限公司 实现语音信号识别的电子设备
CN108922553B (zh) * 2018-07-19 2020-10-09 苏州思必驰信息科技有限公司 用于音箱设备的波达方向估计方法及系统
US10349172B1 (en) * 2018-08-08 2019-07-09 Fortemedia, Inc. Microphone apparatus and method of adjusting directivity thereof
CN112292870A (zh) * 2018-08-14 2021-01-29 阿里巴巴集团控股有限公司 音频信号处理装置及方法
US10622004B1 (en) * 2018-08-20 2020-04-14 Amazon Technologies, Inc. Acoustic echo cancellation using loudspeaker position
US10937443B2 (en) * 2018-09-04 2021-03-02 Babblelabs Llc Data driven radio enhancement
US11574628B1 (en) * 2018-09-27 2023-02-07 Amazon Technologies, Inc. Deep multi-channel acoustic modeling using multiple microphone array geometries
CN109147787A (zh) * 2018-09-30 2019-01-04 深圳北极鸥半导体有限公司 一种智能电视声控识别系统及其识别方法
US10971158B1 (en) * 2018-10-05 2021-04-06 Facebook, Inc. Designating assistants in multi-assistant environment based on identified wake word received from a user
US11043214B1 (en) * 2018-11-29 2021-06-22 Amazon Technologies, Inc. Speech recognition using dialog history
US11170761B2 (en) * 2018-12-04 2021-11-09 Sorenson Ip Holdings, Llc Training of speech recognition systems

Patent Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN102271299A (zh) * 2010-06-01 2011-12-07 索尼公司 声音信号处理装置和声音信号处理方法
CN107742522A (zh) * 2017-10-23 2018-02-27 科大讯飞股份有限公司 基于麦克风阵列的目标语音获取方法及装置
CN109272989A (zh) * 2018-08-29 2019-01-25 北京京东尚科信息技术有限公司 语音唤醒方法、装置和计算机可读存储介质

Also Published As

Publication number Publication date
US20220028404A1 (en) 2022-01-27
CN111627425B (zh) 2023-11-28
CN111627425A (zh) 2020-09-04
US12315527B2 (en) 2025-05-27

Similar Documents

Publication Publication Date Title
JP7434137B2 (ja) 音声認識方法、装置、機器及びコンピュータ読み取り可能な記憶媒体
Diaz-Guerra et al. Robust sound source tracking using SRP-PHAT and 3D convolutional neural networks
Schwartz et al. Multi-microphone speech dereverberation and noise reduction using relative early transfer functions
CN106023996B (zh) 基于十字形声阵列宽带波束形成的声识别方法
CN109509465B (zh) 语音信号的处理方法、组件、设备及介质
TW202008352A (zh) 方位角估計的方法、設備、語音交互系統及儲存介質
CN110556103A (zh) 音频信号处理方法、装置、系统、设备和存储介质
WO2022121184A1 (zh) 声音事件检测与定位方法、装置、设备及可读存储介质
CN114830686B (zh) 声源的改进定位
US12148441B2 (en) Source separation for automatic speech recognition (ASR)
WO2021082547A1 (zh) 语音信号处理方法、声音采集装置和电子设备
WO2016119388A1 (zh) 一种基于语音信号构造聚焦协方差矩阵的方法及装置
JP2023550434A (ja) 改良型音響源測位法
Phapatanaburi et al. Noise robust voice activity detection using joint phase and magnitude based feature enhancement
CN117437930A (zh) 用于多通道语音信号的处理方法、装置、设备和存储介质
Diaz-Guerra et al. Direction of arrival estimation with microphone arrays using SRP-PHAT and neural networks
Barfuss et al. Robust coherence-based spectral enhancement for speech recognition in adverse real-world environments
CN117121104A (zh) 估计用于处理所获取的声音数据的优化掩模
CN115128544A (zh) 一种基于麦克风线性双阵列的声源定位方法、装置及介质
WO2020164397A1 (zh) 一种语音识别方法及系统
CN115802245B (zh) 一种自适应麦克风阵列分离增强方法及系统
US20240249741A1 (en) Guided Speech Enhancement Network
Dehghan Firoozabadi et al. A novel nested circular microphone array and subband processing-based system for counting and DOA estimation of multiple simultaneous speakers
CN117037836A (zh) 基于信号协方差矩阵重构的实时声源分离方法和装置
CN111785282A (zh) 一种语音识别方法及装置和智能音箱

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 20756536

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 20756536

Country of ref document: EP

Kind code of ref document: A1

WWG Wipo information: grant in national office

Ref document number: 17428015

Country of ref document: US