WO2021189979A1 - 语音增强方法、装置、计算机设备及存储介质 - Google Patents
语音增强方法、装置、计算机设备及存储介质 Download PDFInfo
- Publication number
- WO2021189979A1 WO2021189979A1 PCT/CN2020/136364 CN2020136364W WO2021189979A1 WO 2021189979 A1 WO2021189979 A1 WO 2021189979A1 CN 2020136364 W CN2020136364 W CN 2020136364W WO 2021189979 A1 WO2021189979 A1 WO 2021189979A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- voice
- speech
- voice data
- enhancement
- different environments
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/20—Speech recognition techniques specially adapted for robustness in adverse environments, e.g. in noise, of stress induced speech
Definitions
- This application relates to the field of artificial intelligence technology, in particular to a voice enhancement method, device, computer equipment, and storage medium.
- the parameters of the speech enhancement module are usually adjusted according to the surrounding environment and expert experience in order to achieve a better speech recognition effect.
- this method of adjusting speech enhancement parameters based on expert experience can only adapt to the surrounding environment to a certain extent and improve the effect of high speech recognition, but it cannot guarantee that the accuracy of speech recognition reaches the highest rate.
- This application provides a voice enhancement method, device, computer equipment, and storage medium, mainly in that it can automatically select a voice enhancement parameter matching the surrounding environment from a pre-built voice enhancement parameter set, and use the voice enhancement parameter to recognize the voice data After the speech enhancement processing, the accuracy of speech recognition can be maximized, so that the optimal speech recognition effect can be achieved in any environment.
- a speech enhancement method including:
- the speech enhancement parameter set includes speech enhancement parameters in different environments, and the speech enhancement parameters are used to enhance the accuracy of speech recognition in different environments;
- the target voice enhancement parameter perform voice enhancement processing on the voice data to obtain voice data after the voice enhancement processing.
- a voice enhancement device including:
- the acquiring unit is used to acquire the voice data to be processed
- the selecting unit is configured to extract the first voice feature corresponding to the voice data, determine the target environment in which the voice data is located according to the first voice feature, and select the corresponding target environment from a set of pre-built voice enhancement parameters
- the target speech enhancement parameters for the speech enhancement parameter set include speech enhancement parameters in different environments, and the speech enhancement parameters are used to enhance the accuracy of speech recognition in different environments;
- the processing unit is configured to perform voice enhancement processing on the voice data according to the target voice enhancement parameters to obtain voice data after the voice enhancement processing.
- a computer-readable storage medium on which a computer program is stored, and when the program is executed by a processor, the following steps are implemented:
- the speech enhancement parameter set includes speech enhancement parameters in different environments, and the speech enhancement parameters are used to enhance the accuracy of speech recognition in different environments;
- the target voice enhancement parameter perform voice enhancement processing on the voice data to obtain voice data after the voice enhancement processing.
- a computer device including a memory, a processor, and a computer program stored in the memory and capable of running on the processor, and the processor implements the following steps when the program is executed:
- the speech enhancement parameter set includes speech enhancement parameters in different environments, and the speech enhancement parameters are used to enhance the accuracy of speech recognition in different environments;
- the target voice enhancement parameter perform voice enhancement processing on the voice data to obtain voice data after the voice enhancement processing.
- the voice enhancement method, device, computer equipment, and storage medium provided in this application are compared with the current way of adjusting the parameters of the voice enhancement module based on expert experience.
- This application can obtain the voice data to be processed;
- the first voice feature corresponding to the voice data determines the target environment in which the voice data is located according to the first voice feature, and selects the target voice enhancement parameter corresponding to the target environment from a pre-built set of voice enhancement parameters, the
- the speech enhancement parameter set includes speech enhancement parameters in different environments, and the speech enhancement parameters are used to enhance the accuracy of speech recognition in different environments; then, according to the target speech enhancement parameters, speech enhancement processing is performed on the speech data, Obtain the voice data after the voice enhancement process.
- the corresponding target voice enhancement parameter can be automatically selected from the voice enhancement parameter set, and the target voice enhancement parameter can be used to perform the voice data.
- Speech enhancement processing can not only improve the effect of speech enhancement in the target environment, but also ensure the highest accuracy of speech recognition in the target environment.
- Fig. 1 shows a flowchart of a voice enhancement method provided by an embodiment of the present application
- FIG. 2 shows a flowchart of another voice enhancement method provided by an embodiment of the present application
- FIG. 3 shows a schematic structural diagram of a voice enhancement device provided by an embodiment of the present application
- FIG. 4 shows a schematic structural diagram of another voice enhancement device provided by an embodiment of the present application.
- Fig. 5 shows a schematic diagram of the physical structure of a computer device provided by an embodiment of the present application.
- the parameters of the speech enhancement module are usually adjusted according to the surrounding environment and expert experience in order to achieve a better speech recognition effect.
- this method of adjusting the speech enhancement parameters based on expert experience can only adapt to the surrounding environment to a certain extent and improve the effect of high speech recognition, but it cannot guarantee that the accuracy of speech recognition reaches the highest rate.
- an embodiment of the present application provides a credit risk assessment method. As shown in FIG. 1, the method includes:
- the voice data to be processed may be voice sequences collected in different environments, for example, a voice sequence of a certain user collected on the side of a street, or a voice sequence of a certain user collected in a factory.
- the embodiment of the present application constructs a speech enhancement parameter set in advance, and automatically selects the speech enhancement parameter set according to the target environment in which the speech data to be processed is located. Selecting matching speech enhancement parameters can not only improve the speech enhancement effect of speech data in any environment, but also maximize the accuracy of speech recognition.
- the embodiments of the present application are applicable to voice enhancement processing of voice data.
- the execution subject of the embodiments of the present application is a device or device capable of performing voice enhancement processing on voice data, which can be specifically set on the client or server side.
- the voice data needs to be pre-processed, including pre-emphasis processing, framing processing, and windowing function processing.
- the pre-processed voice data is obtained. Furthermore, it is necessary to determine the target environment where the pre-processed voice data is located, and perform voice enhancement processing on the voice data based on the target environment where the voice data is located.
- Extract the first voice feature corresponding to the voice data determine the target environment in which the voice data is located according to the first voice feature, and select the target voice corresponding to the target environment from a set of pre-built voice enhancement parameters Enhanced parameters.
- the speech enhancement parameter set includes speech enhancement parameters in different environments, and the speech enhancement parameters are used to enhance the accuracy of speech recognition in different environments.
- the sample voice data collected in different environments are stored in the preset sample library.
- the sample voice data needs to be clustered to obtain samples in different environments Voice data, and use sample voice data in different environments to train the voice enhancement model, that is, optimize and adjust the initial voice enhancement parameters in the voice enhancement model until the sample voice data after the voice enhancement process is input into the pre-built voice
- speech recognition is performed in the recognition model, the accuracy of speech recognition of speech data can be maximized, so that speech enhancement parameters in different environments can be obtained, and speech enhancement parameter sets can be constructed.
- use and The voice enhancement parameters corresponding to the environment perform voice enhancement processing on the voice data, and input the voice data after the voice enhancement processing into the pre-built voice enhancement model, which can maximize the accuracy of voice recognition of the voice data.
- the target environment where the voice data to be processed is located, specifically, extract the first voice feature corresponding to the voice data to be processed, and extract different clusters respectively.
- the second voice features corresponding to the sample voice data in different cluster categories are then calculated according to the second voice features corresponding to the sample voice data in different cluster categories.
- the voice features corresponding to the voice data collected in the same environment are relatively similar. Therefore, by calculating the distance between the first voice feature and different feature centers, it is determined which cluster category the voice data to be processed should be classified into. , And then be able to determine the target environment where the voice data to be processed is located.
- the target enhancement parameter corresponding to the target environment is selected from the pre-built voice enhancement parameter set, so as to use the target voice enhancement parameter to perform voice enhancement processing on the voice data, and input the voice data after the voice enhancement processing into the pre-built voice
- the speech recognition in the recognition model can maximize the speech recognition efficiency of the speech data, which can determine the target environment of the speech data according to the speech characteristics of the speech data to be processed, and then automatically select the target environment from the speech enhancement parameter set
- the corresponding speech enhancement parameters perform speech enhancement processing on the speech data, which improves the effect of speech enhancement and at the same time ensures that the speech recognition accuracy rate of the speech data after the speech enhancement processing reaches the highest.
- the speech enhancement processing mainly refers to the noise reduction processing of the speech noise in the speech data to be processed.
- VAD Voice Endpoint Detection Al
- the present application can obtain the voice data to be processed; at the same time, extract the first voice data corresponding to the voice data.
- the voice feature determines the target environment in which the voice data is located according to the first voice feature, and selects the target voice enhancement parameter corresponding to the target environment from a pre-built voice enhancement parameter set, and the voice enhancement parameter set includes Speech enhancement parameters in different environments, where the speech enhancement parameters are used to enhance the accuracy of speech recognition in different environments; then according to the target speech enhancement parameters, the speech data is subjected to speech enhancement processing to obtain the speech enhancement processing.
- the voice data by determining the target environment of the voice data to be processed, can automatically select the corresponding target voice enhancement parameter from the voice enhancement parameter set, and use the target voice enhancement parameter to perform voice enhancement processing on the voice data. Improve the speech enhancement effect in the target environment, while also ensuring that the accuracy of speech recognition in the target environment reaches the highest rate.
- an embodiment of the present application provides another voice enhancement method, as shown in FIG. 2.
- the method includes:
- the method includes: using initial speech enhancement parameters to perform speech enhancement processing on sample speech data in different environments to obtain sample speech data after speech enhancement processing in different environments; According to the data, the speech recognition accuracy function under different environments is constructed; according to the accuracy function, the initial speech enhancement parameters are optimized and adjusted to obtain speech enhancement parameters in different environments, and based on the speech enhancement in different environments Parameters to construct the speech enhancement parameter set.
- the constructing a speech recognition accuracy rate function in different environments according to the sample speech data includes: using a pre-built speech recognition model to perform speech recognition on the speech enhancement processed sample speech data to obtain different environments According to the results of speech recognition in different environments, construct a function of the accuracy of speech recognition in different environments.
- the pre-built speech recognition model may specifically be a neural network speech recognition model.
- the initial speech enhancement is given first, and then the initial speech enhancement parameters are used to perform speech enhancement processing on the sample speech data in the factory environment to obtain the sample speech data after the speech enhancement processing in the factory environment.
- the sample voice data is input to the pre-built voice recognition model for voice recognition processing, and the voice recognition results corresponding to the sample voice data in the factory environment are obtained, and then based on the voice recognition results in the factory environment, the accuracy of speech recognition in the factory environment is constructed Function, the function is solved under the condition of the highest speech recognition accuracy.
- the genetic algorithm can be used to search for speech enhancement parameters in different environments.
- the specific formula is:
- T( ⁇ ) is the speech recognition accuracy rate in the factory environment
- ⁇ i is the speech enhancement parameter in the factory environment.
- the speech enhancement parameters ⁇ i and the speech enhancement parameters ⁇ can be obtained. i can maximize the accuracy of speech recognition in the factory environment, so that the speech enhancement parameters in different environments can be obtained according to the above method, and the speech enhancement parameter set ⁇ i ⁇ can be constructed, so as to make the accuracy of speech recognition in different environments Reach the highest.
- the speech data to be processed can be obtained, and the target environment of the speech data to be processed can be determined, and the corresponding speech enhancement parameter can be selected from the speech enhancement parameter set. Perform voice enhancement processing.
- Extract the first voice feature corresponding to the voice data determine the target environment in which the voice data is located according to the first voice feature, and select the target voice corresponding to the target environment from a pre-built voice enhancement parameter set Enhanced parameters.
- the speech enhancement parameter set includes speech enhancement parameters in different environments, and the speech enhancement parameters are used to maximize the accuracy of speech recognition in different environments.
- step 202 specifically includes: obtaining sample voice data in different environments, and extracting the second voice feature corresponding to the sample voice data; 2. Voice features, calculating feature centers corresponding to the sample voice data in the different environments; determining the target environment where the voice data is located according to the feature centers and the first voice feature.
- the determining the target environment where the voice data is located according to the feature center and the first voice feature includes: using a preset Euclidean distance algorithm to calculate the difference between the first voice feature and the different feature centers Euclidean distance between the two; filter out the minimum Euclidean distance from the calculated Euclidean distance, and determine the environment in which the sample voice data corresponding to the minimum Euclidean distance is located as the target environment.
- the preset Mel cepstrum algorithm when extracting the voice features corresponding to the voice data to be processed and the sample voice data, can be used to calculate the Mel cepstrum coefficients corresponding to the sample data to be processed and the sample voice data, and calculate The Mel cepstrum coefficient is determined as the voice feature corresponding to the voice data to be processed and the sample voice data respectively.
- the feature center corresponding to the sample voice data in the street is calculated as A
- the feature center corresponding to the sample voice data in the factory environment is B
- the feature center corresponding to the sample voice data in the airport environment is C.
- the voice data in the same environment The corresponding voice features are similar, and then the Euclidean distances between the first voice feature and feature center A, feature center B and feature center C corresponding to the voice data to be processed are respectively calculated, and the smallest Euclidean distance is selected from the calculated Euclidean distances. Distance. If it is determined that the Euclidean distance between the feature center B and the first voice feature is the smallest, it is determined that the voice data to be processed is relatively close to the sample voice data in the factory environment. Therefore, it is determined that the voice data to be processed is in the factory environment. According to the above method, the target environment of the voice data to be processed can be determined.
- step 203 specifically includes: performing filtering and noise reduction processing on the voice data according to the target filtering noise reduction parameter to obtain the voice data after noise reduction processing.
- the specific method of using the target filter noise reduction parameter to perform noise reduction processing on the voice data is exactly the same as that of step 103, and will not be repeated here.
- a pre-built voice recognition model can be used for voice recognition.
- the voice recognition model may specifically be a neural network voice recognition model.
- the voice data after the voice enhancement process is input to the voice recognition model, and the hidden layer in the voice recognition model can extract the third voice feature corresponding to the voice data, The voice recognition is performed according to the third voice feature to obtain the voice recognition result corresponding to the voice data. At this time, the accuracy of the voice recognition result can reach the highest.
- the another voice enhancement method can obtain the voice data to be processed; at the same time extract the first voice data corresponding to the voice data.
- a voice feature, the target environment where the voice data is located is determined according to the first voice feature, and the target voice enhancement parameter corresponding to the target environment is selected from a pre-built voice enhancement parameter set, the voice enhancement parameter set includes There are speech enhancement parameters in different environments, and the speech enhancement parameters are used to enhance the accuracy of speech recognition in different environments; then, according to the target speech enhancement parameters, the speech data is subjected to speech enhancement processing, and the speech enhancement processing is obtained Therefore, by determining the target environment of the voice data to be processed, the corresponding target voice enhancement parameter can be automatically selected from the voice enhancement parameter set, and the target voice enhancement parameter is used to perform voice enhancement processing on the voice data. It can improve the voice enhancement effect in the target environment, and at the same time can ensure that the accuracy of speech recognition in the target environment reaches the highest
- an embodiment of the present application provides a speech enhancement device.
- the device includes: an acquisition unit 31, a selection unit 32, and a processing unit 33.
- the acquiring unit 31 may be used to acquire voice data to be processed.
- the acquiring unit 31 is the main functional module for acquiring the voice data to be processed in the device.
- the selection unit 32 may be used to extract the first voice feature corresponding to the voice data, determine the target environment in which the voice data is located according to the first voice feature, and select all the voice enhancement parameters from a set of pre-built voice enhancement parameters.
- the target speech enhancement parameters corresponding to the target environment, the speech enhancement parameter set includes speech enhancement parameters in different environments, and the speech enhancement parameters are used to maximize the accuracy of speech recognition in different environments.
- the selection unit 32 extracts the first voice feature corresponding to the voice data in the device, determines the target environment in which the voice data is located according to the first voice feature, and selects all the voice enhancement parameters from a pre-built set of voice enhancement parameters.
- the main function module of the target speech enhancement parameter corresponding to the target environment is also the core module.
- the processing unit 33 may be configured to perform voice enhancement processing on the voice data according to the target voice enhancement parameters to obtain voice data after the voice enhancement processing.
- the processing unit 33 is a main functional module of the device that performs voice enhancement processing on the voice data according to the target voice enhancement parameters to obtain voice data after the voice enhancement processing.
- the selection unit 32 includes an extraction module 321, a calculation module 322, and a determination module 323.
- the extraction module 321 may be used to obtain sample voice data in different environments, and extract the second voice feature corresponding to the sample voice data.
- the calculation module 322 may be used to calculate feature centers corresponding to the sample voice data in the different environments according to the second voice feature.
- the determining module 323 may be used to determine the target environment where the voice data is located according to the feature center and the first voice feature.
- the determination module 323 includes: a calculation sub-module and a determination sub-module.
- the calculation sub-module may be used to calculate the Euclidean distance between the first voice feature and different feature centers by using a preset Euclidean distance algorithm.
- the determining submodule may be used to filter out the minimum Euclidean distance from the calculated Euclidean distance, and determine the environment in which the sample voice data corresponding to the minimum Euclidean distance is located as the target environment.
- the device further includes: a constructing unit 34.
- the processing unit 33 may also be configured to perform voice enhancement processing on the sample voice data in the different environments by using the initial voice enhancement parameters to obtain the sample voice data after the voice enhancement processing in the different environments.
- the construction unit 34 may be used to construct a speech recognition accuracy rate function in different environments according to the sample speech data.
- the construction unit 34 may also be used to optimize and adjust the initial speech enhancement parameters according to the accuracy function to obtain speech enhancement parameters in different environments, and construct based on the speech enhancement parameters in different environments The speech enhancement parameter set.
- the construction unit 34 includes: a recognition module 341 and a construction module 342.
- the recognition module 341 may be used to perform voice recognition on the sample voice data after the voice enhancement processing by using a pre-built voice recognition model to obtain voice recognition results in different environments.
- the construction module 342 may be used to construct a speech recognition accuracy rate function in different environments according to the speech recognition results in the different environments.
- the device further includes: an extracting unit 35 and a determining unit 36.
- the extraction unit 35 may be used to perform feature extraction on the voice data processed by the voice enhancement process to obtain a third voice feature corresponding to the voice data.
- the determining unit 36 may be configured to determine a voice recognition result corresponding to the voice data according to the third voice feature.
- the processing unit 33 may be specifically configured to perform filtering and noise reduction processing on the voice data according to the target filtering noise reduction parameter to obtain the voice data after noise reduction processing.
- an embodiment of the present application also provides a computer-readable storage medium.
- the above-mentioned storage medium may be a non-volatile storage medium or a volatile storage medium.
- a computer program is stored thereon, and when the program is executed by the processor, the following steps are realized: acquiring voice data to be processed; extracting the first voice feature corresponding to the voice data, and determining the location of the voice data according to the first voice feature And select the target speech enhancement parameters corresponding to the target environment from a pre-built speech enhancement parameter set.
- the speech enhancement parameter set contains speech enhancement parameters in different environments, and the speech enhancement parameters are used for enhancement The accuracy of speech recognition in different environments; according to the target speech enhancement parameters, perform speech enhancement processing on the speech data to obtain speech data after the speech enhancement processing.
- the computer device includes: a processor 41, The memory 42 and a computer program that is stored on the memory 42 and can run on the processor, wherein the memory 42 and the processor 41 are both set on the bus 43, the processor 41 implements the following steps when the program is executed: The voice data; extract the first voice feature corresponding to the voice data, determine the target environment in which the voice data is located according to the first voice feature, and select the target environment corresponding to the set of pre-built voice enhancement parameters
- Target speech enhancement parameters the speech enhancement parameter set includes speech enhancement parameters in different environments, the speech enhancement parameters are used to enhance the accuracy of speech recognition in different environments; according to the target speech enhancement parameters, the speech The data undergoes speech enhancement processing to obtain speech data after the speech enhancement processing.
- the present application can obtain the voice data to be processed; at the same time, extract the first voice feature corresponding to the voice data, determine the target environment where the voice data is located according to the first voice feature, and obtain information from The pre-built speech enhancement parameter sets select the target speech enhancement parameters corresponding to the target environment.
- the speech enhancement parameter set contains speech enhancement parameters in different environments, and the speech enhancement parameters are used to enhance the accuracy of speech recognition in different environments.
- the voice data is voice-enhanced to obtain voice-enhanced voice data, which can be automatically enhanced from the voice by determining the target environment where the voice data to be processed is located
- the target voice enhancement parameters corresponding to the parameters are selected in a centralized manner, and the target voice enhancement parameters are used to perform voice enhancement processing on the voice data, which can not only improve the voice enhancement effect in the target environment, but also ensure the highest accuracy rate of speech recognition in the target environment .
- modules or steps of this application can be implemented by a general computing device, and they can be concentrated on a single computing device or distributed in a network composed of multiple computing devices.
- they can be implemented with program codes executable by the computing device, so that they can be stored in the storage device for execution by the computing device, and in some cases, can be executed in a different order than here.
Landscapes
- Engineering & Computer Science (AREA)
- Computational Linguistics (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Human Computer Interaction (AREA)
- Physics & Mathematics (AREA)
- Acoustics & Sound (AREA)
- Multimedia (AREA)
- Navigation (AREA)
Abstract
一种语音增强方法、装置、计算机设备及存储介质,涉及人工智能技术领域,适用于语音数据的语音增强处理。能够自动从预先构建的语音增强参数集中选择与周围环境相匹配的语音增强参数,利用该语音增强参数对待识别语音数据进行语音增强处理后,能够使语音识别准确率达到最高。方法包括:获取待处理的语音数据(101);提取语音数据对应的第一语音特征,根据第一语音特征确定语音数据所处的目标环境,并从预先构建的语音增强参数集中选取目标环境对应的目标语音增强参数,语音增强参数用于增强不同环境下的语音识别准确率(102);根据目标语音增强参数,对语音数据进行语音增强处理,得到语音增强处理后的语音数据(103)。
Description
本申请要求于2020年10月26日提交中国专利局、申请号为202011153521.2,发明名称为“语音增强方法、装置、计算机设备及存储介质”的中国专利申请的优先权,其全部内容通过引用结合在本申请中。
本申请涉及人工智能技术领域,尤其是涉及一种语音增强方法、装置、计算机设备及存储介质。
近年来,随着智能型穿戴设备的快速发展和崛起,通过语音控制的消费类电子产品已成为最新潮流,语音智能需要可靠性强、准确率高的自动语音识别智能系统作为支撑,而前端语音增强技术就是最关键的一环。
目前,在利用前端语音增强技术对噪声进行处理时,通常根据周围环境,依据专家经验对语音增强模块的参数进行调整,以期达到较好的语音识别效果。然而,发明人意识到,这种依据专家经验对语音增强参数进行调整的方式,只能一定程度地适应周围环境,改善高语音识别的效果,但是无法保证语音识别的正确率均达到最高。
本申请提供了一种语音增强方法、装置、计算机设备及存储介质,主要在于能够自动从预先构建的语音增强参数集中选择与周围环境相匹配的语音增强参数,利用该语音增强参数对待识别语音数据进行语音增强处理后,能够使语音识别准确率达到最高,从而能够在任何环境中达到最优的语音识别效果。
根据本申请的第一个方面,提供一种语音增强方法,包括:
获取待处理的语音数据;
提取所述语音数据对应的第一语音特征,根据所述第一语音特征确定所述语音数据所处的目标环境,并从预先构建的语音增强参数集中选取所述目标环境对应的目标语音增强参数,所述语音增强参数集中包含有不同环境下的语音增强参数,所述语音增强参数用于增强不同环境下的语音识别准确率;
根据所述目标语音增强参数,对所述语音数据进行语音增强处理,得到语音增强处理后的语音数据。
根据本申请的第二个方面,提供一种语音增强装置,包括:
获取单元,用于获取待处理的语音数据;
选取单元,用于提取所述语音数据对应的第一语音特征,根据所述第一语音特征确定所述语音数据所处的目标环境,并从预先构建的语音增强参数集中选取所述目标环境对应的目标语音增强参数,所述语音增强参数集中包含有不同环境下的语音增强参数,所述语音增强参数用于增强不同环境下的语音识别准确率;
处理单元,用于根据所述目标语音增强参数,对所述语音数据进行语音增强处理,得到语音增强处理后的语音数据。
根据本申请的第三个方面,提供一种计算机可读存储介质,其上存储有计算机程序,该程序被处理器执行时实现以下步骤:
获取待处理的语音数据;
提取所述语音数据对应的第一语音特征,根据所述第一语音特征确定所述语音数据所处的目标环境,并从预先构建的语音增强参数集中选取所述目标环境对应的目标语音增强 参数,所述语音增强参数集中包含有不同环境下的语音增强参数,所述语音增强参数用于增强不同环境下的语音识别准确率;
根据所述目标语音增强参数,对所述语音数据进行语音增强处理,得到语音增强处理后的语音数据。
根据本申请的第四个方面,提供一种计算机设备,包括存储器、处理器及存储在存储器上并可在处理器上运行的计算机程序,所述处理器执行所述程序时实现以下步骤:
获取待处理的语音数据;
提取所述语音数据对应的第一语音特征,根据所述第一语音特征确定所述语音数据所处的目标环境,并从预先构建的语音增强参数集中选取所述目标环境对应的目标语音增强参数,所述语音增强参数集中包含有不同环境下的语音增强参数,所述语音增强参数用于增强不同环境下的语音识别准确率;
根据所述目标语音增强参数,对所述语音数据进行语音增强处理,得到语音增强处理后的语音数据。
本申请提供的一种语音增强方法、装置、计算机设备及存储介质,与目前依据专家经验对语音增强模块的参数进行调整的方式相比,本申请能够获取待处理的语音数据;同时提取所述语音数据对应的第一语音特征,根据所述第一语音特征确定所述语音数据所处的目标环境,并从预先构建的语音增强参数集中选取所述目标环境对应的目标语音增强参数,所述语音增强参数集中包含有不同环境下的语音增强参数,所述语音增强参数用于增强不同环境下的语音识别准确率;之后根据所述目标语音增强参数,对所述语音数据进行语音增强处理,得到语音增强处理后的语音数据,由此通过确定待处理的语音数据所处的目标环境,能够自动从语音增强参数集中选取与其对应的目标语音增强参数,利用该目标语音增强参数对语音数据进行语音增强处理,不仅能够改善目标环境下的语音增强效果,同时还能够保证目标环境下语音识别的准确率达到最高。
此处所说明的附图用来提供对本申请的进一步理解,构成本申请的一部分,本申请的示意性实施例及其说明用于解释本申请,并不构成对本申请的不当限定。在附图中:
图1示出了本申请实施例提供的一种语音增强方法流程图;
图2示出了本申请实施例提供的另一种语音增强方法流程图;
图3示出了本申请实施例提供的一种语音增强装置的结构示意图;
图4示出了本申请实施例提供的另一种语音增强装置的结构示意图;
图5示出了本申请实施例提供的一种计算机设备的实体结构示意图。
本申请的最佳实施方式
下文中将参考附图并结合实施例来详细说明本申请。需要说明的是,在不冲突的情况下,本申请中的实施例及实施例中的特征可以相互组合。
目前,在利用前端语音增强技术对噪声进行处理时,通常根据周围环境,依据专家经验对语音增强模块的参数进行调整,以期达到较好的语音识别效果。然而,这种依据专家经验对语音增强参数进行调整的方式,只能一定程度地适应周围环境,改善高语音识别的效果,但是无法保证语音识别的正确率均达到最高。
为了解决上述问题,本申请实施例提供了一种信贷风险评估方法,如图1所示,所述方法包括:
101、获取待处理的语音数据。
其中,待处理的语音数据可以为在不同环境中采集到语音序列,例如,在街道旁采集到某用户的一段语音序列,或者在工厂中采集到某用户的一段语音序列,对于本申请实施 例,为了克服现有技术中依据专家经验对语音增强参数进行调整的缺陷,本申请实施例通过预先构建语音增强参数集,并根据待处理的语音数据所处的目标环境,自动从语音增强参数集中选取相匹配的语音增强参数,由此在任何环境中不仅能够改善语音数据的语音增强效果,同时还能够使语音识别准确率达到最高。本申请实施例适用于语音数据的语音增强处理,本申请实施例的执行主体为能够对语音数据进行语音增强处理的装置或者设备,具体可以设置于客户端或者服务器一侧。
具体地,获取用户在某场景下的一段语音数据,在对该语音数据进行语音增强处理之前,需要对该语音数据进行预处理,具体包括预加重处理、分帧处理和加窗函数处理,由此得到预处理后的语音数据,进一步地,需要确定预处理后的语音数据所处的目标环境,基于语音数据所处的目标环境对其进行语音增强处理。
102、提取所述语音数据对应的第一语音特征,根据所述第一语音特征确定所述语音数据所处的目标环境,并从预先构建的语音增强参数集中选取所述目标环境对应的目标语音增强参数。
其中,所述语音增强参数集中包含有不同环境下的语音增强参数,所述语音增强参数用于增强不同环境下的语音识别准确率。对于本申请实施例,预设样本库中存储有在不同环境下采集的样本语音数据,为了确定不同样本语音数据所处的环境,需要对样本语音数据进行聚类处理,得到不同环境下的样本语音数据,并利用不同环境下的样本语音数据对语音增强模型进行训练,即对语音增强模型中的初始语音增强参数进行优化调整,直至经过语音增强处理后的样本语音数据输入至预先构建的语音识别模型中进行语音识别时,能够使语音数据的语音识别准确率达到最高,由此能够得到不同环境下的语音增强参数,并构建语音增强参数集,当语音数据处于某一环境时,利用与该环境对应的语音增强参数对语音数据进行语音增强处理,并将语音增强处理后的语音数据输入至预先构建的语音增强模型,能够使语音数据的语音识别准确率达到最高。
对于本申请实施例,在对语音数据进行语音增强处理之前,需要确定待处理的语音数据所处的目标环境,具体地,提取待处理的语音数据对应的第一语音特征,同时分别提取不同聚类类别(不同环境)下的样本语音数据对应的第二语音特征,之后根据不同聚类类别下样本语音数据对应的第二语音特征,计算不同聚类类别下样本语音数据对应的特征中心,由于相同环境下采集的语音数据对应的语音特征较为相近,因此通过计算第一语音特征与不同特征中心之间的距离,确定待处理的语音数据应归类至哪一聚类类别下的样本语音数据,进而能够确定待处理语音数据所处的目标环境。
进一步地,从预先构建的语音增强参数集中选择目标环境对应的目标增强参数,以便利用该目标语音增强参数对语音数据进行语音增强处理,并将语音增强处理后的语音数据输入至预先构建的语音识别模型中进行语音识别,能够使语音数据的语音识别效率达到最高,由此能够根据待处理的语音数据的语音特征,确定语音数据所处的目标环境,进而自动从语音增强参数集中选择目标环境对应的语音增强参数,对语音数据进行语音增强处理,改善了语音增强效果,同时能够保证经过语音增强处理后的语音数据的语音识别准确率达到最高。
103、根据所述目标语音增强参数,对所述语音数据进行语音增强处理,得到语音增强处理后的语音数据。
对于本申请实施例,语音增强处理主要是指对待处理的语音数据中的语音噪声进行降噪处理,在语音增强处理的过程中可以采用LMS自适应滤波器降噪处理算法对语音数据进行语音增强处理,具体利用该算法进行语音增强处理时,首先通过语音端点检测算法(VAD)对语音信号进行静音剔除处理,得到合适的声音频谱特征序列X=(x
1,x
2,…,x
n),然后再经过多通道的维纳滤波操作,具体包括波束成形处理得到Y=(y
1,y
2,…,y
n),并利用功率谱密度(PSD)估计减少残余噪声分量,得到维纳滤波输入分量
和φ
V(ω,τ), 然后经过维纳滤波计算得到后置滤波器输入参数向量G
Wiener(ω,τ),经过后置滤波器处理得到滤波输出信号Z(ω,τ)=G
Wiener(ω,τ)*Y,再经过信号压缩或膨胀处理后,得到语音增强处理后的语音数据,由此经过语音增强处理后的语音数据能够适配语音识别模型的输入形式。
本申请实施例提供的一种语音增强方法,与目前依据专家经验对语音增强模块的参数进行调整的方式相比,本申请能够获取待处理的语音数据;同时提取所述语音数据对应的第一语音特征,根据所述第一语音特征确定所述语音数据所处的目标环境,并从预先构建的语音增强参数集中选取所述目标环境对应的目标语音增强参数,所述语音增强参数集中包含有不同环境下的语音增强参数,所述语音增强参数用于增强不同环境下的语音识别准确率;之后根据所述目标语音增强参数,对所述语音数据进行语音增强处理,得到语音增强处理后的语音数据,由此通过确定待处理的语音数据所处的目标环境,能够自动从语音增强参数集中选取与其对应的目标语音增强参数,利用该目标语音增强参数对语音数据进行语音增强处理,不仅能够改善目标环境下的语音增强效果,同时还能够保证目标环境下语音识别的准确率达到最高。
进一步的,为了更好的说明上述对语音数据进行语音增强处理的过程,作为对上述实施例的细化和扩展,本申请实施例提供了另一种语音增强方法方法,如图2所示,所述方法包括:
201、获取待处理的语音数据。
对于本申请实施例,为了能够根据待处理的语音数据所处的环境,自动选择与该环境相匹配的语音增强参数,而使语音数据的语音识别准确率达到最高,需要预先构建不同环境下的语音增强参数,基于此,所述方法包括:利用初始语音增强参数对所述不同环境下的样本语音数据进行语音增强处理,得到不同环境下语音增强处理后的样本语音数据;根据所述样本语音数据,构建不同环境下的语音识别准确率函数;根据所述准确率函数,对所述初始语音增强参数进行优化调整,得到不同环境下的语音增强参数,并基于所述不同环境下的语音增强参数,构建所述语音增强参数集。进一步地,所述根据所述样本语音数据,构建不同环境下的语音识别准确率函数,包括:利用预先构建的语音识别模型对所述语音增强处理后的样本语音数据进行语音识别,得到不同环境下的语音识别结果;根据所述不同环境下的语音识别结果,构建不同环境下的语音识别准确率函数。其中,预先构建的语音识别模型具体可以为神经网络语音识别模型。
例如,首先给定初始语音增强,之后利用该初始语音增强参数对工厂环境中的样本语音数据进行语音增强处理,得到在工厂环境下语音增强处理后的样本语音数据,并将该语音增强处理后的样本语音数据输入至预先构建的语音识别模型进行语音识别处理,得到工厂环境中样本语音数据对应的语音识别结果,接着根据该工厂环境中的语音识别结果,构建工厂环境下的语音识别准确率函数,在语音识别准确率最高的条件下求解该函数,具体搜寻最优解时可以利用遗传算法搜寻不同环境的语音增强参数,具体公式为:
θ
i=arg max T(θ)
其中,T(θ)为工厂环境下的语音识别准确率,θ
i为在工厂环境下的语音增强参数,通过不断对初始语音增强参数优化调整,能够得到语音增强参数θ
i,语音增强参数θ
i能够使工厂环境下的语音识别准确率达到最高,由此按照上述方式能够得到不同环境下的语音增强参数,并构建语音增强参数集{θ
i},进而使不同环境下的语音识别准确率达到最高。
对于本申请实施例,在构建完成语音增强参数集后,可以获取待处理的语音数据,并通过确定待处理的语音数据所处的目标环境,从语音增强参数集中选择相应的语音增强参数对其进行语音增强处理。
202、提取所述语音数据对应的第一语音特征,根据所述第一语音特征确定所述语音数据所处的目标环境,并从预先构建的语音增强参数集中选取所述目标环境对应的目标语音增强参数。
其中,所述语音增强参数集中包含有不同环境下的语音增强参数,所述语音增强参数用于最大化不同环境下的语音识别准确率。对于本申请实施例,为了确定待处理的语音数据所处的目标环境,步骤202具体包括:获取不同环境下样本语音数据,并提取所述样本语音数据对应的第二语音特征;根据所述第二语音特征,计算所述不同环境下样本语音数据对应的特征中心;根据所述特征中心和所述第一语音特征,确定所述语音数据所处的目标环境。进一步地,所述根据所述特征中心和所述第一语音特征,确定所述语音数据所处的目标环境,包括:利用预设的欧式距离算法计算所述第一语音特征与不同特征中心之间的欧式距离;从计算的欧式距离中筛选出最小欧式距离,并将所述最小欧式距离对应的样本语音数据所处环境确定为所述目标环境。其中,提取待处理的语音数据和样本语音数据对应的语音特征时,可以采用预设的梅尔倒谱算法计算待处理的样本数据和样本语音数据分别对应的梅尔倒谱系数,并将计算的梅尔倒谱系数确定为待处理的语音数据和样本语音数据分别对应的语音特征。
例如,计算得到街道旁的样本语音数据对应的特征中心为A,工厂环境下的样本语音数据对应的特征中心为B,机场环境下样本语音数据对应的特征中心为C,由于相同环境中语音数据对应的语音特征较为相似,之后分别计算待处理的语音数据对应的第一语音特征与特征中心A,特征中心B和特征中心C之间的欧式距离,并从计算的各个欧式距离中筛选最小欧式距离,如确定特征中心B与第一语音特征之间的欧式距离最小,则确定待处理的语音数据与工厂环境中的样本语音数据较为相近,因此确定待处理的语音数据处于工厂环境中,由此按照上述方式能够确定待处理的语音数据所处的目标环境。
203、根据所述目标语音增强参数,对所述语音数据进行语音增强处理,得到语音增强处理后的语音数据。
对于本申请实施例,为了对语音数据进行语音增强处理,步骤203具体包括:根据所述目标滤波降噪参数,对所述语音数据进行滤波降噪处理,得到降噪处理后的语音数据。具体利用目标滤波降噪参数对语音数据进行降噪处理的方式与步骤103完全相同,在此不再赘述。
204、对所述语音增强处理后的语音数据进行特征提取,得到所述语音数据对应的第三语音特征,并根据所述第三语音特征,确定所述语音数据对应的语音识别结果。
对于本方实施例,在对语音数据进行语音增强处理后,需要进一步对语音增强处理后的语音数据进行语音识别,具体对语音数据进行语音识别时,可以利用预先构建的语音识别模型进行语音识别,该语音识别模型具体可以为神经网络语音识别模型,具体地,将语音增强处理后的语音数据输入至语音识别模型,该语音识别模型中的隐藏层能够提取语音数据对应的第三语音特征,并根据该第三语音特征进行语音识别,从而得到语音数据对应的语音识别结果,此时该语音识别结果的准确率能够达到最高。
本申请实施例提供的另一种语音增强方法,与目前依据专家经验对语音增强模块的参数进行调整的方式相比,本申请能够获取待处理的语音数据;同时提取所述语音数据对应的第一语音特征,根据所述第一语音特征确定所述语音数据所处的目标环境,并从预先构建的语音增强参数集中选取所述目标环境对应的目标语音增强参数,所述语音增强参数集中包含有不同环境下的语音增强参数,所述语音增强参数用于增强不同环境下的语音识别准确率;之后根据所述目标语音增强参数,对所述语音数据进行语音增强处理,得到语音增强处理后的语音数据,由此通过确定待处理的语音数据所处的目标环境,能够自动从语音增强参数集中选取与其对应的目标语音增强参数,利用该目标语音增强参数对语音数据进行语音增强处理,不仅能够改善目标环境下的语音增强效果,同时还能够保证目标环境 下语音识别的准确率达到最高。
进一步地,作为图1的具体实现,本申请实施例提供了一种语音增强装置,如图3所示,所述装置包括:获取单元31、选取单元32和处理单元33。
所述获取单元31,可以用于获取待处理的语音数据。所述获取单元31是本装置中获取待处理的语音数据的主要功能模块。
所述选取单元32,可以用于提取所述语音数据对应的第一语音特征,根据所述第一语音特征确定所述语音数据所处的目标环境,并从预先构建的语音增强参数集中选取所述目标环境对应的目标语音增强参数,所述语音增强参数集中包含有不同环境下的语音增强参数,所述语音增强参数用于最大化不同环境下的语音识别准确率。所述选取单元32是本装置中提取所述语音数据对应的第一语音特征,根据所述第一语音特征确定所述语音数据所处的目标环境,并从预先构建的语音增强参数集中选取所述目标环境对应的目标语音增强参数的主要功能模块,也是核心模块。
所述处理单元33,可以用于根据所述目标语音增强参数,对所述语音数据进行语音增强处理,得到语音增强处理后的语音数据。所述处理单元33是本装置中根据所述目标语音增强参数,对所述语音数据进行语音增强处理,得到语音增强处理后的语音数据的主要功能模块。
进一步地,为了确定所述语音数据所处的目标环境,如图4所示,所述选取单元32,包括提取模块321、计算模块322和确定模块323。
所述提取模块321,可以用于获取不同环境下样本语音数据,并提取所述样本语音数据对应的第二语音特征。
所述计算模块322,可以用于根据所述第二语音特征,计算所述不同环境下样本语音数据对应的特征中心。
所述确定模块323,可以用于根据所述特征中心和所述第一语音特征,确定所述语音数据所处的目标环境。
进一步地,为了确定所述语音数据所处的目标环境,所述确定模块323,包括:计算子模块和确定子模块。
所述计算子模块,可以用于利用预设的欧式距离算法计算所述第一语音特征与不同特征中心之间的欧式距离。
所述确定子模块,可以用于从计算的欧式距离中筛选出最小欧式距离,并将所述最小欧式距离对应的样本语音数据所处环境确定为所述目标环境。
进一步地,为了构建语音增强参数集,所述装置还包括:构建单元34。
所述处理单元33,还可以用于利用初始语音增强参数对所述不同环境下的样本语音数据进行语音增强处理,得到不同环境下语音增强处理后的样本语音数据。
所述构建单元34,可以用于根据所述样本语音数据,构建不同环境下的语音识别准确率函数。
所述构建单元34,还可以用于根据所述准确率函数,对所述初始语音增强参数进行优化调整,得到不同环境下的语音增强参数,并基于所述不同环境下的语音增强参数,构建所述语音增强参数集。
进一步地,为了构建不同环境下的语音识别准确率函数,所述构建单元34,包括:识别模块341和构建模块342。
所述识别模块341,可以用于利用预先构建的语音识别模型对所述语音增强处理后的样本语音数据进行语音识别,得到不同环境下的语音识别结果。
所述构建模块342,可以用于根据所述不同环境下的语音识别结果,构建不同环境下的语音识别准确率函数。
进一步地,为了对语音数据进行语音识别,所述装置还包括:提取单元35和确定单 元36。
所述提取单元35,可以用于对所述语音增强处理后的语音数据进行特征提取,得到所述语音数据对应的第三语音特征。
所述确定单元36,可以用于根据所述第三语音特征,确定所述语音数据对应的语音识别结果。
进一步地,为了对语音数据进行语音增强处理,所述处理单元33,具体可以用于根据所述目标滤波降噪参数,对所述语音数据进行滤波降噪处理,得到降噪处理后的语音数据。
需要说明的是,本申请实施例提供的一种语音增强装置所涉及各功能模块的其他相应描述,可以参考图1所示方法的对应描述,在此不再赘述。
基于上述如图1所示方法,相应的,本申请实施例还提供了一种计算机可读存储介质,上述存储介质可以是非易失性存储介质,也可以是易失性存储介质。其上存储有计算机程序,该程序被处理器执行时实现以下步骤:获取待处理的语音数据;提取所述语音数据对应的第一语音特征,根据所述第一语音特征确定所述语音数据所处的目标环境,并从预先构建的语音增强参数集中选取所述目标环境对应的目标语音增强参数,所述语音增强参数集中包含有不同环境下的语音增强参数,所述语音增强参数用于增强不同环境下的语音识别准确率;根据所述目标语音增强参数,对所述语音数据进行语音增强处理,得到语音增强处理后的语音数据。
基于上述如图1所示方法和如图3所示装置的实施例,本申请实施例还提供了一种计算机设备的实体结构图,如图5所示,该计算机设备包括:处理器41、存储器42、及存储在存储器42上并可在处理器上运行的计算机程序,其中存储器42和处理器41均设置在总线43上所述处理器41执行所述程序时实现以下步骤:获取待处理的语音数据;提取所述语音数据对应的第一语音特征,根据所述第一语音特征确定所述语音数据所处的目标环境,并从预先构建的语音增强参数集中选取所述目标环境对应的目标语音增强参数,所述语音增强参数集中包含有不同环境下的语音增强参数,所述语音增强参数用于增强不同环境下的语音识别准确率;根据所述目标语音增强参数,对所述语音数据进行语音增强处理,得到语音增强处理后的语音数据。
通过本申请的技术方案,本申请能够获取待处理的语音数据;同时提取所述语音数据对应的第一语音特征,根据所述第一语音特征确定所述语音数据所处的目标环境,并从预先构建的语音增强参数集中选取所述目标环境对应的目标语音增强参数,所述语音增强参数集中包含有不同环境下的语音增强参数,所述语音增强参数用于增强不同环境下的语音识别准确率;之后根据所述目标语音增强参数,对所述语音数据进行语音增强处理,得到语音增强处理后的语音数据,由此通过确定待处理的语音数据所处的目标环境,能够自动从语音增强参数集中选取与其对应的目标语音增强参数,利用该目标语音增强参数对语音数据进行语音增强处理,不仅能够改善目标环境下的语音增强效果,同时还能够保证目标环境下语音识别的准确率达到最高。
显然,本领域的技术人员应该明白,上述的本申请的各模块或各步骤可以用通用的计算装置来实现,它们可以集中在单个的计算装置上,或者分布在多个计算装置所组成的网络上,可选地,它们可以用计算装置可执行的程序代码来实现,从而,可以将它们存储在存储装置中由计算装置来执行,并且在某些情况下,可以以不同于此处的顺序执行所示出或描述的步骤,或者将它们分别制作成各个集成电路模块,或者将它们中的多个模块或步骤制作成单个集成电路模块来实现。这样,本申请不限制于任何特定的硬件和软件结合。
以上所述仅为本申请的优选实施例而已,并不用于限制本申请,对于本领域的技术人员来说,本申请可以有各种更改和变化。凡在本申请的精神和原则之内,所作的任何修改、等同替换、改进等,均应包括在本申请的保护范围之内。
Claims (20)
- 一种语音增强方法,其中,包括:获取待处理的语音数据;提取所述语音数据对应的第一语音特征,根据所述第一语音特征确定所述语音数据所处的目标环境,并从预先构建的语音增强参数集中选取所述目标环境对应的目标语音增强参数,所述语音增强参数集中包含有不同环境下的语音增强参数,所述语音增强参数用于增强不同环境下的语音识别准确率;根据所述目标语音增强参数,对所述语音数据进行语音增强处理,得到语音增强处理后的语音数据。
- 根据权利要求1所述的方法,其中,所述根据所述第一语音特征确定所述语音数据所处的目标环境,包括:获取不同环境下样本语音数据,并提取所述样本语音数据对应的第二语音特征;根据所述第二语音特征,计算所述不同环境下样本语音数据对应的特征中心;根据所述特征中心和所述第一语音特征,确定所述语音数据所处的目标环境。
- 根据权利要求2所述的方法,其中,所述根据所述特征中心和所述第一语音特征,确定所述语音数据所处的目标环境,包括:利用预设的欧式距离算法计算所述第一语音特征与不同特征中心之间的欧式距离;从计算的欧式距离中筛选出最小欧式距离,并将所述最小欧式距离对应的样本语音数据所处环境确定为所述目标环境。
- 根据权利要求1所述的方法,其中,在所述获取待处理的语音数据之前,所述方法包括:利用初始语音增强参数对所述不同环境下的样本语音数据进行语音增强处理,得到不同环境下语音增强处理后的样本语音数据;根据所述样本语音数据,构建不同环境下的语音识别准确率函数;根据所述准确率函数,对所述初始语音增强参数进行优化调整,得到不同环境下的语音增强参数,并基于所述不同环境下的语音增强参数,构建所述语音增强参数集。
- 根据权利要求4所述的方法,其中,所述根据所述样本语音数据,构建不同环境下的语音识别准确率函数,包括:利用预先构建的语音识别模型对所述语音增强处理后的样本语音数据进行语音识别,得到不同环境下的语音识别结果;根据所述不同环境下的语音识别结果,构建不同环境下的语音识别准确率函数。
- 根据权利要求1所述的方法,其中,在所述根据所述目标语音增强参数,对所述语音数据进行语音增强处理,得到语音增强处理后的语音数据之后,所述方法还包括:对所述语音增强处理后的语音数据进行特征提取,得到所述语音数据对应的第三语音特征;根据所述第三语音特征,确定所述语音数据对应的语音识别结果。
- 根据权利要求1-6任一项所述的方法,其中,目标语音增强参数为目标滤波降噪参数,所述根据所述目标语音增强参数,对所述语音数据进行语音增强处理,得到语音增强处理后的语音数据,包括:根据所述目标滤波降噪参数,对所述语音数据进行滤波降噪处理,得到降噪处理后的语音数据。
- 一种语音增强装置,其中,包括:获取单元,用于获取待处理的语音数据;选取单元,用于提取所述语音数据对应的第一语音特征,根据所述第一语音特征确定所述语音数据所处的目标环境,并从预先构建的语音增强参数集中选取所述目标环境对应 的目标语音增强参数,所述语音增强参数集中包含有不同环境下的语音增强参数,所述语音增强参数用于增强不同环境下的语音识别准确率;处理单元,用于根据所述目标语音增强参数,对所述语音数据进行语音增强处理,得到语音增强处理后的语音数据。
- 一种计算机可读存储介质,其上存储有计算机程序,其中,所述计算机程序被处理器执行时实现一种语音增强方法的步骤:获取待处理的语音数据;提取所述语音数据对应的第一语音特征,根据所述第一语音特征确定所述语音数据所处的目标环境,并从预先构建的语音增强参数集中选取所述目标环境对应的目标语音增强参数,所述语音增强参数集中包含有不同环境下的语音增强参数,所述语音增强参数用于增强不同环境下的语音识别准确率;根据所述目标语音增强参数,对所述语音数据进行语音增强处理,得到语音增强处理后的语音数据。
- 根据权利要求9所述的计算机可读存储介质,其中,所述根据所述第一语音特征确定所述语音数据所处的目标环境,包括:获取不同环境下样本语音数据,并提取所述样本语音数据对应的第二语音特征;根据所述第二语音特征,计算所述不同环境下样本语音数据对应的特征中心;根据所述特征中心和所述第一语音特征,确定所述语音数据所处的目标环境。
- 根据权利要求10所述的计算机可读存储介质,其中,所述根据所述特征中心和所述第一语音特征,确定所述语音数据所处的目标环境,包括:利用预设的欧式距离算法计算所述第一语音特征与不同特征中心之间的欧式距离;从计算的欧式距离中筛选出最小欧式距离,并将所述最小欧式距离对应的样本语音数据所处环境确定为所述目标环境。
- 根据权利要求9所述的计算机可读存储介质,其中,在所述获取待处理的语音数据之前,所述方法包括:利用初始语音增强参数对所述不同环境下的样本语音数据进行语音增强处理,得到不同环境下语音增强处理后的样本语音数据;根据所述样本语音数据,构建不同环境下的语音识别准确率函数;根据所述准确率函数,对所述初始语音增强参数进行优化调整,得到不同环境下的语音增强参数,并基于所述不同环境下的语音增强参数,构建所述语音增强参数集。
- 根据权利要求12所述的计算机可读存储介质,其中,所述根据所述样本语音数据,构建不同环境下的语音识别准确率函数,包括:利用预先构建的语音识别模型对所述语音增强处理后的样本语音数据进行语音识别,得到不同环境下的语音识别结果;根据所述不同环境下的语音识别结果,构建不同环境下的语音识别准确率函数。
- 根据权利要求9所述的计算机可读存储介质,其中,在所述根据所述目标语音增强参数,对所述语音数据进行语音增强处理,得到语音增强处理后的语音数据之后,所述方法还包括:对所述语音增强处理后的语音数据进行特征提取,得到所述语音数据对应的第三语音特征;根据所述第三语音特征,确定所述语音数据对应的语音识别结果。
- 根据权利要求9-14任一项所述的计算机可读存储介质,其中,目标语音增强参数为目标滤波降噪参数,所述根据所述目标语音增强参数,对所述语音数据进行语音增强处理,得到语音增强处理后的语音数据,包括:根据所述目标滤波降噪参数,对所述语音数据进行滤波降噪处理,得到降噪处理后的 语音数据。
- 一种计算机设备,包括存储器、处理器及存储在存储器上并可在处理器上运行的计算机程序,其中,所述计算机程序被处理器执行时实现一种语音增强方法的步骤:获取待处理的语音数据;提取所述语音数据对应的第一语音特征,根据所述第一语音特征确定所述语音数据所处的目标环境,并从预先构建的语音增强参数集中选取所述目标环境对应的目标语音增强参数,所述语音增强参数集中包含有不同环境下的语音增强参数,所述语音增强参数用于增强不同环境下的语音识别准确率;根据所述目标语音增强参数,对所述语音数据进行语音增强处理,得到语音增强处理后的语音数据。
- 根据权利要求16所述的计算机设备,其中,所述根据所述第一语音特征确定所述语音数据所处的目标环境,包括:获取不同环境下样本语音数据,并提取所述样本语音数据对应的第二语音特征;根据所述第二语音特征,计算所述不同环境下样本语音数据对应的特征中心;根据所述特征中心和所述第一语音特征,确定所述语音数据所处的目标环境。
- 根据权利要求17所述的计算机设备,其中,所述根据所述特征中心和所述第一语音特征,确定所述语音数据所处的目标环境,包括:利用预设的欧式距离算法计算所述第一语音特征与不同特征中心之间的欧式距离;从计算的欧式距离中筛选出最小欧式距离,并将所述最小欧式距离对应的样本语音数据所处环境确定为所述目标环境。
- 根据权利要求16所述的计算机设备,其中,在所述获取待处理的语音数据之前,所述方法包括:利用初始语音增强参数对所述不同环境下的样本语音数据进行语音增强处理,得到不同环境下语音增强处理后的样本语音数据;根据所述样本语音数据,构建不同环境下的语音识别准确率函数;根据所述准确率函数,对所述初始语音增强参数进行优化调整,得到不同环境下的语音增强参数,并基于所述不同环境下的语音增强参数,构建所述语音增强参数集。
- 根据权利要求19所述的计算机设备,其中,所述根据所述样本语音数据,构建不同环境下的语音识别准确率函数,包括:利用预先构建的语音识别模型对所述语音增强处理后的样本语音数据进行语音识别,得到不同环境下的语音识别结果;根据所述不同环境下的语音识别结果,构建不同环境下的语音识别准确率函数。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202011153521.2A CN112151052B (zh) | 2020-10-26 | 2020-10-26 | 语音增强方法、装置、计算机设备及存储介质 |
| CN202011153521.2 | 2020-10-26 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2021189979A1 true WO2021189979A1 (zh) | 2021-09-30 |
Family
ID=73955013
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2020/136364 Ceased WO2021189979A1 (zh) | 2020-10-26 | 2020-12-15 | 语音增强方法、装置、计算机设备及存储介质 |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN112151052B (zh) |
| WO (1) | WO2021189979A1 (zh) |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN114512136A (zh) * | 2022-03-18 | 2022-05-17 | 北京百度网讯科技有限公司 | 模型训练、音频处理方法、装置、设备、存储介质及程序 |
Families Citing this family (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN113539262B (zh) * | 2021-07-09 | 2023-08-22 | 广东金鸿星智能科技有限公司 | 一种用于电动门语音控制的声音增强及收录方法和系统 |
Citations (7)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20090271189A1 (en) * | 2008-04-24 | 2009-10-29 | International Business Machines | Testing A Grammar Used In Speech Recognition For Reliability In A Plurality Of Operating Environments Having Different Background Noise |
| CN101593522A (zh) * | 2009-07-08 | 2009-12-02 | 清华大学 | 一种全频域数字助听方法和设备 |
| CN101710490A (zh) * | 2009-11-20 | 2010-05-19 | 安徽科大讯飞信息科技股份有限公司 | 语音评测的噪声补偿方法及装置 |
| CN103456305A (zh) * | 2013-09-16 | 2013-12-18 | 东莞宇龙通信科技有限公司 | 终端和基于多个声音采集单元的语音处理方法 |
| CN104575509A (zh) * | 2014-12-29 | 2015-04-29 | 乐视致新电子科技(天津)有限公司 | 语音增强处理方法及装置 |
| CN110473568A (zh) * | 2019-08-08 | 2019-11-19 | Oppo广东移动通信有限公司 | 场景识别方法、装置、存储介质及电子设备 |
| CN110648680A (zh) * | 2019-09-23 | 2020-01-03 | 腾讯科技(深圳)有限公司 | 语音数据的处理方法、装置、电子设备及可读存储介质 |
Family Cites Families (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP5731929B2 (ja) * | 2011-08-08 | 2015-06-10 | 日本電信電話株式会社 | 音声強調装置とその方法とプログラム |
| KR20190037867A (ko) * | 2017-09-29 | 2019-04-08 | 주식회사 케이티 | 잡음이 섞인 음성 데이터로부터 잡음을 제거하는 장치, 방법 및 컴퓨터 프로그램 |
| CN111698629B (zh) * | 2019-03-15 | 2021-10-15 | 北京小鸟听听科技有限公司 | 音频重放设备的校准方法、装置及计算机存储介质 |
| CN110503974B (zh) * | 2019-08-29 | 2022-02-22 | 泰康保险集团股份有限公司 | 对抗语音识别方法、装置、设备及计算机可读存储介质 |
-
2020
- 2020-10-26 CN CN202011153521.2A patent/CN112151052B/zh active Active
- 2020-12-15 WO PCT/CN2020/136364 patent/WO2021189979A1/zh not_active Ceased
Patent Citations (7)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20090271189A1 (en) * | 2008-04-24 | 2009-10-29 | International Business Machines | Testing A Grammar Used In Speech Recognition For Reliability In A Plurality Of Operating Environments Having Different Background Noise |
| CN101593522A (zh) * | 2009-07-08 | 2009-12-02 | 清华大学 | 一种全频域数字助听方法和设备 |
| CN101710490A (zh) * | 2009-11-20 | 2010-05-19 | 安徽科大讯飞信息科技股份有限公司 | 语音评测的噪声补偿方法及装置 |
| CN103456305A (zh) * | 2013-09-16 | 2013-12-18 | 东莞宇龙通信科技有限公司 | 终端和基于多个声音采集单元的语音处理方法 |
| CN104575509A (zh) * | 2014-12-29 | 2015-04-29 | 乐视致新电子科技(天津)有限公司 | 语音增强处理方法及装置 |
| CN110473568A (zh) * | 2019-08-08 | 2019-11-19 | Oppo广东移动通信有限公司 | 场景识别方法、装置、存储介质及电子设备 |
| CN110648680A (zh) * | 2019-09-23 | 2020-01-03 | 腾讯科技(深圳)有限公司 | 语音数据的处理方法、装置、电子设备及可读存储介质 |
Cited By (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN114512136A (zh) * | 2022-03-18 | 2022-05-17 | 北京百度网讯科技有限公司 | 模型训练、音频处理方法、装置、设备、存储介质及程序 |
| CN114512136B (zh) * | 2022-03-18 | 2023-09-26 | 北京百度网讯科技有限公司 | 模型训练、音频处理方法、装置、设备、存储介质及程序 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN112151052B (zh) | 2024-06-25 |
| CN112151052A (zh) | 2020-12-29 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| CN110600017B (zh) | 语音处理模型的训练方法、语音识别方法、系统及装置 | |
| CN107742522B (zh) | 基于麦克风阵列的目标语音获取方法及装置 | |
| CN110444214B (zh) | 语音信号处理模型训练方法、装置、电子设备及存储介质 | |
| CN112185352A (zh) | 语音识别方法、装置及电子设备 | |
| CN112466327B (zh) | 语音处理方法、装置和电子设备 | |
| CN114974280B (zh) | 音频降噪模型的训练方法、音频降噪的方法及装置 | |
| CN107564513A (zh) | 语音识别方法及装置 | |
| CN109801635A (zh) | 一种基于注意力机制的声纹特征提取方法及装置 | |
| WO2019237518A1 (zh) | 模型库建立方法、语音识别方法、装置、设备及介质 | |
| CN110610718A (zh) | 一种提取期望声源语音信号的方法及装置 | |
| WO2021189981A1 (zh) | 语音噪声的处理方法、装置、计算机设备及存储介质 | |
| WO2021189979A1 (zh) | 语音增强方法、装置、计算机设备及存储介质 | |
| CN115171716B (zh) | 一种基于空间特征聚类的连续语音分离方法、系统及电子设备 | |
| CN111354372B (zh) | 一种基于前后端联合训练的音频场景分类方法及系统 | |
| CN118899005A (zh) | 一种音频信号处理方法、装置、计算机设备及存储介质 | |
| CN114220430A (zh) | 多音区语音交互方法、装置、设备以及存储介质 | |
| CN116705013B (zh) | 语音唤醒词的检测方法、装置、存储介质和电子设备 | |
| CN119943058B (zh) | 一种基于时频域动态特征矩阵的说话人识别方法和系统 | |
| CN113077779A (zh) | 一种降噪方法、装置、电子设备以及存储介质 | |
| WO2021189980A1 (zh) | 语音数据生成方法、装置、计算机设备及存储介质 | |
| CN118135999A (zh) | 基于边缘设备的离线语音关键词识别方法及装置 | |
| CN117198300A (zh) | 一种基于注意力机制的鸟类声音识别方法及装置 | |
| CN114005459B (zh) | 人声分离方法、装置和电子设备 | |
| WO2023226592A1 (zh) | 噪音信号的处理方法和装置、存储介质及电子装置 | |
| CN117912482A (zh) | 一种基于全卷积神经网络多任务学习的时域语音分离方法 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 20927581 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 20927581 Country of ref document: EP Kind code of ref document: A1 |