WO2022249302A1 - 信号処理装置、信号処理方法及び信号処理プログラム - Google Patents
信号処理装置、信号処理方法及び信号処理プログラム Download PDFInfo
- Publication number
- WO2022249302A1 WO2022249302A1 PCT/JP2021/019874 JP2021019874W WO2022249302A1 WO 2022249302 A1 WO2022249302 A1 WO 2022249302A1 JP 2021019874 W JP2021019874 W JP 2021019874W WO 2022249302 A1 WO2022249302 A1 WO 2022249302A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- speech
- mixed
- signal processing
- sir
- snr
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/20—Speech recognition techniques specially adapted for robustness in adverse environments, e.g. in noise, of stress induced speech
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L21/00—Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
- G10L21/02—Speech enhancement, e.g. noise reduction or echo cancellation
- G10L21/0208—Noise filtering
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L21/00—Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
- G10L21/02—Speech enhancement, e.g. noise reduction or echo cancellation
- G10L21/0272—Voice signal separating
- G10L21/0308—Voice signal separating characterised by the type of parameter measurement, e.g. correlation techniques, zero crossing techniques or predictive techniques
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L25/00—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
- G10L25/78—Detection of presence or absence of voice signals
- G10L25/84—Detection of presence or absence of voice signals for discriminating voice from noise
Definitions
- the present invention relates to a signal processing device, a signal processing method, and a signal processing program.
- Blind sound source separation technology is one of these approaches, and separates the sound originating from each sound source including the mixed sound.
- Blind sound source separation technology enables speech recognition by separating speech, which is difficult to recognize as a mixed sound, into the speech of each speaker (see, for example, Non-Patent Document 1).
- the target speaker extraction technology is a technology that uses the pre-registered utterances of the target speaker as auxiliary information and acquires only the voice of the pre-registered speaker from the mixed sound.
- This target speaker extraction technique has the advantage of not requiring prior information regarding the number of speakers included in the mixed sound, and is a practically useful technique (see, for example, Non-Patent Document 2). Speech recognition is possible because the speech extracted using the target speaker extraction technique contains only the voice of the target speaker.
- speech enhancement technology removes unwanted sounds when the target speaker's speech overlaps with other speakers' speech or overlaps with background noise, This technology makes it possible to extract only voice. However, in removing unwanted sounds, it may also distort the desired target speaker's speech.
- Non-Patent Document 3 speech enhancement is effective when the noise is another speaker's speech (speech noise), but speech enhancement degrades speech recognition when it is not speech (non-speech noise).
- speech enhancement degrades speech recognition when it is not speech (non-speech noise).
- a binary classifier that distinguishes whether the type of noise is the speech of another speaker or other noise, and the type of noise is the speech of another speaker.
- proposed a method for speech enhancement In other words, it is proposed that speech enhancement is not performed when the target speaker's speech does not include another speaker's speech, and that speech enhancement is performed only when the target speaker's speech overlaps.
- the present invention has been made in view of the above. It is an object of the present invention to provide a signal processing device, a signal processing method, and a signal processing program capable of appropriately determining at least one of a mixed voice and an emphasized voice obtained by emphasizing a mixed voice.
- a signal processing apparatus provides a signal processing apparatus that converts a mixed speech in which a speech of a target speaker overlaps a speech of another speaker into a target speech in the mixed speech.
- mixed speech and emphasized speech obtained by emphasizing the mixed speech are used as input speech that can improve the accuracy of speech recognition. At least one of and can be determined appropriately.
- FIG. 1 is a diagram schematically showing an example of the configuration of a signal processing device according to an embodiment.
- FIG. 2 is a diagram schematically showing an example of a configuration of a speech recognition input determination unit shown in FIG. 1;
- FIG. 3 is a diagram showing speech recognition performance for enhanced speech in which mixed speech is emphasized.
- FIG. 4 is a diagram showing speech recognition performance for mixed speech.
- FIG. 5 is a flowchart showing a processing procedure of signal processing according to the embodiment.
- FIG. 6 is a flow chart showing a processing procedure of the speech recognition input determination process shown in FIG.
- FIG. 7 is a diagram showing speech recognition results when the signal processing device according to the embodiment is applied.
- FIG. 1 is a diagram schematically showing an example of the configuration of a signal processing device according to an embodiment.
- FIG. 2 is a diagram schematically showing an example of a configuration of a speech recognition input determination unit shown in FIG. 1
- FIG. 3 is a diagram showing speech recognition performance for enhanced speech in which mixed speech is
- FIG. 8 is a diagram showing the performance improvement rate of the speech recognition result using the signal processing apparatus according to the embodiment with respect to the speech recognition result when using the emphasized speech.
- FIG. 9 is a diagram showing an example of a computer that implements a signal processing device by executing a program.
- mixed speech in which the speech of the target speaker overlaps the speech of another speaker is targeted, and mixed speech and/or enhanced speech that emphasizes the mixed speech.
- SNR Signal to Noise Ratio
- SIR Signal to Noise Ratio
- Speech based on at least one of the mixed speech and the enhanced speech is determined based on at least one of (Signal to Interference Ratio) and the mixed speech.
- speech recognition accuracy can be improved by performing speech recognition processing using speech based on at least one of the determined mixed speech and emphasized speech.
- FIG. 1 is a diagram schematically showing an example of the configuration of a signal processing device according to an embodiment.
- the signal processing device 100 for example, a computer including ROM (Read Only Memory), RAM (Random Access Memory), CPU (Central Processing Unit), etc. loads a predetermined program, and the CPU executes the predetermined program. It is realized by The signal processing device 100 also has a communication interface for transmitting and receiving various information to and from another device connected via a wired connection or a network or the like.
- the signal processing device 100 has a speech enhancement unit 101 , a speech recognition input determination unit 102 and a speech recognition unit 103 .
- the speech enhancement unit 101 uses a known speech enhancement technique to extract desired speech from the input mixed speech.
- a known sound source separation technique or a known target speaker extraction technique can be applied.
- input mixed speech is separated into each speaker, so that the speech of each speaker is independently given to the speech recognition input determination unit 102 .
- auxiliary information about the target speaker must also be given as an input. For example, utterances registered in advance by the target speaker are used as auxiliary information about the target speaker.
- the auxiliary information regarding the target speaker may be held by the signal processing device 100 or may be input from an external device capable of communicating with the signal processing device 100 .
- the speech enhancement unit 101 extracts enhanced speech obtained by enhancing the mixed speech. For example, when mixed speech of the target speaker and other interfering speakers is input to the speech enhancement unit 101 as mixed speech, the speech enhancement unit 101 converts the mixed speech into Extract audio.
- the speech enhancement unit 101 inputs the extracted speech as an enhanced speech to the SIR-SNR acquisition unit 1021 (described later) of the speech recognition input determination unit 102 .
- the input mixed speech and emphasized speech may be obtained by subjecting the speech to some conversion such as feature extraction, or may be speech waveforms.
- the speech recognition input determination unit 102 determines speech based on at least one of the mixed speech and the emphasized speech obtained by emphasizing the mixed speech. , is determined as speech to be used for speech recognition.
- the speech recognition input determination unit 102 inputs the determined speech based on at least one of the mixed speech and the emphasized speech to the speech recognition unit 103 .
- the speech recognition input determination unit 102 converts speech input to speech recognition into speech based on at least one of the emphasized speech output from the speech enhancement unit 101 and the mixed speech input to the signal processing device 100. Judgment is made in units of certain sections such as frames.
- the voice recognition input determination unit 102 performs voice determination by hardware. For example, the speech recognition input determination unit 102 recognizes only one of the emphasized speech output from the speech enhancement unit 101 and the mixed speech input to the signal processing device 100 with "1" and "0". Input to the part 103 .
- the voice recognition input determination unit 102 can softly determine the voice. For example, the speech recognition input determination unit 102 weights one of the emphasized speech output from the speech enhancement unit 101 and the mixed speech input to the signal processing device 100 by, for example, "0.3", and weights the other by "0.7". , and the weighted mixed speech and emphasized speech are input to the speech recognition unit 103 .
- the speech recognition unit 103 performs speech recognition processing using the speech determined by the speech recognition input determination unit 102 based on at least one of the mixed speech and the emphasized speech.
- a known speech recognition technique can be applied to the speech recognition processing of the speech recognition unit 103 .
- FIG. 3 is a diagram showing the speech recognition performance for emphasized speech in which mixed speech is emphasized.
- FIG. 4 is a diagram showing speech recognition performance for mixed speech. In both FIGS. 3 and 4, the character error rate was evaluated as speech recognition performance.
- the speech recognition input determination unit 102 selects emphasized speech for regions of SNR and SIR that can improve the accuracy of speech recognition in FIG. For other regions, mixed speech is selected and input to the speech recognition unit 103 .
- the speech recognition input determination unit 102 weights the emphasized speech more than the mixed speech in the SNR and SIR regions that can improve the accuracy of speech recognition in FIG. , the mixed speech is weighted more heavily than the emphasized speech, and the weighted mixed speech and the emphasized speech are input to the speech recognition unit 103 .
- FIG. 2 is a diagram schematically showing an example of the configuration of the speech recognition input determination section 102 shown in FIG. 1. As shown in FIG.
- the speech recognition input determination unit 102 has an SIR-SNR acquisition unit 1021 (acquisition unit), a determination unit 1022 and a switching unit 1023.
- the SIR-SNR acquisition unit 1021 acquires at least one of the SIR of the mixed speech, the SNR of the mixed speech, and the mixed speech from the mixed speech input to the system.
- the SIR-SNR acquisition unit 1021 may acquire the SIR and SIR in units of a certain time such as frames, or in units of utterances.
- the SIR-SNR acquisition unit 1021 estimates SIR and SNR using, for example, an estimation model for estimating SIR and SNR.
- the estimation model is configured by a neural network or the like.
- the estimation model is pre-learned with paired data of speech and SIR and SNR.
- the speech enhancement unit 101 extracts the target speaker's speech and the other interfering speaker's speech
- the SIR-SNR acquisition unit 1021 extracts the extracted target speaker's speech and the other speech.
- SIR and SIR may be calculated using equations (1) and (2) based on the speech of the interfering speaker.
- the determination unit 1022 Based on at least one of the SIR and SNR obtained by the SIR-SNR obtaining unit 1021 and the mixed speech, the determination unit 1022 obtains speech based on at least one of the mixed speech and the emphasized speech obtained by enhancing the mixed speech. is determined as the voice to be used for voice recognition. Based on at least one of the SIR and SNR estimated or calculated by the SIR-SNR acquisition unit 1021 and the mixed speech, the determination unit 1022 determines which of the mixed speech and the emphasized mixed speech is speech recognition. determine if it is beneficial to
- the determination unit 1022 uses a predetermined rule using SIR and SNR to determine speech based on at least one of mixed speech and enhanced speech as speech to be used for speech recognition.
- the determination unit 1022 uses a rule to select the emphasized voice in the case of the expression (3), for example, when determining the voice by hardware.
- the determination unit 1022 determines whether the mixed speech or the enhanced speech is more advantageous as an input for speech recognition depending on the SIR and SNR (see FIGS. 3 and 4). The determination may be made using a rule in which a weight value for mixed speech and a weight value for emphasized speech are set in advance according to the combination of .
- the judgment unit 1022 uses, as an input of the speech recognition unit 103, a discrimination model for distinguishing speech based on at least one of the mixed speech and the emphasized speech, and distinguishes between the mixed speech and the emphasized speech. Speech based on at least one may be determined as speech to be used for speech recognition.
- a discriminative model consists of a neural network, etc.
- the discriminative model is at least the feature obtained from the mixed speech to which the teacher label indicating which of the mixed speech and the enhanced speech is advantageous as the speech to be used for speech recognition, and the SNR and SIR. One of them is learned in advance as an input.
- the discriminative model takes SNR and SIR as inputs, a teacher label indicating which of mixed speech and enhanced speech is more advantageous as speech to be used for speech recognition, and estimated or calculated SNR and SIR pair data. Based on this, the discriminative model is learned.
- the discriminative model can also directly estimate which of the mixed speech and the enhanced speech is advantageous for speech recognition by using the feature values obtained from the mixed speech as input.
- the training of the discriminant model is based on the teacher label that indicates which of the mixed speech and the emphasized speech is more advantageous as the speech to be used for speech recognition, and the paired data of the feature values extracted from the mixed speech. I do.
- the discrimination model inputs both SNR and SIR and mixed speech, discrimination is performed based on teacher label pair data indicating which of these, mixed speech, and enhanced speech is more advantageous as speech to be used for speech recognition. Train the model.
- the discriminative model outputs weights that indicate the advantages of mixed speech and enhanced speech as speech used for speech recognition. Further, the discrimination model may output a determination result in which either mixed speech or emphasized speech is selected as speech that is advantageous as speech to be used for speech recognition.
- the switching unit 1023 inputs to the voice recognition unit 103 the voice based on at least one of the mixed voice and the emphasized voice.
- speech determination is performed by hardware as shown in FIG.
- the switching unit 1023 weights the mixed speech and the emphasized speech according to the weight output from the judgment unit 1022, and divides the weighted mixed speech and the emphasized speech into weighted speech. is output to the speech recognition unit 103 .
- FIG. 5 is a flowchart showing a processing procedure of signal processing according to the embodiment.
- the speech enhancement unit 101 When input of mixed speech is received (step S1), the speech enhancement unit 101 extracts the speech of the target speaker from the input mixed speech, and inputs the extracted speech as emphasized speech for speech recognition. Speech enhancement processing to be input to the determination unit 102 is performed (step S2).
- the speech recognition input determination unit 102 determines speech based on at least one of the mixed speech and the emphasized speech obtained by emphasizing the mixed speech. , a voice recognition input determination process is performed to determine that the voice is to be used for voice recognition (step S3).
- the speech recognition unit 103 performs speech recognition processing using speech based on at least one of the mixed speech and the emphasized speech determined by the speech recognition input determination unit 102 (step S4), and obtains a speech recognition result. is output (step S5).
- FIG. 6 is a flow chart showing a processing procedure of the speech recognition input determination process shown in FIG.
- the SIR-SNR acquisition unit 1021 performs an acquisition process of acquiring at least one of the SIR of the mixed speech, the SNR of the mixed speech, and the mixed speech from the mixed speech input to the system. (Step S11).
- the determination unit 1022 obtains speech based on at least one of the mixed speech and the emphasized speech obtained by enhancing the mixed speech. is determined as the voice to be used for voice recognition (step S12).
- the switching unit 1023 inputs the speech based on at least one of the mixed speech and the emphasized speech to the speech recognition unit 103 as the speech to be used for speech recognition based on the judgment result of the judgment unit 1022 (step S13).
- the SIR and SNR and the mixed speech Speech based on at least one of the mixed speech and the enhanced speech is determined as the speech to be used for speech recognition.
- FIG. 7 is a diagram showing an example of speech recognition results when the signal processing device 100 according to the embodiment is applied.
- FIG. 8 is a diagram showing an example of the performance improvement rate of the speech recognition result using the signal processing apparatus 100 according to the embodiment with respect to the speech recognition result when using the emphasized speech.
- 7 and 8 show speech recognition results when the speech to be input to the speech recognition unit 103 is selected using the rule that the emphasized speech is selected in the case of expression (3) from the emphasized speech and the mixed speech. be. In both FIGS. 7 and 8, the character error rate was evaluated as speech recognition performance.
- the area of the frame W1 shown in FIG. 7 and the area of the frame W11 shown in FIG. 8 are areas where the input signal is considered to be selected as mixed speech based on the estimated SIR and SNR.
- Signal processing device 100 that selects a signal to speech recognition unit 103 using a rule that emphasized speech is selected in the case of equation (3) from the speech recognition results of frame W1 shown in FIG. 7 and frame W11 shown in FIG. It was found that the speech recognition accuracy is higher in the case of using only the emphasized speech.
- the signal processing device 100 has higher speech recognition accuracy than the case where only the emphasized speech is used for all combinations of SIR and SNR. Furthermore, it is clear that the signal processing device 100 has particularly high speech recognition accuracy in areas where the SIR value is relatively large (see, for example, frames W11 and W12) compared to when only emphasized speech is used. Became.
- the signal processing apparatus 100 can prevent performance deterioration due to distortion of speech enhancement and improve speech recognition accuracy by appropriately determining the input speech that can improve the accuracy of speech recognition.
- the speech enhancement unit 101 does not need to perform speech enhancement processing when mixed speech is selected as input speech that can improve the accuracy of speech recognition. In this way, the speech enhancement unit 101 performs speech enhancement only for a section that requires speech enhancement processing, that is, a section for which the speech recognition input determination unit 102 selects an enhanced speech as an input speech that can improve the accuracy of speech recognition. Perform emphasis processing.
- Each component of the signal processing device 100 is functionally conceptual, and does not necessarily need to be physically configured as illustrated. That is, the specific form of distribution and integration of the functions of the signal processing device 100 is not limited to the illustrated one, and all or part of it can be functionally or physically distributed in arbitrary units according to various loads and usage conditions. can be distributed or integrated.
- each process performed in the signal processing device 100 may be realized by a CPU, a GPU (Graphics Processing Unit), and a program that is analyzed and executed by the CPU and GPU. Further, each process performed in the signal processing device 100 may be realized as hardware by wired logic.
- FIG. 9 is a diagram showing an example of a computer that implements the signal processing device 100 by executing a program.
- the computer 1000 has a memory 1010 and a CPU 1020, for example.
- Computer 1000 also has hard disk drive interface 1030 , disk drive interface 1040 , serial port interface 1050 , video adapter 1060 and network interface 1070 . These units are connected by a bus 1080 .
- the memory 1010 includes a ROM 1011 and a RAM 1012.
- the ROM 1011 stores a boot program such as BIOS (Basic Input Output System).
- BIOS Basic Input Output System
- Hard disk drive interface 1030 is connected to hard disk drive 1090 .
- a disk drive interface 1040 is connected to the disk drive 1100 .
- a removable storage medium such as a magnetic disk or optical disk is inserted into the disk drive 1100 .
- Serial port interface 1050 is connected to mouse 1110 and keyboard 1120, for example.
- Video adapter 1060 is connected to display 1130, for example.
- the hard disk drive 1090 stores an OS (Operating System) 1091, application programs 1092, program modules 1093, and program data 1094, for example. That is, a program that defines each process of the signal processing device 100 is implemented as a program module 1093 in which code executable by the computer 1000 is described. Program modules 1093 are stored, for example, on hard disk drive 1090 .
- the hard disk drive 1090 stores a program module 1093 for executing processing similar to the functional configuration of the signal processing device 100 .
- the hard disk drive 1090 may be replaced by an SSD (Solid State Drive).
- the setting data used in the processing of the above-described embodiment is stored as program data 1094 in the memory 1010 or the hard disk drive 1090, for example. Then, the CPU 1020 reads out the program module 1093 and the program data 1094 stored in the memory 1010 and the hard disk drive 1090 to the RAM 1012 as necessary and executes them.
- the program modules 1093 and program data 1094 are not limited to being stored in the hard disk drive 1090, but may be stored in a removable storage medium, for example, and read by the CPU 1020 via the disk drive 1100 or the like. Alternatively, the program modules 1093 and program data 1094 may be stored in another computer connected via a network (LAN (Local Area Network), WAN (Wide Area Network), etc.). Program modules 1093 and program data 1094 may then be read by CPU 1020 through network interface 1070 from other computers.
- LAN Local Area Network
- WAN Wide Area Network
Landscapes
- Engineering & Computer Science (AREA)
- Computational Linguistics (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Human Computer Interaction (AREA)
- Physics & Mathematics (AREA)
- Acoustics & Sound (AREA)
- Multimedia (AREA)
- Signal Processing (AREA)
- Quality & Reliability (AREA)
- Circuit For Audible Band Transducer (AREA)
- Telephonic Communication Services (AREA)
Abstract
Description
目的話者の音声に別の話者の音声が重複した場合であっても、音声強調を行わない方がよい場合がありうる。例えば、干渉話者の音声が目的話者の音声と比較して振幅が小さい場合、音声強調せずとも音声認識がある程度目的話者の音声を認識することが可能である場合がある。また、干渉話者の音声の他に、非音声雑音が存在する場合、音声強調による歪みの影響は大きくなる傾向が確認されており、このような場合については音声強調の悪影響が大きく、もともとの混合音声を認識したほうが音声認識にとって有利であることがある。
次に、実施の形態に係る信号処理装置について説明する。図1は、実施の形態に係る信号処理装置の構成の一例を模式的に示す図である。
ここで、SIRの演算式を式(1)に示し、SNRの演算式を式(2)に示す。式(1)、式(2)において、Sは、目的話者の音声であり、Iは、干渉話者の音声であり、Nは、雑音である。
次に、音声認識入力判定部102について説明する。図2は、図1に示す音声認識入力判定部102の構成の一例を模式的に示す図である。
次に、信号処理装置100が実行する信号処理について説明する。図5は、実施の形態に係る信号処理の処理手順を示すフローチャートである。
次に、図5の音声認識入力判定処理(ステップS3)について説明する。図6は、図5に示す音声認識入力判定処理の処理手順を示すフローチャートである。
このように、信号処理装置100では、音声認識の入力として、混合音声と強調音声とのいずれかが有利かであるかがSIR及びSNRに依存することに基づいて、SIR及びSNRと、混合音声とのうち少なくとも一つを基に、混合音声と強調音声との少なくとも一方に基づく音声を、音声認識に使用する音声として判定する。
信号処理装置100の各構成要素は機能概念的なものであり、必ずしも物理的に図示のように構成されていることを要しない。すなわち、信号処理装置100の機能の分散及び統合の具体的形態は図示のものに限られず、その全部または一部を、各種の負荷や使用状況などに応じて、任意の単位で機能的または物理的に分散または統合して構成することができる。
図9は、プログラムが実行されることにより、信号処理装置100が実現されるコンピュータの一例を示す図である。コンピュータ1000は、例えば、メモリ1010、CPU1020を有する。また、コンピュータ1000は、ハードディスクドライブインタフェース1030、ディスクドライブインタフェース1040、シリアルポートインタフェース1050、ビデオアダプタ1060、ネットワークインタフェース1070を有する。これらの各部は、バス1080によって接続される。
101 音声強調部
102 音声認識入力判定部
103 音声認識部
1021 SIR-SNR取得部
1022 判定部
1023 スイッチング部
Claims (7)
- 目的話者の音声に別の話者の音声が重複している混合音声から、前記混合音声における、目的音声と干渉話者音声との比率であるSIR(Signal to Interference Ratio)、及び、前記混合音声の、目的音声と雑音との比率であるSNR(Signal to Noise Ratio)と、前記混合音声とのうち少なくとも一つを取得する取得部と、
前記SIR及び前記SNRと、前記混合音声とのうち少なくとも一つを基に、前記混合音声と、前記混合音声を強調した強調音声との少なくとも一方に基づく音声を、音声認識に使用する音声として判定する判定部と、
を有することを特徴とする信号処理装置。 - 前記判定部は、前記SIR及び前記SNRを用いた所定のルールを用いて、前記混合音声と前記強調音声との少なくとも一方に基づく音声を、前記音声認識に使用する音声として判定することを特徴とする請求項1に記載の信号処理装置。
- 前記判定部は、前記混合音声と前記強調音声とのいずれが音声認識に使用する音声として有利であるかを示す教師ラベルが付与された、前記混合音声から取得された特徴量と、前記SNR及び前記SIRと、の少なくとも一方を入力として学習された識別モデルを用いて、前記混合音声と前記強調音声との少なくとも一方に基づく音声を、前記音声認識に使用する音声として判定することを特徴とする請求項1に記載の信号処理装置。
- 前記判定部が判定した前記混合音声と前記強調音声との少なくとも一方に基づく音声を用いて、音声認識処理を行う音声認識部をさらに有することを特徴とする請求項1~3のいずれか一つに記載の信号処理装置。
- 前記混合音声から、前記目的話者の音声を抽出し、抽出した音声を前記強調音声として前記取得部に入力する音声強調部をさらに有することを特徴とする請求項1~4のいずれか一つに記載の信号処理装置。
- 信号処理装置が実行する信号処理方法であって、
目的話者の音声に別の話者の音声が重複している混合音声から、前記混合音声における、目的音声と干渉話者音声との比率であるSIR(Signal to Interference Ratio)、及び、前記混合音声の、目的音声と雑音との比率であるSNR(Signal to Noise Ratio)と、前記混合音声とのうち少なくとも一つを取得する工程と、
前記SIR及び前記SNRと、前記混合音声とのうち少なくとも一つを基に、前記混合音声と、前記混合音声を強調した強調音声との少なくとも一方に基づく音声を、音声認識に使用する音声として判定する工程と、
を含んだことを特徴とする信号処理方法。 - 目的話者の音声に別の話者の音声が重複している混合音声から、前記混合音声における、目的音声と干渉話者音声との比率であるSIR(Signal to Interference Ratio)、及び、前記混合音声の、目的音声と雑音との比率であるSNR(Signal to Noise Ratio)と、前記混合音声とのうち少なくとも一つを取得するステップと、
前記SIR及び前記SNRと、前記混合音声とのうち少なくとも一つを基に、前記混合音声と、前記混合音声を強調した強調音声との少なくとも一方に基づく音声を、音声認識に使用する音声として判定するステップと、
をコンピュータに実行させるための信号処理プログラム。
Priority Applications (3)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/JP2021/019874 WO2022249302A1 (ja) | 2021-05-25 | 2021-05-25 | 信号処理装置、信号処理方法及び信号処理プログラム |
| JP2023523777A JP7635837B2 (ja) | 2021-05-25 | 2021-05-25 | 信号処理装置、信号処理方法及び信号処理プログラム |
| US18/561,727 US20240274149A1 (en) | 2021-05-25 | 2021-05-25 | Signal processing device, signal processing method, and signal processing program |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/JP2021/019874 WO2022249302A1 (ja) | 2021-05-25 | 2021-05-25 | 信号処理装置、信号処理方法及び信号処理プログラム |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2022249302A1 true WO2022249302A1 (ja) | 2022-12-01 |
Family
ID=84228584
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/JP2021/019874 Ceased WO2022249302A1 (ja) | 2021-05-25 | 2021-05-25 | 信号処理装置、信号処理方法及び信号処理プログラム |
Country Status (3)
| Country | Link |
|---|---|
| US (1) | US20240274149A1 (ja) |
| JP (1) | JP7635837B2 (ja) |
| WO (1) | WO2022249302A1 (ja) |
Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2002236497A (ja) * | 2001-02-08 | 2002-08-23 | Alpine Electronics Inc | ノイズリダクションシステム |
| JP2007264327A (ja) * | 2006-03-28 | 2007-10-11 | Matsushita Electric Works Ltd | 浴室装置及びそれに用いる音声操作装置 |
| JP2018205512A (ja) * | 2017-06-02 | 2018-12-27 | 富士通株式会社 | 電子機器及び雑音抑圧プログラム |
-
2021
- 2021-05-25 JP JP2023523777A patent/JP7635837B2/ja active Active
- 2021-05-25 US US18/561,727 patent/US20240274149A1/en not_active Abandoned
- 2021-05-25 WO PCT/JP2021/019874 patent/WO2022249302A1/ja not_active Ceased
Patent Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2002236497A (ja) * | 2001-02-08 | 2002-08-23 | Alpine Electronics Inc | ノイズリダクションシステム |
| JP2007264327A (ja) * | 2006-03-28 | 2007-10-11 | Matsushita Electric Works Ltd | 浴室装置及びそれに用いる音声操作装置 |
| JP2018205512A (ja) * | 2017-06-02 | 2018-12-27 | 富士通株式会社 | 電子機器及び雑音抑圧プログラム |
Also Published As
| Publication number | Publication date |
|---|---|
| JP7635837B2 (ja) | 2025-02-26 |
| JPWO2022249302A1 (ja) | 2022-12-01 |
| US20240274149A1 (en) | 2024-08-15 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| CN109643552B (zh) | 用于可变噪声状况中语音增强的鲁棒噪声估计 | |
| CN105405439B (zh) | 语音播放方法及装置 | |
| CN106486131B (zh) | 一种语音去噪的方法及装置 | |
| KR102410392B1 (ko) | 실행 중 범위 정규화를 이용하는 신경망 음성 활동 검출 | |
| JP6243858B2 (ja) | 音声モデル学習方法、雑音抑圧方法、音声モデル学習装置、雑音抑圧装置、音声モデル学習プログラム及び雑音抑圧プログラム | |
| US10748544B2 (en) | Voice processing device, voice processing method, and program | |
| WO2019062721A1 (zh) | 语音身份特征提取器、分类器训练方法及相关设备 | |
| CN112002307B (zh) | 一种语音识别方法和装置 | |
| CN111524527A (zh) | 话者分离方法、装置、电子设备和存储介质 | |
| KR102275656B1 (ko) | 적대적 학습(adversarial training) 모델을 이용한 강인한 음성 향상 훈련 방법 및 그 장치 | |
| KR100745976B1 (ko) | 음향 모델을 이용한 음성과 비음성의 구분 방법 및 장치 | |
| WO2018176894A1 (zh) | 一种说话人确认方法及装置 | |
| JP6348427B2 (ja) | 雑音除去装置及び雑音除去プログラム | |
| JP7176627B2 (ja) | 信号抽出システム、信号抽出学習方法および信号抽出学習プログラム | |
| KR101618512B1 (ko) | 가우시안 혼합모델을 이용한 화자 인식 시스템 및 추가 학습 발화 선택 방법 | |
| Almajai et al. | Using audio-visual features for robust voice activity detection in clean and noisy speech | |
| Mun et al. | The sound of my voice: Speaker representation loss for target voice separation | |
| WO2020195924A1 (ja) | 信号処理装置および方法、並びにプログラム | |
| CN114299962A (zh) | 基于音频流的对话角色分离方法、系统、设备及存储介质 | |
| Hou et al. | Domain adversarial training for speech enhancement | |
| JP2020071482A (ja) | 語音分離方法、語音分離モデル訓練方法及びコンピュータ可読媒体 | |
| Nasir et al. | A hybrid method for speech noise reduction using log-MMSE | |
| JP7743875B2 (ja) | 音声信号の処理方法、音声信号処理装置、およびプログラム | |
| JP4787979B2 (ja) | 雑音検出装置および雑音検出方法 | |
| Al-Ali et al. | Enhanced forensic speaker verification using multi-run ICA in the presence of environmental noise and reverberation conditions |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 21942956 Country of ref document: EP Kind code of ref document: A1 |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 2023523777 Country of ref document: JP |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 18561727 Country of ref document: US |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 21942956 Country of ref document: EP Kind code of ref document: A1 |

