WO2022249302A1 - 信号処理装置、信号処理方法及び信号処理プログラム - Google Patents

信号処理装置、信号処理方法及び信号処理プログラム Download PDF

Info

Publication number
WO2022249302A1
WO2022249302A1 PCT/JP2021/019874 JP2021019874W WO2022249302A1 WO 2022249302 A1 WO2022249302 A1 WO 2022249302A1 JP 2021019874 W JP2021019874 W JP 2021019874W WO 2022249302 A1 WO2022249302 A1 WO 2022249302A1
Authority
WO
WIPO (PCT)
Prior art keywords
speech
mixed
signal processing
sir
snr
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/JP2021/019874
Other languages
English (en)
French (fr)
Inventor
宏 佐藤
翼 落合
マーク デルクロア
慶介 木下
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
NTT Inc
Original Assignee
Nippon Telegraph and Telephone Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Nippon Telegraph and Telephone Corp filed Critical Nippon Telegraph and Telephone Corp
Priority to PCT/JP2021/019874 priority Critical patent/WO2022249302A1/ja
Priority to JP2023523777A priority patent/JP7635837B2/ja
Priority to US18/561,727 priority patent/US20240274149A1/en
Publication of WO2022249302A1 publication Critical patent/WO2022249302A1/ja
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • G—PHYSICS
    • G10—MUSICAL INSTRUMENTS; ACOUSTICS
    • G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00—Speech recognition
    • G10L15/20—Speech recognition techniques specially adapted for robustness in adverse environments, e.g. in noise, of stress induced speech
    • G—PHYSICS
    • G10—MUSICAL INSTRUMENTS; ACOUSTICS
    • G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L21/00—Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
    • G10L21/02—Speech enhancement, e.g. noise reduction or echo cancellation
    • G10L21/0208—Noise filtering
    • G—PHYSICS
    • G10—MUSICAL INSTRUMENTS; ACOUSTICS
    • G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L21/00—Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
    • G10L21/02—Speech enhancement, e.g. noise reduction or echo cancellation
    • G10L21/0272—Voice signal separating
    • G10L21/0308—Voice signal separating characterised by the type of parameter measurement, e.g. correlation techniques, zero crossing techniques or predictive techniques
    • G—PHYSICS
    • G10—MUSICAL INSTRUMENTS; ACOUSTICS
    • G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L25/00—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
    • G10L25/78—Detection of presence or absence of voice signals
    • G10L25/84—Detection of presence or absence of voice signals for discriminating voice from noise

Definitions

  • the present invention relates to a signal processing device, a signal processing method, and a signal processing program.
  • Blind sound source separation technology is one of these approaches, and separates the sound originating from each sound source including the mixed sound.
  • Blind sound source separation technology enables speech recognition by separating speech, which is difficult to recognize as a mixed sound, into the speech of each speaker (see, for example, Non-Patent Document 1).
  • the target speaker extraction technology is a technology that uses the pre-registered utterances of the target speaker as auxiliary information and acquires only the voice of the pre-registered speaker from the mixed sound.
  • This target speaker extraction technique has the advantage of not requiring prior information regarding the number of speakers included in the mixed sound, and is a practically useful technique (see, for example, Non-Patent Document 2). Speech recognition is possible because the speech extracted using the target speaker extraction technique contains only the voice of the target speaker.
  • speech enhancement technology removes unwanted sounds when the target speaker's speech overlaps with other speakers' speech or overlaps with background noise, This technology makes it possible to extract only voice. However, in removing unwanted sounds, it may also distort the desired target speaker's speech.
  • Non-Patent Document 3 speech enhancement is effective when the noise is another speaker's speech (speech noise), but speech enhancement degrades speech recognition when it is not speech (non-speech noise).
  • speech enhancement degrades speech recognition when it is not speech (non-speech noise).
  • a binary classifier that distinguishes whether the type of noise is the speech of another speaker or other noise, and the type of noise is the speech of another speaker.
  • proposed a method for speech enhancement In other words, it is proposed that speech enhancement is not performed when the target speaker's speech does not include another speaker's speech, and that speech enhancement is performed only when the target speaker's speech overlaps.
  • the present invention has been made in view of the above. It is an object of the present invention to provide a signal processing device, a signal processing method, and a signal processing program capable of appropriately determining at least one of a mixed voice and an emphasized voice obtained by emphasizing a mixed voice.
  • a signal processing apparatus provides a signal processing apparatus that converts a mixed speech in which a speech of a target speaker overlaps a speech of another speaker into a target speech in the mixed speech.
  • mixed speech and emphasized speech obtained by emphasizing the mixed speech are used as input speech that can improve the accuracy of speech recognition. At least one of and can be determined appropriately.
  • FIG. 1 is a diagram schematically showing an example of the configuration of a signal processing device according to an embodiment.
  • FIG. 2 is a diagram schematically showing an example of a configuration of a speech recognition input determination unit shown in FIG. 1;
  • FIG. 3 is a diagram showing speech recognition performance for enhanced speech in which mixed speech is emphasized.
  • FIG. 4 is a diagram showing speech recognition performance for mixed speech.
  • FIG. 5 is a flowchart showing a processing procedure of signal processing according to the embodiment.
  • FIG. 6 is a flow chart showing a processing procedure of the speech recognition input determination process shown in FIG.
  • FIG. 7 is a diagram showing speech recognition results when the signal processing device according to the embodiment is applied.
  • FIG. 1 is a diagram schematically showing an example of the configuration of a signal processing device according to an embodiment.
  • FIG. 2 is a diagram schematically showing an example of a configuration of a speech recognition input determination unit shown in FIG. 1
  • FIG. 3 is a diagram showing speech recognition performance for enhanced speech in which mixed speech is
  • FIG. 8 is a diagram showing the performance improvement rate of the speech recognition result using the signal processing apparatus according to the embodiment with respect to the speech recognition result when using the emphasized speech.
  • FIG. 9 is a diagram showing an example of a computer that implements a signal processing device by executing a program.
  • mixed speech in which the speech of the target speaker overlaps the speech of another speaker is targeted, and mixed speech and/or enhanced speech that emphasizes the mixed speech.
  • SNR Signal to Noise Ratio
  • SIR Signal to Noise Ratio
  • Speech based on at least one of the mixed speech and the enhanced speech is determined based on at least one of (Signal to Interference Ratio) and the mixed speech.
  • speech recognition accuracy can be improved by performing speech recognition processing using speech based on at least one of the determined mixed speech and emphasized speech.
  • FIG. 1 is a diagram schematically showing an example of the configuration of a signal processing device according to an embodiment.
  • the signal processing device 100 for example, a computer including ROM (Read Only Memory), RAM (Random Access Memory), CPU (Central Processing Unit), etc. loads a predetermined program, and the CPU executes the predetermined program. It is realized by The signal processing device 100 also has a communication interface for transmitting and receiving various information to and from another device connected via a wired connection or a network or the like.
  • the signal processing device 100 has a speech enhancement unit 101 , a speech recognition input determination unit 102 and a speech recognition unit 103 .
  • the speech enhancement unit 101 uses a known speech enhancement technique to extract desired speech from the input mixed speech.
  • a known sound source separation technique or a known target speaker extraction technique can be applied.
  • input mixed speech is separated into each speaker, so that the speech of each speaker is independently given to the speech recognition input determination unit 102 .
  • auxiliary information about the target speaker must also be given as an input. For example, utterances registered in advance by the target speaker are used as auxiliary information about the target speaker.
  • the auxiliary information regarding the target speaker may be held by the signal processing device 100 or may be input from an external device capable of communicating with the signal processing device 100 .
  • the speech enhancement unit 101 extracts enhanced speech obtained by enhancing the mixed speech. For example, when mixed speech of the target speaker and other interfering speakers is input to the speech enhancement unit 101 as mixed speech, the speech enhancement unit 101 converts the mixed speech into Extract audio.
  • the speech enhancement unit 101 inputs the extracted speech as an enhanced speech to the SIR-SNR acquisition unit 1021 (described later) of the speech recognition input determination unit 102 .
  • the input mixed speech and emphasized speech may be obtained by subjecting the speech to some conversion such as feature extraction, or may be speech waveforms.
  • the speech recognition input determination unit 102 determines speech based on at least one of the mixed speech and the emphasized speech obtained by emphasizing the mixed speech. , is determined as speech to be used for speech recognition.
  • the speech recognition input determination unit 102 inputs the determined speech based on at least one of the mixed speech and the emphasized speech to the speech recognition unit 103 .
  • the speech recognition input determination unit 102 converts speech input to speech recognition into speech based on at least one of the emphasized speech output from the speech enhancement unit 101 and the mixed speech input to the signal processing device 100. Judgment is made in units of certain sections such as frames.
  • the voice recognition input determination unit 102 performs voice determination by hardware. For example, the speech recognition input determination unit 102 recognizes only one of the emphasized speech output from the speech enhancement unit 101 and the mixed speech input to the signal processing device 100 with "1" and "0". Input to the part 103 .
  • the voice recognition input determination unit 102 can softly determine the voice. For example, the speech recognition input determination unit 102 weights one of the emphasized speech output from the speech enhancement unit 101 and the mixed speech input to the signal processing device 100 by, for example, "0.3", and weights the other by "0.7". , and the weighted mixed speech and emphasized speech are input to the speech recognition unit 103 .
  • the speech recognition unit 103 performs speech recognition processing using the speech determined by the speech recognition input determination unit 102 based on at least one of the mixed speech and the emphasized speech.
  • a known speech recognition technique can be applied to the speech recognition processing of the speech recognition unit 103 .
  • FIG. 3 is a diagram showing the speech recognition performance for emphasized speech in which mixed speech is emphasized.
  • FIG. 4 is a diagram showing speech recognition performance for mixed speech. In both FIGS. 3 and 4, the character error rate was evaluated as speech recognition performance.
  • the speech recognition input determination unit 102 selects emphasized speech for regions of SNR and SIR that can improve the accuracy of speech recognition in FIG. For other regions, mixed speech is selected and input to the speech recognition unit 103 .
  • the speech recognition input determination unit 102 weights the emphasized speech more than the mixed speech in the SNR and SIR regions that can improve the accuracy of speech recognition in FIG. , the mixed speech is weighted more heavily than the emphasized speech, and the weighted mixed speech and the emphasized speech are input to the speech recognition unit 103 .
  • FIG. 2 is a diagram schematically showing an example of the configuration of the speech recognition input determination section 102 shown in FIG. 1. As shown in FIG.
  • the speech recognition input determination unit 102 has an SIR-SNR acquisition unit 1021 (acquisition unit), a determination unit 1022 and a switching unit 1023.
  • the SIR-SNR acquisition unit 1021 acquires at least one of the SIR of the mixed speech, the SNR of the mixed speech, and the mixed speech from the mixed speech input to the system.
  • the SIR-SNR acquisition unit 1021 may acquire the SIR and SIR in units of a certain time such as frames, or in units of utterances.
  • the SIR-SNR acquisition unit 1021 estimates SIR and SNR using, for example, an estimation model for estimating SIR and SNR.
  • the estimation model is configured by a neural network or the like.
  • the estimation model is pre-learned with paired data of speech and SIR and SNR.
  • the speech enhancement unit 101 extracts the target speaker's speech and the other interfering speaker's speech
  • the SIR-SNR acquisition unit 1021 extracts the extracted target speaker's speech and the other speech.
  • SIR and SIR may be calculated using equations (1) and (2) based on the speech of the interfering speaker.
  • the determination unit 1022 Based on at least one of the SIR and SNR obtained by the SIR-SNR obtaining unit 1021 and the mixed speech, the determination unit 1022 obtains speech based on at least one of the mixed speech and the emphasized speech obtained by enhancing the mixed speech. is determined as the voice to be used for voice recognition. Based on at least one of the SIR and SNR estimated or calculated by the SIR-SNR acquisition unit 1021 and the mixed speech, the determination unit 1022 determines which of the mixed speech and the emphasized mixed speech is speech recognition. determine if it is beneficial to
  • the determination unit 1022 uses a predetermined rule using SIR and SNR to determine speech based on at least one of mixed speech and enhanced speech as speech to be used for speech recognition.
  • the determination unit 1022 uses a rule to select the emphasized voice in the case of the expression (3), for example, when determining the voice by hardware.
  • the determination unit 1022 determines whether the mixed speech or the enhanced speech is more advantageous as an input for speech recognition depending on the SIR and SNR (see FIGS. 3 and 4). The determination may be made using a rule in which a weight value for mixed speech and a weight value for emphasized speech are set in advance according to the combination of .
  • the judgment unit 1022 uses, as an input of the speech recognition unit 103, a discrimination model for distinguishing speech based on at least one of the mixed speech and the emphasized speech, and distinguishes between the mixed speech and the emphasized speech. Speech based on at least one may be determined as speech to be used for speech recognition.
  • a discriminative model consists of a neural network, etc.
  • the discriminative model is at least the feature obtained from the mixed speech to which the teacher label indicating which of the mixed speech and the enhanced speech is advantageous as the speech to be used for speech recognition, and the SNR and SIR. One of them is learned in advance as an input.
  • the discriminative model takes SNR and SIR as inputs, a teacher label indicating which of mixed speech and enhanced speech is more advantageous as speech to be used for speech recognition, and estimated or calculated SNR and SIR pair data. Based on this, the discriminative model is learned.
  • the discriminative model can also directly estimate which of the mixed speech and the enhanced speech is advantageous for speech recognition by using the feature values obtained from the mixed speech as input.
  • the training of the discriminant model is based on the teacher label that indicates which of the mixed speech and the emphasized speech is more advantageous as the speech to be used for speech recognition, and the paired data of the feature values extracted from the mixed speech. I do.
  • the discrimination model inputs both SNR and SIR and mixed speech, discrimination is performed based on teacher label pair data indicating which of these, mixed speech, and enhanced speech is more advantageous as speech to be used for speech recognition. Train the model.
  • the discriminative model outputs weights that indicate the advantages of mixed speech and enhanced speech as speech used for speech recognition. Further, the discrimination model may output a determination result in which either mixed speech or emphasized speech is selected as speech that is advantageous as speech to be used for speech recognition.
  • the switching unit 1023 inputs to the voice recognition unit 103 the voice based on at least one of the mixed voice and the emphasized voice.
  • speech determination is performed by hardware as shown in FIG.
  • the switching unit 1023 weights the mixed speech and the emphasized speech according to the weight output from the judgment unit 1022, and divides the weighted mixed speech and the emphasized speech into weighted speech. is output to the speech recognition unit 103 .
  • FIG. 5 is a flowchart showing a processing procedure of signal processing according to the embodiment.
  • the speech enhancement unit 101 When input of mixed speech is received (step S1), the speech enhancement unit 101 extracts the speech of the target speaker from the input mixed speech, and inputs the extracted speech as emphasized speech for speech recognition. Speech enhancement processing to be input to the determination unit 102 is performed (step S2).
  • the speech recognition input determination unit 102 determines speech based on at least one of the mixed speech and the emphasized speech obtained by emphasizing the mixed speech. , a voice recognition input determination process is performed to determine that the voice is to be used for voice recognition (step S3).
  • the speech recognition unit 103 performs speech recognition processing using speech based on at least one of the mixed speech and the emphasized speech determined by the speech recognition input determination unit 102 (step S4), and obtains a speech recognition result. is output (step S5).
  • FIG. 6 is a flow chart showing a processing procedure of the speech recognition input determination process shown in FIG.
  • the SIR-SNR acquisition unit 1021 performs an acquisition process of acquiring at least one of the SIR of the mixed speech, the SNR of the mixed speech, and the mixed speech from the mixed speech input to the system. (Step S11).
  • the determination unit 1022 obtains speech based on at least one of the mixed speech and the emphasized speech obtained by enhancing the mixed speech. is determined as the voice to be used for voice recognition (step S12).
  • the switching unit 1023 inputs the speech based on at least one of the mixed speech and the emphasized speech to the speech recognition unit 103 as the speech to be used for speech recognition based on the judgment result of the judgment unit 1022 (step S13).
  • the SIR and SNR and the mixed speech Speech based on at least one of the mixed speech and the enhanced speech is determined as the speech to be used for speech recognition.
  • FIG. 7 is a diagram showing an example of speech recognition results when the signal processing device 100 according to the embodiment is applied.
  • FIG. 8 is a diagram showing an example of the performance improvement rate of the speech recognition result using the signal processing apparatus 100 according to the embodiment with respect to the speech recognition result when using the emphasized speech.
  • 7 and 8 show speech recognition results when the speech to be input to the speech recognition unit 103 is selected using the rule that the emphasized speech is selected in the case of expression (3) from the emphasized speech and the mixed speech. be. In both FIGS. 7 and 8, the character error rate was evaluated as speech recognition performance.
  • the area of the frame W1 shown in FIG. 7 and the area of the frame W11 shown in FIG. 8 are areas where the input signal is considered to be selected as mixed speech based on the estimated SIR and SNR.
  • Signal processing device 100 that selects a signal to speech recognition unit 103 using a rule that emphasized speech is selected in the case of equation (3) from the speech recognition results of frame W1 shown in FIG. 7 and frame W11 shown in FIG. It was found that the speech recognition accuracy is higher in the case of using only the emphasized speech.
  • the signal processing device 100 has higher speech recognition accuracy than the case where only the emphasized speech is used for all combinations of SIR and SNR. Furthermore, it is clear that the signal processing device 100 has particularly high speech recognition accuracy in areas where the SIR value is relatively large (see, for example, frames W11 and W12) compared to when only emphasized speech is used. Became.
  • the signal processing apparatus 100 can prevent performance deterioration due to distortion of speech enhancement and improve speech recognition accuracy by appropriately determining the input speech that can improve the accuracy of speech recognition.
  • the speech enhancement unit 101 does not need to perform speech enhancement processing when mixed speech is selected as input speech that can improve the accuracy of speech recognition. In this way, the speech enhancement unit 101 performs speech enhancement only for a section that requires speech enhancement processing, that is, a section for which the speech recognition input determination unit 102 selects an enhanced speech as an input speech that can improve the accuracy of speech recognition. Perform emphasis processing.
  • Each component of the signal processing device 100 is functionally conceptual, and does not necessarily need to be physically configured as illustrated. That is, the specific form of distribution and integration of the functions of the signal processing device 100 is not limited to the illustrated one, and all or part of it can be functionally or physically distributed in arbitrary units according to various loads and usage conditions. can be distributed or integrated.
  • each process performed in the signal processing device 100 may be realized by a CPU, a GPU (Graphics Processing Unit), and a program that is analyzed and executed by the CPU and GPU. Further, each process performed in the signal processing device 100 may be realized as hardware by wired logic.
  • FIG. 9 is a diagram showing an example of a computer that implements the signal processing device 100 by executing a program.
  • the computer 1000 has a memory 1010 and a CPU 1020, for example.
  • Computer 1000 also has hard disk drive interface 1030 , disk drive interface 1040 , serial port interface 1050 , video adapter 1060 and network interface 1070 . These units are connected by a bus 1080 .
  • the memory 1010 includes a ROM 1011 and a RAM 1012.
  • the ROM 1011 stores a boot program such as BIOS (Basic Input Output System).
  • BIOS Basic Input Output System
  • Hard disk drive interface 1030 is connected to hard disk drive 1090 .
  • a disk drive interface 1040 is connected to the disk drive 1100 .
  • a removable storage medium such as a magnetic disk or optical disk is inserted into the disk drive 1100 .
  • Serial port interface 1050 is connected to mouse 1110 and keyboard 1120, for example.
  • Video adapter 1060 is connected to display 1130, for example.
  • the hard disk drive 1090 stores an OS (Operating System) 1091, application programs 1092, program modules 1093, and program data 1094, for example. That is, a program that defines each process of the signal processing device 100 is implemented as a program module 1093 in which code executable by the computer 1000 is described. Program modules 1093 are stored, for example, on hard disk drive 1090 .
  • the hard disk drive 1090 stores a program module 1093 for executing processing similar to the functional configuration of the signal processing device 100 .
  • the hard disk drive 1090 may be replaced by an SSD (Solid State Drive).
  • the setting data used in the processing of the above-described embodiment is stored as program data 1094 in the memory 1010 or the hard disk drive 1090, for example. Then, the CPU 1020 reads out the program module 1093 and the program data 1094 stored in the memory 1010 and the hard disk drive 1090 to the RAM 1012 as necessary and executes them.
  • the program modules 1093 and program data 1094 are not limited to being stored in the hard disk drive 1090, but may be stored in a removable storage medium, for example, and read by the CPU 1020 via the disk drive 1100 or the like. Alternatively, the program modules 1093 and program data 1094 may be stored in another computer connected via a network (LAN (Local Area Network), WAN (Wide Area Network), etc.). Program modules 1093 and program data 1094 may then be read by CPU 1020 through network interface 1070 from other computers.
  • LAN Local Area Network
  • WAN Wide Area Network

Landscapes

  • Engineering & Computer Science (AREA)
  • Computational Linguistics (AREA)
  • Health & Medical Sciences (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Human Computer Interaction (AREA)
  • Physics & Mathematics (AREA)
  • Acoustics & Sound (AREA)
  • Multimedia (AREA)
  • Signal Processing (AREA)
  • Quality & Reliability (AREA)
  • Circuit For Audible Band Transducer (AREA)
  • Telephonic Communication Services (AREA)

Abstract

音声認識入力判定部(102)は、目的話者の音声に別の話者の音声が重複している混合音声から、混合音声における、目的音声と干渉話者音声との比率であるSIR(Signal to Interference Ratio)及び混合音声の、目的音声と雑音との比率であるSNR(Signal to Noise Ratio)と、混合音声とのうち少なくとも一つを取得するSIR-SNR取得部(1021)と、SIR及びSNRと、前記混合音声とのうち少なくとも一つを基に、混合音声と、混合音声を強調した強調音声との少なくとも一方に基づく音声を、音声認識に使用する音声として判定する判定部(1022)と、を有する。

Description

信号処理装置、信号処理方法及び信号処理プログラム
 本発明は、信号処理装置、信号処理方法及び信号処理プログラムに関する。
 深層学習技術の発達により音声認識精度は向上し、音声認識の適用範囲は徐々に拡大している。しかしながら、現在でも複数人の声が重複して発せられた場合それを認識するのは困難である。
 そこで、音声の重複に対処するための種々の技術が考案されている。ブラインド音源分離技術はそのうちの一つのアプローチであり、混合音を含まれる各音源に由来する音に分離を行うものである。ブラインド音源分離技術は、混合音のままでは音声認識が困難な音声を、各話者の音声に分離することで音声認識を可能にすることができる(例えば、非特許文献1参照)。
 また、目的話者抽出技術が提案されている。目的話者抽出技術は、目的話者が事前登録した発話を補助的な情報として利用し、事前登録された話者の音声のみを、混合音から取得する技術である。この目的話者抽出技術は、混合音に含まれる話者数に関する事前情報を必要としない利点があり、実用上有用な技術である(例えば、非特許文献2参照)。目的話者抽出技術を用いて抽出した音声は目的話者の声だけを含むことから、音声認識が可能である。
Dong Yu, et al, "PERMUTATION INVARIANT TRAINING OF DEEP MODELS FOR SPEAKER-INDEPENDENT MULTI-TALKER SPEECH SEPARATION", 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2017. Zmolikova Katerina, et al, "SpeakerBeam: Speaker Aware Neural Network for Target Speaker Extraction in Speech Mixtures", IEEE Journal of Selected Topics in Signal Processing 13.4 (2019): 800-814. Quan Wang, et al, "VoiceFilter-Lite: Streaming Targeted Voice Separation for On-Device Speech Recognition", arXiv preprint arXiv:2009.04323 (2020).
 前述のとおり音声強調技術は、目的話者の音声が他の話者の音声と重複している場合や、背景雑音と重複している場合に、それら望ましくない音を除去し、目的話者の音声のみを取り出すことを可能とする技術である。しかしながら、望ましくない音を除去する際に、同時に所望の目的話者音声を歪ませてしまう場合がある。
 音声強調技術を適用した音声に対して、音声認識を行う場合、この目的話者音声の歪みが、音声認識性能の劣化を引き起こすおそれがある。このような音声認識性能の劣化は、他話者の音声が必ずしも重複しない現実の音声に対して目的話者抽出を行う実運用上大きな課題である。
 一方、音声認識は、ある程度の雑音であれば小さな性能劣化で動作することから、雑音の程度や種類によっては、むしろ目的話者抽出を行わない方が、音声認識性能が良い場合もある。すなわち、音声強調技術を適用することによる、雑音除去による効果と、それに伴うひずみの悪影響とを比較して、前者の方が大きい場合雑音除去を行うことが望ましく、後者の方が大きい場合音声強調を行わないことが望ましい。
 例えば、非特許文献3では、雑音が他話者の音声である場合(音声雑音)に音声強調は効果的である一方、音声ではない場合(非音声雑音)に音声強調が音声認識を劣化させる可能性が高いという知見に基づき、雑音の種類が他の話者の音声か、その他の雑音かを識別する二値識別機を導入し、雑音の種類が他の話者の音声であった場合のみ、音声強調を行う手法を提案している。すなわち、目的話者の音声に別の話者の音声が存在しない場合は音声強調をせず、重複した場合にのみ音声強調を行うことを提案している。
 しかしながら、複数話者の音声が重複して発せられた場合であっても、干渉話者音声の存在による性能劣化を、それを除去することによって生じる歪みによって生じる性能劣化が上回れば音声強調が劣化する場合がある。
 本発明は、上記に鑑みてなされたものであって、目的話者の音声に別の話者の音声が重複している場合に、音声認識の精度を高めることができる入力音声として、混合音声と、混合音声を強調した強調音声との少なくとも一方を適切に判定することができる信号処理装置、信号処理方法及び信号処理プログラムを提供することを目的とする。
 上述した課題を解決し、目的を達成するために、本発明に係る信号処理装置は、目的話者の音声に別の話者の音声が重複している混合音声から、混合音声における、目的音声と干渉話者音声との比率であるSIR(Signal to Interference Ratio)、及び、混合音声の、目的音声と雑音との比率であるSNR(Signal to Noise Ratio)と、混合音声とのうちの少なくとも一方を取得する取得部と、SIR及びSNRと、混合音声とのうち少なくとも一つを基に、混合音声と、混合音声を強調した強調音声との少なくとも一方に基づく音声を、音声認識に使用する音声として判定する判定部と、を有することを特徴とする。
 本発明によれば、目的話者の音声に別の話者の音声が重複している場合に、音声認識の精度を高めることができる入力音声として、混合音声と、混合音声を強調した強調音声との少なくとも一方を適切に判定することができる。
図1は、実施の形態に係る信号処理装置の構成の一例を模式的に示す図である。 図2は、図1に示す音声認識入力判定部の構成の一例を模式的に示す図である。 図3は、混合音声を強調した強調音声に対する音声認識性能を示す図である。 図4は、混合音声に対する音声認識性能を示す図である。 図5は、実施の形態に係る信号処理の処理手順を示すフローチャートである。 図6は、図5に示す音声認識入力判定処理の処理手順を示すフローチャートである。 図7は、実施の形態に係る信号処理装置を適用した場合の音声認識結果を示す図である。 図8は、強調音声を用いた場合の音声認識結果に対する、実施の形態に係る信号処理装置を用いた音声認識結果の性能向上率を示す図である。 図9は、プログラムが実行されることにより、信号処理装置が実現されるコンピュータの一例を示す図である。
 以下、図面を参照して、本発明の一実施形態を詳細に説明する。なお、この実施形態により本発明が限定されるものではない。また、図面の記載において、同一部分には同一の符号を付して示している。
[実施の形態]
 目的話者の音声に別の話者の音声が重複した場合であっても、音声強調を行わない方がよい場合がありうる。例えば、干渉話者の音声が目的話者の音声と比較して振幅が小さい場合、音声強調せずとも音声認識がある程度目的話者の音声を認識することが可能である場合がある。また、干渉話者の音声の他に、非音声雑音が存在する場合、音声強調による歪みの影響は大きくなる傾向が確認されており、このような場合については音声強調の悪影響が大きく、もともとの混合音声を認識したほうが音声認識にとって有利であることがある。
 そこで、実施の形態に係る信号処理装置では、目的話者の音声に別の話者の音声が重複している混合音声を対象とし、音声認識の精度を高めることができる入力音声として、混合音声と、混合音声を強調した強調音声との少なくとも一方に基づく音声を判定する。実施の形態に係る信号処理装置では、混合音声における、目的音声と雑音との比率であるSNR(Signal to Noise Ratio)、及び、混合音声の、目的音声と干渉話者音声との比率であるSIR(Signal to Interference Ratio)と、混合音声との少なくとも一つを基に、混合音声と強調音声との少なくとも一方に基づく音声を判定する。実施の形態に係る信号処理装置では、判定した混合音声と強調音声との少なくとも一方に基づく音声を用いて、音声認識処理を行うことで、音声認識精度の向上を図ることができる。
[信号処理装置]
 次に、実施の形態に係る信号処理装置について説明する。図1は、実施の形態に係る信号処理装置の構成の一例を模式的に示す図である。
 信号処理装置100は、例えば、ROM(Read Only Memory)、RAM(Random Access Memory)、CPU(Central Processing Unit)等を含むコンピュータ等に所定のプログラムが読み込まれて、CPUが所定のプログラムを実行することで実現される。また、信号処理装置100は、有線接続、或いは、ネットワーク等を介して接続された他の装置との間で、各種情報を送受信する通信インタフェースを有する。信号処理装置100は、音声強調部101、音声認識入力判定部102及び音声認識部103を有する。
 音声強調部101は、公知の音声強調技術を用いて、入力される混合音声から、所望の音声を抽出する。音声強調部101の処理として、公知の音源分離技術や、公知の目的話者抽出技術を適用することができる。音源分離技術では、入力された混合音声を各話者に分離するため、各話者の音声が独立に音声認識入力判定部102に与えられる。目的話者抽出技術では、入力として、目的話者に関する補助情報も併せて与えられる必要がある。目的話者に関する補助情報として、例えば目的話者が事前に登録した発話が用いられる。なお、目的話者に関する補助情報は、信号処理装置100が保持してもよく、信号処理装置100と通信可能である外部装置から入力されるものであってもよい。
 音声強調部101は、混合音声を強調した強調音声を抽出する。例えば混合音声として、目的話者と他の干渉話者の音声の混合音声が音声強調部101に入力された場合、音声強調部101は、混合音声から、事前登録発話を発した目的話者の音声を抽出する。音声強調部101は、抽出した音声を強調音声として音声認識入力判定部102のSIR-SNR取得部1021(後述)に入力する。入力される混合音声および強調音声としては、音声に対して特徴抽出などの何らかの変換を行ったものでもよいし、音声波形であってもよい。
 音声認識入力判定部102は、混合音声のSNR及び混合音声のSIRと、混合音声とのうち少なくとも一つを基に、混合音声と、混合音声を強調した強調音声との少なくとも一方に基づく音声を、音声認識に使用する音声として判定する。音声認識入力判定部102は、判定した混合音声と強調音声との少なくとも一方に基づく音声を音声認識部103に入力する。音声認識入力判定部102は、音声認識に入力する発話を、音声強調部101の出力である強調音声と、信号処理装置100に入力された混合音声との少なくとも一方に基づく音声を、発話単位ないしはフレームなど一定区間単位で判定する。
 具体的に、音声認識入力判定部102は、音声の判定をハードに行う。例えば、音声認識入力判定部102は、「1」、「0」で、音声強調部101の出力である強調音声と、信号処理装置100に入力された混合音声とのいずれか一方のみを音声認識部103に入力する。
 また、音声認識入力判定部102は、音声の判定をソフトに行うこともできる。例えば、音声認識入力判定部102は、音声強調部101の出力である強調音声と、信号処理装置100に入力された混合音声との一方に例えば「0.3」の重みづけを行い、他方に「0.7」の重みづけを行い、重みづけ後の混合音声と強調音声とを音声認識部103に入力する。
 音声認識部103は、音声認識入力判定部102が判定した、混合音声と強調音声との少なくとも一方に基づく音声を用いて、音声認識処理を行う。音声認識部103の音声認識処理には、公知の音声認識技術を適用することができる。
[SIR、SNR]
 ここで、SIRの演算式を式(1)に示し、SNRの演算式を式(2)に示す。式(1)、式(2)において、Sは、目的話者の音声であり、Iは、干渉話者の音声であり、Nは、雑音である。
Figure JPOXMLDOC01-appb-M000001
Figure JPOXMLDOC01-appb-M000002
 図3は、混合音声を強調した強調音声に対する音声認識性能を示す図である。図4は、混合音声に対する音声認識性能を示す図である。図3及び図4のいずれも、音声認識性能として、文字誤り率(Character Error Rate)を評価した。
 図3及び図4に示すように、音声認識では、入力として、混合音声を強調した強調音声と、もともとの混合音声のいずれかが有利かは、SIR及びSNRに依存することが実験的に判明している。具体的には、SIRからSNRを減じた値が大きいほど、強調音声よりも混合音声を入力とした方が、文字誤り率が低い。
 このため、音声認識入力判定部102は、SIR及びSNRを入力音声から推定することで、図3において音声認識の精度を高めることができるようなSNR及びSIRの領域に対しては強調音声を選択し、そうでない領域については混合音声を選択して、音声認識部103に入力する。或いは、音声認識入力判定部102は、図3において音声認識の精度を高めることができるようなSNR及びSIRの領域に対しては強調音声に、混合音声よりも大きな重みづけを行い、そうでない領域については混合音声に、強調音声よりも大きな重みづけを行い、重みづけ後の混合音声と強調音声とを音声認識部103に入力する。
[音声認識入力判定部]
 次に、音声認識入力判定部102について説明する。図2は、図1に示す音声認識入力判定部102の構成の一例を模式的に示す図である。
 図2に示すように、音声認識入力判定部102は、SIR-SNR取得部1021(取得部)、判定部1022及びスイッチング部1023を有する。
 SIR-SNR取得部1021は、システムに入力された混合音声から、混合音声におけるSIR、及び、混合音声におけるSNRと、混合音声とのうち少なくとも一つを取得する。SIR-SNR取得部1021は、SIR及びSIRを、フレームなど一定時間単位で取得してもよいし、発話単位で取得してもよい。
 SIR-SNR取得部1021は、例えば、SIR及びSNRを推定する推定モデルを用いて、SIR及びSNRを推定する。推定モデルは、ニューラルネットワークなどによって構成される。推定モデルは、音声とSIR及びSNRとのペアデータを予め学習させたものである。
 また、音声強調部101によって、目的話者の音声と、それ以外の干渉話者の音声とが抽出された場合、SIR-SNR取得部1021は、抽出された目的話者の音声と、それ以外の干渉話者の音声を基に、式(1)及び式(2)を用いて、SIR及びSIRを計算してもよい。
 判定部1022は、SIR-SNR取得部1021によって取得されたSIR及びSNRと、混合音声とのうち少なくとも一つを基に、混合音声と、混合音声を強調した強調音声との少なくとも一方に基づく音声を音声認識に使用する音声として判定する。判定部1022は、SIR-SNR取得部1021によって推定或いは計算されたSIR及びSNRと、混合音声とのうち少なくとも一つを基に、混合音声と、混合音声を強調した強調音声とのどちらが音声認識にとって有利かを判定する。
 例えば、判定部1022は、SIR及びSNRを用いた所定のルールを用いて、混合音声と強調音声との少なくとも一方に基づく音声を、音声認識に使用する音声として判定する。判定部1022は、音声の判定をハードに行う場合、例えば、式(3)の場合に強調音声を選択するというルールを用いる。
Figure JPOXMLDOC01-appb-M000003
 また、判定部1022は、音声認識の入力として、混合音声と強調音声とのいずれかが有利かであるかがSIR及びSNRに依存すること(図3,4参照)を基に、SIRとSNRとの組み合わせに応じて、混合音声に対する重み値と強調音声とに対する重み値を予め設定したルールを用いて判定を行ってもよい。
 判定部1022は、音声の判定をソフトに行う場合、音声認識部103の入力として、混合音声と強調音声との少なくとも一方に基づく音声を識別する識別モデルを用いて、混合音声と強調音声との少なくとも一方に基づく音声を、音声認識に使用する音声として判定してもよい。
 識別モデルは、ニューラルネットワークなどから構成される。識別モデルは、混合音声と強調音声とのいずれが音声認識に使用する音声として有利であるかを示す教師ラベルが付与された、混合音声から取得された特徴量と、SNR及びSIRと、の少なくとも一方を入力として予め学習させたものである。
 識別モデルがSNR及びSIRを入力とする場合は、混合音声と強調音声とのいずれが音声認識に使用する音声として有利であるかを示す教師ラベルと、推定あるいは計算されたSNR及びSIRのペアデータとをもとに、識別モデルの学習を行う。
 識別モデルは、混合音声から取得された特徴量を入力として、混合音声と強調音声とのいずれが音声認識にとって有利であるかを直接推定することも可能である。その場合は混合音声と強調音声とのいずれが音声認識に使用する音声として有利であるかを示す教師ラベルと、混合音声から抽出された特徴量のペアデータとをもとに、識別モデルの学習を行う。
 識別モデルがSNR及びSIRと混合音声の両方を入力する場合は、それらと混合音声と強調音声とのいずれが音声認識に使用する音声として有利であるかを示す教師ラベルのペアデータを基に識別モデルの学習を行う。
 識別モデルは、混合音声及び強調音声に対する、音声認識に使用する音声としての有利さを示す重みを出力する。また、識別モデルは、音声認識に使用する音声として有利である音声として、混合音声または強調音声の一方を選択した判定結果を出力するものであってもよい。
 スイッチング部1023は、判定部1022の判定結果をもとに、混合音声と強調音声との少なくとも一方に基づく音声を音声認識部103へ入力する。図4のように音声の判定をハードに行う場合、スイッチング部1023は、判定部1022の選択結果をもとに、強調音声または混合音声を音声認識部103へ入力する。音声の判定をソフトに行う場合(不図示)、スイッチング部1023は、判定部1022が出力した重みに応じて混合音声と強調音声とに重みづけを行い、重みづけ後の混合音声と強調音声とを音声認識部103に出力する。
[信号処理の処理手順]
 次に、信号処理装置100が実行する信号処理について説明する。図5は、実施の形態に係る信号処理の処理手順を示すフローチャートである。
 混合音声の入力を受け付けると(ステップS1)、音声強調部101は、入力された混合音声から、目的話者の音声を抽出し、抽出した音声を、混合音声を強調した強調音声として音声認識入力判定部102に入力する音声強調処理を行う(ステップS2)。
 音声認識入力判定部102は、混合音声のSNR及び混合音声のSIRと、混合音声とのうち少なくとも一つを基に、混合音声と、混合音声を強調した強調音声との少なくとも一方に基づく音声を、音声認識に使用する音声として判定する音声認識入力判定処理を行う(ステップS3)。
 音声認識部103は、音声認識入力判定部102が判定した、混合音声と強調音声との少なくとも一方に基づく音声を用いて、音声認識処理を行う音声認識処理を行い(ステップS4)、音声認識結果を出力する(ステップS5)。
[音声認識入力判定処理の処理手順]
 次に、図5の音声認識入力判定処理(ステップS3)について説明する。図6は、図5に示す音声認識入力判定処理の処理手順を示すフローチャートである。
 図6に示すように、SIR-SNR取得部1021は、システムに入力された混合音声から、混合音声におけるSIR及び混合音声におけるSNRと、混合音声とのうち少なくとも一つを取得する取得処理を行う(ステップS11)。
 判定部1022は、SIR-SNR取得部1021によって取得されたSIR及びSNRと、混合音声とのうち少なくとも一つを基に、混合音声と、混合音声を強調した強調音声との少なくとも一方に基づく音声を、音声認識に使用する音声として判定する判定処理を行う(ステップS12)。スイッチング部1023は、判定部1022の判定結果をもとに、混合音声と強調音声との少なくとも一方に基づく音声を、音声認識に使用する音声として音声認識部103へ入力する(ステップS13)。
[実施の形態の効果]
 このように、信号処理装置100では、音声認識の入力として、混合音声と強調音声とのいずれかが有利かであるかがSIR及びSNRに依存することに基づいて、SIR及びSNRと、混合音声とのうち少なくとも一つを基に、混合音声と強調音声との少なくとも一方に基づく音声を、音声認識に使用する音声として判定する。
 図7は、実施の形態に係る信号処理装置100を適用した場合の音声認識結果の一例を示す図である。図8は、強調音声を用いた場合の音声認識結果に対する、実施の形態に係る信号処理装置100を用いた音声認識結果の性能向上率の一例を示す図である。図7及び図8は、強調音声及び混合音声のうち、式(3)の場合に強調音声を選択するというルールを用いて、音声認識部103へ入力する音声を選択した場合の音声認識結果である。なお、図7及び図8のいずれも、音声認識性能として、文字誤り率(Character Error Rate)を評価した。
 図7に示す枠W1、図8に示す枠W11の領域は、推定したSIR、SNRを基づいて、入力信号が混合音声に選択されていると思われる領域である。図7に示す枠W1、図8に示す枠W11の音声認識結果より、式(3)の場合に強調音声を選択するというルールを用いて音声認識部103への信号を選択する信号処理装置100の方が、強調音声のみを用いた場合と比して、音声認識精度が高いことが分かった。
 SIR、SNRの組み合わせ全体において、平均すると、信号処理装置100の方が、強調音声のみを用いた場合と比して、音声認識精度が高いことが分かった。さらに、SIRの値が比較的大きい領域(例えば、枠W11,W12参照)において、信号処理装置100の方が、強調音声のみを用いた場合と比して、音声認識精度が特に高いことが明らかになった。
 これより、信号処理装置100によれば、目的話者の音声に別の話者の音声が重複している場合であっても、音声認識の精度を高めることができる入力音声として、混合音声と、混合音声を強調した強調音声との少なくとも一方に基づく音声を適切に判定する。そして、信号処理装置100は、音声認識の精度を高めることができる入力音声を適切に判定することで、音声強調の歪みによる性能劣化を防ぎ、音声認識精度を高めることができる。
 なお、本実施の形態では、音声強調部101は、音声認識の精度を高めることができる入力音声として混合音声が選択された場合、音声強調処理を行わなくともよい。このように、音声強調部101は、音声強調処理が必要な区間、すなわち、音声認識入力判定部102によって、音声認識の精度を高めることができる入力音声として強調音声が選択された区間についてのみ音声強調処理を行う。
[実施の形態のシステム構成について]
 信号処理装置100の各構成要素は機能概念的なものであり、必ずしも物理的に図示のように構成されていることを要しない。すなわち、信号処理装置100の機能の分散及び統合の具体的形態は図示のものに限られず、その全部または一部を、各種の負荷や使用状況などに応じて、任意の単位で機能的または物理的に分散または統合して構成することができる。
 また、信号処理装置100においておこなわれる各処理は、全部または任意の一部が、CPU、GPU(Graphics Processing Unit)、及び、CPU、GPUにより解析実行されるプログラムにて実現されてもよい。また、信号処理装置100においておこなわれる各処理は、ワイヤードロジックによるハードウェアとして実現されてもよい。
 また、実施の形態において説明した各処理のうち、自動的におこなわれるものとして説明した処理の全部または一部を手動的に行うこともできる。もしくは、手動的におこなわれるものとして説明した処理の全部または一部を公知の方法で自動的に行うこともできる。この他、上述及び図示の処理手順、制御手順、具体的名称、各種のデータやパラメータを含む情報については、特記する場合を除いて適宜変更することができる。
[プログラム]
 図9は、プログラムが実行されることにより、信号処理装置100が実現されるコンピュータの一例を示す図である。コンピュータ1000は、例えば、メモリ1010、CPU1020を有する。また、コンピュータ1000は、ハードディスクドライブインタフェース1030、ディスクドライブインタフェース1040、シリアルポートインタフェース1050、ビデオアダプタ1060、ネットワークインタフェース1070を有する。これらの各部は、バス1080によって接続される。
 メモリ1010は、ROM1011及びRAM1012を含む。ROM1011は、例えば、BIOS(Basic Input Output System)等のブートプログラムを記憶する。ハードディスクドライブインタフェース1030は、ハードディスクドライブ1090に接続される。ディスクドライブインタフェース1040は、ディスクドライブ1100に接続される。例えば磁気ディスクや光ディスク等の着脱可能な記憶媒体が、ディスクドライブ1100に挿入される。シリアルポートインタフェース1050は、例えばマウス1110、キーボード1120に接続される。ビデオアダプタ1060は、例えばディスプレイ1130に接続される。
 ハードディスクドライブ1090は、例えば、OS(Operating System)1091、アプリケーションプログラム1092、プログラムモジュール1093、プログラムデータ1094を記憶する。すなわち、信号処理装置100の各処理を規定するプログラムは、コンピュータ1000により実行可能なコードが記述されたプログラムモジュール1093として実装される。プログラムモジュール1093は、例えばハードディスクドライブ1090に記憶される。例えば、信号処理装置100における機能構成と同様の処理を実行するためのプログラムモジュール1093が、ハードディスクドライブ1090に記憶される。なお、ハードディスクドライブ1090は、SSD(Solid State Drive)により代替されてもよい。
 また、上述した実施の形態の処理で用いられる設定データは、プログラムデータ1094として、例えばメモリ1010やハードディスクドライブ1090に記憶される。そして、CPU1020が、メモリ1010やハードディスクドライブ1090に記憶されたプログラムモジュール1093やプログラムデータ1094を必要に応じてRAM1012に読み出して実行する。
 なお、プログラムモジュール1093やプログラムデータ1094は、ハードディスクドライブ1090に記憶される場合に限らず、例えば着脱可能な記憶媒体に記憶され、ディスクドライブ1100等を介してCPU1020によって読み出されてもよい。あるいは、プログラムモジュール1093及びプログラムデータ1094は、ネットワーク(LAN(Local Area Network)、WAN(Wide Area Network)等)を介して接続された他のコンピュータに記憶されてもよい。そして、プログラムモジュール1093及びプログラムデータ1094は、他のコンピュータから、ネットワークインタフェース1070を介してCPU1020によって読み出されてもよい。
 以上、本発明者によってなされた発明を適用した実施の形態について説明したが、本実施の形態による本発明の開示の一部をなす記述及び図面により本発明は限定されることはない。すなわち、本実施の形態に基づいて当業者等によりなされる他の実施の形態、実施例及び運用技術等は全て本発明の範疇に含まれる。
 100 信号処理装置
 101 音声強調部
 102 音声認識入力判定部
 103 音声認識部
 1021 SIR-SNR取得部
 1022 判定部
 1023 スイッチング部

Claims (7)

  1.  目的話者の音声に別の話者の音声が重複している混合音声から、前記混合音声における、目的音声と干渉話者音声との比率であるSIR(Signal to Interference Ratio)、及び、前記混合音声の、目的音声と雑音との比率であるSNR(Signal to Noise Ratio)と、前記混合音声とのうち少なくとも一つを取得する取得部と、
     前記SIR及び前記SNRと、前記混合音声とのうち少なくとも一つを基に、前記混合音声と、前記混合音声を強調した強調音声との少なくとも一方に基づく音声を、音声認識に使用する音声として判定する判定部と、
     を有することを特徴とする信号処理装置。
  2.  前記判定部は、前記SIR及び前記SNRを用いた所定のルールを用いて、前記混合音声と前記強調音声との少なくとも一方に基づく音声を、前記音声認識に使用する音声として判定することを特徴とする請求項1に記載の信号処理装置。
  3.  前記判定部は、前記混合音声と前記強調音声とのいずれが音声認識に使用する音声として有利であるかを示す教師ラベルが付与された、前記混合音声から取得された特徴量と、前記SNR及び前記SIRと、の少なくとも一方を入力として学習された識別モデルを用いて、前記混合音声と前記強調音声との少なくとも一方に基づく音声を、前記音声認識に使用する音声として判定することを特徴とする請求項1に記載の信号処理装置。
  4.  前記判定部が判定した前記混合音声と前記強調音声との少なくとも一方に基づく音声を用いて、音声認識処理を行う音声認識部をさらに有することを特徴とする請求項1~3のいずれか一つに記載の信号処理装置。
  5.  前記混合音声から、前記目的話者の音声を抽出し、抽出した音声を前記強調音声として前記取得部に入力する音声強調部をさらに有することを特徴とする請求項1~4のいずれか一つに記載の信号処理装置。
  6.  信号処理装置が実行する信号処理方法であって、
     目的話者の音声に別の話者の音声が重複している混合音声から、前記混合音声における、目的音声と干渉話者音声との比率であるSIR(Signal to Interference Ratio)、及び、前記混合音声の、目的音声と雑音との比率であるSNR(Signal to Noise Ratio)と、前記混合音声とのうち少なくとも一つを取得する工程と、
     前記SIR及び前記SNRと、前記混合音声とのうち少なくとも一つを基に、前記混合音声と、前記混合音声を強調した強調音声との少なくとも一方に基づく音声を、音声認識に使用する音声として判定する工程と、
     を含んだことを特徴とする信号処理方法。
  7.  目的話者の音声に別の話者の音声が重複している混合音声から、前記混合音声における、目的音声と干渉話者音声との比率であるSIR(Signal to Interference Ratio)、及び、前記混合音声の、目的音声と雑音との比率であるSNR(Signal to Noise Ratio)と、前記混合音声とのうち少なくとも一つを取得するステップと、
     前記SIR及び前記SNRと、前記混合音声とのうち少なくとも一つを基に、前記混合音声と、前記混合音声を強調した強調音声との少なくとも一方に基づく音声を、音声認識に使用する音声として判定するステップと、
     をコンピュータに実行させるための信号処理プログラム。
PCT/JP2021/019874 2021-05-25 2021-05-25 信号処理装置、信号処理方法及び信号処理プログラム Ceased WO2022249302A1 (ja)

Priority Applications (3)

Application Number Priority Date Filing Date Title
PCT/JP2021/019874 WO2022249302A1 (ja) 2021-05-25 2021-05-25 信号処理装置、信号処理方法及び信号処理プログラム
JP2023523777A JP7635837B2 (ja) 2021-05-25 2021-05-25 信号処理装置、信号処理方法及び信号処理プログラム
US18/561,727 US20240274149A1 (en) 2021-05-25 2021-05-25 Signal processing device, signal processing method, and signal processing program

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
PCT/JP2021/019874 WO2022249302A1 (ja) 2021-05-25 2021-05-25 信号処理装置、信号処理方法及び信号処理プログラム

Publications (1)

Publication Number Publication Date
WO2022249302A1 true WO2022249302A1 (ja) 2022-12-01

Family

ID=84228584

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/JP2021/019874 Ceased WO2022249302A1 (ja) 2021-05-25 2021-05-25 信号処理装置、信号処理方法及び信号処理プログラム

Country Status (3)

Country Link
US (1) US20240274149A1 (ja)
JP (1) JP7635837B2 (ja)
WO (1) WO2022249302A1 (ja)

Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2002236497A (ja) * 2001-02-08 2002-08-23 Alpine Electronics Inc ノイズリダクションシステム
JP2007264327A (ja) * 2006-03-28 2007-10-11 Matsushita Electric Works Ltd 浴室装置及びそれに用いる音声操作装置
JP2018205512A (ja) * 2017-06-02 2018-12-27 富士通株式会社 電子機器及び雑音抑圧プログラム

Patent Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2002236497A (ja) * 2001-02-08 2002-08-23 Alpine Electronics Inc ノイズリダクションシステム
JP2007264327A (ja) * 2006-03-28 2007-10-11 Matsushita Electric Works Ltd 浴室装置及びそれに用いる音声操作装置
JP2018205512A (ja) * 2017-06-02 2018-12-27 富士通株式会社 電子機器及び雑音抑圧プログラム

Also Published As

Publication number Publication date
JP7635837B2 (ja) 2025-02-26
JPWO2022249302A1 (ja) 2022-12-01
US20240274149A1 (en) 2024-08-15

Similar Documents

Publication Publication Date Title
CN109643552B (zh) 用于可变噪声状况中语音增强的鲁棒噪声估计
CN105405439B (zh) 语音播放方法及装置
CN106486131B (zh) 一种语音去噪的方法及装置
KR102410392B1 (ko) 실행 중 범위 정규화를 이용하는 신경망 음성 활동 검출
JP6243858B2 (ja) 音声モデル学習方法、雑音抑圧方法、音声モデル学習装置、雑音抑圧装置、音声モデル学習プログラム及び雑音抑圧プログラム
US10748544B2 (en) Voice processing device, voice processing method, and program
WO2019062721A1 (zh) 语音身份特征提取器、分类器训练方法及相关设备
CN112002307B (zh) 一种语音识别方法和装置
CN111524527A (zh) 话者分离方法、装置、电子设备和存储介质
KR102275656B1 (ko) 적대적 학습(adversarial training) 모델을 이용한 강인한 음성 향상 훈련 방법 및 그 장치
KR100745976B1 (ko) 음향 모델을 이용한 음성과 비음성의 구분 방법 및 장치
WO2018176894A1 (zh) 一种说话人确认方法及装置
JP6348427B2 (ja) 雑音除去装置及び雑音除去プログラム
JP7176627B2 (ja) 信号抽出システム、信号抽出学習方法および信号抽出学習プログラム
KR101618512B1 (ko) 가우시안 혼합모델을 이용한 화자 인식 시스템 및 추가 학습 발화 선택 방법
Almajai et al. Using audio-visual features for robust voice activity detection in clean and noisy speech
Mun et al. The sound of my voice: Speaker representation loss for target voice separation
WO2020195924A1 (ja) 信号処理装置および方法、並びにプログラム
CN114299962A (zh) 基于音频流的对话角色分离方法、系统、设备及存储介质
Hou et al. Domain adversarial training for speech enhancement
JP2020071482A (ja) 語音分離方法、語音分離モデル訓練方法及びコンピュータ可読媒体
Nasir et al. A hybrid method for speech noise reduction using log-MMSE
JP7743875B2 (ja) 音声信号の処理方法、音声信号処理装置、およびプログラム
JP4787979B2 (ja) 雑音検出装置および雑音検出方法
Al-Ali et al. Enhanced forensic speaker verification using multi-run ICA in the presence of environmental noise and reverberation conditions

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 21942956

Country of ref document: EP

Kind code of ref document: A1

WWE Wipo information: entry into national phase

Ref document number: 2023523777

Country of ref document: JP

WWE Wipo information: entry into national phase

Ref document number: 18561727

Country of ref document: US

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 21942956

Country of ref document: EP

Kind code of ref document: A1