WO2024190063A1 - 情報処理装置、情報処理方法、及び、記録媒体 - Google Patents

情報処理装置、情報処理方法、及び、記録媒体 Download PDF

Info

Publication number
WO2024190063A1
WO2024190063A1 PCT/JP2024/000990 JP2024000990W WO2024190063A1 WO 2024190063 A1 WO2024190063 A1 WO 2024190063A1 JP 2024000990 W JP2024000990 W JP 2024000990W WO 2024190063 A1 WO2024190063 A1 WO 2024190063A1
Authority
WO
WIPO (PCT)
Prior art keywords
noise
speech
enhancement mask
mask
loss
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/JP2024/000990
Other languages
English (en)
French (fr)
Inventor
謙光 遠藤
仁 山本
ホン ヤン
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
NEC Corp
Original Assignee
NEC Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by NEC Corp filed Critical NEC Corp
Priority to JP2025506508A priority Critical patent/JPWO2024190063A5/ja
Publication of WO2024190063A1 publication Critical patent/WO2024190063A1/ja
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L21/00Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
    • G10L21/02Speech enhancement, e.g. noise reduction or echo cancellation
    • G10L21/0272Voice signal separating

Definitions

  • This disclosure relates to the technical fields of information processing devices, information processing methods, and recording media.
  • Non-Patent Document 1 describes a technology that estimates a speech enhancement mask and a noise enhancement mask using a model, calculates an index using each estimated mask, and trains the model using the deviation between the calculated index and the ideal value of the index.
  • the objective of this disclosure is to provide an information processing device, an information processing method, and a recording medium that aim to improve upon the technology described in prior art documents.
  • One aspect of the information processing device includes a constraint loss calculation means for calculating a constraint loss using an estimated speech enhancement mask output by a speech enhancement mask estimation model to which noise-mixed speech is input, and an estimated noise enhancement mask output by a noise enhancement mask estimation model to which the noise-mixed speech is input, and a parameter update means for updating parameters included in the speech enhancement mask estimation model and parameters included in the noise enhancement mask estimation model using a speech enhancement mask loss indicating the difference between the estimated speech enhancement mask and a target speech enhancement mask calculated when the speech and noise in the noise-mixed speech are known, a noise enhancement mask loss indicating the difference between the estimated noise enhancement mask and a target noise enhancement mask, and the constraint loss.
  • One aspect of the information processing method is to calculate a constraint loss using an estimated speech enhancement mask output by a speech enhancement mask estimation model to which noise-mixed speech is input, and an estimated noise enhancement mask output by a noise enhancement mask estimation model to which the noise-mixed speech is input, and to update parameters included in the speech enhancement mask estimation model and parameters included in the noise enhancement mask estimation model using a speech enhancement mask loss indicating the difference between the estimated speech enhancement mask and a target speech enhancement mask calculated when the speech and noise in the noise-mixed speech are known, a noise enhancement mask loss indicating the difference between the estimated noise enhancement mask and a target noise enhancement mask, and the constraint loss.
  • a computer program is recorded to cause a computer to execute an information processing method for calculating a constraint loss using an estimated speech enhancement mask output by a speech enhancement mask estimation model to which noise-mixed speech is input, and an estimated noise enhancement mask output by a noise enhancement mask estimation model to which the noise-mixed speech is input, and updating parameters included in the speech enhancement mask estimation model and parameters included in the noise enhancement mask estimation model using a speech enhancement mask loss indicating the difference between the estimated speech enhancement mask and a target speech enhancement mask calculated when the speech and noise in the noise-mixed speech are known, a noise enhancement mask loss indicating the difference between the estimated noise enhancement mask and a target noise enhancement mask, and the constraint loss.
  • FIG. 1 is a block diagram showing the configuration of an information processing apparatus according to the first embodiment.
  • FIG. 2 is a block diagram showing the configuration of an information processing device according to the second embodiment.
  • FIG. 3 is a flowchart showing the flow of information processing operations of the information processing device in the second embodiment.
  • FIG. 4 is a block diagram showing the configuration of an information processing apparatus according to the third embodiment.
  • FIG. 5 is a block diagram showing the configuration of an information processing apparatus according to the fourth embodiment.
  • FIG. 6 is a block diagram showing the configuration of an information processing device according to the fifth embodiment.
  • FIG. 7 is a flowchart showing the flow of information processing operations of the information processing device according to the fifth embodiment.
  • FIG. 8 is a block diagram showing the configuration of an information processing apparatus according to the sixth embodiment.
  • FIG. 9 is a flowchart showing the flow of information processing operations of the information processing device in the sixth embodiment.
  • a first embodiment of an information processing device, an information processing method, and a recording medium will be described below.
  • the first embodiment of the information processing device, the information processing method, and the recording medium will be described using an information processing device 1 to which the first embodiment of the information processing device, the information processing method, and the recording medium is applied.
  • FIG. 1 is a block diagram showing the configuration of an information processing device 1 in the first embodiment. As shown in FIG. 1, the information processing device 1 includes a constraint loss calculation unit 11 and a parameter update unit 12.
  • the constraint loss calculation unit 11 calculates the constraint loss using the estimated speech emphasis mask output by the speech emphasis mask estimation model to which the noise-mixed speech is input, and the estimated noise emphasis mask output by the noise emphasis mask estimation model to which the noise-mixed speech is input.
  • the parameter update unit 12 updates the parameters included in the speech emphasis mask estimation model and the parameters included in the noise emphasis mask estimation model, using the speech emphasis mask loss indicating the difference between the estimated speech emphasis mask and the target speech emphasis mask, the noise emphasis mask loss indicating the difference between the estimated noise emphasis mask and the target noise emphasis mask, and the constraint loss.
  • the information processing device 1 in the first embodiment introduces a constraint loss calculated using an estimated speech enhancement mask output by a speech enhancement mask estimation model that relatively reduces the time and volume of a frequency band of noise other than the target speech included in the input speech, and an estimated noise enhancement mask output by a noise enhancement mask estimation model that relatively reduces the time and volume of a frequency band of the target speech included in the input speech.
  • the information processing device 1 updates parameters included in the speech enhancement mask estimation model using the difference between the estimated value and target value of the speech enhancement mask, the difference between the estimated value and target value of the noise enhancement mask, and the constraint loss, and can generate a speech enhancement mask estimation model that can perform speech enhancement with high accuracy.
  • audio refers to the target sound signal.
  • audio may be referred to as the "target signal.”
  • noise refers to the non-target sound signal.
  • non-target signal refers to the "non-target signal.”
  • the audio may be a sound that one wishes to focus on.
  • the audio may be a person's voice.
  • the audio may be a specific person's voice.
  • the specific person may be one or more people.
  • the specific person may be a person who is near a mechanism that captures sound, such as a microphone. "Audio" may represent different sound signals depending on the scene.
  • the "noise-mixed speech” is “speech” mixed with “noise.”
  • the “noise-mixed speech” includes “speech,” which is a target sound signal, and “noise,” which is a non-target sound signal.
  • the sound signals included in the "noise-mixed speech” are either “speech” or “noise.”
  • a sound signal that is not “speech” and is included in the “noise-mixed speech” is “noise.”
  • a sound signal that is not “noise” and is included in the “noise-mixed speech” is “speech.”
  • the input signal input to the information processing device 2 is “noise-mixed speech.”
  • the input signal input to the information processing device 2 is composed of "speech,” which is a target signal, and "noise,” which is a non-target signal.
  • Voice enhancement technology is a technology that makes the voice louder relative to the noise from a noisy voice.
  • This technology may be a technology that suppresses noise from a voice mixed with noise, emphasizes the voice, and provides a voice that is easier to hear.
  • This technology may be a technology that emphasizes only the voice from a voice mixed with noise, and provides a voice that is easier to hear.
  • the voice enhancement technology may be capable of removing noise from a telephone call voice in a high-noise situation.
  • the voice enhancement technology may also emphasize the voice of the other party in the call, enabling smoother communication.
  • the voice enhancement technology may also improve the recognition rate of a voice recognizer.
  • the mask type speech enhancement technology estimates a speech enhancement mask that specifies the time, frequency band, and amount of reduction in the volume of the noise to be reduced.
  • the mask type speech enhancement technology may be a technology that makes speech easier to hear by applying the speech enhancement mask to noise-mixed speech that contains noise.
  • the mask type speech enhancement technology may use machine learning to generate a speech enhancement mask estimation model that estimates the speech enhancement mask.
  • the speech enhancement mask estimation model created in this embodiment may be applied to situations where speech is emphasized.
  • the speech enhancement mask estimation model created in this embodiment is trained using a noise-mixed speech sound obtained by superimposing noise on a clean speech sound as an input signal.
  • the clean speech sound may be a sound with very little noise.
  • the machine learning in this embodiment may be deep learning. [2-1-3: Speech enhancement mask and noise enhancement mask]
  • the speech enhancement mask estimation model is a model that outputs a speech enhancement mask when noise-mixed speech is input.
  • the speech enhancement mask estimated by the speech enhancement mask estimation model may be called an estimated speech enhancement mask.
  • the estimated speech enhancement mask may also be called an estimated value.
  • the speech enhancement mask may be expressed as in the following Equation 1. [Formula 1]
  • S(t,f) 2 denotes the power of "speech"
  • N(t,f) 2 denotes the power of "noise”.
  • the speech enhancement mask is expressed as the ratio of "speech" power to the power of "noise-mixed speech.”
  • a speech enhancement mask estimation model is created through training in order to enhance speech, but it is expected that the training of the speech enhancement mask estimation model can be accelerated by training noise enhancement along with speech enhancement.
  • the noise enhancement mask relatively reduces the volume of the time and frequency band of the target voice included in the input voice.
  • the noise enhancement mask estimation model is a model that outputs a noise enhancement mask when noise-mixed voice is input.
  • the noise enhancement mask estimated by the noise enhancement mask estimation model may be called an estimated noise enhancement mask.
  • the estimated noise enhancement mask may also be called an estimated value.
  • the noise enhancement mask may be expressed as in the following Equation 2. [Formula 2]
  • the noise enhancement mask represents the proportion of "noise” power in the “noise-mixed speech” power. [2-1-4: Learning]
  • the information processing device 2 may create a speech enhancement mask estimation model capable of estimating an ideal speech enhancement mask by learning a speech enhancement mask estimation model and a noise enhancement mask estimation model.
  • the speech enhancement mask estimation model and the noise enhancement mask estimation model may be implemented by a neural network (NN).
  • the speech enhancement mask estimation model and the noise enhancement mask estimation model may be implemented by a recurrent neural network (RNN).
  • the information processing device 2 may update parameters included in the speech enhancement mask estimation model and parameters included in the noise enhancement mask estimation model according to estimates by the speech enhancement mask estimation model and estimates by the noise enhancement mask estimation model, and create a speech enhancement mask estimation model capable of estimating an ideal speech enhancement mask.
  • the ideal speech enhancement mask (sometimes referred to as the "target speech enhancement mask”) is a speech enhancement mask that is calculated when the speech and noise in the noise-mixed speech are known.
  • the ideal noise enhancement mask (sometimes referred to as the "target noise enhancement mask”) is a noise enhancement mask that is calculated when the speech and noise in the noise-mixed speech are known.
  • the teacher input signal is an input signal for which an ideal speech emphasis mask and a noise emphasis mask are prepared in advance.
  • the input signal may consist of "speech" which is a target signal and "noise" which is a non-target signal.
  • the ideal speech emphasis mask may be called an ideal speech emphasis mask or an ideal value.
  • the ideal noise emphasis mask may be called an ideal noise emphasis mask or an ideal value.
  • Information including the teacher input signal and the ideal value may be called teacher information.
  • the teacher information may be stored in the teacher information storage unit 222 described later.
  • the ideal value may be a target value (target value) when updating the parameters.
  • the arithmetic device 21 includes, for example, at least one of a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), and an FPGA (Field Programmable Gate Array).
  • the arithmetic device 21 reads a computer program.
  • the arithmetic device 21 may read a computer program stored in the storage device 22.
  • the arithmetic device 21 may read a computer program stored in a computer-readable and non-transient recording medium using a recording medium reading device (e.g., an input device 24 described later) not shown in the figure that is provided in the information processing device 2.
  • a recording medium reading device e.g., an input device 24 described later
  • the arithmetic device 21 may acquire (i.e., download or read) a computer program from a device (not shown) located outside the information processing device 2 via the communication device 23 (or other communication device).
  • the arithmetic device 21 executes the read computer program.
  • a logical functional block for executing the operation to be performed by the information processing device 2 is realized within the calculation device 21.
  • the calculation device 21 can function as a controller for realizing a logical functional block for executing the operation (in other words, processing) to be performed by the information processing device 2.
  • a constraint loss calculation unit 211 which is a specific example of a "constraint loss calculation means” described in the appendix described later
  • a parameter update unit 212 which is a specific example of a "parameter update means” described in the appendix described later
  • a speech enhancement mask estimation unit 213 which is a specific example of a "speech enhancement mask estimation means” described in the appendix described later
  • a noise enhancement mask estimation unit 214 which is a specific example of a "noise enhancement mask estimation means” described in the appendix described later
  • a speech enhancement mask loss calculation unit 215 which is a specific example of a "speech enhancement mask loss calculation means” described in the appendix described later
  • a noise enhancement mask loss calculation unit 216 which is a specific example of a "noise enhancement mask loss calculation means” described in the appendix described later
  • the speech enhancement mask estimation unit 213 estimates the speech enhancement mask using a speech enhancement mask estimation model, which is a specific example of the "speech enhancement mask estimation model" described in the appendix described later.
  • the noise enhancement mask estimation unit 214 estimates the noise enhancement mask using a noise enhancement mask estimation model, which is a specific example of the "noise enhancement mask estimation model” described in the appendix described later.
  • any of the speech enhancement mask estimation unit 213, the noise enhancement mask estimation unit 214, the speech enhancement mask loss calculation unit 215, the noise enhancement mask loss calculation unit 216, the combined loss calculation unit 217, and the noise mixed voice input unit 218 may not be realized in the calculation device 21.
  • the storage device 22 can store desired data.
  • the storage device 22 may temporarily store a computer program executed by the arithmetic device 21.
  • the storage device 22 may temporarily store data that the arithmetic device 21 temporarily uses when the arithmetic device 21 is executing a computer program.
  • the storage device 22 may store data that the information processing device 2 stores for a long period of time.
  • the storage device 22 may include at least one of a RAM (Random Access Memory), a ROM (Read Only Memory), a hard disk device, an optical magnetic disk device, an SSD (Solid State Drive), and a disk array device.
  • the storage device 22 may include a non-temporary recording medium.
  • the storage device 22 may realize a parameter storage unit 221 and a teacher information storage unit 222.
  • the parameter storage unit 221 stores parameters included in the speech enhancement mask estimation model and parameters included in the noise enhancement mask estimation model. However, the storage device 22 does not need to realize either the parameter storage unit 221 or the teacher information storage unit 222.
  • the communication device 23 is capable of communicating with devices external to the information processing device 2 via a communication network (not shown).
  • the communication device 23 may be a communication interface based on standards such as Ethernet (registered trademark), Wi-Fi (registered trademark), Bluetooth (registered trademark), and USB (Universal Serial Bus).
  • the input device 24 is a device that accepts information input to the information processing device 2 from outside the information processing device 2.
  • the input device 24 may include an operating device (e.g., at least one of a keyboard, a mouse, and a touch panel) that can be operated by an operator of the information processing device 2.
  • the input device 24 may include a reading device that can read information recorded as data on a recording medium that can be attached externally to the information processing device 2.
  • the output device 25 is a device that outputs information to the outside of the information processing device 2.
  • the output device 25 may output information as an image. That is, the output device 25 may include a display device (so-called a display) capable of displaying an image showing the information to be output.
  • the output device 25 may output information as sound. That is, the output device 25 may include an audio device (so-called a speaker) capable of outputting sound.
  • the output device 25 may output information on paper. That is, the output device 25 may include a printing device (so-called a printer) capable of printing desired information on paper. [2-3: Information Processing Operation Performed by Information Processing Device 2]
  • FIG. 3 is a flowchart showing the flow of the information processing operation performed by the information processing device 2.
  • the noise-mixed speech input unit 218 acquires the noise-mixed speech and inputs it to the speech enhancement mask estimation unit 213 and the noise enhancement mask estimation unit 214 (step S20).
  • the speech enhancement mask estimation unit 213 estimates a speech enhancement mask using a speech enhancement mask estimation model (step S21).
  • the speech enhancement mask estimation unit 213 outputs, as an estimated value, the estimated speech enhancement mask output by the speech enhancement mask estimation model to which the noise-mixed speech is input.
  • the noise enhancement mask estimation unit 214 estimates the noise enhancement mask using the noise enhancement mask estimation model (step S22).
  • the noise enhancement mask estimation unit 214 outputs, as an estimated value, the estimated noise enhancement mask output by the noise enhancement mask estimation model to which the noise-mixed speech is input. [2-3-1: Speech enhancement mask loss]
  • the speech enhancement mask loss calculation unit 215 calculates a speech enhancement mask loss using the estimated speech enhancement mask estimated by the speech enhancement mask estimation model and an ideal speech enhancement mask (step S23).
  • the speech enhancement mask loss indicates the difference between the estimated speech enhancement mask and the ideal speech enhancement mask.
  • the speech enhancement mask loss may be calculated by a speech enhancement mask loss function L S.
  • the speech enhancement mask loss function L S may be expressed as in the following formula 3. That is, the speech enhancement mask loss function L S may be expressed as the sum of squares of the difference between the estimated value and the ideal value in the frequency bin at each time.
  • M S represents the estimated value of the speech enhancement mask. Also, M S i represents the ideal value of the speech enhancement mask. [Formula 3] In the following, matters in the frequency bin at each time may be expressed using the term "time-frequency". [2-3-2: Noise Enhancement Mask Loss]
  • the noise enhancement mask loss calculation unit 216 calculates the noise enhancement mask loss using the estimated noise enhancement mask estimated by the noise enhancement mask estimation model and the ideal noise enhancement mask (step S24).
  • the noise enhancement mask loss indicates the difference between the estimated noise enhancement mask and the ideal noise enhancement mask.
  • the noise enhancement mask loss may be calculated by a noise enhancement mask loss function L N.
  • the noise enhancement mask loss function L N may be expressed as in the following formula 4. That is, the noise enhancement mask loss function L N may be expressed as the sum of squares of the difference between the estimated value and the ideal value in time-frequency.
  • M N represents the estimated value of the noise enhancement mask. Furthermore, M N i represents the ideal value of the noise enhancement mask. [Formula 4] [2-3-3: Restraint loss]
  • the constraint loss calculation unit 211 calculates the constraint loss using the estimated speech enhancement mask estimated by the speech enhancement mask estimation model and the estimated noise enhancement mask estimated by the noise enhancement mask estimation model (step S25).
  • the constraint loss indicates the difference between the estimated value of the sum (constraint condition) of the speech enhancement mask and the noise enhancement mask and the ideal value.
  • the loss function is designed so that a penalty is imposed when the constraint condition is violated, that is, as the value obtained from the following formula 6 becomes farther from "0.”
  • the constraint loss in this embodiment can be obtained from the following formula 6.
  • a constraint condition is introduced as information related to both speech and noise.
  • the constraint condition may be used as information related to both speech and noise when updating parameters, which will be described later.
  • the parameters of the speech emphasis mask may be updated to make the speech emphasis stronger, and the parameters of the noise emphasis mask may also be updated to make the noise emphasis stronger.
  • the parameters of the speech emphasis mask may be updated to make the speech emphasis weaker, and the parameters of the noise emphasis mask may also be updated to make the noise emphasis weaker.
  • a loss function can be derived based on the amount by which the estimated speech enhancement mask and noise enhancement mask deviate from the constraint condition.
  • a constraint loss function L SN shown in the following formula 7 may be introduced as the constraint condition.
  • the following formula 7 calculates the sum of time-frequency.
  • the combined loss calculation unit 217 calculates a combined loss by combining the speech enhancement mask loss, the noise enhancement mask loss, and the constraint loss (step S26).
  • the combined loss may be calculated by a combined loss function L ALL .
  • the combined loss function L ALL may be expressed as in the following formula 8.
  • the combined loss calculated by the combined loss function L ALL may be called a total loss. [Formula 8]
  • lambda may be a value between “0" and “1". lambda may be a positive number less than 1. lambda may be a value small compared to "1", such as "0.01" or "0.1". lambda may be constant or may vary during the model training process. lambda may be particularly small at the start of the model training process and large at the end.
  • the speech enhancement mask loss function L S and the noise enhancement mask loss function L N are weighted equally in the summation, but different weights may be applied to the speech enhancement mask loss function L S and the noise enhancement mask loss function L N.
  • the total loss may be calculated by the summation loss function L ALL shown in the following formula 9. [Formula 9]
  • the parameter update unit 212 updates the parameters included in the speech enhancement mask estimation model and the parameters included in the noise enhancement mask estimation model according to the calculation result of the combined loss calculation unit 217 (step S27).
  • the parameter update unit 212 updates the parameters included in the speech enhancement mask estimation model and the parameters included in the noise enhancement mask estimation model using the speech enhancement mask loss indicating the difference between the estimated speech enhancement mask and the ideal speech enhancement mask, the noise enhancement mask loss indicating the difference between the estimated noise enhancement mask and the ideal noise enhancement mask, and the constraint loss.
  • the parameter update unit 212 updates the parameters included in the speech enhancement mask estimation model and the parameters included in the noise enhancement mask estimation model stored in the parameter storage unit 221.
  • the parameter update unit 212 may update the parameters of the speech enhancement mask estimation model and the noise enhancement mask estimation model using, for example, a back-error propagation method.
  • the information processing device 2 in this embodiment can proceed with learning of both the speech enhancement mask estimation model and the noise enhancement mask estimation model through the constrained loss function L SN.
  • the speech enhancement mask estimation model can be learned using information obtained by combining information included in the speech enhancement mask estimation model and information included in the noise enhancement mask estimation model. [2-4: Tolerance of constraint conditions]
  • the sum of the speech enhancement mask and the noise enhancement mask is ideally "1", and this embodiment uses this as a constraint.
  • this constraint does not necessarily have to be “1” strictly.
  • "1” may be replaced with “1+ ⁇ ” as shown in the following formula 10.
  • " ⁇ " is a positive number that is very small compared to "1". When “1” is replaced with “1+ ⁇ ”, an error of about “ ⁇ ” can be tolerated. " ⁇ ” may be arbitrarily changed by design.
  • S i and N i in the denominator of the above formula 10 represent ideal values, and S and N in the numerator represent estimated values.
  • the parameter update unit 212 updates the parameters so that at least the sum of the speech emphasis mask and the noise emphasis mask is not less than "1".
  • the audio enhancement mask estimation unit 213 may estimate audio enhancement masks corresponding to each of the multiple people.
  • a audio enhancement mask estimation model corresponding to each person may be prepared, and each audio enhancement mask estimation model may output an audio enhancement mask corresponding to the person.
  • the constraint condition for outputting the speech enhancement mask M SA and the speech enhancement mask M SB may be expressed as in the following formula 12. [Formula 12]
  • the input signal contains speech and noise (as described above, in this embodiment, anything other than speech is called “noise”). Therefore, the sum of the ideal value of the speech enhancement mask and the ideal value of the noise enhancement mask is a constant value of "1". As described above, when the sum of the speech enhancement mask and the noise enhancement mask exceeds "1", at least one of the estimated value of the speech enhancement mask and the estimated value of the noise enhancement mask is larger than the ideal value, and there is a possibility that noise cannot be completely suppressed from the speech.
  • the sum of the speech enhancement mask and the noise enhancement mask is less than "1"
  • at least one of the estimated value of the speech enhancement mask and the estimated value of the noise enhancement mask is smaller than the ideal value, and there is a possibility that too much speech, which is a sound that should be left, is cut off.
  • too much speech which is a sound that should be left
  • the input signal contains noise that changes suddenly in a short time
  • too much speech is often cut off.
  • the estimated speech enhancement mask and the estimated noise enhancement mask are evaluated independently, and it is unclear whether the sum of the speech enhancement mask and the noise enhancement mask is a constant value of "1". Therefore, not only is it not possible to completely suppress noise from the speech, but there is a possibility that too much of the speech that should be preserved will be cut out.
  • the accuracy of the noise enhancement mask is often lower than that of the speech enhancement mask, so if an independent evaluation of the noise enhancement mask is used to update the model parameters, the accuracy of the speech enhancement may deteriorate.
  • the information processing device 2 in the second embodiment introduces a constraint loss calculated using the estimation result by the speech enhancement mask estimation unit 213 and the estimation result by the noise enhancement mask estimation unit 214. Then, the parameters included in the speech enhancement mask estimation model are updated using a combined loss obtained by combining the speech enhancement mask loss, the noise enhancement mask loss, and the constraint loss, so that it is possible to suppress noise from speech and prevent excessive cutting of speech, which is a sound that should be retained.
  • the information processing device 2 can create a speech enhancement mask estimation model that can perform speech enhancement with high accuracy.
  • FIG. 4 is a block diagram showing a configuration of an information processing device 3 in the third embodiment.
  • the information processing device 3 in the third embodiment differs from the information processing device 2 in the second embodiment in the operation of a constraint loss calculation unit 311.
  • the constraint loss calculation unit 311 calculates the constraint loss by multiplying the estimated speech emphasis mask and the estimated noise emphasis mask by the magnitude of the noise mixed speech.
  • the constraint loss calculation unit 311 may calculate the constraint loss by multiplying the estimated speech emphasis mask and the estimated noise emphasis mask in time-frequency by the magnitude of the noise mixed speech.
  • the magnitude of the noise mixed speech may be the absolute value of the amplitude of the noise mixed speech.
  • the magnitude of the noise mixed speech may be the logarithmic power spectrum of the noise mixed speech.
  • the magnitude of the noise mixed speech may be a normalized value. The value may be arbitrarily changed by design.
  • Equation 13 the constraint loss function L SN in the third embodiment may be expressed as in Equation 13 below.
  • Multiplying by the LPS input allows more emphasis to be placed on errors in time-frequency regions where the LPS input is large, which is equivalent to applying magnitude spectrum approximation (MSA).
  • MSA magnitude spectrum approximation
  • Multiplying by the LPS input is effective when learning noise that changes suddenly in a short time, such as the sound of a passing bullet train. It gives more weight to large input signals than to small input signals, making it possible to emphasize errors related to large input signals.
  • the information processing device 3 in the third embodiment multiplies the magnitude of the noise-mixed voice, and therefore can suppress noise related to a particularly large input signal.
  • FIG. 5 is a block diagram showing a configuration of the information processing device 4 according to the fourth embodiment.
  • the information processing device 4 according to the fourth embodiment differs from the information processing device 2 according to the second embodiment and the information processing device 3 according to the third embodiment in the operation of a constraint loss calculation unit 411.
  • the constraint loss calculation unit 411 calculates the constraint loss by raising the estimated speech enhancement mask and the estimated noise enhancement mask to a predetermined exponent.
  • the constraint loss calculation unit 411 may also calculate the constraint loss by raising the estimated speech enhancement mask and the estimated noise enhancement mask in time-frequency to a predetermined exponent.
  • the predetermined exponent may take a value between 0 and 1.
  • the constraint loss function L SN in the fourth embodiment may be expressed as the following formula 14. [Formula 14]
  • the LPS input applied in the third embodiment may be applied in the fourth embodiment as well.
  • the constraint loss function L SN in the fourth embodiment may be expressed as in the following formula 15. [Formula 15] [4-2: Technical Effects of Information Processing Device 4]
  • the information processing device 4 in the fourth embodiment can apply ⁇ to learn small signals with greater emphasis.
  • the information processing device 5 in the fifth embodiment includes a calculation device 21 and a storage device 22, similar to the information processing device 2 in the second embodiment to the information processing device 4 in the fourth embodiment. Furthermore, the information processing device 5 in the fifth embodiment may include a communication device 23, an input device 24, and an output device 25, similar to the information processing device 2 in the second embodiment to the information processing device 4 in the fourth embodiment. However, the information processing device 5 may not include at least one of the communication device 23, the input device 24, and the output device 25. The information processing device 5 in the fifth embodiment differs from the information processing device 2 in the second embodiment to the information processing device 4 in the fourth embodiment in that an acoustic feature extraction unit 519 is further realized in the calculation device 21.
  • FIG. 7 is a flowchart showing the flow of the information processing operation performed by the information processing device 5.
  • the noise mixed speech input unit 518 acquires the noise mixed speech and inputs it to the acoustic feature extraction unit 519 (step S50).
  • the acoustic feature extraction unit 519 extracts noise mixed speech features, which are characteristics of the noise mixed speech, from the noise mixed speech (step S51).
  • the acoustic feature extraction unit 519 may have an acoustic feature extraction model that receives the noise mixed speech as input and outputs the noise mixed speech features.
  • the acoustic feature extraction model may be implemented using an RNN.
  • the speech enhancement mask estimation unit 513 estimates the speech enhancement mask using the speech enhancement mask estimation model (step S52).
  • the speech enhancement mask estimation model may receive noise-mixed speech features and output an estimated speech enhancement mask.
  • the speech enhancement mask estimation unit 513 may output the estimated speech enhancement mask output by the speech enhancement mask estimation model to which the noise-mixed speech features are input as an estimated value.
  • the noise enhancement mask estimation unit 514 estimates the noise enhancement mask using the noise enhancement mask estimation model (step S53).
  • the noise enhancement mask estimation model may receive the noise-mixed speech features and output an estimated noise enhancement mask.
  • the noise enhancement mask estimation unit 514 may output the estimated noise enhancement mask output by the noise enhancement mask estimation model to which the noise-mixed speech features are input as an estimated value.
  • the speech enhancement mask loss calculation unit 515 calculates the speech enhancement mask loss using the estimated speech enhancement mask estimated by the speech enhancement mask estimation model and the ideal speech enhancement mask (step S54).
  • the ideal speech enhancement mask may be a speech enhancement mask that is ideal for the noise-mixed speech features extracted from the teacher input signal.
  • the noise enhancement mask loss calculation unit 516 calculates the noise enhancement mask loss using the estimated noise enhancement mask estimated by the noise enhancement mask estimation model and the ideal noise enhancement mask (step S55).
  • the ideal noise enhancement mask may be an ideal noise enhancement mask for the noise-mixed speech features extracted from the teacher input signal.
  • the constraint loss calculation unit 211 calculates the constraint loss using the estimated speech enhancement mask estimated by the speech enhancement mask estimation model and the estimated noise enhancement mask estimated by the noise enhancement mask estimation model (step S25).
  • the combined loss calculation unit 217 calculates the combined loss by combining the speech enhancement mask loss, the noise enhancement mask loss, and the constraint loss (step S26).
  • the parameter update unit 512 updates the parameters included in the acoustic feature extraction model, as well as the parameters included in the speech enhancement mask estimation model and the parameters included in the noise enhancement mask estimation model, according to the calculation result of the combined loss calculation unit 217 (step S56).
  • the acoustic feature extraction model may be trained using both information about speech and information about noise.
  • the information processing device 5 in the fifth embodiment creates a speech enhancement mask estimation model and an acoustic feature extraction model by machine learning.
  • the acoustic feature extraction model is applied.
  • the information processing device 5 in the fifth embodiment is provided with an acoustic feature extraction unit 519, which allows the input signals to both the speech enhancement mask estimation model and the noise enhancement mask estimation model to be increased, thereby enabling more preferable speech enhancement.
  • the acoustic feature extraction unit 519 is not provided, as in the second embodiment, the operation can be made lighter.
  • the information processing device 6 in the sixth embodiment includes a calculation device 21 and a storage device 22, similar to the information processing device 2 in the second embodiment to the information processing device 5 in the fifth embodiment. Furthermore, the information processing device 6 in the sixth embodiment may include a communication device 23, an input device 24, and an output device 25, similar to the information processing device 2 in the second embodiment to the information processing device 5 in the fifth embodiment. However, the information processing device 6 may not include at least one of the communication device 23, the input device 24, and the output device 25.
  • the information processing device 6 in the sixth embodiment differs from the information processing device 2 in the second embodiment to the information processing device 5 in the fifth embodiment in that the constraint loss calculation unit 611 has a noise constraint loss calculation unit 6111 and a speech constraint loss calculation unit 6112.
  • FIG. 9 is a flowchart showing the flow of the information processing operation performed by the information processing device 6.
  • the noise-mixed speech input unit 218 acquires the noise-mixed speech and inputs it to the speech enhancement mask estimation unit 213 and the noise enhancement mask estimation unit 214 (step S20).
  • the speech enhancement mask estimation unit 213 estimates a speech enhancement mask using a speech enhancement mask estimation model (step S21).
  • the noise enhancement mask estimation unit 214 estimates a noise enhancement mask using the noise enhancement mask estimation model (step S22).
  • the speech enhancement mask loss calculation unit 215 calculates a speech enhancement mask loss using the estimated speech enhancement mask estimated by the speech enhancement mask estimation model and the ideal speech enhancement mask (step S23).
  • the noise enhancement mask loss calculation unit 216 calculates a noise enhancement mask loss using the estimated noise enhancement mask estimated by the noise enhancement mask estimation model and the ideal noise enhancement mask (step S24). [6-2-1: Noise-constrained loss]
  • the noise constrained loss calculation unit 6111 calculates the noise constrained loss using the estimated speech emphasis mask and the ideal noise emphasis mask (step S60).
  • the noise constrained loss calculation unit 6111 calculates the constrained noise emphasis mask based on the estimated speech emphasis mask and the constraint condition expressed by the above formula 5.
  • the constrained noise emphasis mask is expressed as M N *
  • the constrained noise emphasis mask may be expressed as the following formula 16. [Formula 16]
  • the noise-constrained loss calculation unit 6111 calculates a noise-constrained loss indicating a difference between the constrained noise enhancement mask and the ideal noise enhancement mask.
  • the noise-constrained loss is calculated by a noise-constrained loss function L N *
  • the noise-constrained loss function L N * may be expressed as in the following Equation 17. [Formula 17] [6-2-2: Audio Restriction Loss]
  • the speech constraint loss calculation unit 6112 calculates the speech constraint loss using the estimated noise emphasis mask and the ideal noise emphasis mask (step S61).
  • the speech constraint loss calculation unit 6112 calculates the constraint speech emphasis mask based on the estimated noise emphasis mask and the constraint condition expressed by the above formula 5.
  • the constraint speech emphasis mask is expressed as M S *
  • the constraint speech emphasis mask may be expressed as the following formula 18. [Formula 18]
  • the speech constraint loss calculation unit 6112 calculates a speech constraint loss indicating a difference between the constrained speech enhancement mask and the ideal speech enhancement mask.
  • the speech constraint loss is calculated by a speech constraint loss function L S *
  • the speech constraint loss function L S * may be expressed as the following formula 19. [Formula 19] [6-2-3: Combined Losses]
  • the combined loss calculation unit 617 calculates a combined loss by combining the speech enhancement mask loss, the noise enhancement mask loss, the noise constrained loss, and the speech constrained loss (step S62).
  • the combined loss is calculated by a combined loss function L ALL
  • the combined loss function L ALL may be expressed as the following formula 20.
  • M can be considered as an estimate of a mask that also includes a constraint.
  • the information processing device 6 in the sixth embodiment can obtain an effect equivalent to that of the information processing device 2 in the second embodiment.
  • the above formula 20 may be rewritten as the following formula 21.
  • the combined loss in the sixth embodiment can be considered to indicate the difference between the estimated value of the mask including the constraint conditions and the ideal value.
  • the parameter update unit 212 updates the parameters included in the speech enhancement mask estimation model and the parameters included in the noise enhancement mask estimation model according to the calculation result of the combined loss calculation unit 217 (step S27).
  • the information processing device 6 in the sixth embodiment does not employ " ⁇ " as employed in the second embodiment. Therefore, the information processing device 6 in the sixth embodiment can reduce the processing load required for determining " ⁇ " more than the information processing device 2 in the second embodiment.
  • [Appendix 1] a constraint loss calculation means for calculating a constraint loss using an estimated speech enhancement mask output from a speech enhancement mask estimation model to which a noise-mixed speech is input, and an estimated noise enhancement mask output from a noise enhancement mask estimation model to which the noise-mixed speech is input; and a parameter updating means for updating parameters included in the speech enhancement mask estimation model and parameters included in the noise enhancement mask estimation model using a speech enhancement mask loss indicating a difference between the estimated speech enhancement mask and a target speech enhancement mask, a noise enhancement mask loss indicating a difference between the estimated noise enhancement mask and a target noise enhancement mask, and the constraint loss.
  • the parameter update means updates the parameters included in the speech enhancement mask estimation model and the parameters included in the noise enhancement mask estimation model in accordance with a calculation result of the combined loss calculation means.
  • the method further comprises: extracting an acoustic feature from the noise-mixed speech; the speech enhancement mask estimation model receives the noise-mixed speech features and outputs the estimated speech enhancement mask; The information processing apparatus according to claim 1 , wherein the noise enhancement mask estimation model receives the noise-mixed speech feature and outputs the estimated noise enhancement mask.
  • the constraint loss calculation means a speech constraint loss calculation means for calculating a speech constraint loss using the estimated noise emphasis mask and the target speech emphasis mask; a noise constrained loss calculation means for calculating a noise constrained loss using the estimated speech emphasis mask and the target noise emphasis mask; The information processing device according to claim 1 , wherein the constraint loss includes the speech constraint loss and the noise constraint loss.
  • [Appendix 7] Calculating a constraint loss using an estimated speech enhancement mask output from a speech enhancement mask estimation model to which the noise-mixed speech is input, and an estimated noise enhancement mask output from a noise enhancement mask estimation model to which the noise-mixed speech is input; an information processing method for updating parameters included in the speech enhancement mask estimation model and parameters included in the noise enhancement mask estimation model using a speech enhancement mask loss indicating a difference between the estimated speech enhancement mask and a target speech enhancement mask, a noise enhancement mask loss indicating a difference between the estimated noise enhancement mask and a target noise enhancement mask, and the constraint loss.
  • At least some of the components of each of the above-described embodiments can be appropriately combined with at least some of the other components of each of the above-described embodiments. Some of the components of each of the above-described embodiments may not be used.

Landscapes

  • Engineering & Computer Science (AREA)
  • Computational Linguistics (AREA)
  • Quality & Reliability (AREA)
  • Signal Processing (AREA)
  • Health & Medical Sciences (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Human Computer Interaction (AREA)
  • Physics & Mathematics (AREA)
  • Acoustics & Sound (AREA)
  • Multimedia (AREA)
  • Circuit For Audible Band Transducer (AREA)

Abstract

情報処理装置は、雑音混合音声が入力された音声強調マスク推定モデルが出力した推定音声強調マスク、及び雑音混合音声が入力された雑音強調マスク推定モデルが出力した推定雑音強調マスクを用いて拘束ロスを算出する拘束ロス算出部と、推定音声強調マスクと目標音声強調マスクとの違いを示す音声強調マスクロス、推定雑音強調マスクと目標雑音強調マスクとの違いを示す雑音強調マスクロス、及び拘束ロスを用いて、音声強調マスク推定モデルが含むパラメータ、及び雑音強調マスク推定モデルが含むパラメータを更新するパラメータ更新部とを備え、精度よく音声強調をすることができる音声強調マスク推定モデルを生成する。

Description

情報処理装置、情報処理方法、及び、記録媒体
 この開示は、情報処理装置、情報処理方法、及び、記録媒体の技術分野に関する。
 非特許文献1には、モデルにより音声強調マスクと雑音強調マスクとを推定し、推定した各マスクを用いて指標を計算し、計算した指標と指標の理想値とのずれを用いてモデルを学習させる技術が記載されている。
Investigations on Data Augmentation and Loss Functions for Deep Learning Based Speech-Background Separation; Interspeech 2018; P.3499-3503
 この開示は、先行技術文献に記載された技術の改良を目的とする情報処理装置、情報処理方法、及び、記録媒体を提供することを課題とする。
 情報処理装置の一の態様は、雑音混合音声が入力された音声強調マスク推定モデルが出力した推定音声強調マスク、及び前記雑音混合音声が入力された雑音強調マスク推定モデルが出力した推定雑音強調マスクを用いて拘束ロスを算出する拘束ロス算出手段と、前記推定音声強調マスクと雑音混合音声中の音声と雑音が既知な場合に計算される目標音声強調マスクとの違いを示す音声強調マスクロス、前記推定雑音強調マスクと目標雑音強調マスクとの違いを示す雑音強調マスクロス、及び前記拘束ロスを用いて、前記音声強調マスク推定モデルが含むパラメータ、及び前記雑音強調マスク推定モデルが含むパラメータを更新するパラメータ更新手段とを備える。
 情報処理方法の一の態様は、雑音混合音声が入力された音声強調マスク推定モデルが出力した推定音声強調マスク、及び前記雑音混合音声が入力された雑音強調マスク推定モデルが出力した推定雑音強調マスクを用いて拘束ロスを算出し、前記推定音声強調マスクと雑音混合音声中の音声と雑音が既知な場合に計算される目標音声強調マスクとの違いを示す音声強調マスクロス、前記推定雑音強調マスクと目標雑音強調マスクとの違いを示す雑音強調マスクロス、及び前記拘束ロスを用いて、前記音声強調マスク推定モデルが含むパラメータ、及び前記雑音強調マスク推定モデルが含むパラメータを更新する。
 記録媒体の一の態様は、コンピュータに、雑音混合音声が入力された音声強調マスク推定モデルが出力した推定音声強調マスク、及び前記雑音混合音声が入力された雑音強調マスク推定モデルが出力した推定雑音強調マスクを用いて拘束ロスを算出し、前記推定音声強調マスクと雑音混合音声中の音声と雑音が既知な場合に計算される目標音声強調マスクとの違いを示す音声強調マスクロス、前記推定雑音強調マスクと目標雑音強調マスクとの違いを示す雑音強調マスクロス、及び前記拘束ロスを用いて、前記音声強調マスク推定モデルが含むパラメータ、及び前記雑音強調マスク推定モデルが含むパラメータを更新する情報処理方法を実行させるためのコンピュータプログラムが記録されている。
図1は、第1実施形態における情報処理装置の構成を示すブロック図である。 図2は、第2実施形態における情報処理装置の構成を示すブロック図である。 図3は、第2実施形態における情報処理装置の情報処理動作の流れを示すフローチャートである。 図4は、第3実施形態における情報処理装置の構成を示すブロック図である。 図5は、第4実施形態における情報処理装置の構成を示すブロック図である。 図6は、第5実施形態における情報処理装置の構成を示すブロック図である。 図7は、第5実施形態における情報処理装置の情報処理動作の流れを示すフローチャートである。 図8は、第6実施形態における情報処理装置の構成を示すブロック図である。 図9は、第6実施形態における情報処理装置の情報処理動作の流れを示すフローチャートである。
 以下、図面を参照しながら、情報処理装置、情報処理方法、及び、記録媒体の実施形態について説明する。
 [1:第1実施形態]
 情報処理装置、情報処理方法、及び、記録媒体の第1実施形態について説明する。以下では、情報処理装置、情報処理方法、及び記録媒体の第1実施形態が適用された情報処理装置1を用いて、情報処理装置、情報処理方法、及び記録媒体の第1実施形態について説明する。
 [1-1:情報処理装置1の構成]
 図1は、第1実施形態における情報処理装置1の構成を示すブロック図である。図1に示すように、情報処理装置1は、拘束ロス算出部11と、パラメータ更新部12とを備える。
 拘束ロス算出部11は、雑音混合音声が入力された音声強調マスク推定モデルが出力した推定音声強調マスク、及び雑音混合音声が入力された雑音強調マスク推定モデルが出力した推定雑音強調マスクを用いて拘束ロスを算出する。パラメータ更新部12は、推定音声強調マスクと目標音声強調マスクとの違いを示す音声強調マスクロス、推定雑音強調マスクと目標雑音強調マスクとの違いを示す雑音強調マスクロス、及び拘束ロスを用いて、音声強調マスク推定モデルが含むパラメータ、及び雑音強調マスク推定モデルが含むパラメータを更新する。なお、「雑音混合音声」、「音声強調マスク推定モデル」、「推定音声強調マスク」、「雑音強調マスク推定モデル」、「推定雑音強調マスク」、「拘束ロス」、「目標音声強調マスク」、「音声強調マスクロス」、「目標雑音強調マスク」、及び「雑音強調マスクロス」という語句については、後述する他の実施形態において詳しく説明する。
 [1-2:情報処理装置1の技術的効果]
 第1実施形態における情報処理装置1は、入力音声に含まれる目的音声以外の雑音の時刻、周波数帯域の音量を相対的に減少する音声強調マスク推定モデルが出力した推定音声強調マスク、及び入力音声に含まれる目的音声の時刻、周波数帯域の音量を相対的に減少する雑音強調マスク推定モデルが出力した推定雑音強調マスクを用いて算出した拘束ロスを導入する。情報処理装置1は、音声強調マスクの推定値と目標値との違い、雑音強調マスクの推定値と目標値との違い、及び拘束ロスを用いて、音声強調マスク推定モデルが含むパラメータを更新するので、精度よく音声強調をすることができる音声強調マスク推定モデルを生成することができる。
 [2:第2実施形態]
 続いて、情報処理装置、情報処理方法、及び記録媒体の第2実施形態について説明する。以下では、情報処理装置、情報処理方法、及び記録媒体の第2実施形態が適用された情報処理装置2を用いて、情報処理装置、情報処理方法、及び記録媒体の第2実施形態について説明する。
 [2-1-1:音声と雑音]
 本実施形態において、「音声」とは、対象とする音信号を表す。本実施形態において、「音声」を「対象信号」と称する場合がある。また、本実施形態において、「雑音」とは、対象としない音信号を表す。本実施形態において、「雑音」を「対象外信号」と称する場合がある。
 音声は、注目したい音であってもよい。音声は、人物の声であってもよい。音声は、特定の人物の声であってもよい。特定の人物は、1以上の人物であってもよい。特定の人物は、マイク等の音を取得する機構の近くに居る人物であってもよい。「音声」は、場面に応じて異なる音信号を表していてもよい。
 本実施形態において、「雑音混合音声」とは、「雑音」が混合されている「音声」である。「雑音混合音声」は、対象とする音信号である「音声」と、対象としない音信号である「雑音」とを含む。「雑音混合音声」に含まれる音信号は、「音声」、及び「雑音」の何れかである。言い換えると、「雑音混合音声」に含まれる「音声」ではない音信号は「雑音」である。また、「雑音混合音声」に含まれる「雑音」ではない音信号は「音声」である。また、本実施形態において、情報処理装置2に入力される入力信号は、「雑音混合音声」である。すなわち、情報処理装置2に入力される入力信号は、対象信号である「音声」と、対象外信号である「雑音」からなる。
 [2-1-2:音声強調技術]
 音声に混入している音声以外の雑音をマイクロフォンが拾う場合がある。この場合、音声に対して雑音が相対的に大きければ大きいほど、音声は聞き取りにくくなる。音声強調技術は、雑音混合音声から音声を雑音に対して相対的に大きくする技術である。本技術は、雑音が混合されている音声から雑音を抑制して音声を強調し、より聞き取り易い音声を提供する技術であってもよい。雑音が混合されている音声から、音声のみを強調し、より聞き取り易い音声を提供する技術であってもよい。音声強調技術は、高雑音の状況下において、通話音声の雑音を除去することができてもよい。また、音声強調技術は、通話相手の音声を強調して、より円滑なコミュニケーションを可能にしてもよい。また、音声強調技術は、音声認識器の認識率を向上させてもよい。
 また、マスク型音声強調技術は、雑音の音量を下げる時刻、周波数帯、音量の減少量を指定する音声強調マスクを推定する。マスク型音声強調技術は、雑音が混合されている雑音混合音声に当該音声強調マスクを施すことで、音声を聞き取り易くする技術であってもよい。本実施形態では、マスク型音声強調技術においては、音声強調マスクを推定する音声強調マスク推定モデルを機械学習してもよい。本実施形態において作成された音声強調マスク推定モデルは、音声を強調する場面に適用されてもよい。
 本実施形態において作成される音声強調マスク推定モデルは、クリーンな音声に雑音を重畳した雑音混合音声音を入力信号として用いて学習をされる。クリーンな音声とは、雑音が極めて少ない音であってもよい。本実施形態における機械学習は、深層学習であってもよい。
 [2-1-3:音声強調マスク、及び雑音強調マスク]
 音声強調マスク推定モデルは、雑音混合音声が入力されると音声強調マスクを出力するモデルである。音声強調マスク推定モデルが推定する音声強調マスクを推定音声強調マスクと呼んでもよい。推定音声強調マスクを推定値と呼ぶ場合がある。音声強調マスクは、下記式1のように表現してもよい。
 [式1]
Figure JPOXMLDOC01-appb-I000001
 S(t,f)は、「音声」のパワーを示す。また、N(t,f)は、「雑音」のパワーを示す。
 音声強調マスクは、「雑音混合音声」のパワーのうちの「音声」のパワーの割合で表現する。
 上述の通り、本実施形態は、音声を強調することを目的として音声強調マスク推定モデルを学習をさせて作成するが、音声強調と伴に雑音強調も学習させることで、音声強調マスク推定モデルの学習を促進できることが期待できる。
 雑音強調マスクは、入力音声に含まれる目的音声の時刻、周波数帯域の音量を相対的に減少させる。雑音強調マスク推定モデルは、雑音混合音声が入力されると雑音強調マスクを出力するモデルである。雑音強調マスク推定モデルが推定する雑音強調マスクを推定雑音強調マスクと呼んでもよい。推定雑音強調マスクを推定値と呼ぶ場合がある。雑音強調マスクは、下記式2のように表現してもよい。
 [式2]
Figure JPOXMLDOC01-appb-I000002
 雑音強調マスクは、「雑音混合音声」のパワーのうちの「雑音」のパワーの割合を表す。
 [2-1-4:学習]
 情報処理装置2は、音声強調マスク推定モデルと雑音強調マスク推定モデルとを学習させることにより、理想的な音声強調マスクを推定することのできる音声強調マスク推定モデルを作成してもよい。音声強調マスク推定モデル、及び雑音強調マスク推定モデルは、ニューラル・ネットワーク(Neural Network:NN)で実装してもよい。音声強調マスク推定モデル、及び雑音強調マスク推定モデルは、再帰的ニューラル・ネットワーク(Recurrent Neural Network:RNN)で実装してもよい。情報処理装置2は、音声強調マスク推定モデルによる推定値と雑音強調マスク推定モデルによる推定値とに応じて、音声強調マスク推定モデルが含むパラメータと雑音強調マスク推定モデルが含むパラメータとを更新し、理想的な音声強調マスクを推定することのできる音声強調マスク推定モデルを作成してもよい。
 教師入力信号、並びに教師入力信号に対応する理想的な音声強調マスク、及び雑音強調マスクを予め用意する。理想的な音声強調マスク(「目標音声強調マスク」と表すこともある)とは、雑音混合音声中の音声と雑音が既知な場合に計算される音声強調マスクである。また、理想的な雑音強調マスク(「目標雑音強調マスク」と表すこともある)とは、雑音混合音声中の音声と雑音が既知な場合に計算される雑音強調マスクである。
 以降、「理想的な音声強調マスク」および「理想的な雑音強調マスク」という言葉を用いて、情報処理装置2における処理を説明する。
教師入力信号とは、理想的な音声強調マスク、及び雑音強調マスクが予め用意されている入力信号である。入力信号とは、上述の通り、対象信号である「音声」と、対象外信号である「雑音」からなっていてもよい。理想的な音声強調マスクを理想音声強調マスク、又は理想値とよんでもよい。また、理想的な雑音強調マスクを理想雑音強調マスク、又は理想値とよんでもよい。また、教師入力信号、及び理想値を含む情報を、教師情報と呼んでもよい。教師情報は、後に説明する教師情報記憶部222に記憶されていてもよい。理想値は、パラメータの更新において、目標とする値(目標値)であってもよい。
 [2-2:情報処理装置2の構成]
 図2は、第2実施形態における情報処理装置2の構成を示すブロック図である。図2に示すように、情報処理装置2は、演算装置21と、記憶装置22とを備えている。更に、情報処理装置2は、通信装置23と、入力装置24と、出力装置25とを備えていてもよい。但し、情報処理装置2は、通信装置23、入力装置24及び出力装置25のうちの少なくとも一つを備えていなくてもよい。演算装置21と、記憶装置22と、通信装置23と、入力装置24と、出力装置25とは、データバス26を介して接続されていてもよい。
 演算装置21は、例えば、CPU(Central Processing Unit)、GPU(Graphics Proecssing Unit)及びFPGA(Field Programmable Gate Array)のうちの少なくとも一つを含む。演算装置21は、コンピュータプログラムを読み込む。例えば、演算装置21は、記憶装置22が記憶しているコンピュータプログラムを読み込んでもよい。例えば、演算装置21は、コンピュータで読み取り可能であって且つ一時的でない記録媒体が記憶しているコンピュータプログラムを、情報処理装置2が備える図示しない記録媒体読み取り装置(例えば、後述する入力装置24)を用いて読み込んでもよい。演算装置21は、通信装置23(或いは、その他の通信装置)を介して、情報処理装置2の外部に配置される不図示の装置からコンピュータプログラムを取得してもよい(つまり、ダウンロードしてもよい又は読み込んでもよい)。演算装置21は、読み込んだコンピュータプログラムを実行する。その結果、演算装置21内には、情報処理装置2が行うべき動作を実行するための論理的な機能ブロックが実現される。つまり、演算装置21は、情報処理装置2が行うべき動作(言い換えれば、処理)を実行するための論理的な機能ブロックを実現するためのコントローラとして機能可能である。
 図2には、情報処理動作を実行するために演算装置21内に実現される論理的な機能ブロックの一例が示されている。図2に示すように、演算装置21内には、後述する付記に記載された「拘束ロス算出手段」の一具体例である拘束ロス算出部211と、後述する付記に記載された「パラメータ更新手段」の一具体例であるパラメータ更新部212と、後述する付記に記載された「音声強調マスク推定手段」の一具体例である音声強調マスク推定部213と、後述する付記に記載された「雑音強調マスク推定手段」の一具体例である雑音強調マスク推定部214と、後述する付記に記載された「音声強調マスクロス算出手段」の一具体例である音声強調マスクロス算出部215と、後述する付記に記載された「雑音強調マスクロス算出手段」の一具体例である雑音強調マスクロス算出部216と、後述する付記に記載された「合算ロス算出手段」の一具体例である合算ロス算出部217と、雑音混合音声入力部218とが実現される。音声強調マスク推定部213は、後述する付記に記載された「音声強調マスク推定モデル」の一具体例である音声強調マスク推定モデルを用いて音声強調マスクを推定する。雑音強調マスク推定部214は、後述する付記に記載された「雑音強調マスク推定モデル」の一具体例である雑音強調マスク推定モデルを用いて雑音強調マスクを推定する。但し、演算装置21内には、音声強調マスク推定部213、雑音強調マスク推定部214、音声強調マスクロス算出部215、雑音強調マスクロス算出部216、合算ロス算出部217、及び雑音混合音声入力部218の何れかが実現されなくてもよい。拘束ロス算出部211、パラメータ更新部212、音声強調マスク推定部213、雑音強調マスク推定部214、音声強調マスクロス算出部215、雑音強調マスクロス算出部216、合算ロス算出部217、及び雑音混合音声入力部218の各々の動作の詳細については、図3を参照しながら後に説明する。
 記憶装置22は、所望のデータを記憶可能である。例えば、記憶装置22は、演算装置21が実行するコンピュータプログラムを一時的に記憶していてもよい。記憶装置22は、演算装置21がコンピュータプログラムを実行している場合に演算装置21が一時的に使用するデータを一時的に記憶してもよい。記憶装置22は、情報処理装置2が長期的に保存するデータを記憶してもよい。尚、記憶装置22は、RAM(Random Access Memory)、ROM(Read Only Memory)、ハードディスク装置、光磁気ディスク装置、SSD(Solid State Drive)及びディスクアレイ装置のうちの少なくとも一つを含んでいてもよい。つまり、記憶装置22は、一時的でない記録媒体を含んでいてもよい。記憶装置22は、パラメータ記憶部221、及び教師情報記憶部222を実現してもよい。パラメータ記憶部221は、音声強調マスク推定モデルが含むパラメータ、及び雑音強調マスク推定モデルが含むパラメータを記憶する。但し、記憶装置22は、パラメータ記憶部221、及び教師情報記憶部222の何れかを実現しなくてもよい。
 通信装置23は、不図示の通信ネットワークを介して、情報処理装置2の外部の装置と通信可能である。通信装置23は、イーサネット(登録商標)、Wi-Fi(登録商標)、Bluetooth(登録商標)、USB(Universal Serial Bus)等の規格に基づく通信インターフェースであってもよい。
 入力装置24は、情報処理装置2の外部からの情報処理装置2に対する情報の入力を受け付ける装置である。例えば、入力装置24は、情報処理装置2のオペレータが操作可能な操作装置(例えば、キーボード、マウス及びタッチパネルのうちの少なくとも一つ)を含んでいてもよい。例えば、入力装置24は情報処理装置2に対して外付け可能な記録媒体にデータとして記録されている情報を読み取り可能な読取装置を含んでいてもよい。
 出力装置25は、情報処理装置2の外部に対して情報を出力する装置である。例えば、出力装置25は、情報を画像として出力してもよい。つまり、出力装置25は、出力したい情報を示す画像を表示可能な表示装置(いわゆる、ディスプレイ)を含んでいてもよい。例えば、出力装置25は、情報を音声として出力してもよい。つまり、出力装置25は、音声を出力可能な音声装置(いわゆる、スピーカ)を含んでいてもよい。例えば、出力装置25は、紙面に情報を出力してもよい。つまり、出力装置25は、紙面に所望の情報を印刷可能な印刷装置(いわゆる、プリンタ)を含んでいてもよい。
 [2-3:情報処理装置2が行う情報処理動作]
 図3を参照しながら、情報処理装置2が行う情報処理動作について説明する。図3は、情報処理装置2が行う情報処理動作の流れを示すフローチャートである。
 図3に示す様に、雑音混合音声入力部218は、雑音混合音声を取得し、音声強調マスク推定部213、及び雑音強調マスク推定部214に入力する(ステップS20)。
 音声強調マスク推定部213は、音声強調マスク推定モデルを用いて音声強調マスクを推定する(ステップS21)。音声強調マスク推定部213は、雑音混合音声が入力された音声強調マスク推定モデルが出力した推定音声強調マスクを推定値として出力する。
 雑音強調マスク推定部214は、雑音強調マスク推定モデルを用いて雑音強調マスクを推定する(ステップS22)。雑音強調マスク推定部214は、雑音混合音声が入力された雑音強調マスク推定モデルが出力した推定雑音強調マスクを推定値として出力する。
 [2-3-1:音声強調マスクロス]
 音声強調マスクロス算出部215は、音声強調マスク推定モデルが推定した推定音声強調マスクと理想的な理想音声強調マスクとを用いて、音声強調マスクロスを算出する(ステップS23)。音声強調マスクロスは、推定音声強調マスクと理想音声強調マスクとの違いを示す。音声強調マスクロスは、音声強調マスク損失関数Lにより算出されてもよい。音声強調マスク損失関数Lは、下記式3のように表現してもよい。すなわち、音声強調マスク損失関数Lは、各時刻における周波数瓶での推定値と理想値との差の二乗和で表現してもよい。Mは、音声強調マスクの推定値を表現している。また、M は、音声強調マスクの理想値を表現している。
 [式3]
Figure JPOXMLDOC01-appb-I000003
 なお、以下、各時刻における周波数瓶での事項を、「時間-周波数」という文言を用いて表現する場合がある。
 [2-3-2:雑音強調マスクロス]
 雑音強調マスクロス算出部216は、雑音強調マスク推定モデルが推定した推定雑音強調マスクと理想的な理想雑音強調マスクとを用いて、雑音強調マスクロスを算出する(ステップS24)。雑音強調マスクロスは、推定雑音強調マスクと理想雑音強調マスクとの違いを示す。雑音強調マスクロスは、雑音強調マスク損失関数Lにより算出されてもよい。雑音強調マスク損失関数Lは、下記式4のように表現してもよい。すなわち、雑音強調マスク損失関数Lは、時間-周波数での推定値と理想値との差の二乗和で表現してもよい。Mは、雑音強調マスクの推定値を表現している。また、M は、雑音強調マスクの理想値を表現している。
 [式4]
Figure JPOXMLDOC01-appb-I000004
 [2-3-3:拘束ロス]
 上記式1で表現される音声強調マスク、及び上記式2で表現される雑音強調マスクから、下記式5を成り立たせることができる。すなわち、時間-周波数における音声強調マスク、及び雑音強調マスクの和は理想的には一定値「1」となる。本実施形態では、これを拘束条件として用いる。
 [式5]
Figure JPOXMLDOC01-appb-I000005
 拘束ロス算出部211は、音声強調マスク推定モデルが推定した推定音声強調マスク、及び雑音強調マスク推定モデルが推定した推定雑音強調マスクを用いて拘束ロスを算出する(ステップS25)。拘束ロスは、音声強調マスク、及び雑音強調マスクの和(拘束条件)の推定値と理想値との違いを示す。
 拘束条件から外れた場合、すなわち下記式6から求まる値が「0」から離れるに従いペナルティを課すように損失関数を設計する。本実施形態における拘束ロスは、下記式6から求めることができる。
 [式6]
Figure JPOXMLDOC01-appb-I000006
 本実施形態では、音声と雑音との双方に関する情報として拘束条件を導入する。拘束条件は、後述するパラメータの更新の際に、音声と雑音との双方に関する情報として用いられてもよい。例えば、上記式6において、求まる値が「0」より大きい場合、すなわち、M+Mが「1」よりも大きい場合は、音声強調マスクのパラメータは、音声強調がより強くなるように更新され、雑音強調マスクのパラメータも、雑音強調がより強くなるように更新されてもよい。逆に上記式6において求まる値が「0」より小さい場合は、すなわち、M+Mが「1」よりも小さい場合は、音声強調マスクのパラメータは、音声強調がより弱くなるように更新され、雑音強調マスクのパラメータも、雑音強調がより弱くなるように更新されてもよい。
 推定された音声強調マスク及び雑音強調マスクが拘束条件から逸脱した量に基づいて、損失関数を導くことができる。本実施形態では、拘束条件として下記式7で示す拘束損失関数LSNを導入してもよい。下記式7は、時間-周波数の和を算出する。
 [式7]
Figure JPOXMLDOC01-appb-I000007
 拘束損失関数LSNは、「0」超過の場合に損失が存在する。この拘束条件の導入により、音声強調マスク及び雑音強調マスクの一方の推定値が正しくても、他方の推定値が正しくない場合は、損失関数の値は上昇する。音声強調マスク及び雑音強調マスクの各々を独立に学習させる比較例と比較して、正確な推定をすることができるようにモデルを学習させることができる。
 [2-3-4:合算ロス]
 合算ロス算出部217は、音声強調マスクロス、雑音強調マスクロス、及び拘束ロスを合算した合算ロスを算出する(ステップS26)。合算ロスは、合算損失関数LALLにより算出されてもよい。合算損失関数LALLは、下記式8のように表現してもよい。合算損失関数LALLから求まる合算ロスを、全ロスと呼んでもよい。
 [式8]
Figure JPOXMLDOC01-appb-I000008
 λは、「0」から「1」の間の値であってもよい。λは、1未満の正数であってもよい。λは、「0.01」、「0.1」等の「1」と比較して小さい値であってもよい。λは、モデルの学習過程において、一定でもよいし、変化させてもよい。λは、モデルの学習過程において、開始の際には特に小さく、終了の際に大きくなってもよい。
 なお、上記式8では、合算において、音声強調マスク損失関数Lと雑音強調マスク損失関数Lとに同じ重みを付しているが、音声強調マスク損失関数Lと雑音強調マスク損失関数Lとには異なる重みを付してもよい。例えば下記式9に示す合算損失関数LALLにより全ロスを算出してもよい。
 [式9]
Figure JPOXMLDOC01-appb-I000009
 上記「a」及び「b」は異なる値であってもよい。また、上記「a」及び「b」は同じ値であってもよい。上記「a」及び「b」が共に「1」の場合、上記式9は上記式8と等しくなる。
 [2-3-5:パラメータ更新]
 パラメータ更新部212は、合算ロス算出部217の算出結果に応じて、音声強調マスク推定モデルが含むパラメータと雑音強調マスク推定モデルが含むパラメータとを更新する(ステップS27)。パラメータ更新部212は、推定音声強調マスクと理想的な理想音声強調マスクとの違いを示す音声強調マスクロス、推定雑音強調マスクと理想的な理想雑音強調マスクとの違いを示す雑音強調マスクロス、及び拘束ロスを用いて、音声強調マスク推定モデルが含むパラメータと雑音強調マスク推定モデルが含むパラメータとを更新する。パラメータ更新部212は、パラメータ記憶部221に記憶されている音声強調マスク推定モデルが含むパラメータと雑音強調マスク推定モデルが含むパラメータとを更新する。パラメータ更新部212は、例えば、逆誤差伝搬法を用いて、音声強調マスク推定モデルのパラメータと雑音強調マスク推定モデルのパラメータとを更新してもよい。
 本実施形態における情報処理装置2は、拘束損失関数LSNを通じて音声強調マスク推定モデル、及び雑音強調マスク推定モデルの双方の学習をすすめることができる。音声強調マスク推定モデルは、音声強調マスク推定モデルが含む情報、及び雑音強調マスク推定モデルが含む情報を結合させた情報を用いて学習されることができる。
 [2-4:拘束条件の許容範囲]
 上述した通り、音声強調マスクと雑音強調マスクとの和は理想的には「1」であり、本実施形態では、これを拘束条件として用いている。一方、この拘束条件は、厳密に「1」でなくてもよい場合がある。この場合、例えば、下記式10に示す様に、「1」を「1+ε」に置き換えてもよい。
 [式10]
Figure JPOXMLDOC01-appb-I000010
 「ε」は「1」と比較して非常に小さな正数である。「1」を「1+ε」に置き換えた場合、「ε」程度の誤差を許容することができる。「ε」は、設計により任意に変更することができてもよい。上記式10の分母のS、及びNは、理想値を表現しており、分子のS、及びNは、推定値を表現している。
 「1」を「1+ε」に置き換得た場合、拘束損失関数LSNは、下記式11のように表現してもよい。
 [式11]
Figure JPOXMLDOC01-appb-I000011
 拘束損失関数LSNから求まる拘束ロスからは、音声強調マスク、及び雑音強調マスクの何れが理想値と比較して大きいのか又は小さいのかが不明である。音声強調マスクと雑音強調マスクとの和が「1」を超える場合、すなわち、正数の「ε」を採用した場合とは、推定音声強調マスク、及び推定雑音強調マスクの少なくとも一方が理想値と比較して大きい場合である。すなわち、正数の「ε」を採用すると、音声から完全に雑音を抑制できない可能性がある。よって、多少の雑音を許容してもよい場合においては、正数の「ε」を許容してもよい。
 一方で、音声強調マスクと雑音強調マスクとの和が「1」未満の場合、すなわち、負数の「ε」を採用した場合とは、推定音声強調マスク、及び推定雑音強調マスクの少なくとも一方が理想値と比較して小さい場合である。すなわち、負数の「ε」を採用すると、残すべき音である音声を削りすぎてしまう場合がある。音声を削りすぎると音声が聞き取りにくくなるので、負数の「ε」は許容しない。パラメータ更新部212は、少なくとも、音声強調マスクと雑音強調マスクとの和が「1」未満とはならないように、パラメータを更新する。
 [2-5:変形例]
 音声が、複数の人物の声を含む場合、音声強調マスク推定部213は、複数の人物各々に対応する音声強調マスクを推定してもよい。各々の人物に対応する音声強調マスク推定モデルが用意されており、各々の音声強調マスク推定モデルは、人物に対応する音声強調マスクを出力してもよい。
 例えば、音声が人物Aの声と人物Bの声を含む場合、人物Aの声に対応する音声強調マスク推定モデルA、及び人物Bの声に対応する音声強調マスク推定モデルBが学習されてもよい。音声強調マスク推定モデルAは、雑音混合音声が入力されると人物Aの声に対応する音声強調マスクMSAを出力し、音声強調マスク推定モデルBは、雑音混合音声が入力されると人物Bの声に対応する音声強調マスクMSBを出力してもよい。
 音声強調マスクMSA、及び音声強調マスクMSBを出力する場合の拘束条件は、下記式12のように表現してもよい。
 [式12]
Figure JPOXMLDOC01-appb-I000012

Figure JPOXMLDOC01-appb-I000013

Figure JPOXMLDOC01-appb-I000014
 なお、本実施形態では、平均二乗誤差を用いた損失関数を例を挙げたが、平均二乗誤差以外の手法を用いた損失関数を採用してもよい。
 [2-6:情報処理装置2の技術的効果]
 例えば上記非特許文献1には、モデルにより音声強調マスクと雑音強調マスクとを推定し、推定値と理想値との違いを損失として算出し、モデルを学習させる技術が記載されている。上記非特許文献1に記載されている技術を比較例とよぶ。比較例においては、モデルにより推定された音声強調マスクの評価に、モデルにより推定された雑音強調マスクの情報を用いることがない。また、モデルにより推定された雑音強調マスクの評価に、モデルにより推定された音声強調マスクの情報を用いることもない。すなわち、推定された音声強調マスクと推定された雑音強調マスクとは各々独立して評価されている。
 上述の通り、入力信号に含まれるのは、音声と雑音(上述の通り、本実施形態では音声以外を「雑音」とよんでいる)である。したがって、音声強調マスクの理想値と雑音強調マスクの理想値との和は一定値「1」となる。上述の通り、音声強調マスクと雑音強調マスクとの和が「1」を超えた場合、音声強調マスクの推定値、及び雑音強調マスクの推定値の少なくとも一方が理想値と比較して大きく、音声から完全に雑音を抑制できない可能性がある。また、音声強調マスクと雑音強調マスクとの和が「1」未満の場合、音声強調マスクの推定値、及び雑音強調マスクの推定値の少なくとも一方が理想値と比較して小さく、残すべき音である音声を削りすぎてしまう場合がある。例えば、短時間に急変する雑音が入力信号に含まれている場合の推定において、音声を削りすぎてしまう場合が多い。音声を削りすぎると音声が聞き取りにくくなるころを避けるべく、特に音声を削りすぎは抑制することが望ましい。
 比較例では、推定された音声強調マスクと推定された雑音強調マスクとは各々独立して評価しており、音声強調マスクと雑音強調マスクとの和が一定値「1」であるか否かは不明である。したがって、音声から完全に雑音を抑制できないのみならず、残すべき音である音声を削りすぎてしまう場合がある。また、雑音の種類は多様であり、雑音強調マスクの精度は、音声強調マスクと比較して低精度となる場合が多いので、雑音強調マスクの独立した評価をモデルのパラメータの更新に用いると、音声強調の精度が悪くなる可能性がある。
 第2実施形態における情報処理装置2は、音声強調マスク推定部213による推定結果と、雑音強調マスク推定部214による推定結果とを用いて算出される拘束ロスを導入する。そして、音声強調マスクロス、雑音強調マスクロス、及び拘束ロスを合算した合算ロスを用いて、音声強調マスク推定モデルが含むパラメータを更新するので、音声から雑音を抑制し、残すべき音である音声を削りすぎることを防ぐことができる。情報処理装置2は、精度よく音声強調をすることができる音声強調マスク推定モデルを作成することができる。
 [3:第3実施形態]
 続いて、情報処理装置、情報処理方法、及び記録媒体の第3実施形態について説明する。以下では、情報処理装置、情報処理方法、及び記録媒体の第3実施形態が適用された情報処理装置3を用いて、情報処理装置、情報処理方法、及び記録媒体の第3実施形態について説明する。
 図4は、第3実施形態における情報処理装置3の構成を示すブロック図である。第3実施形態における情報処理装置3は、拘束ロス算出部311の動作が第2実施形態における情報処理装置2と異なる。
 [3-1:情報処理装置3が行う情報処理動作]
 第3実施形態において拘束ロス算出部311は、推定音声強調マスク、及び推定雑音強調マスクに雑音混合音声の大きさを乗じて、拘束ロスを算出する。拘束ロス算出部311は、時間-周波数での推定音声強調マスク、及び推定雑音強調マスクに雑音混合音声の大きさを乗じて、拘束ロスを算出してもよい。雑音混合音声の大きさは、雑音混合音声の振幅の絶対値であってもよい。雑音混合音声の大きさは、雑音混合音声の対数パワースペクトラム(Log power spectrum)であってもよい。雑音混合音声の大きさは、正規化した値であってもよい。当該値は、設計により任意に変更することができてもよい。
 雑音混合音声の大きさをLPSinputと表す場合、第3実施形態における拘束損失関数LSNは、下記式13のように表現してもよい。
 [式13]
Figure JPOXMLDOC01-appb-I000015
 LPSinputを乗じると、LPSinputが大きい時間-周波数における誤差をより重視することができる。これは、振幅スペクトラム近似(magnitude spectrum approximation:MSA)の適用に相当する。
 LPSinputを乗じると、例えば新幹線の通過音のような短時間に急変する雑音の学習時に有効である。小さな入力信号よりも大きな入力信号の方に重みをを与えることになり、大きな入力信号に関する誤差をより強調することができる。
 [3-2:情報処理装置3の技術的効果]
 第3実施形態における情報処理装置3は、雑音混合音声の大きさを乗じるので、特に大きな入力信号に関する雑音を抑制することができる。
 [4:第4実施形態]
 続いて、情報処理装置、情報処理方法、及び記録媒体の第4実施形態について説明する。以下では、情報処理装置、情報処理方法、及び記録媒体の第4実施形態が適用された情報処理装置4を用いて、情報処理装置、情報処理方法、及び記録媒体の第4実施形態について説明する。
 図5は、第4実施形態における情報処理装置4の構成を示すブロック図である。第4実施形態における情報処理装置4は、拘束ロス算出部411の動作が第2実施形態における情報処理装置2、及び第3実施形態における情報処理装置3と異なる。
 [4-1:情報処理装置4が行う情報処理動作]
 第4実施形態において拘束ロス算出部411は、推定音声強調マスク、及び推定雑音強調マスクを所定の指数で冪乗して、拘束ロスを算出する。拘束ロス算出部411は、時間-周波数での推定音声強調マスク、及び推定雑音強調マスクを所定の指数で冪乗して、拘束ロスを算出してもよい。所定の指数は、0以上1以下の値を取り得てもよい。
 所定の指数をαと表す場合、第4実施形態における拘束損失関数LSNは、下記式14のように表現してもよい。
 [式14]
Figure JPOXMLDOC01-appb-I000016
 αを冪乗すると、例えば、エアコンの音のような定常的な雑音をより重点的に学習することができる。これは、べき乗圧縮(Power-law compression)の適用に相当する効果を期待することができる。
 なお、第4実施形態においても、第3実施形態において適用したLPSinputを適用してもよい。この場合、第4実施形態における拘束損失関数LSNは、下記式15のように表現してもよい。
 [式15]
Figure JPOXMLDOC01-appb-I000017
 [4-2:情報処理装置4の技術的効果]
 第4実施形態における情報処理装置4は、αを適用することにより、小さい信号をより重点的に学習することができる。
 [5:第5実施形態]
 続いて、情報処理装置、情報処理方法、及び記録媒体の第5実施形態について説明する。以下では、情報処理装置、情報処理方法、及び記録媒体の第5実施形態が適用された情報処理装置5を用いて、情報処理装置、情報処理方法、及び記録媒体の第5実施形態について説明する。
 [5-1:情報処理装置5の構成]
 図6に示すように、第5実施形態における情報処理装置5は、第2実施形態における情報処理装置2から第4実施形態における情報処理装置4と同様に、演算装置21と、記憶装置22とを備えている。更に、第5実施形態における情報処理装置5は、第2実施形態における情報処理装置2から第4実施形態における情報処理装置4と同様に、通信装置23と、入力装置24と、出力装置25とを備えていてもよい。但し、情報処理装置5は、通信装置23、入力装置24及び出力装置25のうちの少なくとも1つを備えていなくてもよい。第5実施形態における情報処理装置5は、演算装置21内に音響特徴抽出部519が更に実現される点で、第2実施形態における情報処理装置2から第4実施形態における情報処理装置4と異なる。情報処理装置5のその他の特徴は、第2実施形態における情報処理装置2から第4実施形態における情報処理装置4の少なくとも1つのその他の特徴と同一であってもよい。このため、以下では、すでに説明した各実施形態と異なる部分について詳細に説明し、その他の重複する部分については適宜説明を省略するものとする。
 [5-2:情報処理装置5が行う情報処理動作]
 図7を参照しながら、情報処理装置5が行う情報処理動作について説明する。図7は、情報処理装置5が行う情報処理動作の流れを示すフローチャートである。
 図7に示す様に、雑音混合音声入力部518は、雑音混合音声を取得し、音響特徴抽出部519に入力する(ステップS50)。音響特徴抽出部519は、雑音混合音声から雑音混合音声の特徴である雑音混合音声特徴を抽出する(ステップS51)。音響特徴抽出部519は、雑音混合音声が入力され、雑音混合音声特徴を出力する音響特徴抽出モデルを有していてもよい。音響特徴抽出モデルは、RNNで実装してもよい。
 音声強調マスク推定部513は、音声強調マスク推定モデルを用いて音声強調マスクを推定する(ステップS52)。第5実施形態において、音声強調マスク推定モデルは、雑音混合音声特徴が入力され、推定音声強調マスクを出力してもよい。音声強調マスク推定部513は、雑音混合音声特徴が入力された音声強調マスク推定モデルが出力した推定音声強調マスクを推定値として出力してもよい。
 雑音強調マスク推定部514は、雑音強調マスク推定モデルを用いて雑音強調マスクを推定する(ステップS53)。第5実施形態において、雑音強調マスク推定モデルは、雑音混合音声特徴が入力され、推定雑音強調マスクを出力してもよい。雑音強調マスク推定部514は、雑音混合音声特徴が入力された雑音強調マスク推定モデルが出力した推定雑音強調マスクを推定値として出力してもよい。
 音声強調マスクロス算出部515は、音声強調マスク推定モデルが推定した推定音声強調マスクと理想音声強調マスクとを用いて、音声強調マスクロスを算出する(ステップS54)。理想音声強調マスクは、教師入力信号から抽出された雑音混合音声特徴に対して理想的な音声強調マスクであってもよい。
 雑音強調マスクロス算出部516は、雑音強調マスク推定モデルが推定した推定雑音強調マスクと理想雑音強調マスクとを用いて、雑音強調マスクロスを算出する(ステップS55)。理想雑音強調マスクは、教師入力信号から抽出された雑音混合音声特徴に対し、理想的な雑音強調マスクであってもよい。
 拘束ロス算出部211は、音声強調マスク推定モデルが推定した推定音声強調マスク、及び雑音強調マスク推定モデルが推定した推定雑音強調マスクを用いて拘束ロスを算出する(ステップS25)。合算ロス算出部217は、音声強調マスクロス、雑音強調マスクロス、及び拘束ロスを合算した合算ロスを算出する(ステップS26)。
 パラメータ更新部512は、合算ロス算出部217の算出結果に応じて、音声強調マスク推定モデルが含むパラメータ、及び雑音強調マスク推定モデルが含むパラメータと伴に、音響特徴抽出モデルが含むパラメータを更新する(ステップS56)。音響特徴抽出モデルは、音声に関する情報と雑音に関する情報との両方を用いて学習されてもよい。
 第5実施形態における情報処理装置5は、音声強調マスク推定モデル、及び音響特徴抽出モデルを機械学習により作成した。第5実施形態において、音声を強調する場面では、作成した音声強調マスク推定モデルに加えて、音響特徴抽出モデルが適用される。
 [5-3:情報処理装置5の技術的効果]
 第5実施形態における情報処理装置5は、音響特徴抽出部519を備えることにより、音声強調マスク推定モデル、及び雑音強調マスク推定モデルの双方への入力信号を大きくすることができるので、より好ましく音声強調をすることができる。一方で、第2実施形態のように、音響特徴抽出部519を備えない場合は、動作を軽くすることができる。
 例えば、上述した比較例の場合、音響特徴抽出部519を有することにより、逆伝搬の際に音声、及び雑音の一方の信号を削りすぎる方向へモデルのパラメータが調整される可能性がある。これに対し、本実施形態では、音声強調マスク推定モデルによる推定結果と、雑音強調マスク推定モデルによる推定結果との情報を用いて算出される拘束ロスを導入するので、音響特徴抽出部519を有していても、逆伝搬の際に音声、及び雑音の一方の信号を削りすぎる方向へモデルのパラメータが調整されることがない。
 [6:第6実施形態]
 続いて、情報処理装置、情報処理方法、及び記録媒体の第6実施形態について説明する。以下では、情報処理装置、情報処理方法、及び記録媒体の第6実施形態が適用された情報処理装置6を用いて、情報処理装置、情報処理方法、及び記録媒体の第6実施形態について説明する。
 [6-1:情報処理装置6の構成]
 図8に示すように、第6実施形態における情報処理装置6は、第2実施形態における情報処理装置2から第5実施形態における情報処理装置5と同様に、演算装置21と、記憶装置22とを備えている。更に、第6実施形態における情報処理装置6は、第2実施形態における情報処理装置2から第5実施形態における情報処理装置5と同様に、通信装置23と、入力装置24と、出力装置25とを備えていてもよい。但し、情報処理装置6は、通信装置23、入力装置24及び出力装置25のうちの少なくとも1つを備えていなくてもよい。第6実施形態における情報処理装置6は、拘束ロス算出部611が雑音拘束ロス算出部6111、及び音声拘束ロス算出部6112を有する点で、第2実施形態における情報処理装置2から第5実施形態における情報処理装置5と異なる。情報処理装置6のその他の特徴は、第2実施形態における情報処理装置2から第5実施形態における情報処理装置5の少なくとも1つのその他の特徴と同一であってもよい。このため、以下では、すでに説明した各実施形態と異なる部分について詳細に説明し、その他の重複する部分については適宜説明を省略するものとする。
 [6-2:情報処理装置6が行う情報処理動作]
 図9を参照しながら、情報処理装置6が行う情報処理動作について説明する。図9は、情報処理装置6が行う情報処理動作の流れを示すフローチャートである。
 図9に示す様に、雑音混合音声入力部218は、雑音混合音声を取得し、音声強調マスク推定部213、及び雑音強調マスク推定部214に入力する(ステップS20)。音声強調マスク推定部213は、音声強調マスク推定モデルを用いて音声強調マスクを推定する(ステップS21)。雑音強調マスク推定部214は、雑音強調マスク推定モデルを用いて雑音強調マスクを推定する(ステップS22)。音声強調マスクロス算出部215は、音声強調マスク推定モデルが推定した推定音声強調マスクと理想音声強調マスクとを用いて、音声強調マスクロスを算出する(ステップS23)。雑音強調マスクロス算出部216は、雑音強調マスク推定モデルが推定した推定雑音強調マスクと理想雑音強調マスクとを用いて、雑音強調マスクロスを算出する(ステップS24)。
 [6-2-1:雑音拘束ロス]
 雑音拘束ロス算出部6111は、推定音声強調マスク、及び理想雑音強調マスクを用いて雑音拘束ロスを算出する(ステップS60)。雑音拘束ロス算出部6111は、推定音声強調マスク、及び上記式5で表現される拘束条件に基づいて、拘束雑音強調マスクを算出する。拘束雑音強調マスクをM と表す場合、拘束雑音強調マスクは、下記式16のように表現してもよい。
 [式16]
Figure JPOXMLDOC01-appb-I000018
 雑音拘束ロス算出部6111は、拘束雑音強調マスクと理想雑音強調マスクとの違いを示す雑音拘束ロスを算出する。雑音拘束ロスが雑音拘束損失関数L により算出される場合、雑音拘束損失関数L は、下記式17のように表現してもよい。
 [式17]
Figure JPOXMLDOC01-appb-I000019
 [6-2-2:音声拘束ロス]
 音声拘束ロス算出部6112は、推定雑音強調マスク、及び理想雑音強調マスクを用いて音声拘束ロスを算出する(ステップS61)。音声拘束ロス算出部6112は、推定雑音強調マスク、及び上記式5で表現される拘束条件に基づいて、拘束音声強調マスクを算出する。拘束音声強調マスクをM と表す場合、拘束音声強調マスクは、下記式18のように表現してもよい。
 [式18]
Figure JPOXMLDOC01-appb-I000020
 音声拘束ロス算出部6112は、拘束音声強調マスクと理想音声強調マスクとの違いを示す音声拘束ロスを算出する。音声拘束ロスが音声拘束損失関数L により算出される場合、音声拘束損失関数L は、下記式19のように表現してもよい。
 [式19]
Figure JPOXMLDOC01-appb-I000021
 [6-2-3:合算ロス]
 合算ロス算出部617は、音声強調マスクロス、雑音強調マスクロス、雑音拘束ロス及び音声拘束ロスを合算した合算ロスを算出する(ステップS62)。合算ロスが合算損失関数LALLにより算出される場合、合算損失関数LALLは、下記式20のように表現してもよい。

 [式20]
Figure JPOXMLDOC01-appb-I000022

Figure JPOXMLDOC01-appb-I000023
 Mは、拘束条件も含んだマスクの推定値と考えることができる。上記式20の第1項と第2項の和はM +M =1を自動的に満たすので、第2実施形態で用いた式5で表現される拘束条件と同等の効果を得ることができる。すなわち、第6実施形態における情報処理装置6は第2実施形態における情報処理装置2と同等の効果を得ることができる。
 上記式20は、下記式21のように書き直してもよい。第6実施形態における合算ロスは、拘束条件も含んだマスクの推定値と理想値との違いを示していると考えることができる。
 [式21]
Figure JPOXMLDOC01-appb-I000024
 パラメータ更新部212は、合算ロス算出部217の算出結果に応じて、音声強調マスク推定モデルが含むパラメータと雑音強調マスク推定モデルが含むパラメータとを更新する(ステップS27)。
 [6-3:情報処理装置6の技術的効果]
 第6実施形態における情報処理装置6は、第2実施形態において採用している「λ」を採用しない。したがって、第6実施形態における情報処理装置6は、「λ」の決定にかかる処理負荷が小さい点において、第2実施形態における情報処理装置2よりも処理負荷を小さくすることができる。
 [7:付記]
 以上説明した実施形態に関して、更に以下の付記を開示する。
 [付記1]
 雑音混合音声が入力された音声強調マスク推定モデルが出力した推定音声強調マスク、及び前記雑音混合音声が入力された雑音強調マスク推定モデルが出力した推定雑音強調マスクを用いて拘束ロスを算出する拘束ロス算出手段と、
 前記推定音声強調マスクと目標音声強調マスクとの違いを示す音声強調マスクロス、前記推定雑音強調マスクと目標雑音強調マスクとの違いを示す雑音強調マスクロス、及び前記拘束ロスを用いて、前記音声強調マスク推定モデルが含むパラメータ、及び前記雑音強調マスク推定モデルが含むパラメータを更新するパラメータ更新手段と
 を備える情報処理装置。
 [付記2]
 前記音声強調マスク推定モデルを用いて前記推定音声強調マスクを推定する音声強調マスク推定手段と、
 前記雑音強調マスク推定モデルを用いて前記推定雑音強調マスクを推定する雑音強調マスク推定手段と、
 前記推定音声強調マスクと前記目標音声強調マスクとを用いて、前記音声強調マスクロスを算出する音声強調マスクロス算出手段と、
 前記推定雑音強調マスクと前記目標雑音強調マスクとを用いて、前記雑音強調マスクロスを算出する雑音強調マスクロス算出手段と、
 前記音声強調マスクロス、前記雑音強調マスクロス、及び前記拘束ロスを合算した合算ロスを算出する合算ロス算出手段と
 を更に備え、
 前記パラメータ更新手段は、前記合算ロス算出手段の算出結果に応じて、前記音声強調マスク推定モデルが含むパラメータと前記雑音強調マスク推定モデルが含むパラメータとを更新する
 請求項1に記載の情報処理装置。
 [付記3]
 前記拘束ロス算出手段は、前記推定音声強調マスク、及び前記推定雑音強調マスクに前記雑音混合音声の大きさを乗じて、前記拘束ロスを算出する
 請求項1又は2に記載の情報処理装置。
 [付記4]
 前記拘束ロス算出手段は、前記推定音声強調マスク、及び前記推定雑音強調マスクを所定の指数で冪乗して、前記拘束ロスを算出する
 請求項1又は2に記載の情報処理装置。
 [付記5]
 前記雑音混合音声から雑音混合音声特徴を抽出する音響特徴抽出手段を更に備え、
 前記音声強調マスク推定モデルは、前記雑音混合音声特徴が入力され、前記推定音声強調マスクを出力し、
 前記雑音強調マスク推定モデルは、前記雑音混合音声特徴が入力され、前記推定雑音強調マスクを出力する
 請求項1又は2に記載の情報処理装置。
 [付記6]
 前記拘束ロス算出手段は、
  前記推定雑音強調マスク、及び前記目標音声強調マスクを用いて音声拘束ロスを算出する音声拘束ロス算出手段と、
  前記推定音声強調マスク、及び前記目標雑音強調マスクを用いて雑音拘束ロスを算出する雑音拘束ロス算出手段とを有し、
 前記拘束ロスは、前記音声拘束ロス、及び前記雑音拘束ロスを含む
 請求項1又は2に記載の情報処理装置。
 [付記7]
 雑音混合音声が入力された音声強調マスク推定モデルが出力した推定音声強調マスク、及び前記雑音混合音声が入力された雑音強調マスク推定モデルが出力した推定雑音強調マスクを用いて拘束ロスを算出し、
 前記推定音声強調マスクと目標音声強調マスクとの違いを示す音声強調マスクロス、前記推定雑音強調マスクと目標雑音強調マスクとの違いを示す雑音強調マスクロス、及び前記拘束ロスを用いて、前記音声強調マスク推定モデルが含むパラメータ、及び前記雑音強調マスク推定モデルが含むパラメータを更新する
 情報処理方法。
 [付記8]
 コンピュータに、
 雑音混合音声が入力された音声強調マスク推定モデルが出力した推定音声強調マスク、及び前記雑音混合音声が入力された雑音強調マスク推定モデルが出力した推定雑音強調マスクを用いて拘束ロスを算出し、
 前記推定音声強調マスクと目標音声強調マスクとの違いを示す音声強調マスクロス、前記推定雑音強調マスクと目標雑音強調マスクとの違いを示す雑音強調マスクロス、及び前記拘束ロスを用いて、前記音声強調マスク推定モデルが含むパラメータ、及び前記雑音強調マスク推定モデルが含むパラメータを更新する
 情報処理方法を実行させるためのコンピュータプログラムが記録された記録媒体。
 上述の各実施形態の構成要件の少なくとも一部は、上述の各実施形態の構成要件の少なくとも他の一部と適宜組み合わせることができる。上述の各実施形態の構成要件のうちの一部が用いられなくてもよい。
 この開示は上記実施形態に限定されるものではない。この開示は、請求の範囲及び明細書全体から読み取るこのできる技術的思想に反しない範囲で適宜変更可能である。そのような変更を伴う情報処理装置、情報処理方法、及び、記録媒体もまた、この開示の技術的思想に含まれる。また、法令で許容される限りにおいて、本願明細書に記載された全ての公開公報及び論文をここに取り込む。
 法令で許容される限りにおいて、この出願は、2023年3月13日に出願された日本出願特願2023-039058を基礎とする優先権を主張し、その開示の全てをここに取り込む。
1,2,3,4,5,6 情報処理装置
11,211,311,411,611 拘束ロス算出部
12,212,512 パラメータ更新部
221 パラメータ記憶部
222,522 教師情報記憶部
213,513 音声強調マスク推定部
214,514 雑音強調マスク推定部
215,515 音声強調マスクロス算出部
216,516 雑音強調マスクロス算出部
217,617 合算ロス算出部
218,518 雑音混合音声入力部
519 音響特徴抽出部
6111 雑音拘束ロス算出部
6112 音声拘束ロス算出部

Claims (8)

  1.  雑音混合音声が入力された音声強調マスク推定モデルが出力した推定音声強調マスク、及び前記雑音混合音声が入力された雑音強調マスク推定モデルが出力した推定雑音強調マスクを用いて拘束ロスを算出する拘束ロス算出手段と、
     前記推定音声強調マスクと雑音混合音声中の音声と雑音が既知な場合に計算される目標音声強調マスクとの違いを示す音声強調マスクロス、前記推定雑音強調マスクと目標雑音強調マスクとの違いを示す雑音強調マスクロス、及び前記拘束ロスを用いて、前記音声強調マスク推定モデルが含むパラメータ、及び前記雑音強調マスク推定モデルが含むパラメータを更新するパラメータ更新手段と
     を備える情報処理装置。
  2.  前記音声強調マスク推定モデルを用いて前記推定音声強調マスクを推定する音声強調マスク推定手段と、
     前記雑音強調マスク推定モデルを用いて前記推定雑音強調マスクを推定する雑音強調マスク推定手段と、
     前記推定音声強調マスクと前記目標音声強調マスクとを用いて、前記音声強調マスクロスを算出する音声強調マスクロス算出手段と、
     前記推定雑音強調マスクと前記目標雑音強調マスクとを用いて、前記雑音強調マスクロスを算出する雑音強調マスクロス算出手段と、
     前記音声強調マスクロス、前記雑音強調マスクロス、及び前記拘束ロスを合算した合算ロスを算出する合算ロス算出手段と
     を更に備え、
     前記パラメータ更新手段は、前記合算ロス算出手段の算出結果に応じて、前記音声強調マスク推定モデルが含むパラメータと前記雑音強調マスク推定モデルが含むパラメータとを更新する
     請求項1に記載の情報処理装置。
  3.  前記拘束ロス算出手段は、前記推定音声強調マスク、及び前記推定雑音強調マスクに前記雑音混合音声の大きさを乗じて、前記拘束ロスを算出する
     請求項1又は2に記載の情報処理装置。
  4.  前記拘束ロス算出手段は、前記推定音声強調マスク、及び前記推定雑音強調マスクを所定の指数で冪乗して、前記拘束ロスを算出する
     請求項1又は2に記載の情報処理装置。
  5.  前記雑音混合音声から雑音混合音声特徴を抽出する音響特徴抽出手段を更に備え、
     前記音声強調マスク推定モデルは、前記雑音混合音声特徴が入力され、前記推定音声強調マスクを出力し、
     前記雑音強調マスク推定モデルは、前記雑音混合音声特徴が入力され、前記推定雑音強調マスクを出力する
     請求項1又は2に記載の情報処理装置。
  6.  前記拘束ロス算出手段は、
      前記推定雑音強調マスク、及び前記目標音声強調マスクを用いて音声拘束ロスを算出する音声拘束ロス算出手段と、
      前記推定音声強調マスク、及び前記目標雑音強調マスクを用いて雑音拘束ロスを算出する雑音拘束ロス算出手段とを有し、
     前記拘束ロスは、前記音声拘束ロス、及び前記雑音拘束ロスを含む
     請求項1又は2に記載の情報処理装置。
  7.  雑音混合音声が入力された音声強調マスク推定モデルが出力した推定音声強調マスク、及び前記雑音混合音声が入力された雑音強調マスク推定モデルが出力した推定雑音強調マスクを用いて拘束ロスを算出し、
     前記推定音声強調マスクと雑音混合音声中の音声と雑音が既知な場合に計算される目標音声強調マスクとの違いを示す音声強調マスクロス、前記推定雑音強調マスクと目標雑音強調マスクとの違いを示す雑音強調マスクロス、及び前記拘束ロスを用いて、前記音声強調マスク推定モデルが含むパラメータ、及び前記雑音強調マスク推定モデルが含むパラメータを更新する
     情報処理方法。
  8.  コンピュータに、
     雑音混合音声が入力された音声強調マスク推定モデルが出力した推定音声強調マスク、及び前記雑音混合音声が入力された雑音強調マスク推定モデルが出力した推定雑音強調マスクを用いて拘束ロスを算出し、
     前記推定音声強調マスクと雑音混合音声中の音声と雑音が既知な場合に計算される目標音声強調マスクとの違いを示す音声強調マスクロス、前記推定雑音強調マスクと目標雑音強調マスクとの違いを示す雑音強調マスクロス、及び前記拘束ロスを用いて、前記音声強調マスク推定モデルが含むパラメータ、及び前記雑音強調マスク推定モデルが含むパラメータを更新する
     情報処理方法を実行させるためのコンピュータプログラムが記録された記録媒体。
PCT/JP2024/000990 2023-03-13 2024-01-16 情報処理装置、情報処理方法、及び、記録媒体 Ceased WO2024190063A1 (ja)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP2025506508A JPWO2024190063A5 (ja) 2024-01-16 情報処理装置、情報処理方法、及び、コンピュータプログラム

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
JP2023039058 2023-03-13
JP2023-039058 2023-03-13

Publications (1)

Publication Number Publication Date
WO2024190063A1 true WO2024190063A1 (ja) 2024-09-19

Family

ID=92754770

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/JP2024/000990 Ceased WO2024190063A1 (ja) 2023-03-13 2024-01-16 情報処理装置、情報処理方法、及び、記録媒体

Country Status (1)

Country Link
WO (1) WO2024190063A1 (ja)

Non-Patent Citations (2)

* Cited by examiner, † Cited by third party
Title
ERDOGAN HAKAN, YOSHIOKA TAKUYA: "Investigations on Data Augmentation and Loss Functions for Deep Learning Based Speech-Background Separation", INTERSPEECH 2018, ISCA, 1 January 2018 (2018-01-01), pages 3499 - 3503, XP093210252, DOI: 10.21437/Interspeech.2018-2441 *
TSUNODA, RYOTA ET AL.: "Pre-training method for audio-visual speech recognition in interference speaker environment", PROCEEDINGS OF THE 2022 SPRING MEETING THE ACOUSTICAL SOCIETY OF JAPAN; MARCH 9 - 11, 2022, vol. 2022, 23 February 2022 (2022-02-23) - 11 March 2022 (2022-03-11), pages 1057 - 1060, XP009558531 *

Also Published As

Publication number Publication date
JPWO2024190063A1 (ja) 2024-09-19

Similar Documents

Publication Publication Date Title
Fu et al. Metricgan+: An improved version of metricgan for speech enhancement
JP5666444B2 (ja) 特徴抽出を使用してスピーチ強調のためにオーディオ信号を処理する装置及び方法
JP5842056B2 (ja) 雑音推定装置、雑音推定方法、雑音推定プログラム及び記録媒体
US12597434B2 (en) Control of speech preservation in speech enhancement
Abdullah et al. Towards more efficient DNN-based speech enhancement using quantized correlation mask
JPWO2020039571A1 (ja) 音声分離装置、音声分離方法、音声分離プログラム、及び音声分離システム
JP7667247B2 (ja) 機械学習を用いたノイズ削減
Xu et al. Deep noise suppression maximizing non-differentiable PESQ mediated by a non-intrusive PESQNet
CN113990343B (zh) 语音降噪模型的训练方法和装置及语音降噪方法和装置
KR20200092501A (ko) 합성 음성 신호 생성 방법, 뉴럴 보코더 및 뉴럴 보코더의 훈련 방법
CN113707167A (zh) 残留回声抑制模型的训练方法和训练装置
US12579961B2 (en) Music enhancement systems
EP1995723A1 (en) Neuroevolution training system
KR102198597B1 (ko) 뉴럴 보코더 및 화자 적응형 모델을 구현하기 위한 뉴럴 보코더의 훈련 방법
JP5994639B2 (ja) 有音区間検出装置、有音区間検出方法、及び有音区間検出プログラム
Kumar et al. Comparative studies of single-channel speech enhancement techniques
Richter et al. Speech signal improvement using causal generative diffusion models
Karthik et al. An optimized convolutional neural network for speech enhancement
Elshamy et al. DNN-based cepstral excitation manipulation for speech enhancement
CN113393852B (zh) 语音增强模型的构建方法及系统、语音增强方法及系统
KR102505653B1 (ko) 심화신경망을 이용한 에코 및 잡음 통합 제거 방법 및 장치
WO2024190063A1 (ja) 情報処理装置、情報処理方法、及び、記録媒体
Biswas et al. Optimal near-end speech intelligibility improvement using CLPSO-based voice transformation in realistic noisy environments
CN114141266A (zh) 基于pesq驱动的强化学习估计先验信噪比的语音增强方法
Vinothkumar Speech enhancement with background noise suppression in various data corpus using Bi-LSTM algorithm

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 24770185

Country of ref document: EP

Kind code of ref document: A1

ENP Entry into the national phase

Ref document number: 2025506508

Country of ref document: JP

Kind code of ref document: A

WWE Wipo information: entry into national phase

Ref document number: 2025506508

Country of ref document: JP

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 24770185

Country of ref document: EP

Kind code of ref document: A1