WO2024190063A1 - 情報処理装置、情報処理方法、及び、記録媒体 - Google Patents
情報処理装置、情報処理方法、及び、記録媒体 Download PDFInfo
- Publication number
- WO2024190063A1 WO2024190063A1 PCT/JP2024/000990 JP2024000990W WO2024190063A1 WO 2024190063 A1 WO2024190063 A1 WO 2024190063A1 JP 2024000990 W JP2024000990 W JP 2024000990W WO 2024190063 A1 WO2024190063 A1 WO 2024190063A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- noise
- speech
- enhancement mask
- mask
- loss
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L21/00—Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
- G10L21/02—Speech enhancement, e.g. noise reduction or echo cancellation
- G10L21/0272—Voice signal separating
Definitions
- This disclosure relates to the technical fields of information processing devices, information processing methods, and recording media.
- Non-Patent Document 1 describes a technology that estimates a speech enhancement mask and a noise enhancement mask using a model, calculates an index using each estimated mask, and trains the model using the deviation between the calculated index and the ideal value of the index.
- the objective of this disclosure is to provide an information processing device, an information processing method, and a recording medium that aim to improve upon the technology described in prior art documents.
- One aspect of the information processing device includes a constraint loss calculation means for calculating a constraint loss using an estimated speech enhancement mask output by a speech enhancement mask estimation model to which noise-mixed speech is input, and an estimated noise enhancement mask output by a noise enhancement mask estimation model to which the noise-mixed speech is input, and a parameter update means for updating parameters included in the speech enhancement mask estimation model and parameters included in the noise enhancement mask estimation model using a speech enhancement mask loss indicating the difference between the estimated speech enhancement mask and a target speech enhancement mask calculated when the speech and noise in the noise-mixed speech are known, a noise enhancement mask loss indicating the difference between the estimated noise enhancement mask and a target noise enhancement mask, and the constraint loss.
- One aspect of the information processing method is to calculate a constraint loss using an estimated speech enhancement mask output by a speech enhancement mask estimation model to which noise-mixed speech is input, and an estimated noise enhancement mask output by a noise enhancement mask estimation model to which the noise-mixed speech is input, and to update parameters included in the speech enhancement mask estimation model and parameters included in the noise enhancement mask estimation model using a speech enhancement mask loss indicating the difference between the estimated speech enhancement mask and a target speech enhancement mask calculated when the speech and noise in the noise-mixed speech are known, a noise enhancement mask loss indicating the difference between the estimated noise enhancement mask and a target noise enhancement mask, and the constraint loss.
- a computer program is recorded to cause a computer to execute an information processing method for calculating a constraint loss using an estimated speech enhancement mask output by a speech enhancement mask estimation model to which noise-mixed speech is input, and an estimated noise enhancement mask output by a noise enhancement mask estimation model to which the noise-mixed speech is input, and updating parameters included in the speech enhancement mask estimation model and parameters included in the noise enhancement mask estimation model using a speech enhancement mask loss indicating the difference between the estimated speech enhancement mask and a target speech enhancement mask calculated when the speech and noise in the noise-mixed speech are known, a noise enhancement mask loss indicating the difference between the estimated noise enhancement mask and a target noise enhancement mask, and the constraint loss.
- FIG. 1 is a block diagram showing the configuration of an information processing apparatus according to the first embodiment.
- FIG. 2 is a block diagram showing the configuration of an information processing device according to the second embodiment.
- FIG. 3 is a flowchart showing the flow of information processing operations of the information processing device in the second embodiment.
- FIG. 4 is a block diagram showing the configuration of an information processing apparatus according to the third embodiment.
- FIG. 5 is a block diagram showing the configuration of an information processing apparatus according to the fourth embodiment.
- FIG. 6 is a block diagram showing the configuration of an information processing device according to the fifth embodiment.
- FIG. 7 is a flowchart showing the flow of information processing operations of the information processing device according to the fifth embodiment.
- FIG. 8 is a block diagram showing the configuration of an information processing apparatus according to the sixth embodiment.
- FIG. 9 is a flowchart showing the flow of information processing operations of the information processing device in the sixth embodiment.
- a first embodiment of an information processing device, an information processing method, and a recording medium will be described below.
- the first embodiment of the information processing device, the information processing method, and the recording medium will be described using an information processing device 1 to which the first embodiment of the information processing device, the information processing method, and the recording medium is applied.
- FIG. 1 is a block diagram showing the configuration of an information processing device 1 in the first embodiment. As shown in FIG. 1, the information processing device 1 includes a constraint loss calculation unit 11 and a parameter update unit 12.
- the constraint loss calculation unit 11 calculates the constraint loss using the estimated speech emphasis mask output by the speech emphasis mask estimation model to which the noise-mixed speech is input, and the estimated noise emphasis mask output by the noise emphasis mask estimation model to which the noise-mixed speech is input.
- the parameter update unit 12 updates the parameters included in the speech emphasis mask estimation model and the parameters included in the noise emphasis mask estimation model, using the speech emphasis mask loss indicating the difference between the estimated speech emphasis mask and the target speech emphasis mask, the noise emphasis mask loss indicating the difference between the estimated noise emphasis mask and the target noise emphasis mask, and the constraint loss.
- the information processing device 1 in the first embodiment introduces a constraint loss calculated using an estimated speech enhancement mask output by a speech enhancement mask estimation model that relatively reduces the time and volume of a frequency band of noise other than the target speech included in the input speech, and an estimated noise enhancement mask output by a noise enhancement mask estimation model that relatively reduces the time and volume of a frequency band of the target speech included in the input speech.
- the information processing device 1 updates parameters included in the speech enhancement mask estimation model using the difference between the estimated value and target value of the speech enhancement mask, the difference between the estimated value and target value of the noise enhancement mask, and the constraint loss, and can generate a speech enhancement mask estimation model that can perform speech enhancement with high accuracy.
- audio refers to the target sound signal.
- audio may be referred to as the "target signal.”
- noise refers to the non-target sound signal.
- non-target signal refers to the "non-target signal.”
- the audio may be a sound that one wishes to focus on.
- the audio may be a person's voice.
- the audio may be a specific person's voice.
- the specific person may be one or more people.
- the specific person may be a person who is near a mechanism that captures sound, such as a microphone. "Audio" may represent different sound signals depending on the scene.
- the "noise-mixed speech” is “speech” mixed with “noise.”
- the “noise-mixed speech” includes “speech,” which is a target sound signal, and “noise,” which is a non-target sound signal.
- the sound signals included in the "noise-mixed speech” are either “speech” or “noise.”
- a sound signal that is not “speech” and is included in the “noise-mixed speech” is “noise.”
- a sound signal that is not “noise” and is included in the “noise-mixed speech” is “speech.”
- the input signal input to the information processing device 2 is “noise-mixed speech.”
- the input signal input to the information processing device 2 is composed of "speech,” which is a target signal, and "noise,” which is a non-target signal.
- Voice enhancement technology is a technology that makes the voice louder relative to the noise from a noisy voice.
- This technology may be a technology that suppresses noise from a voice mixed with noise, emphasizes the voice, and provides a voice that is easier to hear.
- This technology may be a technology that emphasizes only the voice from a voice mixed with noise, and provides a voice that is easier to hear.
- the voice enhancement technology may be capable of removing noise from a telephone call voice in a high-noise situation.
- the voice enhancement technology may also emphasize the voice of the other party in the call, enabling smoother communication.
- the voice enhancement technology may also improve the recognition rate of a voice recognizer.
- the mask type speech enhancement technology estimates a speech enhancement mask that specifies the time, frequency band, and amount of reduction in the volume of the noise to be reduced.
- the mask type speech enhancement technology may be a technology that makes speech easier to hear by applying the speech enhancement mask to noise-mixed speech that contains noise.
- the mask type speech enhancement technology may use machine learning to generate a speech enhancement mask estimation model that estimates the speech enhancement mask.
- the speech enhancement mask estimation model created in this embodiment may be applied to situations where speech is emphasized.
- the speech enhancement mask estimation model created in this embodiment is trained using a noise-mixed speech sound obtained by superimposing noise on a clean speech sound as an input signal.
- the clean speech sound may be a sound with very little noise.
- the machine learning in this embodiment may be deep learning. [2-1-3: Speech enhancement mask and noise enhancement mask]
- the speech enhancement mask estimation model is a model that outputs a speech enhancement mask when noise-mixed speech is input.
- the speech enhancement mask estimated by the speech enhancement mask estimation model may be called an estimated speech enhancement mask.
- the estimated speech enhancement mask may also be called an estimated value.
- the speech enhancement mask may be expressed as in the following Equation 1. [Formula 1]
- S(t,f) 2 denotes the power of "speech"
- N(t,f) 2 denotes the power of "noise”.
- the speech enhancement mask is expressed as the ratio of "speech" power to the power of "noise-mixed speech.”
- a speech enhancement mask estimation model is created through training in order to enhance speech, but it is expected that the training of the speech enhancement mask estimation model can be accelerated by training noise enhancement along with speech enhancement.
- the noise enhancement mask relatively reduces the volume of the time and frequency band of the target voice included in the input voice.
- the noise enhancement mask estimation model is a model that outputs a noise enhancement mask when noise-mixed voice is input.
- the noise enhancement mask estimated by the noise enhancement mask estimation model may be called an estimated noise enhancement mask.
- the estimated noise enhancement mask may also be called an estimated value.
- the noise enhancement mask may be expressed as in the following Equation 2. [Formula 2]
- the noise enhancement mask represents the proportion of "noise” power in the “noise-mixed speech” power. [2-1-4: Learning]
- the information processing device 2 may create a speech enhancement mask estimation model capable of estimating an ideal speech enhancement mask by learning a speech enhancement mask estimation model and a noise enhancement mask estimation model.
- the speech enhancement mask estimation model and the noise enhancement mask estimation model may be implemented by a neural network (NN).
- the speech enhancement mask estimation model and the noise enhancement mask estimation model may be implemented by a recurrent neural network (RNN).
- the information processing device 2 may update parameters included in the speech enhancement mask estimation model and parameters included in the noise enhancement mask estimation model according to estimates by the speech enhancement mask estimation model and estimates by the noise enhancement mask estimation model, and create a speech enhancement mask estimation model capable of estimating an ideal speech enhancement mask.
- the ideal speech enhancement mask (sometimes referred to as the "target speech enhancement mask”) is a speech enhancement mask that is calculated when the speech and noise in the noise-mixed speech are known.
- the ideal noise enhancement mask (sometimes referred to as the "target noise enhancement mask”) is a noise enhancement mask that is calculated when the speech and noise in the noise-mixed speech are known.
- the teacher input signal is an input signal for which an ideal speech emphasis mask and a noise emphasis mask are prepared in advance.
- the input signal may consist of "speech" which is a target signal and "noise" which is a non-target signal.
- the ideal speech emphasis mask may be called an ideal speech emphasis mask or an ideal value.
- the ideal noise emphasis mask may be called an ideal noise emphasis mask or an ideal value.
- Information including the teacher input signal and the ideal value may be called teacher information.
- the teacher information may be stored in the teacher information storage unit 222 described later.
- the ideal value may be a target value (target value) when updating the parameters.
- the arithmetic device 21 includes, for example, at least one of a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), and an FPGA (Field Programmable Gate Array).
- the arithmetic device 21 reads a computer program.
- the arithmetic device 21 may read a computer program stored in the storage device 22.
- the arithmetic device 21 may read a computer program stored in a computer-readable and non-transient recording medium using a recording medium reading device (e.g., an input device 24 described later) not shown in the figure that is provided in the information processing device 2.
- a recording medium reading device e.g., an input device 24 described later
- the arithmetic device 21 may acquire (i.e., download or read) a computer program from a device (not shown) located outside the information processing device 2 via the communication device 23 (or other communication device).
- the arithmetic device 21 executes the read computer program.
- a logical functional block for executing the operation to be performed by the information processing device 2 is realized within the calculation device 21.
- the calculation device 21 can function as a controller for realizing a logical functional block for executing the operation (in other words, processing) to be performed by the information processing device 2.
- a constraint loss calculation unit 211 which is a specific example of a "constraint loss calculation means” described in the appendix described later
- a parameter update unit 212 which is a specific example of a "parameter update means” described in the appendix described later
- a speech enhancement mask estimation unit 213 which is a specific example of a "speech enhancement mask estimation means” described in the appendix described later
- a noise enhancement mask estimation unit 214 which is a specific example of a "noise enhancement mask estimation means” described in the appendix described later
- a speech enhancement mask loss calculation unit 215 which is a specific example of a "speech enhancement mask loss calculation means” described in the appendix described later
- a noise enhancement mask loss calculation unit 216 which is a specific example of a "noise enhancement mask loss calculation means” described in the appendix described later
- the speech enhancement mask estimation unit 213 estimates the speech enhancement mask using a speech enhancement mask estimation model, which is a specific example of the "speech enhancement mask estimation model" described in the appendix described later.
- the noise enhancement mask estimation unit 214 estimates the noise enhancement mask using a noise enhancement mask estimation model, which is a specific example of the "noise enhancement mask estimation model” described in the appendix described later.
- any of the speech enhancement mask estimation unit 213, the noise enhancement mask estimation unit 214, the speech enhancement mask loss calculation unit 215, the noise enhancement mask loss calculation unit 216, the combined loss calculation unit 217, and the noise mixed voice input unit 218 may not be realized in the calculation device 21.
- the storage device 22 can store desired data.
- the storage device 22 may temporarily store a computer program executed by the arithmetic device 21.
- the storage device 22 may temporarily store data that the arithmetic device 21 temporarily uses when the arithmetic device 21 is executing a computer program.
- the storage device 22 may store data that the information processing device 2 stores for a long period of time.
- the storage device 22 may include at least one of a RAM (Random Access Memory), a ROM (Read Only Memory), a hard disk device, an optical magnetic disk device, an SSD (Solid State Drive), and a disk array device.
- the storage device 22 may include a non-temporary recording medium.
- the storage device 22 may realize a parameter storage unit 221 and a teacher information storage unit 222.
- the parameter storage unit 221 stores parameters included in the speech enhancement mask estimation model and parameters included in the noise enhancement mask estimation model. However, the storage device 22 does not need to realize either the parameter storage unit 221 or the teacher information storage unit 222.
- the communication device 23 is capable of communicating with devices external to the information processing device 2 via a communication network (not shown).
- the communication device 23 may be a communication interface based on standards such as Ethernet (registered trademark), Wi-Fi (registered trademark), Bluetooth (registered trademark), and USB (Universal Serial Bus).
- the input device 24 is a device that accepts information input to the information processing device 2 from outside the information processing device 2.
- the input device 24 may include an operating device (e.g., at least one of a keyboard, a mouse, and a touch panel) that can be operated by an operator of the information processing device 2.
- the input device 24 may include a reading device that can read information recorded as data on a recording medium that can be attached externally to the information processing device 2.
- the output device 25 is a device that outputs information to the outside of the information processing device 2.
- the output device 25 may output information as an image. That is, the output device 25 may include a display device (so-called a display) capable of displaying an image showing the information to be output.
- the output device 25 may output information as sound. That is, the output device 25 may include an audio device (so-called a speaker) capable of outputting sound.
- the output device 25 may output information on paper. That is, the output device 25 may include a printing device (so-called a printer) capable of printing desired information on paper. [2-3: Information Processing Operation Performed by Information Processing Device 2]
- FIG. 3 is a flowchart showing the flow of the information processing operation performed by the information processing device 2.
- the noise-mixed speech input unit 218 acquires the noise-mixed speech and inputs it to the speech enhancement mask estimation unit 213 and the noise enhancement mask estimation unit 214 (step S20).
- the speech enhancement mask estimation unit 213 estimates a speech enhancement mask using a speech enhancement mask estimation model (step S21).
- the speech enhancement mask estimation unit 213 outputs, as an estimated value, the estimated speech enhancement mask output by the speech enhancement mask estimation model to which the noise-mixed speech is input.
- the noise enhancement mask estimation unit 214 estimates the noise enhancement mask using the noise enhancement mask estimation model (step S22).
- the noise enhancement mask estimation unit 214 outputs, as an estimated value, the estimated noise enhancement mask output by the noise enhancement mask estimation model to which the noise-mixed speech is input. [2-3-1: Speech enhancement mask loss]
- the speech enhancement mask loss calculation unit 215 calculates a speech enhancement mask loss using the estimated speech enhancement mask estimated by the speech enhancement mask estimation model and an ideal speech enhancement mask (step S23).
- the speech enhancement mask loss indicates the difference between the estimated speech enhancement mask and the ideal speech enhancement mask.
- the speech enhancement mask loss may be calculated by a speech enhancement mask loss function L S.
- the speech enhancement mask loss function L S may be expressed as in the following formula 3. That is, the speech enhancement mask loss function L S may be expressed as the sum of squares of the difference between the estimated value and the ideal value in the frequency bin at each time.
- M S represents the estimated value of the speech enhancement mask. Also, M S i represents the ideal value of the speech enhancement mask. [Formula 3] In the following, matters in the frequency bin at each time may be expressed using the term "time-frequency". [2-3-2: Noise Enhancement Mask Loss]
- the noise enhancement mask loss calculation unit 216 calculates the noise enhancement mask loss using the estimated noise enhancement mask estimated by the noise enhancement mask estimation model and the ideal noise enhancement mask (step S24).
- the noise enhancement mask loss indicates the difference between the estimated noise enhancement mask and the ideal noise enhancement mask.
- the noise enhancement mask loss may be calculated by a noise enhancement mask loss function L N.
- the noise enhancement mask loss function L N may be expressed as in the following formula 4. That is, the noise enhancement mask loss function L N may be expressed as the sum of squares of the difference between the estimated value and the ideal value in time-frequency.
- M N represents the estimated value of the noise enhancement mask. Furthermore, M N i represents the ideal value of the noise enhancement mask. [Formula 4] [2-3-3: Restraint loss]
- the constraint loss calculation unit 211 calculates the constraint loss using the estimated speech enhancement mask estimated by the speech enhancement mask estimation model and the estimated noise enhancement mask estimated by the noise enhancement mask estimation model (step S25).
- the constraint loss indicates the difference between the estimated value of the sum (constraint condition) of the speech enhancement mask and the noise enhancement mask and the ideal value.
- the loss function is designed so that a penalty is imposed when the constraint condition is violated, that is, as the value obtained from the following formula 6 becomes farther from "0.”
- the constraint loss in this embodiment can be obtained from the following formula 6.
- a constraint condition is introduced as information related to both speech and noise.
- the constraint condition may be used as information related to both speech and noise when updating parameters, which will be described later.
- the parameters of the speech emphasis mask may be updated to make the speech emphasis stronger, and the parameters of the noise emphasis mask may also be updated to make the noise emphasis stronger.
- the parameters of the speech emphasis mask may be updated to make the speech emphasis weaker, and the parameters of the noise emphasis mask may also be updated to make the noise emphasis weaker.
- a loss function can be derived based on the amount by which the estimated speech enhancement mask and noise enhancement mask deviate from the constraint condition.
- a constraint loss function L SN shown in the following formula 7 may be introduced as the constraint condition.
- the following formula 7 calculates the sum of time-frequency.
- the combined loss calculation unit 217 calculates a combined loss by combining the speech enhancement mask loss, the noise enhancement mask loss, and the constraint loss (step S26).
- the combined loss may be calculated by a combined loss function L ALL .
- the combined loss function L ALL may be expressed as in the following formula 8.
- the combined loss calculated by the combined loss function L ALL may be called a total loss. [Formula 8]
- lambda may be a value between “0" and “1". lambda may be a positive number less than 1. lambda may be a value small compared to "1", such as "0.01" or "0.1". lambda may be constant or may vary during the model training process. lambda may be particularly small at the start of the model training process and large at the end.
- the speech enhancement mask loss function L S and the noise enhancement mask loss function L N are weighted equally in the summation, but different weights may be applied to the speech enhancement mask loss function L S and the noise enhancement mask loss function L N.
- the total loss may be calculated by the summation loss function L ALL shown in the following formula 9. [Formula 9]
- the parameter update unit 212 updates the parameters included in the speech enhancement mask estimation model and the parameters included in the noise enhancement mask estimation model according to the calculation result of the combined loss calculation unit 217 (step S27).
- the parameter update unit 212 updates the parameters included in the speech enhancement mask estimation model and the parameters included in the noise enhancement mask estimation model using the speech enhancement mask loss indicating the difference between the estimated speech enhancement mask and the ideal speech enhancement mask, the noise enhancement mask loss indicating the difference between the estimated noise enhancement mask and the ideal noise enhancement mask, and the constraint loss.
- the parameter update unit 212 updates the parameters included in the speech enhancement mask estimation model and the parameters included in the noise enhancement mask estimation model stored in the parameter storage unit 221.
- the parameter update unit 212 may update the parameters of the speech enhancement mask estimation model and the noise enhancement mask estimation model using, for example, a back-error propagation method.
- the information processing device 2 in this embodiment can proceed with learning of both the speech enhancement mask estimation model and the noise enhancement mask estimation model through the constrained loss function L SN.
- the speech enhancement mask estimation model can be learned using information obtained by combining information included in the speech enhancement mask estimation model and information included in the noise enhancement mask estimation model. [2-4: Tolerance of constraint conditions]
- the sum of the speech enhancement mask and the noise enhancement mask is ideally "1", and this embodiment uses this as a constraint.
- this constraint does not necessarily have to be “1” strictly.
- "1” may be replaced with “1+ ⁇ ” as shown in the following formula 10.
- " ⁇ " is a positive number that is very small compared to "1". When “1” is replaced with “1+ ⁇ ”, an error of about “ ⁇ ” can be tolerated. " ⁇ ” may be arbitrarily changed by design.
- S i and N i in the denominator of the above formula 10 represent ideal values, and S and N in the numerator represent estimated values.
- the parameter update unit 212 updates the parameters so that at least the sum of the speech emphasis mask and the noise emphasis mask is not less than "1".
- the audio enhancement mask estimation unit 213 may estimate audio enhancement masks corresponding to each of the multiple people.
- a audio enhancement mask estimation model corresponding to each person may be prepared, and each audio enhancement mask estimation model may output an audio enhancement mask corresponding to the person.
- the constraint condition for outputting the speech enhancement mask M SA and the speech enhancement mask M SB may be expressed as in the following formula 12. [Formula 12]
- the input signal contains speech and noise (as described above, in this embodiment, anything other than speech is called “noise”). Therefore, the sum of the ideal value of the speech enhancement mask and the ideal value of the noise enhancement mask is a constant value of "1". As described above, when the sum of the speech enhancement mask and the noise enhancement mask exceeds "1", at least one of the estimated value of the speech enhancement mask and the estimated value of the noise enhancement mask is larger than the ideal value, and there is a possibility that noise cannot be completely suppressed from the speech.
- the sum of the speech enhancement mask and the noise enhancement mask is less than "1"
- at least one of the estimated value of the speech enhancement mask and the estimated value of the noise enhancement mask is smaller than the ideal value, and there is a possibility that too much speech, which is a sound that should be left, is cut off.
- too much speech which is a sound that should be left
- the input signal contains noise that changes suddenly in a short time
- too much speech is often cut off.
- the estimated speech enhancement mask and the estimated noise enhancement mask are evaluated independently, and it is unclear whether the sum of the speech enhancement mask and the noise enhancement mask is a constant value of "1". Therefore, not only is it not possible to completely suppress noise from the speech, but there is a possibility that too much of the speech that should be preserved will be cut out.
- the accuracy of the noise enhancement mask is often lower than that of the speech enhancement mask, so if an independent evaluation of the noise enhancement mask is used to update the model parameters, the accuracy of the speech enhancement may deteriorate.
- the information processing device 2 in the second embodiment introduces a constraint loss calculated using the estimation result by the speech enhancement mask estimation unit 213 and the estimation result by the noise enhancement mask estimation unit 214. Then, the parameters included in the speech enhancement mask estimation model are updated using a combined loss obtained by combining the speech enhancement mask loss, the noise enhancement mask loss, and the constraint loss, so that it is possible to suppress noise from speech and prevent excessive cutting of speech, which is a sound that should be retained.
- the information processing device 2 can create a speech enhancement mask estimation model that can perform speech enhancement with high accuracy.
- FIG. 4 is a block diagram showing a configuration of an information processing device 3 in the third embodiment.
- the information processing device 3 in the third embodiment differs from the information processing device 2 in the second embodiment in the operation of a constraint loss calculation unit 311.
- the constraint loss calculation unit 311 calculates the constraint loss by multiplying the estimated speech emphasis mask and the estimated noise emphasis mask by the magnitude of the noise mixed speech.
- the constraint loss calculation unit 311 may calculate the constraint loss by multiplying the estimated speech emphasis mask and the estimated noise emphasis mask in time-frequency by the magnitude of the noise mixed speech.
- the magnitude of the noise mixed speech may be the absolute value of the amplitude of the noise mixed speech.
- the magnitude of the noise mixed speech may be the logarithmic power spectrum of the noise mixed speech.
- the magnitude of the noise mixed speech may be a normalized value. The value may be arbitrarily changed by design.
- Equation 13 the constraint loss function L SN in the third embodiment may be expressed as in Equation 13 below.
- Multiplying by the LPS input allows more emphasis to be placed on errors in time-frequency regions where the LPS input is large, which is equivalent to applying magnitude spectrum approximation (MSA).
- MSA magnitude spectrum approximation
- Multiplying by the LPS input is effective when learning noise that changes suddenly in a short time, such as the sound of a passing bullet train. It gives more weight to large input signals than to small input signals, making it possible to emphasize errors related to large input signals.
- the information processing device 3 in the third embodiment multiplies the magnitude of the noise-mixed voice, and therefore can suppress noise related to a particularly large input signal.
- FIG. 5 is a block diagram showing a configuration of the information processing device 4 according to the fourth embodiment.
- the information processing device 4 according to the fourth embodiment differs from the information processing device 2 according to the second embodiment and the information processing device 3 according to the third embodiment in the operation of a constraint loss calculation unit 411.
- the constraint loss calculation unit 411 calculates the constraint loss by raising the estimated speech enhancement mask and the estimated noise enhancement mask to a predetermined exponent.
- the constraint loss calculation unit 411 may also calculate the constraint loss by raising the estimated speech enhancement mask and the estimated noise enhancement mask in time-frequency to a predetermined exponent.
- the predetermined exponent may take a value between 0 and 1.
- the constraint loss function L SN in the fourth embodiment may be expressed as the following formula 14. [Formula 14]
- the LPS input applied in the third embodiment may be applied in the fourth embodiment as well.
- the constraint loss function L SN in the fourth embodiment may be expressed as in the following formula 15. [Formula 15] [4-2: Technical Effects of Information Processing Device 4]
- the information processing device 4 in the fourth embodiment can apply ⁇ to learn small signals with greater emphasis.
- the information processing device 5 in the fifth embodiment includes a calculation device 21 and a storage device 22, similar to the information processing device 2 in the second embodiment to the information processing device 4 in the fourth embodiment. Furthermore, the information processing device 5 in the fifth embodiment may include a communication device 23, an input device 24, and an output device 25, similar to the information processing device 2 in the second embodiment to the information processing device 4 in the fourth embodiment. However, the information processing device 5 may not include at least one of the communication device 23, the input device 24, and the output device 25. The information processing device 5 in the fifth embodiment differs from the information processing device 2 in the second embodiment to the information processing device 4 in the fourth embodiment in that an acoustic feature extraction unit 519 is further realized in the calculation device 21.
- FIG. 7 is a flowchart showing the flow of the information processing operation performed by the information processing device 5.
- the noise mixed speech input unit 518 acquires the noise mixed speech and inputs it to the acoustic feature extraction unit 519 (step S50).
- the acoustic feature extraction unit 519 extracts noise mixed speech features, which are characteristics of the noise mixed speech, from the noise mixed speech (step S51).
- the acoustic feature extraction unit 519 may have an acoustic feature extraction model that receives the noise mixed speech as input and outputs the noise mixed speech features.
- the acoustic feature extraction model may be implemented using an RNN.
- the speech enhancement mask estimation unit 513 estimates the speech enhancement mask using the speech enhancement mask estimation model (step S52).
- the speech enhancement mask estimation model may receive noise-mixed speech features and output an estimated speech enhancement mask.
- the speech enhancement mask estimation unit 513 may output the estimated speech enhancement mask output by the speech enhancement mask estimation model to which the noise-mixed speech features are input as an estimated value.
- the noise enhancement mask estimation unit 514 estimates the noise enhancement mask using the noise enhancement mask estimation model (step S53).
- the noise enhancement mask estimation model may receive the noise-mixed speech features and output an estimated noise enhancement mask.
- the noise enhancement mask estimation unit 514 may output the estimated noise enhancement mask output by the noise enhancement mask estimation model to which the noise-mixed speech features are input as an estimated value.
- the speech enhancement mask loss calculation unit 515 calculates the speech enhancement mask loss using the estimated speech enhancement mask estimated by the speech enhancement mask estimation model and the ideal speech enhancement mask (step S54).
- the ideal speech enhancement mask may be a speech enhancement mask that is ideal for the noise-mixed speech features extracted from the teacher input signal.
- the noise enhancement mask loss calculation unit 516 calculates the noise enhancement mask loss using the estimated noise enhancement mask estimated by the noise enhancement mask estimation model and the ideal noise enhancement mask (step S55).
- the ideal noise enhancement mask may be an ideal noise enhancement mask for the noise-mixed speech features extracted from the teacher input signal.
- the constraint loss calculation unit 211 calculates the constraint loss using the estimated speech enhancement mask estimated by the speech enhancement mask estimation model and the estimated noise enhancement mask estimated by the noise enhancement mask estimation model (step S25).
- the combined loss calculation unit 217 calculates the combined loss by combining the speech enhancement mask loss, the noise enhancement mask loss, and the constraint loss (step S26).
- the parameter update unit 512 updates the parameters included in the acoustic feature extraction model, as well as the parameters included in the speech enhancement mask estimation model and the parameters included in the noise enhancement mask estimation model, according to the calculation result of the combined loss calculation unit 217 (step S56).
- the acoustic feature extraction model may be trained using both information about speech and information about noise.
- the information processing device 5 in the fifth embodiment creates a speech enhancement mask estimation model and an acoustic feature extraction model by machine learning.
- the acoustic feature extraction model is applied.
- the information processing device 5 in the fifth embodiment is provided with an acoustic feature extraction unit 519, which allows the input signals to both the speech enhancement mask estimation model and the noise enhancement mask estimation model to be increased, thereby enabling more preferable speech enhancement.
- the acoustic feature extraction unit 519 is not provided, as in the second embodiment, the operation can be made lighter.
- the information processing device 6 in the sixth embodiment includes a calculation device 21 and a storage device 22, similar to the information processing device 2 in the second embodiment to the information processing device 5 in the fifth embodiment. Furthermore, the information processing device 6 in the sixth embodiment may include a communication device 23, an input device 24, and an output device 25, similar to the information processing device 2 in the second embodiment to the information processing device 5 in the fifth embodiment. However, the information processing device 6 may not include at least one of the communication device 23, the input device 24, and the output device 25.
- the information processing device 6 in the sixth embodiment differs from the information processing device 2 in the second embodiment to the information processing device 5 in the fifth embodiment in that the constraint loss calculation unit 611 has a noise constraint loss calculation unit 6111 and a speech constraint loss calculation unit 6112.
- FIG. 9 is a flowchart showing the flow of the information processing operation performed by the information processing device 6.
- the noise-mixed speech input unit 218 acquires the noise-mixed speech and inputs it to the speech enhancement mask estimation unit 213 and the noise enhancement mask estimation unit 214 (step S20).
- the speech enhancement mask estimation unit 213 estimates a speech enhancement mask using a speech enhancement mask estimation model (step S21).
- the noise enhancement mask estimation unit 214 estimates a noise enhancement mask using the noise enhancement mask estimation model (step S22).
- the speech enhancement mask loss calculation unit 215 calculates a speech enhancement mask loss using the estimated speech enhancement mask estimated by the speech enhancement mask estimation model and the ideal speech enhancement mask (step S23).
- the noise enhancement mask loss calculation unit 216 calculates a noise enhancement mask loss using the estimated noise enhancement mask estimated by the noise enhancement mask estimation model and the ideal noise enhancement mask (step S24). [6-2-1: Noise-constrained loss]
- the noise constrained loss calculation unit 6111 calculates the noise constrained loss using the estimated speech emphasis mask and the ideal noise emphasis mask (step S60).
- the noise constrained loss calculation unit 6111 calculates the constrained noise emphasis mask based on the estimated speech emphasis mask and the constraint condition expressed by the above formula 5.
- the constrained noise emphasis mask is expressed as M N *
- the constrained noise emphasis mask may be expressed as the following formula 16. [Formula 16]
- the noise-constrained loss calculation unit 6111 calculates a noise-constrained loss indicating a difference between the constrained noise enhancement mask and the ideal noise enhancement mask.
- the noise-constrained loss is calculated by a noise-constrained loss function L N *
- the noise-constrained loss function L N * may be expressed as in the following Equation 17. [Formula 17] [6-2-2: Audio Restriction Loss]
- the speech constraint loss calculation unit 6112 calculates the speech constraint loss using the estimated noise emphasis mask and the ideal noise emphasis mask (step S61).
- the speech constraint loss calculation unit 6112 calculates the constraint speech emphasis mask based on the estimated noise emphasis mask and the constraint condition expressed by the above formula 5.
- the constraint speech emphasis mask is expressed as M S *
- the constraint speech emphasis mask may be expressed as the following formula 18. [Formula 18]
- the speech constraint loss calculation unit 6112 calculates a speech constraint loss indicating a difference between the constrained speech enhancement mask and the ideal speech enhancement mask.
- the speech constraint loss is calculated by a speech constraint loss function L S *
- the speech constraint loss function L S * may be expressed as the following formula 19. [Formula 19] [6-2-3: Combined Losses]
- the combined loss calculation unit 617 calculates a combined loss by combining the speech enhancement mask loss, the noise enhancement mask loss, the noise constrained loss, and the speech constrained loss (step S62).
- the combined loss is calculated by a combined loss function L ALL
- the combined loss function L ALL may be expressed as the following formula 20.
- M can be considered as an estimate of a mask that also includes a constraint.
- the information processing device 6 in the sixth embodiment can obtain an effect equivalent to that of the information processing device 2 in the second embodiment.
- the above formula 20 may be rewritten as the following formula 21.
- the combined loss in the sixth embodiment can be considered to indicate the difference between the estimated value of the mask including the constraint conditions and the ideal value.
- the parameter update unit 212 updates the parameters included in the speech enhancement mask estimation model and the parameters included in the noise enhancement mask estimation model according to the calculation result of the combined loss calculation unit 217 (step S27).
- the information processing device 6 in the sixth embodiment does not employ " ⁇ " as employed in the second embodiment. Therefore, the information processing device 6 in the sixth embodiment can reduce the processing load required for determining " ⁇ " more than the information processing device 2 in the second embodiment.
- [Appendix 1] a constraint loss calculation means for calculating a constraint loss using an estimated speech enhancement mask output from a speech enhancement mask estimation model to which a noise-mixed speech is input, and an estimated noise enhancement mask output from a noise enhancement mask estimation model to which the noise-mixed speech is input; and a parameter updating means for updating parameters included in the speech enhancement mask estimation model and parameters included in the noise enhancement mask estimation model using a speech enhancement mask loss indicating a difference between the estimated speech enhancement mask and a target speech enhancement mask, a noise enhancement mask loss indicating a difference between the estimated noise enhancement mask and a target noise enhancement mask, and the constraint loss.
- the parameter update means updates the parameters included in the speech enhancement mask estimation model and the parameters included in the noise enhancement mask estimation model in accordance with a calculation result of the combined loss calculation means.
- the method further comprises: extracting an acoustic feature from the noise-mixed speech; the speech enhancement mask estimation model receives the noise-mixed speech features and outputs the estimated speech enhancement mask; The information processing apparatus according to claim 1 , wherein the noise enhancement mask estimation model receives the noise-mixed speech feature and outputs the estimated noise enhancement mask.
- the constraint loss calculation means a speech constraint loss calculation means for calculating a speech constraint loss using the estimated noise emphasis mask and the target speech emphasis mask; a noise constrained loss calculation means for calculating a noise constrained loss using the estimated speech emphasis mask and the target noise emphasis mask; The information processing device according to claim 1 , wherein the constraint loss includes the speech constraint loss and the noise constraint loss.
- [Appendix 7] Calculating a constraint loss using an estimated speech enhancement mask output from a speech enhancement mask estimation model to which the noise-mixed speech is input, and an estimated noise enhancement mask output from a noise enhancement mask estimation model to which the noise-mixed speech is input; an information processing method for updating parameters included in the speech enhancement mask estimation model and parameters included in the noise enhancement mask estimation model using a speech enhancement mask loss indicating a difference between the estimated speech enhancement mask and a target speech enhancement mask, a noise enhancement mask loss indicating a difference between the estimated noise enhancement mask and a target noise enhancement mask, and the constraint loss.
- At least some of the components of each of the above-described embodiments can be appropriately combined with at least some of the other components of each of the above-described embodiments. Some of the components of each of the above-described embodiments may not be used.
Landscapes
- Engineering & Computer Science (AREA)
- Computational Linguistics (AREA)
- Quality & Reliability (AREA)
- Signal Processing (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Human Computer Interaction (AREA)
- Physics & Mathematics (AREA)
- Acoustics & Sound (AREA)
- Multimedia (AREA)
- Circuit For Audible Band Transducer (AREA)
Abstract
Description
[1:第1実施形態]
[1-1:情報処理装置1の構成]
[1-2:情報処理装置1の技術的効果]
[2:第2実施形態]
[2-1-1:音声と雑音]
[2-1-2:音声強調技術]
[2-1-3:音声強調マスク、及び雑音強調マスク]
[式1]
[式2]
[2-1-4:学習]
[2-2:情報処理装置2の構成]
[2-3:情報処理装置2が行う情報処理動作]
[2-3-1:音声強調マスクロス]
[式3]
なお、以下、各時刻における周波数瓶での事項を、「時間-周波数」という文言を用いて表現する場合がある。
[2-3-2:雑音強調マスクロス]
[式4]
[2-3-3:拘束ロス]
[式5]
[式7]
[2-3-4:合算ロス]
[式8]
[式9]
[2-3-5:パラメータ更新]
[2-4:拘束条件の許容範囲]
[式10]
「ε」は「1」と比較して非常に小さな正数である。「1」を「1+ε」に置き換えた場合、「ε」程度の誤差を許容することができる。「ε」は、設計により任意に変更することができてもよい。上記式10の分母のSi、及びNiは、理想値を表現しており、分子のS、及びNは、推定値を表現している。
[2-5:変形例]
[2-6:情報処理装置2の技術的効果]
[3:第3実施形態]
[3-1:情報処理装置3が行う情報処理動作]
[3-2:情報処理装置3の技術的効果]
[4:第4実施形態]
[4-1:情報処理装置4が行う情報処理動作]
[式15]
[4-2:情報処理装置4の技術的効果]
[5:第5実施形態]
[5-1:情報処理装置5の構成]
[5-2:情報処理装置5が行う情報処理動作]
[5-3:情報処理装置5の技術的効果]
[6:第6実施形態]
[6-1:情報処理装置6の構成]
[6-2:情報処理装置6が行う情報処理動作]
[6-2-1:雑音拘束ロス]
[式16]
[式17]
[6-2-2:音声拘束ロス]
[式18]
[式19]
[6-2-3:合算ロス]
[式20]
[6-3:情報処理装置6の技術的効果]
[7:付記]
[付記1]
雑音混合音声が入力された音声強調マスク推定モデルが出力した推定音声強調マスク、及び前記雑音混合音声が入力された雑音強調マスク推定モデルが出力した推定雑音強調マスクを用いて拘束ロスを算出する拘束ロス算出手段と、
前記推定音声強調マスクと目標音声強調マスクとの違いを示す音声強調マスクロス、前記推定雑音強調マスクと目標雑音強調マスクとの違いを示す雑音強調マスクロス、及び前記拘束ロスを用いて、前記音声強調マスク推定モデルが含むパラメータ、及び前記雑音強調マスク推定モデルが含むパラメータを更新するパラメータ更新手段と
を備える情報処理装置。
[付記2]
前記音声強調マスク推定モデルを用いて前記推定音声強調マスクを推定する音声強調マスク推定手段と、
前記雑音強調マスク推定モデルを用いて前記推定雑音強調マスクを推定する雑音強調マスク推定手段と、
前記推定音声強調マスクと前記目標音声強調マスクとを用いて、前記音声強調マスクロスを算出する音声強調マスクロス算出手段と、
前記推定雑音強調マスクと前記目標雑音強調マスクとを用いて、前記雑音強調マスクロスを算出する雑音強調マスクロス算出手段と、
前記音声強調マスクロス、前記雑音強調マスクロス、及び前記拘束ロスを合算した合算ロスを算出する合算ロス算出手段と
を更に備え、
前記パラメータ更新手段は、前記合算ロス算出手段の算出結果に応じて、前記音声強調マスク推定モデルが含むパラメータと前記雑音強調マスク推定モデルが含むパラメータとを更新する
請求項1に記載の情報処理装置。
[付記3]
前記拘束ロス算出手段は、前記推定音声強調マスク、及び前記推定雑音強調マスクに前記雑音混合音声の大きさを乗じて、前記拘束ロスを算出する
請求項1又は2に記載の情報処理装置。
[付記4]
前記拘束ロス算出手段は、前記推定音声強調マスク、及び前記推定雑音強調マスクを所定の指数で冪乗して、前記拘束ロスを算出する
請求項1又は2に記載の情報処理装置。
[付記5]
前記雑音混合音声から雑音混合音声特徴を抽出する音響特徴抽出手段を更に備え、
前記音声強調マスク推定モデルは、前記雑音混合音声特徴が入力され、前記推定音声強調マスクを出力し、
前記雑音強調マスク推定モデルは、前記雑音混合音声特徴が入力され、前記推定雑音強調マスクを出力する
請求項1又は2に記載の情報処理装置。
[付記6]
前記拘束ロス算出手段は、
前記推定雑音強調マスク、及び前記目標音声強調マスクを用いて音声拘束ロスを算出する音声拘束ロス算出手段と、
前記推定音声強調マスク、及び前記目標雑音強調マスクを用いて雑音拘束ロスを算出する雑音拘束ロス算出手段とを有し、
前記拘束ロスは、前記音声拘束ロス、及び前記雑音拘束ロスを含む
請求項1又は2に記載の情報処理装置。
[付記7]
雑音混合音声が入力された音声強調マスク推定モデルが出力した推定音声強調マスク、及び前記雑音混合音声が入力された雑音強調マスク推定モデルが出力した推定雑音強調マスクを用いて拘束ロスを算出し、
前記推定音声強調マスクと目標音声強調マスクとの違いを示す音声強調マスクロス、前記推定雑音強調マスクと目標雑音強調マスクとの違いを示す雑音強調マスクロス、及び前記拘束ロスを用いて、前記音声強調マスク推定モデルが含むパラメータ、及び前記雑音強調マスク推定モデルが含むパラメータを更新する
情報処理方法。
[付記8]
コンピュータに、
雑音混合音声が入力された音声強調マスク推定モデルが出力した推定音声強調マスク、及び前記雑音混合音声が入力された雑音強調マスク推定モデルが出力した推定雑音強調マスクを用いて拘束ロスを算出し、
前記推定音声強調マスクと目標音声強調マスクとの違いを示す音声強調マスクロス、前記推定雑音強調マスクと目標雑音強調マスクとの違いを示す雑音強調マスクロス、及び前記拘束ロスを用いて、前記音声強調マスク推定モデルが含むパラメータ、及び前記雑音強調マスク推定モデルが含むパラメータを更新する
情報処理方法を実行させるためのコンピュータプログラムが記録された記録媒体。
11,211,311,411,611 拘束ロス算出部
12,212,512 パラメータ更新部
221 パラメータ記憶部
222,522 教師情報記憶部
213,513 音声強調マスク推定部
214,514 雑音強調マスク推定部
215,515 音声強調マスクロス算出部
216,516 雑音強調マスクロス算出部
217,617 合算ロス算出部
218,518 雑音混合音声入力部
519 音響特徴抽出部
6111 雑音拘束ロス算出部
6112 音声拘束ロス算出部
Claims (8)
- 雑音混合音声が入力された音声強調マスク推定モデルが出力した推定音声強調マスク、及び前記雑音混合音声が入力された雑音強調マスク推定モデルが出力した推定雑音強調マスクを用いて拘束ロスを算出する拘束ロス算出手段と、
前記推定音声強調マスクと雑音混合音声中の音声と雑音が既知な場合に計算される目標音声強調マスクとの違いを示す音声強調マスクロス、前記推定雑音強調マスクと目標雑音強調マスクとの違いを示す雑音強調マスクロス、及び前記拘束ロスを用いて、前記音声強調マスク推定モデルが含むパラメータ、及び前記雑音強調マスク推定モデルが含むパラメータを更新するパラメータ更新手段と
を備える情報処理装置。 - 前記音声強調マスク推定モデルを用いて前記推定音声強調マスクを推定する音声強調マスク推定手段と、
前記雑音強調マスク推定モデルを用いて前記推定雑音強調マスクを推定する雑音強調マスク推定手段と、
前記推定音声強調マスクと前記目標音声強調マスクとを用いて、前記音声強調マスクロスを算出する音声強調マスクロス算出手段と、
前記推定雑音強調マスクと前記目標雑音強調マスクとを用いて、前記雑音強調マスクロスを算出する雑音強調マスクロス算出手段と、
前記音声強調マスクロス、前記雑音強調マスクロス、及び前記拘束ロスを合算した合算ロスを算出する合算ロス算出手段と
を更に備え、
前記パラメータ更新手段は、前記合算ロス算出手段の算出結果に応じて、前記音声強調マスク推定モデルが含むパラメータと前記雑音強調マスク推定モデルが含むパラメータとを更新する
請求項1に記載の情報処理装置。 - 前記拘束ロス算出手段は、前記推定音声強調マスク、及び前記推定雑音強調マスクに前記雑音混合音声の大きさを乗じて、前記拘束ロスを算出する
請求項1又は2に記載の情報処理装置。 - 前記拘束ロス算出手段は、前記推定音声強調マスク、及び前記推定雑音強調マスクを所定の指数で冪乗して、前記拘束ロスを算出する
請求項1又は2に記載の情報処理装置。 - 前記雑音混合音声から雑音混合音声特徴を抽出する音響特徴抽出手段を更に備え、
前記音声強調マスク推定モデルは、前記雑音混合音声特徴が入力され、前記推定音声強調マスクを出力し、
前記雑音強調マスク推定モデルは、前記雑音混合音声特徴が入力され、前記推定雑音強調マスクを出力する
請求項1又は2に記載の情報処理装置。 - 前記拘束ロス算出手段は、
前記推定雑音強調マスク、及び前記目標音声強調マスクを用いて音声拘束ロスを算出する音声拘束ロス算出手段と、
前記推定音声強調マスク、及び前記目標雑音強調マスクを用いて雑音拘束ロスを算出する雑音拘束ロス算出手段とを有し、
前記拘束ロスは、前記音声拘束ロス、及び前記雑音拘束ロスを含む
請求項1又は2に記載の情報処理装置。 - 雑音混合音声が入力された音声強調マスク推定モデルが出力した推定音声強調マスク、及び前記雑音混合音声が入力された雑音強調マスク推定モデルが出力した推定雑音強調マスクを用いて拘束ロスを算出し、
前記推定音声強調マスクと雑音混合音声中の音声と雑音が既知な場合に計算される目標音声強調マスクとの違いを示す音声強調マスクロス、前記推定雑音強調マスクと目標雑音強調マスクとの違いを示す雑音強調マスクロス、及び前記拘束ロスを用いて、前記音声強調マスク推定モデルが含むパラメータ、及び前記雑音強調マスク推定モデルが含むパラメータを更新する
情報処理方法。 - コンピュータに、
雑音混合音声が入力された音声強調マスク推定モデルが出力した推定音声強調マスク、及び前記雑音混合音声が入力された雑音強調マスク推定モデルが出力した推定雑音強調マスクを用いて拘束ロスを算出し、
前記推定音声強調マスクと雑音混合音声中の音声と雑音が既知な場合に計算される目標音声強調マスクとの違いを示す音声強調マスクロス、前記推定雑音強調マスクと目標雑音強調マスクとの違いを示す雑音強調マスクロス、及び前記拘束ロスを用いて、前記音声強調マスク推定モデルが含むパラメータ、及び前記雑音強調マスク推定モデルが含むパラメータを更新する
情報処理方法を実行させるためのコンピュータプログラムが記録された記録媒体。
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP2025506508A JPWO2024190063A5 (ja) | 2024-01-16 | 情報処理装置、情報処理方法、及び、コンピュータプログラム |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP2023039058 | 2023-03-13 | ||
| JP2023-039058 | 2023-03-13 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2024190063A1 true WO2024190063A1 (ja) | 2024-09-19 |
Family
ID=92754770
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/JP2024/000990 Ceased WO2024190063A1 (ja) | 2023-03-13 | 2024-01-16 | 情報処理装置、情報処理方法、及び、記録媒体 |
Country Status (1)
| Country | Link |
|---|---|
| WO (1) | WO2024190063A1 (ja) |
-
2024
- 2024-01-16 WO PCT/JP2024/000990 patent/WO2024190063A1/ja not_active Ceased
Non-Patent Citations (2)
| Title |
|---|
| ERDOGAN HAKAN, YOSHIOKA TAKUYA: "Investigations on Data Augmentation and Loss Functions for Deep Learning Based Speech-Background Separation", INTERSPEECH 2018, ISCA, 1 January 2018 (2018-01-01), pages 3499 - 3503, XP093210252, DOI: 10.21437/Interspeech.2018-2441 * |
| TSUNODA, RYOTA ET AL.: "Pre-training method for audio-visual speech recognition in interference speaker environment", PROCEEDINGS OF THE 2022 SPRING MEETING THE ACOUSTICAL SOCIETY OF JAPAN; MARCH 9 - 11, 2022, vol. 2022, 23 February 2022 (2022-02-23) - 11 March 2022 (2022-03-11), pages 1057 - 1060, XP009558531 * |
Also Published As
| Publication number | Publication date |
|---|---|
| JPWO2024190063A1 (ja) | 2024-09-19 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| Fu et al. | Metricgan+: An improved version of metricgan for speech enhancement | |
| JP5666444B2 (ja) | 特徴抽出を使用してスピーチ強調のためにオーディオ信号を処理する装置及び方法 | |
| JP5842056B2 (ja) | 雑音推定装置、雑音推定方法、雑音推定プログラム及び記録媒体 | |
| US12597434B2 (en) | Control of speech preservation in speech enhancement | |
| Abdullah et al. | Towards more efficient DNN-based speech enhancement using quantized correlation mask | |
| JPWO2020039571A1 (ja) | 音声分離装置、音声分離方法、音声分離プログラム、及び音声分離システム | |
| JP7667247B2 (ja) | 機械学習を用いたノイズ削減 | |
| Xu et al. | Deep noise suppression maximizing non-differentiable PESQ mediated by a non-intrusive PESQNet | |
| CN113990343B (zh) | 语音降噪模型的训练方法和装置及语音降噪方法和装置 | |
| KR20200092501A (ko) | 합성 음성 신호 생성 방법, 뉴럴 보코더 및 뉴럴 보코더의 훈련 방법 | |
| CN113707167A (zh) | 残留回声抑制模型的训练方法和训练装置 | |
| US12579961B2 (en) | Music enhancement systems | |
| EP1995723A1 (en) | Neuroevolution training system | |
| KR102198597B1 (ko) | 뉴럴 보코더 및 화자 적응형 모델을 구현하기 위한 뉴럴 보코더의 훈련 방법 | |
| JP5994639B2 (ja) | 有音区間検出装置、有音区間検出方法、及び有音区間検出プログラム | |
| Kumar et al. | Comparative studies of single-channel speech enhancement techniques | |
| Richter et al. | Speech signal improvement using causal generative diffusion models | |
| Karthik et al. | An optimized convolutional neural network for speech enhancement | |
| Elshamy et al. | DNN-based cepstral excitation manipulation for speech enhancement | |
| CN113393852B (zh) | 语音增强模型的构建方法及系统、语音增强方法及系统 | |
| KR102505653B1 (ko) | 심화신경망을 이용한 에코 및 잡음 통합 제거 방법 및 장치 | |
| WO2024190063A1 (ja) | 情報処理装置、情報処理方法、及び、記録媒体 | |
| Biswas et al. | Optimal near-end speech intelligibility improvement using CLPSO-based voice transformation in realistic noisy environments | |
| CN114141266A (zh) | 基于pesq驱动的强化学习估计先验信噪比的语音增强方法 | |
| Vinothkumar | Speech enhancement with background noise suppression in various data corpus using Bi-LSTM algorithm |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 24770185 Country of ref document: EP Kind code of ref document: A1 |
|
| ENP | Entry into the national phase |
Ref document number: 2025506508 Country of ref document: JP Kind code of ref document: A |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 2025506508 Country of ref document: JP |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 24770185 Country of ref document: EP Kind code of ref document: A1 |
