WO2004107319A1 - 既知音響信号除去方法及び装置 - Google Patents
既知音響信号除去方法及び装置 Download PDFInfo
- Publication number
- WO2004107319A1 WO2004107319A1 PCT/JP2004/007587 JP2004007587W WO2004107319A1 WO 2004107319 A1 WO2004107319 A1 WO 2004107319A1 JP 2004007587 W JP2004007587 W JP 2004007587W WO 2004107319 A1 WO2004107319 A1 WO 2004107319A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- acoustic signal
- mixed
- signal
- amplitude spectrum
- sound
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L21/00—Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
- G10L21/02—Speech enhancement, e.g. noise reduction or echo cancellation
- G10L21/0272—Voice signal separating
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L21/00—Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
- G10L21/02—Speech enhancement, e.g. noise reduction or echo cancellation
Definitions
- the present invention relates to a known sound signal elimination method and a known sound signal elimination device for removing a component of a known sound signal from a mixed sound signal in which a plurality of sound signals are mixed.
- Non-Patent Document 1 a method called a spectral subtraction method (Non-Patent Document 1) has been known as an acoustic signal processing.
- the conventional spectral subtraction method uses a sound signal (mixed sound) that is a mixture of stationary noise (noise whose spectrum does not change over time and frequency characteristics and volume are almost constant) and a desired sound (target sound). ) To obtain a target sound by removing stationary noise.
- the spectrum of the stationary noise is learned in advance by a simple method such as calculating the average of the stationary spectrum, and the spectrum of the stationary noise is calculated from the spectrum of the input mixed sound. Perform the removal process. That is, the process of subtracting the average of the noise is performed.
- Patent Document 2 Japanese Patent Application Laid-Open No. 2000-201
- Patent Document 3 Japanese Patent Application Laid-Open Publication No. 2000-1922
- Patent Document 4 Japanese Patent Application Laid-Open No. 2000-201
- Patent Document 5 Japanese Patent Application Laid-Open No. Hei 11-0-309
- Patent Document 6 Japanese Patent Application Laid-Open No. H10-24029
- Patent Document 7 Japanese Patent Application Laid-Open No. 08-2221092 Disclosure of the Invention
- the conventional spectral subtraction method presupposes stationary noise and cannot be applied to non-stationary noise (noise whose spectrum changes significantly over time and whose frequency characteristics and volume also change). For example, it was not possible to remove time-varying non-stationary noise, such as music used as background music (BGM). This is because the spectrum of the non-stationary noise changes too much for learning.
- non-stationary noise such as music used as background music (BGM).
- an object of the present invention is to convert a component of a known sound signal (which may be non-stationary or stationary) from a mixed sound signal in which a plurality of sound signals are mixed into a known sound signal from an original sound source corresponding thereto.
- a known acoustic signal that can be removed using the signal It is intended to provide a removing method, a known acoustic signal removing device, and a program used for the device.
- Another object of the present invention is to provide, for example, a method in which a known sound signal is music, and the music sound signal is obtained from a mixed sound used as background music (BGM) for human voices and body sounds.
- BGM background music
- a known sound signal elimination method and a known sound signal that can remove background music using a known sound signal for example, a sound signal of the same music separately obtained from a CD, a record, or the like
- a known sound signal for example, a sound signal of the same music separately obtained from a CD, a record, or the like
- Still another object of the present invention is to remove a component of a known acoustic signal from an acoustic signal (mixed sound) in which a plurality of acoustic signals are mixed, and to accurately detect a known acoustic signal in the mixed sound. It is an object of the present invention to provide a known sound signal removing method and device capable of automatically estimating a position and removing a known sound signal at the position, and a program used for the device.
- Still another object of the present invention is to remove a component of a known acoustic signal from an acoustic signal (mixed sound) in which a plurality of acoustic signals are mixed, and to accurately detect a known acoustic signal in the mixed sound. It is an object of the present invention to provide a known acoustic signal elimination device provided with an interface in which a position can be designated by a human.
- Still another object of the present invention is to remove a component of a known sound signal from a sound signal (mixed sound) in which a plurality of sound signals are mixed.
- Another object of the present invention is to provide a known acoustic signal eliminator provided with an interface that allows a person to specify the expansion and contraction when the expansion and contraction is performed in the frequency axis direction.
- Still another object of the present invention is to remove a plurality of known acoustic signals from acoustic signals mixed with a plurality of acoustic signals, and remove the known acoustic signals one by one.
- An object of the present invention is to provide a method and an apparatus for removing a known acoustic signal and a program used for the apparatus.
- a component of a known acoustic signal (which may be non-stationary or stationary) is mixed with a known acoustic signal from an original sound source from a mixed acoustic signal in which a plurality of acoustic signals are mixed. Remove with.
- the mixed acoustic signal is converted into a time-frequency expression to determine the amplitude spectrum of the mixed acoustic signal and the phase of the mixed acoustic signal (the mixed acoustic signal conversion step).
- a known conversion method such as Fourier transform or wave-rate transform is used.
- a known sound signal (a sound signal of the same music separately obtained from a CD or record) corresponding to (similar to) the known sound signal included in the mixed sound signal is converted into a time-frequency expression. Conversion is performed to obtain the amplitude spectrum of the known sound signal (known sound signal conversion step).
- the time position shift of the amplitude spectrum of the known sound signal with respect to the amplitude spectrum of the mixed sound signal, the time change of the frequency characteristic, the time of the sound volume A corrected amplitude spectrum of the known acoustic signal is obtained by correcting at least one of the change, the expansion and contraction in the time axis direction, and the expansion and contraction in the frequency axis direction (correction step).
- the correction amplitude spectrum of the known sound signal is removed from the amplitude spectrum of the mixed sound signal (removal step). Based on the amplitude spectrum after removal obtained in this removal step and the phase of the mixed acoustic signal, inverse conversion is performed on the time expression to obtain a unit waveform (inverse conversion step).
- the unit waveform is synthesized by using a synthesis method such as an overlapping quadrature method to obtain an audio signal from which components of the known audio signal have been removed (synthesis step).
- the following correction step is executed to shift the temporal position of the torsion width spectrum of the known acoustic signal with respect to the amplitude spectrum of the mixed acoustic signal.
- a corrected amplitude spectrum of the known sound signal in which at least one of the time change of the frequency characteristic, the time change of the sound volume, the expansion and contraction in the time axis direction and the expansion and contraction in the frequency axis direction is corrected is obtained, and the corrected amplitude spectrum is calculated. It is removed from the amplitude spectrum of the mixed sound signal. For this reason, the known acoustic signal included as irregular noise in the mixed acoustic signal can be removed with high accuracy.
- the time shift of the amplitude spectrum of the known acoustic signal the time change of the frequency characteristic, the time change of the volume, the expansion and contraction in the time axis direction and the frequency axis direction It is preferable to correct all occurrences of the phenomenon or change in the mixed acoustic signal.
- the accuracy of removal of the known sound signal can be improved rather than the case where no correction is made. You don't have to do everything. Of course, all necessary corrections may be made.
- the temporal position of the known acoustic signal included in the mixed acoustic signal is estimated, and the temporal position of the amplitude spectrum of the known acoustic signal is shifted based on the estimated temporal position. to correct.
- the estimation method is, for example, obtaining a distance (similarity) between a predetermined section of the amplitude spectrum of the mixed sound signal and a predetermined section of the amplitude spectrum of the known sound signal, and determining a section having the closest distance to the mixed sound signal. It is estimated as the temporal position of the known acoustic signal included in.
- a change in the frequency characteristic of the known acoustic signal included in the mixed acoustic signal is estimated, and the frequency characteristic of the amplitude spectrum of the known acoustic signal is estimated based on the time change of the estimated frequency characteristic. Is corrected over time.
- the estimation of the change in the frequency characteristic is performed, for example, by identifying a section of the mixed acoustic signal that includes only the known acoustic signal, and comparing the frequency characteristic of this section with the frequency characteristic of the known acoustic signal corresponding to this section. From the contrast, the change of the frequency characteristic of the known acoustic signal included in the mixed acoustic signal is estimated.
- a temporal change in the volume of the known acoustic signal included in the mixed acoustic signal is estimated, and a temporal change in the volume of the amplitude spectrum of the known acoustic signal is determined based on the estimated temporal change in the volume.
- Is corrected After estimating the time change of the sound volume, after correcting the frequency characteristics, for example, a frequency band having an amplitude corresponding to a known sound signal included in the mixed sound signal is specified at each time, and the mixing in the frequency band is determined. It is estimated from the contrast between the amplitude of the acoustic signal and the amplitude of the known acoustic signal.
- the expansion and contraction of the known acoustic signal included in the mixed acoustic signal in the time axis direction is estimated, and the time of the amplitude spectrum of the known acoustic signal is determined based on the estimated expansion and contraction in the time axis direction.
- Correct axial expansion and contraction for example, a section containing only a known acoustic signal in the mixed acoustic signal is specified, and a time axis comparison with a section of the known acoustic signal corresponding to this section is performed. Estimate the expansion and contraction in the time axis direction. Or divided the time axis into short sections Estimate by comparing all sections.
- the expansion and contraction of the known acoustic signal included in the mixed acoustic signal in the frequency axis direction is estimated, and the frequency of the amplitude spectrum of the known acoustic signal is determined based on the estimated expansion and contraction in the frequency axis direction.
- Correct axial expansion and contraction for example, a section containing only a known acoustic signal in the mixed acoustic signal is specified, and the section of the frequency axis with the section of the known acoustic signal corresponding to this section is determined by: Estimate expansion and contraction in the frequency axis direction.
- an image display step of displaying an image so that the amplitude spectrum of the mixed acoustic signal and the amplitude spectrum of the known acoustic signal can be visually recognized is further executed. You may do it.
- a human determines a section including a known sound signal in the mixed sound signal based on the image display, and executes a correction step, a removal step, an inverse transformation step, or a synthesis step for this section.
- a sound reproducing step of reproducing the mixed acoustic signal, the known acoustic signal, and the output signal of the synthesis step as sound may be further executed.
- a human determines a section in which the known sound signal is included in the mixed sound signal, and in this section, a correction step, a removal step, an inverse transformation step, and a synthesis step. Perform the steps.
- the section in which the known acoustic signal is included in the mixed acoustic signal is automatically estimated based on the amplitude spectrum of the mixed acoustic signal, and a correction step is performed for this section.
- a removing step, an inverse transforming step, and a combining step may be executed. If the mixed sound signal contains a relatively well-known sound signal (for example, if there is a section where the known sound signal is sounding alone in the mixed sound signal), the section is automatically estimated. Can be identified, and by using automatic estimation, Removal work can be performed quickly. In the case where the existence of a known acoustic signal included in the mixed acoustic signal is not so clear, a human specifies a section.
- the known acoustic signal elimination method of the present invention when there are a plurality of types of known acoustic signals corresponding to the acoustic signals included in the mixed acoustic signal, all of the plurality of known acoustic signals are used. , A post-removal amplitude obtained by performing a known acoustic signal conversion step and a correction step, and performing a removal step of removing all of the corrected amplitude spectra of the plurality of known acoustic signals from the amplitude spectrum of the mixed acoustic signal. The inverse transformation step and the synthesis step are performed using the spectrum. This makes it possible to remove all types of known acoustic signals from the mixed acoustic signal.
- GUI graphic design interface
- the processing module that performs the interface processing removes the component of the known sound signal from the mixed sound signal in which multiple sound signals are mixed, and removes the accurate sound signal of the known sound signal in the mixed sound signal. It is configured so that the position can be specified by a human.
- the processing module that performs the in-plane processing is configured so that when the frequency characteristics of the known acoustic signal change over time in the mixed acoustic signal, these changes can be specified by humans. .
- the processing module that performs the in-plane processing is configured so that when the volume of a known acoustic signal changes with time in a mixed acoustic signal, humans can specify those changes. You.
- the processing module that performs the interface processing is configured such that when a known sound signal in the mixed sound signal expands or contracts in the time axis or frequency axis direction, the expansion and contraction of these can be specified by a human.
- the processing module that performs the interface processing is configured so that a human can specify a section corresponding to the mixed sound signal and the known sound signal.
- the known acoustic signal elimination device further includes a mixed acoustic signal conversion unit that converts the mixed acoustic signal into a time-frequency expression to obtain a wide spectrum of the mixed acoustic signal and a phase of the mixed acoustic signal.
- the temporal position of the amplitude spectrum of the known acoustic signal with respect to the amplitude spectrum of the mixed acoustic signal is shifted, the frequency characteristic changes over time, the volume changes over time, the expansion and contraction in the time axis direction, and the frequency axis direction Correction means for obtaining a corrected amplitude spectrum of a known sound signal in which at least one of expansion and contraction of the known sound signal is corrected, and removing a corrected amplitude spectrum of the known sound signal from the amplitude spectrum of the mixed sound signal
- Removing means for performing a reverse conversion to a time expression based on the amplitude spectrum after removal obtained by the removing means and the phase of the mixed acoustic signal to obtain a unit waveform, and synthesizing the unit waveform to obtain a known unit waveform
- the correction means includes a shift in the temporal position of the amplitude spectrum of the known sound signal with respect to the amplitude spectrum of the mixed sound signal, a time change of the frequency characteristic, a time change of the volume, expansion and contraction in the time axis direction, and frequency.
- Provide a processing module that performs interface processing that allows humans to manually specify at least one correction of axial expansion and contraction.
- the processing module that performs the interface processing includes an image display unit that displays an image so that the amplitude spectrum of the mixed sound signal and the amplitude spectrum of the known sound signal can be visually compared, a mixed sound signal, and a known sound signal. And a sound reproducing unit for generating an output signal of the synthesizing means as sound.
- the amplitude spectrum of the mixed sound signal and the amplitude spectrum of the known sound signal displayed on the image display unit can be displayed from the image display unit and the sound reproduction unit.
- the human being can not only specify the section of the known acoustic signal included in the mixed acoustic signal, but also manually specify the section of the amplitude spectrum of the known acoustic signal in this section manually.
- Displacement At least one correction of the time change of the frequency characteristic, the time change of the volume, the expansion and contraction in the time axis direction, and the expansion and contraction in the frequency axis direction can be specified.
- the known acoustic signal can be removed with high removal accuracy.
- the image display unit displays the amplitude spectrum of the section in the mixed sound signal containing the known sound signal, the displacement of the amplitude spectrum of the section corresponding to the known sound signal with respect to time, and the frequency characteristics. It is configured to be able to display the corrected amplitude spectrum, which is corrected for at least one of the time change of the volume, the time change of the sound volume, the expansion and contraction in the time axis direction, and the expansion and contraction in the frequency axis direction, on the time axis in alignment with each other. Is desirable.
- the state of the corrected amplitude spectrum can be visually checked, and it is possible to estimate how the corrected spectrum can be improved in removal accuracy by looking at the image. Removal work is faster.
- the image display unit is configured to be able to display an image of the amplitude spectrum of the sound signal obtained by removing the corrected amplitude spectrum from the amplitude spectrum of the mixed sound signal.
- the known acoustic signal elimination program further comprises: a mixed acoustic signal conversion step of converting the mixed acoustic signal into a time frequency signal to obtain an amplitude spectrum of the mixed acoustic signal and a phase of the mixed acoustic signal.
- Removal step to remove positive amplitude spectrum an inverse conversion step of performing an inverse conversion to a time expression based on the amplitude spectrum after removal obtained in the removal step and the phase of the mixed acoustic signal to obtain a unit waveform, and synthesizing the unit waveform to obtain a known acoustic signal. And a synthesizing step of obtaining an acoustic signal from which the component has been removed.
- the correction step includes a step of shifting the temporal position of the amplitude spectrum of the known acoustic signal with respect to the amplitude spectrum of the mixed acoustic signal, a temporal change in the frequency characteristic, and a sound volume.
- the known acoustic signal in which at least one of the time change, expansion and contraction in the time axis direction and expansion and contraction in the frequency axis direction is corrected, and the corrected amplitude spectrum is used as the amplitude spectrum of the mixed acoustic signal. Since it is removed from the vector, there is an advantage that the known acoustic signal included as non-stationary noise in the mixed acoustic signal can be removed with high accuracy.
- the sound signal elimination method of the present invention for example, when a sound signal of a TV program or a movie in which BGM is sounding in the background of a human voice or sound is input, the sound signal is separately prepared; It is possible to remove background music in the program using audio signals, and obtain audio signals only of human voices and body sounds. Further, by adding another music as BGM to the sound signal after the BGM removal, it is possible to reuse the music such as a TV program or a movie by replacing the music.
- the known sound signal here may be any sound signal, it can be applied regardless of the genre of the music, regardless of the presence or absence of vocals, and regardless of the presence or absence of accompaniment. In addition to music, it can be applied to any known noise including stationary noise and non-stationary noise.
- FIG. 1 shows a configuration of an example of an embodiment of a known acoustic signal elimination device of the present invention. It is a block diagram shown.
- FIG. 2 is a block diagram showing steps when the known acoustic signal elimination method of the present invention is carried out.
- FIG. 3 is a flowchart showing an example of an algorithm of a program used when a main part of the known acoustic signal elimination device of the present invention is realized by using a computer.
- FIG. 4 is a flowchart showing detailed processing in step ST103 of FIG.
- FIG. 5 is a flowchart showing the details of steps in the case of performing an estimation operation in both estimation involving humans and automatic estimation.
- FIG. 6 is a diagram showing a screen configuration of an editor interface.
- FIG. 7 is a diagram showing a time change of the power of the mixed acoustic signal.
- FIG. 8 is a diagram showing a time change of the amplitude spectrum of the mixed acoustic signal.
- FIG. 9 is a diagram showing a temporal change in power of a known acoustic signal of a sound source that is a source of BGM.
- FIG. 10 is a diagram showing a time change of the amplitude spectrum of a known acoustic signal of a sound source that is a source of BGM.
- FIG. 11 is a diagram showing a temporal change in power of a desired sound signal after removing a known sound signal.
- FIG. 12 is a diagram showing a time change of the amplitude spectrum of a desired sound signal after removing the known sound signal.
- FIG. 1 is a block diagram showing a configuration of an embodiment of a known acoustic signal elimination device for performing the known acoustic signal elimination method of the present invention.
- the known acoustic signal elimination device has a mixed acoustic signal converter It comprises a stage 1, a known acoustic signal conversion means 2, a correction means 3, an interface 4, a removal means 5, an inverse conversion means 6, and a synthesis means 7.
- the mixed sound signal converting means 1 is a mixed sound signal m (t) in which a sound signal b (t) such as BGM is mixed with a sound signal s (t) (t is a time axis) such as a desired sound or a body sound. (At this point, s (t) and b (t) are unknown, and only m (t) is input), converted to a time-frequency representation and mixed with the amplitude spectrum M ( ⁇ , t) of the mixed acoustic signal Find the sound signal phase 0m ( ⁇ , t).
- the known sound signal conversion means 2 converts the known sound signal b '(t) of the sound source, which is the source of the sound signal b (t) to be removed, into a time-frequency expression, and the amplitude spectrum ⁇ , ( ⁇ , t).
- the correction means 3 calculates the amplitude spectrum ⁇ , ( ⁇ , t) of the known acoustic signal with respect to the amplitude spectrum M ( ⁇ , t) of the mixed acoustic signal. )), The corrected amplitude spectrum B ( ⁇ , t) of the known acoustic signal in which the time shift of the frequency characteristic, the time change of the frequency characteristic, the time change of the sound volume, the expansion and contraction in the time axis direction and the expansion and contraction in the frequency axis direction are corrected.
- Correction means for automatically estimating and correcting all of displacement, time change of frequency characteristics, time change of sound volume, expansion and contraction in the time axis direction and expansion and contraction in the frequency axis direction. 3 can be configured.
- the correction means 3 corrects all of the positional deviation in time, the time change of the frequency characteristic, the time change of the volume, the expansion and contraction in the time axis direction and the expansion and contraction in the frequency axis direction. It is configured so that humans can specify it manually by using.
- the input source 4 has an image display unit that displays an image so that the amplitude spectrum of the mixed sound signal and the amplitude spectrum of the known sound signal can be visually compared. It is a processing module that performs interface processing using the graphic design interface (GU I).
- GUI graphic design interface
- Interface 4 uses the input section displayed on the screen to display the amplitude of the mixed acoustic signal. Based on the spectrum and the amplitude spectrum of the known sound signal, the section of the known sound signal included in the mixed sound signal can be designated by a human and the above-mentioned correction can be designated.
- the removing unit 5 removes the corrected amplitude spectrum B ( ⁇ , t) of the known sound signal from the amplitude spectrum M ( ⁇ , t) of the mixed acoustic signal. Then, the inverse transform means 6 performs an inverse transform to a time expression based on the amplitude spectrum S ( ⁇ , t) after removal obtained by the remover 5 and the phase 6> m (w, t) of the mixed acoustic signal. Find the unit waveform s' (t).
- the synthesizing means 7 synthesizes the unit waveform s ′ (t) output from the inverse transform means 6 to obtain an audio signal s (t) from which the components of the known audio signal have been removed.
- the interface 4 displays the post-removal amplitude spectrum S ( ⁇ , t) output from the removing unit 5 on the image display unit (see FIG. 6). Further, the interface 4 has a built-in sound reproducing unit, and reproduces a mixed sound signal, a known sound signal, and a synthesized sound signal output from the synthesizing means 7.
- FIG. 2 is a block diagram showing steps when the known acoustic signal elimination method of the present invention is carried out.
- FIG. 3 is a diagram showing a main part of the known acoustic signal elimination apparatus of the present invention realized by using a computer.
- 9 is a flowchart showing an example of a program algorithm used in the case.
- FIG. 4 is a D-chart showing the detailed processing in step ST103 of FIG. Fig. 5 shows both human-related estimation and automatic estimation. It is a flowchart which shows the detail of a step in performing an estimation process. The operation of removing a known acoustic signal in the method and apparatus for removing a known acoustic signal of the present invention will be described below with reference to FIGS. 1 to 5.
- the known acoustic signal b, (t) often undergoes the following deformation in the mixed sound m (t), so the component corresponding to b (t) is corrected by correction. Is estimated.
- the objects of the correction are mainly a time shift, a frequency characteristic change over time, a volume change over time, and expansion or contraction in the time axis or frequency axis direction, as described below.
- the position at which the known sound signal b 5 (t) is sounding in the mixed sound m (t) is not always from the beginning. Therefore, the known acoustic signal b '(t) is shifted in the time axis direction, It is necessary to subtract the known sound signal from the mixed sound by adjusting the relative position of the person.
- the frequency characteristics often change due to the influence of the graphic equalizer and the like. For example, low and high frequencies may be emphasized and attenuated. Therefore, it is necessary to correct the frequency characteristic of b '(t) by changing it in the same way, and to subtract the known sound signal from the mixed sound.
- step ST1 the mixed sound signal is Fourier-transformed to obtain the phase of the mixed sound signal (step ST2).
- step ST3 the amplitude spectrum (step ST3) of the mixed acoustic signal (step ST3), and a known acoustic signal corresponding to the acoustic signal included in the mixed acoustic signal is extracted in step ST4.
- Fourier transform is performed to determine the amplitude spectrum of the known sound signal (step ST5) (known sound signal conversion step).
- step ST6 based on the amplitude spectrum of the mixed acoustic signal, the temporal position shift of the amplitude spectrum of the known acoustic signal with respect to the amplitude spectrum of the mixed acoustic signal, the time change of the frequency characteristic, and the volume.
- step ST7 a corrected amplitude spectrum of the known acoustic signal in which at least one of the time change, expansion and contraction in the time axis direction, and expansion and contraction in the frequency axis direction is corrected is obtained (correction step).
- step ST8 the corrected amplitude spectrum of the known acoustic signal is removed from the amplitude spectrum of the mixed acoustic signal to obtain a post-removal amplitude spectrum (step ST9) (removal step).
- step ST10 a unit waveform is obtained by performing an inverse Fourier transform on the basis of the amplitude spectrum after removal obtained in the removing step and the phase of the mixed acoustic signal (inverse transforming step).
- step ST11 the unit waveform is synthesized by the overlapped quadrature method to obtain an acoustic signal from which the components of the known acoustic signal have been removed (synthesizing step). As shown in the flowchart of FIG.
- step ST 101 the mixed acoustic signal is subjected to Fourier transform to obtain the mixed acoustic signal. Find the amplitude spectrum and the phase of the mixed acoustic signal.
- step ST102 a known acoustic signal corresponding to the acoustic signal included in the mixed acoustic signal is Fourier-transformed to obtain an amplitude spectrum of the known acoustic signal.
- the temporal position shift of the amplitude spectrum of the known acoustic signal with respect to the amplitude spectrum of the mixed acoustic signal, and the time of the frequency characteristic Determine the corrected amplitude spectrum of the known sound signal by correcting at least one of the change, the volume change over time, the expansion and contraction in the time axis direction, and the expansion and contraction in the frequency axis direction.
- step ST104 the corrected amplitude spectrum of the known sound signal is removed from the amplitude spectrum of the mixed sound signal, and the post-removal amplitude spectrum is obtained.
- step ST105 the post-removal amplitude spectrum obtained in step ST104 is obtained.
- a unit waveform is obtained by performing an inverse Fourier transform on the basis of the phase of the mixed sound signal and the mixed sound signal.
- step ST106 the sound signal obtained by combining the unit waveforms by the overlap-add method to remove the components of the known sound signal Get.
- step ST107 a determination is made as to whether or not the user has evaluated the removed audio signal as being satisfactory. If the determination result is unsatisfactory, the correction is performed again in step ST103. It is. Until the user is satisfied, steps ST103 through ST107 are repeated.
- the subtraction process is performed on the amplitude spectrum in the time frequency domain without performing the subtraction process on the waveform in the time domain.
- S TFT short-time Fourier transform
- the audio signal is A / D-converted at a sampling frequency of 44.lkHz and a quantization bit rate of 16 bits, and a short-time Fourier transform using a 8192-point Haning window as the window function h (t) is performed.
- the frames of the fast Fourier transform (FFT) are shifted by 441 points, so the frame shift time (one frame shift) is 10 ms. This frame shift is used as the processing time unit.
- the amplitude spectrum S ( ⁇ , t) of the desired sound signal s (t) after the removal of the known sound signal is obtained from the amplitude spectrum M ( ⁇ , t), ⁇ ′ ( ⁇ , t) by the following equation.
- B ( ⁇ , t) is the amplitude spectrum after correcting ⁇ , ( ⁇ , t).
- a (t) is a function of an arbitrary shape for finally adjusting the amount of subtracting a component corresponding to the amplitude spectrum of the known sound signal from the amplitude spectrum of the mixed sound.
- a (t) 1 The greater this is, the greater the amount of subtraction It will be good.
- g ( ⁇ , t) is a function for correcting the time change of the frequency characteristic and the time change of the volume.
- P ( ⁇ ) is a function for correcting expansion and contraction in the frequency axis direction.
- ⁇ the frequency axis ⁇ of the amplitude spectrum ⁇ , ( ⁇ , t)
- ⁇ '( ⁇ , t) takes 0 outside the original domain of ⁇ , and interpolates as appropriate when discretized and implemented.
- q (t) is a function for correcting the expansion and contraction in the time axis direction.By converting the time axis t of the amplitude spectrum ⁇ '( ⁇ , t), linear (non-linear) Enables a type of expansion and contraction. Note that ⁇ , ( ⁇ , t) takes 0 outside the original domain of t, and interpolates as appropriate when implementing discretely.
- r (t) is a function for correcting the time position shift, and usually corrects a certain amount of shift by setting a constant. If the shift width changes over time, set a function to correct the width at each time. Note that ⁇ '( ⁇ , t) takes 0 outside the domain of the original t, and interpolates as appropriate when discretized and implemented. Although it is possible to express it as a single function integrating q (t) and (t), here q (t) is set to represent continuous expansion and contraction, and r (t) is a discrete position It is set for the purpose of expressing the deviation of.
- the various parameter overnight functions a (t), g ( ⁇ , t) (g w ( ⁇ , ⁇ ) of the equations (11), (12) and (13) are used.
- t), g t (t), g r (t)), p ( ⁇ ), q (t), r (t), c ( ⁇ , t) May be set manually. Alternatively, it may be corrected by a human after the automatic estimation. In the following, a description will be given of a specific automatic estimation method and a case of using the interface 4 in the known acoustic signal elimination apparatus which enables manual correction by a human.
- step ST201 the various parameter functions g ( ⁇ , t) (g M ( ⁇ , t), g t (t)), p ( ⁇ ) in equations (1 1), (1 2) and (13) , q (t "), r (t)
- step ST201 the set ⁇ of the BGM interval ⁇ is specified and automatically estimated.
- P ( ⁇ ) in the step ST 2 0 2 performs automatic estimation of q (t), gw in step ST 2 03 ( ⁇ , t) , g t (t), for automatic estimation of r (t).
- step ST 205 the correction operation is performed using interface 4.
- BGM section a section containing almost no sound signal s (t) of only human voices or body sounds.
- BGM section a section containing almost no sound signal s (t) of only human voices or body sounds.
- a plurality of BGM sections may be used.
- g w ( ⁇ , t) is estimated by interpolation (interpolation or extrapolation). (If there are BGM sections on both sides, interpolation is performed from both sides.) Finally, g w ( ⁇ , t) is smoothed in the frequency axis direction. Note that the smoothing width can be set arbitrarily. It is not necessary.
- the amplitude spectrum M ( ⁇ , t) is compared with the amplitude at each time of g w ( ⁇ , t) ⁇ and ( ⁇ , t) after frequency characteristic correction.
- ⁇ represents the set of. Any division can be applied to ⁇ . For example, it is good to divide every equal octave of equal temperament used in music (divide at equal intervals on the logarithmic frequency axis).
- g t (t) is min (g :, t (, t)) or Estimate by In the case of min (:, t (, t)), the amplitudes are compared in the frequency band where M ( ⁇ , t) and g w ( ⁇ , t) ⁇ '( ⁇ , t) are closest.
- g t (t) is smoothed in the time axis direction.
- the smoothing width can be arbitrarily set, and the smoothing need not be performed.
- P ( ⁇ ) and q (t) are set so that the distance between M ( ⁇ , t) and ⁇ ( ⁇ , t) (for example, the logarithmic spectral distance) is minimized.
- r (t) In estimating r (t), in principle, the set ⁇ of BGM sections ⁇ is used, and the time axis of the correspondence between M ( ⁇ , t) and ⁇ ( ⁇ , t) in those sections is adjusted. R (t) is obtained. r (t) is often a constant, but if a section of the known sound signal b 3 (t) is not used, but is mixed while being used intermittently, skip that section. Thus, r (t) is a discontinuous function.
- FIG. 5 is a flowchart showing a software algorithm of a program corresponding to a case where a human specifies manually and a case where an automatic estimation is performed.
- steps ST 302 to ST 3 13 in FIG. 5 are executed.
- ⁇ basically, a set of the remaining BGM sections is obtained by using one BGM section ⁇ 1 as a clue.
- the first ⁇ 1 is determined manually by a human or by finely dividing the time axis of the audio signal and determining the correspondence between these short divided sections. If not manually specified by a human, B ( ⁇ , t) is temporarily calculated (step ST 302), and the amplitude spectrum of the time window obtained by dividing M ( ⁇ , t) and B ( ⁇ , t) into small pieces is calculated. Is calculated (corresponding to the degree of similarity) (step ST303). Then, the correspondence of the time window of the minimum distance is checked (step ST 304), and the section including the result is set to ⁇ ⁇ ⁇ 1 as an initial value (step ST 305).
- step ST 306 various parameter overnight functions of B ( ⁇ , t) are estimated (step ST 306 to step S 309 309), and ⁇ ( ⁇ , t) is calculated (step ST310). ).
- step ST310 Check whether the estimated values for each parameter have converged, and if not converged, the distance between the amplitude spectra of M (, t) and B ( ⁇ , t) for the entire interval of (Equivalent to similarity).
- a constant multiple of the maximum value (or the average value) is set as a BGM section determination threshold value (step ST312).
- a section having a distance equal to or less than the BGM section determination threshold is detected and newly added to ⁇ (step ST313).
- an upper limit can be set for addition.
- ⁇ is updated, and various parameter overnight functions are determined appropriately.
- the distance between M ( ⁇ , t) and ⁇ ( ⁇ , t) is, for example, the root mean square log spectrum distance. Is valid.
- This Eddy evening is roughly divided and mixed A sub-window W 3_ for operating the acoustic signal m (t), a sub-window W 2 for operating the known acoustic signal b ′ (t), and a sub-window W 3 for operating the desired acoustic signal s (t) after removing the known acoustic signal. It consists of three subwindows.
- the known acoustic signals b' (t) operated in the sub-window W2 can be switched by the changeover switch W2S. In this evening festival, steps ST 219 to ST 219 shown in FIG. 4 are executed.
- the operation range slider P1 indicates where in the sound signal is currently displayed.
- Cursor P2 indicates the current position of the operation target on the time axis.
- the button P3 When the button P3 is pressed, the subwindow to which the button belongs is temporarily collapsed and becomes smaller. Unused sub-windows other than the one currently being operated can be hidden to effectively use the narrow screen. Pressing the float (enlarge) button P4 temporarily disconnects the subwindow to which the button belongs from the parent window (float), and further enlarges it to facilitate operation and editing. If only the button P 4 is drawn, press this button to float the sub window associated with it and make it appear again.
- a graph E1 of the power of the mixed acoustic signal m (t) and a graph E2 of its amplitude spectrum M ( ⁇ , t) are displayed.
- a graph E3 of the power of the known acoustic signal b, (t) and a graph E4 of the amplitude spectrum ⁇ , ( ⁇ , t) are displayed.
- a graph E5 of the power of the acoustic signal s (t) after the removal of the known acoustic signal and a graph E6 of the amplitude spectrum S ( ⁇ , t) thereof are displayed.
- the amplitude is drawn in shades on the left (the horizontal axis is the time axis, the vertical axis is the frequency axis), and the amplitude at the force-sol position is drawn on the right (The horizontal axis is power and the vertical axis is frequency axis).
- a group of buttons capable of playing, stopping, fast-forwarding, and fast-forwarding the mixed acoustic signal are arranged for human listening and confirmation.
- the interface 4 reproduces the mixed acoustic signal by the built-in acoustic reproducing unit.
- the sub-window W2 for the operation of the known sound signal b '(t) is the main window of the operation, and all the parameter overnight functions a (t), g in Equations (1 2) and (13) -The shape of ( ⁇ , t) (g w ( ⁇ , t), g t (t), g r (t)), p ( ⁇ ), q (t), r (t) can be set freely .
- the following is a description of each operation panel.
- Operation panel C 1 (right side of E7) g w ( ⁇ 3 t) for correcting the time change of the frequency characteristics
- This panel is used to display and operate g a ( ⁇ , t) is drawn (the horizontal axis is the size, the vertical axis is the frequency axis).
- the result of the setting operation is immediately reflected on the display panel E7 of g ( ⁇ , t) (steps ST205 and ST206).
- the magnitude of the value of g ( ⁇ , t) is drawn in shades (the horizontal axis is the time axis, and the vertical axis is the frequency axis).
- g t (t) is displayed on the operation panel.
- the result of the setting operation is immediately reflected on the display panel E 7 of g ( ⁇ , t) (steps ST 207 and ST 208).
- Operation panel C 3 (lower side of E 7) for raising the value of g ( ⁇ , t) as a whole
- g r (t) is displayed.
- the result of the setting operation is immediately reflected on the display panel E 7 of g ( ⁇ , t) (steps ST209 and ST210).
- Display P ( ⁇ ) 'It is a panel for operation. When this panel is operated, the change in ⁇ ( ⁇ ) is immediately reflected on the display (steps ST213 and ST214).
- q (t) is displayed. This is the operation panel. When this panel is operated, the change of q (t) is immediately reflected on the display (step S ⁇ 2 15, ST 2 16).
- the change in r (t) is immediately reflected on the display (steps ST 2 17 and ST 2 18).
- a group of buttons capable of playing, stopping, fast-forwarding, and fast-returning a known sound signal are arranged for human listening and confirmation.
- the interface 4 reproduces a known acoustic signal by a built-in acoustic reproduction unit.
- volume feeder operation panel C 9 (below E 8)
- This panel is used to display and operate the shape of c ( ⁇ , t) in the t direction.
- the result of the setting operation is immediately reflected on the display panel E8 of c ( ⁇ , t).
- a group of buttons capable of playing, stopping, fast-forwarding, and rewinding the synthesized audio signal are arranged side by side for human listening and confirmation. I have.
- the interface 4 reproduces the audio signal synthesized by the built-in audio reproduction unit.
- a program that can determine the unknown s (t) under the condition that the sound signals b and (t) of the sound source are known is provided by various operating systems (Liux 2.4, SG IIRIX 6.5 , Microsoft Windows XP: registered trademark).
- an audio file containing m (t) and b '(t) is provided to this program.
- BGM background music
- FIGS. 7 to 12 show the results of actually processing a mixed sound of classical music being played in the background music of a dialogue between two men and women.
- the mixed sound signal in (t) shown in FIGS. 7 and 8 as input and the known sound signal b '(t) of the original sound source shown in FIGS. 9 and 10; removing the BGM component
- the result is the sound signal s (t) after the removal of the known sound signal shown in FIGS. 11 and 12.
- the mixed sound in the example of this processing result is the audio signal of the dialogue between two men and women extracted from “: RWCP Voice Dialogue Day” and the audio signal extracted from “RWC Research Music Day”.
- the sound signal of Lasik music is added.
- the correction step the temporal position shift of the amplitude spectrum of the known acoustic signal with respect to the amplitude spectrum of the mixed acoustic signal, the time change of the frequency characteristic, A corrected amplitude spectrum of a known sound signal in which at least one of volume change over time, expansion and contraction in the time axis direction, and expansion and contraction in the frequency axis direction is corrected is obtained, and the corrected amplitude spectrum is used as the amplitude spectrum of the mixed acoustic signal. Therefore, the known acoustic signal included as non-stationary noise in the mixed acoustic signal can be removed with high accuracy.
- an audio signal from a TV program or movie with BGM is input in the background of human voice or sound
- the BGM in the program is removed using the music audio signal of BGM prepared separately, and human music is removed. It is possible to obtain an acoustic signal consisting of only voices and body sounds.
- the known sound signal may be any sound signal, it can be applied regardless of the genre of the music, regardless of the presence or absence of vocals, and regardless of the presence of accompaniment. Also, the present invention is applicable not only to music but also to any known noise including stationary noise and non-stationary noise.
Landscapes
- Engineering & Computer Science (AREA)
- Computational Linguistics (AREA)
- Quality & Reliability (AREA)
- Signal Processing (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Human Computer Interaction (AREA)
- Physics & Mathematics (AREA)
- Acoustics & Sound (AREA)
- Multimedia (AREA)
- Stereophonic System (AREA)
- Auxiliary Devices For Music (AREA)
- Measurement Of Mechanical Vibrations Or Ultrasonic Waves (AREA)
Abstract
Description
Claims
Priority Applications (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| GB0526570A GB2418577B (en) | 2003-05-30 | 2004-05-26 | Method and device for removing known acoustic signal |
| US10/558,608 US20070021959A1 (en) | 2003-05-30 | 2004-05-26 | Method and device for removing known acoustic signal |
Applications Claiming Priority (4)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP2003-154964 | 2003-05-30 | ||
| JP2003154964 | 2003-05-30 | ||
| JP2003167118A JP4608650B2 (ja) | 2003-05-30 | 2003-06-11 | 既知音響信号除去方法及び装置 |
| JP2003-167118 | 2003-06-11 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2004107319A1 true WO2004107319A1 (ja) | 2004-12-09 |
Family
ID=33492453
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/JP2004/007587 Ceased WO2004107319A1 (ja) | 2003-05-30 | 2004-05-26 | 既知音響信号除去方法及び装置 |
Country Status (5)
| Country | Link |
|---|---|
| US (1) | US20070021959A1 (ja) |
| JP (1) | JP4608650B2 (ja) |
| KR (1) | KR101008250B1 (ja) |
| GB (1) | GB2418577B (ja) |
| WO (1) | WO2004107319A1 (ja) |
Cited By (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US10043532B2 (en) | 2014-03-17 | 2018-08-07 | Nec Corporation | Signal processing apparatus, signal processing method, and signal processing program |
| CN110970045A (zh) * | 2019-11-15 | 2020-04-07 | 北京达佳互联信息技术有限公司 | 混音处理方法、装置、电子设备和存储介质 |
Families Citing this family (12)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2006243664A (ja) * | 2005-03-07 | 2006-09-14 | Nippon Telegr & Teleph Corp <Ntt> | 信号分離装置、信号分離方法、信号分離プログラム及び記録媒体 |
| CN101300623B (zh) * | 2005-09-02 | 2011-07-27 | 日本电气株式会社 | 用于抑制噪声的方法、设备和计算机程序 |
| ATE458361T1 (de) | 2005-12-13 | 2010-03-15 | Nxp Bv | Einrichtung und verfahren zum verarbeiten eines audio-datenstroms |
| WO2011001589A1 (ja) * | 2009-06-29 | 2011-01-06 | 三菱電機株式会社 | オーディオ信号処理装置 |
| CN102576543B (zh) * | 2010-07-26 | 2014-09-10 | 松下电器产业株式会社 | 多输入噪声抑制装置、多输入噪声抑制方法以及集成电路 |
| US8849199B2 (en) | 2010-11-30 | 2014-09-30 | Cox Communications, Inc. | Systems and methods for customizing broadband content based upon passive presence detection of users |
| US20120136658A1 (en) * | 2010-11-30 | 2012-05-31 | Cox Communications, Inc. | Systems and methods for customizing broadband content based upon passive presence detection of users |
| JP5703807B2 (ja) * | 2011-02-08 | 2015-04-22 | ヤマハ株式会社 | 信号処理装置 |
| WO2013046055A1 (en) * | 2011-09-30 | 2013-04-04 | Audionamix | Extraction of single-channel time domain component from mixture of coherent information |
| US9195431B2 (en) * | 2012-06-18 | 2015-11-24 | Google Inc. | System and method for selective removal of audio content from a mixed audio recording |
| US9373320B1 (en) * | 2013-08-21 | 2016-06-21 | Google Inc. | Systems and methods facilitating selective removal of content from a mixed audio recording |
| US10052494B2 (en) * | 2014-12-23 | 2018-08-21 | Medtronic, Inc. | Hemodynamically unstable ventricular arrhythmia detection |
Citations (7)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPH05197385A (ja) * | 1992-01-20 | 1993-08-06 | Sanyo Electric Co Ltd | 音声認識装置 |
| JPH07199990A (ja) * | 1993-12-28 | 1995-08-04 | Ricoh Co Ltd | 音声認識装置 |
| JPH08107375A (ja) * | 1994-10-06 | 1996-04-23 | Hitachi Ltd | 音響信号記録再生装置 |
| JPH10228296A (ja) * | 1997-02-17 | 1998-08-25 | Nippon Telegr & Teleph Corp <Ntt> | 音響信号分離方法 |
| JPH10307595A (ja) * | 1997-03-07 | 1998-11-17 | Seiko Epson Corp | 入力音声抽出方法および入力音声抽出装置 |
| JP2003022100A (ja) * | 2001-07-09 | 2003-01-24 | Yamaha Corp | 雑音除去方法、雑音除去装置およびプログラム |
| JP2003099085A (ja) * | 2001-09-25 | 2003-04-04 | National Institute Of Advanced Industrial & Technology | 音源の分離方法および音源の分離装置 |
Family Cites Families (10)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US5204969A (en) * | 1988-12-30 | 1993-04-20 | Macromedia, Inc. | Sound editing system using visually displayed control line for altering specified characteristic of adjacent segment of stored waveform |
| US5792971A (en) * | 1995-09-29 | 1998-08-11 | Opcode Systems, Inc. | Method and system for editing digital audio information with music-like parameters |
| US6343268B1 (en) * | 1998-12-01 | 2002-01-29 | Siemens Corporation Research, Inc. | Estimator of independent sources from degenerate mixtures |
| US6446041B1 (en) * | 1999-10-27 | 2002-09-03 | Microsoft Corporation | Method and system for providing audio playback of a multi-source document |
| JP3454206B2 (ja) * | 1999-11-10 | 2003-10-06 | 三菱電機株式会社 | 雑音抑圧装置及び雑音抑圧方法 |
| US6879952B2 (en) * | 2000-04-26 | 2005-04-12 | Microsoft Corporation | Sound source separation using convolutional mixing and a priori sound source knowledge |
| JP4028680B2 (ja) * | 2000-11-01 | 2007-12-26 | インターナショナル・ビジネス・マシーンズ・コーポレーション | 観測データから原信号を復元する信号分離方法、信号処理装置、モバイル端末装置、および記憶媒体 |
| US7076433B2 (en) * | 2001-01-24 | 2006-07-11 | Honda Giken Kogyo Kabushiki Kaisha | Apparatus and program for separating a desired sound from a mixed input sound |
| US7243060B2 (en) * | 2002-04-02 | 2007-07-10 | University Of Washington | Single channel sound separation |
| US6971323B2 (en) * | 2004-03-19 | 2005-12-06 | Peat International, Inc. | Method and apparatus for treating waste |
-
2003
- 2003-06-11 JP JP2003167118A patent/JP4608650B2/ja not_active Expired - Lifetime
-
2004
- 2004-05-26 US US10/558,608 patent/US20070021959A1/en not_active Abandoned
- 2004-05-26 WO PCT/JP2004/007587 patent/WO2004107319A1/ja not_active Ceased
- 2004-05-26 KR KR1020057021034A patent/KR101008250B1/ko not_active Expired - Fee Related
- 2004-05-26 GB GB0526570A patent/GB2418577B/en not_active Expired - Fee Related
Patent Citations (7)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPH05197385A (ja) * | 1992-01-20 | 1993-08-06 | Sanyo Electric Co Ltd | 音声認識装置 |
| JPH07199990A (ja) * | 1993-12-28 | 1995-08-04 | Ricoh Co Ltd | 音声認識装置 |
| JPH08107375A (ja) * | 1994-10-06 | 1996-04-23 | Hitachi Ltd | 音響信号記録再生装置 |
| JPH10228296A (ja) * | 1997-02-17 | 1998-08-25 | Nippon Telegr & Teleph Corp <Ntt> | 音響信号分離方法 |
| JPH10307595A (ja) * | 1997-03-07 | 1998-11-17 | Seiko Epson Corp | 入力音声抽出方法および入力音声抽出装置 |
| JP2003022100A (ja) * | 2001-07-09 | 2003-01-24 | Yamaha Corp | 雑音除去方法、雑音除去装置およびプログラム |
| JP2003099085A (ja) * | 2001-09-25 | 2003-04-04 | National Institute Of Advanced Industrial & Technology | 音源の分離方法および音源の分離装置 |
Cited By (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US10043532B2 (en) | 2014-03-17 | 2018-08-07 | Nec Corporation | Signal processing apparatus, signal processing method, and signal processing program |
| CN110970045A (zh) * | 2019-11-15 | 2020-04-07 | 北京达佳互联信息技术有限公司 | 混音处理方法、装置、电子设备和存储介质 |
| CN110970045B (zh) * | 2019-11-15 | 2022-03-25 | 北京达佳互联信息技术有限公司 | 混音处理方法、装置、电子设备和存储介质 |
Also Published As
| Publication number | Publication date |
|---|---|
| US20070021959A1 (en) | 2007-01-25 |
| KR101008250B1 (ko) | 2011-01-17 |
| GB2418577A (en) | 2006-03-29 |
| KR20060034637A (ko) | 2006-04-24 |
| JP4608650B2 (ja) | 2011-01-12 |
| GB0526570D0 (en) | 2006-02-08 |
| JP2005049364A (ja) | 2005-02-24 |
| GB2418577B (en) | 2007-10-17 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| JP5467098B2 (ja) | オーディオ信号をパラメータ化された表現に変換するための装置および方法、パラメータ化された表現を修正するための装置および方法、オーディオ信号のパラメータ化された表現を合成するための装置および方法 | |
| Le Roux et al. | Explicit consistency constraints for STFT spectrograms and their application to phase reconstruction. | |
| TWI505264B (zh) | 操縱具有瞬變事件的音頻信號的設備和方法以及具有執行該方法之程式碼的電腦程式 | |
| US11410637B2 (en) | Voice synthesis method, voice synthesis device, and storage medium | |
| JP4608650B2 (ja) | 既知音響信号除去方法及び装置 | |
| US20050137729A1 (en) | Time-scale modification stereo audio signals | |
| MX2012009776A (es) | Aparato y metodo para modificar una señal de audio usando bloqueo armonico. | |
| RU2510954C2 (ru) | Способ переозвучивания аудиоматериалов и устройство для его осуществления | |
| CN111739544B (zh) | 语音处理方法、装置、电子设备及存储介质 | |
| WO2002050814A1 (fr) | Systeme et procede d'interpolation de signaux | |
| WO2003003345A1 (en) | Device and method for interpolating frequency components of signal | |
| US20050038534A1 (en) | Fixed-size cross-correlation computation method for audio time scale modification | |
| CN109416911B (zh) | 声音合成装置及声音合成方法 | |
| JP4274419B2 (ja) | 音響信号除去装置、音響信号除去方法及び音響信号除去プログラム | |
| KR20050010927A (ko) | 오디오 신호 처리 장치 | |
| JP2009282536A (ja) | 既知音響信号除去方法及び装置 | |
| JP3849679B2 (ja) | 雑音除去方法、雑音除去装置およびプログラム | |
| JP4274418B2 (ja) | 音響信号除去装置、音響信号除去方法及び音響信号除去プログラム | |
| US20100131276A1 (en) | Audio signal synthesis | |
| JP4272107B2 (ja) | 音響信号除去装置、音響信号除去方法及び音響信号除去プログラム | |
| WO2020179472A1 (ja) | 信号処理装置および方法、並びにプログラム | |
| US11348596B2 (en) | Voice processing method for processing voice signal representing voice, voice processing device for processing voice signal representing voice, and recording medium storing program for processing voice signal representing voice | |
| JPH11143460A (ja) | 音楽演奏に含まれる旋律の分離方法、分離抽出方法および分離除去方法 | |
| Disch et al. | An iterative segmentation algorithm for audio signal spectra depending on estimated local centers of gravity | |
| Wu | Musical pitch shifting based on equalization and bandwidth extension |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| AK | Designated states |
Kind code of ref document: A1 Designated state(s): AE AG AL AM AT AU AZ BA BB BG BR BW BY BZ CA CH CN CO CR CU CZ DE DK DM DZ EC EE EG ES FI GB GD GE GH GM HR HU ID IL IN IS KE KG KP KR KZ LC LK LR LS LT LU LV MA MD MG MK MN MW MX MZ NA NI NO NZ OM PG PH PL PT RO RU SC SD SE SG SK SL SY TJ TM TN TR TT TZ UA UG US UZ VC VN YU ZA ZM ZW |
|
| AL | Designated countries for regional patents |
Kind code of ref document: A1 Designated state(s): GM KE LS MW MZ NA SD SL SZ TZ UG ZM ZW AM AZ BY KG KZ MD RU TJ TM AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HU IE IT LU MC NL PL PT RO SE SI SK TR BF BJ CF CG CI CM GA GN GQ GW ML MR NE SN TD TG |
|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application | ||
| WWE | Wipo information: entry into national phase |
Ref document number: 1020057021034 Country of ref document: KR |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 0526570.7 Country of ref document: GB Ref document number: 0526570 Country of ref document: GB |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 2007021959 Country of ref document: US Ref document number: 10558608 Country of ref document: US |
|
| WWP | Wipo information: published in national office |
Ref document number: 1020057021034 Country of ref document: KR |
|
| 122 | Ep: pct application non-entry in european phase | ||
| WWP | Wipo information: published in national office |
Ref document number: 10558608 Country of ref document: US |



