WO2004107319A1 - 既知音響信号除去方法及び装置 - Google Patents

既知音響信号除去方法及び装置 Download PDF

Info

Publication number
WO2004107319A1
WO2004107319A1 PCT/JP2004/007587 JP2004007587W WO2004107319A1 WO 2004107319 A1 WO2004107319 A1 WO 2004107319A1 JP 2004007587 W JP2004007587 W JP 2004007587W WO 2004107319 A1 WO2004107319 A1 WO 2004107319A1
Authority
WO
WIPO (PCT)
Prior art keywords
acoustic signal
mixed
signal
amplitude spectrum
sound
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/JP2004/007587
Other languages
English (en)
French (fr)
Inventor
Masataka Goto
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
National Institute of Advanced Industrial Science and Technology AIST
Original Assignee
National Institute of Advanced Industrial Science and Technology AIST
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by National Institute of Advanced Industrial Science and Technology AIST filed Critical National Institute of Advanced Industrial Science and Technology AIST
Priority to GB0526570A priority Critical patent/GB2418577B/en
Priority to US10/558,608 priority patent/US20070021959A1/en
Publication of WO2004107319A1 publication Critical patent/WO2004107319A1/ja
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L21/00Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
    • G10L21/02Speech enhancement, e.g. noise reduction or echo cancellation
    • G10L21/0272Voice signal separating
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L21/00Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
    • G10L21/02Speech enhancement, e.g. noise reduction or echo cancellation

Definitions

  • the present invention relates to a known sound signal elimination method and a known sound signal elimination device for removing a component of a known sound signal from a mixed sound signal in which a plurality of sound signals are mixed.
  • Non-Patent Document 1 a method called a spectral subtraction method (Non-Patent Document 1) has been known as an acoustic signal processing.
  • the conventional spectral subtraction method uses a sound signal (mixed sound) that is a mixture of stationary noise (noise whose spectrum does not change over time and frequency characteristics and volume are almost constant) and a desired sound (target sound). ) To obtain a target sound by removing stationary noise.
  • the spectrum of the stationary noise is learned in advance by a simple method such as calculating the average of the stationary spectrum, and the spectrum of the stationary noise is calculated from the spectrum of the input mixed sound. Perform the removal process. That is, the process of subtracting the average of the noise is performed.
  • Patent Document 2 Japanese Patent Application Laid-Open No. 2000-201
  • Patent Document 3 Japanese Patent Application Laid-Open Publication No. 2000-1922
  • Patent Document 4 Japanese Patent Application Laid-Open No. 2000-201
  • Patent Document 5 Japanese Patent Application Laid-Open No. Hei 11-0-309
  • Patent Document 6 Japanese Patent Application Laid-Open No. H10-24029
  • Patent Document 7 Japanese Patent Application Laid-Open No. 08-2221092 Disclosure of the Invention
  • the conventional spectral subtraction method presupposes stationary noise and cannot be applied to non-stationary noise (noise whose spectrum changes significantly over time and whose frequency characteristics and volume also change). For example, it was not possible to remove time-varying non-stationary noise, such as music used as background music (BGM). This is because the spectrum of the non-stationary noise changes too much for learning.
  • non-stationary noise such as music used as background music (BGM).
  • an object of the present invention is to convert a component of a known sound signal (which may be non-stationary or stationary) from a mixed sound signal in which a plurality of sound signals are mixed into a known sound signal from an original sound source corresponding thereto.
  • a known acoustic signal that can be removed using the signal It is intended to provide a removing method, a known acoustic signal removing device, and a program used for the device.
  • Another object of the present invention is to provide, for example, a method in which a known sound signal is music, and the music sound signal is obtained from a mixed sound used as background music (BGM) for human voices and body sounds.
  • BGM background music
  • a known sound signal elimination method and a known sound signal that can remove background music using a known sound signal for example, a sound signal of the same music separately obtained from a CD, a record, or the like
  • a known sound signal for example, a sound signal of the same music separately obtained from a CD, a record, or the like
  • Still another object of the present invention is to remove a component of a known acoustic signal from an acoustic signal (mixed sound) in which a plurality of acoustic signals are mixed, and to accurately detect a known acoustic signal in the mixed sound. It is an object of the present invention to provide a known sound signal removing method and device capable of automatically estimating a position and removing a known sound signal at the position, and a program used for the device.
  • Still another object of the present invention is to remove a component of a known acoustic signal from an acoustic signal (mixed sound) in which a plurality of acoustic signals are mixed, and to accurately detect a known acoustic signal in the mixed sound. It is an object of the present invention to provide a known acoustic signal elimination device provided with an interface in which a position can be designated by a human.
  • Still another object of the present invention is to remove a component of a known sound signal from a sound signal (mixed sound) in which a plurality of sound signals are mixed.
  • Another object of the present invention is to provide a known acoustic signal eliminator provided with an interface that allows a person to specify the expansion and contraction when the expansion and contraction is performed in the frequency axis direction.
  • Still another object of the present invention is to remove a plurality of known acoustic signals from acoustic signals mixed with a plurality of acoustic signals, and remove the known acoustic signals one by one.
  • An object of the present invention is to provide a method and an apparatus for removing a known acoustic signal and a program used for the apparatus.
  • a component of a known acoustic signal (which may be non-stationary or stationary) is mixed with a known acoustic signal from an original sound source from a mixed acoustic signal in which a plurality of acoustic signals are mixed. Remove with.
  • the mixed acoustic signal is converted into a time-frequency expression to determine the amplitude spectrum of the mixed acoustic signal and the phase of the mixed acoustic signal (the mixed acoustic signal conversion step).
  • a known conversion method such as Fourier transform or wave-rate transform is used.
  • a known sound signal (a sound signal of the same music separately obtained from a CD or record) corresponding to (similar to) the known sound signal included in the mixed sound signal is converted into a time-frequency expression. Conversion is performed to obtain the amplitude spectrum of the known sound signal (known sound signal conversion step).
  • the time position shift of the amplitude spectrum of the known sound signal with respect to the amplitude spectrum of the mixed sound signal, the time change of the frequency characteristic, the time of the sound volume A corrected amplitude spectrum of the known acoustic signal is obtained by correcting at least one of the change, the expansion and contraction in the time axis direction, and the expansion and contraction in the frequency axis direction (correction step).
  • the correction amplitude spectrum of the known sound signal is removed from the amplitude spectrum of the mixed sound signal (removal step). Based on the amplitude spectrum after removal obtained in this removal step and the phase of the mixed acoustic signal, inverse conversion is performed on the time expression to obtain a unit waveform (inverse conversion step).
  • the unit waveform is synthesized by using a synthesis method such as an overlapping quadrature method to obtain an audio signal from which components of the known audio signal have been removed (synthesis step).
  • the following correction step is executed to shift the temporal position of the torsion width spectrum of the known acoustic signal with respect to the amplitude spectrum of the mixed acoustic signal.
  • a corrected amplitude spectrum of the known sound signal in which at least one of the time change of the frequency characteristic, the time change of the sound volume, the expansion and contraction in the time axis direction and the expansion and contraction in the frequency axis direction is corrected is obtained, and the corrected amplitude spectrum is calculated. It is removed from the amplitude spectrum of the mixed sound signal. For this reason, the known acoustic signal included as irregular noise in the mixed acoustic signal can be removed with high accuracy.
  • the time shift of the amplitude spectrum of the known acoustic signal the time change of the frequency characteristic, the time change of the volume, the expansion and contraction in the time axis direction and the frequency axis direction It is preferable to correct all occurrences of the phenomenon or change in the mixed acoustic signal.
  • the accuracy of removal of the known sound signal can be improved rather than the case where no correction is made. You don't have to do everything. Of course, all necessary corrections may be made.
  • the temporal position of the known acoustic signal included in the mixed acoustic signal is estimated, and the temporal position of the amplitude spectrum of the known acoustic signal is shifted based on the estimated temporal position. to correct.
  • the estimation method is, for example, obtaining a distance (similarity) between a predetermined section of the amplitude spectrum of the mixed sound signal and a predetermined section of the amplitude spectrum of the known sound signal, and determining a section having the closest distance to the mixed sound signal. It is estimated as the temporal position of the known acoustic signal included in.
  • a change in the frequency characteristic of the known acoustic signal included in the mixed acoustic signal is estimated, and the frequency characteristic of the amplitude spectrum of the known acoustic signal is estimated based on the time change of the estimated frequency characteristic. Is corrected over time.
  • the estimation of the change in the frequency characteristic is performed, for example, by identifying a section of the mixed acoustic signal that includes only the known acoustic signal, and comparing the frequency characteristic of this section with the frequency characteristic of the known acoustic signal corresponding to this section. From the contrast, the change of the frequency characteristic of the known acoustic signal included in the mixed acoustic signal is estimated.
  • a temporal change in the volume of the known acoustic signal included in the mixed acoustic signal is estimated, and a temporal change in the volume of the amplitude spectrum of the known acoustic signal is determined based on the estimated temporal change in the volume.
  • Is corrected After estimating the time change of the sound volume, after correcting the frequency characteristics, for example, a frequency band having an amplitude corresponding to a known sound signal included in the mixed sound signal is specified at each time, and the mixing in the frequency band is determined. It is estimated from the contrast between the amplitude of the acoustic signal and the amplitude of the known acoustic signal.
  • the expansion and contraction of the known acoustic signal included in the mixed acoustic signal in the time axis direction is estimated, and the time of the amplitude spectrum of the known acoustic signal is determined based on the estimated expansion and contraction in the time axis direction.
  • Correct axial expansion and contraction for example, a section containing only a known acoustic signal in the mixed acoustic signal is specified, and a time axis comparison with a section of the known acoustic signal corresponding to this section is performed. Estimate the expansion and contraction in the time axis direction. Or divided the time axis into short sections Estimate by comparing all sections.
  • the expansion and contraction of the known acoustic signal included in the mixed acoustic signal in the frequency axis direction is estimated, and the frequency of the amplitude spectrum of the known acoustic signal is determined based on the estimated expansion and contraction in the frequency axis direction.
  • Correct axial expansion and contraction for example, a section containing only a known acoustic signal in the mixed acoustic signal is specified, and the section of the frequency axis with the section of the known acoustic signal corresponding to this section is determined by: Estimate expansion and contraction in the frequency axis direction.
  • an image display step of displaying an image so that the amplitude spectrum of the mixed acoustic signal and the amplitude spectrum of the known acoustic signal can be visually recognized is further executed. You may do it.
  • a human determines a section including a known sound signal in the mixed sound signal based on the image display, and executes a correction step, a removal step, an inverse transformation step, or a synthesis step for this section.
  • a sound reproducing step of reproducing the mixed acoustic signal, the known acoustic signal, and the output signal of the synthesis step as sound may be further executed.
  • a human determines a section in which the known sound signal is included in the mixed sound signal, and in this section, a correction step, a removal step, an inverse transformation step, and a synthesis step. Perform the steps.
  • the section in which the known acoustic signal is included in the mixed acoustic signal is automatically estimated based on the amplitude spectrum of the mixed acoustic signal, and a correction step is performed for this section.
  • a removing step, an inverse transforming step, and a combining step may be executed. If the mixed sound signal contains a relatively well-known sound signal (for example, if there is a section where the known sound signal is sounding alone in the mixed sound signal), the section is automatically estimated. Can be identified, and by using automatic estimation, Removal work can be performed quickly. In the case where the existence of a known acoustic signal included in the mixed acoustic signal is not so clear, a human specifies a section.
  • the known acoustic signal elimination method of the present invention when there are a plurality of types of known acoustic signals corresponding to the acoustic signals included in the mixed acoustic signal, all of the plurality of known acoustic signals are used. , A post-removal amplitude obtained by performing a known acoustic signal conversion step and a correction step, and performing a removal step of removing all of the corrected amplitude spectra of the plurality of known acoustic signals from the amplitude spectrum of the mixed acoustic signal. The inverse transformation step and the synthesis step are performed using the spectrum. This makes it possible to remove all types of known acoustic signals from the mixed acoustic signal.
  • GUI graphic design interface
  • the processing module that performs the interface processing removes the component of the known sound signal from the mixed sound signal in which multiple sound signals are mixed, and removes the accurate sound signal of the known sound signal in the mixed sound signal. It is configured so that the position can be specified by a human.
  • the processing module that performs the in-plane processing is configured so that when the frequency characteristics of the known acoustic signal change over time in the mixed acoustic signal, these changes can be specified by humans. .
  • the processing module that performs the in-plane processing is configured so that when the volume of a known acoustic signal changes with time in a mixed acoustic signal, humans can specify those changes. You.
  • the processing module that performs the interface processing is configured such that when a known sound signal in the mixed sound signal expands or contracts in the time axis or frequency axis direction, the expansion and contraction of these can be specified by a human.
  • the processing module that performs the interface processing is configured so that a human can specify a section corresponding to the mixed sound signal and the known sound signal.
  • the known acoustic signal elimination device further includes a mixed acoustic signal conversion unit that converts the mixed acoustic signal into a time-frequency expression to obtain a wide spectrum of the mixed acoustic signal and a phase of the mixed acoustic signal.
  • the temporal position of the amplitude spectrum of the known acoustic signal with respect to the amplitude spectrum of the mixed acoustic signal is shifted, the frequency characteristic changes over time, the volume changes over time, the expansion and contraction in the time axis direction, and the frequency axis direction Correction means for obtaining a corrected amplitude spectrum of a known sound signal in which at least one of expansion and contraction of the known sound signal is corrected, and removing a corrected amplitude spectrum of the known sound signal from the amplitude spectrum of the mixed sound signal
  • Removing means for performing a reverse conversion to a time expression based on the amplitude spectrum after removal obtained by the removing means and the phase of the mixed acoustic signal to obtain a unit waveform, and synthesizing the unit waveform to obtain a known unit waveform
  • the correction means includes a shift in the temporal position of the amplitude spectrum of the known sound signal with respect to the amplitude spectrum of the mixed sound signal, a time change of the frequency characteristic, a time change of the volume, expansion and contraction in the time axis direction, and frequency.
  • Provide a processing module that performs interface processing that allows humans to manually specify at least one correction of axial expansion and contraction.
  • the processing module that performs the interface processing includes an image display unit that displays an image so that the amplitude spectrum of the mixed sound signal and the amplitude spectrum of the known sound signal can be visually compared, a mixed sound signal, and a known sound signal. And a sound reproducing unit for generating an output signal of the synthesizing means as sound.
  • the amplitude spectrum of the mixed sound signal and the amplitude spectrum of the known sound signal displayed on the image display unit can be displayed from the image display unit and the sound reproduction unit.
  • the human being can not only specify the section of the known acoustic signal included in the mixed acoustic signal, but also manually specify the section of the amplitude spectrum of the known acoustic signal in this section manually.
  • Displacement At least one correction of the time change of the frequency characteristic, the time change of the volume, the expansion and contraction in the time axis direction, and the expansion and contraction in the frequency axis direction can be specified.
  • the known acoustic signal can be removed with high removal accuracy.
  • the image display unit displays the amplitude spectrum of the section in the mixed sound signal containing the known sound signal, the displacement of the amplitude spectrum of the section corresponding to the known sound signal with respect to time, and the frequency characteristics. It is configured to be able to display the corrected amplitude spectrum, which is corrected for at least one of the time change of the volume, the time change of the sound volume, the expansion and contraction in the time axis direction, and the expansion and contraction in the frequency axis direction, on the time axis in alignment with each other. Is desirable.
  • the state of the corrected amplitude spectrum can be visually checked, and it is possible to estimate how the corrected spectrum can be improved in removal accuracy by looking at the image. Removal work is faster.
  • the image display unit is configured to be able to display an image of the amplitude spectrum of the sound signal obtained by removing the corrected amplitude spectrum from the amplitude spectrum of the mixed sound signal.
  • the known acoustic signal elimination program further comprises: a mixed acoustic signal conversion step of converting the mixed acoustic signal into a time frequency signal to obtain an amplitude spectrum of the mixed acoustic signal and a phase of the mixed acoustic signal.
  • Removal step to remove positive amplitude spectrum an inverse conversion step of performing an inverse conversion to a time expression based on the amplitude spectrum after removal obtained in the removal step and the phase of the mixed acoustic signal to obtain a unit waveform, and synthesizing the unit waveform to obtain a known acoustic signal. And a synthesizing step of obtaining an acoustic signal from which the component has been removed.
  • the correction step includes a step of shifting the temporal position of the amplitude spectrum of the known acoustic signal with respect to the amplitude spectrum of the mixed acoustic signal, a temporal change in the frequency characteristic, and a sound volume.
  • the known acoustic signal in which at least one of the time change, expansion and contraction in the time axis direction and expansion and contraction in the frequency axis direction is corrected, and the corrected amplitude spectrum is used as the amplitude spectrum of the mixed acoustic signal. Since it is removed from the vector, there is an advantage that the known acoustic signal included as non-stationary noise in the mixed acoustic signal can be removed with high accuracy.
  • the sound signal elimination method of the present invention for example, when a sound signal of a TV program or a movie in which BGM is sounding in the background of a human voice or sound is input, the sound signal is separately prepared; It is possible to remove background music in the program using audio signals, and obtain audio signals only of human voices and body sounds. Further, by adding another music as BGM to the sound signal after the BGM removal, it is possible to reuse the music such as a TV program or a movie by replacing the music.
  • the known sound signal here may be any sound signal, it can be applied regardless of the genre of the music, regardless of the presence or absence of vocals, and regardless of the presence or absence of accompaniment. In addition to music, it can be applied to any known noise including stationary noise and non-stationary noise.
  • FIG. 1 shows a configuration of an example of an embodiment of a known acoustic signal elimination device of the present invention. It is a block diagram shown.
  • FIG. 2 is a block diagram showing steps when the known acoustic signal elimination method of the present invention is carried out.
  • FIG. 3 is a flowchart showing an example of an algorithm of a program used when a main part of the known acoustic signal elimination device of the present invention is realized by using a computer.
  • FIG. 4 is a flowchart showing detailed processing in step ST103 of FIG.
  • FIG. 5 is a flowchart showing the details of steps in the case of performing an estimation operation in both estimation involving humans and automatic estimation.
  • FIG. 6 is a diagram showing a screen configuration of an editor interface.
  • FIG. 7 is a diagram showing a time change of the power of the mixed acoustic signal.
  • FIG. 8 is a diagram showing a time change of the amplitude spectrum of the mixed acoustic signal.
  • FIG. 9 is a diagram showing a temporal change in power of a known acoustic signal of a sound source that is a source of BGM.
  • FIG. 10 is a diagram showing a time change of the amplitude spectrum of a known acoustic signal of a sound source that is a source of BGM.
  • FIG. 11 is a diagram showing a temporal change in power of a desired sound signal after removing a known sound signal.
  • FIG. 12 is a diagram showing a time change of the amplitude spectrum of a desired sound signal after removing the known sound signal.
  • FIG. 1 is a block diagram showing a configuration of an embodiment of a known acoustic signal elimination device for performing the known acoustic signal elimination method of the present invention.
  • the known acoustic signal elimination device has a mixed acoustic signal converter It comprises a stage 1, a known acoustic signal conversion means 2, a correction means 3, an interface 4, a removal means 5, an inverse conversion means 6, and a synthesis means 7.
  • the mixed sound signal converting means 1 is a mixed sound signal m (t) in which a sound signal b (t) such as BGM is mixed with a sound signal s (t) (t is a time axis) such as a desired sound or a body sound. (At this point, s (t) and b (t) are unknown, and only m (t) is input), converted to a time-frequency representation and mixed with the amplitude spectrum M ( ⁇ , t) of the mixed acoustic signal Find the sound signal phase 0m ( ⁇ , t).
  • the known sound signal conversion means 2 converts the known sound signal b '(t) of the sound source, which is the source of the sound signal b (t) to be removed, into a time-frequency expression, and the amplitude spectrum ⁇ , ( ⁇ , t).
  • the correction means 3 calculates the amplitude spectrum ⁇ , ( ⁇ , t) of the known acoustic signal with respect to the amplitude spectrum M ( ⁇ , t) of the mixed acoustic signal. )), The corrected amplitude spectrum B ( ⁇ , t) of the known acoustic signal in which the time shift of the frequency characteristic, the time change of the frequency characteristic, the time change of the sound volume, the expansion and contraction in the time axis direction and the expansion and contraction in the frequency axis direction are corrected.
  • Correction means for automatically estimating and correcting all of displacement, time change of frequency characteristics, time change of sound volume, expansion and contraction in the time axis direction and expansion and contraction in the frequency axis direction. 3 can be configured.
  • the correction means 3 corrects all of the positional deviation in time, the time change of the frequency characteristic, the time change of the volume, the expansion and contraction in the time axis direction and the expansion and contraction in the frequency axis direction. It is configured so that humans can specify it manually by using.
  • the input source 4 has an image display unit that displays an image so that the amplitude spectrum of the mixed sound signal and the amplitude spectrum of the known sound signal can be visually compared. It is a processing module that performs interface processing using the graphic design interface (GU I).
  • GUI graphic design interface
  • Interface 4 uses the input section displayed on the screen to display the amplitude of the mixed acoustic signal. Based on the spectrum and the amplitude spectrum of the known sound signal, the section of the known sound signal included in the mixed sound signal can be designated by a human and the above-mentioned correction can be designated.
  • the removing unit 5 removes the corrected amplitude spectrum B ( ⁇ , t) of the known sound signal from the amplitude spectrum M ( ⁇ , t) of the mixed acoustic signal. Then, the inverse transform means 6 performs an inverse transform to a time expression based on the amplitude spectrum S ( ⁇ , t) after removal obtained by the remover 5 and the phase 6> m (w, t) of the mixed acoustic signal. Find the unit waveform s' (t).
  • the synthesizing means 7 synthesizes the unit waveform s ′ (t) output from the inverse transform means 6 to obtain an audio signal s (t) from which the components of the known audio signal have been removed.
  • the interface 4 displays the post-removal amplitude spectrum S ( ⁇ , t) output from the removing unit 5 on the image display unit (see FIG. 6). Further, the interface 4 has a built-in sound reproducing unit, and reproduces a mixed sound signal, a known sound signal, and a synthesized sound signal output from the synthesizing means 7.
  • FIG. 2 is a block diagram showing steps when the known acoustic signal elimination method of the present invention is carried out.
  • FIG. 3 is a diagram showing a main part of the known acoustic signal elimination apparatus of the present invention realized by using a computer.
  • 9 is a flowchart showing an example of a program algorithm used in the case.
  • FIG. 4 is a D-chart showing the detailed processing in step ST103 of FIG. Fig. 5 shows both human-related estimation and automatic estimation. It is a flowchart which shows the detail of a step in performing an estimation process. The operation of removing a known acoustic signal in the method and apparatus for removing a known acoustic signal of the present invention will be described below with reference to FIGS. 1 to 5.
  • the known acoustic signal b, (t) often undergoes the following deformation in the mixed sound m (t), so the component corresponding to b (t) is corrected by correction. Is estimated.
  • the objects of the correction are mainly a time shift, a frequency characteristic change over time, a volume change over time, and expansion or contraction in the time axis or frequency axis direction, as described below.
  • the position at which the known sound signal b 5 (t) is sounding in the mixed sound m (t) is not always from the beginning. Therefore, the known acoustic signal b '(t) is shifted in the time axis direction, It is necessary to subtract the known sound signal from the mixed sound by adjusting the relative position of the person.
  • the frequency characteristics often change due to the influence of the graphic equalizer and the like. For example, low and high frequencies may be emphasized and attenuated. Therefore, it is necessary to correct the frequency characteristic of b '(t) by changing it in the same way, and to subtract the known sound signal from the mixed sound.
  • step ST1 the mixed sound signal is Fourier-transformed to obtain the phase of the mixed sound signal (step ST2).
  • step ST3 the amplitude spectrum (step ST3) of the mixed acoustic signal (step ST3), and a known acoustic signal corresponding to the acoustic signal included in the mixed acoustic signal is extracted in step ST4.
  • Fourier transform is performed to determine the amplitude spectrum of the known sound signal (step ST5) (known sound signal conversion step).
  • step ST6 based on the amplitude spectrum of the mixed acoustic signal, the temporal position shift of the amplitude spectrum of the known acoustic signal with respect to the amplitude spectrum of the mixed acoustic signal, the time change of the frequency characteristic, and the volume.
  • step ST7 a corrected amplitude spectrum of the known acoustic signal in which at least one of the time change, expansion and contraction in the time axis direction, and expansion and contraction in the frequency axis direction is corrected is obtained (correction step).
  • step ST8 the corrected amplitude spectrum of the known acoustic signal is removed from the amplitude spectrum of the mixed acoustic signal to obtain a post-removal amplitude spectrum (step ST9) (removal step).
  • step ST10 a unit waveform is obtained by performing an inverse Fourier transform on the basis of the amplitude spectrum after removal obtained in the removing step and the phase of the mixed acoustic signal (inverse transforming step).
  • step ST11 the unit waveform is synthesized by the overlapped quadrature method to obtain an acoustic signal from which the components of the known acoustic signal have been removed (synthesizing step). As shown in the flowchart of FIG.
  • step ST 101 the mixed acoustic signal is subjected to Fourier transform to obtain the mixed acoustic signal. Find the amplitude spectrum and the phase of the mixed acoustic signal.
  • step ST102 a known acoustic signal corresponding to the acoustic signal included in the mixed acoustic signal is Fourier-transformed to obtain an amplitude spectrum of the known acoustic signal.
  • the temporal position shift of the amplitude spectrum of the known acoustic signal with respect to the amplitude spectrum of the mixed acoustic signal, and the time of the frequency characteristic Determine the corrected amplitude spectrum of the known sound signal by correcting at least one of the change, the volume change over time, the expansion and contraction in the time axis direction, and the expansion and contraction in the frequency axis direction.
  • step ST104 the corrected amplitude spectrum of the known sound signal is removed from the amplitude spectrum of the mixed sound signal, and the post-removal amplitude spectrum is obtained.
  • step ST105 the post-removal amplitude spectrum obtained in step ST104 is obtained.
  • a unit waveform is obtained by performing an inverse Fourier transform on the basis of the phase of the mixed sound signal and the mixed sound signal.
  • step ST106 the sound signal obtained by combining the unit waveforms by the overlap-add method to remove the components of the known sound signal Get.
  • step ST107 a determination is made as to whether or not the user has evaluated the removed audio signal as being satisfactory. If the determination result is unsatisfactory, the correction is performed again in step ST103. It is. Until the user is satisfied, steps ST103 through ST107 are repeated.
  • the subtraction process is performed on the amplitude spectrum in the time frequency domain without performing the subtraction process on the waveform in the time domain.
  • S TFT short-time Fourier transform
  • the audio signal is A / D-converted at a sampling frequency of 44.lkHz and a quantization bit rate of 16 bits, and a short-time Fourier transform using a 8192-point Haning window as the window function h (t) is performed.
  • the frames of the fast Fourier transform (FFT) are shifted by 441 points, so the frame shift time (one frame shift) is 10 ms. This frame shift is used as the processing time unit.
  • the amplitude spectrum S ( ⁇ , t) of the desired sound signal s (t) after the removal of the known sound signal is obtained from the amplitude spectrum M ( ⁇ , t), ⁇ ′ ( ⁇ , t) by the following equation.
  • B ( ⁇ , t) is the amplitude spectrum after correcting ⁇ , ( ⁇ , t).
  • a (t) is a function of an arbitrary shape for finally adjusting the amount of subtracting a component corresponding to the amplitude spectrum of the known sound signal from the amplitude spectrum of the mixed sound.
  • a (t) 1 The greater this is, the greater the amount of subtraction It will be good.
  • g ( ⁇ , t) is a function for correcting the time change of the frequency characteristic and the time change of the volume.
  • P ( ⁇ ) is a function for correcting expansion and contraction in the frequency axis direction.
  • the frequency axis ⁇ of the amplitude spectrum ⁇ , ( ⁇ , t)
  • ⁇ '( ⁇ , t) takes 0 outside the original domain of ⁇ , and interpolates as appropriate when discretized and implemented.
  • q (t) is a function for correcting the expansion and contraction in the time axis direction.By converting the time axis t of the amplitude spectrum ⁇ '( ⁇ , t), linear (non-linear) Enables a type of expansion and contraction. Note that ⁇ , ( ⁇ , t) takes 0 outside the original domain of t, and interpolates as appropriate when implementing discretely.
  • r (t) is a function for correcting the time position shift, and usually corrects a certain amount of shift by setting a constant. If the shift width changes over time, set a function to correct the width at each time. Note that ⁇ '( ⁇ , t) takes 0 outside the domain of the original t, and interpolates as appropriate when discretized and implemented. Although it is possible to express it as a single function integrating q (t) and (t), here q (t) is set to represent continuous expansion and contraction, and r (t) is a discrete position It is set for the purpose of expressing the deviation of.
  • the various parameter overnight functions a (t), g ( ⁇ , t) (g w ( ⁇ , ⁇ ) of the equations (11), (12) and (13) are used.
  • t), g t (t), g r (t)), p ( ⁇ ), q (t), r (t), c ( ⁇ , t) May be set manually. Alternatively, it may be corrected by a human after the automatic estimation. In the following, a description will be given of a specific automatic estimation method and a case of using the interface 4 in the known acoustic signal elimination apparatus which enables manual correction by a human.
  • step ST201 the various parameter functions g ( ⁇ , t) (g M ( ⁇ , t), g t (t)), p ( ⁇ ) in equations (1 1), (1 2) and (13) , q (t "), r (t)
  • step ST201 the set ⁇ of the BGM interval ⁇ is specified and automatically estimated.
  • P ( ⁇ ) in the step ST 2 0 2 performs automatic estimation of q (t), gw in step ST 2 03 ( ⁇ , t) , g t (t), for automatic estimation of r (t).
  • step ST 205 the correction operation is performed using interface 4.
  • BGM section a section containing almost no sound signal s (t) of only human voices or body sounds.
  • BGM section a section containing almost no sound signal s (t) of only human voices or body sounds.
  • a plurality of BGM sections may be used.
  • g w ( ⁇ , t) is estimated by interpolation (interpolation or extrapolation). (If there are BGM sections on both sides, interpolation is performed from both sides.) Finally, g w ( ⁇ , t) is smoothed in the frequency axis direction. Note that the smoothing width can be set arbitrarily. It is not necessary.
  • the amplitude spectrum M ( ⁇ , t) is compared with the amplitude at each time of g w ( ⁇ , t) ⁇ and ( ⁇ , t) after frequency characteristic correction.
  • represents the set of. Any division can be applied to ⁇ . For example, it is good to divide every equal octave of equal temperament used in music (divide at equal intervals on the logarithmic frequency axis).
  • g t (t) is min (g :, t (, t)) or Estimate by In the case of min (:, t (, t)), the amplitudes are compared in the frequency band where M ( ⁇ , t) and g w ( ⁇ , t) ⁇ '( ⁇ , t) are closest.
  • g t (t) is smoothed in the time axis direction.
  • the smoothing width can be arbitrarily set, and the smoothing need not be performed.
  • P ( ⁇ ) and q (t) are set so that the distance between M ( ⁇ , t) and ⁇ ( ⁇ , t) (for example, the logarithmic spectral distance) is minimized.
  • r (t) In estimating r (t), in principle, the set ⁇ of BGM sections ⁇ is used, and the time axis of the correspondence between M ( ⁇ , t) and ⁇ ( ⁇ , t) in those sections is adjusted. R (t) is obtained. r (t) is often a constant, but if a section of the known sound signal b 3 (t) is not used, but is mixed while being used intermittently, skip that section. Thus, r (t) is a discontinuous function.
  • FIG. 5 is a flowchart showing a software algorithm of a program corresponding to a case where a human specifies manually and a case where an automatic estimation is performed.
  • steps ST 302 to ST 3 13 in FIG. 5 are executed.
  • basically, a set of the remaining BGM sections is obtained by using one BGM section ⁇ 1 as a clue.
  • the first ⁇ 1 is determined manually by a human or by finely dividing the time axis of the audio signal and determining the correspondence between these short divided sections. If not manually specified by a human, B ( ⁇ , t) is temporarily calculated (step ST 302), and the amplitude spectrum of the time window obtained by dividing M ( ⁇ , t) and B ( ⁇ , t) into small pieces is calculated. Is calculated (corresponding to the degree of similarity) (step ST303). Then, the correspondence of the time window of the minimum distance is checked (step ST 304), and the section including the result is set to ⁇ ⁇ ⁇ 1 as an initial value (step ST 305).
  • step ST 306 various parameter overnight functions of B ( ⁇ , t) are estimated (step ST 306 to step S 309 309), and ⁇ ( ⁇ , t) is calculated (step ST310). ).
  • step ST310 Check whether the estimated values for each parameter have converged, and if not converged, the distance between the amplitude spectra of M (, t) and B ( ⁇ , t) for the entire interval of (Equivalent to similarity).
  • a constant multiple of the maximum value (or the average value) is set as a BGM section determination threshold value (step ST312).
  • a section having a distance equal to or less than the BGM section determination threshold is detected and newly added to ⁇ (step ST313).
  • an upper limit can be set for addition.
  • is updated, and various parameter overnight functions are determined appropriately.
  • the distance between M ( ⁇ , t) and ⁇ ( ⁇ , t) is, for example, the root mean square log spectrum distance. Is valid.
  • This Eddy evening is roughly divided and mixed A sub-window W 3_ for operating the acoustic signal m (t), a sub-window W 2 for operating the known acoustic signal b ′ (t), and a sub-window W 3 for operating the desired acoustic signal s (t) after removing the known acoustic signal. It consists of three subwindows.
  • the known acoustic signals b' (t) operated in the sub-window W2 can be switched by the changeover switch W2S. In this evening festival, steps ST 219 to ST 219 shown in FIG. 4 are executed.
  • the operation range slider P1 indicates where in the sound signal is currently displayed.
  • Cursor P2 indicates the current position of the operation target on the time axis.
  • the button P3 When the button P3 is pressed, the subwindow to which the button belongs is temporarily collapsed and becomes smaller. Unused sub-windows other than the one currently being operated can be hidden to effectively use the narrow screen. Pressing the float (enlarge) button P4 temporarily disconnects the subwindow to which the button belongs from the parent window (float), and further enlarges it to facilitate operation and editing. If only the button P 4 is drawn, press this button to float the sub window associated with it and make it appear again.
  • a graph E1 of the power of the mixed acoustic signal m (t) and a graph E2 of its amplitude spectrum M ( ⁇ , t) are displayed.
  • a graph E3 of the power of the known acoustic signal b, (t) and a graph E4 of the amplitude spectrum ⁇ , ( ⁇ , t) are displayed.
  • a graph E5 of the power of the acoustic signal s (t) after the removal of the known acoustic signal and a graph E6 of the amplitude spectrum S ( ⁇ , t) thereof are displayed.
  • the amplitude is drawn in shades on the left (the horizontal axis is the time axis, the vertical axis is the frequency axis), and the amplitude at the force-sol position is drawn on the right (The horizontal axis is power and the vertical axis is frequency axis).
  • a group of buttons capable of playing, stopping, fast-forwarding, and fast-forwarding the mixed acoustic signal are arranged for human listening and confirmation.
  • the interface 4 reproduces the mixed acoustic signal by the built-in acoustic reproducing unit.
  • the sub-window W2 for the operation of the known sound signal b '(t) is the main window of the operation, and all the parameter overnight functions a (t), g in Equations (1 2) and (13) -The shape of ( ⁇ , t) (g w ( ⁇ , t), g t (t), g r (t)), p ( ⁇ ), q (t), r (t) can be set freely .
  • the following is a description of each operation panel.
  • Operation panel C 1 (right side of E7) g w ( ⁇ 3 t) for correcting the time change of the frequency characteristics
  • This panel is used to display and operate g a ( ⁇ , t) is drawn (the horizontal axis is the size, the vertical axis is the frequency axis).
  • the result of the setting operation is immediately reflected on the display panel E7 of g ( ⁇ , t) (steps ST205 and ST206).
  • the magnitude of the value of g ( ⁇ , t) is drawn in shades (the horizontal axis is the time axis, and the vertical axis is the frequency axis).
  • g t (t) is displayed on the operation panel.
  • the result of the setting operation is immediately reflected on the display panel E 7 of g ( ⁇ , t) (steps ST 207 and ST 208).
  • Operation panel C 3 (lower side of E 7) for raising the value of g ( ⁇ , t) as a whole
  • g r (t) is displayed.
  • the result of the setting operation is immediately reflected on the display panel E 7 of g ( ⁇ , t) (steps ST209 and ST210).
  • Display P ( ⁇ ) 'It is a panel for operation. When this panel is operated, the change in ⁇ ( ⁇ ) is immediately reflected on the display (steps ST213 and ST214).
  • q (t) is displayed. This is the operation panel. When this panel is operated, the change of q (t) is immediately reflected on the display (step S ⁇ 2 15, ST 2 16).
  • the change in r (t) is immediately reflected on the display (steps ST 2 17 and ST 2 18).
  • a group of buttons capable of playing, stopping, fast-forwarding, and fast-returning a known sound signal are arranged for human listening and confirmation.
  • the interface 4 reproduces a known acoustic signal by a built-in acoustic reproduction unit.
  • volume feeder operation panel C 9 (below E 8)
  • This panel is used to display and operate the shape of c ( ⁇ , t) in the t direction.
  • the result of the setting operation is immediately reflected on the display panel E8 of c ( ⁇ , t).
  • a group of buttons capable of playing, stopping, fast-forwarding, and rewinding the synthesized audio signal are arranged side by side for human listening and confirmation. I have.
  • the interface 4 reproduces the audio signal synthesized by the built-in audio reproduction unit.
  • a program that can determine the unknown s (t) under the condition that the sound signals b and (t) of the sound source are known is provided by various operating systems (Liux 2.4, SG IIRIX 6.5 , Microsoft Windows XP: registered trademark).
  • an audio file containing m (t) and b '(t) is provided to this program.
  • BGM background music
  • FIGS. 7 to 12 show the results of actually processing a mixed sound of classical music being played in the background music of a dialogue between two men and women.
  • the mixed sound signal in (t) shown in FIGS. 7 and 8 as input and the known sound signal b '(t) of the original sound source shown in FIGS. 9 and 10; removing the BGM component
  • the result is the sound signal s (t) after the removal of the known sound signal shown in FIGS. 11 and 12.
  • the mixed sound in the example of this processing result is the audio signal of the dialogue between two men and women extracted from “: RWCP Voice Dialogue Day” and the audio signal extracted from “RWC Research Music Day”.
  • the sound signal of Lasik music is added.
  • the correction step the temporal position shift of the amplitude spectrum of the known acoustic signal with respect to the amplitude spectrum of the mixed acoustic signal, the time change of the frequency characteristic, A corrected amplitude spectrum of a known sound signal in which at least one of volume change over time, expansion and contraction in the time axis direction, and expansion and contraction in the frequency axis direction is corrected is obtained, and the corrected amplitude spectrum is used as the amplitude spectrum of the mixed acoustic signal. Therefore, the known acoustic signal included as non-stationary noise in the mixed acoustic signal can be removed with high accuracy.
  • an audio signal from a TV program or movie with BGM is input in the background of human voice or sound
  • the BGM in the program is removed using the music audio signal of BGM prepared separately, and human music is removed. It is possible to obtain an acoustic signal consisting of only voices and body sounds.
  • the known sound signal may be any sound signal, it can be applied regardless of the genre of the music, regardless of the presence or absence of vocals, and regardless of the presence of accompaniment. Also, the present invention is applicable not only to music but also to any known noise including stationary noise and non-stationary noise.

Landscapes

  • Engineering & Computer Science (AREA)
  • Computational Linguistics (AREA)
  • Quality & Reliability (AREA)
  • Signal Processing (AREA)
  • Health & Medical Sciences (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Human Computer Interaction (AREA)
  • Physics & Mathematics (AREA)
  • Acoustics & Sound (AREA)
  • Multimedia (AREA)
  • Stereophonic System (AREA)
  • Auxiliary Devices For Music (AREA)
  • Measurement Of Mechanical Vibrations Or Ultrasonic Waves (AREA)

Abstract

 複数の音響信号が混合された音響信号を入力とし、そのうち一つの音響信号に類似した既知音響信号が与えられたときに、その既知音響信号を除去することを可能にする既知音響信号除去装置が提供される。この既知音響信号除去装置は、入力混合音響信号m(t)と既知音響信号b'(t)をそれぞれ時間周波数領域での振幅スペクトルM(ω,t),B'(ω,t)に変換し、M(ω,t)中のB'(ω,t)に対応する成分を減算して除去することにより、除去後の振幅スペクトルS(ω,t)を得る。その際、M(ω,t)中のB'(ω,t)に対応する成分は、時間的な位置のずれ、周波数特性の時間変化、音量の時間変化等の要因により変形しているため、それらを補正したB(ω,t)を減算する。最後に、m(t)の位相と、S(ω,t)とを用いて時間領域に逆変換し、所望の除去後の音響信号s(t)を得る。

Description

明細書 既知音響信号除去方法及び装置 技術分野
本発明は、 複数の音響信号が混合された混合音響信号の中から、 既知の 音響信号の成分を除去する既知音響信号除去方法および既知音響信号除去装置 に関するものである。 背景技術
従来から、 音響信号処理として、 スペクトルサブトラクシヨン法 (非特 許文献 1) と呼ばれる方法が知られている。 従来のスペクトルサブトラクショ ン法は、 定常雑音 (スペクトルが時間的に変化せず、 周波数特性や音量等がほ ぼ一定な雑音) と所望の音 (ターゲット音) が混合された音響信号 (混合音) から定常雑音を除去してターゲット音を得る方法である。 この方法では、 事前 に定常なスぺクトルの平均を求める等の簡易な方法で定常雑音のスぺクトルを 学習しておき、 入力された混合音のスぺクトルから定常雑音のスぺクトルを引 き去る処理を行う。 つまり、 雑音の平均を引き去る処理を行う。
一般に、 音響信号除去に関しては、 複数のマイクロホンからの入力を用 いる方法が多数提案されている。 また、 スペクトルサブトラクシヨン法には、 特許文献 1〜 7に開示されているように、 様々な改良がなされている。
[非特許文献 1 ]
St e en Bo l l, "Suppr e s s i on oェ Ac ous t i c No i s e in Spe e ch Us i ng Spe c t ral S ubt rac t i on" , IEEE Trans ac t i ons on Ac o u s t i c s , S e e c , and S i gnal Pr o c e s s i n g, Vo l. AS S P- 27 , No. 2 , Apr i l 1979. [特許文献 1 ] 特開 2 0 0 2一 1 7 5 0 9 9号公報
[特許文献 2 ] 特開 2 0 0 2一 0 1 4 6 9 4号公報
[特許文献 3 ] 特開 2 0 0 1 - 2 2 8 8 9 2号公報
[特許文献 4 ] 特開 2 0 0 1一 2 1 5 9 9 2号公報
[特許文献 5 ] 特開平 1 1― 0 0 3 0 9 号公報
[特許文献 6 ] 特開平 1 0— 2 4 0 2 9 号公報
[特許文献 7 ] 特開平 0 8— 2 2 1 0 9 2号公報 発明の開示
従来におけるスペクトルサブトラクシヨン法は、 定常雑音を前提として おり、 非定常雑音 (スペクトルが時間的に大きく変化し、 周波数特性や音量等 も変化する雑音) には適用できなかった。 例えば、 バックグラウンドミュ一ジ ック (B G M ) として使用されている音楽のような時間的に大きく変化する非 定常雑音を除去することは不可能であった。 これは、 非定常雑音のスペクトル の変化が大きすぎて学習ができないからである。
また、 仮に、 従来の方法により非定常雑音が事前に与えられた条件を扱 おうとしても、 非定常雑音の周波数特性、 音量、 振幅スペクトルの時間軸方向 の伸縮及び周波数軸方向の伸縮等の変化の影響で、 引き去る処理を適切に行う ことはできなかった。 複数のマイクロホンからの入力を用いる方法では、 モノ ラル音響信号には適用することができなかった。 改良された従来のスぺクトル サブトラクション法のいずれの方法も、 主に音声認識の前処理を目的としてい るため、 非定常雑音が事前に与えられ、 その非定常雑音を除去する用途には利 用できなかった。 したがって、 本発明の目的は、 複数の音響信号が混合された混合音響信 号の中から、 既知の音響信号 (非定常でも定常でもよい) の成分を、 それに対 応する元音源からの既知音響信号を用いて除去することができる既知音響信号 除去方法及び既知音響信号除去装置並びに該装置に用いるプログラムを提供す しとに る。
また、 本発明の他の目的は、 例えば、 既知の音響信号が音楽であり、 そ の音楽音響信号が、 人間の音声や物音に対するバックグラウンドミュージック ( B G M) として使用されている混合音から、 既知の音響信号に対応する元音 源である既知音響信号 (例えば C Dやレコード等から同一音楽の音響信号を別 途入手したもの) を用いて B G Mを除去することができる既知音響信号除去方 法及ぴ装置並びに該装置に用いるプログラムを提供することにある。
本発明の更に別の目的は、 複数の音響信号が混合された音響信号 (混合 音) の中から、 既知音響信号の成分を除去する際に、 混合音中での既知音響信 号の正確な位置を自動推定し、 その位置の既知の音響信号を除去することがで きる既知音響信号除去方法及び装置並びに該装置に用いるプログラムを提供す ることにある。
本発明の更に別の目的は、 複数の音響信号が混合された音響信号 (混合 音) の中から、 既知音響信号の成分を除去する際に、 混合音中での既知音響信 号の正確な位置を人間が指定できるイン夕フエ一スを備えた既知音響信号除去 装置を提供することにある。
本発明の更に別の目的は、 複数の音響信号が混合された音響信号 (混合 音) の中から、 既知音響信号の成分を除去する際に、 混合音中で既知音響信号 の周波数特性や音量が時間的に変化しているときに、 それらの変化を自動推定 して補正しながら除去することができる既知音響信号除去方法及び装置並びに 該装置に用いるプログラムを提供することにある。 ' -' 本発明の更に別の目的は、 複数の音響信号が混合された音響信号 (混合 音) の中から、 既知音響信号の成分を除去する際に、 混合音中で既知音響信号 の周波数特性や音量が時間的に変化しているときに、 それらの変化を人間が指 定することができるィン夕フェースを備えた既知音響信号除去装置を提供する ことにある。 本発明の更に別の目的は、 複数の音響信号が混合された音響信号 (混合 音) の中から、 既知音響信号の成分を除去する際に、 混合音中で既知の音響信 号が時間軸あるいは周波数軸方向に伸縮しているときに、 それらの伸縮を自動 推定して補正しながら除去することができる既知音響信号除去方法及び装置並 びに該装置に用いるプログラムを提供することにある。
本発明の更に別の目的は、 複数の音響信号が混合された音響信号 (混合 音) の中から、 既知音響信号の成分を除去する際に、 混合音中で既知の音響信 号が時間軸あるいは周波数軸方向に伸縮しているときに、 それらの伸縮を人間 が指定することができるイン夕フェースを備えた既知音響信号除去装置を提供 することにある。
本発明の更に別の目的は、 複数の音響信号が混合された音響信号の中か ら、 複数の既知音響信号の成分を除去する際に、 既知の音響信号を一つずつ繰 り返し除去できるようにした既知音響信号除去方法及び装置並びに該装置に用 いるプログラムを提供することにある。 本発明による既知音響信号除去方法においては、 複数の音響信号が混合 された混合音響信号から、 既知の音響信号 (非定常でも定常でもよい) の成分 を、 それに対応する元音源からの既知音響信号を用いて除去する。
このため、 本発明の既知音響信号除去方法では、 まず、 混合音響信号を 時間周波数表現に変換して混合音響信号の振幅スぺクトルと混合音響信号の位 相とを求める (混合音響信号変換ステップ) 。 ここでの音響信号を時間周波数 表現に変換する方法としては、 フ一リエ変換やウェーブレヅト変換など公知の 変換方法を用いる。
次に、 混合音響信号中に含まれている既知の音響信号に対応 (類似) し ている既知音響信号 (C Dやレコード等から同一音楽の音響信号を別途入手し たもの) を時間周波数表現に変換して既知音響信号の振幅スぺクトルを求める (既知音響信号変換ステップ) 。 続いて、 求めた混合音響信号の振幅スペクトルに基づいて、 混合音響信 号の振幅スぺクトルに対する既知音響信号の振幅スぺクトルの時間的な位置の ずれ、 周波数特性の時間変化、 音量の時間変化、 時間軸方向の伸縮及び周波数 軸方向の伸縮の少なくとも 1つを補正した前記既知音響信号の補正振幅スぺク トルを求める (補正ステヅプ) 。
次に、 混合音響信号の振幅スぺクトルから既知音響信号の補正振幅スぺ クトルを除去する (除去ステップ) 。 この除去ステップにより得た除去後振幅 スぺクトルと混合音響信号の位相とに基づいて時間表現に逆変換を行って単位 波形を求める (逆変換ステップ) 。
最後に、 単位波形をオーバ一ラップ ' ァド法等の合成方法を用いて合成 して既知音響信号の成分を除去した音響信号を得る (合成ステップ) 。 また、 本発明の既知音響信号除去方法においては、 次のような補正ステ ヅプを実行することにより、 混合音響信号の振幅スぺクトルに対する既知音響 信号の捩幅スペクトルの時間的な位置のずれ、 周波数特性の時間変化、 音量の 時間変化、 時間軸方向の伸縮及ぴ周波数軸方向の伸縮の少なくとも 1つを補正 した既知音響信号の補正振幅スぺクトルを求め、 この補正振幅スぺクトルを混 合音響信号の振幅スペクトルから除去する。 このため、 混合音響信号中に非定 常雑音として含まれている既知音響信号を高い精度で除去することができる。
原則的には、 既知音響信号の振幅スペク トルの時間的な位置のずれ、 周 波数特性の時間変化、 音量の時間変化、 時間軸方向の伸縮及び周波数軸方向の 伸縮め中そ'、 実際に混合音響信号中でその現象または変化が起きていたものを 全て補正するのが好ましい。
しかしながら、 何も補正しない場合よりも、 実際に混合音響信号中でそ の現象または変化が起きているものの 1つでも補正すれば、 既知音響信号の除 去精度を高めることができるので、 補正のすべてを行わなくてもよい。 もちろ ん必要な補正のすべてを行ってもよい。 補正ステップの実行では、 例えば、 混合音響信号に含まれる既知音響信 号の時間的な位置を推定し、 推定した時間的な位置に基づいて既知音響信号の 振幅スペクトルの時間的な位置のずれを補正する。 推定方法は、 例えば、 混合 音響信号の振幅スぺクトルの所定の区間と既知音響信号の振幅スぺクトルの所 定の区間の距離 (類似度) を求め、 距離が最も近い区間を混合音響信号に含ま れる既知音響信号の時間的な位置と推定する。
また、 補正ステップの実行では、 例えば、 混合音響信号に含まれる既知 音響信号の周波数特性の変化を推定し、 推定した周波数特性の時間変化に基づ いて既知音響信号の振幅スぺクトルの周波数特性の時間変化を補正する。 周波 数特性の変化の推定は、 例えば、 混合音響信号中の既知の音響信号だけが含ま れている区間を特定し、 この区間の周波数特性とこの区間に対応する既知音響 信号の周波数特性との対比から、 混合音響信号に含まれる既知音響信号の周波 数特性の変化を推定する。
また、 補正ステップの実行では、 例えば、 混合音響信号に含まれる既知 音響信号の音量の時間変化を推定し、 推定した音量の時間変化に基づいて既知 音響信号の振幅スぺクトルの音量の時間変化を補正する。 音量の時間変化の推 定は、 周波数特性の補正を行った後に、 例えば、 混合音響信号に含まれる既知 音響信号に相当する振幅を持つ周波数帯域を各時刻において特定し、 その周波 数帯域における混合音響信号の振幅と既知音響信号の振幅との対比から推定す る。
また、 補正ステップの'実行では、 例えば、 混合音響信号に含まれる既知 音響信号の時間軸方向の伸縮を推定し、 推定した時間軸方向の伸縮に基づいて 既知音響信号の振幅スぺクトルの時間軸方向の伸縮を補正する。 時間軸方向の 伸縮の推定には、 例えば、 混合音響信号中の既知の音響信号だけが含まれてい る区間を特定し、 この区間に対応する既知音響信号の区間との時間軸の対比に より、 時間軸方向の伸縮を推定する。 あるいは、 時間軸を短い区間に分割した 全区間の対比によって推定する。
また、 補正ステップの実行では、 例えば、 混合音響信号に含まれる既知 音響信号の周波数軸方向の伸縮を推定し、 推定した周波数軸方向の伸縮に基づ いて既知音響信号の振幅スぺクトルの周波数軸方向の伸縮を補正する。 周波数 軸方向の伸縮の推定には、 例えば、 混合音響信号中の既知の音響信号だけが含 まれている区間を特定し、 この区間に対応する既知音響信号の区間との周波数 軸の対比により、 周波数軸方向の伸縮を推定する。 また、 本発明の既知音響信号除去方法においては、 混合音響信号の振幅 スぺクトルと既知音響信号の振幅スぺクトルを視覚により認識できるように画 像表示する画像表示ステップを、 更に実行するようにしても良い。 この場合に は、 画像表示に基づいて人間が混合音響信号中における既知の音響信号が含ま れている区間を定め、 この区間について補正ステップ、 除去ステップ、 逆変換 ステップまたは合成ステヅプを実行する。
また、 本発明の既知音響信号除去方法においては、 混合音響信号、 既知 音響信号及び合成ステツプの出力信号を音響として再生する音響再生ステツプ を、 更に実行するようにしても良い。 この場合には、 音響再生ステップからの 再生音に基づいて人間が混合音響信号中における既知の音響信号が含まれてい る区間を定め、 この区間について補正ステヅプ、 除去ステヅプ、 逆変換ステヅ プ及び合成ステップを実行する。
また、 本発明の既知音響信号除去方法においては、 混合音響信号の振幅 スぺグトルに基づいて、 混合音響信号中における既知の音響信号が含まれてい る区間を自動推定し、 この区間について補正ステップ、 除去ステップ、 逆変換 ステヅプ及び合成ステツプを実行するようにしても良い。 混合音響信号中に比 較的はっきりと既知の音響信号が含まれている場合 (例えば、 混合音響信号中 で既知の音響信号が単独で鳴っている区間がある場合) には、 自動推定により 区間を特定することができ、 自動推定を利用することにより、 既知音響信号の 除去作業を速く実施できる。 なお、 混合音響信号中に含まれる既知音響信号の 存在があまりはっきりとしていない場合においては、 人間が区間を指定する。
更に、 本発明の既知音響信号除去方法においては、 混合音響信号中に含 まれている音響信号に相当する既知音響信号が複数種類存在する場合には、 そ れら複数の既知音響信号のすべてに関して、 既知音響信号変換ステップ及び補 正ステップを実行し、 混合音響信号の振幅スぺクトルから複数の既知音響信号 の補正振幅スペクトルをすベて除去する除去ステップを実行して得た除去後振 幅スペクトルを用いて、 逆変換ステップ及び合成ステップを実行する。 これに より、 混合音響信号中から複数種類のすべての既知音響信号を除去することが できる。
また、 補正ステップを実行する際、 時間的な位置のずれ、 周波数特性の 時間変化、 音量の時間変化、 時間軸方向の伸縮及び周波数軸方向の伸縮の少な くとも 1つの補正を人間が手作業で指定することができるインタフヱ一ス処理 を行うグラフィヅクュ一ザイン夕フェース ( G U I ) を用いる。
イン夕フェース処理を行う処理モジュールは、 複数の音響信号が混合さ れた混合音響信号の中から、 既知音響信号の成分を除去する際に、 混合音響信 号中での既知音響信号の正確な位置を人間が指定できるように構成される。
また、 イン夕フエ一ス処理を行う処理モジュールは、 混合音響信号中で 既知音響信号の周波数特性が時間的に変化しているときに、 それらの変化を人 間が指定できるように構成される。 また、 イン夕フエ一ス処理を行う処理モジ ' ユールは、 混合音響信号中で既知音響信号の音量が時間的に変化しているとき に、 それらの変化を人間が指定できる'ように構成される。
更に、 インタフェース処理を行う処理モジュールは、 混合音響信号中で 既知の音響信号が時間軸または周波数軸方向に伸縮しているときに、 それらの 伸縮を人間が指定できるように構成される。
また、 イン夕フェース処理を行う処理モジュールは、 混合音響信号と既 知音響信号の対応する区間を人間が指定できるように構成される。 また、 本発明による既知音響信号除去装置は、 混合音響信号を時間周波 数表現に変換して混合音響信号の搌幅スぺクトルと混合音響信号の位相とを求 める混合音響信号変換手段と、 混合音響信号中に含まれている音響信号に相当 する既知音響信号を時間周波数表現に変換して既知音響信号の振幅スぺクトル を求める既知音響信号変換手段と、混合音響信号の振幅スぺクトルに基づいて、 混合音響信号の振幅スぺクトルに対する既知音響信号の振幅スぺクトルの時間 的な位置のずれ、 周波数特性の時間変化、 音量の時間変化、 時間軸方向の伸縮 及び周波数軸方向の伸縮の少なくとも 1つを補正した既知音響信号の補正振幅 スぺクトルを求める補正手段と、 混合音響信号の振幅スぺクトルから既知音響 信号の補正振幅スペクトルを除去する除去手段と、 除去手段により得た除去後 振幅スぺクトルと混合音響信号の位相とに基づいて時間表現に逆変換を行って 単位波形を求める逆変換手段と、 単位波形を合成して既知音響信号の成分を除 去した音響信号を得る合成手段とから構成される。
ここでの補正手段には、 混合音響信号の振幅スぺクトルに対する既知音 響信号の振幅スペクトルの時間的な位置のずれ、 周波数特性の時間変化、 音量 の時間変化、 時間軸方向の伸縮及び周波数軸方向の伸縮の少なくとも 1つの補 正の指定を人間が手作業で行うことができるインタフヱ一ス処理を行う処理モ ジュールを設ける。
ィンタフヱ一ス処理を行う処理モジュールは、 混合音響信号の振幅スぺ クトルと既知音響信号の振幅スぺクトルとを視覚により対比できるように画像 表示する画像表示部と、 混合音響信号、 既知音響信号及び合成手段の出力信号 を音響として苒生する音響再生部とを備える。
ィン夕フエ一ス処理を行う処理モジュールを用いると、 画像表示部に表 示された混合音響信号の振幅スぺクトル及び既知音響信号の振幅スぺクトルの 画像表示及びノまたは音響再生部からの再生音に基づいて、 混合音響信号中に 含まれている既知音響信号の区間を人間が指定できるだけでなく、 この区間に ついて人間が手作業で既知音響信号の振幅スぺクトルの時間的な位置のずれ、 周波数特性の時間変化、 音量の時間変化、 時間軸方向の伸縮及び周波数軸方向 の伸縮の少なくとも 1つの補正を指定できる。 その結果、 混合音響信号中に含 まれている既知音響信号の態様が多少複雑であっても、 高い除去精度で既知音 響信号を除去することができる。
なお、 画像表示部は、 既知の音響信号が含まれている混合音響信号中の 区間の振幅スぺクトルと、 既知音響信号の対応区間の振幅スぺクトルの時間的 な位置のずれ、 周波数特性の時間変化、 音量の時間変化、 時間軸方向の伸縮及 び周波数軸方向の伸縮の少なくとも 1つを補正した補正振幅スぺクトルとを時 間軸上で位置を合わせて表示できるように構成されていることが望ましい。
このような構成にすると、 補正振幅スペクトルの状態を視覚で確認でき るので、 補正スペクトルをどのようにすれば、 除去精度を高めることができる のかを、 画像を見ながら推測することができるので、 除去作業が速くなる。
また、 画像表示部は、 前記混合音響信号の前記振幅スペクトルから前記 補正振幅スぺクトルを除去した音響信号の振幅スぺクトルを画像表示できるよ うに構成することが望ましい。 このような構成にすると、 補正の効果を画像で 確認できるので、 カットアンドトライ方式で補正を行いながら、 混合音響信号 中から既知音響信号を最大限除去することができる。
また、 本発明による既知音響信号除去プログラムは、 混合音響信号を時 間周波数奉現に変換して混合音響信号の振幅スぺクトルと混合音響信号の位相 とを求める混合音響信号変換ステップと、 混合音響信号中に含まれている音響 信号に相当する既知音響信号を時間周波数表現に変換して既知音響信号の振幅 スぺクトルを求める既知音響信号変換ステップと、 混合音響信号の振幅スぺク トルに基づいて、 混合音響信号の振幅スぺクトルに対する既知音響信号の振幅 スペクトルの時間的な位置のずれ、 周波数特性の時間変化、 音量の時間変化、 時間軸方向の伸縮及び周波数軸方向の伸縮の少なくとも 1つを補正した前記既 知音響信号の補正振幅スぺクトルを求める補正ステップと、 混合音響信号の振 幅スぺクトルから既知音響信号の補正振幅スぺクトルを除去する除去ステップ と、 除去ステップにより得た除去後振幅スぺクトルと混合音響信号の位相とに 基づいて時間表現に逆変換を行って単位波形を求める逆変換ステツプと、 単位 波形を合成して既知音響信号の成分を除去した音響信号を得る合成ステップと の処理をコンピュータにより実行させるように構成されている。
本発明の既知音響信号除去方法によれば、 補正ステップにより、 混合音 響信号の振幅スぺクトルに対する既知音響信号の振幅スぺクトルの時間的な位 置のずれ、 周波数特性の時間変化、 音量の時間変化、 時間軸方向の伸縮及び周 波数軸方向の伸縮の少なくとも 1つを補正した既知音響信号の補正振幅スぺク トルを求め、 この補正振幅スぺクトルを混合音響信号の振幅スぺクトルから除 去するため、 混合音響信号中に非定常な雑音として含まれている既知音響信号 を高い精度で除去することができる利点が得られる。
また、 本発明の既知音響信号除去方法によれば、 例えば、 人間の声や物 音の背景に B G Mが鳴っているテレビ番組や映画等の音響信号を入力とすると、 別途用意した; B G Mの音楽音響信号を用いて番組中の B GMを除去し、 人間の 声や物音だけの音響信号を得ることが可能となる。 更に、 B G M除去後の音響 信号に、 別の音楽を B G Mとして付与することにより、 テレビ番組や映画等の 音楽を差し換えた再利用が可能となる。
ここでの既知音響信号は、 任意の音響信号でよいため、 音楽のジャンル を問わず、ボ一カルの有無を問わず、伴奏の有無を問わずに適用できる。また、 音楽に限らず、 定常雑音及び非定常雑音を含めた、 任意の既知の雑音に適用で
& ^ o
また、既知音響信号除去装置におけるユーザィン夕フェースを使用して、 人間が手作業で修正することで、 実務の現場でより高品質な除去作業が実現で きる。 図面の簡単な説明
第 1図は、 本発明の既知音響信号除去装置の実施の形態の一例の構成を 示すブロック図である。
第 2図は、 本発明の既知音響信号除去方法を実施する場合のステツプを 示すブロック図である。
第 3図は、 本発明の既知音響信号除去装置の主要部を、 コンピュータを 用いて実現する場合に用いるプログラムのアルゴリズムの一例を示すフローチ ャ一トである。
第 4図は、 第 3図のステヅプ S T 1 0 3内の詳細な処理を示すフローチ ャ一トである。
第 5図は、 人間がかかわる推定と自動推定のいずれでも推定動作をする 場合のステップの詳細を示すフローチャートである。
第 6図は、 エディタのィン夕フェースの画面構成を示す図である。
第 7図は、 混合音響信号のパワーの時間変化を示す図である。
第 8図は、 混合音響信号の振幅スぺクトルの時間変化を示す図である。 第 9図は、 B G Mの元となる音源の既知音響信号のパワーの時間変化を 示す図である。
第 1 0図は、 B G Mの元となる音源の既知音響信号の振幅スぺクトルの 時間変化を示す図である。
第 1 1図は、 既知音響信号除去後の所望の音響信号のパワーの時間変化 を示す図である。
第 1 2図は、 既知音響信号除去後の所望の音響信号の振幅スぺクトルの 時間変化を示す図である。 発明を実施するための最良の形態
以下、 図面を参照して本発明の実施の形態の一例を詳細に説明する。 第 1図は、 本発明の既知音響信号除去方法を実施する既知音響信号除去装置の一 実施の形態の構成を示すプロックである。
既知音響信号除去装置は、 システム構成としては、 混合音響信号変換手 段 1と、 既知音響信号変換手段 2と、 補正手段 3と、 イン夕フエ一ス 4と、 除 去手段 5と、 逆変換手段 6と、 合成手段 7とから構成される。
混合音響信号変換手段 1は、所望の音声や物音等の音響信号 s ( t ) (t は時間軸) に、 B GM等の音響信号 b (t ) が混合された混合音響信号 m (t ) を(この時点では s (t ) と b (t )は未知であり m (t )のみが入力される)、 時間周波数表現に変換して混合音響信号の振幅スペクトル M (ω, t ) と混合 音響信号の位相 0m (ω, t ) とを求める。
また、 既知音響信号変換手段 2は、 除去すべき音響信号 b (t ) の元と なる音源の既知音響信号 b' (t ) を時間周波数表現に変換して既知音響信号 の振幅スペクトル Β, (ω, t) を求める。
そして、 補正手段 3は、 混合音響信号の振幅スペクトル M (ω, t ) に 基づいて、 混合音響信号の振幅スペク トル M (ω, t ) に対する既知音響信号 の振幅スペク トル Β, (ω, t ) の時間的な位置のずれ、 周波数特性の時間変 ィ匕、 音量の時間変化、 時間軸方向の伸縮及び周波数軸方向の伸縮を補正した既 知音響信号の補正振幅スペクトル B (ω, t ) を求める。 自動化のためには、 . 自動で位置のずれ、 周波数特性の時間変化、 音量の時間変化、 時間軸方向の伸 縮及び周波数軸方向の伸縮のすべてを自動で推定して補正するように補正手段 3を構成することができる。
この実施の形態では、 補正手段 3は、 時間的な位置のずれ、 周波数特性 の時間変化、 音量の時間変化、 時間軸方向の伸縮及び周波数軸方向の伸縮のす ベての補正を、 ィンタフエース 4を用いて人間が手作業で指定することができ るように構成されている。
このイン夕フヱ一ス 4は、 後に詳しく説明するように、 混合音響信号の 振幅スぺクトルと既知音響信号の振幅スぺクトルとを視覚により対比できるよ うに画像表示する画像表示部を備えており、 グラフィックュ一ザイン夕フエ一 ス (GU I) によりインタフェース処理を行う処理モジュールである。
ィン夕フェース 4は、 画面表示された入力部により混合音響信号の振幅 スぺクトルと既知音響信号の振幅スぺクトルとに基づいて混合音響信号中に含 まれている既知音響信号の区間を人間が指定でき且つ前述の補正を指定できる ように構成されている。
除去手段 5は、 混合音響信号の振幅スペクトル M ( ω , t ) から既知音 響信号の補正振幅スぺクトル B ( ω , t ) を除去する。そして逆変換手段 6は、 除去手段 5により得た除去後振幅スペクトル S ( ω , t ) と混合音響信号の位 相 6> m ( w , t ) とに基づいて時間表現に逆変換を行って単位波形 s ' ( t ) を求める。
最後に、 合成手段 7は、 逆変換手段 6から出力される単位波形 s ' ( t ) を合成して既知音響信号の成分を除去した音響信号 s ( t ) を得る。 インタフ エース 4は、 除去手段 5から出力された除去後振幅スペクトル S ( ω , t ) を 画像表示部 (第 6図参照) に表示する。 また、 イン夕フェース 4は音響再生部 を内蔵しており、 混合音響信号、 既知音響信号及び合成手段 7から出力された 合成された音響信号を再生する。
この構成によれば、 補正の効果を画像表示部で視覚により確認し、 また 内蔵された音響再生部で聴覚によっても確認できるので、 カットアンドトライ 方式で補正を行いながら、 ィン夕フエース 4の画像表示部の画面表示を見なが ら、 人間が必要な補正を指定することにより、 混合音響信号中から既知音響信 号を最大限除去することができる。 次に、 第 2図及び第 3図を用いて、 本発明の既知音響信号除去装置の詳 細な実施の形態の一例を説明する。 第 2図は、 本発明の既知音響信号除去方法 を実施する場合のステップを示すプロック図であり、 第 3図は本発明の既知音 響信号除去装置の主要部を、 コンピュータを用いて実現する場合に用いるプロ グラムのァルゴリズムの一例を示すフローチャートである。
第 4図は、 第 3図のステヅプ S T 1 0 3内の詳細な処理を示すフ D—チ ャ一トである。 また、 第 5図は、 人間がかかわる推定と自動推定のいずれでも 推定処理を実行する場合のステップの詳細を示すフローチャートである。 以下 これらの第 1図乃至第 5図を参照しながら、 本発明の既知音響信号除去方法及 び装置における既知音響信号除去の動作を説明する。
まず、 以下の説明では、 所望の音声や物音等の音響信号 s (t ) (tは 時間軸) に、 除去する既知音響信号である BGM等の音響信号 b (t ) が混合 された混合音響信号 m (t) が観測されるものとする。 (t) = s(t) + b(t) (1) ここでは、 b (t ) の元となる音源の音響信号 b' (t ) が既知という条件下 で、 m (t ) が与えられたときに、 未知の s (t ) を求める問題を解く。 例え ば、 人間の声や物音と共に BGMが鳴っているテレビ番組等の音響信号 m (t ) を入力とし、 その BGMの楽曲が既知であり、 その音響信号 b, (t ) が別途 用意できるときに、 その; B GMの音楽音響信号を用いて番組中の B GMを除去 し、 人間の声や物音だけの音響信号 s ( t ) を得る処理を実現する。 ここで、 b (t) とわ' (t) は完全には一致しないため、 s(t) = m(t) - b(t) (2) の減算に相当する処理では、 b, (t ) から b (t ) に相当する成分を推定し て、 s (t ) を求める必要がある。 具体的には、 既知音響信号 b, (t ) は、 混合音 m (t ) 中では、 以下のような変形を伴うことが多いため、 補正するこ とにより、 b (t ) に相当する成分を推定する。 補正の対象は、 主として以下 に説明するように、 時間的な位置のずれ、 周波数特性の時間変化、 音量の時間 変化、 時間軸あるいは周波数軸方向の伸縮である。
(時間的な位置のずれ)
混合音 m (t ) 中で既知音響信号 b 5 (t) が鳴っている位置は先頭か らとは限らない。 そこで、 既知音響信号 b' (t ) を時間軸方向にずらし、 両 者の相対位置を合わせて、 混合音から既知音響信号を減算する必要がある。
(周波数特性の時間変化)
混合音 m (t) 中で既知音響信号 b, (t ) が鳴る際には、 グラフイツ クイコライザ等の影響で周波数特性が変化することが多い。 例えば、 低域や高 域が強調 ·減衰されることがある。 そこで、 b' (t ) の周波数特性を同様に 変化させて補正し、 混合音から既知音響信号を減算する必要がある。
(音量の時間変化)
混合音 m (t ) 中で既知音響信号 b' (t) が鳴る際には、 混合音作成 時のミキサーのフエーダー等の操作で混合比率が変更され、 音量が時間変化す ることが多い。 そこで、 b' (t ) の音量を同様に時間変化させて補正し、 混 合音から既知音響信号を減算する必要がある。 (時間軸あるいは周波数軸方向の伸縮)
混合音 m (t) 中で既知音響信号 b' (t ) が鳴る際には、 レコード等 の回転数の違いにより、 時間軸あるいは周波数軸方向に伸縮されることがある.。 そこで、 b' (t ) を時間軸あるいは周波数軸方向に伸縮して補正し、 混合音 から既知音響信号を減算する必要がある。 本発明の既知音響信号除去方法においては、 基本的な処理として、 第 2 図に示すように、 ステップ ST 1において、 まず、 混合音響信号をフーリエ変 換して、 混合音響信号の位相 (ステップ ST2) と混合音響信号の振幅スぺク トル (ステヅプ ST3) を求める (混合音響信号変換ステップ) と共に、 ステ ップ S T 4で混合音響信号中に含まれている音響信号に相当する既知音響信号 をフ一リエ変換して、 既知音響信号の振幅スペクトル (ステップ ST5) を求 める (既知音響信号変換ステップ) 。 そして、 ステップ S T 6により、 混合音響信号の振幅スペクトルに基づ いて、 混合音響信号の振幅スぺクトルに対する既知音響信号の振幅スぺクトル の時間的な位置のずれ、 周波数特性の時間変化、 音量の時間変化、 時間軸方向 の伸縮及び周波数軸方向の伸縮の少なくとも 1つを補正した既知音響信号の補 正振幅スペクトル (ステップ S T 7 ) を求める (補正ステップ) 。 次に、 ステ ップ S T 8で、 混合音響信号の振幅スぺクトルから既知音響信号の補正振幅ス ぺクトルを除去して除去後振幅スペクトル (ステップ S T 9 ) を求めて (除去 ステップ) 、 次のステップ S T 1 0により、 除去ステップにより得た除去後振 幅スペクトルと混合音響信号の位相とに基づいて逆フーリエ変換を行って単位 波形を求める (逆変換ステップ) 。 最後に、 ステップ S T 1 1で、 単位波形を ォ—バーラップ ·ァド法により合成して既知音響信号の成分を除去した音響信 号を得る (合成ステップ) 。 これらの処理をコンピュータを用いて実現する場合に用いるプログラム のアルゴリズムでは、 第 3図フロ一チャートに示すように、 まず、 ステップ S T 1 0 1で、 混合音響信号をフーリエ変換して混合音響信号の振幅スぺクトル と混合音響信号の位相とを求める。 次に、 ステップ S T 1 0 2で、 混合音響信 号中に含まれている音響信号に相当する既知音響信号をフ一リエ変換して既知 音響信号の振幅スぺクトルを求める。
次のステップ S T 1 0 3では、 混合音響信号の振幅スぺクトルに基づい て、 混合音響信号の振幅スぺクトルに対する既知音響信号の振幅スぺクトルの 時間的な位置のずれ、 周波数特性の時間変化、 音量の時間変化、 時間軸方向の 伸縮及び周波数軸方向の伸縮の少なくとも 1つを補正した既知音響信号の補正 振幅スぺクトルを求める。
その後、 ステヅプ S T 1 0 4で、 混合音響信号の振幅スぺクトルから既 知音響信号の補正振幅スぺクトルを除去して除去後振幅スぺクトルを求める。 次にステップ S T 1 0 5で、 ステヅプ S T 1 0 4で得た除去後振幅スぺクトル と混合音響信号の位相とに基づいて逆フ一リエ変換を行って単位波形を求め、 ステップ S T 1 06で単位波形をオーバーラップ 'アド法により合成して既知 音響信号の成分を除去した音響信号を得る。
その後、 ステップ S T 1 0 7で、 除去後の音響信号をュ一ザが満足した と評価したか否かの判定が加わり、 判定結果が不満足であれば、 ステップ S T 1 03へと つて補正がやり直される。 ュ一ザが満足するまでは、 ステップ S T 1 03からステヅプ S T 1 07が繰り返される。
以下、 更に、 各ステップで実行される内容を詳細に説明する。 本発明の 実施の形態の方法では、 時間領域で波形を減算処理をせずに、 時間周波数領域 での振幅スぺクトル上で減算処理を行う。
例えば、 音響信号 m (t) , b, (t) に対する窓関数 h (t) を用い た時刻 tにおける短時間フーリエ変換(S TFT) Xm (w, t) , Xb, (ω, t ) が
Xm{^t) = Γ m(r)h{r - t)e~j 3Tdr (3)
J— OO
= am(w,t)+jpm(w, ) (4)
Figure imgf000020_0001
で定義されるとき、 それらの振幅スぺク小ル Μ (ω, t ) , Β' (ω, t ) は、
M( ) = ( m(w,i)| (7)
Figure imgf000020_0002
(9)
= V (",i〉+ ¾(o,i) (10) で求まる。
現在の実装では、 音響信号を標本化周波数 44. lkHz、 量子化ビッ ト数 16 b i tで A/D変換し、 窓関数 h (t ) として窓幅 8192点のハニ ング窓を用いた短時間フーリエ変換 (STFT) を、 高速フーリエ変換 (FF T) によって計算する。 その際、 高速フーリエ変換 (FFT) のフレームを 4 41点ずつシフ卜するため、 フレ一ムシフ 卜時間 ( 1フレームシフト) は 10 msとなる。 このフレ一ムシフトを、 処理の時間単位とする。
既知音響信号除去後の所望の音響信号 s (t )の振幅スぺクトル S (ω, t ) は、 振幅スペクトル M (ω, t) , Β' (ω, t) から以下の式によって 求める。 ここで、 B (ω, t) は Β, (ω, t) を補正した後の振幅スぺクト ルである。
, t) (Μ(ω, t), Β{ω, i)) if Μ(ω, t) > Β(ω, t)
S(w,t) = (ll)
otherwise ait) 9(ω,ί) Bf{p( ),q{t) ^T{t)) (12)
上記の式における各種パラメ一夕閧数 a (t ) , g (ω, t), ρ (ω), q (t) , r (t) , c (ω, t) を順に説明する。 ここでの a ( t ) は、 混合音の振幅スペクトルから既知音響信号の振幅 スぺクトルに相当する成分を減算する分量を最終的に調整するための任意の形 状の関数であり、 通常、 a (t ) 1とする。 これが大きいほど、 減算量が大 きくなる。
g (ω, t ) は、 周波数特性の時間変化と音量の時間変化を補正するた めの関数であり、
9( ),t) = (ω, + (13) のように定義する。 ここで、 §ω (ω, t ) は、 周波数特性の時間変化を表し、 周波数特性の変化がないときは gw (ω, t ) - 1となる。一方、 gt (t ) は、 音量の時間変化を表し、 音量の変化がないときは定数となる。 M (ω, t ) と Β' (ω, t) との音量差は、 基本的に gt (t ) で補正される。 gr (t) は、 主に g (ω, t ) の値を全体的に持ち上げるための関数で、 補正時の微調整に 使用される。 使用しない場合には、 gr (t) = 0とする。
P (ω) は、 周波数軸方向の伸縮を補正するための関数であり、 振幅ス ぺクトル Β, (ω, t ) の周波数軸 ωを変換することで、 周波数軸方向の線形 - 非線型な伸縮を可能にする。 なお、 Β' (ω, t ) は本来の ωの定義域外では 0をとり、 離散化して実装する際には適宜補間することとする。
q (t ) は、 時間軸方向の伸縮を補正するための閧数であり、 振幅スぺ クトル Β' (ω, t ) の時間軸 tを変換することで、 時間軸方向の線形 -非線 型な伸縮を可能にする。 なお、 Β, (ω, t) は本来の tの定義域外では 0を とり、 離散化して実装する際には適宜補間することとする。
r (t ) は、 時間的な位置のずれを補正するための関数であり、 通常は 定数を設定することで、 一定のずれ幅を補正する。 ずれ幅が時間変化するとき には、 各時刻での幅を補正する関数を設定する。 なお、 Β' (ω, t) は本来 の tの定義域外では 0をとり、 離散化して実装する際には適宜補間することと する。 q (t ) と (t ) を統合した一つの関数で表現することも可能だが、 ここでは、 q (t) は連続的な伸縮を表す目的で設定し、 r (t ) は不連続な 位置のずれを表す目的で設定することとする。
c (ω, t ) は、 振幅スペクトルに対するィコライジング処理及びフエ ーダー操作処理のための任意の形状の関数である。 ω方向の形状により、 グラ フィックイコライザのように、 既知音響信号除去後の周波数特性を調整するこ とができる。 また、 t方向の形状により、 ミキサ一のボリュームフエーダ一操 作のように、 既知音響信号除去後の音量変化を調整することができる。 使用し ない場合には、 c {ω, t ) = 1とする。 こうして求めた振幅スペクトル S (ω, t ) と、 混合音 m (t ) の位相 ( ω , - 1 ) を用いて Xs (ω, t ) を求め、 それを逆フ一リエ変換 ( I F F することで、 単位波形 s ' (t ) を得る。
9m(^,t) = arctan [ , , )
am(w,i)ノ
Figure imgf000023_0001
Sf(t) = 厂 X9( ,t)e^ (16)
この単位波形 s, (t ) を、 ォ一バーラヅプ 'アド (Ove r 1 ap Ad d) 法によって配置することにより、 既知音響信号除去後の所望の音響信号 s (t ) を合成する。
以上では、 混合音響信号 m (t ) の中に、 既知音響信号 b' (t ) が一 種類含まれている場合を説明じたが、 b' x (t ) 5 b' 2 (t ) , …, b, N (t)のように複数含まれている場合には、それらの振幅スぺクトル B' 丄 ( , t) 3 B' 2 (ω, t ) , …, Β, Ν (ω, t ) からそれぞれに応じたパラメ一 夕関数の設定で式 ( 1 2) によってそれそれ求めた: B! (ω, t) , Β2 (ω, t ) , …, ΒΝ (ω, t ) を用いて、 ' ' ' 0 otherwise
(17) のように S (ω, t ) を求める処理へ拡張できる。 その際には、 Βη (ω, t ) の各種パラメ一夕関数を順に設定するか、 全体のバランスを取りながら、 複数 の Βη (ω, t ) の各種パラメ一夕関数を平行して設定する。 また、 以上では、 モノラル信号を対象に説明したが、 ステレオ信号は、 左右を混合してモノラル信号に変換して適用してもよいし、 ステレオ信号の左 右の各信号に対して適用してもよい。 また、 ステレオ信号中の音源方向を利用 して、 適用してもよい。 上記各種パラメ一夕閧数の設定について説明する。 本発明の方法を適用 する際に、 式( 1 1)、 式( 1 2)、 式( 1 3) の各種パラメ一夕関数 a ( t ) , g (ω, t ) (gw (ω, t ) , gt ( t ) , gr ( t ) ) , p (ω) , q ( t ) , r (t) , c (ω, t ) の形状は、 自動推定してもよいし、 人間が手作業で設 定してもよい。 あるいは、 自動推定後に人間が修正してもよい。 以下では、 具 体的な自動推定方法と、 人間の手作業による修正を可能にする既知音響信号除 去装置におけるインタフエース 4を用いる場合について説明する。
最初に、 式 ( 1 1) 、 式 ( 1 2) 、 式 ( 1 3) の各種パラメータ関数 g (ω, t ) (gM (ω, t ) , gt (t ) ) , p (ω) , q (t") , r (t) め 形状を推定する方法を、 第 4図を用いて以下に説明する。 まず、 ステップ S T 20 1で BGM区間^の集合 Φの指定 · 自動推定を行い、 ステップ S T 2 0 2 で P (ω), q ( t ) の自動推定を行い、 ステップ S T 2 03で gw (ω, t ), gt (t ) , r (t) の自動推定を行う。 そして推定結果のパラメ一夕関数が収 束するまでこれらのステップが継続される (ステヅプ S T 2 04) 。 ステップ ST 205以降では、 補正動作がインタフヱ一ス 4を用いて実行される。
g (ω, t) の推定では、 まず、 周波数特性の時間変化 gw (ω, t) を 推定し、 次に、 音量の時間変化 gt (t ) を推定する。 ただし、 g (ω, t) の 推定に先立ち、 P (ω) , q (t) , r (t) は決定されている必要がある。 ここでは、 便宜上、 (ρ (ω) , q (t ) + r (t ) ) を B, (a), t ) と記述する。
周波数特性の時間変化 gw (ω, t ) の推定では、 原則として、 人間の声 や物音だけの音響信号 s (t ) がほとんど含まれていない区間 (以下、 BGM 区間と呼ぶ) を用いる。 B GM区間は、 複数用いてもよい。 BGM区間では、 混合音 m (t ) の振幅スペク トル Μ (ω, t ) は、 既知音響信号 b' (t ) に よる BGMに相当する振幅スペクトル B' (a), t ) に由来の成分がほとんど となる。 そこで、 周波数特性が時間変化せずに定常、 すなわち、 (ω, t ) =g, ω (ω) と仮定できるときには、 g' ω (ω) を ) '—-… ™
Figure imgf000025_0001
により推定する。 ただし、 は一つの BGM区間 (時間軸上の領域) を表し、 は、 の集合とする。一方、周波数特性が時間変化していくときには、 ω(ω, t ) の時刻 tに近い BGM区間 から
Figure imgf000025_0002
を求め、補間(内挿あるいは外揷)することにより gw (ω, t ) を推定する (両 側に BGM区間があるときには、 両側から内挿する) 。 最後に、 gw (ω, t) を周波数軸方向に平滑化する。 なお、 平滑化幅は任意に設定でき、 平滑化をし なくてもよい。
音量の時間変化 gt ( t ) の推定では、 振幅スペクトル M (ω, t ) と、 周波数特性補正後の gw (ω, t) Β, (ω, t) の各時刻における振幅を比較 する。 しかし、 振幅スペクトル M (ω, t ) には、 Β, (ω, t ) に由来の成 分以外に、 音響信号 s (t ) に由来の成分も含まれる。 そこで、 周波数軸 を 複数の周波数帯域 Φに分割し、 それぞれの帯域 {φ Φ) ごとに αί(φΛ) = — _ 1J- _ - _ ( €Φ) (20)
}φ9^) B^t)d
を求める (Φは の集合を表す) 。 Φとして任意の分割が適用できるが、 例え ば、 音楽で用いる平均律の 1オクターブごとに分割 (対数周波数軸上で等間隔 に分割) するとよい。 そして、 gt (t ) は、 mi n (g:, t ( , t ) ) ある いは
Figure imgf000026_0001
により推定する。 mi n ( :, t ( , t ) ) の場合には、 M (ω, t) と gw (ω, t ) Β' (ω, t) が一番近い周波数帯域において振幅が比較されるこ とになる。 最後に、 gt (t ) を時間軸方向に平滑化する。 なお、 平滑化幅は任 意に設定で'き、 平滑化をしなくてもよい。 p (ω) , q ( t ) の推定では、 M (ω, t) と Β (ω, t ) との距離 (例えば、 対数スペクトル距離等) が最小となるように、 P (ω) と q (t ) を変更する。 その際、 B (ω, t) =a (t ) g (ω, t ) B3 (p (ω) , q (t ) +r (t ) ) の右辺のうち、 a (t) = 1とし、 1. (推定途中の) ρ (ω) と q (t ) を仮に固定した上で、 g (ω, t) と r (t) を推定する、
2. (推定途中の) g (ω, t) と r> (t) を仮に固定した上で、 p (ω) と q (t) を推定する
との二つの推定を反復的に繰り返して、 適切な p (ω), q (t ) を推定する。 これは、 音響信号の全区間に対して一度に実行せず、 時間軸を分割して、 区分 的に行うとよい。 初期値は前後の区間の連続性を考慮して定める。 また、 BG M区間 øの集合 を用いて、それらの複数の区間における] νί(ω, t )と Β {ω, t ) との対応関係の時間軸を合わせるように、 P (ω) , q (t) を推定する とよい。
r (t ) の推定では、 原則として、 B GM区間^の集合 Ψを用いて、 そ れらの区間における M (ω, t) と Β (ω, t) との対応関係の時間軸を合わ せるように、 ; r (t ) を求める。 r (t ) は定数であることが多いが、 既知音 響信号 b3 (t ) の一部区間が使われずに、 飛び飛びで使用されながら混合さ れていたとき等には、 その区間を飛ばすように r (t ) が不連続関数となる。
上記の g (ω, t ) や (t ) 等の推定では、 BGM区間 の集合 を 用いていた。 これは、 人間が手作業で指定してもよい。 あるいは、 手作業で指 定した BGM区間の集合に自動推定で追加してもよい。 第 5図は、 人間が手作 業で指定する場合と自動推定する場合のいずれでも対応するプログラムのソフ トウエアのアルゴリズムを示すフローチャートである。 自動推定する場合には、 第 5図のステップ ST 302 ~S T 3 1 3を実行する。 Ψの自動推定では、 基 本的に、 どこか一箇所の BGM区間^ 1を手掛かりとして、 残りの BGM区間 の集合を求める。 まず、 最初の^ 1は、 人間が手作業で指定するか、 音響信号 の時間軸を細かく分割して、 それらの短い分割区間の対応関係を判定して求め る。 人間が手作業で指定しない場合、 B (ω, t) を仮に計算し (ステップ S T 302 )、 M (ω, t) と B (ω, t ) を細かく分割した時間窓の振幅スぺ クトル間の距離 (類似度に相当) を計算する (ステップ S T 303 ) 。 そして、その最小距離の時間窓の対応関係を調べ(ステップ ST 304)、 その結果を含む区間を^ 1に設定して初期の とする(ステヅプ S T 305 )。 次に、 01を含む Φに基づいて、 B (ω, t) の各種パラメ一夕関数を推定し (ステップ S T 306乃至ステップ S Τ 309)、 Β (ω , t ) を計算する (ス テツプ ST310) 。 各パラメ一夕の推定値が収束しているかを調べ、 収束し ていない場合には、 の全区間に対して、 M ( , t ) と B (ω, t ) との振 幅スペクトル間の距離 (類似度に相当) を求める。 ここでその最大値 (もしく は平均値) の定数倍を BGM区間判定用閾値とする (ステップ ST 3 12) 。 そして、 BGM区間判定用閾値以下の距離を持つ区間を検出し、 新たに Φに追 加する (ステヅプ ST313)。 ただし、 追加には上限を設けることもできる。 この推定と追加を繰り返すことで、 Ψが更新され、 各種パラメ一夕関数が適切 に求まっていく。 ここで、 M (ω, t ) と Β (ω, t ) との距離としては、 例 えば、 二乗平均対数スペクトル距離
Figure imgf000028_0001
が有効である。
次に、 既知音響信号除去エディ夕上のィンタフエースによる各種パラメ —夕関数の調整について説明する。
式 ( 1 1) 〜式 ( 13) のすベてのパラメ一夕関数 a ( t ) , g (ω, t) { Εω (ω, t) , gt (t) , gr (t) ) , p (ω) , q (t) 3 r (t) , c (ω, t) の形状を、 人間が手作業で設定するための既知音響信号除去装置 のユーザィンフェースであるエディタを以下に説明する。エディ夕のュ一ザは、 最初から任意の関数形状を描いて指定してもよいし、 最初はまず自動推定をし て、 その結果を修正してもよい。 エディ夕の画面構成を第 6図に示す。 このエディ夕は、 大別して、 混合 音響信号 m (t ) 操作用のサブウィンドウ W 3_、 既知音響信号 b' (t ) 操作 用のサブウィンドウ W 2、 既知音響信号除去後の所望の音響信号 s (t ) 操作 用のサブウインドウ W 3の三つのサブウインドウで構成されている。 既知音響 信号 b' (t ) が複数種類ある場合には、 切り替えスイッチ W2 Sにより、 サ ブウィンドウ W2で操作する既知音響信号 b' (t ) を切り替えることができ る。 このイン夕フエ一スでは、 第 4図に示したステヅプ S T 2 0 5からステヅ プ S T 2 1 9が実行される。
まず、 全サブウィンドウに共通の機能を述べる。 操作範囲スライダー P 1は、 音響信号中のどこを現在表示しているかを表す。 カーソル P 2は、 現在 の操作対象の時間軸上の位置を表す。 アイコン化 (折り畳み) ボタン P 3は、 これを押すと一時的にそのボタンの属するサブウィンドウが折り畳まれ、 小さ くなる。 現在操作対象以外の未使用のサブウィンドウを隠して、 狭い画面を有 効活用できる。 フロート化 (拡大) ボタン P 4は、 これを押すと一時的にその ボタンの属するサブウインドウが、親ウインドウから切り離され(フロート化)、 さらに拡大されて操作 '編集が容易になる。 フロ一卜化 (拡大) ボタン P 4し か描かれていない場合には、 このポ夕ンを押すと、 それに閧連づけられたサブ ウインドウがフロート化されて新たに出現する。 サブウインドウ W 1には、 混合音響信号 m ( t ) のパワーのグラフ E 1 とその振幅スペクトル M (ω, t ) のグラフ E 2が表示されている。 サブウイ ンドウ W2には、 既知音響信号 b, ( t ) のパワーのグラフ E 3とその振幅ス ぺクトル Β, (ω, t ) のグラフ E 4が表示されている。 サブウィンドウ W3 には、 既知音響信号除去後の音響信号 s (t ) のパワーのグラフ E 5とその振 幅スペクトル S (ω, t ) のグラフ E 6が表示されている。 各振幅スペクトル のグラフ (E 1, E 2, E 3) では、 左側に濃淡で振幅が描かれ (横軸が時間 軸、 縦軸が周波数軸) 、 右側に力一ソル位置での振幅が描かれている (横軸が パワー、 縦軸が周波数軸) 。 また、 再生制御操作パネル P 5 1には、 人間が聞いて確認するために、 混合音響信号の再生、 停止、 早送り、 早戾しが可能なボタン群が並んでいる。 再生制御操作パネル P 5 1の操作により、 イン夕フェース 4は、 内蔵する音響 再生部によって混合音響信号を再生する。
既知音響信号 b' ( t ) 操作用のサブウィンドウ W2が、 操作の中心と なるウィンドウであり、 式 ( 1 2) 、 式 ( 1 3) のすベてのパラメ一夕関数 a ( t ) , g- (ω, t ) ( gw (ω, t ) , gt (t) , gr ( t ) ) , p (ω) , q (t ) , r (t ) の形状を、 自由に設定できる。 以下、 各操作パネルの説明 を述べる。
1. 周波数特性の時間変化の補正用操作パネル C 1 (E 7の右側) gw3 t) を表示 ·操作するためのパネルで、 力一ソル位置の時刻 での ga (ω, t) が描かれている (横軸が大きさ、 縦軸が周波数軸) 。 設定操作結 果は、 g (ω, t ) の表示パネル E 7に即座に反映される (ステップ S T 2 0 5 , S T 2 0 6) 。 E 7には、 濃淡で g (ω, t ) の値の大きさが描かれてい る (横軸が時間軸、 縦軸が周波数軸) 。
2. 音量の時間変化の補正用操作パネル C 2 (E 7の下側)
gt (t) を表示 '操作するためのパネルで、 設定操作結果は、 g (ω, t ) の表示パネル E 7に即座に反映される (ステップ ST 207 , ST 20 8) 。
3. g (ω, t) の値を全体的に持ち上げるための操作パネル C 3 (E 7の下側)
gr (t) を表示 .操作するためのパネルで、 設定操作結果は、 g (ω, t ) の表示パネル E 7に即座に反映される (ステップ ST 2 09, S T 2 1 0) 。
4. 混合音の振幅スぺクトルから既知音響信号の振幅スぺクトルに相当 する成分を減算する分量を最終的に調整するための操作パネル C 4 a (t ) を表示 ·操作するためのパネルである。 このパネルを操作すると a ( t ) の変更が即座に表示に反映する (ステップ ST 2 1 1, ST 2 1 2) 。 5. 周波数軸方向の伸縮を補正するための操作パネル C 5
P (ω) を表示 '操作するためのパネルである。 このパネルを操作すると ρ (ΐ) の変更が即座に表示に反映する (ステップ ST 2 13 , S T 2 14) 。
6. 時間軸方向の伸縮を補正するための操作パネル C 6
q ( t ) を表示 .操作するためのパネルである。 このパネルを操作すると q ( t ) の変更が即座に表示に反映する (ステップ S Ύ 2 1 5 , ST 2 1 6) 。
7. 時間的な位置のずれを補正するための操作パネル C 7
r (t ) を表示 ·操作するためのパネルである。 このパネルを操作すると r (t ) の変更が即座に表示に反映する (ステップ ST 2 1 7, ST 2 1 8) 。 また、 再生制御操作パネル P 52には、 人間が聞いて確認するために、 既知音響信号の再生、 停止、 早送り、 早戻しが可能なボタン群が並んでいる。 再生制御操作パネル P 52の操作により、 インタフヱ一ス 4は、 内蔵する音響 再生部によって既知音響信号を再生する。 次に、 既知音響信号除去後の音響信号 s (t ) 操作用のサブウィンドウ W3では、 式 ( 1 1) のパラメ一夕関数 c (ω, t ) の形状を、 自由に設定で きる。 以下、 各操作パネルを説明する。
1. グラフィックイコライザ (GEQ) 操作パネル C 8 (E 8の右側) c (ω, t ) の ω方向の形状を表示 .操作するためのパネルで、 力一ソル位 置の時刻 tでの c (ω, t)が描かれている(横軸が大きさ、縦軸が周波数軸)。 設定操作結果は、 c (ω, t ) の表示パネル E 8に即座に反映される。 E 8に は、 濃淡で c (ω, t) の値の大きさが描かれている (横軸が時間軸、 縦軸が 周波数軸) 。
2. ボリュ一ムフエーダー操作パネル C 9 ( E 8の下側)
c (ω, t ) の t方向の形状を表示 '操作するためのパネルで、 設定操作結 果は、 c (ω, t ) の表示パネル E 8に即座に反映される。
また再生制御操作パネル P 5 3には、 人間が聞いて確認するために、 合成し た音響信号 (合成手段 7の出力) の再生、 停止、 早送り、 早戻しが可能なボタ ン群が並んでいる。 再生制御操作パネル P 5 3の操作により、 インタフェース 4は、 内蔵する音響再生部によって合成した音響信号を再生する。 次に、 本実施の形態の実装について説明する。 まず、 音声や物音等の音 響信号 s (t ) に BGM等の音響信号 b (t ) が加えられている混合音響信号 m (t ) が観測されたときに、 b (t ) の元となる音源の音響信号 b, (t ) が既知という条件下で、 未知の s ( t ) を求めることが可能なプログラムを、 各種ォペレ一ティングシステム(L i ux 2. 4 , SG I I R I X 6. 5 , Mi c r o s o f t Wi nd ows X P :登録商標) 上に実装した。 本プ ログラムに、 m (t) と b' (t) が収録されたオーディオファイルを与える と、 s (t) のオーディオファイルを得ることができる。 人間の音声や物音にバヅクグラウンドミュージック (BGM) が加えら れた様々な混合音に対して実験した結果、 その: B GMの原曲の音響信号を用い て、混合音中の BGMを除去し、人間の音声や物音が得られることを確認した。 ドラムスの鳴っている曲や鳴っていない曲、 ポピユラ一音楽やクラシヅク音楽 等の様々なジャンルの曲が B GMとして含まれていても、 除去が可能であった。 実験結果の例として、 二人の男女の対話の B GMにクラシック音楽が鳴 つている混合音を実際に処理した結果を、 第 7図〜第 1 2図に示す。 第 7図、 第 8図に示す混合音響信号 in (t ) を入力として、 第 9図、 第 1 0図に示す元 音源の既知音響信号 b' (t ) を用いて; B GM成分を除去した結果が、 第 1 1 図、 第 1 2図に示す既知音響信号除去後の音響信号 s (t ) となる。 この処理 結果の例の混合音は、 「: RWCP音声対話デ一夕べ一ス」 から抜粋した二人の 男女の対話の音響信号に、 「RWC研究用音楽デ一夕べ一ス」 から抜粋したク ラシック音楽の音響信号が加えられたものである。 以上に説明したように、 本発明によれば、 特に、 補正ステップにより、 混合音響信号の振幅スぺクトルに対する既知音響信号の振幅スぺクトルの時間 的な位置のずれ、 周波数特性の時間変化、 音量の時間変化、 時間軸方向の伸縮 及び周波数軸方向の伸縮の少なくとも 1つを補正した既知音響信号の補正振幅 スぺクトルを求め、 この補正振幅スぺクトルを混合音響信号の振幅スぺクトル から除去するため、 混合音響信号中に非定常な雑音として含まれている既知音 響信号を高い精度で除去することができる利点が得られる。
また、 人間の声や物音の背景に B G Mが鳴っているテレビ番組や映画等 の音響信号を入力とすると、 別途用意した BGMの音楽音響信号を用いて番組 中の B GMを除去し、 人間の声や物音だけの音響信号を得ることが可能となる。
更に、 BGM除去後の音響信号に、 別の音楽を BGMとして付与するこ とで、 テレビ番組や映画等の音楽を差し換えた再利用が可能となる。
既知音響信号は、任意の音響信号でよいため、音楽のジャンルを問わず、 ボーカルの有無を問わず、 伴奏の有無を問わずに適用できる。 また、 音楽に限 らず、 定常雑音及び非定常雑音を含めた、 任意の既知の雑音に適用できる。

Claims

請求の範囲
1 . 複数の音響信号が混合された混合音響信号から、 既知音響信号の成分を 除去する既知音響信号除去方法であって、
5 前記混合音響信号を時間周波数表現に変換して前記混合音響信号の振幅スぺ クトルと前記混合音響信号の位相とを求める混合音響信号変換ステツプと、 前記混合音響信号中に含まれている既知の音響信号に相当する既知音響信号 を時間周波数表現に変換して前記既知音響信号の振幅スぺクトルを求める既知 音響信号変換ステップと、
10 前記混合音響信号の振幅スペク トルに基づいて、 前記混合音響信号の振幅ス ぺクトルに対する前記既知音響信号の振幅スぺクトルの時間的な位置のずれ、 周波数特性の時間変化、 音量の時間変化、 時間軸方向の伸縮及び周波数軸方向 の伸縮の少なくとも 1つを補正した前記既知音響信号の補正振幅スぺクトルを 求める補正ステップと、
15 前記混合音響信号の振幅スぺクトルから前記既知音響信号の補正振幅スぺク トルを除去する除去ステップと、
前記除去ステップにより得た除去後振幅スペクトルと前記混合音響信号の位 相とに基づいて時間表現に逆変換を行って単位波形を求める逆変換ステップと、 前記単位波形を合成して前記既知音響信号の成分を除去し fe音響信号を得る Z0 合成ステップと
からなる既知音響信号除去方法。
2 . 前記補正ステップでは、 前記混合音響信号に含まれる前記既知音響信号 の時間的な位置を推定し、
25 推定した前記時間的な位置に基づいて前記既知音響信号の振幅スぺクトルの 時間的な位置のずれを補正する
ことを特徴とする請求の範囲第 1項に記載の既知音響信号除去方法。
3 . 前記補正ステップでは、 前記混合音響信号に含まれる前記既知音響信号 の周波数特性の時間変化を推定し、
推定した前記周波数特性の時間変化に基づいて前記既知音響信号の振幅スぺ クトルの周波数特性の時間変化を補正する
ことを特徴とする請求の範囲第 1項に記載の既知音響信号除去方法。
4 . 前記補正ステツプでは、 前記混合音響信号に含まれる前記既知音響信号 の音量の時間変化を推定し、
推定した前記音量の時間変化に基づいて前記既知音響信号の振幅スぺクトル の音量の時間変化を補正する
ことを特徴とする請求の範囲第 1項に記載の既知音響信号除去方法。
5 . 前記補正ステツプでは、 前記混合音響信号に含まれる前記既知音響信号 の時間軸方向の伸縮を推定し、
推定した前記時間軸方向の伸縮に基づいて前記既知音響信号の振幅スぺク卜 ルの時間軸方向の伸縮を補正する
ことを特徴とする請求の範囲第 1項に記載の既知音響信号除去方法。
6 . 前記補正ステップでは、 前記混合音響信号に含まれる前記既知音響信号 の周波数軸方向の伸縮を推定し、
推定した前記周波数軸方向の伸縮に基づいて前記既知音響信号の振幅スぺク トルの周波数軸方向の伸縮を補正する '
ことを特徴とする請求の範囲第 1項に記載の既知音響信号除去方法。
7 . 前記混合音響信号の振幅スぺクトルと前記既知音響信号の振幅スぺクト ルとを視覚により対比できるように画像表示する画像表示ステップと、 前記混合音響信号、 前記既知音響信号及び前記合成ステップの出力信号を音 として音響再生する音響再生ステツプとを更に備え、
前記画像表示と前記音響再生とに基づいて人間が前記混合音響信号中におけ る前記既知の音響信号が含まれている区間を定め、
前記区間について前記補正ステップ、 前記除去ステップ、 前記逆変換ステツ プ及び前記合成ステツプを実行する
ことを特徴とする請求の範囲第 1項に記載の既知音響信号除去方法。
8 . 前記混合音響信号の振幅スぺクトルに基づいて前記混合音響信号中にお ける前記既知の音響信号が含まれている区間を自動推定し、
前記区間について前記補正ステップ、 前記除去ステップ、 前記逆変換ステツ プ及び前記合成ステ、リプを実行する
ことを特徴とする請求の範囲第 1項に記載の既知音響信号除去方法。
9 . 前記混合音響信号中に含まれている前記既知の音響信号に相当する複数 の前記既知音響信号が存在する場合に、
前記複数の既知音響信号のすべてに関して前記既知音響信号変換ステツプ及 び前記補正ステップを実行し、
前記混合音響信号の振幅スぺクトルから前記複数の既知音響信号の補正振幅 スぺクトルをすベて除去する除去ステップによって得た除去後振幅スぺクトル を用いて、 前記逆変換ステップ及び前記合成ステップを実行する
ことを特徴とする請求の範囲第 1項に記載の既知音響信号除去方法。
1 0 . 前記補正ステップを実行する際に、 前記時間的な位置のずれ、 前記周 波数特性の時間変ィヒ、 前記音量の時間変化、 前記時間軸方向の伸縮及ぴ前記周 波数軸方向の伸縮の少なくとも 1つの補正を人間が手作業で指定するインタフ ヱ一スを用いる ことを特徴とする請求の範囲第 1項に記載の既知音響信号除去方法。
1 1 . 前記イン夕フェースは、 前記混合音響信号の振幅スぺクトルと前記既 知音響信号の振幅スぺクトルを視覚により対比できるように画像表示する画像 表示部を備え、
前記画像表示部の表示に基づいて前記補正を前記人間が手作業で指定する ことを特徴とする請求の範囲第 1 0項に記載の既知音響信号除去方法。
1 2 . 前記ィン夕フヱ一スは、 前記混合音響信号、 前記既知音響信号及び前 記合成ステップの出力信号を音響として再生する音響再生部を備え、
前記音響再生部の表示に基づいて前記補正を前記人間が手作業で指定する ことを特徴とする請求の範囲第 1 0項に記載の既知音響信号除去方法。
1 3 . 前記インタフヱ一スは、 前記混合音響信号の振幅スぺクトルと前記既 知音響信号の振幅スペクトルを視覚により対比できるように画像表示する画像 表示部と、 前記混合音響信号、 前記既知音響信号及び前記合成ステップの出力 信号を音響として再生する音響再生部とを備え、
前記画像表示部の表示及び前記音響再生部の再生音に基づいて前記補正を前 記人間が手作業で指定する
ことを特徴とする請求の範囲第 1 0項に記載の既知音響信号除去方法。
1 4 . 複数の音響信号が混合された混合音響信号から、 既知音響信'号の成分 を除去する既知音響信号除去装置であって、
前記混合音響信号を時間周波数表現に変換して前記混合音響信号の振幅スぺ クトルと前記混合音響信号の位相とを求める混合音響信号変換手段と、
前記混合音響信号中に含まれている既知の音響信号に相当する既知音響信号 を時間周波数表現に変換して前記既知音響信号の振幅スぺクトルを求める既知 音響信号変換手段と、
前記混合音響信号の振幅スぺクトルに基づいて、 前記混合音響信号の振幅ス ぺクトルに対する前記既知音響信号の振幅スぺクトルの時間的な位置のずれ、 周波数特性の時間変化、 音量の時間変化、 時間軸方向の伸縮及び周波数軸方向 の伸縮の少なくとも 1つを補正した前記既知音響信号の補正振幅スぺクトルを 求める補正手段と、
前記混合音響信号の振幅スぺクトルから前記既知音響信号の補正振幅スぺク トルを除去する除去手段と、
前記除去手段により得た除去後振幅スぺクトルと前記混合音響信号の位相と に基づいて時間表現に逆変換を行って単位波形を求める逆変換手段と、
前記単位波形を合成して前記既知音響信号の成分を除去した音響信号を得る 合成手段とからなる既知音響信号除去装置。
1 5 . 前記補正手段は、 前記時間的な位置のずれ、 前記周波数特性の時間変 化、 前記音量の時間変化、 前記時間軸方向の伸縮及び前記周波数軸方向の伸縮 の少なくとも 1つの補正を人間が手作業で指定することを可能にするイン夕フ エースを備えている
ことを特徴とする請求の範囲第 1 4項に記載の既知音響信号除去装置。
1 6 . 前記インタフェースは、 前記混合音響信号の振幅スペク トルと前記既 知音響信号の振幅スぺクトルとを視覚により対比できるように画像表示する画 像表示部と、 前記混合音響信号、 前記既知音響信号及び前記合成手段の出力信 号を音響として再生する音響再生部とを備え、
前記画像表示部に表示された前記混合音響信号の振幅スぺクトルと前記既知 音響信号の振幅スペクトルと、 前記音響再生部からの再生音とに基づいて、 前 記混合音響信号中に含まれている前記既知音響信号の区間の指定と、 前記既知 音響信号の振幅スぺクトルの前記時間的な位置のずれ、 前記周波数特性の時間 変化、 前記音量の時間変化、 前記時間軸方向の伸縮及び前記周波数軸方向の伸 縮の少なくとも 1つの補正の指定を人間が手作業で行えるように構成されてい る
ことを特徴とする請求の範囲第 1 5項に記載の既知音響信号除去装置。
1 7 . 前記画像表示部は、 前記既知の音響信号が含まれている前記混合音響 信号中の区間の前記振幅スぺクトルと、 前記混合音響信号中に含まれている前 記既知音響信号の対応区間の前記既知音響信号の振幅スぺクトルの時間的な位 置のずれ、 周波数特性の時間変化、 音量の時間変化、 時間軸方向の伸縮及び周 波数軸方向の伸縮の少なくとも 1つを補正した補正振幅スぺクトルとを時間軸 上で位置を合わせて表示できる構成である
ことを特徴とする請求の範囲第 1 6項に記載の既知音響信号除去装置。
1 8 . 前記画像表示部は、 前記混合音響信号の前記振幅スぺクトルから前記 補正振幅スぺクトルを除去した音響信号の振幅スぺクトルを画像表示できる構 成である
ことを特徴とする請求の範囲第 1 6または第 1 7項に記載の既知音響信号除去 装置。
1 9 . 複数の音響信号が混合された混合音響信号から既知音響信号の成分を 除去する処理をコンピュ一夕により実行するためのプログラムであって、 前記混合音響信号を時間周波数表現に変換して前記混合音響信号の振幅スぺ クトルと前記混合音響信号の位相とを求める混合音響信号変換ステツプと、 前記混合音響信号中に含まれている既知の音響信号に相当する既知音響信号 を時間周波数表現に変換して前記既知音響信号の振幅スぺクトルを求める既知 音響信号変換ステップと、
前記混合音響信号の振幅スぺクトルに基づいて、 前記混合音響信号の振幅ス ぺクトルに対する前記既知音響信号の振幅スぺクトルの時間的な位置のずれ、 周波数特性の時間変化、 音量の時間変化、 時間軸方向の伸縮及び周波数軸方向 の伸縮の少なくとも 1つを補正した前記既知音響信号の補正振幅スぺクトルを 求める補正ステヅプと、
前記混合音響信号の振幅スぺクトルから前記既知音響信号の補正振幅スぺク トルを除去する除去ステップと、
前記除去ステップにより得た除去後振幅スぺクトルと前記混合音響信号の位 相とに基づいて時間表現に逆変換を行って単位波形を求める逆変換ステップと、 前記単位波形を合成して前記既知音響信号の成分を除去した音響信号を得る 合成ステップとの処理をコンピュータにより実行させる
ことを特徴とするプログラム。
2 0 . 前記補正ステップで、 前記混合音響信号に含まれる前記既知音響信号 の時間的な位置を推定し、
推定した前記時間的な位置に基づいて前記既知音響信号の振幅スぺクトルの 時間的な位置のずれを補正することをコンピュータに実行させる
ことを特徴とする請求の範囲第 1 9項に記載のプログラム。
2 1 . 前記補正ステップで、 前記混合音響信号に含まれる前記既知音響信号 の周波数特性の時間変化を推定し、
推定した前記周波数特性の時閬変化に基づいて前記既知音響信号の振幅スぺ クトルの周波数特性の時間変化を補正することをコンビュ一夕に実行させる · · ことを特徴とする請求の範囲第 1 9項に記載のプログラム。
2 2 . 前記補正ステップで、 前記混合音響信号に含まれる前記既知音響信号 の音量の時間変化を推定し、
推定した前記音量の時間変化に基づいて前記既知音響信号の振幅スぺクトル の音量の時間変化を補正することをコンピュータに実行させる
ことを特徴とする請求の範囲第 1 9項に記載のプログラム。
2 3 . 前記補正ステップで、 前記混合音響信号に含まれる前記既知音響信号 の時間軸方向の伸縮を推定し、
推定した前記時間軸方向の伸縮に基づいて前記既知音響信号の振幅スぺクト ルの時間軸方向の伸縮を補正することをコンビュ一夕に実行させる
ことを特徴とする請求の範囲第 1 9項に記載のプログラム。
2 4 . 前記補正ステップで、 前記混合音響信号に含まれる前記既知音響信号 の周波数軸方向の伸縮を推定し、
推定した前記周波数軸方向の伸縮に基づいて前記既知音響信号の振幅スぺク トルの周波数軸方向の伸縮を補正することをコンピュ一タに実行させる ことを特徴とする請求の範囲第 1 9項に記載のプログラム。
2 5 . 前記混合音響信号の振幅スぺクトルと前記既知音響信号の振幅スぺク トルを視覚により対比できるように画像表示する画像表示ステップ
を更にコンピュータに実行させる
ことを特徴とする請求の範囲第 1 9項に記載のプログラム。
2 6 . 前記混合音響信号、 前記既知音響信号及び前記合成ステツプの出力信 号を音響として再生する音響再生ステップ
を更にコンビュ一夕に実行させる
ことを特徴とする請求の範囲第 1 9項に記載のプログラム。
2 7 . 前記混合音響信号の振幅スぺクトルに基づいて前記混合音響信号中に おける前記既知の音響信号が含まれている区間を自動推定するステップの処理 をコンピュータに実行させ、
前記区間について前記補正ステップ、 前記除去ステップ、 前記逆変換ステツ プ及び前記合成ステップと
の処理をコンピュータに実行させる
ことを特徴とする請求の範囲第 1 9項に記載のプログラム。
2 8 . 前記混合音響信号中に含まれている前記既知の音響信号に相当する複 数の前記既知音響信号が存在する場合に、
前記複数の既知音響信号のすべてに関して前記既知音響信号変換ステップ及 び前記補正ステップを前記コンビュ一夕に実行させ、
前記混合音響信号の振幅スぺクトルから前記複数の既知音響信号の補正振幅 スぺクトルをすベて除去する除去ステップによって得た除去後振幅スぺクトル を用いて、 前記逆変換ステツプ及び前記合成ステツプを前記コンピュ一タに実 行させる
ことを特徴とする請求の範囲第 1 9項に記載のプログラム。
PCT/JP2004/007587 2003-05-30 2004-05-26 既知音響信号除去方法及び装置 Ceased WO2004107319A1 (ja)

Priority Applications (2)

Application Number Priority Date Filing Date Title
GB0526570A GB2418577B (en) 2003-05-30 2004-05-26 Method and device for removing known acoustic signal
US10/558,608 US20070021959A1 (en) 2003-05-30 2004-05-26 Method and device for removing known acoustic signal

Applications Claiming Priority (4)

Application Number Priority Date Filing Date Title
JP2003-154964 2003-05-30
JP2003154964 2003-05-30
JP2003167118A JP4608650B2 (ja) 2003-05-30 2003-06-11 既知音響信号除去方法及び装置
JP2003-167118 2003-06-11

Publications (1)

Publication Number Publication Date
WO2004107319A1 true WO2004107319A1 (ja) 2004-12-09

Family

ID=33492453

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/JP2004/007587 Ceased WO2004107319A1 (ja) 2003-05-30 2004-05-26 既知音響信号除去方法及び装置

Country Status (5)

Country Link
US (1) US20070021959A1 (ja)
JP (1) JP4608650B2 (ja)
KR (1) KR101008250B1 (ja)
GB (1) GB2418577B (ja)
WO (1) WO2004107319A1 (ja)

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US10043532B2 (en) 2014-03-17 2018-08-07 Nec Corporation Signal processing apparatus, signal processing method, and signal processing program
CN110970045A (zh) * 2019-11-15 2020-04-07 北京达佳互联信息技术有限公司 混音处理方法、装置、电子设备和存储介质

Families Citing this family (12)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2006243664A (ja) * 2005-03-07 2006-09-14 Nippon Telegr & Teleph Corp <Ntt> 信号分離装置、信号分離方法、信号分離プログラム及び記録媒体
CN101300623B (zh) * 2005-09-02 2011-07-27 日本电气株式会社 用于抑制噪声的方法、设备和计算机程序
ATE458361T1 (de) 2005-12-13 2010-03-15 Nxp Bv Einrichtung und verfahren zum verarbeiten eines audio-datenstroms
WO2011001589A1 (ja) * 2009-06-29 2011-01-06 三菱電機株式会社 オーディオ信号処理装置
CN102576543B (zh) * 2010-07-26 2014-09-10 松下电器产业株式会社 多输入噪声抑制装置、多输入噪声抑制方法以及集成电路
US8849199B2 (en) 2010-11-30 2014-09-30 Cox Communications, Inc. Systems and methods for customizing broadband content based upon passive presence detection of users
US20120136658A1 (en) * 2010-11-30 2012-05-31 Cox Communications, Inc. Systems and methods for customizing broadband content based upon passive presence detection of users
JP5703807B2 (ja) * 2011-02-08 2015-04-22 ヤマハ株式会社 信号処理装置
WO2013046055A1 (en) * 2011-09-30 2013-04-04 Audionamix Extraction of single-channel time domain component from mixture of coherent information
US9195431B2 (en) * 2012-06-18 2015-11-24 Google Inc. System and method for selective removal of audio content from a mixed audio recording
US9373320B1 (en) * 2013-08-21 2016-06-21 Google Inc. Systems and methods facilitating selective removal of content from a mixed audio recording
US10052494B2 (en) * 2014-12-23 2018-08-21 Medtronic, Inc. Hemodynamically unstable ventricular arrhythmia detection

Citations (7)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH05197385A (ja) * 1992-01-20 1993-08-06 Sanyo Electric Co Ltd 音声認識装置
JPH07199990A (ja) * 1993-12-28 1995-08-04 Ricoh Co Ltd 音声認識装置
JPH08107375A (ja) * 1994-10-06 1996-04-23 Hitachi Ltd 音響信号記録再生装置
JPH10228296A (ja) * 1997-02-17 1998-08-25 Nippon Telegr & Teleph Corp <Ntt> 音響信号分離方法
JPH10307595A (ja) * 1997-03-07 1998-11-17 Seiko Epson Corp 入力音声抽出方法および入力音声抽出装置
JP2003022100A (ja) * 2001-07-09 2003-01-24 Yamaha Corp 雑音除去方法、雑音除去装置およびプログラム
JP2003099085A (ja) * 2001-09-25 2003-04-04 National Institute Of Advanced Industrial & Technology 音源の分離方法および音源の分離装置

Family Cites Families (10)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US5204969A (en) * 1988-12-30 1993-04-20 Macromedia, Inc. Sound editing system using visually displayed control line for altering specified characteristic of adjacent segment of stored waveform
US5792971A (en) * 1995-09-29 1998-08-11 Opcode Systems, Inc. Method and system for editing digital audio information with music-like parameters
US6343268B1 (en) * 1998-12-01 2002-01-29 Siemens Corporation Research, Inc. Estimator of independent sources from degenerate mixtures
US6446041B1 (en) * 1999-10-27 2002-09-03 Microsoft Corporation Method and system for providing audio playback of a multi-source document
JP3454206B2 (ja) * 1999-11-10 2003-10-06 三菱電機株式会社 雑音抑圧装置及び雑音抑圧方法
US6879952B2 (en) * 2000-04-26 2005-04-12 Microsoft Corporation Sound source separation using convolutional mixing and a priori sound source knowledge
JP4028680B2 (ja) * 2000-11-01 2007-12-26 インターナショナル・ビジネス・マシーンズ・コーポレーション 観測データから原信号を復元する信号分離方法、信号処理装置、モバイル端末装置、および記憶媒体
US7076433B2 (en) * 2001-01-24 2006-07-11 Honda Giken Kogyo Kabushiki Kaisha Apparatus and program for separating a desired sound from a mixed input sound
US7243060B2 (en) * 2002-04-02 2007-07-10 University Of Washington Single channel sound separation
US6971323B2 (en) * 2004-03-19 2005-12-06 Peat International, Inc. Method and apparatus for treating waste

Patent Citations (7)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH05197385A (ja) * 1992-01-20 1993-08-06 Sanyo Electric Co Ltd 音声認識装置
JPH07199990A (ja) * 1993-12-28 1995-08-04 Ricoh Co Ltd 音声認識装置
JPH08107375A (ja) * 1994-10-06 1996-04-23 Hitachi Ltd 音響信号記録再生装置
JPH10228296A (ja) * 1997-02-17 1998-08-25 Nippon Telegr & Teleph Corp <Ntt> 音響信号分離方法
JPH10307595A (ja) * 1997-03-07 1998-11-17 Seiko Epson Corp 入力音声抽出方法および入力音声抽出装置
JP2003022100A (ja) * 2001-07-09 2003-01-24 Yamaha Corp 雑音除去方法、雑音除去装置およびプログラム
JP2003099085A (ja) * 2001-09-25 2003-04-04 National Institute Of Advanced Industrial & Technology 音源の分離方法および音源の分離装置

Cited By (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US10043532B2 (en) 2014-03-17 2018-08-07 Nec Corporation Signal processing apparatus, signal processing method, and signal processing program
CN110970045A (zh) * 2019-11-15 2020-04-07 北京达佳互联信息技术有限公司 混音处理方法、装置、电子设备和存储介质
CN110970045B (zh) * 2019-11-15 2022-03-25 北京达佳互联信息技术有限公司 混音处理方法、装置、电子设备和存储介质

Also Published As

Publication number Publication date
US20070021959A1 (en) 2007-01-25
KR101008250B1 (ko) 2011-01-17
GB2418577A (en) 2006-03-29
KR20060034637A (ko) 2006-04-24
JP4608650B2 (ja) 2011-01-12
GB0526570D0 (en) 2006-02-08
JP2005049364A (ja) 2005-02-24
GB2418577B (en) 2007-10-17

Similar Documents

Publication Publication Date Title
JP5467098B2 (ja) オーディオ信号をパラメータ化された表現に変換するための装置および方法、パラメータ化された表現を修正するための装置および方法、オーディオ信号のパラメータ化された表現を合成するための装置および方法
Le Roux et al. Explicit consistency constraints for STFT spectrograms and their application to phase reconstruction.
TWI505264B (zh) 操縱具有瞬變事件的音頻信號的設備和方法以及具有執行該方法之程式碼的電腦程式
US11410637B2 (en) Voice synthesis method, voice synthesis device, and storage medium
JP4608650B2 (ja) 既知音響信号除去方法及び装置
US20050137729A1 (en) Time-scale modification stereo audio signals
MX2012009776A (es) Aparato y metodo para modificar una señal de audio usando bloqueo armonico.
RU2510954C2 (ru) Способ переозвучивания аудиоматериалов и устройство для его осуществления
CN111739544B (zh) 语音处理方法、装置、电子设备及存储介质
WO2002050814A1 (fr) Systeme et procede d&#39;interpolation de signaux
WO2003003345A1 (en) Device and method for interpolating frequency components of signal
US20050038534A1 (en) Fixed-size cross-correlation computation method for audio time scale modification
CN109416911B (zh) 声音合成装置及声音合成方法
JP4274419B2 (ja) 音響信号除去装置、音響信号除去方法及び音響信号除去プログラム
KR20050010927A (ko) 오디오 신호 처리 장치
JP2009282536A (ja) 既知音響信号除去方法及び装置
JP3849679B2 (ja) 雑音除去方法、雑音除去装置およびプログラム
JP4274418B2 (ja) 音響信号除去装置、音響信号除去方法及び音響信号除去プログラム
US20100131276A1 (en) Audio signal synthesis
JP4272107B2 (ja) 音響信号除去装置、音響信号除去方法及び音響信号除去プログラム
WO2020179472A1 (ja) 信号処理装置および方法、並びにプログラム
US11348596B2 (en) Voice processing method for processing voice signal representing voice, voice processing device for processing voice signal representing voice, and recording medium storing program for processing voice signal representing voice
JPH11143460A (ja) 音楽演奏に含まれる旋律の分離方法、分離抽出方法および分離除去方法
Disch et al. An iterative segmentation algorithm for audio signal spectra depending on estimated local centers of gravity
Wu Musical pitch shifting based on equalization and bandwidth extension

Legal Events

Date Code Title Description
AK Designated states

Kind code of ref document: A1

Designated state(s): AE AG AL AM AT AU AZ BA BB BG BR BW BY BZ CA CH CN CO CR CU CZ DE DK DM DZ EC EE EG ES FI GB GD GE GH GM HR HU ID IL IN IS KE KG KP KR KZ LC LK LR LS LT LU LV MA MD MG MK MN MW MX MZ NA NI NO NZ OM PG PH PL PT RO RU SC SD SE SG SK SL SY TJ TM TN TR TT TZ UA UG US UZ VC VN YU ZA ZM ZW

AL Designated countries for regional patents

Kind code of ref document: A1

Designated state(s): GM KE LS MW MZ NA SD SL SZ TZ UG ZM ZW AM AZ BY KG KZ MD RU TJ TM AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HU IE IT LU MC NL PL PT RO SE SI SK TR BF BJ CF CG CI CM GA GN GQ GW ML MR NE SN TD TG

121 Ep: the epo has been informed by wipo that ep was designated in this application
WWE Wipo information: entry into national phase

Ref document number: 1020057021034

Country of ref document: KR

WWE Wipo information: entry into national phase

Ref document number: 0526570.7

Country of ref document: GB

Ref document number: 0526570

Country of ref document: GB

WWE Wipo information: entry into national phase

Ref document number: 2007021959

Country of ref document: US

Ref document number: 10558608

Country of ref document: US

WWP Wipo information: published in national office

Ref document number: 1020057021034

Country of ref document: KR

122 Ep: pct application non-entry in european phase
WWP Wipo information: published in national office

Ref document number: 10558608

Country of ref document: US