WO2014054918A1 - Imdct 입력신호를 이용한 오디오 신호의 음정 및 속도 가변 장치 및 방법 - Google Patents

Imdct 입력신호를 이용한 오디오 신호의 음정 및 속도 가변 장치 및 방법 Download PDF

Info

Publication number
WO2014054918A1
WO2014054918A1 PCT/KR2013/008905 KR2013008905W WO2014054918A1 WO 2014054918 A1 WO2014054918 A1 WO 2014054918A1 KR 2013008905 W KR2013008905 W KR 2013008905W WO 2014054918 A1 WO2014054918 A1 WO 2014054918A1
Authority
WO
WIPO (PCT)
Prior art keywords
frequency
imdct
pitch
audio signal
input data
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/KR2013/008905
Other languages
English (en)
French (fr)
Inventor
박주성
이동훈
허경철
정승표
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
University Industry Cooperation Foundation of Pusan National University
Original Assignee
University Industry Cooperation Foundation of Pusan National University
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by University Industry Cooperation Foundation of Pusan National University filed Critical University Industry Cooperation Foundation of Pusan National University
Publication of WO2014054918A1 publication Critical patent/WO2014054918A1/ko
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L21/00Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
    • G10L21/04Time compression or expansion
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L19/00Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
    • GPHYSICS
    • G11INFORMATION STORAGE
    • G11BINFORMATION STORAGE BASED ON RELATIVE MOVEMENT BETWEEN RECORD CARRIER AND TRANSDUCER
    • G11B20/00Signal processing not specific to the method of recording or reproducing; Circuits therefor
    • G11B20/10Digital recording or reproducing

Definitions

  • the present invention relates to the pitch and speed of an audio signal, and specifically, to transform the pitch by processing the IMDCT input data X (k) before converting the signal into a time domain through an inverse modified discrete cosine transform (IMDCT).
  • IMDCT inverse modified discrete cosine transform
  • the present invention relates to an apparatus and method for varying the pitch and speed of an audio signal using an IMDCT input signal to reduce the amount of computation and memory.
  • Audio data is compressed to store a large amount of audio data on a CD, a hard disk, a mobile storage medium, or to be transmitted in a wired or wireless manner.
  • Audio data compression methods include compression in the time domain and compression in the frequency domain.
  • the compression method in the frequency domain not only has a high compression ratio but also a good sound quality, so that the audio signal in the time domain is converted into the frequency domain and compressed using a psychoacoustic model and other methods.
  • MP3 MPEG 3
  • AAC Advanced Audio Coding
  • the pitch converted signal When the audio signal converted to the time domain is played back faster than the normal speed, the tone increases, and when played slowly, the pitch decreases. Therefore, in order to reproduce faster or slower than normal speed without changing the pitch, the pitch is changed by using a method such as Synchronous Overlap and Add (SOLA).
  • SOLA Synchronous Overlap and Add
  • the pitch converted signal When the pitch converted signal is played back at normal speed, the converted pitch is reproduced as it is.
  • the pitch-converted audio signal can be played without changing the pitch according to the playback speed, or the pitch and playback speed can be changed simultaneously.
  • the audio signal is decomposed into frequency components of various bands through a filter bank 21, and the components decomposed in the filter bank 21 are modified discrete cosine transform (MDCT) in the MDCT block 22.
  • MDCT discrete cosine transform
  • the signal converted into the frequency domain is quantized in the quantization unit 23 and compressed in a form of low noise and low loss through coding in the coding unit 24.
  • the compressed data is bitstreamed along with side information in the bitstream encoding unit 25 and stored or transmitted.
  • the sidestream information and the compressed data are separated from the bitstream encoded by the bitstream decoding unit 31.
  • the data compressed in the frequency domain may be transformed in the IMDCT block 32 into the time domain through an inverse modified discrete cosine transform (IMDCT).
  • IMDCT inverse modified discrete cosine transform
  • the restored data is reconstructed through the synthesis filter bank 33 to be converted into an audio signal in the time domain.
  • the audio signal is completely converted into a signal in the time domain, and then the pitch change and speed of the audio signal are changed.
  • Equation 1 The process of converting the audio signal x (n) in the time domain into the frequency information X (k) through the MDCT process is shown in Equation 1.
  • w (n) is a window function and is expressed as Equation 2
  • N in Equations 1 and 2 means a window size to be analyzed.
  • N denotes the size of the analysis window.
  • MDCT is performed on the signal x (n) of the time domain within the analysis window to obtain the data X (k) of the frequency domain as shown in Equation (1).
  • Equation 3 An IMDCT process for converting the information X (k) transformed into the frequency domain into an audio signal in the time domain is shown in Equation 3.
  • the frequency information X (k) may be processed to change the frequency of the audio signal, thereby obtaining an effect of changing the pitch.
  • the present invention is to solve the problem that a large amount of computation and memory is required in the process of varying the playback speed without changing the pitch and pitch of the audio signal converted to the frequency domain in the prior art, it is easy to pitch the audio signal in the IMDCT process
  • the purpose of the present invention is to provide an apparatus and method for varying the pitch and speed of an audio signal using an IMDCT input signal by extracting various frequency components, amplitude and phase of each frequency from the IMDCT input data X (k) .
  • the present invention processes an input signal X (k) of an inverse modified discrete cosine transform (IMDCT) process that transforms a signal in a frequency domain into a signal in a time domain. Its purpose is to provide a pitch and speed variable device and method.
  • IMDCT inverse modified discrete cosine transform
  • the present invention relates to a pitch and speed variable method of an audio signal converted into a frequency domain, and a preprocessing apparatus and method for appropriately converting an IMDCT input signal to enable frequency conversion in an IMDCT process, which is converted into a signal in a time domain through an IMCDT.
  • the purpose of the present invention is to provide an apparatus and method for varying the pitch and speed of an audio signal using the IMDCT input signal by interpolating the playback speed of the audio with the IMDCT input signal preprocessing step so that the playback speed can be changed without changing the pitch. have.
  • the present invention extracts various frequency components before converting the frequency domain data of the audio signal converted into the frequency domain through the MDCT into the signal in the time domain through the IMDCT, and enables frequency conversion using the amplitude and phase of each frequency component. It is an object of the present invention to provide an apparatus and method for varying the pitch and speed of an audio signal using an IMDCT input signal by regenerating IMDCT input data.
  • the present invention provides an audio signal using an IMDCT input signal which can omit the process of pitch conversion in the time domain by changing the pitch and the speed by using the frequency information of the signal converted into the frequency domain in order to increase the compression ratio of the audio signal. Its purpose is to provide a pitch and speed variable device and method.
  • a device for changing a pitch of an audio signal using an IMDCT input signal may include: a window unit configured to determine a window size of a sample to be processed; the IMDCT input data X (k) at a window size determined by the window unit; A frequency extractor and a phase extractor to extract a frequency and a phase of the amplifier; an amplitude extractor to extract an amplitude; a frequency converter to convert the extracted frequency; IMDCT input data using the converted frequency and the extracted phase and amplitude An input data reconstruction unit configured to reconstruct an IMDCT block for converting the reconstructed input data into a time domain through IMDCT; And a synthesis filter bank for synthesizing the various subband components transformed into the time domain through the IMDCT.
  • the window unit in order to determine a window required for analyzing and reconstructing X (k), which is the IMDCT input data, the window unit has a relatively small window size in comparison with other regions.
  • the window unit In order to determine a window necessary for analyzing and reconstructing X (k), which is the IMDCT input data, the window unit overlaps and sets an analysis window at a server band, a server band, a frame and a frame boundary, and sets a spectrum within the analysis window. Finding the integer frequency (k in ) through the calculation is characterized by configuring the analysis window around the frequency.
  • a method of varying a pitch of an audio signal using an IMDCT input signal is performed by processing data X (k) input to the IMDCT before converting the signal into a signal in a time domain through the IMDCT. Determining a number of samples for frequency extraction; extracting the frequency (k), phase, and amplitude of the input data X (k) required for the IMDCT process with the selected window size; Reconstructing the IMDCT input data using the converted frequency and the extracted phase and amplitude; an IMDCT step of converting the data in the frequency domain into the time domain; synthesizing the time domain signals of various frequencies generated in the IMDCT process Synthesis filter bank step; characterized in that it comprises a.
  • phase of the IMDCT data X (k) It is characterized by calculating using the extracted IMDCT integer frequency component (k in ) to extract.
  • the amplitude A k of the IMDCT input data is obtained by using the integer frequency component Kin and the fractional frequency component ⁇ .
  • reconstruct the IMDCT input X '(k) using the converted frequency f shift f (1 + R f ).
  • a device for varying the pitch and speed of an audio signal using an IMDCT input signal including: a window unit configured to determine a window size of a sample to be processed; the IMDCT input data X ( k) a frequency extractor for extracting the frequency and phase and a phase extractor, an amplitude extractor for extracting the amplitude; a frequency converter for converting the extracted frequency; IMDCT using the converted frequency, the extracted phase and amplitude
  • An input data reconstruction unit configured to reconstruct input data; an IMDCT block for converting input data reconstructed through IMDCT into a time domain; a synthesis filter bank for synthesizing various subband components transformed into a time domain through IMDCT; in the synthesis filter bank
  • An interpolator for adjusting a sampling interval of an output audio signal to change a reproduction speed and a pitch; Characterized in that it comprises a.
  • the signal in order to change the pitch of the audio signal in the IMDCT block, the signal is decomposed and extracted from the IMDCT input data X (k) into sine wave components, and the frequency is changed as much as desired to reconstruct the IMDCT input.
  • the window unit may reduce the window size of the frequency domain to be analyzed relatively smaller than other regions.
  • the analysis window is superimposed on the server band, server band, frame, and frame boundary, and the spectrum within the analysis window. It is characterized by finding the integer frequency (k in ) through the calculation to configure the analysis window around the frequency.
  • a method of varying the pitch and speed of an audio signal using an IMDCT input signal processes data X (k) input to IMDCT before converting it into a signal in a time domain through IMDCT. Determining a number of samples for frequency extraction to change the pitch; extracting frequency k, phase, and amplitude of input data X (k) required for the IMDCT process with the selected window size; Converting the extracted frequency and reconstructing the IMDCT input data using the converted frequency and the extracted phase and amplitude; outputting a coded audio signal by performing IMDCT processing and interpolation.
  • the process of separating the frequency, phase, amplitude of the IMDCT input data X (k) in the form of an equation or storing the equation in the form of a look-up table in order to change the frequency as desired in the IMDCT process is characterized by including.
  • the determination of whether to use the one from the ⁇ and ⁇ is, the larger the absolute value of the frequency components X (k in) in the window and in the ratio k of the spectral values (S k) If the ratio is small compared to a specific threshold ⁇ 0 , ⁇ is selected, otherwise ⁇ is obtained to obtain a fraction ( ⁇ ) of the frequency component.
  • phase of cosine function of IMDCT data X (k) It is characterized by calculating using the extracted IMDCT integer frequency component (k in ) to extract.
  • an amplitude A k of the IMDCT input data may be obtained by using the integer frequency component Kin and the fractional frequency component ⁇ .
  • the variable speed, the sampling interval of the original signal (t s ), and the new signal are newly generated when the original speed is 1 in conjunction with the IMDCT processing and interpolation.
  • the apparatus and method for changing the pitch and speed of an audio signal using the IMDCT input signal according to the present invention have the following effects.
  • the frequency IMDCT input data X (k) may be processed to change the pitch before converting the audio signal compressed in the frequency domain into a signal in the time domain through IMDCT.
  • the frequency component and phase and amplitude of the input data X (k) are separated and displayed in the form of an equation, so that the frequency can be easily converted in the IMDCT process, thereby reducing the amount of calculation.
  • the process of converting the pitch in the time domain may be omitted.
  • 1 is a configuration diagram for varying the pitch or speed of an audio signal compressed in the frequency domain of the prior art
  • FIG. 2 is a configuration diagram for converting an audio signal in a time domain into a signal in a frequency domain in the prior art
  • FIG. 3 is a block diagram illustrating a process of converting a signal converted into a frequency domain of the prior art into a time domain
  • FIG. 4 is a configuration diagram of a variable pitch device of an audio signal using IMDCT according to the present invention
  • FIG. 5 is a block diagram of a pitch and speed variable device of an audio signal using the IMDCT according to the present invention
  • FIG. 6 is a conceptual diagram of a window configuration according to a frequency band within one frame of IMDCT input data X (k).
  • FIG. 7 is a flowchart illustrating a frequency component, phase, and amplitude extraction process in a window according to the present invention.
  • FIG. 4 is a block diagram of a device for changing the pitch of an audio signal using the IMDCT according to the present invention
  • Figure 5 is a block diagram of a device for changing the pitch and speed of an audio signal using the IMDCT according to the present invention.
  • the present invention provides a preprocessing step for processing an input signal to be used in the IMDCT step to vary the pitch and speed.
  • the apparatus for varying the pitch of an audio signal using the IMDCT according to the present invention as shown in FIG.
  • An input data reconstruction unit 46 for reconstructing, an IMDCT block 47 for converting input data reconstructed through IMDCT into a time domain, and a plurality of subband components converted into time domain through IMDCT Synthesis filter bank 48 is included.
  • the apparatus for varying the pitch and speed of an audio signal using IMDCT includes a window unit 41 for determining a window size (number of samples to be processed) of a sample to be processed, and the window unit (as shown in FIG. 5).
  • a frequency extractor 42 and a phase extractor 43 for extracting a frequency and a phase of the IMDCT input data X (k) at a window size determined by 41), an amplitude extractor 44 for extracting an amplitude, and the extracted
  • a frequency converter 45 for converting a frequency
  • an input data reconstruction unit 46 for reconstructing IMDCT input data using the converted frequency, an extracted phase, and an amplitude, and a time domain for the input data reconstructed through IMDCT
  • An IMDCT block 47 for converting the data
  • a synthesis filter bank 48 for synthesizing various subband components transformed in the time domain through IMDCT
  • an interpolation unit for changing the playback speed by adjusting the sampling interval of the audio signal (Int). erpolator 49).
  • the pitch variable method of the audio signal using the IMDCT input signal has a pre-processing step of processing the input signal to be used in the IMDCT step to change the pitch, to extract the size of the frequency component to finely extract the frequency component Adjusting, extracting a frequency of the IMDCT input data X (k), extracting a phase, extracting an amplitude, converting the extracted frequency according to a pitch change ratio, extracting frequency, Reconstructing the IMDCT input using phase and amplitude, an IMDCT step of converting data in the frequency domain into the time domain, and a synthesis filter bank step of synthesizing time domain signals of various frequencies generated in the IMDCT process.
  • the method of varying the pitch and speed of an audio signal using the IMDCT includes a preprocessing step of processing an input signal to be used in the IMDCT step, so as to vary the pitch and the speed, and to extract frequency components in detail. Adjusting the number of samples to be processed; extracting the frequency of the IMDCT input data X (k); extracting the phase; extracting the amplitude; and converting the extracted frequency according to the pitch change ratio. Reconstructing the IMDCT input using the extracted frequency, phase, and amplitude, the IMDCT step of converting the data in the frequency domain to the time domain, and synthesizing the time domain signals of various frequencies generated in the IMDCT process.
  • the calculation amount is reduced and the frequency, phase, and amplitude of the input data X (k) are separated and displayed in the form of an equation or look-up in the previous stage of IMDCT in proportion to the frequency change amount.
  • the data X (k) of the encoded frequency domain without undergoing the window process, frequency extraction, amplitude extraction, input data reconstruction, and frequency conversion process proposed by the present invention Can be decoded through IMDCT and synthesis filter bank.
  • the windowing process of the window portion 41 is as follows.
  • the frequency component of the IMDCT input data X (k) in the frequency domain should be extracted as closely as possible to reduce the loss of the frequency component.
  • the accuracy of the extraction is determined by the size (N) of the window, which determines how many (N) X (k) samples are collected and extracted.
  • the method of extracting frequency components is to extract the largest frequency component within N sample windows. It is good to reduce the amount of computation to make the window as large as possible within the range of overlapping frequencies. If the window size is reduced, the frequency components can be extracted finely, but there is a problem that the amount of calculation increases.
  • the frequency band of the audio signal is known as 20 KHz
  • the frequency component of the low frequency region is higher than the frequency component of the high frequency region.
  • the window size of the low frequency region is reduced and the window of the high frequency region is increased.
  • the input data X (k) used in the IMDCT step is a complex form as shown in Equation 1, but the present invention ,
  • f and Is the specific frequency and phase extracted from the IMDCT input data k is the frequency index in the MDCT, A k is the amplitude (frequency) of the frequency index.
  • the frequency f is divided into the integer part k in and the fractional part ⁇ , and three frequency components X (k) neighboring in the window are obtained.
  • X (k), X (k + 1) is used to calculate the spectral value using Equation 4, and k of X (k) having the largest spectral value in the window is an integer frequency component k in . do.
  • the fractional frequency component is compared with the ratio of the absolute value of the integer frequency (k in ) to the spectral value with the threshold value ( ⁇ 0 ) and X (k in ⁇ 1) or X (k in Calculate the fractional part using ⁇ 2).
  • the amplitude of the extracted frequency is obtained using + ⁇ ), X (k in ), X (k in -1), fractional frequency ( ⁇ ), and window size (N).
  • the speed change, the pitch change, the pitch and the play speed can be simultaneously changed by interlocking the pitch change through the frequency conversion in the step of IMDCT and the sampling interval of the interpolator.
  • the window is set by overlapping the server band and the frame in order to extract the integer frequency components present in the server band and server band, and the frame and frame boundary.
  • the spectrum is calculated using Equation 4 in the server band or frame, and the integer frequency component is found and an analysis window composed of a certain number of samples centered on several integer frequencies (k in ) having a large value.
  • the problem of how many integer frequency components are selected in one frame or subband depends on the number of samples constituting the frame or subband.
  • the window is determined by the integer frequency k in within 5 subbands within one subband.
  • X (k) to be used as an IMDCT input is expressed as the sum of terms equal to the window size (N) as shown in Equation 1. If the window is large, such as MP3 or AAC, MDCT sine wave of a single frequency (f) as shown in Equation 5, X (k) can be approximated to one term as shown in Equation 6.
  • the IMDCT input is the MDCT frequency index (k) and the phase of the frequency index ( ), Amplitude (A k ), and the single frequency (f) of the signal used as the input of the MDCT. Since even a complex audio signal in the time domain can be thought of as a sine wave sum of several frequencies, it extracts and extracts information (amplitude, frequency, and phase) of each frequency component of X (k), that is, index (k), to be used as an IMDCT input. The pitch can be changed by changing the frequency to the IMDCT stage.
  • the frequency, phase, and amplitude extraction processes of the frequency extractor 42, the phase extractor 43, and the amplitude extractor 44 are the same as those of FIG. 7.
  • f k in + ⁇ .
  • all the spectral values 61 existing in the window to be analyzed are obtained and the largest spectral value is obtained.
  • the frequency index k to make (S k ) is defined as the integer frequency component k in . (62)
  • Phase of cosine function of integer part IMDCT input data X (k) ) Is obtained using Equation 9 using the component X (k in ⁇ 1) and the fractional frequency component ⁇ immediately adjacent to the integer frequency component X (k in ). (70)
  • Equation 10 Substituting k in and k in -1 into X (k) of Equation 6, respectively, Equations 10 and 11 are obtained. Using the equations (10) and (11), an amplitude A k as shown in equation (12) can be obtained through simple manipulation. In Equations 10 and 11, f uses extracted frequency information (k in + ⁇ ).
  • the frequency, phase, and amplitude representing the window in one window were obtained.
  • various frequency (f), phase ( ), The amplitude (A k ) can be analyzed.
  • the process of regenerating IMDCT input data of the input data reconstruction unit 46 is as follows.
  • Frequency component (f k in + ⁇ ) extracted through the above process, frequency change rate (R f ) according to pitch change, phase ( ),
  • the IMDCT input data X '(k) is regenerated using the integer frequency component k in and the amplitude information A k as shown in Equation (13).
  • k is a frequency component of the IMDCT region.
  • FIG. 8 is the original signal with no change in pitch.
  • T 0 / T sh is called a frequency conversion ratio R f . If R f is greater than 1, the frequency will be higher than the original audio signal, resulting in higher pitch.
  • t 's ratio t s / t of a' s are defined as ratio R t of the sampling interval.
  • the ratio R t of the sampling interval may be thought of as the ratio of the reproduction speed (original speed / variable speed).
  • the slower playback speed is when (original speed / variable speed) is greater than 1.
  • the frequency before the interpolation step becomes (R f xf). If the ratio between the sampling interval of the converted signal and the sampling interval of the interpolation step is R t, the frequency of the final signal that has undergone the interpolation step is (R f xf) x R t .
  • the pitch change rate R f is set to the frequency change rate R f of the IMDCT preprocessing step, and the sampling interval rate is not changed.
  • the apparatus and method for varying the pitch and speed of an audio signal using the IMDCT input signal are to convert the frequency data X (k) input before converting the signal into a time domain through an Inverse Modified Discrete Cosine Transform (IMDCT).
  • IMDCT Inverse Modified Discrete Cosine Transform
  • the frequency component and amplitude of the input data X (k) are separated and displayed in the form of an equation. It includes an interpolation process that enables the simultaneous adjustment of the pitch, playback speed, pitch and playback speed by adjusting the sampling interval of the audio signal changed into the time domain.
  • An apparatus and method for varying the pitch and speed of an audio signal using an IMDCT input signal process the frequency data X (k) input before converting the signal into a time domain through an inverse modified discrete cosine transform (IMDCT). It is possible to change the pitch. It is possible to simultaneously adjust the pitch, playback speed, pitch and playback speed by adjusting the sampling interval of the audio signal that is changed in the time domain through the IMDCT process.
  • IMDCT inverse modified discrete cosine transform

Landscapes

  • Engineering & Computer Science (AREA)
  • Signal Processing (AREA)
  • Computational Linguistics (AREA)
  • Health & Medical Sciences (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Human Computer Interaction (AREA)
  • Physics & Mathematics (AREA)
  • Acoustics & Sound (AREA)
  • Multimedia (AREA)
  • Quality & Reliability (AREA)
  • Compression, Expansion, Code Conversion, And Decoders (AREA)

Abstract

본 발명은 오디오 신호의 음정 및 속도 가변시에 IMDCT(Inverse Modified Discrete Cosine Transform)를 통하여 시간의 영역의 신호로 변환하기 전에 입력 데이터 X(k)를 가공하여 음정과 속도를 변화시킬 수 있도록 하여 계산량 및 메모리의 사용을 줄일 수 있도록 한 IMDCT 입력신호를 이용한 오디오 신호의 음정 및 속도 가변 장치 및 방법에 관한 것으로, 주파수 추출을 위한 샘플의 수를 결정하는 윈도우 크기 결정 단계; 선택된 윈도우 크기로 IMDCT 과정에 필요한 입력 데이터 X(k)의 주파수(k), 위상, 진폭을 추출하는 단계; 추출된 주파수를 변환하고, 변환된 주파수와 추출된 위상과 진폭을 이용하여 IMDCT 입력 데이터를 재구성하는 단계;IMDCT 처리 및 보간을 하여 코딩된 오디오 신호를 출력하는 단계;를 포함한다.

Description

IMDCT 입력신호를 이용한 오디오 신호의 음정 및 속도 가변 장치 및 방법
본 발명은 오디오 신호의 음정 및 속도가변에 관한 것으로, 구체적으로 IMDCT(Inverse Modified Discrete Cosine Transform)를 통하여 시간영역의 신호로 변환하기 전에 IMDCT 입력 데이터 X(k)를 가공하여 음정을 변화시킬 수 있도록 하여 계산량 및 메모리의 사용을 줄일 수 있도록 한 IMDCT 입력신호를 이용한 오디오 신호의 음정 및 속도 가변 장치 및 방법에 관한 것이다.
일반적으로 많은 양의 오디오 데이터를 CD, 하드디스크, 이동저장매체에 저장하거나 유무선 방식으로 전송하기 위해서 오디오 데이터를 압축한다. 오디오 데이터 압축방법에는 시간영역에서 압축하는 방식과 주파수 영역에서 압축하는 방식이 있다.
주파수 영역에서 압축하는 방식은 압축율이 높을 뿐만 아니라 음질도 좋으므로 시간영역의 오디오 신호를 주파수영역으로 변환하여 심리음향모델과 기타의 방식을 이용하여 압축한다.
MP3(MPEG 3)나 AAC(Advanced Audio Coding) 방식은 오디오 신호를 주파수 영역에서 압축하는 방식을 사용하고 있다. 사람이 오디오 신호를 청취하기 위해서는 압축된 데이터를 풀어서 압축되기 전 주파수영역의 신호로 복원하고 다시 시간영역의 신호로 변환해야 한다.
시간영역으로 변환된 오디오 신호를 정상속도보다 빨리 재생하면 음정(tone)이 높아지고, 느리게 재생하면 음정이 낮아진다. 따라서 음정변화 없이 정상속도보다 빠르거나 느리게 재생하기 위하여 SOLA(Synchronous Overlap and Add)와 같은 방법을 이용하여 음정을 변화시킨다. 음정이 변환된 신호를 정상속도로 재생하면 변환된 음정이 그대로 재생된다. 음정이 변환된 오디오 신호를 재생속도에 따라 음정변화 없이 재생 시키거나, 음정과 재생속도를 동시에 가변시킬 수 있다.
종래 기술의 경우 주파수영역에서 압축된 오디오 신호의 음정이나 속도를 가변시키고자 하는 경우에는 도 1에서와 같이, 시간영역의 신호로 일단 변환시킨 후 음정이나 속도를 가변 시킨다. 이러한 과정에서 시간영역에서 음정이나 속도를 가변 시키기 때문에 추가적인 계산이 요구되고 계산과정의 데이터를 저장하기 위하여 많은 메모리가 필요하게 된다.
MP3와 AAC방식에서 시간영역의 오디오 신호를 주파수영역의 신호로 압축하는 과정은 도 2에서와 같다.
이 방식들에서 오디오 신호는 필터 뱅크(filter bank)(21)를 통하여 여러 대역의 주파수 성분으로 분해되고, 필터 뱅크(21)에서 분해된 성분은 MDCT 블록(22)에서 MDCT(Modified Discrete Cosine Transform)을 통하여 시간영역에서 주파수영역으로 변환된다.
주파수영역으로 변환된 신호는 양자화부(23)에서 양자화(quantization)되고 코딩부(24)에서 코딩(coding)을 통하여 노이즈가 적고 손실이 적은 형태로 압축된다. 압축된 데이터는 비트스트림 엔코딩부(25)에서 사이드 정보(side information)와 함께 비트스트림(bitstream)으로 만들어져 저장되거나 전송된다.
주파수영역으로 압축된 신호를 시간영역으로 변환하는 일반적인 과정은 도 3에서와 같다.
비트스트림 디코딩부(31)에서 부호화된 비트스트림으로부터 사이드 정보와 압축된 데이터를 분리한다.
사이드 정보는 복호화 방법에 대한 정보를 포함하고 있으므로 주파수 영역으로 압축된 데이터를 IMDCT 블록(32)에서 IMDCT(Inverse Modified Discrete Cosine Transform)를 통하여 시간 영역으로 변환할 수 있다.
MP3나 AAC 방식으로 압축된 오디오 데이터는 여러 주파수 대역으로 나누어 압축하고 복원하므로 복원된 데이터를 합성 필터 뱅크(33)를 통하여 재구성하여 시간영역의 오디오 신호로 변환된다. 종래 기술의 경우에는 도 3에서와 같은 과정을 통하여 시간영역의 신호로 완전하게 변환시킨 후 오디오 신호의 음정 변화와 속도를 변화시키는 단계를 거치게 된다.
시간영역의 오디오 신호 x(n)를 MDCT과정을 거쳐 주파수 정보 X(k)로 변환시켜는 과정은 수학식 1과 같다. 수학식 1에서 w(n)은 윈도우 함수(window function)이고 수학식 2과 같이 표시되며, 수학식 1, 2에서 N은 분석하는 윈도우 크기를 의미한다. (이하 모든 수학식에서 N은 분석 윈도우의 크기를 의미한다.) 분석 윈도우 내에 있는 시간영역의 신호 x(n)에 MDCT을 하면 수학식 1과 같은 주파수 영역의 데이터 X(k)를 얻을 수 있다.
주파수 영역으로 변환된 정보 X(k)를 시간 영역의 오디오 신호로 변환하는 IMDCT 과정은 수학식 3과 같다. 이러한 과정에서 주파수 정보 X(k)를 가공하여 오디오 신호의 주파수를 변화시켜 음정을 변화시키는 효과를 얻을 수 있다.
수학식 1
Figure PCTKR2013008905-appb-M000001
수학식 2
Figure PCTKR2013008905-appb-M000002
수학식 3
Figure PCTKR2013008905-appb-M000003
종래기술에서는 주파수영역의 오디오 신호를 시간영역의 신호로 변환시킨 후 음정이나 속도를 변화시키기 때문에 많은 계산량이 요구되고 계산과정의 데이터를 저장하기 위하여 많은 메모리가 필요하게 된다.
본 발명은 종래의 기술에서 주파수 영역으로 변환된 오디오 신호의 음정과 음정 변화없이 재생속도를 가변시키는 과정에서 많은 계산량과 메모리가 요구되는 문제점을 해결하기 위한 것으로, IMDCT 과정에서 오디오 신호의 음정을 용이하게 변화시키기 위해서 IMDCT 입력 데이터 X(k)에서 다양한 주파수성분, 각 주파수의 진폭과 위상을 추출하여 IMDCT 입력신호를 이용한 오디오 신호의 음정 및 속도 가변 장치 및 방법을 제공하는데 그 목적이 있다.
본 발명은 주파수 영역의 신호를 시간영역의 신호를 변환하는 단계인 IMDCT(Inverse Modified Discrete Cosine Transform) 과정의 입력 데이터 X(k)를 가공하여 음정을 변화시킬 수 있도록 한 IMDCT 입력신호를 이용한 오디오 신호의 음정 및 속도 가변 장치 및 방법을 제공하는데 그 목적이 있다.
본 발명은 주파수 영역으로 변환된 오디오 신호의 음정 및 속도가변 방법에 있어, IMDCT 과정에서 주파수 변환이 가능하게 IMDCT 입력신호를 적절하게 변환시키는 전처리 장치 및 방법, IMCDT를 통하여 시간영역의 신호로 변환된 오디오의 재생속도를 IMDCT 입력신호 전처리 단계와 연동시켜 음정과 음정변화 없이 재생속도를 변화시킬 수 있게 보간(Interpolation)하여 IMDCT 입력신호를 이용한 오디오 신호의 음정 및 속도 가변 장치 및 방법을 제공하는데 목적이 있다.
본 발명은 MDCT를 통하여 주파수영역으로 변환된 오디오 신호의 주파수 영역 데이터를 IMDCT를 통하여 시간영역의 신호로 변환하기 전에 다양한 주파수 성분을 추출하고 각 주파수 성분의 진폭과 위상을 이용하여 주파수 변환이 가능하게 IMDCT 입력 데이터를 재생성하여 IMDCT 입력신호를 이용한 오디오 신호의 음정 및 속도 가변 장치 및 방법을 제공하는데 그 목적이 있다.
본 발명은 오디오 신호의 압축율을 높이기 위하여 주파수영역으로 변환된 신호의 주파수 정보를 활용하여 음정과 속도를 변화시키는 방법으로 시간영역에서 음정 변환시키는 과정을 생략할 수 있도록 한 IMDCT 입력신호를 이용한 오디오 신호의 음정 및 속도 가변 장치 및 방법을 제공하는데 그 목적이 있다.
본 발명의 목적들은 이상에서 언급한 목적들로 제한되지 않으며, 언급되지 않은 또 다른 목적들은 아래의 기재로부터 당업자에게 명확하게 이해될 수 있을 것이다.
이와 같은 목적을 달성하기 위한 본 발명에 따른 IMDCT 입력신호를 이용한 오디오 신호의 음정 가변 장치는 처리할 샘플의 윈도우 크기를 결정하는 윈도우부;상기 윈도우부에서 결정된 윈도우 크기로 IMDCT 입력 데이터 X(k)의 주파수와 위상을 추출하는 주파수 추출부 및 위상 추출부, 진폭을 추출하는 진폭 추출부;상기 추출된 주파수를 변환하는 주파수 변환부;상기 변환된 주파수, 추출된 위상과 진폭을 이용하여 IMDCT 입력 데이터를 재구성하는 입력 데이터 재구성부;IMDCT를 통하여 상기 재구성된 입력 데이터를 시간 영역으로 변환하는 IMDCT 블록; 상기 IMDCT를 통하여 시간영역으로 변환된 여러 서브밴드 성분을 합성하는 합성 필터뱅크;를 포함하는 것을 특징으로 한다.
여기서, 상기 윈도우부에서 IMDCT 입력 데이터인 X(k)를 분석하고 재구성하는 데 필요한 윈도우를 결정하기 위하여, 분석 대상이 되는 주파수 영역의 윈도우 크기를 다른 영역에 비하여 상대적으로 작게 하는 것을 특징으로 한다.
그리고 상기 윈도우부에서 IMDCT 입력 데이터인 X(k)를 분석하고 재구성 하는데 필요한 윈도우를 결정하기 위하여, 서버밴드와 서버밴드, 프레임과 프레임 경계에서 분석윈도우를 중첩시켜 설정하고, 분석윈도우 내에서 스펙트럼을 계산을 통하여 정수주파수(kin)를 찾아 그 주파수를 중심으로 분석 윈도우를 구성하는 것을 특징으로 한다.
다른 목적을 달성하기 위한 본 발명에 따른 IMDCT 입력신호를 이용한 오디오 신호의 음정 가변 방법은 IMDCT를 통하여 시간의 영역의 신호로 변환하기 전에 IMDCT에 입력되는 데이터(X(k))를 처리하여 음정을 변화시키기 위하여, 주파수 추출을 위한 샘플의 수를 결정하는 윈도우 크기 결정 단계;선택된 윈도우 크기로 IMDCT 과정에 필요한 입력 데이터 X(k)의 주파수(k), 위상, 진폭을 추출하는 단계;추출된 주파수를 변환하고, 변환된 주파수와 추출된 위상과 진폭을 이용하여 IMDCT 입력 데이터를 재구성하는 단계;주파수 영역의 데이터를 시간영역으로 변환하는 IMDCT 단계;IMDCT 과정에서 만들어진 다양한 주파수의 시간영역 신호를 합성하는 합성필터뱅크 단계;를 포함하는 것을 특징으로 한다.
여기서, 상기 IMDCT 입력 데이터 X(k)의 주파수 추출 단계에서,IMDCT 입력 데이터 X(k)의 주파수를 정수부(kin)와 소수부(ε)로 나누어 f = kin+ε로 표시하고, 이웃하는 세 개의 주파수 성분 X(kin-1), X(kin), X(kin+1)을 이용하여 분석하는 윈도우 내에 존재하는 모든 스펙트럼 값을 구하여 그 중 가장 큰 스펙트럼 값(Sk)을 만드는 kin를 정수부 주파수 성분(kin)으로 하는 것을 특징으로 한다.
그리고 상기 주파수 성분의 소수부분 ε을,
Figure PCTKR2013008905-appb-I000001
라고 두면,
Figure PCTKR2013008905-appb-I000002
인 경우에 대해서
Figure PCTKR2013008905-appb-I000003
Figure PCTKR2013008905-appb-I000004
라고 두면,
Figure PCTKR2013008905-appb-I000005
인 경우에 대해서
Figure PCTKR2013008905-appb-I000006
2 종류를 구하고, α와 β 중에서 어느 것을 사용한 것인지의 결정은, 윈도우 내의 가장 큰 주파수 성분 X(kin)의 절대값과 kin의 스펙트럼 값(Sk)의 비율을
Figure PCTKR2013008905-appb-I000007
로 정의하고, 그 비율이 특정 문턱값(threshold) λ0 과 비교하여 작으면 α, 그 외의 경우엔 β를 선택하여 주파수 성분의 소수부분인(ε)을 얻는 것을 특징으로 한다.
그리고 IMDCT 데이터 X(k)의 위상
Figure PCTKR2013008905-appb-I000008
를 추출하기 위해서 추출한 IMDCT 정수부 주파수 성분(kin)을 이용하여 계산하는 것을 특징으로 한다.
그리고 상기 진폭을 추출하는 단계에서, 정수부 주파수 성분(Kin)과 소수부(ε) 주파수 성분을 이용하여 IMDCT 입력 데이터의 진폭 Ak를 구하는 것을 특징으로 한다.
그리고 상기 IMDCT 입력 데이터를 재구성하는 단계에서, 상기 윈도우 선택과정, 주파수 추출과정, 위상 추출과정, 주파수 변환 과정으로부터 얻은 윈도우 크기(N), 주파수(f = kin+ε), 위상(
Figure PCTKR2013008905-appb-I000009
), 변환된 주파수 fshift = f(1+Rf)을 이용하여 IMDCT 입력 X'(k)를 재구성하는 것을 특징으로 한다.
또 다른 목적을 달성하기 위한 본 발명에 따른 IMDCT 입력신호를 이용한 오디오 신호의 음정 및 속도 가변 장치는 처리할 샘플의 윈도우 크기를 결정하는 윈도우부;상기 윈도우부에서 결정된 윈도우 크기로 IMDCT 입력 데이터 X(k)의 주파수와 위상을 추출하는 주파수 추출부 및 위상 추출부, 진폭을 추출하는 진폭 추출부;상기 추출된 주파수를 변환하는 주파수 변환부;상기 변환된 주파수, 추출된 위상과 진폭을 이용하여 IMDCT 입력 데이터를 재구성하는 입력 데이터 재구성부;IMDCT를 통하여 재구성된 입력 데이터를 시간 영역으로 변환하는 IMDCT 블록;IMDCT를 통하여 시간영역으로 변환된 여러 서브밴드 성분을 합성하는 합성 필터뱅크;상기 합성 필터뱅크에서 출력되는 오디오 신호의 샘플링 간격을 조절하여 재생속도와 음정을 변화시키는 보간부(Interpolator);를 포함하는 것을 특징으로 한다.
여기서, 상기 IMDCT 블록에서 오디오 신호의 음정을 변화시키기 위하여 IMDCT 입력 데이터 X(k)로 부터 정현파 성분으로 분해하여 추출한 후, 원하는 만큼 주파수를 변환하여 IMDCT 입력을 재구성하는 것을 특징으로 한다.
그리고 상기 윈도우부에서 IMDCT 입력 데이터인 X(k)를 분석하고 재구성 하는데 필요한 윈도우를 결정하기 위하여, 분석 대상이 되는 주파수 영역의 윈도우 크기를 다른 영역에 비하여 상대적으로 작게 하는 것을 특징으로 한다.
그리고 상기 윈도우부에서의 IMDCT 입력 데이터인 X(k)를 분석하고 재구성 하는데 필요한 윈도우를 결정하는 데 있어 서버밴드와 서버밴드, 프레임과 프레임 경계에서 분석윈도우를 중첩시켜 설정하고, 분석윈도우 내에서 스펙트럼을 계산을 통하여 정수주파수(kin)를 찾아 그 주파수를 중심으로 분석 윈도우를 구성하는 것을 특징으로 한다.
또 다른 목적을 달성하기 위한 본 발명에 따른 IMDCT 입력신호를 이용한 오디오 신호의 음정 및 속도 가변 방법은 IMDCT를 통하여 시간의 영역의 신호로 변환하기 전에 IMDCT에 입력되는 데이터(X(k))를 처리하여 음정을 변화시키기 위하여, 주파수 추출을 위한 샘플의 수를 결정하는 윈도우 크기 결정 단계;선택된 윈도우 크기로 IMDCT 과정에 필요한 입력 데이터 X(k)의 주파수(k), 위상, 진폭을 추출하는 단계;추출된 주파수를 변환하고, 변환된 주파수와 추출된 위상과 진폭을 이용하여 IMDCT 입력 데이터를 재구성하는 단계;IMDCT 처리 및 보간을 하여 코딩된 오디오 신호를 출력하는 단계;를 포함하는 것을 특징으로 한다.
여기서, 상기 IMDCT 처리 과정에서 원하는 만큼 주파수를 변화시키기 위해서 IMDCT 입력 데이터 X(k)의 주파수, 위상, 진폭을 분리하여 방정식 형태로 표시하거나 그 방정식을 look-up 테이블 형태로 저장해두고 사용하는 과정을 포함하는 것을 특징으로 한다.
그리고 IMDCT 입력 데이터 X(k)의 주파수 추출 단계에서, IMDCT 입력 데이터 X(k)의 주파수를 정수부(kin)와 소수부(ε)로 나누어 f = kin+ε로 표시하고, 이웃하는 세 개의 주파수 성분 X(kin-1), X(kin), X(kin+1)을 이용하여 분석하는 윈도우 내 존재하는 모든 스펙트럼 값을 구하여 그 중 가장 큰 스펙트럼 값(Sk)을 만드는 kin를 정수부 주파수 성분(kin)으로 하는 것을 특징으로 한다.
그리고 상기 주파수 성분의 소수부분 ε을,
Figure PCTKR2013008905-appb-I000010
라고 두면,
Figure PCTKR2013008905-appb-I000011
인 경우에 대해서
Figure PCTKR2013008905-appb-I000012
Figure PCTKR2013008905-appb-I000013
라고 두면,
Figure PCTKR2013008905-appb-I000014
인 경우에 대해서
Figure PCTKR2013008905-appb-I000015
2 종류를 구하고, α와 β중에서 어느 것을 사용한 것인지의 결정은, 윈도우 내의 가장 큰 주파수 성분 X(kin)의 절대값과 kin의 스펙트럼 값(Sk)의 비율을
Figure PCTKR2013008905-appb-I000016
로 정의하고, 그 비율이 특정 문턱값(threshold) λ0 과 비교하여 작으면 α, 그 외의 경우엔 β를 선택하여 주파수 성분의 소수부분인(ε)을 얻는 것을 특징으로 한다.
그리고 IMDCT 데이터 X(k)의 cosine 함수의 위상
Figure PCTKR2013008905-appb-I000017
를 추출하기 위해서 추출한 IMDCT 정수부 주파수 성분(kin)을 이용하여 계산하는 것을 특징으로 한다.
그리고 상기 진폭을 추출하는 단계에서,정수부 주파수 성분(Kin)과 소수부(ε) 주파수 성분을 이용하여 IMDCT 입력 데이터의 진폭 Ak를 구하는 것을 특징으로 한다.
그리고 음정변화 없이 재생속도를 변화시키기 위하여, 상기 IMDCT 처리 및 보간을 하여 코딩된 오디오 신호를 출력하는 단계와 연계하여 원래 속도를 1로 할 때 가변속도, 원신호의 샘플링 간격(ts), 새롭게 만들 신호의 샘플링 간격(t's), (원래속도/가변속도) = ts/t's = Rt 관계를 이용하여 Rt를 구한 후 (Rf x Rt) = 1 되게 Rf를 결정한 다음, 상기 추출한 IMDCT 입력 데이터 X(k)의 주파수 성분(k)을 fshift = f(1+Rf) 변화시키는 것을 특징으로 한다.
그리고 음정과 재생속도를 동시에 변화시키려는 경우에는 재생속도로부터 (원래속도/가변속도) = Rt로부터 Rt를 구하고, 변화시키고 싶은 반음의 수 n에 따라 주파수 변화비율 Rfinal = (1±0.06n)을 결정하고, Rfinal = Rf x Rt 관계로부터 IMCDT 전처리 단계의 주파수 변화율 Rf를 결정하여 fshift = f(1+Rf)을 이용하여 주파수를 변화시키는 것을 특징으로 한다.
그리고 상기 IMDCT 입력 데이터를 재구성하는 단계에서, 상기 윈도우 선택과정, 주파수 추출과정, 위상 추출과정, 주파수 변환 과정으로부터 얻은 윈도우 크기(N), 주파수(f = kin+ε), 위상(
Figure PCTKR2013008905-appb-I000018
), 변환된 주파수 fshift = f(1+Rf), 을 이용하여 IMDCT 입력 X'(k)를 구하는 것을 특징으로 한다.
그리고 가변속도에 따라 주파수 변환부의 주파수 변화량을 조절하고, 주파수 변화량에 따라 보간 단계의 샘플링 간격을 조절하는 것을 특징으로 한다.
그리고 상기 샘플링 간격의 조절은, 원래속도/가변속도, 원 신호의 샘플링 간격(ts), 보간에 의해 재생성되는 오디오 신호의 샘플링 간격(t's) 사이에 Rt=(원래속도/가변속도)= ts/t's 관계식이 성립하고, IMDCT 전 단계의 주파수 변환부의 주파수 변화량이 Rf 이라면 최종 음정이 가변속도와 (Rf x f)x Rt 로 결정되는 것을 특징으로 한다.
이와 같은 본 발명에 따른 IMDCT 입력신호를 이용한 오디오 신호의 음정 및 속도 가변 장치 및 방법은 다음과 같은 효과를 갖는다.
첫째, 주파수 영역에서 압축된 오디오 신호를 IMDCT를 통하여 시간의 영역의 신호로 변환하기 전에 주파수 IMDCT 입력 데이터 X(k)를 가공하여 음정을 변화시킬 수 있다.
둘째, 시간영역의 신호로 변환하기 전에 IMDCT 입력 데이터 X(k)를 가공하여 음정을 변화시킴으로써 계산량을 줄이게 되어 시스템의 CPU 부담을 줄여줄 수 있을 뿐만 아니라 소비전력을 줄일 수 있다.
셋째, 시간영역의 음정변환 과정이 불필요하게 됨에 따라 데이터를 저장하는 메모리를 줄일 수 있어 하드웨어 시스템을 값싸게 구성할 수 있다.
넷째, 입력 데이터 X(k)의 주파수 성분 및 위상과 진폭을 분리하여 방정식 형태로 표시하여 IMDCT 과정에서 주파수 변환이 용이하여 계산량을 줄일 수 있다.
다섯째, 시간영역의 신호를 주파수 영역의 신호로 변환하는 과정에서 주파수 영역의 신호에 포함되는 정보를 활용하여 음정을 변화시키는 방법으로 시간영역에서 음정 변환하는 과정을 생략할 수 있다.
여섯째, 주파수 추출 윈도우 크기를 변화시킴으로써 다양한 주파수를 세밀하게 추출할 수 있어 음정변환 음질을 높일 수 있다.
일곱째, 스펙트럼성분이 큰 주파수를 중심으로 윈도우를 구성함으로써 불필요한 윈도우를 제거할 수 있어 계산량을 줄일 수 있다.
여덟째, 서버밴드나 프레임 가장자리에서 윈도우를 중첩시킴으로써 그들 가장자리에 있는 주파수를 추출할 수 있어 음질을 개선할 수 있다.
아홉째, IMDCT 앞 단계에서의 주파수 변환비율과 보간부의 샘플링 간격을 연동함으로써 오디오신호의 음정변화, 속도변화, 음정과 속도를 동시에 변화시킬 수 있다.
도 1은 종래 기술의 주파수영역에서 압축된 오디오 신호의 음정이나 속도를 가변 시키기 위한 구성도
도 2는 종래 기술의 시간영역의 오디오 신호를 주파수 영역의 신호로 변환하기 위한 구성도
도 3은 종래 기술의 주파수 영역으로 변환된 신호를 시간영역으로 변환하는 과정을 나타낸 구성도
도 4는 본 발명에 따른 IMDCT를 이용한 오디오 신호의 음정 가변 장치의 구성도
도 5는 본 발명에 따른 IMDCT를 이용한 오디오 신호의 음정 및 속도 가변 장치의 구성도
도 6은 IMDCT 입력 데이터 X(k)의 한 프레임 내에서 주파수 대역에 따른 윈도우 구성 개념도
도 7은 본 발명에 따른 윈도우에서 주파수 성분, 위상, 진폭 추출과정을 나타내는 플로우 차트
도 8 내지 도 10은 보간(interpolation) 주파수에 따른 원신호의 음정변환 개념도
이하, 본 발명에 따른 IMDCT 입력신호를 이용한 오디오 신호의 음정 및 속도 가변 장치 및 방법의 바람직한 실시 예에 관하여 상세히 설명하면 다음과 같다.
본 발명에 따른 IMDCT 입력신호를 이용한 오디오 신호의 음정 및 속도 가변 장치 및 방법의 특징 및 이점들은 이하에서의 각 실시 예에 대한 상세한 설명을 통해 명백해질 것이다.
도 4는 본 발명에 따른 IMDCT를 이용한 오디오 신호의 음정 가변 장치의 구성도이고, 도 5는 본 발명에 따른 IMDCT를 이용한 오디오 신호의 음정 및 속도 가변 장치의 구성도이다.
본 발명은 IMDCT 단계에서 사용될 입력 신호를 가공하는 전처리 단계를 두어 음정과 속도를 가변시키는 것으로, 본 발명에 따른 IMDCT를 이용한 오디오 신호의 음정 가변 장치는 도 4에서와 같이, 처리할 샘플의 윈도우 크기(처리할 샘플의 개수)를 결정하는 윈도우부(41)와, 상기 윈도우부(41)에서 결정된 윈도우 크기로 IMDCT 입력 데이터 X(k)의 주파수와 위상을 추출하는 주파수 추출부(42) 및 위상 추출부(43), 진폭을 추출하는 진폭 추출부(44)와, 상기 추출된 주파수를 변환하는 주파수 변환부(45)와, 상기 변환된 주파수, 추출된 위상과 진폭을 이용하여 IMDCT 입력 데이터를 재구성하는 입력 데이터 재구성부(46)와, IMDCT를 통하여 재구성된 입력 데이터를 시간 영역으로 변환하는 IMDCT 블록(47)와, IMDCT를 통하여 시간영역으로 변환된 여러 서브밴드 성분을 합성하는 합성 필터뱅크(48)를 포함한다.
그리고 본 발명에 따른 IMDCT를 이용한 오디오 신호의 음정 및 속도 가변 장치는 도 5에서와 같이, 처리할 샘플의 윈도우 크기(처리할 샘플의 개수)를 결정하는 윈도우부(41)와, 상기 윈도우부(41)에서 결정된 윈도우 크기로 IMDCT 입력 데이터 X(k)의 주파수와 위상을 추출하는 주파수 추출부(42) 및 위상 추출부(43), 진폭을 추출하는 진폭 추출부(44)와, 상기 추출된 주파수를 변환하는 주파수 변환부(45)와, 상기 변환된 주파수, 추출된 위상과 진폭을 이용하여 IMDCT 입력 데이터를 재구성하는 입력 데이터 재구성부(46)와, IMDCT를 통하여 재구성된 입력 데이터를 시간 영역으로 변환하는 IMDCT 블록(47)와, IMDCT를 통하여 시간영역으로 변환된 여러 서브밴드 성분을 합성하는 합성 필터뱅크(48)와, 오디오 신호의 샘플링 간격을 조절하여 재생속도를 변화시키는 보간부(Interpolator)(49)를 포함한다.
이와 같은 본 발명에 따른 IMDCT 입력신호를 이용한 오디오 신호의 음정 가변 방법은 IMDCT 단계에서 사용될 입력 신호를 가공하는 전처리 단계를 두어 음정을 가변시키기 위하여, 주파수 성분을 세밀하게 추출하기 위하여 분석 윈도우의 크기를 조절하는 단계와, IMDCT 입력 데이터 X(k)의 주파수를 추출하는 단계와 위상을 추출하는 단계, 진폭을 추출하는 단계와, 음정변화 비율에 따라 추출된 주파수를 변환하는 단계와, 추출된 주파수, 위상, 진폭을 이용하여 IMDCT 입력을 재생성하는 단계와, 주파수 영역의 데이터를 시간영역으로 변환하는 IMDCT 단계와, IMDCT 과정에서 만들어진 다양한 주파수의 시간영역 신호를 합성하는 합성필터뱅크 단계를 포함한다.
그리고 본 발명에 따른 IMDCT를 이용한 오디오 신호의 음정 및 속도 가변 방법은 IMDCT 단계에서 사용될 입력 신호를 가공하는 전처리 단계를 두어 음정과 속도를 가변시키기 위하여, 주파수 성분을 세밀하게 추출하기 위하여 분석 윈도우의 크기(처리할 샘플의 개수)를 조절하는 단계와, IMDCT 입력 데이터 X(k)의 주파수를 추출하는 단계와 위상을 추출하는 단계, 진폭을 추출하는 단계와, 음정변화 비율에 따라 추출된 주파수를 변환하는 단계와, 추출된 주파수, 위상, 진폭을 이용하여 IMDCT 입력을 재구성하는 단계와, 주파수 영역의 데이터를 시간영역으로 변환하는 IMDCT 단계와, IMDCT 과정에서 만들어진 다양한 주파수의 시간영역 신호를 합성하는 합성필터뱅크 단계와, 재생속도 변화비율에 따라 샘플링 간격을 조절하여 재생속도를 변화시키는 단계를 포함한다.
여기서, IMDCT 과정에서 원하는 만큼 주파수를 변화시키는데 있어, 계산량을 줄이고 주파수 변화량에 비례하게 IMDCT 전 단계에서 입력 데이터 X(k)의 주파수, 위상, 진폭을 분리하여 방정식 형태로 표시하거나 룩업(look-up) 테이블 형태로 저장하는 과정을 포함한다.
그리고 IMDCT를 이용하여 압축된 데이터를 음정 변환 없이 단순히 복호화하는 경우에는 본 발명에서 제안하는 윈도우 과정, 주파수 추출, 진폭 추출, 입력 데이터 재구성, 주파수 변환 과정을 거치지 않고 부호화된 주파수 영역의 데이터 X(k)를 IMDCT 과정과 합성필터뱅크를 거쳐 복호화하면 된다.
윈도우부(41)의 윈도우잉(windowing) 과정은 다음과 같다.
입력신호의 음정을 변화시키기 위해서는 주파수 영역의 IMDCT 입력 데이터 X(k)의 주파수 성분을 가능한 세밀하게 추출하여 주파수 성분의 손실을 줄여야 한다. 추출의 정확도는 몇(N) 개의 X(k) 샘플을 모아서 추출할 것인가를 결정하는 윈도우의 크기(N)에 의하여 결정된다.
주파수 성분을 추출하는 방식은 N개의 샘플 윈도우 내에서 가장 큰 주파수 성분을 추출하는 방식이다. 추출하는 주파수가 중복되지 않은 범위 내에서 가능한 윈도우를 크게 잡는 것이 계산량을 줄이는 측면에서 좋다. 윈도우 크기를 작게 하면 주파수 성분을 세밀하게 추출할 수 있으나 계산량이 많아지는 문제점이 있다.
오디오 신호의 주파수 대역은 20 KHz로 알려져 있으나, 오디오 신호의 주파수 특성을 분석해보면 저주파 영역의 주파수 성분이 고주파 영역의 주파수 성분보다 많다. 계산량도 줄이고 세밀하게 주파수 성분을 추출하기 위하여 도 6에서와 같이 IMDCT 입력 데이터 X(k)에서 주파수 추출할 때 저주파 영역의 윈도우 크기를 작게 하고 고주파 영역의 윈도우를 크게 한다.
상기 IMDCT 단계에서 사용되는 입력 데이터 X(k)는 수학식 1과 같이 복잡한 형태이지만, 본 발명은
Figure PCTKR2013008905-appb-I000019
을 사용하며, 여기서, f와
Figure PCTKR2013008905-appb-I000020
는 IMDCT 입력 데이터로부터 추출한 특정 주파수와 위상을 의미하며, k는 MDCT에서 주파수 인덱스, Ak는 주파수 인덱스의 진폭(amplitude)이다.
그리고 상기 IMDCT 처리 과정에서 원하는 만큼 주파수를 변화시키기 위해서 IMDCT 입력 데이터 X(k)의 주파수 성분과 그 주파수의 진폭과 위상을 분리하여 간단한 방정식 형태로 표시하는 과정을 포함한다.
그리고 상기 주파수 추출단계에서, IMDCT 입력 데이터의 주파수 성분을 f = kin+ε 으로 하여 주파수 f를 정수부(kin)와 소수부(ε)로 나누고, 윈도우 내에서 이웃하는 세 개의 주파수 성분 X(k-1), X(k), X(k+1)을 이용하여 스펙트럼 값을 수학식 4를 이용하여 구하여 윈도우 내에서 가장 큰 스펙트럼 값을 가지는 X(k)의 k를 정수 주파수 성분 kin으로 한다.
수학식 4
Figure PCTKR2013008905-appb-M000004
보다 정확한 주파수 성분의 값을 알기 위해 소수부 주파수 성분을 정수부 주파수(kin) 성분의 절대값과 스펙트럼 값의 비를 문턱 값(λ0)과 비교하여 X(kin±1) 이나 X(kin±2) 를 사용하여 소수부를 계산한다.
그리고 상기에서 추출한 주파수의 위상을 정수부 주파수(kin)와 정수부 주파수 성분 X(kin)과 이웃하는 주파수 성분 X(kin-1)을 이용하여 구하고, 상기 단계에서 추출한 주파수(f = kin+ε), X(kin), X(kin-1), 소수부 주파수 (ε), 윈도우 크기(N)을 이용하여 추출한 주파수의 진폭을 구한다.
그리고 상기와 같이 추출한 IMDCT 입력 데이터를 구성하는 각 주파수와 그 주파수의 위상과 진폭을 사용하여 음정을 가변하기 위해서 음정변화에 대응되게 주파수 f를 변화시켜 변환주파수(fshift)를 fshift = f(1+Rf) 으로 표시하고, 여기서 Rf은 주파수 변환비율이며 양의 값은 음정을 높이고 음의 값은 음정을 낮추는 경우이다.
본 발명은 IMDCT 앞 단계에서의 주파수 변환을 통한 음정 변화와 보간부의 샘플링 간격을 연동시켜 속도가변, 음정가변, 음정과 재생속도 동시가변을 수행할 수 있다.
그리고 윈도우를 구성함에 있어 서버밴드와 서버밴드, 프레임과 프레임 경계에 존재하는 정수주파수 성분을 추출하기 위해서 서버밴드와 프레임을 중첩시켜 윈도우를 설정한다.
서버밴드나 프레임에서 수학식 4를 이용하여 스펙트럼을 계산하여 정수주파수 성분을 찾아 그 값이 큰 몇 개의 정수주파수(kin)를 중심으로 일정수의 샘플로 구성된 분석 윈도우를 구성한다.
그리고 하나의 프레임이나 서브밴드에서 몇 개의 정수주파수 성분을 선택할 것인가 하는 문제는 프레임이나 서브밴드를 구성하는 샘플의 수에 따라 다르다.
MP3 방식과 같은 경우에는 하나의 서브밴드 내에서 5개 이내의 정수주파수 kin 로 윈도우를 결정한다.
시간영역의 오디오 신호를 MDCT를 통하여 주파수 영역의 신호로 변환하면 IMDCT 입력으로 사용될 X(k)는 수학식 1과 같이 윈도우 크기(N) 만큼의 항들의 합으로 표현된다. MP3나 AAC와 같은 방식과 같이 윈도우가 클 경우 수학식 5와 같은 단일 주파수(f)의 정현파를 MDCT 하면 수학식 6과 같이 X(k)는 하나의 항으로 근사화 할 수 있다.
수학식 6를 분석해보면 IMDCT 입력은 MDCT 주파수 인덱스(k)와 그 주파수 인덱스의 위상(
Figure PCTKR2013008905-appb-I000021
)과 진폭(Ak), MDCT의 입력으로 사용된 신호의 단일 주파수(f)로 표현됨을 알 수 있다. 시간영역의 복잡한 모양의 오디오 신호도 결국 여러 주파수의 정현파 합으로 생각할 수 있으므로 IMDCT 입력으로 사용될 X(k)의 각 주파수 성분 즉 인덱스(k)에 대한 정보(진폭, 주파수, 위상)를 추출하여 추출된 주파수를 변화 시켜 IMDCT 단계를 거치면 음정을 변화시킬 수 있다.
수학식 5
Figure PCTKR2013008905-appb-M000005
수학식 6
Figure PCTKR2013008905-appb-M000006
그리고 주파수 추출부(42) 및 위상 추출부(43), 진폭 추출부(44) 에서의 주파수, 위상, 진폭 추출과정은 도 7에서와 같다.
IMDCT 입력 데이터 X(k)의 주파수를 정수부(kin)와 소수부(ε)로 나누어 f = kin+ε 으로 표시할 수 있다. 이웃하는 세 개의 주파수 성분 X(kin-1), X(kin), X(kin+1)을 이용하여 분석할 윈도우 내 존재하는 모든 스펙트럼 값(61)을 구하여 그 중 가장 큰 스펙트럼 값(Sk)을 만드는 주파수 인덱스 k를 정수부 주파수 성분 kin으로 한다.(62)
소수부 주파수 성분을 정수부 주파수 성분 X(kin)의 바로 이웃하는 성분 X(kin±1)을 이용하여 구할 것인가, 아니면 그 다음 성분 X(kin±2)를 이용하여 구할 것인 가를 정하기 위하여 정수부 주파수 인덱스 kin과 kin을 중심으로 한 스펙트럼 값의 비
Figure PCTKR2013008905-appb-I000022
를 구한다.(63)
정수부 주파수 성분과 바로 인접한 성분 X(kin±1)을 이용하여 수학식 7을 이용하여 소수부 주파수성분(ε1)을 계산한다.(64, 65) 정수부 주파수 성분 X(kin)에서 2만큼 떨어진 주파수 성분 X(kin±2)를 이용하여 수학식 8을 이용하여 또 다른 소수부 주파수 성분(ε2)를 구한다. (66, 67)
수학식 7
Figure PCTKR2013008905-appb-M000007
수학식 8
Figure PCTKR2013008905-appb-M000008
정수부 주파수 kin과 kin을 중심으로 한 스펙트럼 값의 비
Figure PCTKR2013008905-appb-I000023
가 특정 문턱 값보다 작으면 수학식 7을 이용하여 구한 소수부 주파수성분(ε1)을 선택하고, 클 경우는 수학식 8을 이용하여 구한 소수부(ε2)를 선택한다. (68, 69)
IMDCT 입력 데이터 X(k)의 추출된 정수부 주파수(kin)와 소수부 주파수(ε)을 이용하여 주파수(f)를 f = kin +ε 와 같이 구한다. (71)
정수부 IMDCT 입력 데이터 X(k)의 cosine 함수의 위상(
Figure PCTKR2013008905-appb-I000024
)은 정수부 주파수 성분 X(kin)과 바로 인접한 성분 X(kin±1)과 소수부 주파수 성분(ε)을 이용하여 수학식 9를 이용하여 구한다. (70)
수학식 9
Figure PCTKR2013008905-appb-M000009
그리고 본 발명에 따른 진폭 추출부에서의 진폭 추출 과정은 다음과 같다. (72) 수학식 6의 X(k)에 kin과 kin-1을 각 각 대입하여 수학식 10, 수학식 11을 얻는다. 수학식 10과 수학식 11을 이용하여 간단한 조작을 통하여 수학식 12와 같은 진폭(Ak)를 얻을 수 있다. 수학식 10, 11에서 f는 추출된 주파수 정보 (kin+ε)를 사용한다.
수학식 10
Figure PCTKR2013008905-appb-M000010
수학식 11
Figure PCTKR2013008905-appb-M000011
수학식 12
Figure PCTKR2013008905-appb-M000012
이상의 과정은 하나의 윈도우에서 그 윈도우를 대표하는 주파수, 위상, 진폭을 구하였다. MDCT 방식을 이용하는 MP3, AAC, AC-3 방식에 따라 하나의 프레임에는 위와 같은 분석이 필요한 윈도우가 여러 개 있다. 본 발명의 개념을 활용하여 방식에 따라 분석 윈도우의 크기를 적절하게 조절하여 하나의 프레임을 구성하는 다양한 주파수(f), 위상(
Figure PCTKR2013008905-appb-I000025
), 진폭(Ak)을 분석할 수 있다.
하나의 프레임에 대한 주파수 성분의 추출이 완료되면, 재생속도 변화 없이 음정을 변화시키는 경우 즉 원래 오디오 신호의 샘플링 간격과 보간부의 샘플링 간격이 동일한 경우에는 추출된 주파수를 fshift = f(1+Rf)와 같이 변화시킨다. 여기서 Rf는 주파수 변화율이다. 일반적으로 반음을 올리거나 내릴 경우 원래 주파수의 6% 만큼 주파수의 변화가 있으므로 n-반음정을 올리거나 내리기 위해서는 fshift = f{1±(0.06)n}와 같이 주파수를 변화시키면 된다.
음정과 재생속도를 동시에 가변 시켜는 경우는 재생속도로부터 (원래속도/가변속도) = Rt 로 부터 Rt를 구하고, 변화시키고 싶은 반음의 수 n에 따라 주파수 변화비율 Rfinal = (1±0.06n)을 결정하고, Rfinal = Rf x Rt로 부터 주파수 변화율 Rf = (1±0.06n)/ Rt 결정하여 최종적으로 fshift = f(1+Rf)을 구한다.
그리고 입력 데이터 재구성부(46)의 입력 데이터 재구성(regenerating IMDCT input data) 과정은 다음과 같다.
상기 과정을 거쳐 추출한 주파수 성분(f = kin+ε), 음정변화에 따른 주파수 변화율(Rf), 위상(
Figure PCTKR2013008905-appb-I000026
), 정수 주파수 성분(kin), 진폭 정보(Ak)를 이용하여 수학식 13과 같이 IMDCT 입력 데이터 X'(k)를 다시 생성한다. 수학식 13에서 k는 IMDCT 영역의 주파수 성분이다.
수학식 13
Figure PCTKR2013008905-appb-M000013
여기서
Figure PCTKR2013008905-appb-I000027
상기 수학식 13을 이용하여 여러 윈도우로 구성된 프레임을 각 각의 윈도우에 대하여 IMDCT 과정을 거치면 그 프레임에 대하여 음정이 변화된 시간영역의 오디오신호를 얻을 수 있다. 이러한 개념을 사용하여 시간영역의 오디오신호를 필요에 따라 적절하게 보간을 하면 음정가변, 속도 가변, 음정 및 속도 동시가변 효과를 얻을 수 있다.
보간 단계에서 속도 및 음정변화 변화 개념을 도 8 내지 도 10을 이용하여 상세히 설명하면 다음과 같다.
도 8을 음정 변화가 안된 원래의 신호라고 가정한다.
도 9는 상기 IMDCT 과정을 통하여 주파수변환 즉 음정변환이 된 신호이다.
도 8과 도 9에서 T0/Tsh를 주파수 변환비율 Rf라고 한다. Rf가 1보다 크게 되면 원래 오디오 신호보다 주파수가 올라가 음정이 높아진다.
도 9와 도 10에서 신호를 샘플링 하는 간격 ts와, t's의 비 ts/ t's를 샘플링 간격의 비 Rt라고 정의한다. 샘플링 간격의 비 Rt는 재생속도의 비 (원래속도/가변속도)와 같이 생각해도 된다.
여기서 재생속도가 느린 경우는 (원래속도/가변속도)가 1보다 큰 경우이다.
샘플링 간격을 짧게 하면 주어진 시간에 많은 샘플 데이터를 얻을 수 있어, 음정을 변화시키지 않은 신호에 대하여 샘플링 간격의 비 Rt 1보다 크게하면 느리게 재생되고 음정이 낮아진다. 이러한 개념을 사용하면 샘플링 간격의 비를 조절함으로써 재생속도와 음정을 바꿀 수 있다.
원래의 주파수(f)가 상기 IMDCT 전처리과정과 IMDCT를 통하여 Rf 비율로 주파수 변환이 일어났다면 보간단계 전의 주파수는 (Rf x f)가 된다. 변환된 신호의 샘플링 간격과 보간 단계의 샘플링 간격의 비가 Rt 이라면 보간 단계를 거친 최종 신호의 주파수는 (Rf x f) x Rt가 된다.
상기 IMDCT를 이용한 오디오 신호의 음정 및 속도 가변장치에서 음정변화 없이 재생속도를 변화시켜려면 (원래속도/가변속도) = ts/t's= Rt로 부터 Rt를 구한 후 (Rf x Rt) = 1 되게 IMDCT 전처리 단계의 주파수 변화율 Rf를 정하면 된다. 그리고 원신호의 샘플링 간격 ts는 이미 알고 있는 값이므로 음정변화 없이 재생속도를 변화시킬 수 있는 t's를 구할 수 있다.
상기 IMDCT를 이용한 오디오 신호의 음정 및 속도 가변장치에서 음정을 변화시키고자 하는 경우에는 음정변화비율 Rf를 IMDCT 전처리 단계의 주파수 변화율 Rf로 정하고 샘플링 간격 비율은 변화시키지 않고 재생하면 된다.
Rf가 1보다 크면 음정이 높아지고 1보다 작으면 음정이 낮아진다. 일반적으로 반음은 주파수 측면에서 ±6% 변화를 가져오므로 변화시키려는 반음의 수(n)에 따라 주파수 변환율을 Rf = (1±0.06n) 형태로 결정하여 수학식 13을 이용하여 IMDCT 입력 X'(k)을 재생성한다.
상기 IMDCT를 이용한 오디오 신호의 음정 및 속도 가변장치에서 음정과 재생속도를 동시에 변화시키려는 경우에는 재생속도로부터 (원래속도/가변속도) = Rt 로부터 Rt를 구하고, 변화시키고 싶은 반음의 수 n에 따라 주파수 변화비율 Rfinal = (1±0.06n)을 결정하고 Rfinal = Rf x Rt로부터 IMCDT 전처리 단계의 주파수 변화율 Rf를 결정하여 수학식 13을 이용하여 IMDCT 입력 X'(k)을 재생성하고 재생성할 신호의 샘플링 간격(t's)은 (원래속도/가변속도) = ts/t's = Rt 을 이용하여 구한다.
이와 같은 본 발명에 따른 IMDCT 입력신호를 이용한 오디오 신호의 음정 및 속도 가변 장치 및 방법은 IMDCT(Inverse Modified Discrete Cosine Transform)를 통하여 시간의 영역의 신호로 변환하기 전에 입력되는 주파수 데이터 X(k)를 가공하여 음정을 변화시킬 수 있도록 한 것으로, IMDCT 과정에서 원하는 만큼 주파수를 용이하게 변화시키기 위해서는 입력 데이터 X(k)의 주파수 성분과 진폭을 분리하여 방정식 형태로 표시하는 과정을 포함하고, IMDCT 과정을 거쳐 시간영역으로 변화된 오디오 신호의 샘플링 간격을 조절하여 음정, 재생속도, 음정 및 재생속도 동시 가변이 가능하게 하는 보간 과정을 포함하여, 계산량 및 메모리의 사용을 줄일 수 있도록 한 것이다.
상기 개념을 이용하면 음정변화뿐만 아니라 임의의 주파수와 임의의 재생속도를 얻을 수 있다.
이상에서의 설명에서와 같이 본 발명의 본질적인 특성에서 벗어나지 않는 범위에서 변형된 형태로 본 발명이 구현되어 있음을 이해할 수 있을 것이다.
그러므로 명시된 실시 예들은 한정적인 관점이 아니라 설명적인 관점에서 고려되어야 하고, 본 발명의 범위는 전술한 설명이 아니라 특허청구 범위에 나타나 있으며, 그와 동등한 범위 내에 있는 모든 차이점은 본 발명에 포함된 것으로 해석되어야 할 것이다.
본 발명에 따른 IMDCT 입력신호를 이용한 오디오 신호의 음정 및 속도 가변 장치 및 방법은 IMDCT(Inverse Modified Discrete Cosine Transform)를 통하여 시간의 영역의 신호로 변환하기 전에 입력되는 주파수 데이터 X(k)를 가공하여 음정을 변화시킬 수 있도록 한 것으로, IMDCT 과정을 거쳐 시간영역으로 변화된 오디오 신호의 샘플링 간격을 조절하여 음정, 재생속도, 음정 및 재생속도 동시 가변이 가능하다.

Claims (24)

  1. 처리할 샘플의 개수를 결정하는 윈도우부;
    상기 윈도우부에서 결정된 윈도우 크기로 IMDCT 입력 데이터 X(k)의 주파수와 위상을 추출하는 주파수 추출부 및 위상 추출부, 진폭을 추출하는 진폭 추출부;
    상기 추출된 주파수를 변환하는 주파수 변환부;
    상기 변환된 주파수, 추출된 위상과 진폭을 이용하여 IMDCT 입력 데이터를 재구성하는 입력 데이터 재구성부;
    IMDCT(Inverse Modified Discrete Cosine Transform)를 통하여 상기 재구성된 입력 데이터를 시간 영역으로 변환하는 IMDCT 블록;
    상기 IMDCT를 통하여 시간영역으로 변환된 여러 서브밴드 성분을 합성하는 합성 필터뱅크;를 포함하는 것을 특징으로 하는 IMDCT 입력신호를 이용한 오디오 신호의 음정 가변 장치.
  2. 제 1 항에 있어서, 상기 윈도우부에서 IMDCT 입력 데이터인 X(k)를 분석하고 재구성하는 데 필요한 윈도우를 결정하기 위하여,
    분석 대상이 되는 주파수 영역의 윈도우 크기를 다른 영역에 비하여 상대적으로 작게 하는 것을 특징으로 하는 IMDCT 입력신호를 이용한 오디오 신호의 음정 가변 장치.
  3. 제 1 항에 있어서, 상기 윈도우부에서 IMDCT 입력 데이터인 X(k)를 분석하고 재구성 하는데 필요한 윈도우를 결정하기 위하여,
    서버밴드와 서버밴드, 프레임과 프레임 경계에서 분석윈도우를 중첩시켜 설정하고,
    분석윈도우 내에서 스펙트럼을 계산을 통하여 정수주파수(kin)를 찾아 그 주파수를 중심으로 분석 윈도우를 구성하는 것을 특징으로 하는 IMDCT 입력신호를 이용한 오디오 신호의 음정 가변 장치.
  4. IMDCT를 통하여 시간의 영역의 신호로 변환하기 전에 IMDCT에 입력되는 데이터(X(k))를 처리하여 음정을 변화시키기 위하여,
    주파수 추출을 위한 샘플의 수를 결정하는 윈도우 크기 결정 단계;
    선택된 윈도우 크기로 IMDCT 과정에 필요한 입력 데이터 X(k)의 주파수(k), 위상, 진폭을 추출하는 단계;
    추출된 주파수를 변환하고, 변환된 주파수와 추출된 위상과 진폭을 이용하여 IMDCT 입력 데이터를 재구성하는 단계;
    주파수 영역의 데이터를 시간영역으로 변환하는 IMDCT 단계;
    IMDCT 과정에서 만들어진 다양한 주파수의 시간영역 신호를 합성하는 합성필터뱅크 단계;를 포함하는 것을 특징으로 하는 IMDCT 입력신호를 이용한 오디오 신호의 음정 가변 방법.
  5. 제 4 항에 있어서, 상기 IMDCT 입력 데이터 X(k)의 주파수 추출 단계에서,
    IMDCT 입력 데이터 X(k)의 주파수를 정수부(kin)와 소수부(ε)로 나누어 f = kin+ε로 표시하고, 이웃하는 세 개의 주파수 성분 X(kin-1), X(kin), X(kin+1)을 이용하여 분석하는 윈도우 내 존재하는 모든 스펙트럼 값을 구하여 그 중 가장 큰 스펙트럼 값(Sk)을 만드는 kin를 정수부 주파수 성분(kin)으로 하는 것을 특징으로 하는 IMDCT 입력신호를 이용한 오디오 신호의 음정 가변 방법.
  6. 제 5 항에 있어서, 상기 주파수 성분의 소수부분 ε을,
    Figure PCTKR2013008905-appb-I000028
    라고 두면,
    Figure PCTKR2013008905-appb-I000029
    인 경우에 대해서
    Figure PCTKR2013008905-appb-I000030
    Figure PCTKR2013008905-appb-I000031
    라고 두면,
    Figure PCTKR2013008905-appb-I000032
    인 경우에 대해서
    Figure PCTKR2013008905-appb-I000033
    2 종류를 구하고, α와 β 중에서 어느 것을 사용한 것인지의 결정은, 윈도우 내의 가장 큰 주파수 성분 X(kin)의 절대값과 kin의 스펙트럼 값(Sk)의 비율을
    Figure PCTKR2013008905-appb-I000034
    로 정의하고, 그 비율이 특정 문턱값(threshold) λ0 과 비교하여 작으면 α, 그 외의 경우엔 β를 선택하여 주파수 성분의 소수부분인(ε)을 얻는 특징으로 하는 IMDCT 입력신호를 이용한 오디오 신호의 음정 가변 방법.
  7. 제 4 항에 있어서, IMDCT 데이터 X(k)의 위상
    Figure PCTKR2013008905-appb-I000035
    를 추출하기 위해서 추출한 IMDCT 정수부 주파수 성분(kin)을 이용하여 계산하는 것을 특징으로 하는 IMDCT 입력신호를 이용한 오디오 신호의 음정 가변 방법.
  8. 제 4 항에 있어서, 상기 진폭을 추출하는 단계에서,
    정수부 주파수 성분(Kin)과 소수부(ε) 주파수 성분을 이용하여 IMDCT 입력 데이터의 진폭 Ak를 구하는 것을 특징으로 IMDCT 입력신호를 이용한 오디오 신호의 음정 가변 방법.
  9. 제 4 항에 있어서, 상기 IMDCT 입력 데이터를 재구성하는 단계에서,
    상기 윈도우 선택과정, 주파수 추출과정, 위상 추출과정, 주파수 변환 과정으로부터 얻은 윈도우 크기(N), 주파수(f = kin+ε), 위상(
    Figure PCTKR2013008905-appb-I000036
    ), 변환된 주파수 fshift = f(1+Rf), 을 이용하여 IMDCT 입력 X'(k)를 구하는 것을 특징으로 하는 IMDCT 입력신호를 이용한 오디오 신호의 음정 가변 방법.
  10. 처리할 샘플의 윈도우 크기를 결정하는 윈도우부;
    상기 윈도우부에서 결정된 윈도우 크기로 IMDCT 입력 데이터 X(k)의 주파수와 위상을 추출하는 주파수 추출부 및 위상 추출부, 진폭을 추출하는 진폭 추출부;
    상기 추출된 주파수를 변환하는 주파수 변환부;
    상기 변환된 주파수, 추출된 위상과 진폭을 이용하여 IMDCT 입력 데이터를 재구성하는 입력 데이터 재구성부;
    IMDCT를 통하여 재구성된 입력 데이터를 시간 영역으로 변환하는 IMDCT 블록;
    IMDCT를 통하여 시간영역으로 변환된 여러 서브밴드 성분을 합성하는 합성 필터뱅크;
    상기 합성 필터뱅크에서 출력되는 오디오 신호의 샘플링 간격을 조절하여 재생속도를 변화시키는 보간부(Interpolator);를 포함하는 것을 특징으로 하는 IMDCT 입력신호를 이용한 오디오 신호의 음정 및 속도 가변 장치.
  11. 제 10 항에 있어서, 상기 IMDCT 블록에서 오디오 신호의 음정을 변화시키기 위하여 IMDCT 입력 데이터 X(k)로 부터 정현파 성분으로 분해하여 추출한 후, 원하는 만큼 주파수를 변환하여 IMDCT 입력을 재구성하는 것을 특징으로 하는 IMDCT 입력신호를 이용한 오디오 신호의 음정 및 속도 가변 장치.
  12. 제 10 항에 있어서, 상기 윈도우부에서 IMDCT 입력 데이터인 X(k)를 분석하고 재구성 하는데 필요한 윈도우를 결정하기 위하여,
    분석 대상이 되는 주파수 영역의 윈도우 크기를 다른 영역에 비하여 상대적으로 작게 하는 것을 특징으로 하는 IMDCT 입력신호를 이용한 오디오 신호의 음정 및 속도 가변 장치.
  13. 제 10 항에 있어서, 상기 윈도우부에서의 IMDCT 입력 데이터인 X(k)를 분석하고 재구성 하는데 필요한 윈도우를 결정하는 데 있어 서버밴드와 서버밴드, 프레임과 프레임 경계에서 분석윈도우를 중첩시켜 설정하고,
    분석윈도우 내에서 스펙트럼을 계산을 통하여 정수주파수(kin)를 찾아 그 주파수를 중심으로 분석 윈도우를 구성하는 것을 특징으로 하는 IMDCT 입력신호를 이용한 오디오 신호의 음정 및 속도 가변 장치.
  14. IMDCT를 통하여 시간의 영역의 신호로 변환하기 전에 IMDCT에 입력되는 데이터(X(k))를 처리하여 음정을 변화시키기 위하여,
    주파수 추출을 위한 샘플의 수를 결정하는 윈도우 크기 결정 단계;
    선택된 윈도우 크기로 IMDCT 과정에 필요한 입력 데이터 X(k)의 주파수(k), 위상, 진폭을 추출하는 단계;
    추출된 주파수를 변환하고, 변환된 주파수와 추출된 위상과 진폭을 이용하여 IMDCT 입력 데이터를 재구성하는 단계;
    IMDCT 처리 및 보간을 하여 코딩된 오디오 신호를 출력하는 단계;를 포함하는 것을 특징으로 하는 IMDCT 입력신호를 이용한 오디오 신호의 음정 및 속도 가변 방법.
  15. 제 14 항에 있어서, 상기 IMDCT 처리 과정에서 원하는 만큼 주파수를 변화시키기 위해서 IMDCT 입력 데이터 X(k)의 주파수, 위상, 진폭을 분리하여 방정식 형태로 표시하거나 그 방정식을 look-up 테이블 형태로 저장해두고 사용하는 과정을 포함하는 것을 특징으로 하는 IMDCT 입력신호를 이용한 오디오 신호의 음정 및 속도 가변 방법.
  16. 제 14 항에 있어서, IMDCT 입력 데이터 X(k)의 주파수 추출 단계에서,
    IMDCT 입력 데이터 X(k)의 주파수를 정수부(kin)와 소수부(ε)로 나누어 f = kin+ε로 표시하고, 이웃하는 세 개의 주파수 성분 X(kin-1), X(kin), X(kin+1)을 이용하여 분석하는 윈도우 내 존재하는 모든 스펙트럼 값을 구하여 그 중 가장 큰 스펙트럼 값(Sk)을 만드는 kin를 정수부 주파수 성분(kin)으로 하는 것을 특징으로 하는 IMDCT 입력신호를 이용한 오디오 신호의 음정 및 속도 가변 방법.
  17. 제 16 항에 있어서, 상기 주파수 성분의 소수부분 ε을,
    Figure PCTKR2013008905-appb-I000037
    라고 두면,
    Figure PCTKR2013008905-appb-I000038
    인 경우에 대해서
    Figure PCTKR2013008905-appb-I000039
    Figure PCTKR2013008905-appb-I000040
    라고 두면,
    Figure PCTKR2013008905-appb-I000041
    인 경우에 대해서
    Figure PCTKR2013008905-appb-I000042
    2 종류를 구하고, α와 β중에서 어느 것을 사용한 것인지의 결정은, 윈도우 내의 가장 큰 주파수 성분 X(kin)의 절대값과 kin의 스펙트럼 값의 비율을
    Figure PCTKR2013008905-appb-I000043
    로 정의하고, 그 비율이 특정 문턱값(threshold) λ0 과 비교하여 작으면 α, 그 외의 경우엔 β를 선택하여 주파수 성분의 소수부분인(ε)을 얻는 것을 특징으로 하는 IMDCT 입력신호를 이용한 오디오 신호의 음정 및 속도 가변 방법.
  18. 제 14 항에 있어서, IMDCT 데이터 X(k)의 위상
    Figure PCTKR2013008905-appb-I000044
    를 추출하기 위해서 추출한 IMDCT 정수부 주파수 성분(kin)을 이용하여 계산하는 것을 특징으로 하는 IMDCT 입력신호를 이용한 오디오 신호의 음정 및 속도 가변 방법.
  19. 제 14 항에 있어서, 상기 진폭을 추출하는 단계에서,
    정수부 주파수 성분(Kin)과 소수부(ε) 주파수 성분을 이용하여 IMDCT 입력 데이터의 진폭 Ak를 구하는 것을 특징으로 IMDCT 입력신호를 이용한 오디오 신호의 음정 및 속도 가변 방법.
  20. 제 14 항에 있어서,
    가변속도와 주파수 변화량에 따라 주파수 변환부와 보간부를 연동하여, 주파수 변화량과 샘플링 간격을 조절하는 것을 특징으로 하는 IMDCT 입력신호를 이용한 오디오 신호의 음정 및 속도 가변 방법.
  21. 제 20 항에 있어서, 음정변화 없이 재생속도를 변화시키기 위하여,
    상기 IMDCT 처리 및 보간을 하여 코딩된 오디오 신호를 출력하는 단계와 연계하여 원래 속도를 1로 할 때 가변속도, 원신호의 샘플링 간격(ts), 새롭게 만들 신호의 샘플링 간격(t's), (원래속도/가변속도) = ts/t's = Rt 관계를 이용하여 Rt를 구한 후 (Rf x Rt) = 1 되게 Rf를 결정한 다음,
    상기 추출한 IMDCT 입력 데이터 X(k)의 주파수 성분(k)을 fshift = f(1+Rf) 변화시키는 것을 특징으로 하는 IMDCT 입력신호를 이용한 오디오 신호의 음정 및 속도 가변 방법.
  22. 제 20 항에 있어서, 음정과 재생속도를 동시에 변화시키려는 경우에는 재생속도로부터 (원래속도/가변속도) = Rt로부터 Rt를 구하고, 변화시키고 싶은 반음의 수 n에 따라 주파수 변화비율 Rfinal = (1±0.06n)을 결정하고, Rfinal = Rf x Rt 관계로부터 IMCDT 전처리 단계의 주파수 변화율 Rf를 결정하여 fshift = f(1+Rf)을 이용하여 주파수를 변화시키는 것을 특징으로 하는 IMDCT 입력신호를 이용한 오디오 신호의 음정 및 속도 가변 장치 및 방법.
  23. 제 20 항에 있어서, 상기 IMDCT 입력 데이터를 재구성하는 단계에서,
    상기 윈도우 선택과정, 주파수 추출과정, 위상 추출과정, 주파수 변환 과정으로부터 얻은 윈도우 크기(N), 주파수(f = kin+ε), 위상(
    Figure PCTKR2013008905-appb-I000045
    ), 변환된 주파수 fshift = f(1+Rf), 을 이용하여 IMDCT 입력 X'(k)를 구하는 것을 특징으로 하는 IMDCT 입력신호를 이용한 오디오 신호의 음정 및 속도 가변 방법.
  24. 제 20 항에 있어서, 상기 샘플링 간격의 조절은,
    원래속도/가변속도, 원 신호의 샘플링 간격(ts), 보간에 의해 재생성되는 오디오 신호의 샘플링 간격(t's) 사이에 Rt=(원래속도/가변속도)= ts/t's 관계식이 성립하고, IMDCT 전 단계의 주파수 변환부의 주파수 변화량이 Rf 이라면 최종 음정이 가변속도와 (Rf x f)x Rt 로 결정되는 것을 특징으로 하는 IMDCT 입력신호를 이용한 오디오 신호의 음정 및 속도 가변 방법.
PCT/KR2013/008905 2012-10-04 2013-10-04 Imdct 입력신호를 이용한 오디오 신호의 음정 및 속도 가변 장치 및 방법 Ceased WO2014054918A1 (ko)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
KR1020120110337A KR101333162B1 (ko) 2012-10-04 2012-10-04 Imdct 입력신호를 이용한 오디오 신호의 음정 및 속도 가변 장치 및 방법
KR10-2012-0110337 2012-10-04

Publications (1)

Publication Number Publication Date
WO2014054918A1 true WO2014054918A1 (ko) 2014-04-10

Family

ID=49858505

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/KR2013/008905 Ceased WO2014054918A1 (ko) 2012-10-04 2013-10-04 Imdct 입력신호를 이용한 오디오 신호의 음정 및 속도 가변 장치 및 방법

Country Status (2)

Country Link
KR (1) KR101333162B1 (ko)
WO (1) WO2014054918A1 (ko)

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN115052241A (zh) * 2022-06-17 2022-09-13 厦门理工学院 基于时频域峰谷特征学习的扬声器质量检测方法

Families Citing this family (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN111739544B (zh) * 2019-03-25 2023-10-20 Oppo广东移动通信有限公司 语音处理方法、装置、电子设备及存储介质

Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
KR970050862A (ko) * 1995-12-28 1997-07-29 슈즈이 다케오 음 피치 변환 장치
JP2001102932A (ja) * 1999-09-28 2001-04-13 Sanyo Electric Co Ltd オーディオ信号再生装置およびオーディオ信号再生方法
JP2009501353A (ja) * 2005-07-14 2009-01-15 コーニンクレッカ フィリップス エレクトロニクス エヌ ヴィ オーディオ信号合成

Patent Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
KR970050862A (ko) * 1995-12-28 1997-07-29 슈즈이 다케오 음 피치 변환 장치
JP2001102932A (ja) * 1999-09-28 2001-04-13 Sanyo Electric Co Ltd オーディオ信号再生装置およびオーディオ信号再生方法
JP2009501353A (ja) * 2005-07-14 2009-01-15 コーニンクレッカ フィリップス エレクトロニクス エヌ ヴィ オーディオ信号合成

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN115052241A (zh) * 2022-06-17 2022-09-13 厦门理工学院 基于时频域峰谷特征学习的扬声器质量检测方法

Also Published As

Publication number Publication date
KR101333162B1 (ko) 2013-11-27

Similar Documents

Publication Publication Date Title
RU2487429C2 (ru) Устройство и метод для обработки аудиосигнала, содержащего переходный сигнал
WO2013183928A1 (ko) 오디오 부호화방법 및 장치, 오디오 복호화방법 및 장치, 및 이를 채용하는 멀티미디어 기기
KR101264486B1 (ko) 오디오 신호 스펙트럼의 복수개의 로컬 무게 중심 주파수들을 결정하는 장치 및 방법
JP4031909B2 (ja) 時間領域エイリアシングを効率的に除去する装置及び方法
WO2010005272A2 (ko) 멀티 채널 부호화 및 복호화 방법 및 장치
JP2976860B2 (ja) 再生装置
CN108461081B (zh) 语音控制的方法、装置、设备和存储介质
AU2013366642A1 (en) Generation of a comfort noise with high spectro-temporal resolution in discontinuous transmission of audio signals
US7580761B2 (en) Fixed-size cross-correlation computation method for audio time scale modification
TR201904282T4 (tr) Bağımsız gürültü-doldurma kullanarak iyileştirilmiş bir ses sinyali üretmek için cihaz ve yöntem.
WO2015102452A1 (en) Method and apparatus for improved ambisonic decoding
WO2011055982A2 (ko) 멀티 채널 오디오 신호의 부호화/복호화 장치 및 방법
JP2004198485A (ja) 音響符号化信号復号化装置及び音響符号化信号復号化プログラム
KR101008250B1 (ko) 기지 음향신호 제거방법 및 장치
KR20070070174A (ko) 스케일러블 부호화 장치, 스케일러블 복호 장치 및스케일러블 부호화 방법
WO2014054918A1 (ko) Imdct 입력신호를 이용한 오디오 신호의 음정 및 속도 가변 장치 및 방법
US20050010397A1 (en) Phase locking method for frequency domain time scale modification based on a bark-scale spectral partition
Maher A method for extrapolation of missing digital audio data
WO2015034115A1 (ko) 오디오 신호의 부호화, 복호화 방법 및 장치
WO2021172053A1 (ja) 信号処理装置および方法、並びにプログラム
CN121483262B (zh) 一种基于伪标签信号生成的弱监督目标说话人提取方法和系统
JP3365908B2 (ja) データ変換装置
BR0006912A (pt) Terminal de informação portátil, método de processar dados de áudio, meio de gravação e programa
KR100359988B1 (ko) 실시간 화속 변환 장치
JP2011133568A (ja) 音声処理装置、音声処理方法および音声処理プログラム

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 13843376

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 13843376

Country of ref document: EP

Kind code of ref document: A1