WO2019216192A1 - ピッチ強調装置、その方法、およびプログラム - Google Patents
ピッチ強調装置、その方法、およびプログラム Download PDFInfo
- Publication number
- WO2019216192A1 WO2019216192A1 PCT/JP2019/017155 JP2019017155W WO2019216192A1 WO 2019216192 A1 WO2019216192 A1 WO 2019216192A1 JP 2019017155 W JP2019017155 W JP 2019017155W WO 2019216192 A1 WO2019216192 A1 WO 2019216192A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- pitch
- signal
- time
- time interval
- emphasis
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L21/00—Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
- G10L21/003—Changing voice quality, e.g. pitch or formants
- G10L21/007—Changing voice quality, e.g. pitch or formants characterised by the process used
- G10L21/013—Adapting to target pitch
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L19/00—Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
- G10L19/04—Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis using predictive techniques
- G10L19/26—Pre-filtering or post-filtering
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L21/00—Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
- G10L21/02—Speech enhancement, e.g. noise reduction or echo cancellation
- G10L21/0316—Speech enhancement, e.g. noise reduction or echo cancellation by changing the amplitude
- G10L21/0324—Details of processing therefor
- G10L21/034—Automatic adjustment
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L21/00—Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
- G10L21/02—Speech enhancement, e.g. noise reduction or echo cancellation
- G10L21/0316—Speech enhancement, e.g. noise reduction or echo cancellation by changing the amplitude
- G10L21/0364—Speech enhancement, e.g. noise reduction or echo cancellation by changing the amplitude for improving intelligibility
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L25/00—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
- G10L25/90—Pitch determination of speech signals
Definitions
- the present invention relates to a technique for analyzing and enhancing a pitch component of a sample sequence derived from a sound signal in a signal processing technique such as a sound signal encoding technique.
- Non-Patent Document 1 A technique for performing processing for enhancing pitch components by adding samples and converting the sound into a sound with less sense of incongruity is widely used (for example, Non-Patent Document 1).
- the pitch component is There is a technique in which the process of emphasizing is performed, and in the case of “non-speech”, the process of enhancing the pitch component is not performed.
- Non-Patent Document 1 feels unnatural when listening to the consonant part by performing a process of enhancing the pitch component even for the consonant part having no clear pitch structure. There is a problem of being able to.
- the technique described in Patent Document 1 since the processing for enhancing the pitch component is not performed at all even when the pitch component is present as a signal in the consonant portion, when the consonant portion is heard, There is a problem that it feels unnatural. Further, the technique described in Patent Document 1 frequently causes discontinuity in the sound signal by switching the presence / absence of the pitch emphasis processing between the time interval of the vowel and the time interval of the consonant. There is also a problem of increasing.
- the present invention is for solving these problems, and is a pitch emphasis process with little sense of incongruity even in a consonant time interval, where the consonant time interval and other time intervals are frequently switched. Even if it exists, it aims at implement
- the consonant includes a frictional sound, a plosive sound, a semi-vowel, a nasal sound, and a rubbing sound (see Reference Document 1 and Reference Document 2).
- a pitch emphasizing apparatus performs pitch emphasis processing for each time interval on a signal derived from an input sound signal to obtain an output signal.
- the pitch emphasizing process sets ⁇ to a value larger than 1, and for each time n in the time interval, the number of samples T 0 corresponding to the pitch period of the time interval is past the time n.
- a signal including a signal obtained by adding a signal obtained by multiplying a signal obtained by multiplying a signal obtained by the time ⁇ to the ⁇ th power of the pitch gain ⁇ 0 of the time interval and a predetermined constant B 0 and the signal at the time n as an output signal.
- Including a pitch emphasis unit is included in the pitch emphasizing process.
- the pitch enhancement process when the pitch enhancement process is performed on the audio signal obtained by the decoding process, there is little discomfort even in the time period of the consonant, and the time period of the consonant and other time periods are frequent. Even in the case of switching to, there is an effect that it is possible to realize pitch emphasis processing with less discomfort during listening based on discontinuity.
- the functional block diagram of the pitch emphasis apparatus which concerns on 1st embodiment, 2nd embodiment, 3rd embodiment, and those modifications.
- the figure which shows the example of the processing flow of the pitch emphasis apparatus which concerns on 1st embodiment, 2nd embodiment, 3rd embodiment, and those modifications.
- the functional block diagram of the pitch emphasis apparatus which concerns on another modification.
- FIG. 1 is a functional block diagram of the speech pitch emphasizing apparatus according to the first embodiment, and FIG. 2 shows a processing flow thereof.
- the speech pitch emphasizing apparatus analyzes a signal to obtain a pitch period and a pitch gain, and emphasizes the pitch based on the pitch period and the pitch gain.
- the pitch component is not the pitch gain itself. Multiply the pitch gain by the power of ⁇ .
- the consonant has a property that the periodicity is smaller than that of the vowel, and the pitch gain obtained by analyzing the input signal is smaller in the time interval of the consonant than in the time interval of the vowel.
- this property is used to multiply the pitch component by the ⁇ th power of the pitch gain instead of the pitch gain itself, thereby emphasizing the pitch component in the time interval of the consonant. The degree is made smaller than the time interval of vowels.
- the speech pitch enhancement apparatus includes an autocorrelation function calculation unit 110, a pitch analysis unit 120, a pitch enhancement unit 130, and a signal storage unit 140, and further includes a pitch information storage unit 150 and an autocorrelation function storage.
- a unit 160 and an attenuation coefficient storage unit 180 may be provided.
- the voice pitch emphasis device is, for example, a special configuration in which a special program is read into a known or dedicated computer having a central processing unit (CPU: Central Processing Unit), a main memory (RAM: Random Access Memory), Device.
- the voice pitch emphasizing apparatus executes each process under the control of the central processing unit, for example. Data input to the voice pitch emphasis device and data obtained in each process are stored in, for example, a main storage device, and the data stored in the main storage device is read out to the central processing unit as necessary. Used for other processing.
- At least a part of each processing unit of the speech pitch emphasizing apparatus may be configured by hardware such as an integrated circuit.
- Each storage unit included in the voice pitch emphasizing device can be configured by a main storage device such as RAM (Random Access Memory) or middleware such as a relational database or a key value store.
- a main storage device such as RAM (Random Access Memory) or middleware such as a relational database or a key value store.
- each storage unit is not necessarily provided in the voice pitch emphasizing device, but is constituted by an auxiliary storage device constituted by a semiconductor memory element such as a hard disk, an optical disk, or a flash memory (Flash Memory), and the voice pitch is set. It is good also as a structure provided in the exterior of an emphasis apparatus.
- the main processes performed by the speech pitch emphasizing apparatus of the first embodiment are an autocorrelation function calculation process (S110), a pitch analysis process (S120), and a pitch emphasis process (S130) (see FIG. 2). Since a plurality of hardware resources provided in the pitch emphasizing device are performed in cooperation, the following description will be given for each of the autocorrelation function calculation process (S110), the pitch analysis process (S120), and the pitch emphasis process (S130). This will be described together with the processing.
- the autocorrelation function calculation unit 110 receives a time-domain sound signal (input signal).
- This sound signal is a signal obtained by, for example, compressing and encoding an acoustic signal such as an audio signal with an encoding device to obtain a code, and decoding the code with a decoding device corresponding to the encoding device.
- the autocorrelation function calculation unit 110 receives a sample sequence of sound signals in the time domain of the current frame input to the speech pitch emphasizing device in units of frames (time intervals) having a predetermined time length. When a positive integer indicating the length of the sample sequence of one frame is N, the autocorrelation function calculation unit 110 has N time domain sound signals constituting the sample sequence of the time domain sound signal of the current frame.
- the autocorrelation function calculation unit 110 includes an autocorrelation function R 0 with a time difference of 0 in the sample sequence of the latest L (L is a positive integer) sound signal samples including the input N time domain sound signal samples.
- the autocorrelation function calculated by the autocorrelation function calculation unit 110 in the processing of the current frame that is, the autocorrelation function in the sample sequence by the latest sound signal sample including the sound signal sample in the time domain of the current frame.
- the autocorrelation function of the current frame is also referred to as “the autocorrelation function of the current frame”.
- the autocorrelation function calculated by the autocorrelation function calculation unit 110 in the processing of frame F that is, at the time of frame F including the sound signal sample in the time domain of frame F.
- the autocorrelation function in the sample sequence of the latest sound signal samples is also referred to as “frame F autocorrelation function”.
- the “autocorrelation function” may be simply referred to as “autocorrelation”.
- the speech pitch emphasizing apparatus includes a signal storage unit 140, and the previous frame.
- the latest L ⁇ N sound signal samples input up to now can be stored.
- the autocorrelation function calculation unit 110 receives the latest L ⁇ N sound signal samples stored in the signal storage unit 140 when N time domain sound signal samples of the current frame are input.
- the autocorrelation function calculation unit 110 uses the latest L sound signal samples X 0 , X 1 ,..., X L ⁇ 1 to generate an autocorrelation function R 0 with a time difference of 0 and a plurality of predetermined time differences.
- the autocorrelation functions R ⁇ (1) ,..., R ⁇ (M) for ⁇ (1) are used as a time difference such as ⁇ (1),..., ⁇ (M) or 0 is ⁇
- the autocorrelation function calculating unit 110 calculates the autocorrelation function R ⁇ by the following equation (1), for example.
- the autocorrelation function calculation unit 110 outputs the calculated autocorrelation functions R 0 , R ⁇ (1) ,..., R ⁇ (M) to the pitch analysis unit 120.
- the time differences ⁇ (1),..., ⁇ (M) are candidates for the pitch period T 0 of the current frame obtained by the pitch analysis unit 120 described later.
- ⁇ (1),..., ⁇ (M) are set as integer values from 75 to 320 suitable as sound pitch period candidates.
- R ⁇ in equation (1) a normalized autocorrelation function R ⁇ / R 0 obtained by dividing R ⁇ in equation (1) by R 0 may be obtained.
- the autocorrelation function R ⁇ may be calculated by the equation (1) itself, but the same value as that obtained by the equation (1) may be calculated by another calculation method.
- the autocorrelation function previously frame autocorrelation function obtained by the process of calculating the autocorrelation function of the previous frame (previous frame) with the autocorrelation function storage unit 160 provided in the speech pitch enhancement apparatus.
- R ⁇ (1) ,..., R ⁇ (M) are stored, and the autocorrelation function calculation unit 110 obtains an autocorrelation function (immediately before) obtained by processing the previous frame read from the autocorrelation function storage unit 160.
- the L sound signal samples instead of using the latest L sound signal samples of the input sound signal itself, a signal whose number of samples has been reduced by down-sampling or thinning samples is used for the L sound signal samples.
- the calculation amount may be saved by calculating the autocorrelation function by the same process as described above.
- the M time differences ⁇ (1),..., ⁇ (M) are expressed by half the number of samples when the number of samples is halved, for example. For example, when 8192 sound signal samples with a sampling frequency of 32 kHz are downsampled to 4096 samples with a sampling frequency of 16 kHz, ⁇ (1),..., ⁇ (M) that are candidates for the pitch period T are 37 to 160, which is about half of 75 to 320.
- the signal storage unit 140 stores the latest L ⁇ N sound signal samples at the time after the voice pitch emphasizing apparatus finishes the processing of the pitch emphasizing unit 130 described later for the current frame. Update the stored contents. Specifically, for example, when L> 2N, the signal storage unit 140 sets the oldest N sound signal samples X 0 , X 1 ,... Among the stored L ⁇ N sound signal samples.
- X N-1 is deleted, X N , X N + 1 , ..., X L-N-1 is made X 0 , X 1 , ..., X L-2N-1 and N of the current frame inputted
- the sound signal samples in the time domain are newly stored as X L ⁇ 2N , X L ⁇ 2N + 1 ,..., X L ⁇ N ⁇ 1 . If L ⁇ 2N, the signal storage unit 140 deletes the stored L ⁇ N sound signal samples X 0 , X 1 ,..., X L ⁇ N ⁇ 1 and inputs the current frame that has been input.
- the latest L ⁇ N sound signal samples of the N time domain sound signal samples are newly stored as X 0 , X 1 ,..., X L ⁇ N ⁇ 1 .
- L ⁇ N it is not necessary to include the signal storage unit 140 in the audio pitch emphasizing device.
- the autocorrelation function storage unit 160 calculates the autocorrelation function R ⁇ (1),.
- the stored contents are updated so as to store ⁇ (M) .
- the autocorrelation function storage unit 160 deletes the stored R ⁇ (1) ,..., R ⁇ (M) and calculates the calculated autocorrelation function R ⁇ (1) ,. , R ⁇ (M) is newly stored.
- the autocorrelation function calculation unit 110 uses the L consecutive sound signal samples X 0 , X 1 ,..., X L ⁇ 1 included in the N frames of the current frame to generate an autocorrelation function with a time difference of 0.
- the pitch analysis unit 120 receives the autocorrelation functions R 0 , R ⁇ (1) ,..., R ⁇ (M) of the current frame output from the autocorrelation function calculation unit 110.
- the pitch analysis unit 120 obtains a maximum value among the autocorrelation functions R ⁇ (1) ,..., R ⁇ (M) of the current frame with respect to a predetermined time difference, and self-correlates between the maximum value of the autocorrelation function and the time difference 0.
- the ratio of the correlation function R 0 is obtained as the pitch gain ⁇ 0 of the current frame, and the time difference at which the autocorrelation function is the maximum value is obtained as the pitch period T 0 of the current frame.
- the pitch emphasizing unit 130 receives the pitch period and pitch gain output from the pitch analysis unit 120, and the time domain sound signal (input signal) of the current frame input to the voice pitch emphasizing device, and receives the sound signal of the current frame.
- a sample sequence of the output signal obtained by emphasizing the pitch component corresponding to the pitch period T 0 of the current frame with a degree of emphasis proportional to the pitch gain ⁇ 0 to the power of ⁇ ( ⁇ > 1) with respect to the sample sequence.
- the pitch emphasizing unit 130 performs pitch emphasis processing on the sample sequence of the sound signal of the current frame, using the pitch gain ⁇ 0 of the input current frame and the pitch period T 0 of the input current frame. Specifically, the pitch emphasizing unit 130 applies the following formula (4) to each sample X n (L ⁇ N ⁇ n ⁇ L ⁇ 1) constituting the sample sequence of the input sound signal of the current frame. ) To obtain the output signal X new n , the sample sequence of the output signal of the current frame by N samples X new L ⁇ N ,..., X new L ⁇ 1 .
- ⁇ is a predetermined value larger than 1.
- a in the equation (4) is an amplitude correction coefficient obtained by the following equation (5).
- B 0 is a predetermined value, for example, 3/4.
- the pitch gain ⁇ 0 is usually a value smaller than 1 except in exceptional cases. Further, when a value greater than exceptionally 1 had been determined as pitch gain sigma 0 may be performed pitch enhancement processing of the above formula (4) by replacing the pitch gain sigma 0 to 1. Therefore, the pitch emphasis process of Equation (4) is a process for emphasizing the pitch component considering not only the pitch period but also the pitch gain, and for the pitch component of the frame having a small pitch gain, the pitch of the frame having a large pitch gain is used. This is a process for emphasizing the pitch component by reducing the degree of emphasis over the component.
- the pitch emphasizing unit 130 for each time n in the frame (time interval), the signal X at the time nT 0 that is past the time n by the number of samples T 0 corresponding to the pitch period of the frame including the signal X n.
- a signal including the signal (X n + B 0 ⁇ 0 ⁇ X n ⁇ T — 0 ) obtained by adding the above and the output signal X new n a signal (B 0 ⁇ 0 ⁇ X n-T_0 ) obtained by multiplying n-T_0 , the ⁇ th power ⁇ 0 ⁇ of the pitch gain ⁇ 0 of the frame, and a predetermined constant B 0, and a signal X n at time n
- This pitch emphasis process reduces the sense of discomfort even for consonant frames, and changes in the degree of pitch component emphasis between frames even when the consonant frame and other frames are frequently switched. The effect of reducing the sense of incongruity due to can be obtained.
- the voice pitch emphasizing device of the first modification further includes a pitch information storage unit 150.
- the pitch emphasizing unit 130 receives the pitch period and pitch gain output from the pitch analysis unit 120 and the sound signal in the time domain of the current frame input to the audio pitch emphasizing device, and outputs the sound signal sample sequence of the current frame.
- a sample train of output signals obtained by emphasizing the pitch component corresponding to the pitch period T 0 of the current frame and the pitch component corresponding to the pitch period of the past frame is output.
- the pitch component corresponding to the pitch period T 0 of the current frame is emphasized at a degree of emphasis proportional to the ⁇ power ( ⁇ > 1) of the pitch gain ⁇ 0 of the current frame.
- the pitch period and the pitch gain of s frames before the current frame (s past frames) are expressed as T ⁇ s and ⁇ ⁇ s , respectively.
- the pitch information storage unit 150 stores pitch periods T ⁇ 1 ,..., T ⁇ and pitch gains ⁇ ⁇ 1 ,..., ⁇ ⁇ from the previous frame to ⁇ past frames.
- ⁇ is a predetermined positive integer, for example, 1.
- the pitch emphasizing unit 130 inputs the pitch gain ⁇ 0 of the input current frame, the pitch gain ⁇ - ⁇ of ⁇ past frames read from the pitch information storage unit 150, and the pitch period T of the input current frame. Using 0 and the pitch period T- ⁇ of ⁇ past frames read from the pitch information storage unit 150, the pitch emphasis processing is performed on the sample sequence of the sound signal of the current frame.
- the pitch emphasizing unit 130 applies the following formula to each sample X n (L ⁇ N ⁇ n ⁇ L ⁇ 1) constituting the sample sequence of the input sound signal of the current frame.
- X new n By obtaining the output signal X new n by (6), a sample sequence of the output signal of the current frame by N samples X new L ⁇ N ,..., X new L ⁇ 1 is obtained.
- a in equation (6) is an amplitude correction coefficient obtained from equation (7) below.
- B 0 and B ⁇ are smaller than a predetermined value, for example, 3/4 and 1/4.
- the pitch emphasizing unit 130 applies the following formula to each sample X n (L ⁇ N ⁇ n ⁇ L ⁇ 1) constituting the sample sequence of the input sound signal of the current frame.
- the sample sequence of the output signal of the current frame by N samples X new L ⁇ N ,..., X new L ⁇ 1 is obtained.
- a in equation (8) is an amplitude correction coefficient obtained by the following equation (9).
- B 0 and B ⁇ are smaller than a predetermined value, for example, 3/4 and 1/4.
- the pitch emphasizing process of the first modification is a process of emphasizing the pitch component considering not only the pitch period but also the pitch gain, and the pitch component of the frame having a small pitch gain is more than the pitch component of the frame having a large pitch gain.
- Equation (6) and (8) it is preferable to satisfy B 0 > B ⁇ . However, even if B 0 ⁇ B ⁇ in equations (6) and (8), the pitch period varies between frames. The effect of reducing discontinuity due to is exhibited.
- the amplitude correction coefficient A obtained by the equations (7) and (9) assumes that the pitch period T 0 of the current frame and the pitch period T ⁇ of ⁇ past frames are sufficiently close to each other. Sometimes the energy of the pitch component is preserved before and after pitch enhancement.
- the pitch information storage unit 150 stores the current frame pitch period and pitch gain as the pitch period and pitch gain of the past frame in the processing of the pitch emphasizing unit 130 of the subsequent frame. Update.
- the pitch component corresponding to the pitch period T 0 of the current frame and the pitch component corresponding to the pitch period of one past frame are emphasized with respect to the sound signal sample sequence of the current frame.
- the sample sequence of the output signal is obtained.
- the pitch component corresponding to the pitch period of a plurality of (two or more) frames in the past may be emphasized.
- emphasizing a pitch component corresponding to a pitch period of a plurality of past frames an example of emphasizing a pitch component corresponding to a pitch period of two past frames will be described as different from the first modification. To do.
- the pitch information storage unit 150 stores pitch periods T ⁇ 1 ,..., T ⁇ ,..., T ⁇ and pitch gains ⁇ ⁇ 1 ,. , ⁇ ⁇ ,..., ⁇ ⁇ are stored.
- ⁇ is a predetermined positive integer larger than ⁇ .
- ⁇ is 1 and ⁇ is 2.
- the pitch emphasizing unit 130 inputs the pitch gain ⁇ 0 of the current frame, ⁇ pitch gains ⁇ ⁇ of the past frames read from the pitch information storage unit 150, and ⁇ pieces of pitch gain read from the pitch information storage unit 150.
- Pitch gain ⁇ - ⁇ of the past frame, pitch period T 0 of the input current frame, ⁇ pitch periods T - ⁇ of the past frames read from the pitch information storage unit 150, and pitch information storage unit 150 Is used to perform pitch emphasis processing on the sample sequence of the sound signal of the current frame.
- the pitch component corresponding to the pitch period T 0 of the current frame is emphasized with an emphasis degree proportional to the ⁇ power ( ⁇ > 1) of the pitch gain ⁇ 0 of the current frame, and ⁇ past
- the pitch component corresponding to the pitch period T- ⁇ of the frame is emphasized with a degree of emphasis proportional to the pitch gain ⁇ - ⁇ of the ⁇ past frames, and corresponds to the pitch period T- ⁇ of the ⁇ past frames.
- the pitch emphasizing unit 130 applies the following formula to each sample X n (L ⁇ N ⁇ n ⁇ L ⁇ 1) constituting the sample sequence of the input sound signal of the current frame.
- X new n By obtaining the output signal X new n by (10), a sample sequence of the output signal of the current frame by N samples X new L ⁇ N ,..., X new L ⁇ 1 is obtained.
- a in the equation (10) is an amplitude correction coefficient obtained by the following equation (11).
- B 0 , B ⁇ and B ⁇ are smaller than a predetermined value of 1, for example, 3/4, 3/16 and 1/16.
- the pitch component corresponding to the pitch period T 0 of the current frame is emphasized with a degree of emphasis proportional to the pitch gain ⁇ 0 of the current frame to the ⁇ power ( ⁇ > 1), and ⁇ past
- the pitch component corresponding to the pitch period T ⁇ of the frame is emphasized with the degree of enhancement proportional to the pitch gain ⁇ ⁇ of the ⁇ past frames to the ⁇ th power, and the pitch period T ⁇ of the ⁇ past frames.
- the pitch component corresponding to ⁇ is emphasized with a degree of emphasis proportional to the pitch gain ⁇ ⁇ of the ⁇ past frames to the ⁇ th power.
- the pitch emphasizing unit 130 applies the following formula to each sample X n (L ⁇ N ⁇ n ⁇ L ⁇ 1) constituting the sample sequence of the input sound signal of the current frame.
- the sample sequence of the output signal of the current frame by N samples X new L ⁇ N ,..., X new L ⁇ 1 is obtained.
- a in the equation (12) is an amplitude correction coefficient obtained by the following equation (13).
- B 0 , B ⁇ and B ⁇ are smaller than a predetermined value of 1, for example, 3/4, 3/16 and 1/16.
- the pitch enhancement process of the second modification is a process of emphasizing a pitch component considering not only the pitch period but also the pitch gain, and a frame having a small consonant pitch gain.
- the pitch component is emphasized by lowering the degree of emphasis than the pitch component of a frame with a large pitch gain that is not a consonant, and the pitch component corresponding to the pitch period T 0 of the current frame is emphasized.
- the pitch component corresponding to the pitch period in the past frame is emphasized with a slightly lower degree of emphasis than the pitch component. Even if the pitch emphasis process is performed for each short time interval (frame) by the pitch emphasis process of the second modified example, an effect of reducing discontinuity due to the variation of the pitch period between frames can be obtained.
- B 0 > B ⁇ > B ⁇ is preferable, but in formulas (10) and (12), B 0 ⁇ B ⁇ and B 0 ⁇ B ⁇ Even if ⁇ and B ⁇ ⁇ B ⁇ , the effect of reducing discontinuity due to the variation of the pitch period between frames is exhibited.
- the amplitude correction coefficient A obtained by the equations (11) and (13) includes the pitch period T 0 of the current frame, the pitch period T ⁇ of the past frame ⁇ , and the pitch period T ⁇ of the past frame of the ⁇ number.
- ⁇ is a sufficiently close value
- the energy of the pitch component is stored before and after pitch emphasis.
- the amplitude correction coefficient A is not a value obtained from Equation (5), Equation (7), Equation (9), Equation (11), Equation (11), or Equation (13), but one or more predetermined values. May be used.
- the pitch emphasizing unit 130 may obtain the output signal X new n using an expression that does not include the 1 / A term in the above expression.
- a sample before each pitch period in the sound signal that has passed through the low-pass filter may be used, A process equivalent to a low-pass filter may be performed.
- pitch emphasis processing that does not include the pitch component may be performed. For example, when the pitch gain ⁇ 0 of the current frame is smaller than a predetermined threshold, the pitch component corresponding to the pitch period T 0 of the current frame is not included in the output signal, and the pitch gain of the past frame is the predetermined threshold. If it is smaller, the pitch signal corresponding to the pitch period of the past frame may not be included in the output signal.
- the speech pitch enhancement apparatus is configured as shown in FIG.
- the pitch may be emphasized based on the period and the pitch gain.
- FIG. 4 shows the processing flow.
- the pitch emphasis process (S130) may be performed using the pitch period and pitch gain input to the speech pitch emphasizing apparatus instead of the pitch period and pitch gain output by the pitch analysis unit 120.
- the audio pitch emphasizing device of the first embodiment and the modification thereof can obtain the pitch period and the pitch gain without depending on the frequency of obtaining the pitch period and the pitch gain outside the audio pitch emphasizing device. It is possible to perform pitch emphasis processing in units of frames with a short time length. In the case of the above sampling frequency of 32 kHz, if N is set to 32, for example, pitch emphasis processing can be performed in units of 1 ms frames.
- the present invention may be applied as pitch enhancement processing for linear prediction residuals in a configuration that performs linear prediction synthesis. That is, the present invention may be applied not to the sound signal itself but to a signal derived from a sound signal such as a signal obtained by analyzing or processing the sound signal.
- the program describing the processing contents can be recorded on a computer-readable recording medium.
- a computer-readable recording medium for example, any recording medium such as a magnetic recording device, an optical disk, a magneto-optical recording medium, and a semiconductor memory may be used.
- this program is distributed by selling, transferring, or lending a portable recording medium such as a DVD or CD-ROM in which the program is recorded. Further, the program may be distributed by storing the program in a storage device of the server computer and transferring the program from the server computer to another computer via a network.
- a computer that executes such a program first stores a program recorded on a portable recording medium or a program transferred from a server computer in its storage unit. When executing the process, this computer reads the program stored in its own storage unit and executes the process according to the read program.
- a computer may read a program directly from a portable recording medium and execute processing according to the program. Further, each time a program is transferred from the server computer to the computer, processing according to the received program may be executed sequentially.
- the program is not transferred from the server computer to the computer, and the above-described processing is executed by a so-called ASP (Application Service Provider) type service that realizes a processing function only by an execution instruction and result acquisition. It is good.
- the program includes information provided for processing by the electronic computer and equivalent to the program (data that is not a direct command to the computer but has a property that defines the processing of the computer).
- each device is configured by executing a predetermined program on a computer, at least a part of these processing contents may be realized by hardware.
Landscapes
- Engineering & Computer Science (AREA)
- Quality & Reliability (AREA)
- Human Computer Interaction (AREA)
- Signal Processing (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Computational Linguistics (AREA)
- Physics & Mathematics (AREA)
- Acoustics & Sound (AREA)
- Multimedia (AREA)
- Compression, Expansion, Code Conversion, And Decoders (AREA)
- Stereophonic System (AREA)
- Electrophonic Musical Instruments (AREA)
Abstract
子音の時間区間であっても違和感が少ないピッチ強調処理であり、子音の時間区間とそれ以外の時間区間とが頻繁に切り替わる場合であっても不連続に基づく受聴時の違和感の少ないピッチ強調処理を実現する。ピッチ強調装置は、入力された音信号に由来する信号に対して時間区間毎にピッチ強調処理を施して出力信号を得る。ピッチ強調装置は、ピッチ強調処理として、ηを1より大きい値とし、当該時間区間の各時刻nについて、当該時間区間のピッチ周期に対応するサンプル数T0だけ、当該時刻nよりも過去の時刻の信号と、当該時間区間のピッチ利得σ0のη乗と、所定の定数B0と、を乗算した信号と、当該時刻nの信号と、を加算した信号を含む信号を出力信号として得る処理を行うピッチ強調部を含む。
Description
この発明は、音信号の符号化技術などの信号処理技術において、音信号に由来するサンプル列に対して、そのピッチ成分を分析し、強調する技術に関連する。
一般的に、時系列信号などのサンプル列を非可逆に圧縮符号化した場合、復号時に得られるサンプル列は元のサンプル列とは違った、歪のあるサンプル列となる。特に音信号の符号化においては、この歪が自然音にはないようなパターンを含むことが多く、復号した音信号を受聴した際に不自然に感じられることがある。そこで、自然音の多くがある一定区間で観測した際に音に応じた周期成分、つまりピッチを含むことに着目し、復号により得た音信号の各サンプルに対して、ピッチ周期分だけ過去のサンプルを加算することにより、ピッチ成分を強調する処理を行い、より違和感の少ない音に変換する技術が広く用いられている(例えば非特許文献1)。
また、例えば特許文献1に記載されているように、復号により得た音信号が「音声」であるか「非音声」であるかの情報に基づき、「音声」である場合にはピッチ成分を強調する処理を行い、「非音声」である場合にはピッチ成分を強調する処理を行わない技術もある。
ITU-T Recommendation G.723.1 (05/2006) pp.16-18, 2006
しかしながら、非特許文献1に記載された技術には、明確なピッチ構造をもたない子音部についてもピッチ成分を強調する処理を行ってしまうことにより、子音部を受聴した際に不自然に感じられるという課題がある。一方、特許文献1に記載された技術では、子音部に信号としてはピッチ成分が存在している場合であってもピッチ成分を強調する処理を全く行わないことから、子音部を受聴した際に不自然に感じられるという課題がある。また、特許文献1に記載された技術には、母音の時間区間と子音の時間区間とでピッチ強調処理の有無が切り替わることによって音信号に不連続が頻繁に生じてしまい、受聴時の違和感が増してしまう、という課題もある。
本発明は、これらの課題を解決するためのものであり、子音の時間区間であっても違和感が少ないピッチ強調処理であり、子音の時間区間とそれ以外の時間区間とが頻繁に切り替わる場合であっても不連続に基づく受聴時の違和感の少ないピッチ強調処理を実現することを目的とする。なお、子音は、摩擦音、破裂音、半母音、鼻音、および破擦音を含む(参考文献1、参考文献2参照)。
(参考文献1)古井貞煕著、「音響・音声工学」、近代科学社、1992年、p.99
(参考文献2)斎藤収三、中田和男、「音声情報処理の基礎」、オーム社、1981年、p.38-39
(参考文献1)古井貞煕著、「音響・音声工学」、近代科学社、1992年、p.99
(参考文献2)斎藤収三、中田和男、「音声情報処理の基礎」、オーム社、1981年、p.38-39
上記の課題を解決するために、本発明の一態様によれば、ピッチ強調装置は、入力された音信号に由来する信号に対して時間区間毎にピッチ強調処理を施して出力信号を得る。ピッチ強調装置は、ピッチ強調処理として、ηを1より大きい値とし、当該時間区間の各時刻nについて、当該時間区間のピッチ周期に対応するサンプル数T0だけ、当該時刻nよりも過去の時刻の信号と、当該時間区間のピッチ利得σ0のη乗と、所定の定数B0と、を乗算した信号と、当該時刻nの信号と、を加算した信号を含む信号を出力信号として得る処理を行うピッチ強調部を含む。
本発明によれば、復号処理により得られた音声信号に対してピッチ強調処理を施す場合に、子音の時間区間であっても違和感が少なく、子音の時間区間とそれ以外の時間区間とが頻繁に切り替わる場合であっても不連続に基づく受聴時の違和感の少ないピッチ強調処理を実現することができるという効果を奏する。
以下、本発明の実施形態について、説明する。なお、以下の説明に用いる図面では、同じ機能を持つ構成部や同じ処理を行うステップには同一の符号を記し、重複説明を省略する。以下の説明において、ベクトルや行列の各要素単位で行われる処理は、特に断りが無い限り、そのベクトルやその行列の全ての要素に対して適用されるものとする。
<第一実施形態>
図1は第一実施形態に係る音声ピッチ強調装置の機能ブロック図を、図2はその処理フローを示す。
図1は第一実施形態に係る音声ピッチ強調装置の機能ブロック図を、図2はその処理フローを示す。
図1を参照して、第一実施形態の音声ピッチ強調装置の処理手続きを説明する。第一実施形態の音声ピッチ強調装置は、信号を分析してピッチ周期とピッチ利得を得て、そのピッチ周期とピッチ利得に基づきピッチを強調するものである。本実施形態では、時間区間ごとの入力された音信号に対してピッチ周期に対応するピッチ成分にピッチ利得を乗算したものを用いてピッチ強調処理を施す際に、ピッチ成分にピッチ利得そのものではなく、ピッチ利得のη乗を乗算する。ただし、η>1である。子音には母音に比べて周期性が小さいという性質があり、入力された信号を分析して得られるピッチ利得は、子音の時間区間のほうが母音の時間区間より小さな値となる。なお、このピッチの利得は、例外的な場合を除き、通常1より小さい値である。本実施形態では、上述の課題を解決するために、この性質を利用し、ピッチ成分にピッチ利得そのものではなく、ピッチ利得のη乗を乗算することで、子音の時間区間のピッチ成分の強調の度合いを母音の時間区間よりも小さくする。
第一実施形態の音声ピッチ強調装置は、自己相関関数算出部110とピッチ分析部120とピッチ強調部130と信号記憶部140とを備えるものであり、更にピッチ情報記憶部150と自己相関関数記憶部160と減衰係数記憶部180とを備えてもよい。
音声ピッチ強調装置は、例えば、中央演算処理装置(CPU: Central Processing Unit)、主記憶装置(RAM: Random Access Memory)などを有する公知又は専用のコンピュータに特別なプログラムが読み込まれて構成された特別な装置である。音声ピッチ強調装置は、例えば、中央演算処理装置の制御のもとで各処理を実行する。音声ピッチ強調装置に入力されたデータや各処理で得られたデータは、例えば、主記憶装置に格納され、主記憶装置に格納されたデータは必要に応じて中央演算処理装置へ読み出されて他の処理に利用される。音声ピッチ強調装置の各処理部は、少なくとも一部が集積回路等のハードウェアによって構成されていてもよい。音声ピッチ強調装置が備える各記憶部は、例えば、RAM(Random Access Memory)などの主記憶装置、またはリレーショナルデータベースやキーバリューストアなどのミドルウェアにより構成することができる。ただし、各記憶部は、必ずしも音声ピッチ強調装置がその内部に備える必要はなく、ハードディスクや光ディスクもしくはフラッシュメモリ(Flash Memory)のような半導体メモリ素子により構成される補助記憶装置により構成し、音声ピッチ強調装置の外部に備える構成としてもよい。
第一実施形態の音声ピッチ強調装置が行う主な処理は自己相関関数算出処理(S110)とピッチ分析処理(S120)とピッチ強調処理(S130)であり(図2参照)、これらの処理は音声ピッチ強調装置が備える複数のハードウェア資源が連携して行うものであるので、以下では、自己相関関数算出処理(S110)とピッチ分析処理(S120)とピッチ強調処理(S130)のそれぞれについて、関連する処理と共に説明する。
[自己相関関数算出処理(S110)]
まず、音声ピッチ強調装置が行う自己相関関数算出処理とこれに関連する処理について説明する。
まず、音声ピッチ強調装置が行う自己相関関数算出処理とこれに関連する処理について説明する。
自己相関関数算出部110には、時間領域の音信号(入力信号)が入力される。この音信号は、例えば音声信号などの音響信号を符号化装置で圧縮符号化して符号を得て、その符号化装置に対応する復号装置で符号を復号して得た信号である。自己相関関数算出部110には、所定の時間長のフレーム(時間区間)単位で、音声ピッチ強調装置に入力された現在のフレームの時間領域の音信号のサンプル列が入力される。1フレームのサンプル列の長さを示す正の整数をNとすると、自己相関関数算出部110には、現在のフレームの時間領域の音信号のサンプル列を構成するN個の時間領域の音信号サンプルが入力される。自己相関関数算出部110は、入力されたN個の時間領域の音信号サンプルを含む最新のL個(Lは正の整数)の音信号サンプルによるサンプル列における時間差0の自己相関関数R0及び複数個(M個、Mは正の整数)の所定の時間差τ(1),…,τ(M)それぞれに対する自己相関関数Rτ(1),…,Rτ(M)を算出する。すなわち、自己相関関数算出部110は、現在のフレームの時間領域の音信号サンプルを含む最新の音信号サンプルによるサンプル列における自己相関関数を算出する。
なお、以降では、現在のフレームの処理において自己相関関数算出部110が算出した自己相関関数、すなわち、現在のフレームの時間領域の音信号サンプルを含む最新の音信号サンプルによるサンプル列における自己相関関数、のことを「現在のフレームの自己相関関数」とも呼ぶ。同様に、過去のあるフレームをフレームFとしたとき、フレームFの処理において自己相関関数算出部110が算出した自己相関関数、すなわち、フレームFの時間領域の音信号サンプルを含むフレームFの時点での最新の音信号サンプルによるサンプル列における自己相関関数、のことを「フレームFの自己相関関数」とも呼ぶ。また、「自己相関関数」は単に「自己相関」と呼ぶこともある。LがNより大きい値である場合には、自己相関関数の算出に最新のL個の音信号サンプルを用いるために、音声ピッチ強調装置内には信号記憶部140を備え、1つ前のフレームまでに入力された最新の少なくともL‐N個の音信号サンプルを記憶できるようにしておく。そして、自己相関関数算出部110は、現在のフレームのN個の時間領域の音信号サンプルが入力された際には、信号記憶部140に記憶された最新のL‐N個の音信号サンプルをX0,X1,…,XL-N-1として読み出し、入力されたN個の時間領域の音信号サンプルをXL-N,XL-N+1,…,XL-1とすることにより、最新のL個の音信号サンプルX0,X1,…,XL-1を得る。
そして、自己相関関数算出部110は、最新のL個の音信号サンプルX0,X1,…,XL-1を用いて、時間差0の自己相関関数R0、及び複数個の所定の時間差τ(1),…,τ(M)それぞれに対する自己相関関数Rτ(1),…,Rτ(M)を算出する。τ(1),…,τ(M)や0などの時間差をτとすると、自己相関関数算出部110は、自己相関関数Rτを例えば以下の式(1)で算出する。
自己相関関数算出部110は算出した自己相関関数R0,Rτ(1),…,Rτ(M)をピッチ分析部120に出力する。
なお、この時間差τ(1),…,τ(M)は後述するピッチ分析部120が求める現在のフレームのピッチ周期T0の候補である。例えば、サンプリング周波数32kHzの音声信号を主とする音信号の場合には、音声のピッチ周期の候補として好適な75から320までの整数値をτ(1),…,τ(M)とするなどの実装が考えられる。なお、式(1)のRτに代えて、式(1)のRτをR0で除算した正規化自己相関関数Rτ/R0を求めてもよい。ただし、Lを8192などのピッチ周期T0の候補である75から320に対して十分に大きな値とした場合などには、自己相関関数Rτに代えて正規化自己相関関数Rτ/R0を求めるよりも、以下で説明する演算量を抑えた方法で自己相関関数Rτを算出するほうがよい。
自己相関関数Rτは、式(1)そのもので算出してもよいが、式(1)で求まるのと同じ値を別の算出方法で算出してもよい。例えば、音声ピッチ強調装置内に自己相関関数記憶部160を備えて1つ前のフレーム(直前のフレーム)の自己相関関数を算出する処理で得られた自己相関関数(直前のフレーム自己相関関数)Rτ(1),…,Rτ(M)を記憶しておき、自己相関関数算出部110は、自己相関関数記憶部160から読み出した直前のフレームの処理で得られた自己相関関数(直前のフレーム自己相関関数)Rτ(1),…,Rτ(M)それぞれに、新たに入力された現在のフレームの音信号サンプルの寄与分の加算と、最も過去のフレームの寄与分の減算と、を行うことにより現在のフレームの自己相関関数Rτ(1),…,Rτ(M)を算出するようにしてもよい。これにより、式(1)そのもので算出するよりも自己相関関数の算出に要する演算量を抑えることが可能である。この場合、τ(1),…,τ(M)のそれぞれをτとすると、自己相関関数算出部110は、現在のフレームの自己相関関数Rτを、直前のフレームの処理で得られた自己相関関数Rτ(直前のフレームの自己相関関数Rτ)に対して、以下の式(2)で得られる差分ΔRτ
+を加算し、式(3)で得られる差分ΔRτ
-を減算することにより得る。
また、入力された音信号の最新のL個の音信号サンプルそのものではなく、当該L個の音信号サンプルに対してダウンサンプリングやサンプルの間引きなどを行うことによりサンプル数を減らした信号を用いて、上記と同様の処理により自己相関関数を算出することで演算量を節約してもよい。この場合、M個の時間差τ(1),…,τ(M)は、例えばサンプル数を半分にした際には半分のサンプル数で表現する。例えば、上述したサンプリング周波数32kHzの8192個の音信号サンプルをサンプリング周波数16kHzの4096個のサンプルにダウンサンプリングした場合には、ピッチ周期Tの候補であるτ(1),…,τ(M)は、75から320の約半分である37から160とすればよい。
信号記憶部140は、音声ピッチ強調装置が現在のフレームについての後述するピッチ強調部130の処理までを終えた後に、その時点で最新のL‐N個の音信号サンプルを記憶しておくように記憶内容を更新する。具体的には、例えば、L>2Nの場合、信号記憶部140は、記憶されているL‐N個の音信号サンプルのうちの一番古いN個の音信号サンプルX0,X1,…,XN-1を削除し、XN,XN+1,…,XL-N-1をX0,X1,…,XL-2N-1とし、入力された現在のフレームのN個の時間領域の音信号サンプルをXL-2N,XL-2N+1,…,XL-N-1として新たに記憶する。また、L≦2Nの場合、信号記憶部140は、記憶されているL‐N個の音信号サンプルX0,X1,…,XL-N-1を削除し、入力された現在のフレームのN個の時間領域の音信号サンプルのうちの最新のL‐N個の音信号サンプルをX0,X1,…,XL-N-1として新たに記憶する。なお、L≦Nである場合には、音声ピッチ強調装置内には信号記憶部140を備える必要はない。
また、自己相関関数記憶部160は、自己相関関数算出部110が現在のフレームについての自己相関関数の算出を終えた後に、算出した現在のフレームの自己相関関数Rτ(1),…,Rτ(M)を記憶しておくように記憶内容を更新する。具体的には、自己相関関数記憶部160は、記憶されているRτ(1),…,Rτ(M)を削除し、算出した現在のフレームの自己相関関数Rτ(1),…,Rτ(M)を新たに記憶する。
なお、上述の説明では、最新のL個の音信号サンプルが現在のフレームのN個の音信号サンプルを含む(つまりL≧N)ことを前提としているが、必ずしもL≧Nである必要はなく、L<Nであってもよい。この場合、自己相関関数算出部110は、現在のフレームのN個に含まれる連続したL個の音信号サンプルX0,X1,…,XL-1を用いて、時間差0の自己相関関数R0、及び複数個の所定の時間差τ(1),…,τ(M)それぞれに対する自己相関関数Rτ(1),…,Rτ(M)を算出すればよい。
[ピッチ分析処理(S120)]
次に、音声ピッチ強調装置が行うピッチ分析処理について説明する。
次に、音声ピッチ強調装置が行うピッチ分析処理について説明する。
ピッチ分析部120には、自己相関関数算出部110が出力した現在のフレームの自己相関関数R0,Rτ(1),…,Rτ(M)が入力される。
ピッチ分析部120は、所定の時間差に対する現在のフレームの自己相関関数Rτ(1),…,Rτ(M)の中での最大値を求め、自己相関関数の最大値と時間差0の自己相関関数R0の比を現在のフレームのピッチ利得σ0として得て、また、自己相関関数が最大値となる時間差を現在のフレームのピッチ周期T0として得て、それぞれをピッチ強調部130へ出力する。
[ピッチ強調処理(S130)]
次に、音声ピッチ強調装置が行うピッチ強調処理について説明する。
次に、音声ピッチ強調装置が行うピッチ強調処理について説明する。
ピッチ強調部130は、ピッチ分析部120が出力したピッチ周期とピッチ利得、及び音声ピッチ強調装置に入力された現在のフレームの時間領域の音信号(入力信号)を受け取り、現在のフレームの音信号サンプル列に対し、現在のフレームのピッチ周期T0に対応するピッチ成分を、ピッチ利得σ0のη乗(η>1)に比例した強調の度合いで強調して得た出力信号のサンプル列を出力する。
以下、具体例を説明する。
ピッチ強調部130は、入力された現在のフレームのピッチ利得σ0と、入力された現在のフレームのピッチ周期T0とを用い、現在のフレームの音信号のサンプル列に対するピッチ強調処理を行う。具体的には、ピッチ強調部130は、入力された現在のフレームの音信号のサンプル列を構成する各サンプルXn(L-N≦n≦L-1)に対して、以下の式(4)により出力信号Xnew
nを得ることにより、N個のサンプルXnew
L―N, …, Xnew
L―1による現在のフレームの出力信号のサンプル列を得る。
ただし、ηは1より大きい所定の値である。なお、式(4)のAは、下記の式(5)により求まる振幅補正係数である。
また、B0は予め定めた値であり、例えば3/4である。ピッチ利得σ0は、例外的な場合を除き、通常は1より小さい値である。また、例外的に1より大きな値がピッチ利得σ0として求まってしまった場合には、ピッチ利得σ0を1に置き換えてから上記式(4)のピッチ強調処理を行えばよい。従って、式(4)のピッチ強調処理は、ピッチ周期だけではなくピッチ利得も考慮したピッチ成分を強調する処理であり、かつ、ピッチ利得が小さいフレームのピッチ成分についてはピッチ利得が大きいフレームのピッチ成分よりも強調の度合いを落としてピッチ成分を強調する処理である。
つまり、ピッチ強調部130では、フレーム(時間区間)中の各時刻nについて、信号Xnを含むフレームのピッチ周期に対応するサンプル数T0だけ、時刻nよりも過去の時刻n-T0の信号Xn-T_0と、そのフレームのピッチ利得σ0のη乗σ0
ηと、所定の定数B0と、を乗算した信号(B0σ0
ηXn-T_0)と、時刻nの信号Xnと、を加算した信号(Xn+B0σ0
ηXn-T_0)を含む信号を出力信号Xnew
nとして得る。
このピッチ強調処理により、子音のフレームであっても違和感を低減し、また、子音のフレームとそれ以外のフレームとが頻繁に切り替わる場合であっても、フレーム間におけるピッチ成分の強調の度合いの変動による違和感を低減する効果を得ることができる。
[ピッチ強調処理(S130)の第1変形例]
次に、音声ピッチ強調装置が行うピッチ強調処理の第1変形例とこれに関連する処理について説明する。
次に、音声ピッチ強調装置が行うピッチ強調処理の第1変形例とこれに関連する処理について説明する。
第1変形例の音声ピッチ強調装置は、更にピッチ情報記憶部150を備える。
ピッチ強調部130は、ピッチ分析部120が出力したピッチ周期とピッチ利得、及び音声ピッチ強調装置に入力された現在のフレームの時間領域の音信号を受け取り、現在のフレームの音信号サンプル列に対し、現在のフレームのピッチ周期T0に対応するピッチ成分と、過去のフレームのピッチ周期に対応するピッチ成分と、を強調して得た出力信号のサンプル列を出力する。その際、現在のフレームのピッチ周期T0に対応するピッチ成分については、現在のフレームのピッチ利得σ0のη乗(η>1)に比例した強調の度合いで、強調する。なお、以下の説明において、現在のフレームからみてs個前のフレーム(s個過去のフレーム)のピッチ周期及びピッチ利得をそれぞれT-s及びσ-sと表記する。
ピッチ情報記憶部150には、1つ前のフレームからα個過去のフレームまでのピッチ周期T-1, ..., T-αとピッチ利得σ-1, ...,σ-αとを記憶しておく。ただし、αは、予め定めた正の整数であり、例えば1である。
ピッチ強調部130は、入力された現在のフレームのピッチ利得σ0と、ピッチ情報記憶部150から読み出したα個過去のフレームのピッチ利得σ-αと、入力された現在のフレームのピッチ周期T0と、ピッチ情報記憶部150から読み出したα個過去のフレームのピッチ周期T-αとを用い、現在のフレームの音信号のサンプル列に対するピッチ強調処理を行う。
以下、具体例を説明する。
(ピッチ強調処理の第1変形例の具体例1)
具体例1は、現在のフレームのピッチ周期T0に対応するピッチ成分については、現在のフレームのピッチ利得σ0のη乗(η>1)に比例した強調の度合いで強調し、α個過去のフレームのピッチ周期T-αに対応するピッチ成分については、α個過去のフレームのピッチ利得σ-αに比例した強調の度合いで強調する例である。
(ピッチ強調処理の第1変形例の具体例1)
具体例1は、現在のフレームのピッチ周期T0に対応するピッチ成分については、現在のフレームのピッチ利得σ0のη乗(η>1)に比例した強調の度合いで強調し、α個過去のフレームのピッチ周期T-αに対応するピッチ成分については、α個過去のフレームのピッチ利得σ-αに比例した強調の度合いで強調する例である。
すなわち、この具体例では、ピッチ強調部130は、入力された現在のフレームの音信号のサンプル列を構成する各サンプルXn(L-N≦n≦L-1)に対して、以下の式(6)により出力信号Xnew
nを得ることにより、N個のサンプルXnew
L―N, …, Xnew
L―1による現在のフレームの出力信号のサンプル列を得る。
なお、式(6)のAは、下記の式(7)により求まる振幅補正係数である。
また、B0とB-αは、予め定めた1より小さい値であり、例えば3/4と1/4である。
(ピッチ強調処理の第1変形例の具体例2)
具体例2は、現在のフレームのピッチ周期T0に対応するピッチ成分については、現在のフレームのピッチ利得σ0のη乗(η>1)に比例した強調の度合いで強調し、α個過去のフレームのピッチ周期T-αに対応するピッチ成分については、α個過去のフレームのピッチ利得σ-αのη乗に比例した強調の度合いで強調する例である。
具体例2は、現在のフレームのピッチ周期T0に対応するピッチ成分については、現在のフレームのピッチ利得σ0のη乗(η>1)に比例した強調の度合いで強調し、α個過去のフレームのピッチ周期T-αに対応するピッチ成分については、α個過去のフレームのピッチ利得σ-αのη乗に比例した強調の度合いで強調する例である。
すなわち、この具体例では、ピッチ強調部130は、入力された現在のフレームの音信号のサンプル列を構成する各サンプルXn(L-N≦n≦L-1)に対して、以下の式(8)により出力信号Xnew
nを得ることにより、N個のサンプルXnew
L―N, …, Xnew
L―1による現在のフレームの出力信号のサンプル列を得る。
なお、式(8)のAは、下記の式(9)により求まる振幅補正係数である。
また、B0とB-αは、予め定めた1より小さい値であり、例えば3/4と1/4である。
第1変形例のピッチ強調処理は、ピッチ周期だけではなくピッチ利得も考慮したピッチ成分を強調する処理であり、かつ、ピッチ利得が小さいフレームのピッチ成分についてはピッチ利得が大きいフレームのピッチ成分よりも強調の度合いを落としてピッチ成分を強調する処理であり、かつ、現在のフレームのピッチ周期T0に対応するピッチ成分を強調しつつ、そのピッチ成分より少し強調の度合いを落として過去のフレームでのピッチ周期T-αに対応するピッチ成分も強調する処理である。第1変形例のピッチ強調処理により、短い時間区間(フレーム)ごとにピッチ強調処理を施す場合であっても、フレーム間におけるピッチ周期の変動による不連続性を低減する効果も得ることができる。
なお、式(6),(8)においてはB0>B-αとするのが好ましいが、式(6),(8)においてB0≦B-αとしても、フレーム間におけるピッチ周期の変動による不連続性を低減する効果は奏される。
また、式(7)と式(9)により求まる振幅補正係数Aは、現在のフレームのピッチ周期T0とα個過去のフレームのピッチ周期T-αとが十分に近い値であると仮定したときに、ピッチ成分のエネルギーがピッチ強調前後で保存されるようにするものである。
なお、ピッチ情報記憶部150は、現在のフレームのピッチ周期とピッチ利得を、以降のフレームのピッチ強調部130の処理において過去のフレームのピッチ周期とピッチ利得として用いることができるように、記憶内容を更新する。
[ピッチ強調処理(S130)の第2変形例]
第1変形例では、現在のフレームの音信号サンプル列に対し、現在のフレームのピッチ周期T0に対応するピッチ成分と、過去の1つのフレームのピッチ周期に対応するピッチ成分と、を強調して出力信号のサンプル列を得たが、過去の複数(2つ以上)のフレームのピッチ周期に対応するピッチ成分を強調するようにしてもよい。以下では、過去の複数のフレームのピッチ周期に対応するピッチ成分を強調する一例として、過去の2つのフレームのピッチ周期に対応するピッチ成分を強調する例について、第1変形例と異なる点を説明する。
第1変形例では、現在のフレームの音信号サンプル列に対し、現在のフレームのピッチ周期T0に対応するピッチ成分と、過去の1つのフレームのピッチ周期に対応するピッチ成分と、を強調して出力信号のサンプル列を得たが、過去の複数(2つ以上)のフレームのピッチ周期に対応するピッチ成分を強調するようにしてもよい。以下では、過去の複数のフレームのピッチ周期に対応するピッチ成分を強調する一例として、過去の2つのフレームのピッチ周期に対応するピッチ成分を強調する例について、第1変形例と異なる点を説明する。
ピッチ情報記憶部150には、現在のフレームよりβ個過去のフレームまでのピッチ周期T-1, ..., T-α, ..., T-βとピッチ利得σ-1, ...,σ-α, ...,σ-βとを記憶しておく。ただし、βは、αより大きい予め定めた正の整数である。例えば、αは1であり、βは2である。
ピッチ強調部130は、入力された現在のフレームのピッチ利得σ0と、ピッチ情報記憶部150から読み出したα個過去のフレームのピッチ利得σ-αと、ピッチ情報記憶部150から読み出したβ個過去のフレームのピッチ利得σ-βと、入力された現在のフレームのピッチ周期T0と、ピッチ情報記憶部150から読み出したα個過去のフレームのピッチ周期T-αと、ピッチ情報記憶部150から読み出したβ個過去のフレームのピッチ周期T-βとを用い、現在のフレームの音信号のサンプル列に対するピッチ強調処理を行う。
以下、具体例を説明する。
(ピッチ強調処理の第2変形例の具体例1)
具体例1は、現在のフレームのピッチ周期T0に対応するピッチ成分については、現在のフレームのピッチ利得σ0のη乗(η>1)に比例した強調の度合いで強調し、α個過去のフレームのピッチ周期T-αに対応するピッチ成分については、α個過去のフレームのピッチ利得σ-αに比例した強調の度合いで強調し、β個過去のフレームのピッチ周期T-βに対応するピッチ成分については、β個過去のフレームのピッチ利得σ-βに比例した強調の度合いで強調する例である。
(ピッチ強調処理の第2変形例の具体例1)
具体例1は、現在のフレームのピッチ周期T0に対応するピッチ成分については、現在のフレームのピッチ利得σ0のη乗(η>1)に比例した強調の度合いで強調し、α個過去のフレームのピッチ周期T-αに対応するピッチ成分については、α個過去のフレームのピッチ利得σ-αに比例した強調の度合いで強調し、β個過去のフレームのピッチ周期T-βに対応するピッチ成分については、β個過去のフレームのピッチ利得σ-βに比例した強調の度合いで強調する例である。
すなわち、この具体例では、ピッチ強調部130は、入力された現在のフレームの音信号のサンプル列を構成する各サンプルXn(L-N≦n≦L-1)に対して、以下の式(10)により出力信号Xnew
nを得ることにより、N個のサンプルXnew
L―N, …, Xnew
L―1による現在のフレームの出力信号のサンプル列を得る。
なお、式(10)のAは、下記の式(11)により求まる振幅補正係数である。
また、B0とB-αとB-βは、予め定めた1より小さい値であり、例えば3/4と3/16と1/16である。
(ピッチ強調処理の第2変形例の具体例2)
具体例2は、現在のフレームのピッチ周期T0に対応するピッチ成分については、現在のフレームのピッチ利得σ0のη乗(η>1)に比例した強調の度合いで強調し、α個過去のフレームのピッチ周期T-αに対応するピッチ成分については、α個過去のフレームのピッチ利得σ-αのη乗に比例した強調の度合いで強調し、β個過去のフレームのピッチ周期T-βに対応するピッチ成分については、β個過去のフレームのピッチ利得σ-βのη乗に比例した強調の度合いで強調する例である。
具体例2は、現在のフレームのピッチ周期T0に対応するピッチ成分については、現在のフレームのピッチ利得σ0のη乗(η>1)に比例した強調の度合いで強調し、α個過去のフレームのピッチ周期T-αに対応するピッチ成分については、α個過去のフレームのピッチ利得σ-αのη乗に比例した強調の度合いで強調し、β個過去のフレームのピッチ周期T-βに対応するピッチ成分については、β個過去のフレームのピッチ利得σ-βのη乗に比例した強調の度合いで強調する例である。
すなわち、この具体例では、ピッチ強調部130は、入力された現在のフレームの音信号のサンプル列を構成する各サンプルXn(L-N≦n≦L-1)に対して、以下の式(12)により出力信号Xnew
nを得ることにより、N個のサンプルXnew
L―N, …, Xnew
L―1による現在のフレームの出力信号のサンプル列を得る。
なお、式(12)のAは、下記の式(13)により求まる振幅補正係数である。
また、B0とB-αとB-βは、予め定めた1より小さい値であり、例えば3/4と3/16と1/16である。
第2変形例のピッチ強調処理も、第1変形例のピッチ強調処理と同様に、ピッチ周期だけではなくピッチ利得も考慮したピッチ成分を強調する処理であり、かつ、子音のピッチ利得が小さいフレームのピッチ成分については子音でないピッチ利得が大きいフレームのピッチ成分よりも強調の度合いを落としてピッチ成分を強調する処理であり、かつ、現在のフレームのピッチ周期T0に対応するピッチ成分を強調しつつ、そのピッチ成分より少し強調の度合いを落として過去のフレームでのピッチ周期に対応するピッチ成分も強調する処理である。第2変形例のピッチ強調処理により、短い時間区間(フレーム)ごとにピッチ強調処理を施す場合であっても、フレーム間におけるピッチ周期の変動による不連続性を低減する効果も得ることができる。
なお、式(10),(12)においてはB0>B-α>B-βとするのが好ましいが、式(10),(12)においてB0≦B-αやB0≦B-βやB-α≦B-βとしても、フレーム間におけるピッチ周期の変動による不連続性を低減する効果は奏される。
また、式(11)と式(13)により求まる振幅補正係数Aは、現在のフレームのピッチ周期T0とα個過去のフレームのピッチ周期T-αとβ個過去のフレームのピッチ周期T-βとが十分に近い値であると仮定したときに、ピッチ成分のエネルギーがピッチ強調前後で保存されるようにするものである。
(ピッチ強調処理のその他の変形例)
なお、振幅補正係数Aは、式(5)や式(7)や式(9)や式(11)や式(11)や式(13)により求まる値ではなく、予め定めた1以上の値を用いてもよい。振幅補正係数Aを1とする場合には、ピッチ強調部130は、上記の式中の1/Aの項を含まないようにした式により出力信号Xnew nを得るようにしてもよい。
なお、振幅補正係数Aは、式(5)や式(7)や式(9)や式(11)や式(11)や式(13)により求まる値ではなく、予め定めた1以上の値を用いてもよい。振幅補正係数Aを1とする場合には、ピッチ強調部130は、上記の式中の1/Aの項を含まないようにした式により出力信号Xnew nを得るようにしてもよい。
また、入力された音信号の各サンプルに加算する各ピッチ周期分前のサンプルに基づく値に代えて、例えばローパスフィルタを通した音信号における各ピッチ周期分前のサンプルを用いてもよいし、ローパスフィルタと等価な処理を行ってもよい。
また、ピッチ利得が所定の閾値より小さい場合には、そのピッチ成分を含まないピッチ強調処理を行うようにしてもよい。例えば、現在のフレームのピッチ利得σ0が所定の閾値より小さい場合には、現在のフレームのピッチ周期T0に対応するピッチ成分を出力信号に含めず、過去のフレームのピッチ利得が所定の閾値より小さい場合には、その過去のフレームのピッチ周期に対応するピッチ成分を出力信号に含めない構成としてもよい。
<その他の変形例>
音声ピッチ強調装置外で行われる復号処理などにより各フレームのピッチ周期とピッチ利得を得られている場合には、音声ピッチ強調装置を図3の構成として、音声ピッチ強調装置外で得られたピッチ周期とピッチ利得に基づきピッチを強調してもよい。図4はその処理フローを示す。この場合には、第一実施形態、およびその変形例の音声ピッチ強調装置が備える自己相関関数算出部110やピッチ分析部120や自己相関関数記憶部160を備える必要はなく、ピッチ強調部130が、ピッチ分析部120が出力したピッチ周期とピッチ利得ではなく、音声ピッチ強調装置に入力されたピッチ周期とピッチ利得を用いてピッチ強調処理(S130)を行うようにすればよい。このような構成とすれば、音声ピッチ強調装置自体の演算処理量は第一実施形態、およびその変形例よりも少なくすることが可能である。ただし、第一実施形態、およびその変形例の音声ピッチ強調装置は、音声ピッチ強調装置外のピッチ周期やピッチ利得を得る頻度に依存せずにピッチ周期やピッチ利得を得ることができることから、非常に短い時間長のフレーム単位でのピッチ強調処理を行うことが可能である。上記のサンプリング周波数32kHzの例であれば、Nを例えば32とすれば、1msのフレーム単位でピッチ強調処理を行うことができる。
音声ピッチ強調装置外で行われる復号処理などにより各フレームのピッチ周期とピッチ利得を得られている場合には、音声ピッチ強調装置を図3の構成として、音声ピッチ強調装置外で得られたピッチ周期とピッチ利得に基づきピッチを強調してもよい。図4はその処理フローを示す。この場合には、第一実施形態、およびその変形例の音声ピッチ強調装置が備える自己相関関数算出部110やピッチ分析部120や自己相関関数記憶部160を備える必要はなく、ピッチ強調部130が、ピッチ分析部120が出力したピッチ周期とピッチ利得ではなく、音声ピッチ強調装置に入力されたピッチ周期とピッチ利得を用いてピッチ強調処理(S130)を行うようにすればよい。このような構成とすれば、音声ピッチ強調装置自体の演算処理量は第一実施形態、およびその変形例よりも少なくすることが可能である。ただし、第一実施形態、およびその変形例の音声ピッチ強調装置は、音声ピッチ強調装置外のピッチ周期やピッチ利得を得る頻度に依存せずにピッチ周期やピッチ利得を得ることができることから、非常に短い時間長のフレーム単位でのピッチ強調処理を行うことが可能である。上記のサンプリング周波数32kHzの例であれば、Nを例えば32とすれば、1msのフレーム単位でピッチ強調処理を行うことができる。
なお、以上の説明では、音信号そのものに対してピッチ強調処理を施すことを前提としていたが、非特許文献1に記載されているような線形予測残差に対してピッチ強調処理を行ってから線形予測合成をするような構成における、線形予測残差に対するピッチ強調処理として本発明を適用してもよい。すなわち、本発明を、音信号そのものではなく、音信号に対して分析や加工をして得た信号などの音信号に由来する信号に対して適用してもよい。
本発明は上記の実施形態及び変形例に限定されるものではない。例えば、上述の各種の処理は、記載に従って時系列に実行されるのみならず、処理を実行する装置の処理能力あるいは必要に応じて並列的にあるいは個別に実行されてもよい。その他、本発明の趣旨を逸脱しない範囲で適宜変更が可能である。
<プログラム及び記録媒体>
また、上記の実施形態及び変形例で説明した各装置における各種の処理機能をコンピュータによって実現してもよい。その場合、各装置が有すべき機能の処理内容はプログラムによって記述される。そして、このプログラムをコンピュータで実行することにより、上記各装置における各種の処理機能がコンピュータ上で実現される。
また、上記の実施形態及び変形例で説明した各装置における各種の処理機能をコンピュータによって実現してもよい。その場合、各装置が有すべき機能の処理内容はプログラムによって記述される。そして、このプログラムをコンピュータで実行することにより、上記各装置における各種の処理機能がコンピュータ上で実現される。
この処理内容を記述したプログラムは、コンピュータで読み取り可能な記録媒体に記録しておくことができる。コンピュータで読み取り可能な記録媒体としては、例えば、磁気記録装置、光ディスク、光磁気記録媒体、半導体メモリ等どのようなものでもよい。
また、このプログラムの流通は、例えば、そのプログラムを記録したDVD、CD-ROM等の可搬型記録媒体を販売、譲渡、貸与等することによって行う。さらに、このプログラムをサーバコンピュータの記憶装置に格納しておき、ネットワークを介して、サーバコンピュータから他のコンピュータにそのプログラムを転送することにより、このプログラムを流通させてもよい。
このようなプログラムを実行するコンピュータは、例えば、まず、可搬型記録媒体に記録されたプログラムもしくはサーバコンピュータから転送されたプログラムを、一旦、自己の記憶部に格納する。そして、処理の実行時、このコンピュータは、自己の記憶部に格納されたプログラムを読み取り、読み取ったプログラムに従った処理を実行する。また、このプログラムの別の実施形態として、コンピュータが可搬型記録媒体から直接プログラムを読み取り、そのプログラムに従った処理を実行することとしてもよい。さらに、このコンピュータにサーバコンピュータからプログラムが転送されるたびに、逐次、受け取ったプログラムに従った処理を実行することとしてもよい。また、サーバコンピュータから、このコンピュータへのプログラムの転送は行わず、その実行指示と結果取得のみによって処理機能を実現する、いわゆるASP(Application Service Provider)型のサービスによって、上述の処理を実行する構成としてもよい。なお、プログラムには、電子計算機による処理の用に供する情報であってプログラムに準ずるもの(コンピュータに対する直接の指令ではないがコンピュータの処理を規定する性質を有するデータ等)を含むものとする。
また、コンピュータ上で所定のプログラムを実行させることにより、各装置を構成することとしたが、これらの処理内容の少なくとも一部をハードウェア的に実現することとしてもよい。
Claims (5)
- 入力された音信号に由来する信号に対して時間区間毎にピッチ強調処理を施して出力信号を得るピッチ強調装置であって、
前記ピッチ強調処理として、
ηを1より大きい値とし、当該時間区間の各時刻nについて、当該時間区間のピッチ周期に対応するサンプル数T0だけ当該時刻nよりも過去の時刻の前記信号と、当該時間区間のピッチ利得σ0のη乗と、所定の定数B0と、を乗算した信号と、
当該時刻nの前記信号と、を加算した信号を含む信号を出力信号として得る処理を行うピッチ強調部を含む、
ピッチ強調装置。 - 請求項1に記載のピッチ強調装置であって、
前記ピッチ強調部は、
当該時間区間の各時刻nについて、
前記加算した信号に、
当該時間区間よりもα個過去の時間区間のピッチ周期に対応するサンプル数T-αだけ当該時刻nよりも過去の時刻の前記信号と、当該時間区間よりもα個過去の時間区間のピッチ利得σ-αと、所定の定数B-αと、を乗算した信号
も加算した信号を含む信号を出力信号として得る処理を行うものである
ピッチ強調装置。 - 請求項1に記載のピッチ強調装置であって、
前記ピッチ強調部は、
当該時間区間の各時刻nについて、
前記加算した信号に、
当該時間区間よりもα個過去の時間区間のピッチ周期に対応するサンプル数T-αだけ当該時刻nよりも過去の時刻の前記信号と、当該時間区間よりもα個過去の時間区間のピッチ利得σ-αのη乗と、所定の定数B-αと、を乗算した信号
も加算した信号を含む信号を出力信号として得る処理を行うものである
ピッチ強調装置。 - 入力された音信号に由来する信号に対して時間区間毎にピッチ強調処理を施して出力信号を得るピッチ強調方法であって、
前記ピッチ強調処理として、
ηを1より大きい値とし、当該時間区間の各時刻nについて、当該時間区間のピッチ周期に対応するサンプル数T0だけ当該時刻nよりも過去の時刻の前記信号と、当該時間区間のピッチ利得σ0のη乗と、所定の定数B0と、を乗算した信号と、
当該時刻nの前記信号と、を加算した信号を含む信号を出力信号として得る処理を行うピッチ強調ステップを含む、
ピッチ強調方法。 - 請求項1から請求項3の何れかのピッチ強調装置としてコンピュータを機能させるためのプログラム。
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US17/053,711 US11302340B2 (en) | 2018-05-10 | 2019-04-23 | Pitch emphasis apparatus, method and program for the same |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP2018-091201 | 2018-05-10 | ||
| JP2018091201A JP6962269B2 (ja) | 2018-05-10 | 2018-05-10 | ピッチ強調装置、その方法、およびプログラム |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2019216192A1 true WO2019216192A1 (ja) | 2019-11-14 |
Family
ID=68467446
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/JP2019/017155 Ceased WO2019216192A1 (ja) | 2018-05-10 | 2019-04-23 | ピッチ強調装置、その方法、およびプログラム |
Country Status (3)
| Country | Link |
|---|---|
| US (1) | US11302340B2 (ja) |
| JP (1) | JP6962269B2 (ja) |
| WO (1) | WO2019216192A1 (ja) |
Families Citing this family (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP6962268B2 (ja) * | 2018-05-10 | 2021-11-05 | 日本電信電話株式会社 | ピッチ強調装置、その方法、およびプログラム |
Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPH10143195A (ja) * | 1996-11-14 | 1998-05-29 | Olympus Optical Co Ltd | ポストフィルタ |
| JP2002268690A (ja) * | 2001-03-09 | 2002-09-20 | Mitsubishi Electric Corp | 音声符号化装置、音声符号化方法、音声復号化装置及び音声復号化方法 |
| WO2011086923A1 (ja) * | 2010-01-14 | 2011-07-21 | パナソニック株式会社 | 符号化装置、復号装置、スペクトル変動量算出方法及びスペクトル振幅調整方法 |
Family Cites Families (12)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPH11119800A (ja) * | 1997-10-20 | 1999-04-30 | Fujitsu Ltd | 音声符号化復号化方法及び音声符号化復号化装置 |
| US6141638A (en) * | 1998-05-28 | 2000-10-31 | Motorola, Inc. | Method and apparatus for coding an information signal |
| US7117146B2 (en) * | 1998-08-24 | 2006-10-03 | Mindspeed Technologies, Inc. | System for improved use of pitch enhancement with subcodebooks |
| US7072832B1 (en) * | 1998-08-24 | 2006-07-04 | Mindspeed Technologies, Inc. | System for speech encoding having an adaptive encoding arrangement |
| US6871176B2 (en) * | 2001-07-26 | 2005-03-22 | Freescale Semiconductor, Inc. | Phase excited linear prediction encoder |
| US7065485B1 (en) * | 2002-01-09 | 2006-06-20 | At&T Corp | Enhancing speech intelligibility using variable-rate time-scale modification |
| CN100369111C (zh) * | 2002-10-31 | 2008-02-13 | 富士通株式会社 | 话音增强装置 |
| WO2006098274A1 (ja) * | 2005-03-14 | 2006-09-21 | Matsushita Electric Industrial Co., Ltd. | スケーラブル復号化装置およびスケーラブル復号化方法 |
| US8326614B2 (en) * | 2005-09-02 | 2012-12-04 | Qnx Software Systems Limited | Speech enhancement system |
| KR101475724B1 (ko) * | 2008-06-09 | 2014-12-30 | 삼성전자주식회사 | 오디오 신호 품질 향상 장치 및 방법 |
| US8600737B2 (en) * | 2010-06-01 | 2013-12-03 | Qualcomm Incorporated | Systems, methods, apparatus, and computer program products for wideband speech coding |
| MX352092B (es) * | 2013-06-21 | 2017-11-08 | Fraunhofer Ges Forschung | Aparato y método para mejorar el ocultamiento del libro de códigos adaptativo en la ocultación similar a acelp empleando una resincronización de pulsos mejorada. |
-
2018
- 2018-05-10 JP JP2018091201A patent/JP6962269B2/ja active Active
-
2019
- 2019-04-23 WO PCT/JP2019/017155 patent/WO2019216192A1/ja not_active Ceased
- 2019-04-23 US US17/053,711 patent/US11302340B2/en active Active
Patent Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPH10143195A (ja) * | 1996-11-14 | 1998-05-29 | Olympus Optical Co Ltd | ポストフィルタ |
| JP2002268690A (ja) * | 2001-03-09 | 2002-09-20 | Mitsubishi Electric Corp | 音声符号化装置、音声符号化方法、音声復号化装置及び音声復号化方法 |
| WO2011086923A1 (ja) * | 2010-01-14 | 2011-07-21 | パナソニック株式会社 | 符号化装置、復号装置、スペクトル変動量算出方法及びスペクトル振幅調整方法 |
Also Published As
| Publication number | Publication date |
|---|---|
| US20210090586A1 (en) | 2021-03-25 |
| JP6962269B2 (ja) | 2021-11-05 |
| JP2019197150A (ja) | 2019-11-14 |
| US11302340B2 (en) | 2022-04-12 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| Giacobello et al. | Sparse linear prediction and its applications to speech processing | |
| RU2677453C2 (ru) | Способы, кодер и декодер для линейного прогнозирующего кодирования и декодирования звуковых сигналов после перехода между кадрами, имеющими различные частоты дискретизации | |
| JP2007526691A (ja) | 信号解析及び合成のための適応型混合変換 | |
| KR19980042556A (ko) | 음성부호화방법, 음성복호화방법, 음성부호화장치, 음성복호화장치, 전화장치, 피치변환방법 및 매체 | |
| Dendani et al. | Speech enhancement based on deep AutoEncoder for remote Arabic speech recognition | |
| JP2019090930A (ja) | 音源強調装置、音源強調学習装置、音源強調方法、プログラム | |
| US12106767B2 (en) | Pitch emphasis apparatus, method and program for the same | |
| Kumar et al. | Performance evaluation of a ACF-AMDF based pitch detection scheme in real-time | |
| KR20250169167A (ko) | 비자기회귀 디코딩을 사용한 오디오 생성 | |
| EP2571170B1 (en) | Encoding method, decoding method, encoding device, decoding device, program, and recording medium | |
| CN112088404B (zh) | 基音强调装置、其方法、以及记录介质 | |
| WO2019216192A1 (ja) | ピッチ強調装置、その方法、およびプログラム | |
| CN114495977A (zh) | 语音翻译和模型训练方法、装置、电子设备以及存储介质 | |
| JP6911939B2 (ja) | ピッチ強調装置、その方法、およびプログラム | |
| JP5361565B2 (ja) | 符号化方法、復号方法、符号化器、復号器およびプログラム | |
| Mineo et al. | Improving sign-algorithm convergence rate using natural gradient for lossless audio compression | |
| JPH09127987A (ja) | 信号符号化方法及び装置 | |
| JP2019531505A (ja) | オーディオコーデックにおける長期予測のためのシステム及び方法 | |
| Lee et al. | Speech Enhancement Using Phase‐Dependent A Priori SNR Estimator in Log‐Mel Spectral Domain | |
| Sakka et al. | Using geometric spectral subtraction approach for feature extraction for DSR front-end Arabic system | |
| JP6220610B2 (ja) | 信号処理装置、信号処理方法、プログラム、記録媒体 | |
| HK40057033B (zh) | 在声音信号编码器和解码器中使用的方法、设备和存储器 | |
| JPWO2018225412A1 (ja) | 符号化装置、復号装置、平滑化装置、逆平滑化装置、それらの方法、およびプログラム |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 19798963 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 19798963 Country of ref document: EP Kind code of ref document: A1 |





