EP0801788B1 - Procede de codage de parole a analyse par synthese - Google Patents
Procede de codage de parole a analyse par synthese Download PDFInfo
- Publication number
- EP0801788B1 EP0801788B1 EP96901008A EP96901008A EP0801788B1 EP 0801788 B1 EP0801788 B1 EP 0801788B1 EP 96901008 A EP96901008 A EP 96901008A EP 96901008 A EP96901008 A EP 96901008A EP 0801788 B1 EP0801788 B1 EP 0801788B1
- Authority
- EP
- European Patent Office
- Prior art keywords
- frame
- delays
- open
- delay
- loop
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Expired - Lifetime
Links
- 238000004458 analytical method Methods 0.000 title claims abstract description 67
- 238000003786 synthesis reaction Methods 0.000 title claims abstract description 40
- 230000015572 biosynthetic process Effects 0.000 title claims abstract description 39
- 238000000034 method Methods 0.000 title claims abstract description 32
- 230000007774 longterm Effects 0.000 claims abstract description 60
- 230000001934 delay Effects 0.000 claims abstract description 59
- 230000005284 excitation Effects 0.000 claims abstract description 42
- 230000001419 dependent effect Effects 0.000 claims 1
- 239000013598 vector Substances 0.000 description 28
- 230000004044 response Effects 0.000 description 26
- 238000011002 quantification Methods 0.000 description 24
- 239000011159 matrix material Substances 0.000 description 19
- 238000004364 calculation method Methods 0.000 description 18
- 238000012360 testing method Methods 0.000 description 16
- 230000005540 biological transmission Effects 0.000 description 15
- 238000013139 quantization Methods 0.000 description 14
- 230000008569 process Effects 0.000 description 13
- 230000006870 function Effects 0.000 description 11
- 238000012546 transfer Methods 0.000 description 10
- 238000000354 decomposition reaction Methods 0.000 description 8
- 238000005457 optimization Methods 0.000 description 8
- 230000008901 benefit Effects 0.000 description 7
- 230000003111 delayed effect Effects 0.000 description 7
- 230000003044 adaptive effect Effects 0.000 description 6
- 101100176198 Caenorhabditis elegans nst-1 gene Proteins 0.000 description 4
- 230000000717 retained effect Effects 0.000 description 4
- 230000001755 vocal effect Effects 0.000 description 4
- 101100148606 Caenorhabditis elegans pst-1 gene Proteins 0.000 description 3
- 241000135309 Processus Species 0.000 description 3
- 238000010586 diagram Methods 0.000 description 3
- 238000011160 research Methods 0.000 description 3
- 230000003595 spectral effect Effects 0.000 description 3
- 238000013459 approach Methods 0.000 description 2
- 150000001875 compounds Chemical class 0.000 description 2
- 230000007423 decrease Effects 0.000 description 2
- 239000006185 dispersion Substances 0.000 description 2
- 238000001914 filtration Methods 0.000 description 2
- 238000012545 processing Methods 0.000 description 2
- 230000035945 sensitivity Effects 0.000 description 2
- 230000017105 transposition Effects 0.000 description 2
- 238000011144 upstream manufacturing Methods 0.000 description 2
- UWNXGZKSIKQKAH-UHFFFAOYSA-N Cc1cc(CNC(CO)C(O)=O)c(OCc2cccc(c2)C#N)cc1OCc1cccc(c1C)-c1ccc2OCCOc2c1 Chemical compound Cc1cc(CNC(CO)C(O)=O)c(OCc2cccc(c2)C#N)cc1OCc1cccc(c1C)-c1ccc2OCCOc2c1 UWNXGZKSIKQKAH-UHFFFAOYSA-N 0.000 description 1
- 101000822695 Clostridium perfringens (strain 13 / Type A) Small, acid-soluble spore protein C1 Proteins 0.000 description 1
- 101000655262 Clostridium perfringens (strain 13 / Type A) Small, acid-soluble spore protein C2 Proteins 0.000 description 1
- 101000655256 Paraclostridium bifermentans Small, acid-soluble spore protein alpha Proteins 0.000 description 1
- 101000655264 Paraclostridium bifermentans Small, acid-soluble spore protein beta Proteins 0.000 description 1
- 241000897276 Termes Species 0.000 description 1
- 230000001174 ascending effect Effects 0.000 description 1
- 238000004422 calculation algorithm Methods 0.000 description 1
- 230000008859 change Effects 0.000 description 1
- 238000006243 chemical reaction Methods 0.000 description 1
- 230000006835 compression Effects 0.000 description 1
- 238000007906 compression Methods 0.000 description 1
- 238000012937 correction Methods 0.000 description 1
- 230000003247 decreasing effect Effects 0.000 description 1
- 230000000593 degrading effect Effects 0.000 description 1
- 238000001514 detection method Methods 0.000 description 1
- 230000009977 dual effect Effects 0.000 description 1
- 238000012986 modification Methods 0.000 description 1
- 230000004048 modification Effects 0.000 description 1
- 230000001575 pathological effect Effects 0.000 description 1
- 230000009467 reduction Effects 0.000 description 1
- 238000012552 review Methods 0.000 description 1
- 238000005070 sampling Methods 0.000 description 1
- 238000007493 shaping process Methods 0.000 description 1
- 230000011664 signaling Effects 0.000 description 1
- 238000004088 simulation Methods 0.000 description 1
- 238000006467 substitution reaction Methods 0.000 description 1
- 230000009897 systematic effect Effects 0.000 description 1
- 238000013519 translation Methods 0.000 description 1
- 238000011282 treatment Methods 0.000 description 1
Images
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L19/00—Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
- G10L19/04—Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis using predictive techniques
- G10L19/08—Determination or coding of the excitation function; Determination or coding of the long-term prediction parameters
- G10L19/09—Long term prediction, i.e. removing periodical redundancies, e.g. by using adaptive codebook or pitch predictor
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L25/00—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
- G10L25/93—Discriminating between voiced and unvoiced parts of speech signals
Definitions
- the present invention relates to speech coding using synthetic analysis.
- a linear prediction of the speech signal is carried out to obtain the coefficients of a short-term synthesis filter modeling the transfer function of the vocal tract. These coefficients are transmitted to the decoder, as well as parameters characterizing an excitation to be applied to the short-term synthesis filter.
- further research is carried out on the longer-term correlations of the speech signal in order to characterize a long-term synthesis filter accounting for the pitch of the speech.
- the excitation indeed has a predictable component which can be represented by the past excitation, delayed by TP samples of the speech signal and affected by a gain g P.
- the remaining, unpredictable part of the excitation is called stochastic excitation.
- CELP Code Excited Linear Prediction
- MPLPC Multi-Pulse Linear Prediction Coding
- the stochastic excitation comprises a certain number of pulses whose positions are sought by the coder.
- CELP coders are preferred for low transmission rates, but they are more complex to implement than MPLPC coders.
- an open loop analysis an analysis in closed loop or a combination of both.
- the analysis in open loop requires little computational volume, but its accuracy is limited.
- loop analysis closed requires a lot of calculations, but it's more reliable because it directly contributes to minimizing the difference perceptually balanced between the speech signal and the synthetic signal.
- a loop analysis open is first performed to limit the interval in which the closed loop analyzer will look for the delay prediction. This search interval must nevertheless remain relatively wide because we have to take into account that the delay can vary quickly.
- the invention aims in particular to find a good compromise between the quality of the modeling of the long-term part of the explanation and the complexity of finding the delay correspondent in a speech coder.
- the invention thus proposes a coding method using analysis by synthesis of a speech signal digitized in frames successive divided into nst subframes, including the next steps: linear prediction analysis of the signal speech to determine parameters of a filter short-term synthesis; open loop signal analysis speech to detect voiced signal frames and to determine, for each voiced frame, a degree of voicing signal and a delay search interval of long-term prediction; closed loop predictive analytics of the speech signal to select, for some at minus of the frames of the voiced frames, a delay of long-term prediction contained in the range of research and constituting a parameter of a synthesis filter long-term ; and determination of a stochastic excitation for each subframe, so as to minimize a weighted difference perceptually between the speech signal and the excitation stochastic filtered by long-term synthesis filters and in the short term.
- the open loop analysis step we determines the search interval for each frame voiced so that it contains a number of delays depending on the degree of voicing of said frame.
- the number of delays that are to be tested in closed loop is adaptable to the voicing mode of the frame.
- the width of the search interval will be more weak for the most voiced frames in order to take into account of their greater harmonic stability.
- a speech coder implementing the invention is applicable in various types of speech transmission and / or storage systems using a digital compression technique.
- the speech coder 16 is part of a mobile radio station.
- the speech signal S is a digital signal sampled at a frequency typically equal to 8 kHz.
- the signal S comes from an analog-digital converter 18 receiving the amplified and filtered output signal from a microphone 20.
- the converter 18 puts the speech signal S in the form of successive frames themselves subdivided into nst sub-frames lst samples.
- the speech signal S can also be subjected to conventional shaping treatments such as Hamming filtering.
- the speech coder 16 delivers a binary sequence with a significantly lower bit rate than that of the speech signal S, and addresses this sequence to a channel coder 22 whose function is to introduce redundancy bits into the signal in order to allow detection and / or a correction of any transmission errors.
- the output signal from the channel encoder 22 is then modulated on a carrier frequency by the modulator 24, and the modulated signal is transmitted on the air interface.
- the wall coder 16 is a coder with analysis by synthesis.
- the encoder 16 determines on the one hand parameters characterizing a short-term synthesis filter modeling the vocal tract of the speaker, and on the other hand a sequence excitation which, applied to the short synthesis filter term, provides a synthetic signal constituting a estimation of the speech signal S according to a criterion of perceptual weighting.
- the short-term synthesis filter has a transfer function of the form 1 / A (z), with:
- the coefficients a i are determined by a module 26 for short-term linear prediction analysis of the speech signal S.
- the a i are the linear prediction coefficients of the speech signal S.
- the order q of the linear prediction is typically of the order of 10.
- the methods applicable by module 26 for short-term linear prediction are well known in the field of speech coding.
- Module 26, for example, implements the Durbin-Levinson algorithm (see J. Makhoul: "Linear Prediction: A tutorial review", Proc. IEEE, Vol.63, N ° 4, April 1975, p. 561-580 ).
- the coefficients a i obtained are supplied to a module 28 which converts them into spectral line parameters (LSP).
- the representation of the prediction coefficients a i by LSP parameters is frequently used in speech coders with analysis by synthesis.
- LST t (nst-1) LSP t for sub -frames 0,1,2, ..., nst-1 of the frame t.
- the coefficients a i of the filter 1 / A (z) are then determined, sub-frame by sub-frame from the interpolated LSP parameters.
- the non-quantified LSP parameters are supplied by the module 28 to a module 32 for calculating the coefficients of a perceptual weighting filter 34.
- the coefficients of the perceptual weighting filter are calculated by the module 32 for each subframe after interpolation of the LSP parameters received from the module 28.
- the perceptual weighting filter 34 receives the speech signal S and delivers a perceptually weighted SW signal which is analyzed by modules 36, 38, 40 for determine the excitation sequence.
- the excitation sequence of the short-term filter consists of an excitation predictable by a long-term synthetic filter modeling the pitch of the speech, and an excitement unpredictable stochastic, or innovation sequence.
- Module 36 performs long-term prediction (LTP) in open loop, i.e. it does not contribute directly to the minimization of the weighted error.
- LTP long-term prediction
- the weighting filter 34 intervenes in upstream of the open loop analysis module, but it could otherwise: module 36 could operate directly on the speech signal S or on the signal S cleared of its short-term correlations by a filter transfer function A (z).
- modules 38 and 40 operate in a closed loop, i.e. they directly contribute to minimizing the error perceptually weighted.
- Long-term prediction lag is determined in two steps.
- the analysis module 36 Open loop LTP detects voiced frames from the speech signal and determines, for each voiced frame, a degree of voicing MV and a delay search interval long-term prediction.
- the search interval is defined by a central value represented by its quantification index ZP and by a width in the field of quantification indexes, depending on the degree of voicing MV.
- the module 30 operates the quantization of the LSP parameters which have previously been determined for this frame.
- This quantization is for example vectorial, that is to say it consists in selecting, from one or more predetermined quantization tables, a set of quantized parameters LSP Q which has a minimum distance from the set of parameters LSP provided by the module 28.
- the quantification tables differ according to the degree of voicing MV provided to the quantization module 30 by the open-loop analyzer 36.
- a set of quantization tables for a degree of voicing MV is determined, during prior tests, so as to be statistically representative of frames having this degree MV. These sets are stored both in the coders and in the decoders implementing the invention.
- the module 30 delivers the set of quantized parameters LSP Q as well as its index Q in the applicable quantification tables.
- the speech coder 16 further comprises a module 42 for calculating the impulse response of the compound filter short-term summary filter and weighting filter perceptual.
- This compound filter has the function of transfer W (z) / A (z).
- module 42 takes for the weighting filter perceptual W (z) that corresponding to the LSP parameters interpolated but not quantified, i.e. the one whose coefficients were calculated by module 32, and for the synthesis filter 1 / A (z) the one corresponding to the parameters Quantified and interpolated LSP, i.e. the one that will actually reconstructed by the decoder.
- the TP delay index is ZP + DP.
- closed-loop LTP analysis consists in determining, in the search interval for long-term prediction delays T, the delay TP which maximizes, for each sub-frame of a voiced frame, the normalized correlation : where x (i) denotes the weighted speech signal SW of the subframe from which the memory of the weighted synthesis filter has been subtracted (i.e. the response to a zero signal, due to its initial states, of the filter whose impulse response has been calculated by module 42), and y T (i) denotes the convolution product: u (jT) designating the predictable component of the delayed excitation sequence of T samples, estimated by the well-known technique of the adaptive codebook.
- the missing values of u (jT) can be extrapolated from the previous values.
- Fractional delays are taken into account by oversampling the signal u (jT) in the adaptive repertoire.
- An oversampling of a factor m is obtained by means of polyphase interpolating filters.
- the gain g P of long-term prediction could be determined by the module 38 for each sub-frame, by applying the known formula: However, in a preferred version of the invention, the gain g P is calculated by the stochastic analysis module 40.
- the stochastic excitation determined for each subframe by the module 40 is of the multi-pulse type.
- the positions and gains calculated by the analysis module 40 stochastics are quantified by a module 44.
- a module 48 is thus provided in the encoder which receives the different parameters and which adds to some of them redundancy bits to detect and / or correct any transmission errors.
- redundancy bits are added to this parameter by module 48.
- bit rate per 20 ms frame is for example that indicated in table I.
- the channel coder 22 is that used in the pan-European system of radiocommunication with mobiles (GSM).
- GSM pan-European system of radiocommunication with mobiles
- This channel coder described in detail in Recommendation GSM 05.03, was developed for a 13 kbit / s speech coder of RPE-LTP type which also produces 260 bits per 20 ms frame. The sensitivity of each of the 260 bits was determined from listening tests.
- the bits from the source encoder have been grouped into three categories. The first of these categories IA groups 50 bits which are coded convolutionally on the basis of a generator polynomial giving a half redundancy with a constraint length equal to 5. Three parity bits are calculated and added to the 50 bits of the category IA before convolutional coding.
- the second category (IB) has 132 bits which are protected at a rate of a half by the same polynomial as the previous category.
- the third category (II) contains 78 unprotected bits. After application of the convolutional code, the bits (456 per frame) are subjected to interleaving.
- a mobile radio station capable of receiving the speech signal processed by the source encoder 16 is shown schematically in Figure 2.
- the radio signal received is first processed by a demodulator 50 then by a channel 52 decoder which performs dual operations of those of modulator 24 and channel encoder 22.
- the decoder channel 52 provides the speech decoder 54 with a sequence binary which, in the absence of transmission errors or when any errors have been corrected by the decoder channel 52, corresponds to the binary sequence that delivered the scheduling module 46 at the encoder 16.
- the decoder 54 comprises a module 56 which receives this binary sequence and which identifies the parameters relating to different frames and subframes.
- the module 56 performs in in addition to some checks on the parameters received. In particular, module 56 examines the redundancy bits introduced by the encoder module 48, to detect and / or correct errors affecting the parameters associated with these redundancy bits.
- a module 58 of the decoder receives the degree of voicing MV and the index of Q for quantizing the LSP parameters.
- the module 58 finds the quantized LSP parameters in the tables corresponding to the value of MV, and, after interpolation, converts them into coefficients a i for the short-term synthesis filter 60.
- a pulse generator 62 receives the positions p (n) of the np pulses of the stochastic excitation.
- the generator 62 delivers pulses of unit amplitude which are each multiplied by 64 by the associated gain g (n).
- the output of amplifier 64 is addressed to the long-term synthesis filter 66.
- This filter 66 has an adaptive directory structure.
- the output samples u of the filter 66 are stored in the adaptive directory 68 so as to be available for the subsequent subframes.
- the delay TP relative to a sub-frame, calculated from the quantization indices ZP and DP, is supplied to the adaptive repertoire 68 to produce the signal u suitably delayed.
- the amplifier 70 multiplies the signal thus delayed by the gain g P of long-term prediction.
- the long-term filter 66 finally comprises an adder 72 which adds the outputs of amplifiers 64 and 70 to provide the excitation sequence u.
- the excitation sequence is addressed to the short-term synthesis filter 60, and the resulting signal can also, in known manner, be subjected to a post-filter 74 whose coefficients depend on the synthesis parameters received, to form the signal of synthetic speech S '.
- the output signal S 'of the decoder 54 is then converted into analog by the converter 76 before being amplified to control a loudspeaker 78.
- the module 36 also determines, for each sub-frame st, the entire delay K st which maximizes the open loop estimation P st (k) at the long-term prediction gain on the sub-frame st, excluding the delays k for which the autocorrelation C st (k) is negative or smaller than a small fraction ⁇ of the energy R0 st of the subframe.
- step 94 the degree of voicing MV of the current frame is taken equal to 0 in step 94, which in this case ends the operations performed by the module 36 on this frame. If on the contrary the threshold S0 is exceeded in step 92, the current frame is detected as voiced and the degree MV will be equal to 1, 2 or 3. The module 36 then calculates, for each subframe st, a list I st containing candidate delays to constitute the ZP center of the search interval for long-term prediction delays.
- the module 36 determines the basic delay rbf in full resolution for the rest of the processing. This basic delay could be taken equal to the integer K st obtained in step 90. The fact of finding the basic delay in fractional resolution around K st however makes it possible to gain in precision.
- Step 100 thus consists in finding, around the integer delay K st obtained in step 90, the fractional delay which maximizes the expression C st 2 / G st .
- This search can be carried out at the maximum resolution of the fractional delays (1/6 in the example described here) even if the entire delay K st is not in the domain where this maximum resolution applies.
- the autocorrelations C st (T) and the delayed energies G st (T) are obtained by interpolation from the values stored in step 90 for the whole delays.
- the basic delay relating to a sub-frame could also be determined in fractional resolution from step 90 and taken into account in the first estimation of the overall prediction gain on the frame.
- step 102 the address j in the list I st and the index m of the submultiple are initialized to 0 and 1, respectively.
- a comparison 104 is made between the submultiple rbf / m and the minimum delay rmin. The submultiple rbf / m is to be examined if it is greater than rmin.
- step 110 If P st (r i ) ⁇ SE st , the delay r i is not taken into account, and we go directly to step 110 of incrementing the index m before carrying out the comparison 104 again for the next submultiple. If test 108 shows that P st (r i ) ⁇ SE st , the delay r i is retained and step 112 is executed before incrementing the index m in step 110. In step 112, we stores the index i at the address j in the list I st , we give the value m to the integer m0 intended to be equal to the index of the smallest submultiple retained, then we increment by one unit l 'address j.
- the examination of the sub-multiples of the basic delay is finished when the comparison 104 shows rbf / m ⁇ rmin.
- a comparison 116 is made between the multiple n.rbf / m0 and the maximum delay rmax. If n.rbf / m0> rmax, test 118 is carried out to determine whether the index m0 of the smallest sub-multiple is an integer multiple of n.
- step 120 the delay n.rbf / m0 has already been examined when examining the sub-multiples of rbf, and we go directly to step 120 of incrementing the index n before carrying out again comparison 116 for the next multiple. If test 118 shows that m0 is not an integer multiple of n, the multiple n.rbf / m0 is to be examined. We then take for the integer i the value of the index of the quantized delay r i closest to n.rbf / m0 (step 122), then we compare, at 124, the estimated value of the prediction gain P st ( r i ) at the selection threshold SE st .
- step 120 If P st (r i ) ⁇ SE st , the delay r i is not taken into account, and we go directly to step 120 of incrementing the index n. If test 124 shows that P st (r i ) ⁇ SE st , the delay r i is retained and step 126 is executed before incrementing the index n in step 120. In step 126, we stores the index i at address j in the list I st , then the address j is incremented by one.
- the list I st contains j candidate delay index. If we wish to limit the maximum length of the list I st to jmax for the following steps, we can take the length j st of this list equal to min (j, jmax) (step 128) and then, in step 130, order the list I st in the order of gains C st 2 (r Ist (j) ) / G st 2 (r Ist (j) ) decreasing for 0 ⁇ j ⁇ j st so as to keep only the j st delays providing the largest gain values.
- the value of jmax is chosen according to the compromise sought between the efficiency of the search for LTP delays and the complexity of this search. Typical values of jmax range from 3 to 5.
- the analysis module 36 calculates a quantity Ymax determining a second open-loop estimate of the prediction gain at long term over the entire frame, as well as indexes ZP, ZP0 and ZP1 in a phase 132, the progress of which is detailed in FIG. 6.
- This phase 132 consists in testing search intervals of length N1 to determine which one maximizes a second estimate of the overall prediction gain on the frame. The intervals tested are those whose centers are the candidate delays contained in the list I st calculated during phase 101.
- Phase 132 begins with a step 136 where the address j in the list I st is initialized to 0.
- step 138 we check if the index I st (j) has already been encountered by testing a previous interval centered on I st' (j ') with st' ⁇ st and 0 ⁇ j ' ⁇ j st' , in order to d '' Avoid testing the same interval twice. If test 138 reveals that I st (j) already appeared in a list I st , with st ' ⁇ st, we directly increment the address j in step 140, then we compare it to the length j st of the list I st . If the comparison 142 shows that j ⁇ j st , we return to step 138 for the new value of the address j.
- the quantity Y determining the second estimate of the overall prediction gain for the interval centered on I st (j) is calculated according to: then compared to Ymax, where Ymax represents the value to be maximized.
- This value Ymax is for example initialized to 0 at the same time as the index st in step 96. If Y ⁇ Ymax, we go directly to step 140 for incrementing the index j. If the comparison 150 shows that Y> Ymax, step 152 is executed before incrementing the address j in step 140. At this step 152, the index ZP is taken equal to I st (j) and the indices ZP0 and ZP1 are respectively taken equal to the smallest and the largest of the indices i st ' determined in step 148.
- the index st is incremented by one (step 154) then compared, in step 156, to the number nst of subframes per frame. If st ⁇ nst, we return to step 98 to perform the operations relating to the following sub-frame.
- the index ZP denotes the center of the search interval that will be provided to the module 38 closed loop LTP analysis
- ZP0 and ZP1 are index whose difference is representative of the dispersion of optimal delays per subframe in the interval centered on ZP.
- Gp 20.log 10 (RO / RO-Y max ).
- Two other thresholds S1 and S2 are used. If Gp ⁇ S1, the degree of voicing MV is taken equal to 1 for the current frame.
- ZP + DP index of TP delay ultimately determined may therefore in some cases be more small than 0 or larger than 255.
- This allows analysis LTP in closed loop to also carry on some delays TP smaller than rmin or larger than rmax.
- Reducing the delay search interval for very closely spaced frames reduces the complexity of the closed loop LTP analysis performed by the module 38 by reducing the number of convolutions y T (i) to be calculated according to formula (1).
- Another possibility is to provide a parity bit for the delay TP and / or the gain g P , making it possible to detect possible errors affecting these parameters.
- the first optimizations carried out in step 90 relative to the different subframes are replaced by a single optimization relating to the entire frame.
- the autocorrelations C (k) and the delayed energies G (k) for the entire frame are also calculated:
- nz basic delays K 1 ', ..., K nz ' in full resolution.
- the voiced / unvoiced decision (step 92) is taken on the basis of that of the basic delays K i 'which provides the greatest value for the first open-loop estimate at the long-term prediction gain.
- the basic delays in fractional resolution are determined by the same process as in step 100, but only allowing the quantized delay values. Examination 101 of the sub-multiples and multiples is not performed. For the phase 132 of calculating the second estimate of the prediction gain, the nz basic delays previously determined are taken as candidate delays. This second variant makes it possible to dispense with the systematic examination of the submultiples and of the multiples which are generally taken into account by virtue of the subdivision of the domain of possible delays.
- phase 132 is modified in that, in the optimization steps 148, the index i st ' which maximizes C st' 2 (r i ) / G st ' (r i ) for I st (j) -N1 / 2 ⁇ i ⁇ I st (j) + N1 / 2 and 0 ⁇ i ⁇ N, and on the other hand, during the same maximization loop, the index k st ' which maximizes this same quantity over a reduced interval I st (j) -N3 / 2 ⁇ i ⁇ I st (j) + N3 / 2 and 0 ⁇ i ⁇ N.
- Step 152 is also modified: the indexes ZP0 and ZP1 are no longer stored, but a quantity Ymax 'defined in the same way as Ymax but with reference to the reduced length interval:
- Gp' 20.log 10 [R0 / (R0-Ymax ')].
- the sub-frames for which the prediction gain is negative or negligible can be identified by consulting the nst pointers. If necessary, the module 38 is deactivated for the corresponding sub-frames. This does not affect the quality of the LTP analysis since the prediction gain corresponding to these subframes will be almost zero anyway.
- Another aspect of the invention relates to the module 42 for calculating the impulse response of the weighted synthesis filter.
- the closed loop LTP analysis module 38 needs this impulse response h over the duration of a subframe to calculate the convolutions y T (i) according to formula (1).
- the stochastic analysis module 40 also needs it to calculate convolutions as will be seen below.
- the operations performed by the module 42 are for example in accordance with the flowchart of FIG. 7.
- the truncated energies of the impulse response are also calculated:
- the coefficients a k are those involved in the perceptual weighting filter, i.e. the linear prediction coefficients interpolated but not quantified
- the coefficients a k are those applied to the synthesis filter, i.e. the quantized and interpolated linear prediction coefficients.
- the module 42 determines the shortest length L ⁇ such that the energy Eh (L ⁇ -1) of the impulse response truncated at L ⁇ samples is at least equal to a proportion ⁇ of its total energy Eh (pst-1) estimated over pst samples.
- a typical value of ⁇ is 98%.
- the number L ⁇ is initialized to pst in step 162 and decremented by unit as 166 as Eh (L ⁇ -2)> ⁇ .Eh (pst-1) (test 164).
- the length L ⁇ sought is obtained when test 164 shows that Eh (L ⁇ -2) ⁇ .Eh (pst-1).
- a term corrector ⁇ (MV) is added to the value of L ⁇ which has been obtained (step 168).
- This corrector term is preferably an increasing function of the degree of voicing.
- ⁇ (0) - 5
- ⁇ (3) + 7.
- the truncation length Lh of the response impulse is taken equal to L ⁇ if L ⁇ nst and to nst if not.
- a third aspect of the invention relates to the module 40 of stochastic analysis used to model the unpredictable part of the excitement.
- the stochastic excitation considered here is of the multi-pulse type.
- the stochastic excitation relating to a subframe is represented by np pulses of positions p (n) and of amplitudes, or gains, g (n) (1 ⁇ n ⁇ np).
- the gain g P of long-term prediction can also be calculated during the same process.
- the excitation sequence relating to a sub-frame comprises nc contributions associated respectively with nc gains.
- the contributions are lst sample vectors which, weighted by the associated and summed gains correspond to the excitation sequence of the short-term synthesis filter.
- np vectors comprising only 0 except an impulse of amplitude 1.
- the vectors F p (n) are simply constituted by the vector of the impulse response h shifted by p (n) samples. Truncating the impulse response as described above therefore makes it possible to significantly reduce the number of operations useful for calculating the scalar products involving these vectors F p (n) .
- the gains g nc-1 (i) are the selected gains and the minimized quadratic error E is equal to the energy at the target vector e nc-1 .
- the decomposition of Cholesky and the inversion of the matrix M n however require to carry out divisions and calculations of square roots which are operations demanding in terms of computation complexity.
- Different constraints can be brought to the domain of maximization of the quantity above included in the interval [0, lst [.
- the maximization is carried out in step 182 on the set of possible positions excluding the segments in which the positions p (1), ..., p (n have been found respectively) -1) pulses during previous iterations.
- the module 40 proceeds to the calculation 184 of the line n of the matrices L, R and K involved in the decomposition of the matrix B, which makes it possible to complete the matrices L n , R n and K n defined above.
- the column index j is first initialized at 0, in step 186.
- the variable tmp is first initialized at the value of component B (n, j), that is:
- step 188 the integer k is also initialized to 0.
- a comparison 190 is then made between the integers k and j. If k ⁇ j, we add the term L (n, k). R (j, k) to the variable tmp, then we increment the whole k by one unit (step 192) before re-performing the comparison 190.
- step 196 If j ⁇ n, the component R (n, j) is taken equal to tmp and the component L (n, j) to tmp.K (j) in step 196, then the column index j is incremented d 'a unit before returning to step 188 to calculate the following components.
- K (n) is taken equal to 1 / tmp if tmp ⁇ 0 (step 198) and to 0 otherwise.
- the calculation 184 requires at most one division 198, to obtain K (n).
- any singularity of the matrix B n does not cause instabilities since we avoid divisions by 0.
- the inversion 200 then begins with an initialization 202 of the column index j 'at n-1.
- the term Linv (j ') is initialized to -L (n, j') and the integer k 'to j' + 1.
- a comparison 206 is then carried out between the integers k ′ and n.
- the inversion 200 is followed by the calculation 214 of the reoptimized gains and of the target vector E for the following iteration.
- the computation of the reoptimized gains is also very simplified by the decomposition retained for the matrix B.
- One can indeed compute the vector g n (g n (0), ..., g n (n)) solution of g n .
- B n b n according to: and
- g n (i ') g n-1 (i') + L -1 (n, i ').
- the calculation 214 is detailed in FIG. 11.
- b (n) serves as the initialization value for the variable tmq.
- index i is also initialized to 0.
- the comparison 218 is then carried out between the integers i and n. If i ⁇ n, we add the term b (i). Linv (i) to the variable tmq and we increment i by one unit (step 220) before returning to the comparison 218.
- Step 226 also includes the incrementation of the index i 'before returning to the comparison 224.
- Segmental pulse search significantly decreases the number of pulse positions to be evaluated during steps 182 of the search for stochastic excitation. It also allows efficient quantification of the positions found.
- ns> np also has the advantage that good robustness to transmission errors can be obtained with regard to the positions of the pulses, by virtue of a separate quantification of the sequence numbers of the occupied segments and of the relative positions pulses in each occupied segment.
- the possible binary words are stored in a quantification table in which the reading addresses are the quantization indexes received.
- the order in this table, determined once for all, can be optimized so that an error of transmission affecting a bit of the index (the error case the more frequent, especially when interlacing is used work in the channel encoder 22) has, on average, minimal consequences according to a neighborhood criterion.
- the neighborhood criterion is for example that a word of ns bits does not can be replaced only by words "neighbors", distant a Hamming distance at most equal to an np-2 ⁇ threshold, so as to keep all the pulses except ⁇ of them at valid positions in case of transmission error the single-bit index.
- Other criteria would be usable in substitution or in addition, for example that two words are considered neighbors if the replacement of one by the other does not change the order of assignment of gains associated with pulses.
- the order in the table word quantification can be determined from arithmetic considerations or, if this is insufficient, in simulating error scenarios on a computer (so exhaustive or by statistical sampling of the type Monte-Carlo according to the number of possible error cases).
- the module scheduling 46 can put in the category of minimum protection, or in the unprotected category, a nx number of index bits which, if they are affected by a transmission error, give rise to a wrong word but checking the neighborhood criterion with a probability considered satisfactory, and put in a category more protected the other bits of the index. This way of proceed uses another word order in the quantification table.
- This scheduling can also be optimized using simulations if you want maximize the number nx of the index bits assigned to the least protected category.
- One possibility is to start by constituting a list of words of ns bits by counting in Gray code from 0 to 2 ns -1, and to obtain the ordered quantification table by deleting from this list the words having no weight of Hamming of np.
- the table thus obtained is such that two consecutive words have a Hamming distance of np-2. If the indexes in this table have a binary representation in Gray code, any error on the least significant bit causes the index to vary by ⁇ 1 and therefore causes the replacement of the actual occupancy word by a neighboring word in the sense of the np-2 threshold on the Hamming distance, and an error on the i-th least significant bit also varies the index by ⁇ 1 with a probability of approximately 2 1-i .
- nx By placing the nx least significant bits of the index in Gray code in an unprotected category, a possible transmission error affecting one of these bits leads to the replacement of the busy word by a neighboring word with a probability at least equal. to (1 + 1/2 + ... + 1/2 nx-1 ) / nx. This minimum probability decreases from 1 to (2 / nb) (1-1 / 2 nb ) for nx increasing from 1 to nb.
- the errors affecting the nb-nx most significant bits of the index will most often be corrected thanks to the protection applied to them by the channel coder.
- the value of nx is in this case chosen according to a compromise between robustness to small value errors) and a reduced bulk of the protected categories (large values).
- the possible binary words to represent the occupation of the segments are arranged in ascending order in a search table.
- An indexing table associates with each address the serial number, in the quantification table stored at the decoder, of the binary word having this address in the search table.
- the content of the search table and of the indexing table is given in table III (in decimal values).
- the quantification of the occupancy word of the segments deduced from the np positions provided by the analysis module stochastic 40 is performed in two stages by the module 44.
- a dichotomous search is first performed in the lookup table to determine the address in this table of the word to be quantified.
- the index of quantification is then obtained at the address determined in the indexing table then supplied to the scheduling module 46 bits.
- the module 44 also performs the quantification of the gains calculated by the module 40.
- the quantization bits of Gs are placed in a category protected by the channel 22 encoder, as well as most significant bits of the gain quantification indexes relative.
- the relative gain quantization bits are ordered to allow assignment to impulses associated belonging to the segments localized by the word of occupation. Segmental research according to the invention also allows effective protection of positions relative pulses associated with the largest values gain.
- the decoder 54 To reconstruct impulse contributions of excitation, the decoder 54 first locates the segments by means of the occupation word received; he then assigns the associated earnings; then he assigns the positions relative to pulses based on the order of importance of the gains.
- the 13 kbit / s speech coder requires order 15 million comma instructions per second (Mips) fixed. So we will typically do this and program a commercial digital signal processor (DSP) as well as the decoder which requires only about 5 Mips.
- DSP digital signal processor
Landscapes
- Engineering & Computer Science (AREA)
- Computational Linguistics (AREA)
- Signal Processing (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Human Computer Interaction (AREA)
- Physics & Mathematics (AREA)
- Acoustics & Sound (AREA)
- Multimedia (AREA)
- Compression, Expansion, Code Conversion, And Decoders (AREA)
- Investigating Or Analysing Materials By The Use Of Chemical Reactions (AREA)
Description
- la figure 1 est un schéma synoptique d'une station de radiocommunication incorporant un codeur de parole mettant en oeuvre l'invention ;
- la figure 2 est un schéma synoptique d'une station de radiocommunication apte à recevoir un signal produit par celle de la figure 1 ;
- les figures 3 à 6 sont des organigrammes Illustrant un processus d'analyse LTP en boucle ouverte appliqué dans le codeur de parole de la figure 1 ;
- la figure 7 est un organigramme illustrant un processus de détermination de la réponse impulsionnelle du filtre de synthèse pondéré appliqué dans le codeur de parole de la figure 1 ;
- les figures 8 à 11 sont des organigrammes illustrant un processus de recherche de l'excitation stochastique appliqué dans le codeur de parole de la figure 1.
- l'index Q des paramètres LSP quantifiés pour chaque trame ;
- le degré MV de voisement de chaque trame ;
- l'index ZP du centre de l'intervalle de recherche des retards LTP pour chaque trame voisée ;
- l'index différentiel DP du retard LTP pour chaque sous-trame d'une trame voisée, et le gain associé gP ;
- les positions p(n) et les gains g(n) des impulsions de l'excitation stochastique pour chaque sous-trame.
| paramètres quantifiés | MV=0 | MV=1 ou 2 | MV=3 |
| LSP | 34 | 34 | 34 |
| MV + redondance | 6 | 6 | 6 |
| ZP | - | 8 | 8 |
| DP | - | 20 | 16 |
| gTP | - | 20 | 24 |
| positions impulsions | 80 | 72 | 72 |
| gains impulsions | 140 | 100 | 100 |
| Total | 260 | 260 | 260 |
- X désigne un vecteur-cible initial composé des lst échantillons du signal de parole pondéré SW sans mémoire : X=(x(0),x(1),...,x(lst-1)), les x(i) ayant été calculés comme indiqué précédemment lors de l'analyse LTP en boucle fermée ;
- g désigne le vecteur ligne composé des np+1 gains : g=(g(0)=gP, g(1),...,g(np)) ;
- les vecteurs-ligne Fp(n) (0≤n≤nc) sont des contributions pondérées ayant pour composantes i(0≤i<lst) les produits de convolution entre la contribution n à la séquence d'excitation et la réponse impulsionnelle h du filtre de synthèse pondéré ;
- b désigne le vecteur ligne composé des nc produits scalaires entre le vecteur X et les vecteurs ligne Fp(n) ;
- B désigne une matrice symétrique à nc lignes et nc colonnes dont le terme Bi,j=Fp(i).Fp(j) T (0≤i,j<nc) est égal au produit scalaire entre les vecteurs Fp(i) et Fp(j) précédemment définis ;
- (.)T désigne la transposition matricielle.
| index de quantification | mot d'occupation des segments | ||
| décimal | binaire naturel | binaire naturel | décimal |
| 0 | 000 | 0011 | 3 |
| 1 | 001 | 0101 | 5 |
| 2 | 010 | 1001 | 9 |
| 3 | 011 | 1100 | 12 |
| 4 | 100 | 1010 | 10 |
| 5 | 101 | 0110 | 6 |
| (6) | (110) | (1001 ou 1010) | (9 ou 10) |
| (7) | (111) | (1100 ou 0110) | (12 ou 6) |
| Adresse | Table de recherche | Table d'indexage |
| 0 | 3 | 0 |
| 1 | 5 | 1 |
| 2 | 6 | 5 |
| 3 | 9 | 2 |
| 4 | 10 | 4 |
| 5 | 12 | 3 |
Claims (10)
- Procédé de codage à analyse par synthèse d'un signal de parole (S) numérisé en trames successives divisées en nst sous-trames, comprenant les étapes suivantes :caractérisé en ce qu'à l'étape d'analyse en boucle ouverte, on détermine l'intervalle de recherche relatif à chaque trame voisée de façon qu'il contienne un nombre de retards (N1, N3) dépendant du degré de voisement de ladite trame.analyse par prédiction linéaire du signal de parole pour déterminer des paramètres d'un filtre (60) de synthèse à court terme ;analyse en boucle ouverte du signal de parole pour détecter les trames voisées du signal et pour déterminer, pour chaque trame voisée, un degré de voisement du signal (MV) et un intervalle de recherche d'un retard de prédiction à long terme ;analyse prédictive en boucle fermée du signal de parole pour sélectionner, pour certaines au moins des sous-trames des trames voisées, un retard de prédiction à long terme contenu dans l'intervalle de recherche et constituant un paramètre d'un filtre de synthèse à long terme (66); etdétermination d'une excitation stochastique pour chaque sous-trame, de façon à minimiser un écart pondéré perceptuellement entre le signal de parole et l'excitation stochastique filtrée par les filtres de synthèse à long terme et à court terme,
- Procédé selon la revendication 1, caractérisé en ce que l'intervalle de recherche du retard de prédiction à long terme contient moins de retards pour les trames ayant le plus grand degré de voisement que pour les autres trames voisées.
- Procédé selon la revendication 1 ou 2, caractérisé en ce que l'analyse en boucle ouverte relative à une trame comprend la détermination de nst retards de base (Kst) qui maximisent chacun une estimation en boucle ouverte du gain de prédiction à long terme sur une sous-trame respective de ladite trame, puis la comparaison entre un premier seuil prédéterminé (S0) et une première estimation en boucle ouverte du gain de prédiction à long terme sur la trame obtenue sur la base des nst retards de base relativement aux sous-trames correspondantes pour détecter si la trame est voisée, en ce que, si la trame est détectée comme voisée, l'analyse en boucle ouverte comprend en outre pour chaque sous-trame la détermination d'une liste (Ist) de retards candidats pour lesquels l'estimation en boucle ouverte au gain de prédiction sur la sous-trame est supérieure à une fraction déterminée (β) de l'estimation relative au retard de base pour la sous-trame, en ce qu'on sélectionne dans lesdites listes le retard candidat pour lequel une seconde estimation en boucle ouverte du gain de prédiction à long terme sur la trame est maximale, la seconde estimation en boucle ouverte sur la trame associée à un retard candidat étant obtenue sur la base de nst retards optimaux, compris dans un intervalle de N1 retards centré sur ledit retard candidat, qui, respectivement, maximisent sur ledit intervalle l'estimation en boucle ouverte du gain de prédiction sur les nst sous-trames, en ce que la détermination au degré de voisement de la trame comprend une comparaison entre la seconde estimation maximisée du gain de prédiction sur la trame et au moins un autre seuil prédéterminé (S1,S2), et en ce que l'intervalle de recherche déterminé à l'issue de l'analyse en boucle ouverte est centré sur ledit retard sélectionné.
- Procédé selon la revendication 1 ou 2, caractérisé en ce que l'analyse en boucle ouverte relative à une trame comprend la détermination d'un retard de base (K) qui maximise une première estimation en boucle ouverte du gain de prédiction à long terme sur ladite trame, puis la comparaison entre un premier seuil prédéterminé (S0) et la première estimation maximisée du gain de prédiction à long terme sur la trame pour détecter si la trame est voisée, en ce que, si la trame est détectée comme voisée, l'analyse en boucle ouverte comprend en outre la détermination d'une liste (I) de retards candidats pour lesquels l'estimation en boucle ouverte du gain de prédiction sur la trame est supérieure à une fraction déterminée (β) de l'estimation relative au retard de base, en ce qu'on sélectionne dans ladite liste le retard candidat pour lequel une seconde estimation en boucle ouverte du gain de prédiction à long terme sur la trame est maximale, la seconde estimation en boucle ouverte sur la trame associée à un retard candidat étant obtenue sur la base de nst retards optimaux, compris dans un intervalle de N1 retards centré sur ledit retard candidat, qui, respectivement, maximisent sur ledit intervalle l'estimation en boucle ouverte du gain de prédiction sur les nst sous-trames, en ce que la détermination du degré de voisement de la trame comprend une comparaison entre la seconde estimation maximisée du gain de prédiction sur la trame et au moins un autre seuil prédéterminé (S1,S2), et en ce que l'intervalle de recherche déterminé à l'issue de l'analyse en boucle ouverte est centré sur ledit retard sélectionné.
- Procédé selon la revendication 1 ou 2, caractérisé en ce que l'analyse en boucle ouverte relative à une trame comprend la détermination d'un nombre nz de retards de base (K1' ,..., Knz') qui maximisent chacun, sur un sous-intervalle respectif de valeurs de retard possibles, une première estimation en boucle ouverte du gain de prédiction à long terme sur ladite trame, puis la comparaison entre un premier seuil prédéterminé (S0) et la plus grande des nz premières estimations maximisées du gain de prédiction à long terme sur la trame pour détecter si la trame est voisée, en ce que, si la trame est détectée comme voisée, on sélectionne, parmi nz retards candidats obtenus à partir des nz retards de base, le retard candidat pour lequel une seconde estimation en boucle ouverte du gain de prédiction à long terme sur la trame est maximale, la seconde estimation en boucle ouverte sur la trame associée à un retard candidat étant obtenue sur la base de nst retards optimaux, compris dans un intervalle de N1 retards centré sur ledit retard candidat, qui, respectivement, maximisent sur ledit intervalle l'estimation en boucle ouverte au gain de prédiction sur les nst sous-trames, en ce que la détermination du degré de voisement de la trame comprend une comparaison entre la seconde estimation maximisée du gain de prédiction sur la trame et au moins un autre seuil prédéterminé (S1,S2), et en ce que l'intervalle de recherche déterminé à l'issue de l'analyse en boucle ouverte est centré sur ledit retard sélectionné.
- Procédé selon l'une quelconque des revendications 3 à 5, caractérisé en ce que si la seconde estimation maximisée du gain de prédiction sur une trame voisée est supérieure à un des seuils (S2), on détermine si les nst retards optimaux sont compris dans un intervalle centré sur le retard sélectionné et contenant un nombre de retards N3 inférieur à N1 et, dans l'affirmative, on attribue à la trame un degré de voisement pour lequel l'intervalle de recherche du retard de prédiction à long terme contient N3 retards, l'intervalle de recherche contenant N1 retards pour au moins un autre degré de voisement.
- Procédé selon l'une quelconque des revendications 3 à 5, caractérisé en ce que lors de la maximisation de la seconde estimation en boucle ouverte du gain de prédiction à long terme sur une trame voisée, on calcule en outre une troisième estimation en boucle ouverte du gain sur la trame sur la base de nst retards, compris dans un intervalle centré sur le retard sélectionné et contenant un nombre N3 de retards inférieur à N1, qui, respectivement, maximisent sur ledit intervalle de N3 retards l'estimation en boucle ouverte du gain de prédiction sur les nst sous-trames, et en ce qu'on attribue à la trame un degré de voisement pour lequel l'intervalle de recherche contient N3 retards si ladite troisième estimation dépasse un seuil prédéterminé (S2), l'intervalle de recherche contenant N1 retards pour au moins un autre degré de voisement.
- Procédé selon la revendication 3 ou 4, caractérisé en ce que les retards candidats d'une liste sont choisis parmi les sous-multiples du retard de base associé à ladite liste et parmi les multiples du plus petit desdits sous-multiples pour lequel l'estimation en boucle ouverte au gain de prédiction est supérieure à ladite fraction déterminée de l'estimation relative au retard de base.
- Procédé selon la revendication 8, caractérisé en ce que les retards de prédiction à long terme peuvent correspondre à des nombres entiers ou fractionnaires d'échantillons du signal de parole, en ce qu'on détermine les retards de base (rbf) en résolution fractionnaire pour chercher les sous-multiples et les multiples à inclure dans une liste des retards candidats, et en ce que les retards de base sont déterminés en résolution entière pour évaluer les premières estimations en boucle ouverte du gain de prédiction sur une trame.
- Procédé selon l'une quelconque des revendications 3 à 9, caractérisé en ce qu'on n'effectue pas l'analyse prédictive en boucle fermée relativement à chaque sous-trame pour laquelle l'autocorrélation (Cst) du signal de parole associée au retard optimal pour ladite sous-trame est négative.
Applications Claiming Priority (3)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| FR9500134 | 1995-01-06 | ||
| FR9500134A FR2729246A1 (fr) | 1995-01-06 | 1995-01-06 | Procede de codage de parole a analyse par synthese |
| PCT/FR1996/000004 WO1996021218A1 (fr) | 1995-01-06 | 1996-01-03 | Procede de codage de parole a analyse par synthese |
Publications (2)
| Publication Number | Publication Date |
|---|---|
| EP0801788A1 EP0801788A1 (fr) | 1997-10-22 |
| EP0801788B1 true EP0801788B1 (fr) | 1999-06-09 |
Family
ID=9474931
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP96901008A Expired - Lifetime EP0801788B1 (fr) | 1995-01-06 | 1996-01-03 | Procede de codage de parole a analyse par synthese |
Country Status (9)
| Country | Link |
|---|---|
| US (1) | US5974377A (fr) |
| EP (1) | EP0801788B1 (fr) |
| CN (1) | CN1145143C (fr) |
| AT (1) | ATE181170T1 (fr) |
| AU (1) | AU704229B2 (fr) |
| CA (1) | CA2209384C (fr) |
| DE (1) | DE69602822T2 (fr) |
| FR (1) | FR2729246A1 (fr) |
| WO (1) | WO1996021218A1 (fr) |
Families Citing this family (21)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| EP1553564A3 (fr) | 1996-08-02 | 2005-10-19 | Matsushita Electric Industrial Co., Ltd. | Codec vocal, support sur lequel est enregistré un programme codec vocal, et appareil mobile de télécommunications |
| JP3166697B2 (ja) * | 1998-01-14 | 2001-05-14 | 日本電気株式会社 | 音声符号化・復号装置及びシステム |
| US6192335B1 (en) * | 1998-09-01 | 2001-02-20 | Telefonaktieboiaget Lm Ericsson (Publ) | Adaptive combining of multi-mode coding for voiced speech and noise-like signals |
| FI116992B (fi) * | 1999-07-05 | 2006-04-28 | Nokia Corp | Menetelmät, järjestelmä ja laitteet audiosignaalin koodauksen ja siirron tehostamiseksi |
| US7272553B1 (en) * | 1999-09-08 | 2007-09-18 | 8X8, Inc. | Varying pulse amplitude multi-pulse analysis speech processor and method |
| JP3372908B2 (ja) * | 1999-09-17 | 2003-02-04 | エヌイーシーマイクロシステム株式会社 | マルチパルス探索処理方法と音声符号化装置 |
| KR100324204B1 (ko) * | 1999-12-24 | 2002-02-16 | 오길록 | 예측분할벡터양자화 및 예측분할행렬양자화 방식에 의한선스펙트럼쌍 양자화기의 고속탐색방법 |
| US6965640B2 (en) * | 2001-08-08 | 2005-11-15 | Octasic Inc. | Method and apparatus for generating a set of filter coefficients providing adaptive noise reduction |
| US6999509B2 (en) * | 2001-08-08 | 2006-02-14 | Octasic Inc. | Method and apparatus for generating a set of filter coefficients for a time updated adaptive filter |
| US6957240B2 (en) * | 2001-08-08 | 2005-10-18 | Octasic Inc. | Method and apparatus for providing an error characterization estimate of an impulse response derived using least squares |
| US6970896B2 (en) | 2001-08-08 | 2005-11-29 | Octasic Inc. | Method and apparatus for generating a set of filter coefficients |
| CA2365203A1 (fr) * | 2001-12-14 | 2003-06-14 | Voiceage Corporation | Methode de modification de signal pour le codage efficace de signaux de la parole |
| ES2291939T3 (es) * | 2003-09-29 | 2008-03-01 | Koninklijke Philips Electronics N.V. | Codificacion de señales de audio. |
| US7792670B2 (en) * | 2003-12-19 | 2010-09-07 | Motorola, Inc. | Method and apparatus for speech coding |
| US8329884B2 (en) | 2004-12-17 | 2012-12-11 | Roche Molecular Systems, Inc. | Reagents and methods for detecting Neisseria gonorrhoeae |
| CN101320565B (zh) * | 2007-06-08 | 2011-05-11 | 华为技术有限公司 | 感知加权滤波方法及感知加权滤波器 |
| US9626982B2 (en) * | 2011-02-15 | 2017-04-18 | Voiceage Corporation | Device and method for quantizing the gains of the adaptive and fixed contributions of the excitation in a CELP codec |
| FR2987931A1 (fr) * | 2012-03-12 | 2013-09-13 | France Telecom | Modification des caracteristiques spectrales d'un filtre de prediction lineaire d'un signal audionumerique represente par ses coefficients lsf ou isf. |
| EP3011561B1 (fr) | 2013-06-21 | 2017-05-03 | Fraunhofer-Gesellschaft zur Förderung der angewandten Forschung e.V. | Appareil et procédé pour l'affaiblissement graduel amélioré de signal dans différents domaines pendant un masquage d'erreur |
| CN107452390B (zh) * | 2014-04-29 | 2021-10-26 | 华为技术有限公司 | 音频编码方法及相关装置 |
| CN114036779B (zh) * | 2021-11-30 | 2025-04-01 | 南方电网科学研究院有限责任公司 | 一种电网多重时间区间同步仿真方法、装置、介质及设备 |
Family Cites Families (44)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| NL8302985A (nl) * | 1983-08-26 | 1985-03-18 | Philips Nv | Multipulse excitatie lineair predictieve spraakcodeerder. |
| CA1223365A (fr) * | 1984-02-02 | 1987-06-23 | Shigeru Ono | Methode et appareil de codage de paroles |
| NL8500843A (nl) * | 1985-03-22 | 1986-10-16 | Koninkl Philips Electronics Nv | Multipuls-excitatie lineair-predictieve spraakcoder. |
| US4868867A (en) * | 1987-04-06 | 1989-09-19 | Voicecraft Inc. | Vector excitation speech or audio coder for transmission or storage |
| US4831624A (en) * | 1987-06-04 | 1989-05-16 | Motorola, Inc. | Error detection method for sub-band coding |
| US4802171A (en) * | 1987-06-04 | 1989-01-31 | Motorola, Inc. | Method for error correction in digitally encoded speech |
| CA1337217C (fr) * | 1987-08-28 | 1995-10-03 | Daniel Kenneth Freeman | Codage vocal |
| US5359696A (en) * | 1988-06-28 | 1994-10-25 | Motorola Inc. | Digital speech coder having improved sub-sample resolution long-term predictor |
| WO1990013891A1 (fr) * | 1989-05-11 | 1990-11-15 | Telefonaktiebolaget Lm Ericsson | Procede de positionnement d'impulsions d'excitation dans un codeur de parole a prediction lineaire |
| US5060269A (en) * | 1989-05-18 | 1991-10-22 | General Electric Company | Hybrid switched multi-pulse/stochastic speech coding technique |
| US5097508A (en) * | 1989-08-31 | 1992-03-17 | Codex Corporation | Digital speech coder having improved long term lag parameter determination |
| SG47028A1 (en) * | 1989-09-01 | 1998-03-20 | Motorola Inc | Digital speech coder having improved sub-sample resolution long-term predictor |
| ATE177867T1 (de) * | 1989-10-17 | 1999-04-15 | Motorola Inc | Digitaler sprachdekodierer unter verwendung einer nachfilterung mit einer reduzierten spektralverzerrung |
| US5073940A (en) * | 1989-11-24 | 1991-12-17 | General Electric Company | Method for protecting multi-pulse coders from fading and random pattern bit errors |
| US5307441A (en) * | 1989-11-29 | 1994-04-26 | Comsat Corporation | Wear-toll quality 4.8 kbps speech codec |
| US5097507A (en) * | 1989-12-22 | 1992-03-17 | General Electric Company | Fading bit error protection for digital cellular multi-pulse speech coder |
| US5265219A (en) * | 1990-06-07 | 1993-11-23 | Motorola, Inc. | Speech encoder using a soft interpolation decision for spectral parameters |
| JPH04264597A (ja) * | 1991-02-20 | 1992-09-21 | Fujitsu Ltd | 音声符号化装置および音声復号装置 |
| FI98104C (fi) * | 1991-05-20 | 1997-04-10 | Nokia Mobile Phones Ltd | Menetelmä herätevektorin generoimiseksi ja digitaalinen puhekooderi |
| EP1998319B1 (fr) * | 1991-06-11 | 2010-08-11 | Qualcomm Incorporated | Vocodeur à débit variable |
| DK0556354T3 (da) * | 1991-09-05 | 2001-12-17 | Motorola Inc | Fejlbeskyttelse til multitilstandstalekodere |
| US5253269A (en) * | 1991-09-05 | 1993-10-12 | Motorola, Inc. | Delta-coded lag information for use in a speech coder |
| TW224191B (fr) * | 1992-01-28 | 1994-05-21 | Qualcomm Inc | |
| US5495555A (en) * | 1992-06-01 | 1996-02-27 | Hughes Aircraft Company | High quality low bit rate celp-based speech codec |
| US5317595A (en) * | 1992-06-30 | 1994-05-31 | Nokia Mobile Phones Ltd. | Rapidly adaptable channel equalizer |
| US5717824A (en) * | 1992-08-07 | 1998-02-10 | Pacific Communication Sciences, Inc. | Adaptive speech coder having code excited linear predictor with multiple codebook searches |
| FI95086C (fi) * | 1992-11-26 | 1995-12-11 | Nokia Mobile Phones Ltd | Menetelmä puhesignaalin tehokkaaksi koodaamiseksi |
| FR2702590B1 (fr) * | 1993-03-12 | 1995-04-28 | Dominique Massaloux | Dispositif de codage et de décodage numériques de la parole, procédé d'exploration d'un dictionnaire pseudo-logarithmique de délais LTP, et procédé d'analyse LTP. |
| IT1264766B1 (it) * | 1993-04-09 | 1996-10-04 | Sip | Codificatore della voce utilizzante tecniche di analisi con un'eccitazione a impulsi. |
| IT1270438B (it) * | 1993-06-10 | 1997-05-05 | Sip | Procedimento e dispositivo per la determinazione del periodo del tono fondamentale e la classificazione del segnale vocale in codificatori numerici della voce |
| US5784532A (en) * | 1994-02-16 | 1998-07-21 | Qualcomm Incorporated | Application specific integrated circuit (ASIC) for performing rapid speech compression in a mobile telephone system |
| US5751903A (en) * | 1994-12-19 | 1998-05-12 | Hughes Electronics | Low rate multi-mode CELP codec that encodes line SPECTRAL frequencies utilizing an offset |
| FR2729245B1 (fr) * | 1995-01-06 | 1997-04-11 | Lamblin Claude | Procede de codage de parole a prediction lineaire et excitation par codes algebriques |
| FR2734389B1 (fr) * | 1995-05-17 | 1997-07-18 | Proust Stephane | Procede d'adaptation du niveau de masquage du bruit dans un codeur de parole a analyse par synthese utilisant un filtre de ponderation perceptuelle a court terme |
| US5732389A (en) * | 1995-06-07 | 1998-03-24 | Lucent Technologies Inc. | Voiced/unvoiced classification of speech for excitation codebook selection in celp speech decoding during frame erasures |
| US5699485A (en) * | 1995-06-07 | 1997-12-16 | Lucent Technologies Inc. | Pitch delay modification during frame erasures |
| US5664055A (en) * | 1995-06-07 | 1997-09-02 | Lucent Technologies Inc. | CS-ACELP speech compression system with adaptive pitch prediction filter gain based on a measure of periodicity |
| US5710863A (en) * | 1995-09-19 | 1998-01-20 | Chen; Juin-Hwey | Speech signal quantization using human auditory models in predictive coding systems |
| US5790759A (en) * | 1995-09-19 | 1998-08-04 | Lucent Technologies Inc. | Perceptual noise masking measure based on synthesis filter frequency response |
| JP4005154B2 (ja) * | 1995-10-26 | 2007-11-07 | ソニー株式会社 | 音声復号化方法及び装置 |
| JP3680380B2 (ja) * | 1995-10-26 | 2005-08-10 | ソニー株式会社 | 音声符号化方法及び装置 |
| FR2742568B1 (fr) * | 1995-12-15 | 1998-02-13 | Catherine Quinquis | Procede d'analyse par prediction lineaire d'un signal audiofrequence, et procedes de codage et de decodage d'un signal audiofrequence en comportant application |
| US5729694A (en) * | 1996-02-06 | 1998-03-17 | The Regents Of The University Of California | Speech coding, reconstruction and recognition using acoustics and electromagnetic waves |
| US5708757A (en) * | 1996-04-22 | 1998-01-13 | France Telecom | Method of determining parameters of a pitch synthesis filter in a speech coder, and speech coder implementing such method |
-
1995
- 1995-01-06 FR FR9500134A patent/FR2729246A1/fr active Granted
-
1996
- 1996-01-03 EP EP96901008A patent/EP0801788B1/fr not_active Expired - Lifetime
- 1996-01-03 CA CA002209384A patent/CA2209384C/fr not_active Expired - Fee Related
- 1996-01-03 US US08/860,673 patent/US5974377A/en not_active Expired - Lifetime
- 1996-01-03 AT AT96901008T patent/ATE181170T1/de not_active IP Right Cessation
- 1996-01-03 AU AU44901/96A patent/AU704229B2/en not_active Ceased
- 1996-01-03 CN CNB961917946A patent/CN1145143C/zh not_active Expired - Fee Related
- 1996-01-03 WO PCT/FR1996/000004 patent/WO1996021218A1/fr not_active Ceased
- 1996-01-03 DE DE69602822T patent/DE69602822T2/de not_active Expired - Fee Related
Also Published As
| Publication number | Publication date |
|---|---|
| CN1173939A (zh) | 1998-02-18 |
| ATE181170T1 (de) | 1999-06-15 |
| DE69602822D1 (de) | 1999-07-15 |
| CN1145143C (zh) | 2004-04-07 |
| US5974377A (en) | 1999-10-26 |
| CA2209384C (fr) | 2001-05-29 |
| AU704229B2 (en) | 1999-04-15 |
| WO1996021218A1 (fr) | 1996-07-11 |
| FR2729246B1 (fr) | 1997-03-07 |
| DE69602822T2 (de) | 1999-12-23 |
| FR2729246A1 (fr) | 1996-07-12 |
| AU4490196A (en) | 1996-07-24 |
| CA2209384A1 (fr) | 1996-07-11 |
| EP0801788A1 (fr) | 1997-10-22 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| EP0801790B1 (fr) | Procede de codage de parole a analyse par synthese | |
| EP0801789B1 (fr) | Procede de codage de parole a analyse par synthese | |
| EP0801788A1 (fr) | Procede de codage de parole a analyse par synthese | |
| EP1994531B1 (fr) | Codage ou decodage perfectionnes d'un signal audionumerique, en technique celp | |
| EP0749626B1 (fr) | Procede de codage de parole a prediction lineaire et excitation par codes algebriques | |
| EP1692689B1 (fr) | Procede de codage multiple optimise | |
| CN1124589C (zh) | 码激励线性预测(celp)编码器中搜索激励代码簿的方法和装置 | |
| EP2080194B1 (fr) | Attenuation du survoisement, notamment pour la generation d'une excitation aupres d'un decodeur, en absence d'information | |
| EP0490740A1 (fr) | Procédé et dispositif pour l'évaluation de la périodicité et du voisement du signal de parole dans les vocodeurs à très bas débit. | |
| EP0616315A1 (fr) | Dispositif de codage et de décodage numérique de la parole, procédé d'exploration d'un dictionnaire pseudo-logarithmique de délais LTP, et procédé d'analyse LTP | |
| EP1192619B1 (fr) | Codage et decodage audio par interpolation | |
| WO2002029786A1 (fr) | Procede et dispositif de codage segmental d'un signal audio | |
| JP4007730B2 (ja) | 音声符号化装置、音声符号化方法および音声符号化アルゴリズムを記録したコンピュータ読み取り可能な記録媒体 | |
| Jung et al. | Efficient implementation of ITU-t g. 723.1 speech coder for multichannel voice transmission and storage. | |
| EP1194923B1 (fr) | Procedes et dispositifs d'analyse et de synthese audio | |
| EP1192618B1 (fr) | Codage audio avec liftrage adaptif | |
| EP1192621B1 (fr) | Codage audio avec composants harmoniques | |
| JP2000330594A (ja) | 音声符号化装置及び方法並びに音声符号化プログラムを記録した記憶媒体 | |
| FR2980620A1 (fr) | Traitement d'amelioration de la qualite des signaux audiofrequences decodes | |
| EP1192620A1 (fr) | Codage et decodage audio incluant des composantes non harmoniques du signal | |
| JPH0675597A (ja) | 音声符号化装置 | |
| WO2013135997A1 (fr) | Modification des caractéristiques spectrales d'un filtre de prédiction linéaire d'un signal audionumérique représenté par ses coefficients lsf ou isf |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| 17P | Request for examination filed |
Effective date: 19970725 |
|
| AK | Designated contracting states |
Kind code of ref document: A1 Designated state(s): AT BE CH DE GB IT LI LU NL SE |
|
| RAP1 | Party data changed (applicant data changed or rights of an application transferred) |
Owner name: MATRA NORTEL COMMUNICATIONS |
|
| GRAG | Despatch of communication of intention to grant |
Free format text: ORIGINAL CODE: EPIDOS AGRA |
|
| GRAG | Despatch of communication of intention to grant |
Free format text: ORIGINAL CODE: EPIDOS AGRA |
|
| GRAH | Despatch of communication of intention to grant a patent |
Free format text: ORIGINAL CODE: EPIDOS IGRA |
|
| 17Q | First examination report despatched |
Effective date: 19981029 |
|
| GRAH | Despatch of communication of intention to grant a patent |
Free format text: ORIGINAL CODE: EPIDOS IGRA |
|
| GRAA | (expected) grant |
Free format text: ORIGINAL CODE: 0009210 |
|
| AK | Designated contracting states |
Kind code of ref document: B1 Designated state(s): AT BE CH DE GB IT LI LU NL SE |
|
| REF | Corresponds to: |
Ref document number: 181170 Country of ref document: AT Date of ref document: 19990615 Kind code of ref document: T |
|
| REG | Reference to a national code |
Ref country code: CH Ref legal event code: EP |
|
| REG | Reference to a national code |
Ref country code: CH Ref legal event code: NV Representative=s name: KELLER & PARTNER PATENTANWAELTE AG |
|
| GBT | Gb: translation of ep patent filed (gb section 77(6)(a)/1977) |
Effective date: 19990622 |
|
| REF | Corresponds to: |
Ref document number: 69602822 Country of ref document: DE Date of ref document: 19990715 |
|
| PG25 | Lapsed in a contracting state [announced via postgrant information from national office to epo] |
Ref country code: LU Free format text: LAPSE BECAUSE OF NON-PAYMENT OF DUE FEES Effective date: 20000103 Ref country code: AT Free format text: LAPSE BECAUSE OF NON-PAYMENT OF DUE FEES Effective date: 20000103 |
|
| PG25 | Lapsed in a contracting state [announced via postgrant information from national office to epo] |
Ref country code: LI Free format text: LAPSE BECAUSE OF NON-PAYMENT OF DUE FEES Effective date: 20000131 Ref country code: CH Free format text: LAPSE BECAUSE OF NON-PAYMENT OF DUE FEES Effective date: 20000131 Ref country code: BE Free format text: LAPSE BECAUSE OF NON-PAYMENT OF DUE FEES Effective date: 20000131 |
|
| PLBE | No opposition filed within time limit |
Free format text: ORIGINAL CODE: 0009261 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: NO OPPOSITION FILED WITHIN TIME LIMIT |
|
| 26N | No opposition filed | ||
| BERE | Be: lapsed |
Owner name: MATRA NORTEL COMMUNICATIONS Effective date: 20000131 |
|
| PG25 | Lapsed in a contracting state [announced via postgrant information from national office to epo] |
Ref country code: NL Free format text: LAPSE BECAUSE OF NON-PAYMENT OF DUE FEES Effective date: 20000801 |
|
| REG | Reference to a national code |
Ref country code: CH Ref legal event code: PL |
|
| NLV4 | Nl: lapsed or anulled due to non-payment of the annual fee |
Effective date: 20000801 |
|
| PGFP | Annual fee paid to national office [announced via postgrant information from national office to epo] |
Ref country code: SE Payment date: 20001218 Year of fee payment: 6 |
|
| REG | Reference to a national code |
Ref country code: GB Ref legal event code: IF02 |
|
| PG25 | Lapsed in a contracting state [announced via postgrant information from national office to epo] |
Ref country code: SE Free format text: LAPSE BECAUSE OF NON-PAYMENT OF DUE FEES Effective date: 20020104 |
|
| EUG | Se: european patent has lapsed |
Ref document number: 96901008.1 |
|
| PGFP | Annual fee paid to national office [announced via postgrant information from national office to epo] |
Ref country code: DE Payment date: 20040130 Year of fee payment: 9 |
|
| PGFP | Annual fee paid to national office [announced via postgrant information from national office to epo] |
Ref country code: GB Payment date: 20041210 Year of fee payment: 10 |
|
| PG25 | Lapsed in a contracting state [announced via postgrant information from national office to epo] |
Ref country code: IT Free format text: LAPSE BECAUSE OF NON-PAYMENT OF DUE FEES;WARNING: LAPSES OF ITALIAN PATENTS WITH EFFECTIVE DATE BEFORE 2007 MAY HAVE OCCURRED AT ANY TIME BEFORE 2007. THE CORRECT EFFECTIVE DATE MAY BE DIFFERENT FROM THE ONE RECORDED. Effective date: 20050103 |
|
| PG25 | Lapsed in a contracting state [announced via postgrant information from national office to epo] |
Ref country code: DE Free format text: LAPSE BECAUSE OF NON-PAYMENT OF DUE FEES Effective date: 20050802 |
|
| PG25 | Lapsed in a contracting state [announced via postgrant information from national office to epo] |
Ref country code: GB Free format text: LAPSE BECAUSE OF NON-PAYMENT OF DUE FEES Effective date: 20060103 |
|
| GBPC | Gb: european patent ceased through non-payment of renewal fee |
Effective date: 20060103 |
































