EP4693281A1 - Frame recovery after frame loss using lsf stabilization - Google Patents
Frame recovery after frame loss using lsf stabilizationInfo
- Publication number
- EP4693281A1 EP4693281A1 EP24193936.2A EP24193936A EP4693281A1 EP 4693281 A1 EP4693281 A1 EP 4693281A1 EP 24193936 A EP24193936 A EP 24193936A EP 4693281 A1 EP4693281 A1 EP 4693281A1
- Authority
- EP
- European Patent Office
- Prior art keywords
- frame
- audio
- signal
- energy
- parameters
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L19/00—Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
- G10L19/005—Correction of errors induced by the transmission channel, if related to the coding algorithm
Definitions
- the present examples refer to an audio decoder.
- some examples refer to frame recovery after frame loss using LSF stabilization. Techniques for energy compensation are also disclosed.
- Frame Recovery is a strategy used in Audio Codecs to offer a smooth transition between a concealed frame and a received frame. If a frame gets lost, packet loss concealment (PLC) is applied, where data from the last good frame is used to render an artificial frame. If a received frame relies on information of the last frame, which is a concealed frame, then there might be a mismatch between concealed and received CELP parameters, like the Linear Prediction Coefficients (LPC) or the adaptive excitation. So, there is need to improve this transition between a lost frame and a received frame.
- PLC packet loss concealment
- LPC Linear Prediction Coefficients
- One option is to transfer side information from encoder to decoder like energy or phase information for a guided recovery.
- Another option is blind recovery. In case of blind recovery, the implantation is only done on decoder side.
- the filter is used to calculate the residual of the signal, but also to synthesize a signal.
- s(n) is the speech signal, which is usually pre-emphasized
- L ist the time length of the current frame
- x(n) is the residual.
- a weighted speech filter is used to obtain a weighted speech signal, to find the optimum pitch and innovation parameters, by minimizing the squared error between this weighted speech signal and the input signal.
- Further steps are the open-loop pitch search and closed loop pitch search to find the optimum pitch lag and pitch gain, also known as the Long-Term Prediction or adaptive codebook and the innovation excitation search.
- the optimum parameters on encoder side are quantized and the codebook indices are sent to the decoder.
- the PLC on decoder side is enabled to conceal a frame based on the data from the last good frame. If a new good frame is decoded after frame loss, decoder and encoder are not synchronized anymore and noticeable artifacts can occur in the speech.
- phase or periodicity information are transmitted from encoder to decoder to improve the robustness of the frame loss concealment. [3].
- the phase control is particularly important for the recovery.
- the decoder memory is desynchronized with the encoder memory.
- a rough position of the first glottal pulse is send.
- the precision the encoder the position of the pulse depends on the closed-loop pitch value for the first subframe pitch lag T0.
- the sample with the maximum amplitude is searched to find the position of the first glottal pulse, where T0 is also the length of the search interval.
- the highest precision for the glottal pulse is achieved with a low pass filtered residual signal.
- a periodicity parameter or voicing parameter can be transmitted.
- the voicing is estimated based on the normalized correlation of the signal. It can be encoded precisely using 4 bits.
- the voicing information is needed for highly voiced frames and better voicing resolution.
- the disadvantage of sending information from encoder to decoder is that bits need to be reserved (this is unwanted, and it will be shown that in the examples proposed here the decoder will not use this information, so that the encoder skips it).
- the most important part of the recovery is the energy control of the signal due to the strong prediction used in modern speech codecs. The energy needs to control in that way, that the energy at the beginning of the first good, received frame after frame erasure matches the energy at the end of the concealed frame.
- the signal is scaled to prevent a strong energy increase in the signal.
- EVS a scaling gain is applied to the decoded speech signal [3].
- the scaling is done in the excitation domain to serve the long-term prediction memory for the following frame.
- the synthesis is done again to achieve a smooth transition from concealed frame to received frame.
- L is the frame time length (in number of samples) and g(n) is the gain applied to the samples of the excitation (it is anticipated that some optional examples will be shown in which a particular g(n) is generated, different from the prior a).
- E q E 1 ⁇ E LP 0 E LP 1
- E LP 0 the linear prediction filter gain of the last good frame before erasure
- E LP 1 the linear prediction filter gain of the current frame after erasure. This is done to compensate a possible energy mismatch between excitation signal energy and the LP filter gain.
- go is set to 0.5 * g1.
- x old ( n ) is the excitation of the concealed frame
- L old is the length of the concealed excitation
- x in ( n ) is the excitation of the received frame
- T 0 is the estimated lag.
- T 0 is limited by L trans .
- Energy limitation is applied to the signal x trans ( n ) in case that the energy of the signal x trans ( n ) is higher than x old ( n ).
- the excitation is the residual of the LPC analysis filter.
- the LPC represents short-term characteristic of the signal and the LPC coefficients are interpolated for each subframe through their Linear Spectral Pairs (LSP) from the previous frame and the current frame.
- LSP Linear Spectral Pairs
- LSFs Linear Spectral Frequencies
- the excitation is determined based on the quantized and interpolated set of LPCs and the codebook indexes are determined. On decoder side the LSF are de-quantized.
- GANs generative adversarial networks
- an audio decoder for synthesizing an audio signal from a bitstream which represents the audio signal, the audio decoder including:
- a method for synthesizing an audio signal from a bitstream which represents the audio signal including:
- Fig. 5 can be taken into account, when reading below, for distinguishing between the different frames and subframes. Fig. 5 distinguishes between different scenarios and shows formulas which are discussed here-below.
- the current frame is normally referred to as a q-th frame of the sequence of frames. If the (q-1)-th frame has been a non-properly decoded frame (and therefore has been concealed), then the q-th frame is a recovery frame (provided that the q-th frame is also properly decoded). Otherwise, if the (q-1)-th frame has been a properly decoded frame (and therefore has not been concealed), and the q-th frame is a properly-received frame as well, then the q-th frame is a non-recovery frame. Finally, if the q-th frame is non-properly decoded, then it is a lost frame, and is substituted by a concealed frame synthesized by the concealment unit 40.
- the q-th frame is in general substituted into N subframes.
- the subframes are in general indexed with "k" (with 1 ⁇ k ⁇ N). Even in that case, the use of complicated pedices like "k,q" will be preferably avoided.
- pedex "end” will be used. That pedex "end” is to be understood as globally valid for a whole frame, and not for a particular subframe.
- LSF and LSP parameters are often indicated with pedex "end" (e.g. LSF end , LSP end , LSP' end ) despite being often intended as global values for a particular frame. The reason is that it is intended that, at the encoder side, these values are only calculated on a particular, final subframe (i.e. the N-th subframe of the frame) and, rigorously speaking, they should be intended as parameters for that particular N-th, final subframe. Notwithstanding, it is hypothesized that those values are globally valid for the entire frame, at least according to a first approximation. It will be shown that, in some cases (e.g.
- an interpolation (or another weighting function) will be carried out, e.g. through interpolation weights w k with w k increasing (e.g. linearly) for k increasing (e.g. it may be w 1 being a value between 0 and e.g. 0.3 or 0.2, and w N being a value between 0.7 or 0.8 and 1).
- Fig. 6 shows an example of an audio decoder 100 to synthesize an audio signal from a bitstream which represents the audio signal.
- the audio decoder 100 may include a bitstream receiver 5, which receives the bitstream (the bitstream receiver 5 may include a bitstream reader, not shown, which e.g. reads the bitstream e.g. from remote or from a storage unit).
- the bitstream may have, encoded therein, a set of encoded audio parameters (e.g., encoded versions of residual audio parameters, here indicated with ⁇ LSF q for each q-th frame) for each frame of the audio signal.
- Fig. 6 shows that the bitstream receiver 5 includes a dequantization block 602 (e.g.
- the dequantization block 602 may provide the encoded audio parameters as residual values ( ⁇ LSF q ) of linear spectral frequencies (LSFs).
- the audio decoder 100 may comprise a first audio parameter decoding unit 10 which can be understood substantially as operating as in the prior art.
- the first audio parameter decoding unit 10 may decode, for a current properly received frame, a first set of decoded audio parameters (a i1 ) from at least the set of the encoded audio parameters ( ⁇ LSF q ). As shown by Fig.
- the first audio parameter decoding unit 10 may include a block 604 for computation of first set of parameters LSPs (linear spectral pairs), which provides in output linear spectral pairs (LSP end ).
- the first audio parameter decoding unit 10 may include a block 606 for computation of a first set of LPC parameters which may be LPC coefficients indicated here as a i1 .
- the first audio parameter decoding unit 10 may normally operate for decoding properly-received frames (good frames).
- the audio decoder 100 may include (not shown in Fig. 6 , but shown for example in Fig. 1 ), a concealment unit 40.
- the concealment unit 40 may conceal at least one non-properly received frame (e.g. a (q-1)-th frame) based on at least one previously properly received frame (e.g. a (q-2)-th frame).
- the concealment unit 40 may therefore generate at least one set of concealment audio parameters (here indicated as LSP PLC ).
- the concealment unit 40 may synthesize a concealed frame from the at least one concealment audio parameters.
- the way how the concealment unit 40 operates is left general: we are not really interested in how the concealment unit 40 performs the concealment of the (q-1)-th non-properly received frame, but, instead, in how the LPC parameters (a i1 ) are obtained for the q-th properly received frame which immediately follows the (q-1)-th concealed frame.
- the properly received q-th frame which immediately follows the (q-1)-th concealed frame (substituting the non-properly received frame) is here called "recovery frame”. It is mostly intended to discuss the way of how to obtain the synthesis signal (and also the parameters for obtaining the synthesis signal) for the recovery frame.
- the first audio parameter decoding unit 10 may obtain the LPC parameters an by taking into account the LPC parameters used by the concealment unit 40 for performing the concealment for the preceding, (q-1)-th non-properly decoded frame.
- Block 604 may be inputted by a product of a predefined prediction factor m (e.g. a value between 0 and 1/3 (e.g. 0.333333) or a value between 1/4 and 1/2, or another natural number larger than 1/10 and less than 1) with concealment audio parameters ⁇ LSF q-1 (e.g. LSF parameters, which may be differential parameters).
- a predefined prediction factor m e.g. a value between 0 and 1/3 (e.g. 0.333333) or a value between 1/4 and 1/2, or another natural number larger than 1/10 and less than 1
- concealment audio parameters ⁇ LSF q-1 e.g. LSF parameters, which may be differential parameters.
- block 606 may be also inputted with LSP PLC (concealment parameters, which may be LSP parameters, taken from the concealed frame).
- LSP PLC concealment parameters, which may be LSP parameters, taken from the concealed frame.
- the blocks 604 and 606 operate normally by taking into account parameters from the immediately preceding (properly decoded) (q-1)-th frame (those parameters being also indicated with ⁇ LSF q-1 ).
- the LPC parameters a i1 are obtained as usual, by taking into account the parameters of the immediately preceding, properly decoded frame.
- the audio decoder 100 may include a synthesizing unit 50.
- the synthesizing unit 50 (in particular in block 608 for a first synthesis) may generate a first synthesis signal s 1 (n) from the LPC parameters a i1 .
- the synthesizing operation would be finished (apart from possible other operations such as post-processing and/or scaling by a gain g(n)) and the synthesized output audio signal s 1 (n) would be provided as the output of the synthesizer unit 50.
- the synthesizing unit 50 (in particular in block 608 for computation of first synthesis, inputted from by the excitation x(n) obtained from the previous frame and the LPC parameters a i1 from the first audio parameter decoding unit 10) may therefore provide a first synthesis signal s 1 (n).
- the audio decoder 100 also comprises a second audio parameter decoding unit 20.
- the second audio parameter decoding unit 20 may decode, for the currently q-th properly received frame which immediately follows the at least one non-properly received frame, a second set of audio parameters (e.g. a second set of LCP parameters) a i2 .
- the second audio parameter decoding unit 20 may include a block 612, for computation of second set of a LSP parameter (LSP' end ) for each frame, block 612 not being inputted with the audio parameters of the immediately preceding audio frame (or, in some examples, which is inputted by a number of concealment parameters less than the concealment parameters used inputted into the first audio parameter decoding unit 10).
- the linear spectral pair LSP' end may be provided to a block 614 for computation of second set of LPC parameters (which are indicated with a i2 ).
- Block 614 is not inputted with concealment parameters LSP PLC (or any other parameter taken from the concealment unit 40) but, rather, with LSP mean (which may be obtained from a table).
- the second audio parameter decoding unit 20, (in particular block 614) may output or derive the second set of decoded audio parameters (a i2 ) which is different (and in particular a i2 is different from the first set of decoded audio parameters a i1 , and in particular is not controlled by the concealment parameters of the immediately preceding concealed frame).
- the second audio parameter decoding unit 20 is deactivated in the case of the q-th current frame not being a recovery frame: in case of the current q-th frame being a non-properly received frame, then the synthesis is performed by the concealment unit 40, while in the case of the q-th current frame being a non-recovery frame (i.e. a properly decoded frame which immediately follows another, (q-1)-th properly decoded frame), then the synthesis is performed by the first audio parameter decoding unit 10 by keeping into account the parameters of the (q-1)-th immediately preceding audio frame (which is properly decoded), and the second audio parameter decoding unit 20 is deactivated.
- a non-recovery frame i.e. a properly decoded frame which immediately follows another, (q-1)-th properly decoded frame
- the synthesizing unit 50 (in particular in block 616 for computation of second synthesis, inputted by the excitation x(n), obtained from the previous frame, and the LPC parameters a i2 from block 614) may therefore provide a second synthesis signal s 2 (n).
- the synthesis signal s 2 (n) will compete with the synthesis signal s 1 (n) for being selected as the output or derived signal s(n) for the current frame.
- Block 604 of the first audio parameter decoding unit 10 and block 612 of the second audio parameter decoding unit 20 may each provide, in examples, a parameter which is valid for the whole current q-th frame, while block 606 of the first audio parameter decoding unit 10 and block 614 of the second audio parameter decoding unit 20 may each provide a respective LPC set of parameters (a i1 , a i2 ) for each subframe (indeed, we use the wording "set of parameters" both because the parameters vary with the particular subframe, and also because they vary with the index i).
- Fig. 6 also shows an excitation decoding block 622 which outputs an excitation x(n) (or z(n) if expressed as z-transform).
- the first operational step may also interest the first audio parameter decoding unit 10 (e.g. in at least one of blocks 604 and 606) and/or the second audio parameter decoding unit 20 (e.g. in at least one of blocks 612 and 614), because, in order to further save computational power, it is possible to only compute those audio parameters which are in the first subframe(s) of the first and/or second synthesis signals, without computing those audio parameters in the final subframe of the first and/or second synthesis signals (in practice, initially only decoding a first subset, which is a proper subset, of the first set of decoded audio parameters, and only decoding a second subset, which is a proper subset, of the second set of decoded audio parameters): as soon as one of the two synthesis versions of the synthesis signal is selected is selected, then only the first audio parameter decoding unit 10 (in case the first synthesis is selected) or the second audio parameter decoding unit 20 (in case the second synthesis is selected) will perform the decoding of the audio parameters of
- the non-void subset of the remaining decoded audio parameters which were initially not decoded is only subsequently decoded, and the synthesis of the selected version of the synthesis signal is processed using the remaining decoded audio parameters of the selected synthesis, disregarding the remaining decoded audio parameters of the non-selected synthesis).
- the audio decoder 100 may include (e.g., within the synthesizing unit 50), a selector 60 which selects between the first synthesis signal s 1 (n) (or the first set of parameters an) and a synthesis signal s 2 (n) (or the second set of parameters a i2 ).
- the selector 60 may include a block 610 for energy computation and/or envelope evaluation for the first synthesis, which may provide a first energy-related measurement.
- the selector 60 may include block 618 for energy-related measurement computation (e.g. energy computation and/or envelope computation) for the second synthesis which provides an energy-related measurement (e.g. energy computation and/or envelope computation) on the second synthesis signal s 2 (n).
- a block 620 for decision of LPC set and synthesis may decide which synthesis signal, among s 1 (n) and s 2 (n) to be used as a synthesis signal s(n) (output version or derived version of the synthesis signal).
- a block 623 synthesis smoothing/scaling may also be used, inputted with the selected signal s(n), to thereby process the output version or derived version of the synthesis signal and render it.
- Fig. 1 also shows elements of the audio decoder 100 in terms of block scheme.
- a decision 41 is made between determining whether the current frame is a non-properly decoded frame or a properly decoded frame. If the current frame is a non-properly decoded frame (e.g. by a determination based on cyclical recurrent calculations, such as cyclic redundancy check, CRC, or the like) then the concealment unit 40 may be activated. Otherwise (if the frame is determined as valid), in block 42, it is evaluated whether the previous frame was lost. If the previous frame was lost, then the current frame is a recovery frame. Therefore, at block 43, both the first audio parameter decoding unit 10 and the second audio parameter decoding unit 20 are activated.
- the synthesizing unit 50 is invoked.
- the first audio parameter decoding unit 10 only is activated and the synthesizing unit is activated as well in block 44, and there is an updating 45 of the memory for the next frame.
- first synthesis signal s 1 (n) and the second synthesis signal s 2 (n) may be made based on the comparison between the energies of the signals, and in particular on the behavior of the envelope of the signals.
- Fig. 7 shows the example 100 with other blocks, which may be optional in some examples.
- Fig. 7 shows, in particular, the block 620 for decision of LPC set and synthesis.
- Block 620 may provide both LPC parameters a i (chosen between the LPC parameters a i1 of the first synthesis and a i2 of the second synthesis) and provide those parameters to a zero input response block 702. (This can be a outringing filter. The last e.g. 16 samples (or another amount, e.g. less than 32 samples) of the previous (concealed) frame are inputted and also a zero input excitation. Then an outringing synthesis is performed).
- the zero input response block 702 may therefore provide a zero input response (ZIR) indicated with s ZIR (n).
- Block 620 may also output or derive the chosen synthesis signal s(n) between s 1 (n) and s 2 (n) to the subtractor block 704.
- the synthesis signal s(n) may therefore be subtracted with the zero impulse response signal s ZIR (n) in the subtractor block 704.
- the energy computation block 706 may compute an energy of the synthesis signal s(n).
- a gain computation block 708 may be used to obtain a first gain g 1 and a second gain g 2 .
- Values of the first gain g 1 and of the second gain g 2 may be provided to a scaler 360 to provide a gain g(n), which may be defined sample-by-sample, and may, for each sample, take a value between the first gain g 1 and the second gain g 2 .
- the scaler 360 may be conditioned by the concealment unit 40 and, in particular, by the energy E PLC () of the concealed frame (it will be shown, in particular, that the scaler 360 may be conditioned by the values of the energy in a final portions (e.g. final half portion) of the concealed frame).
- the output of the scaler 360 may be added with the zero input response s ZIR (n) at adder 710.
- a value s s (n) is to be provided to a residual computation block 712.
- a post-processing block 716 may be used from the output s s (n) of the adder 710.
- a block 718 of updating excitation for the next frame permits to obtain the residual computation from the output x s (n).
- a block 620 for decision is also used to provide a coefficient ⁇ value (proportionality coefficient) to the gain computation block 708 and the parameters to the residual computation block 712.
- the audio signal s s (n) may therefore be post-processed and used and rendered as audio signal.
- m is a prediction factor (e.g. a value between 0 and 1/3 (e.g. 0.333333) or a value between 1/4 and 1/2, or another natural number larger than 1/10 and less than 1)
- ⁇ LSF q -1 is the LSF prediction residual decoded for the previous (q-1)-th frame.
- LSP end cos 2 * ⁇ * LSF end f s where f s is the sampling frequency.
- LSP end -1 (which a priori should be called LSP q -1, N since it is the LSP value of the last, N-th subframe of the immediately preceding, (q-1)-th concealed frame) are the old end LSP coefficients of the end subframe (i.e.
- LSP end is the LSP coefficient of the end subframe (i.e. of the N-th subframe) of the current frame (as explained above, used globally for the whole current q-th frame).
- LSF end which a priority should be written LSF q,N
- LPC does not change to much about frame, because it is a short-term representation.
- the LSP k (which a priority should be written LSP q,k ) are converted to the LPC k (which are notwithstanding written a i1 ).
- LSP PLC which could be written LSP end-1
- LSP PLC have been inferred by taking into account at least the previously correctly received frame, e.g. the properly-received (q-2)-th frame before the (q-1)-th concealed frame.
- LSF end LSF mean + m ⁇ ⁇ LSF PLC + ⁇ LSF q
- ⁇ LSF PLC which can also be written as ⁇ LSF q -1
- ⁇ LSF q -1 is a concealed version of the residual in the previous, (q-1)-th, concealed frame. It is not of interest how that delta or residual is concealed, but it is important to know that the residual of the previous concealed frame is used for the recovery frame.
- ⁇ LSF PLC of the concealed frame and the transmitted ⁇ LSF q -1 in case the frame is not lost and also if the LSP PLC of the concealed frame (i.e., the LPC coefficients which have been inferred during concealment) and the LSP end -1 , in case the frame is not lost, distinguish from each other, here can be a high deviation of the LPC filter response between erroneous signal and clean signal.
- a big difference in ⁇ LSF PLC in a previous concealed frame and ⁇ LSF q -1 in a previous received frame can also means a big difference in ⁇ LSF PLC and ⁇ LSF q in the recovery frame and the same for LSP PLC in the previous concealed frame and LSP end in the recovery.
- Figs. 10 shows that it is possible (e.g., by relying on the second audio parameter decoding unit 20) to obtain a more stable output signal (or derived signal).
- the presented example proposes, inter alia, a technique introducing a second set of LPC coefficients to synthesize a second signal s 2 (n) in the decoder 100.
- An idea is to compare (e.g. at selector 620) the energy (E 1k , E 2k ) of two synthesis signals s 1 (n) and s 2 (n) (or of initial portions of s 1 (n) and s 2 (n)):
- the second set of LSP parameters may be obtained by using a mean LSP (LSP mean ) of the recovery frame instead of LSP PLC of the previous, concealed frame.
- LSP mean mean LSP
- m being pre-defined (e.g. a value between 0 and 1/3 (e.g. 0.333333) or a value between 1/4 and 1/2, or another natural number larger than 1/10 and less than 1), so that the prediction contains only the mean LSF.
- LSF end LSF mean + ⁇ LSF PLC + ⁇ lsf q with LSF mean obtained from a table, ⁇ LSF PLC obtained from the concealment unit 40, and ⁇ lsf q read in the bitstream
- LSF end ′ LSF mean + ⁇ lsf q with LSF mean obtained from a table and ⁇ lsf q read in the bitstream
- LSP ′ k LSP mean ⁇ 1 ⁇ w k + LSP ′ end ⁇ w k
- w k are interpolation weights (e.g., increasing linearly with the increase of k, e.g. it may be w 1 being a value between 0 and e.g. 0.3 or 0.2, and w N being a value between 0.7 or 0.8 and 1) and N is the number of subframes for the frame.
- LSP k cos 2 * ⁇ * LSF PLC f s ⁇ 1 ⁇ w k + cos 2 * ⁇ * LSF mean + m ⁇ ⁇ LSF PLC + ⁇ LSF q f s ⁇ w k
- Second set (at the second decoding unit 20):
- LSP ′ k cos 2 * ⁇ * LS F mean f s ⁇ 1 ⁇ w k + cos 2 * ⁇ * LSF mean + ⁇ LSF q f s ⁇ w k
- LPC Linear Prediction Coefficients
- the second set of LSP parameters LSP' k is converted (in block 614 of the second audio parameter decoding unit 20) to a second set of LPC parameters.
- the two synthesis signals s i1 (n) and s i2 (n) are therefore synthesized.
- the first synthesis signal s i1 (n) is synthesized (at block 608 of the synthesizing unit 50) by using the first set of LPC coefficients
- the subframe energies of both first and second synthesis signals s 1 ( n ) and s 2 ( n ) are compared.
- the signal (among signals s 1 ( n ) and s 2 ( n )) with the smoother energy envelope compared to the energy in the last subframe of the concealed signal is chosen.
- L 1 L_initial
- the energies may be scaled or divided by the length of the subframe or the length on which the energy is calculated on to be comparable in case the frame, half frame or subframe length are changing from the concealed frame to the received frame.
- At least one condition can be taken into account.
- some conditions may be used alone, but in some examples more than one condition are evaluated. For this reason, it is sometimes written in terms like "if the ... condition is fulfilled, then the first/second synthesis signal is preferentially selected", where "preferentially” means that the particular synthesis signal is selected either tout-court (e.g. in the examples in which only one condition is present) or that the particular synthesis signal is selected provided that other conditions are fulfilled.
- E 1,1 is a measurement of the energy of the first subframe of the first synthesis signal s 1 ( n ) and E 2,1 is a measurement of the energy of the first subframe of the second synthesis signal s 2 ( n ), E 1,2 is a measurement of the energy of the second subframe of the first synthesis signal s 1 ( n ) and E 2,2 is a measurement of the energy of the second subframe of the second synthesis signal s 2 ( n ).
- th 1 may have a value of 1.455 (or more in general between 1.4 and 1.5, or even more in general a value which is 1 or larger than 1) and th 2 may have a value of 1.5 (or more in general between 1.45 and 1.55, or even more in general a value which is 1 or larger than 1), and it may be preferably th 1 ⁇ th 2 , and it may be th 1 > 0 and th 2 > 0.
- the preferred values were determined empirically. If the first condition is fulfilled, then the second synthesis signal s 2 ( n ) (and the second set of LPC) is selected to be used, subjected to the second and/or third conditions, in examples.
- the first condition may be generalized even more as: if, along a number N MAX ⁇ 2 (with N MAX ⁇ N or N MAX ⁇ N) of the first consecutive subframes, the energy of the first synthesis signal s 1 ( n ) evolves coherently over (e.g. by at least a threshold e.g. of at least 40%) the energy of the first synthesis signal, then the second synthesis signal is preferentially selected (subjected to the other conditions, if present). Otherwise, if in at least one subframe of the first N MAX ⁇ 2 (with N MAX ⁇ N or N MAX ⁇ N) consecutive subframes the energy of the first synthesis signal does not evolve coherently over the energy of the first synthesis signal (e.g.
- the first synthesis signal s 1 ( n ) is chosen (in some examples, even without checking other conditions, even if present).
- the energy of the first two subframes (E 2,1 and E 2,2 ) of the second synthesis signal may be compared to the energy ( E PLC sub ) of the last subframe in the concealed frame, which is derived in same way like the subframe energy of the received frame: E 2,1 E PLC sub > th PLC & & E 2,2 E PLC sub > th PLC where "&&" means the logical operator "AND”.
- Each ratio E 2,1 E PLC sub and E 2,2 E PLC sub has to be higher than the threshold th PLC to fulfil the second condition. If the second condition is fulfilled, the second synthesis signal is preferentially chosen (in some examples, subjected to the fulfilment of other conditions, e.g. the first condition and/or the third condition).
- the first synthesis signal s 1 ( n ) is selected.
- the subframe energy (E 1,1 and E 1,2 ) of the first and second subframe of the first synthesis signal compared to the energy ( E PLC ) of the last subframe of the concealed signal has to be higher than a certain threshold (e.g., the same th PLC ): E 1,1 E PLC sub > th PLC & & E 1,2 E PLC sub > th PLC where "&&" means the logical operator "AND”.
- a certain threshold e.g., the same th PLC
- "&&" means the logical operator "AND”.
- the second condition may be generalized in that: if the energies (E 1,1 , E 1,2 , E 2,1 , E 2,2 ) of both the first subframe and the second subframe of both the first synthesis signal s 1 ( n ) and the second synthesis signal s 2 ( n ) are larger than the energy of the last subframe of the concealment signal by a predefined amount (e.g. 50%), then the second synthesis signal s 2 ( n ) is preferentially chosen (in some examples, subjected to the fulfilment of other conditions, e.g. the first condition and/or the third condition).
- a predefined amount e.g. 50%
- the first synthesis signal s 1 ( n ) is preferentially selected (in some examples, without examining the fulfilment of other conditions, if present).
- the second condition may be generalized eve more in that: if the energies (e.g. E 1,1 , E 1,2 , E 2,1 , E 2,2 ) of a number N max (with N max between 2 and N or between 2 and N-1) of initial consecutive subframes of both the first synthesis signal s 1 ( n ) and the second synthesis signal s 2 ( n ) are all larger than the energy of the last subframe of the concealment signal by a predefined amount, then the second synthesis signal s 2 ( n ) is preferentially chosen (in some examples, subjected to the fulfilment of other conditions, e.g. the first condition and/or the third condition). Otherwise, if at least one of the energies (e.g.
- the first synthesis signal s 1 ( n ) is preferentially chosen (in some examples, without examining the fulfilment of other conditions, if present).
- th 3 / th 1 >1.36 (in some examples by at least 1.2). At least one of these conditions must fulfilled to allow the second synthesis signal to be used. However, this third condition is optional and can be dropped.
- the second synthesis signal is used, when: E 1,1 E 2,1 > th 1 & & E 1,2 E 2,2 > th 2 & & E 2,1 E PLC sub > th PLC & & E 2,2 E PLC sub > th PLC & & E 1,1 E PLC sub > th PLC & & E 1,2 E PLC sub > th PLC & & & &
- an energy compensating function g(n) is applied to the selected synthesis signal which is controlled by two gains g 1 and g 2 .
- the first gain g1 may be determined by the ratio of the energy of the last half frame (or at least one last portion) of the concealed signal in the concealed frame and the first half frame (or at least one first portion) of the recovery frame.
- g 1 may be limited to a ceiling value of 1.2 (or more in general a value D with 1 ⁇ D ⁇ 2, e.g. a value between 1.1 and 1.3): if E PLC half E frame 2 is larger than 1.2, g 1 will be set to 1.2 (or D), so that g 1 is between 0 and 1.2 (or between 0 and D).
- g(n) permits to scale, sample by sample, the output or derived version s(n) of the synthesized audio signal in the properly received frame immediately following the at least one previously non-properly received frame by an energy compensating gain greater than 0, the energy compensating gain evolving, monotonically (e.g. strictly monotonically) or constantly, from the first value g 1 towards the second value g 2 .
- g 1 may be:
- g 1 may be conditioned by a conditioning term ( ⁇ * ( E frame 2 - E PLC half )) which may be proportional to the difference between the energy ( E frame 2 ) of a last portion of the of the current frame and the energy ( E PLC half ) of the last portion of the concealed frame.
- g 2 may be:
- g 2 may be the same of the ratio E PLC half + ⁇ ⁇ E frame 2 ⁇ E PLC half E frame 2 .
- g 2 may be defined so that, the higher the ratio, the higher g 2 (e.g. with a ceiling and/or a floor)).
- g 2 is higher (e.g. closer to 1) than if E PLC half and E frame 2 are distant from each other.
- E frame 2 when E frame 2 is smaller than or equal to E PLC half , (i.e. if E frame 2 E PLC half ⁇ 1 ), then g 2 may be set to be constantly 1 (ceiling value).
- E frame 2 > E PLC half i.e. E frame 2 E PLC half > 1
- g 2 lies between 0 and 1 (and the higher E frame 2 E PLC half , the lower g 2 , apart from a possible floor).
- the term ⁇ * ( E frame 2 - E PLC half ) is a conditioning term which ensures that the energy compensation is not too strong if there is a big difference between E PLC half and E frame 2 .
- the gain g(n) can be between 0 and 1 (and therefore an attenuation is performed) or (at least for some samples) larger than 1 (and in that case being amplifying). In some cases, the gain amplifies in a first part of the frame and attenuates in a last part of the frame.
- the gain g(n) causes an energy compensation: the more the energy of the last portion of the concealed signal in the concealed frame is different from the energy of the initial portion of the output or derived synthesis signal in the current q-th frame, the larger the conditioning caused by g(n).
- weak monotonicity may be also possible (e.g., following possible quantization of values of g(n) and/or in the case that a ceiling value or floor value is taken for an interval of samples).
- the proportionality coefficient ⁇ is a proportionality factor which depends on the energy in the first subframes of the chosen synthesis.
- m may have a value like -2.14e-05 or -2e-05 ⁇ m ⁇ -3e-05 (or another value, which may be a negative value) and c a value 0.1999 ⁇ c ⁇ 0.20001 (or another value, which may be a positive value).
- ⁇ may have a constant value like 0.2, or 0.1999 ⁇ ⁇ ⁇ 0.20001. More in general, the coefficient ⁇ may be obtained as a linear combination of E rel , e.g. with a negative angular coefficient m and/or positive constant term.
- E rel max E 1,1 E PLC sub E 1,2 E PLC sub
- AGC is the active gain control with a value of 0.98 (or more in general a value between 0.9 and 0.999, e.g. between 0.97 and 0.99).
- the gain g(n) is computed in block 708.
- the energy compensation is an attenuation, because the energy compensating gain is between 0 and 1.
- the higher the ratio E frame 2 E PLC half the higher the attenuation (i.e., the closer g(n) is to 0, g(n) being greater than 0).
- the lower the alpha the higher the attenuation (i.e., the closer g(n) is to 0, g(n) being greater than 0).
- the smoothing is done directly in the synthesis domain, so that the impact of the previous filter memory is not considered. This could in principle lead to discontinuities at the frame border.
- Fig. 10 shows the result of using the second set of LPC (second audio parameter decoder 20). Compared to the signal after a frame loss in Fig. 4 , the second synthesis signal s 2 (n) is clearly stable and similar to the clean signal.
- a first subframe for the first synthesis signal could have length L 1,1 while a second subframe for the second synthesis signal could have length L 1,2 ).
- it could be possible to modify the thresholds e.g.
- the energy-related measurements are not uniquely energy measurements.
- envelope measurements (which are also energy-related measurements) may be performed.
- the envelope can be described as the change (progression) of the energy over time (e.g., over the samples).
- blocks 608 and 610 may obtain an envelope of the first synthesis signal s 1 (n) and the second synthesis signal s 2 (n), respectively.
- the envelopes may be evaluated so as:
- the stability of the signals may be evaluated, for example, by measuring the fluctuations of the signal over time.
- Fig. 4 shown, for example, a highly fluctuating signal ("Synthesis signal obtained using the parameters of the concealed frame: set 1"). Measurements of fluctuations of an envelope are known in the art, and can be based, for example, on variance measurements, etc. In general terms, however, where the envelope of the second synthesis signal has a stability which overwhelms the stability of the first synthesis signal by at least the predetermined extent (by the predetermined threshold), then the second synthesis signal is enough stable and can be used instead of the first synthesis signal as selected synthesis signal (and as derived or output synthesis signal).
- first and second conditions described above may be used, in some examples, as evaluating the stability of the first and second synthesis signals.
- Figs. 2 and 3 show the LPC frequency response of the concealed frame (Aq_dec) vs. the clean signal (Aq_dec_clean).
- Fir. 4 shows the clean signal compared to the synthesis signal (e.g. s 1 (n)) obtained using the parameters of the concealed frame (e.g., like in the prior art, or as outputted by the first audio parameter decoding unit 10), and shows evident oscillations.
- Fig. 9 shows:
- Fig. 10 shows the clean signal compared to the synthesis signal (s 2 (n)) obtained without using the parameters of the concealed frame (as outputted by the second audio parameter decoding unit 10), and shows a stable behavior as compared to that of Fig. 4 .
- the current q-th frame and the (q-1)-th concealed frame are partitioned both: according to a first partitioning which partitions the current q-th frame and the (q-1)-th concealed frame among a sequence of subframes in a number N of subframes which is 3 or more than 3 (examples of these subframes are E 1,1 , E 2,1 , E 1,2 , E 2,2 and E PLC sub etc.), the decoded audio parameters being decoded for each subframe, the selection at 60 being performed only based on an initial subframe or a group (sequence) of initial subframes of the current q-th frame and on the last subframe of the (q-1)-th concealed frame (further, in case of partial synthesis before the selection 60, the synthesis being originally only performed for on single initial subframe or a particular sequence of initial subframes; further and E rel being calculated based on the first subframe only); according to a second partitioning which partitions the current frame and the
- half frames in a number of portions which is 2 (in the case of half frames) or more than 2, but the number of portions being less than the number N of subframes, and each portion having larger time length than any subframe (this partitioning is preferably used for calculating the energy compensating gain g(n); examples are E PLC half , E frame 2 , E frame 1 ).
- the second audio parameter decoding unit 20 mainly refrains from adopting the parameters from the concealed signal, it is notwithstanding noted that the excitation x(n) can be obtained from the (q-1)-th concealed audio frame.
- the encoded parameters are differential parameters and the audio signal is synthesized by taking into account an excitation (which can be obtained, or at least inferred, in some examples, from the concealed audio signal in the (q-1)-th concealed frame).
- a major advantage of aspects of the proposed technique is the consideration of the mismatch between LPC parameters of a concealed loss frame and a first well-received, aka recovery, q-th frame.
- the prior art methods propose smoothing operations of the excitation domain signal but don't treat the mismatch sets of LPC into account. Further, the proposed technique operates solely on decoder side and requires no extra side information to be transmitted.
- Fig. 11 An implementation based on a neural network is illustrated in Fig. 11 , where the synthesizing unit 50 of Fig. 1 includes a neural network, NN-based synthesizing unit 90, which implements a NN.
- block 620 (and the selector 60) is represented as being part of the block "Recovery and Computation/Decision of speech parameters" 91 (which implements blocks 10 and 20, in particular), while block 623 is part of the NN-based synthesizing unit 90.
- Transmitted speech parameters like LPC or any representation of it may be used as input for a neural network NN (in a NN processor 90) with at least one learnable layer (e.g. with a plurality of learnable layers).
- the learnable layer may be a generative adversarial neural (GAN) layer.
- concealed parameters generated by the concealment unit 40, which is not necessary part of the NN-based synthesizing unit
- LCP parameters a i1 and a i2 are fed to the NN-based synthesizing unit 90.
- the received speech parameters are used (i.e. the first set of parameters is used, and the first audio parameter decoding unit 10 is activated, while the second audio parameter decoding unit 20 is deactivated).
- recovery frame i.e.
- the block 620 (which may be a deterministic block) decides which synthesis signal between s 1 (n) and s 2 (n) is to be used.
- Speech (audio) parameters are in this case used by the NN-based synthesizing unit 90 (also implementing blocks 612, 614, and 616) to generate the synthesis signal, e.g. based on the excitation (e.g. previously obtained).
- a NN may be used.
- the proposed blocks 604, 608, 610, 614, 616, 618 may be put in front of a learnable layer of the NN-based synthesizing unit 90 where the prediction is created from the LPC, which are converted from LSP from the first set or second set in recovery case, LSP plc in case of concealment or LSP from the first set in clean channel case.
- an audio decoder for synthesizing an audio signal (e.g. s) from a bitstream which represents the audio signal, the audio decoder (e.g. 100) including:
- the first audio parameter decoding unit (e.g. 10) is configured, in the case the current frame is the properly received frame which immediately follows the at least one previously non-properly received frame, to decode the first set of decoded audio parameters from the at least one set of concealment parameters of the immediately preceding non-properly received frame and the set of encoded audio parameters of the current frame.
- the first audio parameter decoding unit (e.g. 10) is configured, in the case the current frame is the properly received frame which immediately follows the at least one previously non-properly received frame, to decode the first set of decoded audio parameters through a first prediction (e.g. 604) of the first set of audio parameters obtained from the version of the set of concealment parameters of the immediately preceding non-properly received frame and a pre-defined value.
- a first prediction e.g. 604
- the second audio parameter decoding unit (e.g. 20) is configured, in the case the current frame is the properly received frame which immediately follows the at least one previously non-properly received frame, to decode (e.g. 612) the second set of decoded audio parameters from the encoded audio parameters of the current frame.
- the second audio parameter decoding unit (e.g. 20) is configured, in the case the current frame is the properly received frame which immediately follows the at least one previously non-properly received frame, to decode the second set of decoded audio parameters from encoded audio parameters of the current frame but not from the set of concealment parameters of the immediately preceding non-properly received frame.
- the first audio parameter decoding unit is configured, in the case the current frame is the properly received frame which immediately follows the at least one previously non-properly received frame, to decode the first set of decoded audio parameters, from a first number of concealment parameters of the set of concealment parameters of the immediately preceding non-properly received frame and from the encoded audio parameters of the current frame
- the second audio parameter decoding unit is configured, in the case the current frame is the properly received frame which immediately follows the at least one previously non-properly received frame, to decode the second set of decoded audio parameters from the encoded audio parameters of the current frame and from a second number of concealment parameters of the set of concealment parameters of the immediately preceding non-properly received frame which is smaller than the first number of concealment parameters.
- the audio synthesizer may be configured, in the case the current frame is the properly decoded frame which immediately follows the at least one previously non-properly received frame, to perform the selection (e.g. 620) between the first synthesis signal (e.g. s1) and the second synthesis signal (e.g. s2) based on a comparison between at least energy-related measurements on the first version (e.g. s1) of the synthesized audio signal with energy-related measurements on the second version (e.g. s2) of the synthesized audio signal, so as to output or derive, as the output or derived version (e.g. s) of the synthesized audio signal, the second version (e.g. s2) of the synthesized audio signal in case at least one of the following condition or a combination of at least two of the following conditions is satisfied:
- the audio synthesizer may be configured, in the case the current frame is the properly decoded frame which immediately follows the at least one previously non-properly received frame, to perform the selection (e.g. 620) between the first synthesis signal (e.g. s1) and the second synthesis signal (e.g. s2) at least based on a comparison between at least energy-related measurements on the first version (e.g. s1) of the synthesized audio signal with energy-related measurements on the second version (e.g. s2) of the synthesized audio signal, so as to output or derive, as the output or derived version (e.g. s) of the synthesized audio signal, the second version (e.g. s2) of the synthesized audio signal in case of at least one of the following conditions, or a combination of at least one of the following conditions, is satisfied:
- the audio synthesizer may be configured, in the case the current frame is the properly decoded frame which immediately follows the at least one previously non-properly received frame, to perform the selection (e.g. 620) between the first synthesis signal (e.g. s1) and the second synthesis signal (e.g. s2) based on a comparison between an envelope of the first synthesis signal and an envelope of the second synthesis signal, so as to select the second synthesis signal (e.g. s2) in case the envelope of the second synthesis signal (e.g. s2) is, at least in the initial subframe or in a sequence of initial subframes, more stable, by at least one predetermined extent, than the envelope of the first synthesis signal (e.g. s1), and to select the first synthesis signal otherwise.
- the selection e.g. 620
- the first synthesis signal e.g. s1
- the second synthesis signal e.g. s2
- the second synthesis signal e.g. s2
- the audio decoder is configured to scale (e.g. 360), sample by sample, the output or derived version (e.g. s) of the synthesized audio signal in the properly received frame immediately following the at least one previously non-properly received frame by an energy compensating gain greater than 0, the energy compensating gain reducing, in at least one portion of the current frame, the energy gap between the concealed signal in a last portion of the concealed frame and the output or derived version of the synthesized audio signal in the at least one portion of the current frame.
- the energy compensating gain evolves, monotonically or constantly, from a first value towards a second value, the first value being:
- the current frame and the concealed frame are partitioned both according to a first partitioning which partitions the current frame and the concealed frame among a sequence of subframes in a number of subframes which is 3 or more than 3, and according to a second partitioning which partitions the current frame and the concealed frame among a sequence of portions in a number of portions which is 2 or more than 2, but the number of portions being less than the number of subframes, and each portion having larger time length than any subframe.
- the conditioning term has a proportionality coefficient ⁇ is linearly dependent on a maximum value between a first ratio and a second ratio, where the first ratio is a ratio between the energy of the initial subframe of the output or derived version (e.g. s) of the synthesized audio signal and the energy of the last subframe of the concealed frame, and the second ratio is a ratio between the energy of the second subframe of the output or derived version (e.g. s) of the synthesized audio signal and the energy of the last subframe of the concealed frame.
- the first ratio is a ratio between the energy of the initial subframe of the output or derived version (e.g. s) of the synthesized audio signal and the energy of the last subframe of the concealed frame
- the second ratio is a ratio between the energy of the second subframe of the output or derived version (e.g. s) of the synthesized audio signal and the energy of the last subframe of the concealed frame.
- the energy compensating gain is comparatively close to 1, in at least one portion of the current frame, in case the distance between the energy of the output or derived version (e.g. s) of the synthesized audio signal in the at least one portion of the current frame and the energy of concealed signal in the end portion of the concealed frame is comparatively low, and the energy compensating gain is comparatively distant from 1, in the at least one portion of the current frame, in case the distance between the energy of the output or derived version (e.g. s) of the synthesized audio signal in the at least one portion of the current frame and the energy of concealed signal in the end portion of the concealed frame is comparatively high.
- the energy compensating gain is defined recursively by cross-fading the energy compensating gain for an immediately preceding sample with the second gain value.
- the synthesizing unit (e.g. 50) is configured to, initially, partially synthesize only an initial subframe, or a group of initial subframes, of the of the first synthesis signal (e.g. s 1 ) and, partially synthesize only an initial subframe, or a group of initial subframes, of the second synthesis signal (e.g. s 2 ), so that the selection (e.g. 60) is based on energy-related measurements on the partially synthesized version of the first synthesis signal (e.g. s 1 ) and the partially synthesized version of the second synthesis signal (e.g. s 2 ), so that, only after the selection (e.g. 60), the remaining subframe or subframe of the selected synthesis signal is or are synthesized, without synthesizing the remaining subframe or subframe of the non-selected synthesis signal.
- the first audio parameter decoding unit (e.g. 10) is configured to, initially, decode, respectively, only a first subset of the first set of decoded audio parameters and only a second subset of the second decoded audio parameters, the first subset and second subset corresponding to the initial subframe, or the group of initial subframes, so that the selection (e.g. 60) is based on energy-related measurements of the partially synthesized version of the first synthesis signal (e.g. s 1 ) obtained from the first subset and the partially synthesized version of the second synthesis signal (e.g. s 2 ) obtained from the second subset, so that, only after the selection (e.g. 60), the remaining audio parameters of the set of audio parameter associated with the selected synthesis signal are decoded, and the remaining audio parameters of the set of audio parameter associated with the non-selected synthesis signal are not decoded.
- the selection e.g. 60
- a neural network processor using at least one learnable layer, to be inputted with the concealment parameters as well as the decoded parameters and/or the first and second synthesis signals or the output or derived version of the synthesis signal, so as to process the output or derived version of the synthesis signal through the at least one learnable layer.
- the at least one learnable layers is a generative adversarial network, GAN, learnable layer.
- the encoded audio parameter include, or provide information on, linear spectral frequencies.
- a method for synthesizing an audio signal e.g. s
- the method including:
- a non-transitory storage unit storing instructions which, when executed by a processor, cause the processor to perform the method above (or any of the methods above and below).
- examples may be implemented in hardware.
- the implementation may be performed using a digital storage medium, for example a floppy disk, a Digital Versatile Disc (DVD), a Blu-Ray Disc, a Compact Disc (CD), a Read-only Memory (ROM), a Programmable Read-only Memory (PROM), an Erasable and Programmable Read-only Memory (EPROM), an Electrically Erasable Programmable Read-Only Memory (EEPROM) or a flash memory, having electronically readable control signals stored thereon, which cooperate (or are capable of cooperating) with a programmable computer system such that the respective method is performed. Therefore, the digital storage medium may be computer readable.
- DVD Digital Versatile Disc
- CD Compact Disc
- ROM Read-only Memory
- PROM Programmable Read-only Memory
- EPROM Erasable and Programmable Read-only Memory
- EEPROM Electrically Erasable Programmable Read-Only Memory
- flash memory having electronically readable control signals stored thereon, which cooperate (or are capable of
- examples may be implemented as a computer program product with program instructions, the program instructions being operative for performing one of the methods when the computer program product runs on a computer.
- the program instructions may for example be stored on a machine readable medium.
- Examples comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier.
- an example of method is, therefore, a computer program having program instructions for performing one of the methods described herein, when the computer program runs on a computer.
- a further example of the methods is, therefore, a data carrier medium (or a digital storage medium, or a computer-readable medium) comprising, recorded thereon, the computer program for performing one of the methods described herein.
- the data carrier medium, the digital storage medium or the recorded medium are tangible and/or non-transitionary, rather than signals which are intangible and transitory.
- a further example comprises a processing unit, for example a computer, or a programmable logic device performing one of the methods described herein.
- a further example comprises a computer having installed thereon the computer program for performing one of the methods described herein.
- a further example comprises an apparatus or a system transferring (for example, electronically or optically) a computer program for performing one of the methods described herein to a receiver.
- the receiver may, for example, be a computer, a mobile device, a memory device or the like.
- the apparatus or system may, for example, comprise a file server for transferring the computer program to the receiver.
- a programmable logic device for example, a field programmable gate array
- a field programmable gate array may cooperate with a microprocessor in order to perform one of the methods described herein.
- the methods may be performed by any appropriate hardware apparatus.
Landscapes
- Engineering & Computer Science (AREA)
- Computational Linguistics (AREA)
- Signal Processing (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Human Computer Interaction (AREA)
- Physics & Mathematics (AREA)
- Acoustics & Sound (AREA)
- Multimedia (AREA)
- Compression, Expansion, Code Conversion, And Decoders (AREA)
Abstract
There is provided an Audio decoder (100) for synthesizing an audio signal from a bitstream which represents the audio signal, the audio decoder (100) including:
a bitstream receiver (5), to receive the bitstream representative of the audio signal, the bitstream having, encoded therein, a set of encoded audio parameters (ΔLSFq ) for each frame of the audio signal,
a first audio parameter decoding unit (10), to decode, for a current properly received frame, a first set of decoded audio parameters (ai1) from at least the set of encoded audio parameters,
a concealment unit (40) to conceal at least one non-properly received frame based on at least one previously properly received frame, so as to generate at least one set of concealment audio parameters;
a second audio parameter decoding unit (20), to decode, for a current properly received frame which immediately follows the at least one previously non-properly received frame, a second set of decoded audio parameters (ai2) from at least the set of encoded audio parameters, the second set of decoded audio parameters (ai2) being different from the first set of decoded audio parameters (ai1);
a synthesizing unit (50), to output, or derive, an output or derived version of the synthesized audio signal in such way that, if the current properly received frame immediately follows the at least one previously non-properly received frame, a selection (60) is made between:
a first version (s1) of the synthesized audio signal, synthesized from at least the first set of decoded audio parameters; and
a second version (s2) of the synthesized audio signal, synthesized from the second set of decoded audio parameters.
a bitstream receiver (5), to receive the bitstream representative of the audio signal, the bitstream having, encoded therein, a set of encoded audio parameters (ΔLSFq ) for each frame of the audio signal,
a first audio parameter decoding unit (10), to decode, for a current properly received frame, a first set of decoded audio parameters (ai1) from at least the set of encoded audio parameters,
a concealment unit (40) to conceal at least one non-properly received frame based on at least one previously properly received frame, so as to generate at least one set of concealment audio parameters;
a second audio parameter decoding unit (20), to decode, for a current properly received frame which immediately follows the at least one previously non-properly received frame, a second set of decoded audio parameters (ai2) from at least the set of encoded audio parameters, the second set of decoded audio parameters (ai2) being different from the first set of decoded audio parameters (ai1);
a synthesizing unit (50), to output, or derive, an output or derived version of the synthesized audio signal in such way that, if the current properly received frame immediately follows the at least one previously non-properly received frame, a selection (60) is made between:
a first version (s1) of the synthesized audio signal, synthesized from at least the first set of decoded audio parameters; and
a second version (s2) of the synthesized audio signal, synthesized from the second set of decoded audio parameters.
Description
- The present examples refer to an audio decoder. In particular, some examples refer to frame recovery after frame loss using LSF stabilization. Techniques for energy compensation are also disclosed.
- Frame Recovery is a strategy used in Audio Codecs to offer a smooth transition between a concealed frame and a received frame. If a frame gets lost, packet loss concealment (PLC) is applied, where data from the last good frame is used to render an artificial frame. If a received frame relies on information of the last frame, which is a concealed frame, then there might be a mismatch between concealed and received CELP parameters, like the Linear Prediction Coefficients (LPC) or the adaptive excitation. So, there is need to improve this transition between a lost frame and a received frame. One option is to transfer side information from encoder to decoder like energy or phase information for a guided recovery. Another option is blind recovery. In case of blind recovery, the implantation is only done on decoder side.
- Modern speech codecs like EVS [1] or the open-source codec iLBC [2] using linear prediction to get a short-term representation of the speech signal. The Linear Prediciton Coefficients ai with i = 1, ..., M, where M is the filter order, are obtained by performing algorithms like the Levinson-Durbin Recursion. The Linear Prediction filter is given by the formula:
- The filter is used to calculate the residual of the signal, but also to synthesize a signal. The residual or the excitation is obtained by:
where s(n) is the speech signal, which is usually pre-emphasized, L ist the time length of the current frame and x(n) is the residual. The speech signal can be re-obtained by the synthesis filter: - E.g. in analysis-by-synthesis codecs usually a weighted speech filter is used to obtain a weighted speech signal, to find the optimum pitch and innovation parameters, by minimizing the squared error between this weighted speech signal and the input signal. Further steps are the open-loop pitch search and closed loop pitch search to find the optimum pitch lag and pitch gain, also known as the Long-Term Prediction or adaptive codebook and the innovation excitation search. The optimum parameters on encoder side are quantized and the codebook indices are sent to the decoder.
- If a frame gets lost, the transmitted parameters are not received by the decoder. The PLC on decoder side is enabled to conceal a frame based on the data from the last good frame. If a new good frame is decoded after frame loss, decoder and encoder are not synchronized anymore and noticeable artifacts can occur in the speech.
- In case of a guided concealment and recover speech parameters, for example energy information, phase or periodicity information are transmitted from encoder to decoder to improve the robustness of the frame loss concealment. [3]. The phase control is particularly important for the recovery. After one erased frame or a block of erased frames, the decoder memory is desynchronized with the encoder memory. To resynchronize the decoder a rough position of the first glottal pulse is send. The precision the encoder the position of the pulse depends on the closed-loop pitch value for the first subframe pitch lag T0. Then the sample with the maximum amplitude is searched to find the position of the first glottal pulse, where T0 is also the length of the search interval. The highest precision for the glottal pulse is achieved with a low pass filtered residual signal. In case of a high bandwidth, a periodicity parameter or voicing parameter can be transmitted. The voicing is estimated based on the normalized correlation of the signal. It can be encoded precisely using 4 bits. The voicing information is needed for highly voiced frames and better voicing resolution. The disadvantage of sending information from encoder to decoder is that bits need to be reserved (this is unwanted, and it will be shown that in the examples proposed here the decoder will not use this information, so that the encoder skips it). The most important part of the recovery is the energy control of the signal due to the strong prediction used in modern speech codecs. The energy needs to control in that way, that the energy at the beginning of the first good, received frame after frame erasure matches the energy at the end of the concealed frame. Also, the signal is scaled to prevent a strong energy increase in the signal.
- In EVS a scaling gain is applied to the decoded speech signal [3]. The scaling is done in the excitation domain to serve the long-term prediction memory for the following frame. The synthesis is done again to achieve a smooth transition from concealed frame to received frame. The excitation signal is scaled as follows:
where n is a sample, x(n) is the excitation and xs (n) is the scaled excitation. L is the frame time length (in number of samples) and g(n) is the gain applied to the samples of the excitation (it is anticipated that some optional examples will be shown in which a particular g(n) is generated, different from the prior a). The gain g(n) in the prior art starts from an initial gain g0 and converts recursively to g1 (presenting an exponential-like behavior): with g(-1) = g 0 and the attenuation factor fAGC , which was determined experimentally. The gains go and g1 are defined as: where E -1 is the energy at the end subframe of the previous frame, E 0 is the energy at the beginning subframe of the current frame and E 1 is the energy at the end subframe of the current frame. Eq is the quantized transmitted energy parameter, which is signalled in the bitstream (while E -1 , E 0, and E 1 are simply calculated by the decoder). If Eq cannot be transmitted it is set to E 1 and therefore g1 = 1. Cases, where the last good frame before erasure and the first good frame after erasure is classified as a VOICED, VOICED_TRANSITION or ONSET [4], Eq is calculated using: where E LP0 is the linear prediction filter gain of the last good frame before erasure and E LP1 is the linear prediction filter gain of the current frame after erasure. This is done to compensate a possible energy mismatch between excitation signal energy and the LP filter gain. Further, there are more exceptional cases, where in case of using an artificial onset frame, go is set to 0.5 * g1. If the last good frame is classified as VOICED, VOICED_TRANSITION or ONSET and the first good frame after erasure is UNVOICED, go is set to g 1. Finally, the synthesis is redone with the synthesis filter: - If Eq is not transmitted, the recovery is done only on decoder side and no side information needs to be transmitted.
- In the internet low bitrate codec (iLBC) an overlap-add procedure is performed to merge the previous excitation smoothly into the current block's excitation to avoid discontinuity at the frame border. Therefore, a correlation between the excitation of the received frame and the excitation of the concealed frame is done to find the best phase match. First, a closer correlation estimation of the input signal near to the estimated lag is done. If the new correlation is higher than the correlation at the position of the old lag, the new lag is used for the phase match. Then, a portion of the excitation of the previous concealed frame and a portion of excitation of the received frame are copied to a new buffer with a specific length of Ltrans. This buffer is specified by:
- xold (n) is the excitation of the concealed frame, Lold is the length of the concealed excitation, xin (n) is the excitation of the received frame and T 0 is the estimated lag. T 0 is limited by Ltrans .
- Energy limitation is applied to the signal xtrans (n) in case that the energy of the signal xtrans (n) is higher than xold (n). The signal xold (n) and xtrans (n) are than merged doing an overlap-add operation:
- Drawback of the prior art include the missing or insufficient consideration of a mismatch between LPC coefficients and the excitation in case of a frame loss. In CELP, the excitation is the residual of the LPC analysis filter. The LPC represents short-term characteristic of the signal and the LPC coefficients are interpolated for each subframe through their Linear Spectral Pairs (LSP) from the previous frame and the current frame. To send the LPC to the decoder, they (the LPC coefficients) are converted into Linear Spectral Frequencies (LSFs) before being quantized. The excitation is determined based on the quantized and interpolated set of LPCs and the codebook indexes are determined. On decoder side the LSF are de-quantized.
- Recently, deep neural network models like WaveNet are used as speech synthesizer. The advantage is compared to classic speech coding algorithm is the the improvement of speech quality without increasing the bitrate. In [5] generative adversarial networks (GANs) are used to create a speech signal in combination with the LPC to calculate a glottal excitation from a speech input which is fed to the neural network. The example does not provide concealment of frame losses.
- Independent definitions of the examples are found in the independent claims.
- According to an aspect, there is provided an audio decoder for synthesizing an audio signal from a bitstream which represents the audio signal, the audio decoder including:
- a bitstream receiver, to receive the bitstream representative of the audio signal, the bitstream having, encoded therein, a set of encoded audio parameters for each frame of the audio signal,
- a first audio parameter decoding unit, to decode, for a current properly received frame, a first set of decoded audio parameters from at least the set of encoded audio parameters,
- a concealment unit to conceal at least one non-properly received frame based on at least one previously properly received frame, so as to generate at least one set of concealment audio parameters;
- a second audio parameter decoding unit, to decode, for a current properly received frame which immediately follows the at least one previously non-properly received frame, a second set of decoded audio parameters from at least the set of encoded audio parameters, the second set of decoded audio parameters being different from the first set of decoded audio parameters;
- a synthesizing unit, to output, or derive, an output or derived version of the synthesized audio signal in such way that, if the current properly received frame immediately follows the at least one previously non-properly received frame, a selection is made between:
- a first version of the synthesized audio signal, synthesized from at least the first set of decoded audio parameters; and
- a second version of the synthesized audio signal, synthesized from the second set of decoded audio parameters.
- According to an aspect, there is provided a method for synthesizing an audio signal from a bitstream which represents the audio signal, the method including:
- receiving the bitstream representative of the audio signal, the bitstream having, encoded therein, a set of encoded audio parameters for each frame of the audio signal,
- decoding, for a current properly received frame, a first set of decoded audio parameters from at least the set of encoded audio parameters,
- concealing at least one previously non-properly received frame based on at least one previously properly received frame, so as to generate at least one set of concealment audio parameters, and to synthesize a concealed frame from the at least one set of concealment audio parameters;
- decoding, for a current properly received frame which immediately follows the at least one previously non-properly received frame, a second set of decoded audio parameters from at least the set of encoded audio parameters, the second set of audio parameters being different from the first set of audio parameters,
- outputting or deriving an output or derived version of the synthesized audio signal in such way that, if the current properly received frame immediately follows the at least one previously non-properly received frame, a selection is made between:
- a first version of the synthesized audio signal, synthesized from at least the set of encoded audio parameters
- a second version of the synthesized audio signal, synthesized from the second set of decoded audio parameters.
-
-
Fig. 1 shows an example according to the present disclosure. -
Figs. 2-5 show behaviors in the prior art. -
Figs. 6-12 show examples according to the present disclosure. -
Fig. 5 can be taken into account, when reading below, for distinguishing between the different frames and subframes.Fig. 5 distinguishes between different scenarios and shows formulas which are discussed here-below. - First of all, it is here indicated that the current frame is normally referred to as a q-th frame of the sequence of frames. If the (q-1)-th frame has been a non-properly decoded frame (and therefore has been concealed), then the q-th frame is a recovery frame (provided that the q-th frame is also properly decoded). Otherwise, if the (q-1)-th frame has been a properly decoded frame (and therefore has not been concealed), and the q-th frame is a properly-received frame as well, then the q-th frame is a non-recovery frame. Finally, if the q-th frame is non-properly decoded, then it is a lost frame, and is substituted by a concealed frame synthesized by the concealment unit 40.
- Since many of the pedices (subscripts) of the signs would in principle carry "q" for the current frame and "q-1" for the immediately preceding frame, these pedices will not be used to avoid reading burdens.
- Further, it will be understood that the q-th frame is in general substituted into N subframes. For each frame, the subframes are in general indexed with "k" (with 1≤k≤N). Even in that case, the use of complicated pedices like "k,q" will be preferably avoided.
- When referring to the LPC parameters ai1, ai2 the index i will refer to the index i of the filter
, while the indices 1 and 2 will refer to the first and second synthesis (see below). - Further, in some cases the pedex "end" will be used. That pedex "end" is to be understood as globally valid for a whole frame, and not for a particular subframe.
- It will be noted that LSF and LSP parameters are often indicated with pedex "end" (e.g. LSFend, LSPend, LSP'end) despite being often intended as global values for a particular frame. The reason is that it is intended that, at the encoder side, these values are only calculated on a particular, final subframe (i.e. the N-th subframe of the frame) and, rigorously speaking, they should be intended as parameters for that particular N-th, final subframe. Notwithstanding, it is hypothesized that those values are globally valid for the entire frame, at least according to a first approximation. It will be shown that, in some cases (e.g. for calculating the LSP parameter LSPk for each k-th subframe between the 1st (k=1) and the penultimate (k=N) subframe), an interpolation (or another weighting function) will be carried out, e.g. through interpolation weights wk with wk increasing (e.g. linearly) for k increasing (e.g. it may be w1 being a value between 0 and e.g. 0.3 or 0.2, and wN being a value between 0.7 or 0.8 and 1).
-
Fig. 6 shows an example of an audio decoder 100 to synthesize an audio signal from a bitstream which represents the audio signal. The audio decoder 100 may include a bitstream receiver 5, which receives the bitstream (the bitstream receiver 5 may include a bitstream reader, not shown, which e.g. reads the bitstream e.g. from remote or from a storage unit). The bitstream may have, encoded therein, a set of encoded audio parameters (e.g., encoded versions of residual audio parameters, here indicated with ΔLSFq for each q-th frame) for each frame of the audio signal. For example,Fig. 6 shows that the bitstream receiver 5 includes a dequantization block 602 (e.g. inputted with the bitstream which has been read by the non-shown bitstream reader). The dequantization block 602 may provide the encoded audio parameters as residual values (ΔLSFq) of linear spectral frequencies (LSFs). The audio decoder 100 may comprise a first audio parameter decoding unit 10 which can be understood substantially as operating as in the prior art. The first audio parameter decoding unit 10 may decode, for a current properly received frame, a first set of decoded audio parameters (ai1) from at least the set of the encoded audio parameters (ΔLSFq). As shown byFig. 6 , the first audio parameter decoding unit 10 may include a block 604 for computation of first set of parameters LSPs (linear spectral pairs), which provides in output linear spectral pairs (LSPend). The first audio parameter decoding unit 10 may include a block 606 for computation of a first set of LPC parameters which may be LPC coefficients indicated here as ai1. - The first audio parameter decoding unit 10 may normally operate for decoding properly-received frames (good frames). The audio decoder 100 may include (not shown in
Fig. 6 , but shown for example inFig. 1 ), a concealment unit 40. The concealment unit 40 may conceal at least one non-properly received frame (e.g. a (q-1)-th frame) based on at least one previously properly received frame (e.g. a (q-2)-th frame). The concealment unit 40 may therefore generate at least one set of concealment audio parameters (here indicated as LSPPLC). The concealment unit 40 may synthesize a concealed frame from the at least one concealment audio parameters. - Here, the way how the concealment unit 40 operates is left general: we are not really interested in how the concealment unit 40 performs the concealment of the (q-1)-th non-properly received frame, but, instead, in how the LPC parameters (ai1) are obtained for the q-th properly received frame which immediately follows the (q-1)-th concealed frame. The properly received q-th frame which immediately follows the (q-1)-th concealed frame (substituting the non-properly received frame) is here called "recovery frame". It is mostly intended to discuss the way of how to obtain the synthesis signal (and also the parameters for obtaining the synthesis signal) for the recovery frame.
- In particular, the first audio parameter decoding unit 10 (in case of the current q-th frame being a recovery frame) may obtain the LPC parameters an by taking into account the LPC parameters used by the concealment unit 40 for performing the concealment for the preceding, (q-1)-th non-properly decoded frame. Block 604 may be inputted by a product of a predefined prediction factor m (e.g. a value between 0 and 1/3 (e.g. 0.333333) or a value between 1/4 and 1/2, or another natural number larger than 1/10 and less than 1) with concealment audio parameters ΔLSFq-1 (e.g. LSF parameters, which may be differential parameters). Also, block 606 (downstream to block 604) may be also inputted with LSPPLC (concealment parameters, which may be LSP parameters, taken from the concealed frame). In the case that the current q-th frame is a non-recovery frame (i.e. a properly received frame which does not follow a non-properly received frame), then the blocks 604 and 606 operate normally by taking into account parameters from the immediately preceding (properly decoded) (q-1)-th frame (those parameters being also indicated with ΔLSFq-1). In this case, the LPC parameters ai1 are obtained as usual, by taking into account the parameters of the immediately preceding, properly decoded frame.
- The audio decoder 100 may include a synthesizing unit 50. The synthesizing unit 50 (in particular in block 608 for a first synthesis) may generate a first synthesis signal s1(n) from the LPC parameters ai1. In the prior art, the synthesizing operation would be finished (apart from possible other operations such as post-processing and/or scaling by a gain g(n)) and the synthesized output audio signal s1(n) would be provided as the output of the synthesizer unit 50.
- The synthesizing unit 50 (in particular in block 608 for computation of first synthesis, inputted from by the excitation x(n) obtained from the previous frame and the LPC parameters ai1 from the first audio parameter decoding unit 10) may therefore provide a first synthesis signal s1(n).
- However, the audio decoder 100 also comprises a second audio parameter decoding unit 20. The second audio parameter decoding unit 20 may decode, for the currently q-th properly received frame which immediately follows the at least one non-properly received frame, a second set of audio parameters (e.g. a second set of LCP parameters) ai2. In particular, the second audio parameter decoding unit 20 may include a block 612, for computation of second set of a LSP parameter (LSP'end) for each frame, block 612 not being inputted with the audio parameters of the immediately preceding audio frame (or, in some examples, which is inputted by a number of concealment parameters less than the concealment parameters used inputted into the first audio parameter decoding unit 10). Therefore, the linear spectral pair LSP'end may be provided to a block 614 for computation of second set of LPC parameters (which are indicated with ai2). Block 614 is not inputted with concealment parameters LSPPLC (or any other parameter taken from the concealment unit 40) but, rather, with LSPmean (which may be obtained from a table). The second audio parameter decoding unit 20, (in particular block 614) may output or derive the second set of decoded audio parameters (ai2) which is different (and in particular ai2 is different from the first set of decoded audio parameters ai1, and in particular is not controlled by the concealment parameters of the immediately preceding concealed frame). The second audio parameter decoding unit 20 is deactivated in the case of the q-th current frame not being a recovery frame: in case of the current q-th frame being a non-properly received frame, then the synthesis is performed by the concealment unit 40, while in the case of the q-th current frame being a non-recovery frame (i.e. a properly decoded frame which immediately follows another, (q-1)-th properly decoded frame), then the synthesis is performed by the first audio parameter decoding unit 10 by keeping into account the parameters of the (q-1)-th immediately preceding audio frame (which is properly decoded), and the second audio parameter decoding unit 20 is deactivated.
- The synthesizing unit 50 (in particular in block 616 for computation of second synthesis, inputted by the excitation x(n), obtained from the previous frame, and the LPC parameters ai2 from block 614) may therefore provide a second synthesis signal s2(n). The synthesis signal s2(n) will compete with the synthesis signal s1(n) for being selected as the output or derived signal s(n) for the current frame.
- Block 604 of the first audio parameter decoding unit 10 and block 612 of the second audio parameter decoding unit 20, may each provide, in examples, a parameter which is valid for the whole current q-th frame, while block 606 of the first audio parameter decoding unit 10 and block 614 of the second audio parameter decoding unit 20 may each provide a respective LPC set of parameters (ai1, ai2) for each subframe (indeed, we use the wording "set of parameters" both because the parameters vary with the particular subframe, and also because they vary with the index i).
- Summarizing:
- 1) If the current q-th frame is a non-properly received frame, then the synthesis is carried out through the packet concealment unit 40.
- 2) If the current q-th frame is a non-recovery frame (i.e. a properly-received frame which immediately follows a properly-received frame, i.e. both the q-th frame and the (q-1)-th frame are properly-received frames) the synthesis (to derive the output or derived synthesis signal) is carried out through the following path:
- a. Block 604, computing the LSPend from parameters ΔLSFq-1 taken from the (q-1)-th immediately preceding frame.
- b. Block 606, computing the LPC parameters ai1, from the LSPend and the LSPend-1 (which are the non-concealed LSP from the previous properly-received frame)
- c. Blocks 716 and 718 (blocks 60, 702, 706, 708, 360, 712 being bypassed).
- 3) If the current q-th frame is a recovery frame (i.e. a properly-received frame which immediately follows a concealed frame, i.e. the q-th frame is properly received but the (q-1)-th frame is non-properly received and, therefore, has been concealed) the synthesis is carried out through the following paths:
- a. A first path, with:
- i. Block 604, computing LSPend taking into account parameters (e.g. prediction residual parameters) (or more in general concealment parameters) ΔLSFq-1 taken from the immediately preceding frame (concealed frame) and ΔLSFq from the bitstream for the current q-th frame.
- ii. Block 606, computing the first set of LPC parameters ai1 from LSPend and concealment parameters LSPPLC.
- iii. Block 608, processing the first synthesis signal s1(n) from LPC parameters ai1 and the excitation x(n).
- b. A second path, with
- i. Block 612, computing LSP'end taking into account parameters (e.g. prediction residual parameters) ΔLSFq taken from the bitstream for the current q-th frame, but not from concealment parameters (e.g. residual parameters) of the concealed (q-1)-th frame (or in some alternative examples, by taking into account less concealment parameters than the concealment parameters taken into account by the block 606).
- ii. Block 614, computing the second set of LPC parameters ai1, from LSP'end and LSPmean.
- iii. Block 616, processing the second synthesis signal s2(n) from the second set of LPC parameters ai2 and the excitation x(n).
- c. A selection (60) including:
- i. A first energy computation block (or more in general first energy-related measurement block, such as a first envelope measurement block) 610 on the first synthesis signal s1(n) (or only an initial subframe or a group of initial subframes of the first synthesis signal s1(n))
- ii. A second energy computation block (or more in general second energy-related measurement block, such as a second envelope measurement block) 618 on the second synthesis signal s2(n) (or only an initial subframe or a group of initial subframes of the second synthesis signal s2(n))
- iii. A decision (e.g. based on energy-related measurements such as those performed at first/second blocks 610/618 and/or on energy related measurements on the concealed frame or on the final frame or final frames of the concealed frame) on whether the first synthesis signal s1(n) or the second synthesis signal s2(n) is to become the synthesis signal s(n) for the current frame.
- iv. Further block 623 with synthesis smoothing/scaling (see also below).
- a. A first path, with:
-
Fig. 6 also shows an excitation decoding block 622 which outputs an excitation x(n) (or z(n) if expressed as z-transform). - In order to reduce the computational burden of the two syntheses, it is possible that the following operational steps are performed:
- 1) First operational step:
- a. Block 608 performs a partial synthesis of the first synthesis signal s1(n), by only synthesizing the initial subframe or a group of initial subframes of the first synthesis signal s1(n) (hence skipping a last subframe or a group of last subframes of the first synthesis signal s1(n)) and
- b. Block 616 performs a partial synthesis of the second synthesis signal s2(n), by only synthesizing the initial subframe or a group of initial subframes of the second synthesis signal s2(n) (hence skipping a last subframe or a group of last subframes of the second synthesis signal s2(n)).
- 2) Second operational step (which could precede or follow or be simultaneous with the first operational step):
- a. Block 610 performs an energy-related measurement on the partially synthesized version of the first synthesis signal s1(n), e.g. by measuring the energy and/or the envelope (or another energy-related measurement) only in the synthesized initial subframe or synthesized group of initial subframes of the first synthesis signal s1(n) (hence skipping measurements on the last subframe or a group of last subframes of the first synthesis signal si(n)) and
- b. Block 618 performs an energy-related measurement on the partially synthesized version of the second synthesis signal s2(n), e.g. by measuring the energy and/or the envelope (or another energy-related measurement) only in the synthesized initial subframe or synthesized group of initial subframes of the second synthesis signal s2(n) (hence skipping measurements on the last subframe or a group of last subframes of the second synthesis signal s2(n)).
- 3) Third operational step (which follows the first operational step and the second operational step):
- a. Decision block 620 performs the decision based on the energy-related measurements (e.g. energy measurements and/or envelope measurements) obtained from blocks 610 and 618 only on the partially synthesized version of the first synthesis signal s1(n) and the partially synthesized version of the second synthesis signal s2(n) (hence without considering the evolution of the synthesis signals s1(n) and s2(n) in the last subframe or a group of last subframes of the first synthesis signal).
- 4) Fourth operational step (which follows the third operational step):
- a. Once the selector 60 has decided, from the partially synthesized version of the first synthesis signal s1(n) and the partially synthesized version of the second synthesis signal s2(n), which (between s1(n) and s2(n)) is the synthesis signal to be selected as output or derived synthesis signal s(n), then the block 608 or 616 (in accordance to the selected synthesis signal) is reactivated to complete the synthesis of the selected synthesis signal through a second, partial synthesis of the selected signal (the non-selected synthesis signal being therefore disregarded without completing its synthesis).
- It is noted that the first operational step may also interest the first audio parameter decoding unit 10 (e.g. in at least one of blocks 604 and 606) and/or the second audio parameter decoding unit 20 (e.g. in at least one of blocks 612 and 614), because, in order to further save computational power, it is possible to only compute those audio parameters which are in the first subframe(s) of the first and/or second synthesis signals, without computing those audio parameters in the final subframe of the first and/or second synthesis signals (in practice, initially only decoding a first subset, which is a proper subset, of the first set of decoded audio parameters, and only decoding a second subset, which is a proper subset, of the second set of decoded audio parameters): as soon as one of the two synthesis versions of the synthesis signal is selected is selected, then only the first audio parameter decoding unit 10 (in case the first synthesis is selected) or the second audio parameter decoding unit 20 (in case the second synthesis is selected) will perform the decoding of the audio parameters of the remaining, final subframe(s) of the current q-th frame and/or of calculating parameters of the last portions of the selected signal for the current q-th frame (e.g. for calculating the energy compensation gain g(n), see below) (in practice, the non-void subset of the remaining decoded audio parameters which were initially not decoded is only subsequently decoded, and the synthesis of the selected version of the synthesis signal is processed using the remaining decoded audio parameters of the selected synthesis, disregarding the remaining decoded audio parameters of the non-selected synthesis).
- However, these operational steps are not strictly necessary, and in some (less preferred) examples the complete two synthesis signals s1(n) and s2(n) are synthesized before the selection (60). In another example, the complete first and second sets of decoded audio parameters are decoded, but only the first and second partial syntheses are initially performed.
- The audio decoder 100 may include (e.g., within the synthesizing unit 50), a selector 60 which selects between the first synthesis signal s1(n) (or the first set of parameters an) and a synthesis signal s2(n) (or the second set of parameters ai2). In particular, the selector 60 may include a block 610 for energy computation and/or envelope evaluation for the first synthesis, which may provide a first energy-related measurement. The selector 60 may include block 618 for energy-related measurement computation (e.g. energy computation and/or envelope computation) for the second synthesis which provides an energy-related measurement (e.g. energy computation and/or envelope computation) on the second synthesis signal s2(n). A block 620 for decision of LPC set and synthesis may decide which synthesis signal, among s1(n) and s2(n) to be used as a synthesis signal s(n) (output version or derived version of the synthesis signal). A block 623 synthesis smoothing/scaling may also be used, inputted with the selected signal s(n), to thereby process the output version or derived version of the synthesis signal and render it.
-
Fig. 1 also shows elements of the audio decoder 100 in terms of block scheme. In particular, a decision 41 is made between determining whether the current frame is a non-properly decoded frame or a properly decoded frame. If the current frame is a non-properly decoded frame (e.g. by a determination based on cyclical recurrent calculations, such as cyclic redundancy check, CRC, or the like) then the concealment unit 40 may be activated. Otherwise (if the frame is determined as valid), in block 42, it is evaluated whether the previous frame was lost. If the previous frame was lost, then the current frame is a recovery frame. Therefore, at block 43, both the first audio parameter decoding unit 10 and the second audio parameter decoding unit 20 are activated. Subsequently (block 44), the synthesizing unit 50 is invoked. In case at block 42 it is determined that the previous frame was not lost, then the first audio parameter decoding unit 10 only is activated and the synthesizing unit is activated as well in block 44, and there is an updating 45 of the memory for the next frame. This happens both in the case where the present frame is concealed frame (and the concealment unit is therefore activated), or whether the present frame is a properly decoded, recovery frame (and therefore blocks 43 and 44 are activated, therefore using both the first and second audio parameter decoding units (10, 20), and in the case that the current frame is a properly received frame which follows a properly received frame (i.e., the current frame is a non-recovery frame) and therefore only the decoding block 44 is activated but not block 43. After block 45, a new instance of block 41 is activated. - It will be shown subsequently that the choice between the first synthesis signal s1(n) and the second synthesis signal s2(n) may be made based on the comparison between the energies of the signals, and in particular on the behavior of the envelope of the signals.
-
Fig. 7 shows the example 100 with other blocks, which may be optional in some examples.Fig. 7 shows, in particular, the block 620 for decision of LPC set and synthesis. Block 620 may provide both LPC parameters ai (chosen between the LPC parameters ai1 of the first synthesis and ai2 of the second synthesis) and provide those parameters to a zero input response block 702. (This can be a outringing filter. The last e.g. 16 samples (or another amount, e.g. less than 32 samples) of the previous (concealed) frame are inputted and also a zero input excitation. Then an outringing synthesis is performed). The zero input response block 702 may therefore provide a zero input response (ZIR) indicated with sZIR(n). Block 620 may also output or derive the chosen synthesis signal s(n) between s1(n) and s2(n) to the subtractor block 704. The synthesis signal s(n) may therefore be subtracted with the zero impulse response signal sZIR(n) in the subtractor block 704. The energy computation block 706 may compute an energy of the synthesis signal s(n). A gain computation block 708 may be used to obtain a first gain g1 and a second gain g2. Values of the first gain g1 and of the second gain g2 may be provided to a scaler 360 to provide a gain g(n), which may be defined sample-by-sample, and may, for each sample, take a value between the first gain g1 and the second gain g2. It is to be noted that the scaler 360 may be conditioned by the concealment unit 40 and, in particular, by the energy EPLC () of the concealed frame (it will be shown, in particular, that the scaler 360 may be conditioned by the values of the energy in a final portions (e.g. final half portion) of the concealed frame). The output of the scaler 360 may be added with the zero input response sZIR(n) at adder 710. Therefore, a value ss(n) is to be provided to a residual computation block 712. A post-processing block 716 may be used from the output ss(n) of the adder 710. A block 718 of updating excitation for the next frame permits to obtain the residual computation from the output xs(n). As can be seen fromFig. 7 , a block 620 for decision is also used to provide a coefficient α value (proportionality coefficient) to the gain computation block 708 and the parameters to the residual computation block 712. The audio signal ss(n) may therefore be post-processed and used and rendered as audio signal. - First, a description of the behavior of the decoder 100 in the first audio parameter decoding unit 10 (in particular in blocks 604 and 606 of
Fig. 6 ). This section discloses technique which, as such, can be already present in the prior art. Here, it is assumed that both the current frame and the immediately preceding frame are properly decoded (no concealment, no recovery). It is noted that from the bitstream, only parameters like ΔLSF q-1 and ΔLSFq (residual) are received. No voicing information (e.g., "voiced" vs "unvoiced") are received from the bitstream. It is noted that an LSF value (and also the residual ΔLSF q-1) is common for the whole frame, while LSP values are different for different subframe of a frame. - First, a prediction of the LSF for the current frame may be determined (at block 604) by:
where m is a prediction factor (e.g. a value between 0 and 1/3 (e.g. 0.333333) or a value between 1/4 and 1/2, or another natural number larger than 1/10 and less than 1) and ΔLSF q-1 is the LSF prediction residual decoded for the previous (q-1)-th frame. To get the decoded end LSF, the prediction LSF pred and the received ΔLSFq (possibly in dequantized form) are added: - The term m * ΔLSF q-1 + ΔLSFq = (1 + m z -1) * ΔLSF corresponds to a moving average.
- The so-obtained LSF coefficient LSF end is converted to the LSP coefficient (the following formula is valid for the whole current frame, without any distinction between the subframes):
where fs is the sampling frequency. - To obtain a set of LSP coefficients (precursors of the set of decoded audio parameters) for each subframe of the current frame, the LSP coefficients are interpolated from the LSP coefficients of the previous frame to the LSP coefficients of the current received frame for each k-th subframe out of the N subframes of the current frame:
where wk are interpolation weights (e.g., increasing linearly with the increase of k, e.g. it may be w1 being a value between 0 and e.g. 0.3 or 0.2, and wN being a value between 0.7 or 0.8 and 1) and N is the number of subframes for the frame. LSP end-1 (which a priori should be called LSP q-1,N since it is the LSP value of the last, N-th subframe of the immediately preceding, (q-1)-th concealed frame) are the old end LSP coefficients of the end subframe (i.e. of the N-th subframe) of the previous frame, while the LSPend is the LSP coefficient of the end subframe (i.e. of the N-th subframe) of the current frame (as explained above, used globally for the whole current q-th frame). (It is reminded that LSFend (which a priority should be written LSFq,N ) can be used to approximate the whole frame, even if the encoder had only calculated it on the end of the frame. We assume that the LPC does not change to much about frame, because it is a short-term representation.). The LSPk (which a priority should be written LSPq,k ) are converted to the LPCk (which are notwithstanding written a i1 ). This is a so-called Conversion of LSP parameters to LP coefficients (see ETSI TS 126 445 V16.2.0 (2022-03), currently available at the web page https://www.etsi.org/deliver/etsi_ts/126400_126499/126445/12.06.00_60/ts_126445v120600p.pdf). The signal is than synthesized subframe-wise using the LPC synthesis filter with the excitation as input. See, in particular,Fig. 5 . - This section explains techniques which, as such, are also in the prior art. In the prior art, however, there is not the backup provided by the second audio parameter decoding unit 20 (explained below). It is reminded that, while the ΔLSF q-1 (residual) is received from the bitstream, no voicing information (e.g., "voiced" vs "unvoiced") are received from the bitstream. Further, an LSF value (and also the residual ΔLSF q-1) is common for the whole frame, while LSP values are different for different subframes of a frame.
- It is here irrelevant which technique is used for concealment (packet lost concealment, PLC). What is important now is how to get the parameters of the recovery frame, i.e. the first properly-decoded frame after a concealed frame (non-properly decoded frame).
- The LSP coefficients of the recovery frame are obtained by:
where LSPPLC (which could be written LSPend-1) are the LSP coefficients which are used for the concealed frame (i.e. LSPPLC have been inferred by taking into account at least the previously correctly received frame, e.g. the properly-received (q-2)-th frame before the (q-1)-th concealed frame). The LSP end are calculated by: with fs being the sampling frequency. - The LSFend is calculated by:
where ΔLSFPLC , which can also be written as ΔLSF q-1, is a concealed version of the residual in the previous, (q-1)-th, concealed frame. It is not of interest how that delta or residual is concealed, but it is important to know that the residual of the previous concealed frame is used for the recovery frame. - There can be a great big difference between s1(n) and s2(n) in some cases, and in particular in the cases in which ΔLSFPLC differs from the lost parameters ΔLSF q-1, and also in case of LSP PLC (generated by the concealment unit 40) of the concealed frame being different from the parameters LSP end-1 of the lost frame. If the ΔLSFPLC of the concealed frame and the transmitted ΔLSF q-1, in case the frame is not lost and also if the LSP PLC of the concealed frame (i.e., the LPC coefficients which have been inferred during concealment) and the LSP end-1, in case the frame is not lost, distinguish from each other, here can be a high deviation of the LPC filter response between erroneous signal and clean signal. A big difference in ΔLSFPLC in a previous concealed frame and ΔLSF q-1 in a previous received frame can also means a big difference in ΔLSFPLC and ΔLSFq in the recovery frame and the same for LSP PLC in the previous concealed frame and LSPend in the recovery. See for example
Figs. 2 and3 . This might result in an unstable synthesized signal in the recovery due to the mismatch of the interpolated LPC and the excitation. A strong oscillation and energy increase is visible. These fast changes of LPC might appear especially in frames, where an onset occurs, i.e. a change from unvoiced to voiced frames. In sharp contrast,Fig. 10 (obtained using the present techniques) shows that it is possible (e.g., by relying on the second audio parameter decoding unit 20) to obtain a more stable output signal (or derived signal). - This section discloses techniques which are not present in the prior art, to the best knowledge of the inventors. Here, for a recovery frame, audio parameters are generated without taking into account parameters from the (q-1)-th concealed frame (and indeed ΔLSF q-1 is not an input to bclos 612 and 614). Subsequnelty, a choice (at block 620) will be performed on whether to use the signal s1(n) (synthesized by taking into account the parameters from the (q-1)-th concealed frame) and the signal s2(n) (synthesized without taking into account the parameters from the (q-1)-th concealed frame, or taking into account less parameters from the (q-1)-th concealed frame than for s1(n)).
- The presented example proposes, inter alia, a technique introducing a second set of LPC coefficients to synthesize a second signal s2(n) in the decoder 100. An idea is to compare (e.g. at selector 620) the energy (E1k, E2k) of two synthesis signals s1(n) and s2(n) (or of initial portions of s1(n) and s2(n)):
- the first signal s1(n) based on a first set of LPC coefficients (an), as described above (e.g. obtained by taking into account parameters of the concealed frame, e.g. through LSPPLC), and
- the second signal s2(n) based on a second set of LPC (e.g. without taking into account parameters of the concealed frame, but taking into account parameters of the recovery frame) and to choose (in block 620) the signal (s1(n) or s2(n)) which, for example, provides a smoother transition in the recovery.
- The second set of LSP parameters may be obtained by using a mean LSP (LSPmean) of the recovery frame instead of LSPPLC of the previous, concealed frame.
- A motivation behind this technique is that the shape of s2(n) might be closer to the LSP in clean channel conditions. Additionally, the parameters (ai2) of s2(n) are independent from the LSP coefficients of the concealed frame. To avoid instability at the end of the recovery frame which could be caused by the previous LSF coefficients, at block 604 (in the first audio parameter decoding unit 10) the prediction of LSPend preferably doesn't include the moving average (previously formulated as m * ΔLSF q-1 + ΔLSFq = (1 + m z -1) * ΔLSF) with the previous delta LSF (ΔLSF q-1) and with m being pre-defined (e.g. a value between 0 and 1/3 (e.g. 0.333333) or a value between 1/4 and 1/2, or another natural number larger than 1/10 and less than 1), so that the prediction contains only the mean LSF.
- The second set of LSF parameters, LSF' end is calculated (in block 612, in the second audio parameter decoding unit 20) by:
where LSFmean is obtained by from a table with pre-defined values and ΔLSFq are the LSF coefficients read in (or derived from) the bitstream for the recovery frame (using the formula LSP end = ). -
First Set of End LSF (block 604) Second Set of End LSF (block 612) LSFend = LSFmean + Δ LSFPLC + Δlsfq with LSFmean obtained from a table, ΔLSFPLC obtained from the concealment unit 40, and Δlsfq read in the bitstream with LSFmean obtained from a table and Δlsfq read in the bitstream - The
of the second set are converted to the (in block 612) through - To avoid more impact of the previous (q-1)-th frame, for the interpolation of the LSP for each subframe, in the second set of end LSF, the LSF q-1 are no longer used. Instead, the LSP are obtained by interpolating from the LSPmean , which are obtained by conversion from the LSFmean , to the new LSFend'.
where wk are interpolation weights (e.g., increasing linearly with the increase of k, e.g. it may be w1 being a value between 0 and e.g. 0.3 or 0.2, and wN being a value between 0.7 or 0.8 and 1) and N is the number of subframes for the frame. - So the final LSPk for each k-th subframe of each frame are:
-
-
- and LSP'k are converted into Linear Prediction Coefficients (LPC), named as a i1 and a i2. It is remembered that N is the number of subframes in the recovery frame, and k indicates the k-th subframe in the frame. It is noted that there is a LSP' end for each k-th subframe and it could therefore be written as LSP' end,k.
- In order to get the second synthesis signal si2(n), in block 614 the second set of LSP parameters LSP'k is converted (in block 614 of the second audio parameter decoding unit 20) to a second set of LPC parameters. The two synthesis signals si1(n) and si2(n) are therefore synthesized. The first synthesis signal si1(n) is synthesized (at block 608 of the synthesizing unit 50) by using the first set of LPC coefficients, and the second synthesis signal si2(n) is synthesized (at block 616 of the synthesizing unit 50) by using the second set of LPC coefficients:
where x(n) is the excitation (e.g. obtained from block 622), M is the filter order, L is the length of the frame (in samples). The excitation x(n) may be obtained from the concealment signal of the (q-1)-th concealment frame. - (As explained above, it is actually possible to initially partially synthesize the initial portions s i1_initial(n) and s i2 _initial(n) of the first and second signals synthesis signals si1(n) and si2(n), e.g. with
, with n = 0, ... , L_initial - 1, i = 1, ... , M and , with n = 0, ... , L_initial - 1, i = 1, ... , M, (with L_initial being the number of samples in the initial frames which are synthesized), and to synthesize the remaining part s i1_final(n) or s i2final (n) (according to the selection) as i), or , with n = Linitial , ... , L - 1.) - In general:
- in the first audio parameter decoding unit 10 (and in particular in block 604) a first set of linear pairs LSPend are obtained by taking into account the parameters from the concealed frame (which subsequently allow, in block 606, to generate the LPC parameters ai1 by also taking into account the concealment parameters LSPPLC and the first synthesis signal si1(n))
- in the second audio parameter decoding unit 20 (and in particular in block 612) a second set of linear pairs LSP'end are obtained by taking into account the parameters written in the bitstream for the recovery frame but not from the concealed frame (and will subsequently allow, in block 614, to generate the LPC parameters ai2 by without taking into account the parameters LSPPLC or ΔLSF q-1 from the concealed frame, and the second synthesis signal si2(n)).
- a competition between the first synthesis signal si1(n) and the second synthesis signal si2(n) for becoming the output or derived version of the synthesized signal s(n) is then based on comparing the energies of the two synthesis signals or of the first portions of the two synthesis signals (see selector 60, and in particular blocks 610, 618, and 620). The explanation is below.
- Due to the realization that the mismatching LPC coefficients in the recovery might yield in a strong energy increase, the subframe energies of both first and second synthesis signals s 1(n) and s 2(n) are compared. In general terms, the signal (among signals s 1(n) and s 2(n)) with the smoother energy envelope compared to the energy in the last subframe of the concealed signal is chosen. The subframe energy of the last subframe in the concealed signal and the energies E 1,k and E 2,k of the first three (k=1, 2, 3) subframes of the synthesis signals s 1(n) and s 2(n) are determined:
with L 1 being the number of samples for each subframe (e.g. such that 3 * L 1 = L_initial), e.g. L 1 ≅ L/N (N being the number of subframes, e.g. with N=3 we have L 1 ≅ L/3 in such a way that L 1 + L 2 + L 3 = N; the symbol "≅" being used instead of "=", for example, for keeping into account the possibility that some subframes have not exactly the same number of samples, e.g. by virtue of the total number of samples N not being divisible by 3), and k being the index of the subframe. The energies may be scaled or divided by the length of the subframe or the length on which the energy is calculated on to be comparable in case the frame, half frame or subframe length are changing from the concealed frame to the received frame. However, in some examples the formulas above may be substituted by and E 2,k = . - At least one condition (e.g. two conditions) can be taken into account. Here below, some conditions may be used alone, but in some examples more than one condition are evaluated. For this reason, it is sometimes written in terms like "if the ... condition is fulfilled, then the first/second synthesis signal is preferentially selected", where "preferentially" means that the particular synthesis signal is selected either tout-court (e.g. in the examples in which only one condition is present) or that the particular synthesis signal is selected provided that other conditions are fulfilled.
- To check whether the second synthesis signal s 2(n) is to be used, the following first conditions may be considered:
where "&&" means logical operator "AND". E 1,1 is a measurement of the energy of the first subframe of the first synthesis signal s 1(n) and E 2,1 is a measurement of the energy of the first subframe of the second synthesis signal s 2(n), E 1,2 is a measurement of the energy of the second subframe of the first synthesis signal s 1(n) and E 2,2 is a measurement of the energy of the second subframe of the second synthesis signal s 2(n). th 1 may have a value of 1.455 (or more in general between 1.4 and 1.5, or even more in general a value which is 1 or larger than 1) and th 2 may have a value of 1.5 (or more in general between 1.45 and 1.55, or even more in general a value which is 1 or larger than 1), and it may be preferably th 1 < th 2, and it may be th 1 > 0 and th 2 > 0. The preferred values were determined empirically. If the first condition is fulfilled, then the second synthesis signal s 2(n) (and the second set of LPC) is selected to be used, subjected to the second and/or third conditions, in examples. - In practice, the first condition may be generalized as: if the energy measurement E 1,1 of the first subframe of the first synthesis signal s 1(n) is larger than the energy measurement E 2,1 of the first subframe of the second synthesis signal s 2(n) by a predefined amount (e.g. 45% in the case of th 1 = 1.45), and if the energy measurement E 1,2 of the second subframe of the first synthesis signal s 1(n) is larger than the energy measurement E 2,2 of the second subframe of the second synthesis signal s 2(n) by a predefined amount (e.g. 50% in the case of th 2 = 1.5) then the second synthesis signal s 2(n) is preferentially selected (subjected to the other conditions, if present). Otherwise, if the energy of the first subframe of the first synthesis signal is not larger than the energy of the first subframe of the second synthesis signal by the predefined amount (e.g. 45% in the case of th 1 = 1.45), or if the energy E 1,2 of the second subframe of the first synthesis signal s 1(n) is not larger than the energy of the second subframe E 2,2 of the second synthesis signal s 2(n) by the predefined amount (e.g. 50% in the case of th 2 = 1.5) then the first synthesis signal s 1(n) is preferentially selected (in some examples, even without checking other conditions, even if present).
- The first condition may be generalized even more as: if, along a number NMAX≥2 (with NMAX<N or NMAX≤N) of the first consecutive subframes, the energy of the first synthesis signal s 1(n) evolves coherently over (e.g. by at least a threshold e.g. of at least 40%) the energy of the first synthesis signal, then the second synthesis signal is preferentially selected (subjected to the other conditions, if present). Otherwise, if in at least one subframe of the first NMAX≥2 (with NMAX<N or NMAX≤N) consecutive subframes the energy of the first synthesis signal does not evolve coherently over the energy of the first synthesis signal (e.g. if in at least one of the first NMAX consecutive subframes the energy of the first synthesis signal is not larger than the energy of the first synthesis signal in the corresponding subframe) then the first synthesis signal s 1(n) is chosen (in some examples, even without checking other conditions, even if present).
- To avoid an energy decrease by the second set, the energy of the first two subframes (E2,1 and E2,2) of the second synthesis signal may be compared to the energy (EPLC
sub ) of the last subframe in the concealed frame, which is derived in same way like the subframe energy of the received frame: where "&&" means the logical operator "AND". EPLCsub might be calculated the same way like E 1,k and E 2,k , (but in some examples it could be EPLCsub = ) on the last subframe, where k is the index of the last subframe of the concealed frame and sPLC [n] being the concealed signal. thPLC may have a value of 1.5 (or more in general between 1.45 and 1.55, or even more in general a value which is 1 or larger than 1; in some examples, thPLC = th 2, and/or thPLC > th 2) which was conducted experimentally. Each ratio and has to be higher than the threshold thPLC to fulfil the second condition. If the second condition is fulfilled, the second synthesis signal is preferentially chosen (in some examples, subjected to the fulfilment of other conditions, e.g. the first condition and/or the third condition). - According to the second condition, if at least one of both
and is not verified (e.g. if and/or ), the first synthesis signal s 1(n) is selected. - Further, the subframe energy (E1,1 and E1,2) of the first and second subframe of the first synthesis signal compared to the energy (EPLC ) of the last subframe of the concealed signal has to be higher than a certain threshold (e.g., the same thPLC ):
where "&&" means the logical operator "AND". This condition also prevents that the second synthesis signal s 2(n) is used when the energy difference of the first set compared to the concealed signal (or at least to the last subframe of the concealed signal) is small. - Therefore, the second condition may be generalized in that: if the energies (E1,1, E1,2, E2,1, E2,2) of both the first subframe and the second subframe of both the first synthesis signal s 1(n) and the second synthesis signal s 2(n) are larger than the energy of the last subframe of the concealment signal by a predefined amount (e.g. 50%), then the second synthesis signal s 2(n) is preferentially chosen (in some examples, subjected to the fulfilment of other conditions, e.g. the first condition and/or the third condition). Otherwise, if a least one of the energies (E1,1, E1,2, E2,1, E2,2) of the first subframe and the second subframe of at least one of the first synthesis signal s 1(n) and the second synthesis signal s 2(n) is not larger than the energy of the last subframe of the concealment signal by the predefined amount, then the first synthesis signal s 1(n) is preferentially selected (in some examples, without examining the fulfilment of other conditions, if present).
- The second condition may be generalized eve more in that: if the energies (e.g. E1,1, E1,2, E2,1, E2,2) of a number Nmax (with Nmax between 2 and N or between 2 and N-1) of initial consecutive subframes of both the first synthesis signal s 1(n) and the second synthesis signal s 2(n) are all larger than the energy of the last subframe of the concealment signal by a predefined amount, then the second synthesis signal s 2(n) is preferentially chosen (in some examples, subjected to the fulfilment of other conditions, e.g. the first condition and/or the third condition). Otherwise, if at least one of the energies (e.g. E1,1, E1,2, E2,1, E2,2) of the first Nmax consecutive subframes of at least one of the first synthesis signal s 1(n) and the second synthesis signal s 2(n) is not larger than the energy of the last subframe of the concealment signal by the predefined amount, then the first synthesis signal s 1(n) is preferentially chosen (in some examples, without examining the fulfilment of other conditions, if present).
- Further some heuristic checks are introduced to guaranty that the second synthesis signal is only chosen if the energy envelope of the first synthesis is changing or increasing fast:
where "II" means logical operator "OR". The Threshold th 3 may have a value of 2 (or more in general between 1.5 and 2.5, or even more in general a value which is 1 or greater than 1; it may be th 3 = thPLC and/or th 3 = th 2 and/or th 3 > th 1), th 4 may have a value of 300 (or more in general larger than 100 or between 100 and 500, more in particular between 200 and 400, and even more in particular between 250 and 350; it may be th 4 > th 3 and/or th 4 > th 2 and/or th 4 > th 1 and/or th 4 > thPLC ) and th 5 may be 500 (or more in general more than 100 or between 300 and 700, more in particular between 400 and 600, and even more in particular between 450 and 550; it may be th 5 > th 4 and/or th 5 > th 3 and/or th 5 > th 2 and/or th 5 > th 1 and/or th 5 > thPLC ). In examples, it may be that th 4/ th 3 is 150, or more in general between 200 and 300; and/or th 5/th 4=1.666667 (or more in general between 1.5 and 1.8); and/or th 5/th 3 = 250 (or more in general between 200 and 300). In examples, it may be that th 3/th 1 >1.36 (in some examples by at least 1.2). At least one of these conditions must fulfilled to allow the second synthesis signal to be used. However, this third condition is optional and can be dropped. - Finally, the second synthesis signal is used, when:
- In the case that all the three conditions are used, then the second synthesis is used when:
- Since the second set of LPC does not work for every frame especially where the energy differences are small an additional smoothing is done to improve the performance of the recovery.
- Therefore, an energy compensating function g(n) is applied to the selected synthesis signal which is controlled by two gains g1 and g2. The first gain g1 may be determined by the ratio of the energy of the last half frame (or at least one last portion) of the concealed signal in the concealed frame and the first half frame (or at least one first portion) of the recovery frame.
- For the gains the half frame energies may be used, so the energy calculation is:
where L 2 is the length of a half frame (L2 = L/2) and k are the indices of the half frames (in some examples, it could be ). - The first gain g 1 may be obtained by:
which is the start gain of the scaling function g(n). g 1 may be limited to a ceiling value of 1.2 (or more in general a value D with 1≤D<2, e.g. a value between 1.1 and 1.3): if is larger than 1.2, g 1 will be set to 1.2 (or D), so that g 1 is between 0 and 1.2 (or between 0 and D). EPLChalf may be calculated in the same way like Eframe2 , but on the second half frame of the concealed frame before recovery, i.e. (but in some examples it may be EPLChalf = ). - The second gain g2 may be calculated as follows:
- In some cases, however, this formula may have a ceiling in e.g. g 2 = 1 in case
). Therefore, g 2 may be defined as being always 1 or less than 1, and therefore it may be guaranteed that g(n), at least in its final portion, has an attenuating (or at least a non-amplifying) effect. - In practice, g(n) permits to scale, sample by sample, the output or derived version s(n) of the synthesized audio signal in the properly received frame immediately following the at least one previously non-properly received frame by an energy compensating gain greater than 0, the energy compensating gain evolving, monotonically (e.g. strictly monotonically) or constantly, from the first value g 1 towards the second value g2. g 1 may be:
- comparatively high (e.g. larger than 1) in case of a ratio, between the energy of the concealed signal in a last portion of the concealed frame and the energy of the output or derived version (s) of the synthesized audio signal in an initial portion of the current frame, is comparatively high; and
- comparatively low (e.g. closer to 0) in case of the ratio, between the energy (EPLC
half ) of the concealed signal in the last portion of the concealed frame and the energy of the output or derived version (s(n)) of the synthesized audio signal in an initial portion of the current frame, is comparatively low. - For example, if EPLC
half >> E frame1 , then g1 is also high (e.g. higher than 1, e.g. reaching the ceiling value), while if EPLChalf << E frame1 , then g1 is also low (e.g. closer to 0).g 2 may be conditioned by a conditioning term (α * (E frame2 - EPLChalf )) which may be proportional to the difference between the energy (E frame2 ) of a last portion of the of the current frame and the energy (EPLChalf ) of the last portion of the concealed frame. g 2 may be: - comparatively high (e.g. closer to 1) in case of a ratio
between the energy (EPLChalf ) of the concealed signal in the last portion of the concealed frame, added with the conditioning term (α * (E frame2 - EPLChalf )) and the energy (E frane2 ) of the output or derived version (s) of the synthesized audio signal in the last portion of the current frame, is comparatively high; and - comparatively low (e.g. closer to 0) in case of the ratio
between the energy (EPLChalf ) of the concealed signal in the last portion of the concealed frame, added by the conditioning term (α * (E frame2 - EPLChalf )), and the energy (E frame2 ) of the output or derived version (s) of the synthesized audio signal in the last portion of the current frame, is comparatively low. - (Notably, g2 may be the same of the ratio
. Or, g2 may be defined so that, the higher the ratio, the higher g2 (e.g. with a ceiling and/or a floor)). - E.g., if EPLC
half and E frame2 are closer to each other, then g 2 is higher (e.g. closer to 1) than if EPLChalf and E frame2 are distant from each other. - With reference to
Fig. 9 (chart (b)), when E frame2 is smaller than or equal to EPLChalf , (i.e. if ), then g 2 may be set to be constantly 1 (ceiling value). For E frame2 > EPLChalf (i.e. ), then g 2 lies between 0 and 1 (and the higher , the lower g2, apart from a possible floor). The term α * (E frame2 - EPLChalf ) is a conditioning term which ensures that the energy compensation is not too strong if there is a big difference between EPLChalf and E frame2 . - The gain g(n) can be between 0 and 1 (and therefore an attenuation is performed) or (at least for some samples) larger than 1 (and in that case being amplifying). In some cases, the gain amplifies in a first part of the frame and attenuates in a last part of the frame.
- In general terms, the gain g(n) causes an energy compensation: the more the energy of the last portion of the concealed signal in the concealed frame is different from the energy of the initial portion of the output or derived synthesis signal in the current q-th frame, the larger the conditioning caused by g(n). In case, for example, In particular:
- 1) If EPLC
half < E frame1 , then the gain g(n) is, at least at the start of the q-th frame, attenuating (because ), meaning that the energy of the first portion of the q-th frame is attenuated, to avoid an unwanted step from the last portion of the (q-1)-th concealed frame (which has less energy) (and the higher the distance, the higher the attenuation);- a. (In particular, if EPLC
half << E frame1 , then the gain g(n), at least one at the start of the q-th, goes towards 0 or to a floor value of g1)
- a. (In particular, if EPLC
- 2) If EPLC
half > E frame1 , then the gain g(n) is, at least at the start of the q-th frame, amplifying (because ), meaning that the energy of the first portion of the q-th frame is amplified, to avoid an unwanted step from the last portion of the (q-1)-th concealed frame (which has more energy) (and the higher the distance, the higher the amplification);- a. (In particular, if EPLC
half >> E frame1 , then the gain g(n), at least one at the start of the q-th, goes towards a value greater than 1 or to a ceiling value of g1)
- a. (In particular, if EPLC
- 3) If EPLC
half = E frame1 , then the gain g(n) is unitary (because ) at least at the start of the q-th frame (e.g. at the very first sample of the q-th frame), meaning that the energy of the first portion of the q-th frame does not need to be compensated because is the same of the energy of the last portion of the (q-1)-th concealed frame (which has more energy); - 4) If EPLChalf + α * (E frame
2 - EPLChalf ) < E frame2 , then then the gain g(n) is, at least at the end of the q-th frame, attenuating (because of ) (and the higher the distance, the higher the attenuation, despite the attenuation being to the conditioning term) - 5) If EPLC
half + α * (E frame2 - EPLChalf ) > E frame2 , then then the gain g(n) is, at least at the end of the q-th frame, set to a ceiling value, e.g. 1; - 6) If EPLC
half = E frame2 , then the gain g(n) is, at least at the end of the q-th frame, unitary (because )- a. (in particular, if EPLC
half and E frame2 close with each other, then the gain g(n) goes towards 1 (ceiling value), at least at the end of the q-th frame) - b. (if EPLChalf << E frame
2 ) then the gain g(n) goes towards a value less than 1, at least at the end of the q-th frame)
- a. (in particular, if EPLC
- It is to be noted that g(n) may evolve constantly (e.g. if g 1 = g 2) or monotonically (e.g. strictly monotonically). In some examples, weak monotonicity may be also possible (e.g., following possible quantization of values of g(n) and/or in the case that a ceiling value or floor value is taken for an interval of samples).
- The proportionality coefficient α is a proportionality factor which depends on the energy in the first subframes of the chosen synthesis. The dependency may be described by a linear function α = m * Erel + c. (See
Fig. 12 ), where Erel may be the maximum of the ratio between the first subframe and the last subframe of the concealed frame and the second subframe and the last subframe of the concealed frame of the chosen synthesis. m may have a value like -2.14e-05 or -2e-05 < m < -3e-05 (or another value, which may be a negative value) and c a value 0.1999 < c < 0.20001 (or another value, which may be a positive value). For Erel smaller than 1, which means energy in PLC is bigger in the last subframe, α may have a constant value like 0.2, or 0.1999 < α < 0.20001. More in general, the coefficient α may be obtained as a linear combination of Erel , e.g. with a negative angular coefficient m and/or positive constant term. - If the first synthesis signal (si(n)) is chosen, Erel may be obtained by:
- For the case the second synthesis signal (s2(n)) is chosen, Erel may be obtained by:
- For Erel bigger than a value Emax like 7000 or for example 6900 < Erel < 7100, a constant value α is chosen, like α = 0.05. So α may be obtained by:
- An example of the function α = m * Erel + c is provided by
Fig. 12 , having in abscissa or and in ordinate the value α = m * Erel + c (with c=0.2). α may be not negative, and therefore we don't move from the first quadrant. - In practice, α may be obtained from a linear combination (e.g. with negative coefficient m) of Erel , but it may have a ceiling value (e.g. where 1 ≤ Erel ≤ Emax ) and/or a floor value (e.g. α = 0.05 if Erel > Emax ).
- The final applied scaling gain g(n) (here below being represented as "g[n]" without any distinction from "g(n)") is:
where L is the length of the frame and is the initial value of g. AGC is the active gain control with a value of 0.98 (or more in general a value between 0.9 and 0.999, e.g. between 0.97 and 0.99). The gain g(n) is computed in block 708. The function can also be written as where AGC is a base and n+1 is an exponent. -
Fig. 9 shows in chart (a) the energy compensating factor g(n) evolving along the L samples of the q-th frame for three different values of alpha (0.2, 0.01, and 0.0005), in the case second synthesis being selected and . In this case, the energy compensation is an attenuation, because the energy compensating gain is between 0 and 1. In this case, the higher the ratio , the higher the attenuation (i.e., the closer g(n) is to 0, g(n) being greater than 0). Further, the lower the alpha, the higher the attenuation (i.e., the closer g(n) is to 0, g(n) being greater than 0). The smoothing is done directly in the synthesis domain, so that the impact of the previous filter memory is not considered. This could in principle lead to discontinuities at the frame border. To prevent these discontinuities the zero input response sZIR (n) based on the synthesis memory of the erased frame is removed from the signal first (in particular in correspondence of the first subframe). The ZIR output is derived (at 702) by using the synthesis filter with a zero-signal input: where L 1 is the length of a subframe (e.g. with N subframes, e.g. with N>2, such as N=3 or N=4 or N=5, e.g. L1=L/N, eg. L/3, or L/4 or L/5), x 0(n) are just zeros (determined at 702), M is the filter order, ai are the LPC coefficients. The ZIR s'(n) is removed (at 704) from the unscaled synthesis s(n) (e.g. as outputted by block 620): - Then, the scaling function is applied at 360:
- Additionally, the memory of the synthesis (e.g. in block 360) for the next frame is also scaled:
- Finally, the ZIR is added (at 710) on top of the scaled synthesis:
- The excitation is updated based on the new scaled synthesis by using the analysis filter at 712:
-
Fig. 10 shows the result of using the second set of LPC (second audio parameter decoder 20). Compared to the signal after a frame loss inFig. 4 , the second synthesis signal s2(n) is clearly stable and similar to the clean signal. - Here above and below, reference is normally made to energy-related measurements in particular the form of average energies in subframes, such as
and E 2,k = , with L 1 being the length of the subframe for which the energy is calculated. However, in some cases it is also possible to simply measure the integral value of the energy, such as and . In particular, it may be unnecessary to calculate the ratios and if, for example, the evaluation of the condition is to be performed, because the ratio would notwithstanding cancel the L 1 values at the numerator and the denominator. Notwithstanding, it is at least theoretically possible to have that the subframes for which the energy-related measurements are calculated are different. E.g. a first subframe for the first synthesis signal could have length L 1,1 while a second subframe for the second synthesis signal could have length L 1,2). In this case, it could be opportune to average the energy-related measurements e.g. through and . Or, it could be possible to modify the thresholds (e.g. th 1) to keep into account the different lengths, while using the integral values, such as and E 2,k = . It is notwithstanding here supposed that all the subframes have the same length (apart for the possibility that some subframes have not exactly the same number of samples, e.g. by virtue of the total number of samples N not being divisible by 3). - Further, it is noted that the energy-related measurements are not uniquely energy measurements. For example, envelope measurements (which are also energy-related measurements) may be performed. The envelope can be described as the change (progression) of the energy over time (e.g., over the samples). For example, blocks 608 and 610 may obtain an envelope of the first synthesis signal s1(n) and the second synthesis signal s2(n), respectively. As a condition evaluated by block 620 of the selector 60, the envelopes may be evaluated so as:
- to select the second synthesis signal (s2) in case the envelope of the second synthesis signal (s2) is, at least in the initial subframe or in a sequence of initial subframes, more stable, by at least one predetermined extent (e.g. based on at least a predetermined threshold), than the envelope of the first synthesis signal (s1), and
- to select the first synthesis signal otherwise.
- The stability of the signals may be evaluated, for example, by measuring the fluctuations of the signal over time.
Fig. 4 shown, for example, a highly fluctuating signal ("Synthesis signal obtained using the parameters of the concealed frame: set 1"). Measurements of fluctuations of an envelope are known in the art, and can be based, for example, on variance measurements, etc. In general terms, however, where the envelope of the second synthesis signal has a stability which overwhelms the stability of the first synthesis signal by at least the predetermined extent (by the predetermined threshold), then the second synthesis signal is enough stable and can be used instead of the first synthesis signal as selected synthesis signal (and as derived or output synthesis signal). - Notably, the first and second conditions described above may be used, in some examples, as evaluating the stability of the first and second synthesis signals.
-
Figs. 2 and3 show the LPC frequency response of the concealed frame (Aq_dec) vs. the clean signal (Aq_dec_clean). - Fir. 4 shows the clean signal compared to the synthesis signal (e.g. s1(n)) obtained using the parameters of the concealed frame (e.g., like in the prior art, or as outputted by the first audio parameter decoding unit 10), and shows evident oscillations.
-
Fig. 9 shows: - In chart (a), the behavior of g(n) along the samples of the current q-th frame parametrized for different proportionality coefficients α (in this case, the higher the α, the higher the attenuation along the samples, in case of
) - In chart (b), the behavior of g(n) along the samples of the current q-th frame parametrized for different ratios E frame
2 /EPLChalf (notably, the higher the ratio, the higher the attenuation along the samples, in this case). -
Fig. 10 shows the clean signal compared to the synthesis signal (s2(n)) obtained without using the parameters of the concealed frame (as outputted by the second audio parameter decoding unit 10), and shows a stable behavior as compared to that ofFig. 4 . - It has been explained above that the current q-th frame and the (q-1)-th concealed frame are partitioned both:
according to a first partitioning which partitions the current q-th frame and the (q-1)-th concealed frame among a sequence of subframes in a number N of subframes which is 3 or more than 3 (examples of these subframes are E 1,1, E 2,1, E 1,2, E 2,2 and EPLCsub etc.), the decoded audio parameters being decoded for each subframe, the selection at 60 being performed only based on an initial subframe or a group (sequence) of initial subframes of the current q-th frame and on the last subframe of the (q-1)-th concealed frame (further, in case of partial synthesis before the selection 60, the synthesis being originally only performed for on single initial subframe or a particular sequence of initial subframes; further and Erel being calculated based on the first subframe only);
according to a second partitioning which partitions the current frame and the concealed frame among a sequence of portions (e.g. half frames) in a number of portions which is 2 (in the case of half frames) or more than 2, but the number of portions being less than the number N of subframes, and each portion having larger time length than any subframe (this partitioning is preferably used for calculating the energy compensating gain g(n); examples are EPLChalf ,E frame2 ,E frame1 ). - It has been noted that, in this way, diversity is increased.
- It is not strictly necessary to perform the decoding of the audio parameters to obtain LFS parameters, then LSP parameters, and finally LPC parameters. While the examples above have been mainly directed to that technique, other techniques may be implemented.
- Also other representations like ISF (ImmittanceSpectral Frequency) of the LPC or any other could be used with this approach (see [6]).
- While the second audio parameter decoding unit 20 mainly refrains from adopting the parameters from the concealed signal, it is notwithstanding noted that the excitation x(n) can be obtained from the (q-1)-th concealed audio frame. In several examples, however, the encoded parameters are differential parameters and the audio signal is synthesized by taking into account an excitation (which can be obtained, or at least inferred, in some examples, from the concealed audio signal in the (q-1)-th concealed frame).
- A major advantage of aspects of the proposed technique is the consideration of the mismatch between LPC parameters of a concealed loss frame and a first well-received, aka recovery, q-th frame. The prior art methods propose smoothing operations of the excitation domain signal but don't treat the mismatch sets of LPC into account. Further, the proposed technique operates solely on decoder side and requires no extra side information to be transmitted.
- An implementation based on a neural network is illustrated in
Fig. 11 , where the synthesizing unit 50 ofFig. 1 includes a neural network, NN-based synthesizing unit 90, which implements a NN. In particular, in the example ofFig. 11 block 620 (and the selector 60) is represented as being part of the block "Recovery and Computation/Decision of speech parameters" 91 (which implements blocks 10 and 20, in particular), while block 623 is part of the NN-based synthesizing unit 90. - Transmitted speech parameters like LPC or any representation of it may be used as input for a neural network NN (in a NN processor 90) with at least one learnable layer (e.g. with a plurality of learnable layers). The learnable layer may be a generative adversarial neural (GAN) layer.
- In case of concealment, concealed parameters (generated by the concealment unit 40, which is not necessary part of the NN-based synthesizing unit) such as LCP parameters ai1 and ai2 (see above) are fed to the NN-based synthesizing unit 90. In case of a properly received frame which is not a recovery frame (i.e. the immediately preceding frame is also a properly received frame), the received speech parameters are used (i.e. the first set of parameters is used, and the first audio parameter decoding unit 10 is activated, while the second audio parameter decoding unit 20 is deactivated). In case of recovery frame (i.e. properly received frame after a concealed frame), the block 620 (which may be a deterministic block) decides which synthesis signal between s1(n) and s2(n) is to be used. Speech (audio) parameters are in this case used by the NN-based synthesizing unit 90 (also implementing blocks 612, 614, and 616) to generate the synthesis signal, e.g. based on the excitation (e.g. previously obtained). Here, a NN may be used. If The proposed blocks 604, 608, 610, 614, 616, 618 may be be put in front of a learnable layer of the NN-based synthesizing unit 90 where the prediction is created from the LPC, which are converted from LSP from the first set or second set in recovery case, LSPplc in case of concealment or LSP from the first set in clean channel case.
-
-
- According to an aspect, there is provided an audio decoder (e.g. 100) for synthesizing an audio signal (e.g. s) from a bitstream which represents the audio signal, the audio decoder (e.g. 100) including:
- a bitstream receiver (e.g. 5), to receive the bitstream representative of the audio signal, the bitstream having, encoded therein, a set of encoded audio parameters (e.g. ΔLSFq ) for each frame of the audio signal,
- a first audio parameter decoding unit (e.g. 10), to decode, for a current properly received frame, a first set of decoded audio parameters (e.g. ai1) from at least the set of encoded audio parameters,
- a concealment unit (e.g. 40) to conceal at least one non-properly received frame based on at least one previously properly received frame, so as to generate at least one set of concealment audio parameters;
- a second audio parameter decoding unit (e.g. 20), to decode, for a current properly received frame which immediately follows the at least one previously non-properly received frame, a second set of decoded audio parameters (e.g. ai2) from at least the set of encoded audio parameters, the second set of decoded audio parameters (e.g. ai2) being different from the first set of decoded audio parameters (e.g. ai1);
- a synthesizing unit (e.g. 50), to output, or derive, an output or derived version (e.g. s) of the synthesized audio signal in such way that, if the current properly received frame immediately follows the at least one previously non-properly received frame, a selection (e.g. 60) is made between:
- a first version (e.g. s1) of the synthesized audio signal (e.g. s), synthesized from at least the first set of decoded audio parameters; and
- a second version (e.g. s2) of the synthesized audio signal (e.g. s), synthesized from the second set of decoded audio parameters.
- According to a further aspect, the first audio parameter decoding unit (e.g. 10) is configured, in the case the current frame is the properly received frame which immediately follows the at least one previously non-properly received frame, to decode the first set of decoded audio parameters from the at least one set of concealment parameters of the immediately preceding non-properly received frame and the set of encoded audio parameters of the current frame.
- According to a further aspect, the first audio parameter decoding unit (e.g. 10) is configured, in the case the current frame is the properly received frame which immediately follows the at least one previously non-properly received frame, to decode the first set of decoded audio parameters through a first prediction (e.g. 604) of the first set of audio parameters obtained from the version of the set of concealment parameters of the immediately preceding non-properly received frame and a pre-defined value.
- According to a further aspect, the second audio parameter decoding unit (e.g. 20) is configured, in the case the current frame is the properly received frame which immediately follows the at least one previously non-properly received frame, to decode (e.g. 612) the second set of decoded audio parameters from the encoded audio parameters of the current frame.
- According to a further aspect, the second audio parameter decoding unit (e.g. 20) is configured, in the case the current frame is the properly received frame which immediately follows the at least one previously non-properly received frame, to decode the second set of decoded audio parameters from encoded audio parameters of the current frame but not from the set of concealment parameters of the immediately preceding non-properly received frame.
according to a further aspect, the first audio parameter decoding unit is configured, in the case the current frame is the properly received frame which immediately follows the at least one previously non-properly received frame, to decode the first set of decoded audio parameters, from a first number of concealment parameters of the set of concealment parameters of the immediately preceding non-properly received frame and from the encoded audio parameters of the current frame,
wherein the second audio parameter decoding unit is configured, in the case the current frame is the properly received frame which immediately follows the at least one previously non-properly received frame, to decode the second set of decoded audio parameters from the encoded audio parameters of the current frame and from a second number of concealment parameters of the set of concealment parameters of the immediately preceding non-properly received frame which is smaller than the first number of concealment parameters. - According to a further aspect, the audio synthesizer may be configured, in the case the current frame is the properly decoded frame which immediately follows the at least one previously non-properly received frame, to perform the selection (e.g. 620) between the first synthesis signal (e.g. s1) and the second synthesis signal (e.g. s2) based on a comparison between at least energy-related measurements on the first version (e.g. s1) of the synthesized audio signal with energy-related measurements on the second version (e.g. s2) of the synthesized audio signal, so as to output or derive, as the output or derived version (e.g. s) of the synthesized audio signal, the second version (e.g. s2) of the synthesized audio signal in case at least one of the following condition or a combination of at least two of the following conditions is satisfied:
- the energy of the first version (e.g. s1) of the synthesized audio signal is, at least in an initial subframe or sequence of initial subframes, larger than the energy of the second version (e.g. s2) of the synthesized audio signal by at least one threshold ratio (e.g. th1) greater than or equal to 1;
- the energy of the first version (e.g. s1) of the synthesized audio signal is larger, by at least one threshold ratio equal to or greater than 1, in at least one initial subframe or sequence of initial subframes of the first version (e.g. s1) of the synthesized audio signal, than the energy of the concealed signal in at least a final subframe or sequence of final subframes of the concealed frame;
- the energy of the second version (e.g. s2) of the synthesized audio signal is larger, by at least one threshold ratio equal to or greater than 1, in at least one initial subframe or sequence of initial subframes of the second version (e.g. s2) of the synthesized audio signal, than the energy of the concealed signal in at least a final subframe or sequence of final subframes of the concealed frame; and
- the energy of the first version (e.g. s1) of the synthesized audio signal is larger by at least one threshold ratio, in at least one initial subframe or sequence of initial subframes of the first version (e.g. s1) of the synthesized audio signal, than the energy of the second version (e.g. s2) of the synthesized audio signal in at least one initial subframe or sequence of initial subframes of the second audio signal (e.g. s2) corresponding to the at least one initial subframe or sequence of initial subframes of the first version (e.g. s1), and,
- in case the condition is not satisfied, to output or derive, as the output or derived version (e.g. s) of the synthesized audio signal, the first version (e.g. s1) of the synthesized audio signal.
- According to a further aspect, the audio synthesizer may be configured, in the case the current frame is the properly decoded frame which immediately follows the at least one previously non-properly received frame, to perform the selection (e.g. 620) between the first synthesis signal (e.g. s1) and the second synthesis signal (e.g. s2) at least based on a comparison between at least energy-related measurements on the first version (e.g. s1) of the synthesized audio signal with energy-related measurements on the second version (e.g. s2) of the synthesized audio signal, so as to output or derive, as the output or derived version (e.g. s) of the synthesized audio signal, the second version (e.g. s2) of the synthesized audio signal in case of at least one of the following conditions, or a combination of at least one of the following conditions, is satisfied:
- the energy of at least one initial subframe, or of a sequence of initial subframes (e.g. E1,1, E1,2), of the first version (e.g. s1) of the synthesized audio signal is, subframe by subframe, larger than the energy of at least one initial subframe, or of a sequence of initial subframes (e.g. E2,1, E2,2), of the second version (e.g. s2) of the synthesized audio signal according to at least one predetermined ratio threshold (e.g. th1, th2) equal to or greater than 1;
- the energy of at least one initial subframe, or of each subframe of a sequence of initial subframes (e.g. E1,1, E1,2), of the first version (e.g. s1) of the synthesized audio signal is larger than the energy (e.g. EPLCsub) of a final subframe of the audio signal in the concealed frame according to at least one predetermined ratio threshold (e.g. thPLC) equal to or greater than 1; and
- the energy of at least one initial subframe, or of each subframe of a sequence of initial subframes (e.g. E2,1, E2,2), of the second version (e.g. s2) of the synthesized audio signal is larger than the energy (e.g. EPLCsub) of a final subframe of the audio signal in the concealed frame according to at least one predetermined ratio threshold (e.g. thPLC) equal to or greater than 1; and
in case the at least one condition or combination of conditions is not satisfied, to output or derive, as the output or derived version (e.g. s) of the synthesized audio signal, the first version (e.g. s1) of the synthesized audio signal. - According to a further aspect, the audio synthesizer may be configured, in the case the current frame is the properly decoded frame which immediately follows the at least one previously non-properly received frame, to perform the selection (e.g. 620) between the first synthesis signal (e.g. s1) and the second synthesis signal (e.g. s2) based on a comparison between an envelope of the first synthesis signal and an envelope of the second synthesis signal,
so as to select the second synthesis signal (e.g. s2) in case the envelope of the second synthesis signal (e.g. s2) is, at least in the initial subframe or in a sequence of initial subframes, more stable, by at least one predetermined extent, than the envelope of the first synthesis signal (e.g. s1), and to select the first synthesis signal otherwise. - According to a further aspect, the audio decoder is configured to scale (e.g. 360), sample by sample, the output or derived version (e.g. s) of the synthesized audio signal in the properly received frame immediately following the at least one previously non-properly received frame by an energy compensating gain greater than 0, the energy compensating gain reducing, in at least one portion of the current frame, the energy gap between the concealed signal in a last portion of the concealed frame and the output or derived version of the synthesized audio signal in the at least one portion of the current frame.
- According to a further aspect, the energy compensating gain evolves, monotonically or constantly, from a first value towards a second value, the first value being:
- comparatively high in case of a ratio, between the energy of the concealed signal in a last portion of the concealed frame and the energy of the output or derived version (e.g. s) of the synthesized audio signal in an initial portion of the current frame, is comparatively high; and
- comparatively low in case of the ratio, between the energy of the concealed signal in a last portion of the concealed frame and the energy of the output or derived version (e.g. s) of the synthesized audio signal in an initial portion of the current frame, is comparatively low,
- comparatively high in case of a ratio between the energy of the concealed signal in the last portion of the concealed frame, added with the conditioning term, and the energy of the output or derived version (e.g. s) of the synthesized audio signal in the last portion of the current frame, is comparatively high; and
- comparatively low in case of the ratio between the energy of the concealed signal in the last portion of the concealed frame, added by the conditioning term, and the energy of the output or derived version (e.g. s) of the synthesized audio signal in the last portion of the current frame, is comparatively low.
- According to a further aspect, the current frame and the concealed frame are partitioned both according to a first partitioning which partitions the current frame and the concealed frame among a sequence of subframes in a number of subframes which is 3 or more than 3, and according to a second partitioning which partitions the current frame and the concealed frame among a sequence of portions in a number of portions which is 2 or more than 2, but the number of portions being less than the number of subframes, and each portion having larger time length than any subframe.
- According to a further aspect, the conditioning term has a proportionality coefficient α is linearly dependent on a maximum value between a first ratio and a second ratio, where the first ratio is a ratio between the energy of the initial subframe of the output or derived version (e.g. s) of the synthesized audio signal and the energy of the last subframe of the concealed frame, and the second ratio is a ratio between the energy of the second subframe of the output or derived version (e.g. s) of the synthesized audio signal and the energy of the last subframe of the concealed frame.
- According to a further aspect, the energy compensating gain is comparatively close to 1, in at least one portion of the current frame, in case the distance between the energy of the output or derived version (e.g. s) of the synthesized audio signal in the at least one portion of the current frame and the energy of concealed signal in the end portion of the concealed frame is comparatively low, and
the energy compensating gain is comparatively distant from 1, in the at least one portion of the current frame, in case the distance between the energy of the output or derived version (e.g. s) of the synthesized audio signal in the at least one portion of the current frame and the energy of concealed signal in the end portion of the concealed frame is comparatively high. - According to a further aspect, the energy compensating gain is defined recursively by cross-fading the energy compensating gain for an immediately preceding sample with the second gain value.
- According to a further aspect, the synthesizing unit (e.g. 50) is configured to, initially, partially synthesize only an initial subframe, or a group of initial subframes, of the of the first synthesis signal (e.g. s1) and, partially synthesize only an initial subframe, or a group of initial subframes, of the second synthesis signal (e.g. s2), so that the selection (e.g. 60) is based on energy-related measurements on the partially synthesized version of the first synthesis signal (e.g. s1) and the partially synthesized version of the second synthesis signal (e.g. s2), so that, only after the selection (e.g. 60), the remaining subframe or subframe of the selected synthesis signal is or are synthesized, without synthesizing the remaining subframe or subframe of the non-selected synthesis signal.
- According to a further aspect, the first audio parameter decoding unit (e.g. 10) is configured to, initially, decode, respectively, only a first subset of the first set of decoded audio parameters and only a second subset of the second decoded audio parameters, the first subset and second subset corresponding to the initial subframe, or the group of initial subframes, so that the selection (e.g. 60) is based on energy-related measurements of the partially synthesized version of the first synthesis signal (e.g. s1) obtained from the first subset and the partially synthesized version of the second synthesis signal (e.g. s2) obtained from the second subset, so that, only after the selection (e.g. 60), the remaining audio parameters of the set of audio parameter associated with the selected synthesis signal are decoded, and the remaining audio parameters of the set of audio parameter associated with the non-selected synthesis signal are not decoded.
- According to a further aspect, further comprising a neural network processor using at least one learnable layer, to be inputted with the concealment parameters as well as the decoded parameters and/or the first and second synthesis signals or the output or derived version of the synthesis signal, so as to process the output or derived version of the synthesis signal through the at least one learnable layer.
- According to a further aspect, the at least one learnable layers is a generative adversarial network, GAN, learnable layer.
- According to a further aspect, the encoded audio parameter include, or provide information on, linear spectral frequencies.
- According to an aspect, there is provided a method for synthesizing an audio signal (e.g. s) from a bitstream which represents the audio signal, the method including:
- receiving the bitstream representative of the audio signal, the bitstream having, encoded therein, a set of encoded audio parameters (e.g. ΔLSFq ) for each frame of the audio signal,
- decoding, for a current properly received frame, a first set of decoded audio parameters (e.g. ai1, LSPend) from at least the set of encoded audio parameters,
- concealing at least one previously non-properly received frame based on at least one previously properly received frame, so as to generate at least one set of concealment audio parameters, and to synthesize a concealed frame from the at least one set of concealment audio parameters;
- decoding, for a current properly received frame which immediately follows the at least one previously non-properly received frame, a second set of decoded audio parameters from at least the set of encoded audio parameters, the second set of audio parameters being different from the first set of audio parameters,
- outputting or deriving an output or derived version (e.g. s) of the synthesized audio signal in such way that, if the current properly received frame immediately follows the at least one previously non-properly received frame, a selection is made between:
- a first version (e.g. s1) of the synthesized audio signal (e.g. s), synthesized from at least the set of encoded audio parameters
- a second version (e.g. s2) of the synthesized audio signal (e.g. s), synthesized from the second set of decoded audio parameters.
- According to an aspect, there is provided a non-transitory storage unit storing instructions which, when executed by a processor, cause the processor to perform the method above (or any of the methods above and below).
- Depending on certain implementation requirements, examples may be implemented in hardware. The implementation may be performed using a digital storage medium, for example a floppy disk, a Digital Versatile Disc (DVD), a Blu-Ray Disc, a Compact Disc (CD), a Read-only Memory (ROM), a Programmable Read-only Memory (PROM), an Erasable and Programmable Read-only Memory (EPROM), an Electrically Erasable Programmable Read-Only Memory (EEPROM) or a flash memory, having electronically readable control signals stored thereon, which cooperate (or are capable of cooperating) with a programmable computer system such that the respective method is performed. Therefore, the digital storage medium may be computer readable.
- Generally, examples may be implemented as a computer program product with program instructions, the program instructions being operative for performing one of the methods when the computer program product runs on a computer. The program instructions may for example be stored on a machine readable medium.
- Other examples comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier. In other words, an example of method is, therefore, a computer program having program instructions for performing one of the methods described herein, when the computer program runs on a computer.
- A further example of the methods is, therefore, a data carrier medium (or a digital storage medium, or a computer-readable medium) comprising, recorded thereon, the computer program for performing one of the methods described herein. The data carrier medium, the digital storage medium or the recorded medium are tangible and/or non-transitionary, rather than signals which are intangible and transitory.
- A further example comprises a processing unit, for example a computer, or a programmable logic device performing one of the methods described herein.
- A further example comprises a computer having installed thereon the computer program for performing one of the methods described herein.
- A further example comprises an apparatus or a system transferring (for example, electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may, for example, be a computer, a mobile device, a memory device or the like. The apparatus or system may, for example, comprise a file server for transferring the computer program to the receiver.
- In some examples, a programmable logic device (for example, a field programmable gate array) may be used to perform some or all of the functionalities of the methods described herein. In some examples, a field programmable gate array may cooperate with a microprocessor in order to perform one of the methods described herein. Generally, the methods may be performed by any appropriate hardware apparatus.
- While this invention has been described in terms of several examples, there are alterations, permutations, and equivalents which fall within the scope of this invention. It should also be noted that there are many alternative ways of implementing the methods and compositions of the present invention. It is therefore intended that the following appended claims be interpreted as including all such alterations, permutations and equivalents as fall within the true spirit and scope of the present invention.
-
- [1] https://www.etsi.org/dliver/etsi_ts/126400_126499/126445/12.06.00_60/ts_126445v120600p.pdf
- [2] internet Low Bitrate Codec, WEB RTC, https ://webrtc. github .io/webrtc-org/license/ilbc-free-ware/ilbc-extra-documentation/
- [3] Method and Decive for efficient frame erasure concealment in linear predictive based speech codecs, Voice Age Company, https://www.voiceageevs.com/documents/patents/usa/VAEVS%201000%20-%20US%207%20693%20710%20B2.PDF
- [4] Method and Device for efficient frame erasure concealment in speech codecs https://www.voiceageevs.com/documents/patents/usa/VAEVS%202200%20-%20US%208%20255%20207%20B2.PDF
- [5] Analysis by Adversarial Synthesis - A Novel Approach for Speech Vocodinghttp://arxiv.org/pdf/1907.00772
- [6] http://www.jcomputers.us/vol2/jcp0207-09.pdf
Claims (20)
- Audio decoder (100) for synthesizing an audio signal from a bitstream which represents the audio signal, the audio decoder (100) including:a bitstream receiver (5), to receive the bitstream representative of the audio signal, the bitstream having, encoded therein, a set of encoded audio parameters (ΔLSFq ) for each frame of the audio signal,a first audio parameter decoding unit (10), to decode, for a current properly received frame, a first set of decoded audio parameters (an) from at least the set of encoded audio parameters,a concealment unit (40) to conceal at least one non-properly received frame based on at least one previously properly received frame, so as to generate at least one set of concealment audio parameters;a second audio parameter decoding unit (20), to decode, for a current properly received frame which immediately follows the at least one previously non-properly received frame, a second set of decoded audio parameters (ai2) from at least the set of encoded audio parameters, the second set of decoded audio parameters (ai2) being different from the first set of decoded audio parameters (ai1);a synthesizing unit (50), to output, or derive, an output or derived version of the synthesized audio signal in such way that, if the current properly received frame immediately follows the at least one previously non-properly received frame, a selection (60) is made between:a first version (s1) of the synthesized audio signal, synthesized from at least the first set of decoded audio parameters; anda second version (s2) of the synthesized audio signal, synthesized from the second set of decoded audio parameters.
- The audio synthesizer of claim 1, wherein the first audio parameter decoding unit (10) is configured, in the case the current frame is the properly received frame which immediately follows the at least one previously non-properly received frame, to decode the first set of decoded audio parameters from the at least one set of concealment parameters of the immediately preceding non-properly received frame and the set of encoded audio parameters of the current frame.
- The audio synthesizer of claim 2, wherein the first audio parameter decoding unit (10) is configured, in the case the current frame is the properly received frame which immediately follows the at least one previously non-properly received frame, to decode the first set of decoded audio parameters through a first prediction (604) of the first set of audio parameters obtained from the version of the set of concealment parameters of the immediately preceding non-properly received frame and a pre-defined value.
- The audio synthesizer of any of the preceding claims, wherein the second audio parameter decoding unit (20) is configured, in the case the current frame is the properly received frame which immediately follows the at least one previously non-properly received frame, to decode (612) the second set of decoded audio parameters from the encoded audio parameters of the current frame.
- The audio synthesizer of any of the preceding claims, wherein the second audio parameter decoding unit (20) is configured, in the case the current frame is the properly received frame which immediately follows the at least one previously non-properly received frame, to decode the second set of decoded audio parameters from encoded audio parameters of the current frame but not from the set of concealment parameters of the immediately preceding non-properly received frame.
- The audio synthesizer of any of the preceding claims, configured, in the case the current frame is the properly decoded frame which immediately follows the at least one previously non-properly received frame, to perform the selection (620) between the first synthesis signal (s1) and the second synthesis signal (s2) based on a comparison between at least energy-related measurements on the first version (s1) of the synthesized audio signal with energy-related measurements on the second version (s2) of the synthesized audio signal, so as to output or derive, as the output or derived version of the synthesized audio signal, the second version (s2) of the synthesized audio signal in case at least one of the following condition or a combination of at least two of the following conditions is satisfied:the energy of the first version (s1) of the synthesized audio signal is, at least in an initial subframe or sequence of initial subframes, larger than the energy of the second version (s2) of the synthesized audio signal by at least one threshold ratio (th1) greater than or equal to 1;the energy of the first version (s1) of the synthesized audio signal is larger, by at least one threshold ratio equal to or greater than 1, in at least one initial subframe or sequence of initial subframes of the first version (s1) of the synthesized audio signal, than the energy of the concealed signal in at least a final subframe or sequence of final subframes of the concealed frame;the energy of the second version (s2) of the synthesized audio signal is larger, by at least one threshold ratio equal to or greater than 1, in at least one initial subframe or sequence of initial subframes of the second version (s2) of the synthesized audio signal, than the energy of the concealed signal in at least a final subframe or sequence of final subframes of the concealed frame; andthe energy of the first version (s1) of the synthesized audio signal is larger by at least one threshold ratio, in at least one initial subframe or sequence of initial subframes of the first version (s1) of the synthesized audio signal, than the energy of the second version (s2) of the synthesized audio signal in at least one initial subframe or sequence of initial subframes of the second audio signal (s2) corresponding to the at least one initial subframe or sequence of initial subframes of the first version (s1), and,in case the condition is not satisfied, to output or derive, as the output or derived version of the synthesized audio signal, the first version (s1) of the synthesized audio signal.
- The audio synthesizer of any of the preceding claims, configured, in the case the current frame is the properly decoded frame which immediately follows the at least one previously non-properly received frame, to perform the selection (620) between the first synthesis signal (s1) and the second synthesis signal (s2) at least based on a comparison between at least energy-related measurements on the first version (s1) of the synthesized audio signal with energy-related measurements on the second version (s2) of the synthesized audio signal, so as to output or derive, as the output or derived version of the synthesized audio signal, the second version (s2) of the synthesized audio signal in case of at least one of the following conditions, or a combination of at least one of the following conditions, is satisfied:- the energy of at least one initial subframe, or of a sequence of initial subframes (E1,1, E1,2), of the first version (s1) of the synthesized audio signal is, subframe by subframe, larger than the energy of at least one initial subframe, or of a sequence of initial subframes (E2,1, E2,2), of the second version (s2) of the synthesized audio signal according to at least one predetermined ratio threshold (th1, th2) equal to or greater than 1;- the energy of at least one initial subframe, or of each subframe of a sequence of initial subframes (E1,1, E1,2), of the first version (s1) of the synthesized audio signal is larger than the energy (EPLCsub) of a final subframe of the audio signal in the concealed frame according to at least one predetermined ratio threshold (thPLC) equal to or greater than 1; and- the energy of at least one initial subframe, or of each subframe of a sequence of initial subframes (E2,1, E2,2), of the second version (s2) of the synthesized audio signal is larger than the energy (EPLCsub) of a final subframe of the audio signal in the concealed frame according to at least one predetermined ratio threshold (thPLC) equal to or greater than 1; andin case the at least one condition or combination of conditions is not satisfied, to output or derive, as the output or derived version of the synthesized audio signal, the first version (s1) of the synthesized audio signal.
- The audio synthesizer of any of the preceding claims, configured, in the case the current frame is the properly decoded frame which immediately follows the at least one previously non-properly received frame, to perform the selection (620) between the first synthesis signal (s1) and the second synthesis signal (s2) based on a comparison between an envelope of the first synthesis signal and an envelope of the second synthesis signal,so as to select the second synthesis signal (s2) in case the envelope of the second synthesis signal (s2) is, at least in the initial subframe or in a sequence of initial subframes, more stable, by at least one predetermined extent, than the envelope of the first synthesis signal (s1), andto select the first synthesis signal otherwise.
- The audio decoder of any of the preceding claims, wherein the audio decoder is configured to scale (360), sample by sample, the output or derived version of the synthesized audio signal in the properly received frame immediately following the at least one previously non-properly received frame by an energy compensating gain greater than 0, the energy compensating gain reducing, in at least one portion of the current frame, the energy gap between the concealed signal in a last portion of the concealed frame and the output or derived version of the synthesized audio signal in the at least one portion of the current frame.
- The audio decoder of claim 9, wherein the energy compensating gain evolves, monotonically or constantly, from a first value towards a second value, the first value being:comparatively high in case of a ratio, between the energy of the concealed signal in a last portion of the concealed frame and the energy of the output or derived version of the synthesized audio signal in an initial portion of the current frame, is comparatively high; andcomparatively low in case of the ratio, between the energy of the concealed signal in a last portion of the concealed frame and the energy of the output or derived version of the synthesized audio signal in an initial portion of the current frame, is comparatively low,the second value, conditioned by a conditioning term which is proportional to the difference between the energy of a last portion of the of the current frame and the energy of the last portion of the concealed frame, wherein the second value is:comparatively high in case of a ratio between the energy of the concealed signal in the last portion of the concealed frame, added with the conditioning term, and the energy of the output or derived version of the synthesized audio signal in the last portion of the current frame, is comparatively high; andcomparatively low in case of the ratio between the energy of the concealed signal in the last portion of the concealed frame, added by the conditioning term, and the energy of the output or derived version of the synthesized audio signal in the last portion of the current frame, is comparatively low.
- The audio decoder of claim 10, wherein the current frame and the concealed frame are partitioned both according to a first partitioning which partitions the current frame and the concealed frame among a sequence of subframes in a number of subframes which is 3 or more than 3, and according to a second partitioning which partitions the current frame and the concealed frame among a sequence of portions in a number of portions which is 2 or more than 2, but the number of portions being less than the number of subframes, and each portion having larger time length than any subframe.
- The audio decoder of any of claims 9-11, wherein the conditioning term has a proportionality coefficient α is linearly dependent on a maximum value between a first ratio and a second ratio, where the first ratio is a ratio between the energy of the initial subframe of the output or derived version of the synthesized audio signal and the energy of the last subframe of the concealed frame, and the second ratio is a ratio between the energy of the second subframe of the output or derived version of the synthesized audio signal and the energy of the last subframe of the concealed frame.
- The audio decoder of any of claims 9-12, wherein the energy compensating gain is comparatively close to 1, in at least one portion of the current frame, in case the distance between the energy of the output or derived version of the synthesized audio signal in the at least one portion of the current frame and the energy of concealed signal in the end portion of the concealed frame is comparatively low, and
the energy compensating gain is comparatively distant from 1, in the at least one portion of the current frame, in case the distance between the energy of the output or derived version of the synthesized audio signal in the at least one portion of the current frame and the energy of concealed signal in the end portion of the concealed frame is comparatively high. - The audio decoder of any of claims 9-13, wherein the energy compensating gain is defined recursively by cross-fading the energy compensating gain for an immediately preceding sample with the second gain value.
- The audio decoder of any of the preceding claims, wherein the synthesizing unit (50) is configured to, initially, partially synthesize only an initial subframe, or a group of initial subframes, of the of the first synthesis signal (s1) and, partially synthesize only an initial subframe, or a group of initial subframes, of the second synthesis signal (s2), so that the selection (60) is based on energy-related measurements on the partially synthesized version of the first synthesis signal (s1) and the partially synthesized version of the second synthesis signal (s2), so that, only after the selection (60), the remaining subframe or subframe of the selected synthesis signal is or are synthesized, without synthesizing the remaining subframe or subframe of the non-selected synthesis signal.
- The audio decoder of claim 14 or 15,
wherein the first audio parameter decoding unit (10) is configured to, initially, decode, respectively, only a first subset of the first set of decoded audio parameters and only a second subset of the second decoded audio parameters, the first subset and second subset corresponding to the initial subframe, or the group of initial subframes, so that the selection (60) is based on energy-related measurements of the partially synthesized version of the first synthesis signal (s1) obtained from the first subset and the partially synthesized version of the second synthesis signal (s2) obtained from the second subset, so that, only after the selection (60), the remaining audio parameters of the set of audio parameter associated with the selected synthesis signal are decoded, and the remaining audio parameters of the set of audio parameter associated with the non-selected synthesis signal are not decoded. - The audio decoder of any of the preceding claims, further comprising a neural network processor using at least one learnable layer, to be inputted with the concealment parameters as well as the decoded parameters and/or the first and second synthesis signals or the output or derived version of the synthesis signal, so as to process the output or derived version of the synthesis signal through the at least one learnable layer.
- The audio decoder of claim 17, wherein the at least one learnable layers is a generative adversarial network, GAN, learnable layer.
- The audio decoder of any of the preceding claims, wherein the encoded audio parameter include, or provide information on, linear spectral frequencies.
- A method for synthesizing an audio signal from a bitstream which represents the audio signal, the method including:receiving the bitstream representative of the audio signal, the bitstream having, encoded therein, a set of encoded audio parameters (ΔLSFq ) for each frame of the audio signal,decoding, for a current properly received frame, a first set of decoded audio parameters (ai1, LSPend) from at least the set of encoded audio parameters,concealing at least one previously non-properly received frame based on at least one previously properly received frame, so as to generate at least one set of concealment audio parameters, and to synthesize a concealed frame from the at least one set of concealment audio parameters;decoding, for a current properly received frame which immediately follows the at least one previously non-properly received frame, a second set of decoded audio parameters from at least the set of encoded audio parameters, the second set of audio parameters being different from the first set of audio parameters,outputting or deriving an output or derived version of the synthesized audio signal in such way that, if the current properly received frame immediately follows the at least one previously non-properly received frame, a selection is made between:a first version (s1) of the synthesized audio signal, synthesized from at least the set of encoded audio parametersa second version (s2) of the synthesized audio signal, synthesized from the second set of decoded audio parameters.
Priority Applications (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| EP24193936.2A EP4693281A1 (en) | 2024-08-09 | 2024-08-09 | Frame recovery after frame loss using lsf stabilization |
| PCT/EP2025/072601 WO2026033017A1 (en) | 2024-08-09 | 2025-08-06 | Frame recovery after frame loss using lsf stabilization |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| EP24193936.2A EP4693281A1 (en) | 2024-08-09 | 2024-08-09 | Frame recovery after frame loss using lsf stabilization |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP4693281A1 true EP4693281A1 (en) | 2026-02-11 |
Family
ID=92301014
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP24193936.2A Pending EP4693281A1 (en) | 2024-08-09 | 2024-08-09 | Frame recovery after frame loss using lsf stabilization |
Country Status (2)
| Country | Link |
|---|---|
| EP (1) | EP4693281A1 (en) |
| WO (1) | WO2026033017A1 (en) |
Citations (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2014130087A1 (en) * | 2013-02-21 | 2014-08-28 | Qualcomm Incorporated | Systems and methods for mitigating potential frame instability |
-
2024
- 2024-08-09 EP EP24193936.2A patent/EP4693281A1/en active Pending
-
2025
- 2025-08-06 WO PCT/EP2025/072601 patent/WO2026033017A1/en active Pending
Patent Citations (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2014130087A1 (en) * | 2013-02-21 | 2014-08-28 | Qualcomm Incorporated | Systems and methods for mitigating potential frame instability |
Non-Patent Citations (1)
| Title |
|---|
| "IEEE Standard for Advanced Mobile Speech and Audio", IEEE STANDARD, IEEE, PISCATAWAY, NJ, USA, 30 January 2017 (2017-01-30), pages 1 - 156, XP068113050 * |
Also Published As
| Publication number | Publication date |
|---|---|
| WO2026033017A1 (en) | 2026-02-12 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US9153237B2 (en) | Audio signal processing method and device | |
| US10964334B2 (en) | Audio decoder and method for providing a decoded audio information using an error concealment modifying a time domain excitation signal | |
| US7593852B2 (en) | Speech compression system and method | |
| US7590525B2 (en) | Frame erasure concealment for predictive speech coding based on extrapolation of speech waveform | |
| EP1363273B1 (en) | A speech communication system and method for handling lost frames | |
| US7260522B2 (en) | Gain quantization for a CELP speech coder | |
| EP1979895B1 (en) | Method and device for efficient frame erasure concealment in speech codecs | |
| EP1922718B1 (en) | Method and apparatus for coding an information signal using pitch delay contour adjustment | |
| EP3063760B1 (en) | Audio decoder and method for providing a decoded audio information using an error concealment based on a time domain excitation signal | |
| US7478042B2 (en) | Speech decoder that detects stationary noise signal regions | |
| US7324937B2 (en) | Method for packet loss and/or frame erasure concealment in a voice communication system | |
| EP2259255A1 (en) | Speech encoding method and system | |
| CN102985966B (en) | Audio coder and decoder and the method for the coding of audio signal and decoding | |
| KR20140005277A (en) | Apparatus and method for error concealment in low-delay unified speech and audio coding | |
| US6564182B1 (en) | Look-ahead pitch determination | |
| EP4693281A1 (en) | Frame recovery after frame loss using lsf stabilization |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE APPLICATION HAS BEEN PUBLISHED |
|
| AK | Designated contracting states |
Kind code of ref document: A1 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR |