US7983904B2 - Scalable decoding apparatus and scalable encoding apparatus - Google Patents
Scalable decoding apparatus and scalable encoding apparatus Download PDFInfo
- Publication number
- US7983904B2 US7983904B2 US11/718,437 US71843705A US7983904B2 US 7983904 B2 US7983904 B2 US 7983904B2 US 71843705 A US71843705 A US 71843705A US 7983904 B2 US7983904 B2 US 7983904B2
- Authority
- US
- United States
- Prior art keywords
- spectrum
- section
- decoding
- band
- decoded
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Expired - Fee Related, expires
Links
- 238000001228 spectrum Methods 0.000 claims abstract description 506
- 238000012545 processing Methods 0.000 claims description 42
- 238000001914 filtration Methods 0.000 claims description 29
- 230000005236 sound signal Effects 0.000 abstract description 4
- 230000015556 catabolic process Effects 0.000 abstract description 3
- 238000006731 degradation reaction Methods 0.000 abstract description 3
- 239000010410 layer Substances 0.000 description 107
- 238000000034 method Methods 0.000 description 41
- 238000010586 diagram Methods 0.000 description 35
- 239000013598 vector Substances 0.000 description 23
- 230000000873 masking effect Effects 0.000 description 21
- 238000004891 communication Methods 0.000 description 15
- 230000005540 biological transmission Effects 0.000 description 14
- 239000000872 buffer Substances 0.000 description 9
- 230000006870 function Effects 0.000 description 5
- 238000013139 quantization Methods 0.000 description 5
- 238000005516 engineering process Methods 0.000 description 4
- 238000010295 mobile communication Methods 0.000 description 4
- 230000002194 synthesizing effect Effects 0.000 description 4
- 238000004364 calculation method Methods 0.000 description 3
- 230000010354 integration Effects 0.000 description 3
- 239000012792 core layer Substances 0.000 description 2
- 230000006866 deterioration Effects 0.000 description 2
- 238000004458 analytical method Methods 0.000 description 1
- 230000015572 biosynthetic process Effects 0.000 description 1
- 230000003247 decreasing effect Effects 0.000 description 1
- 238000011161 development Methods 0.000 description 1
- 230000000694 effects Effects 0.000 description 1
- 238000004519 manufacturing process Methods 0.000 description 1
- 238000012986 modification Methods 0.000 description 1
- 230000004048 modification Effects 0.000 description 1
- 230000008447 perception Effects 0.000 description 1
- 239000004065 semiconductor Substances 0.000 description 1
- 238000003786 synthesis reaction Methods 0.000 description 1
- 238000012546 transfer Methods 0.000 description 1
Images
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L21/00—Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
- G10L21/02—Speech enhancement, e.g. noise reduction or echo cancellation
- G10L21/038—Speech enhancement, e.g. noise reduction or echo cancellation using band spreading techniques
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L27/00—Modulated-carrier systems
- H04L27/02—Amplitude-modulated carrier systems, e.g. using on-off keying; Single sideband or vestigial sideband modulation
- H04L27/06—Demodulator circuits; Receiver circuits
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L19/00—Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
- G10L19/04—Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis using predictive techniques
- G10L19/16—Vocoder architecture
- G10L19/18—Vocoders using multiple modes
- G10L19/24—Variable rate codecs, e.g. for generating different qualities using a scalable representation such as hierarchical encoding or layered encoding
-
- H—ELECTRICITY
- H03—ELECTRONIC CIRCUITRY
- H03M—CODING; DECODING; CODE CONVERSION IN GENERAL
- H03M7/00—Conversion of a code where information is represented by a given sequence or number of digits to a code where the same, similar or subset of information is represented by a different sequence or number of digits
- H03M7/30—Compression; Expansion; Suppression of unnecessary data, e.g. redundancy reduction
Definitions
- the present invention relates to a scalable decoding apparatus and scalable encoding apparatus used for carrying out communication using speech signals and audio signals in a mobile communication system and a packet communication system using Internet protocol.
- a speech coding scheme is required that can flexibly support communication between different networks, communication between terminals utilizing different services, communication between terminals having different processing performance, and conversational communication at multipoints as well as communication between two parties.
- a speech coding scheme is required to be robust against transmission path errors (in particular, packet loss in packet switching networks typified by IP networks).
- the bandwidth scalable coding scheme is a coding scheme that encodes speech signals in a layered way, and a coding scheme where coding quality increases in accordance with an increase in the number of coding layers.
- the bit rate can be set variable by increasing or decreasing the number of coding layers, so that it is possible to effectively use transmission path capacity.
- the bandwidth scalable speech coding scheme it is only necessary to receive at least the data coded by a base layer at a decoder side, and it is possible to allow to some extent information coded by additional layers being lost on the transmission path, and therefore the bandwidth scalable speech coding scheme provides robustness against transmission path errors.
- the frequency bandwidth of speech signals to be encoded also becomes wider in accordance with an increase in the number of coding layers. For example, for a base layer (i.e. core layer), a coding scheme for telephone band speech of the related art is used. Further, in additional layers (i.e. enhancement layers), layers are configured so that wideband speech which has a bandwidth such as 7 kHz can be encoded.
- the band scalable speech coding scheme telephone band speech signals are encoded in the core layer, and high-quality wideband signals are encoded in the enhancement layers, so that it is possible to utilize the bandwidth scalable speech coding scheme for both telephone band speech service terminals and high-quality wideband speech service terminals and support multipoint communication including the two kinds of terminals.
- the coded information is layered, so that it is possible to increase error robustness by devising a transmission method, and readily control the bit rate on the encoding side or on the transmission path. Therefore, the bandwidth scalable speech coding scheme draws attention as a speech coding scheme for future communication.
- non-patent document 1 The method disclosed in non-patent document 1 is given as an example of the bandwidth scalable speech coding scheme described above.
- MDCT coefficients are encoded using a scale factor and fine structure information for each band.
- the scale factor is Huffman encoded, and the fine structure is subjected to vector quantization. Auditory weighting of each band is calculated using a scale factor decoding result, and the bit allocation to each band is decided.
- the bandwidth of each band is non-uniform and set in advance so as to become wider for a higher band.
- transmission information is classified into four groups as described below.
- ⁇ Case 3> When information for B is received in addition to the information for A, a high band is generated by mirroring the decoded signal for the core codec and a decoded signal having a wider bandwidth than the decoded signal of the core codec is generated. Decoded information for B is used in generation of high band spectrum shapes. Mirroring is carried out at a voiced frame, and is carried out so that the harmonic structure does not collapse. The high band is generated at an unvoiced frame using random noise. ⁇ Case 4> When information for C is received in addition to information for A and B, the same decoding processing as in case 3 is carried out using only information for A and B.
- ⁇ Case 5> When information for D is received in addition to the information for A, B and C, complete decoding processing is carried out at bands where all information for A to D is received, and a fine spectrum is decoded by mirroring a decoded signal spectrum on the low band side at bands where information for D is not received. Even if the information for D is not received, it is possible to receive the information for B and C, and this information for B and C is utilized in decoding of spectrum envelope information. Mirroring is carried out at a voiced frame, and is carried out so that the harmonic structure does not collapse. The high band is generated at an unvoiced frame using random noise.
- Non-patent document 1 B. Kovesi et al, “A scalable speech and audio coding scheme with continuous bit rate flexibility,” in proc. IEEE ICASSP2004, pp. I-273--I-276.
- a high band is generated by mirroring. At this time, mirroring is carried out so that a harmonic structure does not collapse, so that this harmonic structure is maintained.
- the low band harmonic structure appears in the high band as a mirror image.
- a harmonic structure is more likely to collapse in the higher band, and therefore the harmonic structure does not appear more markedly at the high band than the low band.
- valleys of harmonics are deep at the low band, at the high band, valleys of harmonics are shallow, or, depending on the case, the harmonic structure itself becomes less defined. Therefore, with the technique of the related art described above, a harmonic structure excessively appears more easily at high band components, and therefore, the quality of the decoded speech signal deteriorates.
- a scalable decoding apparatus of the present invention adopts a configuration including: a first decoding section that decodes low frequency band coding information and obtains a low frequency band decoded signal; a second decoding section that obtains a high frequency band decoded signal from the low frequency band decoded signal and high frequency band coding information, wherein the second decoding section includes: a transform section that transforms the low frequency band decoded signal and obtains a low frequency band spectrum; an adjusting section that carries out amplitude adjustment on the low frequency band spectrum; and a generating section that generates a high frequency band spectrum in a pseudo manner using the amplitude-adjusted low frequency band spectrum and the high frequency band coding information.
- the preset invention it is possible to obtain a high quality decoded speech (or audio) signal with little deterioration in the high band spectrum even when the speech (or audio) signal is decoded by generating a high band spectrum using a low band spectrum.
- FIG. 1 is a block diagram showing a configuration of a scalable decoding apparatus according to Embodiment 1 of the present invention
- FIG. 2 is a block diagram showing a configuration of a scalable encoding apparatus according to Embodiment 1 of the present invention
- FIG. 3 is a block diagram showing a configuration of a second layer decoding section according to Embodiment 1 of the present invention.
- FIG. 4 is a block diagram showing a configuration of a second layer encoding section according to Embodiment 1 of the present invention.
- FIG. 5 is a block diagram showing a configuration of a spectrum decoding section according to Embodiment 1 of the present invention.
- FIG. 6 is a further block diagram showing a configuration of the spectrum decoding section according to Embodiment 1 of the present invention.
- FIG. 7 is another block diagram showing a configuration of the spectrum decoding section according to Embodiment 1 of the present invention.
- FIG. 8 is a further block diagram showing a configuration of the spectrum decoding section according to Embodiment 1 of the present invention.
- FIG. 9 is a still further block diagram showing a configuration of the spectrum decoding section according to Embodiment 1 of the present invention.
- FIG. 10 is yet another block diagram showing a configuration of the spectrum decoding section according to Embodiment 1 of the present invention.
- FIG. 11 is a schematic diagram showing processing of generating a high band component at a high band spectrum decoding section according to Embodiment 1 of the present invention.
- FIG. 12 is yet another block diagram showing a configuration of the spectrum decoding section according to Embodiment 1 of the present invention.
- FIG. 13 is another block diagram showing a configuration of the spectrum decoding section according to Embodiment 1 of the present invention.
- FIG. 14 is a block diagram showing a configuration of a second layer decoding section according to Embodiment 2 of the present invention.
- FIG. 15 is a block diagram showing a configuration of a second layer encoding section according to Embodiment 2 of the present invention.
- FIG. 16 is another block diagram showing a configuration of the spectrum decoding section according to Embodiment 2 of the present invention.
- FIG. 17 is yet another block diagram showing a configuration of the spectrum decoding section according to Embodiment 2 of the present invention.
- FIG. 18 is a block diagram showing a configuration of a first spectrum encoding section according to Embodiment 2 of the present invention.
- FIG. 19 is a block diagram showing a configuration of an extended band decoding section according to Embodiment 2 of the present invention.
- FIG. 20 is a further block diagram showing a configuration of the extended band decoding section according to Embodiment 2 of the present invention.
- FIG. 21 is a yet further block diagram showing a configuration of the extended band decoding section according to Embodiment 2 of the present invention.
- FIG. 22 is another block diagram showing a configuration of the extended band decoding section according to Embodiment 2 of the present invention.
- FIG. 23 is a schematic diagram showing processing of generating a high band component at a second extended band decoding section according to Embodiment 1 of the present invention.
- FIG. 24 is a block diagram showing a configuration of an extended band encoding section according to Embodiment 2 of the present invention.
- FIG. 25 is a schematic diagram showing the content of a bitstream received by the separating section of the scalable decoding apparatus according to Embodiment 2 of the present invention.
- FIG. 26 is a block diagram showing a configuration of an extended band decoding section according to Embodiment 3 of the present invention.
- FIG. 1 is a block diagram showing a configuration of a scalable decoding apparatus for forming, for example, a bandwidth scalable speech (or audio) signal decoding apparatus.
- Scalable decoding apparatus 100 includes separating section 101 , first layer decoding section 102 and second layer decoding section 103 .
- Separating section 101 receives a bitstream transmitted from the scalable encoding apparatus described later, separates the bitstream into a first layer coding parameter and a second layer coding parameter, and outputs the parameters respectively to first layer decoding section 102 and second layer decoding section 103 .
- First layer decoding section 102 decodes the first layer coding parameter inputted from separating section 101 and outputs a first layer decoded signal. This first layer decoded signal is also outputted to second layer decoding section 103 .
- Second layer decoding section decodes the second layer coding parameter inputted from separating section 101 using the first layer decoded signal inputted from first layer decoding section 102 and outputs a second layer decoded signal.
- FIG. 2 An example of a configuration of scalable encoding apparatus 200 corresponding to scalable decoding apparatus 100 of FIG. 1 is shown in FIG. 2 .
- first layer encoding section 201 encodes the inputted speech signal (i.e. original signal) and outputs the obtained parameters coded to first layer decoding section 202 and multiplexing section 203 .
- First layer decoding section 201 implements bandwidth scalability for the first and second layers by carrying out downsampling processing and low pass filtering processing for encoding.
- First layer decoding section 202 then generates a first layer decoded signal from coded parameters inputted from first layer encoding section 201 and outputs the first layer decoded signal to second layer encoding section 204 .
- Second layer encoding section 204 then encodes the inputted speech signal (i.e. original signal) using the first layer decoded signal inputted from first layer decoding section 202 and outputs the obtained parameter coded to multiplexing section 203 .
- Second layer encoding section 204 carries out upsampling processing of the first layer decoded signal and phase adjustment processing in order to match the phases of the first decoded signal and the inputted speech signal according to processing carried out at first layer encoding section 201 (such processing includes downsampling processing or low pass filtering processing) for encoding.
- Multiplexing section 203 then multiplexes the coding parameter inputted from first layer encoding section 201 and the coded parameter inputted from second layer encoding section 204 , and outputs the result as a bitstream.
- FIG. 3 is a block diagram showing a configuration of second layer decoding section 103 .
- Second layer decoding section 103 includes separating section 301 , scaling coefficient decoding section 302 , fine spectrum decoding section 303 , frequency domain transform section 304 , spectrum decoding section 305 and time domain transform section 306 .
- Separating section 301 separates the inputted second layer coded parameter into a coding parameter indicating scaling coefficients (i.e. scaling coefficient parameter) and a coding parameter indicating fine spectrum structure (i.e. fine spectrum parameter) and outputs the coding parameters to scaling coefficient decoding section 302 and fine spectrum decoding section 303 , respectively.
- a coding parameter indicating scaling coefficients i.e. scaling coefficient parameter
- a coding parameter indicating fine spectrum structure i.e. fine spectrum parameter
- Scaling coefficient decoding section 302 decodes the inputted scaling coefficient parameter so as to obtain low band scaling coefficients and high band scaling coefficients, and outputs the decoded scaling coefficients to spectrum decoding section 305 and fine spectrum decoding section 303 .
- Fine spectrum decoding section 303 calculates auditory weighting of each band using the scaling coefficients inputted from scaling coefficient decoding section 302 and obtains the number of bits allocated to fine spectrum information of each band. Fine spectrum decoding section 303 then decodes the fine spectrum parameters inputted from separating section 301 and obtains decoded fine spectrum information of each band, and outputs the result to spectrum decoding section 305 . It is also possible to use information for the first layer decoded signal in calculation of auditory weighting, and in this case, the output of frequency domain transform section 304 is also inputted to fine spectrum decoding section 303 .
- Frequency domain transform section 304 transforms the inputted first layer decoded signal to a frequency domain spectrum parameter (for example, MDCT coefficients), and outputs the result to spectrum decoding section 305 .
- a frequency domain spectrum parameter for example, MDCT coefficients
- Spectrum decoding section 305 decodes the second decoded signal from the first layer decoded signal which is inputted from frequency domain transform section 304 and transformed to a frequency domain, the decoding scaling coefficients (for low band and high band) inputted from scaling coefficient decoding section 302 , and the decoding fine spectrum information inputted from fine spectrum decoding section 303 , and outputs the result to time domain transform section 306 .
- Time domain transform section 306 transforms a spectrum of the second layer decoded signal inputted from spectrum decoding section 305 to a time domain signal and outputs the result as a second layer decoded signal.
- FIG. 4 An example of a configuration of second layer encoding section 204 corresponding to second layer decoding section 103 of FIG. 3 is shown in FIG. 4 .
- the inputted speech signal is inputted to auditory masking calculating section 401 and frequency domain transform section 402 A.
- Auditory masking calculating section 401 calculates auditory masking for each subband having a pre-defined bandwidth and outputs this auditory masking to scaling coefficient encoding section 403 and fine spectrum encoding section 404 .
- human auditory perception has auditory masking characteristics that, when a given signal is being heard, even if sound having a frequency close to that signal comes to the ear, the sound is difficult to be heard. It is therefore possible to implement efficient spectrum encoding by, using the auditory masking based on this auditory masking characteristic, allocating a small number of quantization bits to a frequency spectrum where quantization distortion is difficult to be perceived and allocating a large number of quantization bits to a frequency spectrum where quantization distortion is easy to be perceived.
- Frequency domain transform section 402 A transforms the inputted speech signal to a frequency domain spectrum parameter (for example, MDCT coefficients) and outputs the result to scaling coefficient encoding section 403 and spectrum encoding section 404 .
- Frequency domain transform section 402 B transforms the inputted first layer decoded signal to a frequency domain spectrum parameter (for example, MDCT coefficients) and outputs the result to scaling coefficient encoding section 403 and spectrum encoding section 404 .
- Scaling coefficient encoding section 403 encodes a differential spectrum between the spectrum parameter inputted from frequency domain transform section 402 A and the first layer decoded spectrum inputted from frequency domain transform section 402 using the auditory masking information inputted from auditory masking calculating section 401 , obtains a scaling coefficient parameter, and outputs the scaling coefficient parameter to coded parameter multiplexing section 405 and fine spectrum encoding section 404 .
- a high band spectrum scaling coefficient parameter and a low band spectrum scaling coefficient parameter are outputted separately.
- Fine spectrum encoding section 404 decodes the scaling coefficient parameter (for low band and high band) inputted from scaling coefficient encoding section 403 , obtains decoded scaling coefficients (for low band and high band), and normalizes a differential spectrum between the spectrum parameter inputted from frequency domain transform section 402 A and the first layer decoded spectrum inputted from frequency domain transform section 402 B using decoded scaling coefficients (for low band and high band).
- Fine spectrum encoding section 404 encodes the normalized differential spectrum, and outputs the differential spectrum after encoding (i.e. fine spectrum coding parameters) to coding parameter multiplexing section 405 .
- fine spectrum encoding section 404 calculates perceptual importance of fine spectrum in each band using decoded scaling coefficients (for low band and high band) and defines bit allocation according to the perceptual importance. It is also possible to calculate this perceptual importance using the first layer decoded spectrum.
- Coded parameter multiplexing section 405 multiplexes the high band spectrum scaling coefficient parameter and the low band spectrum scaling coefficient parameter inputted from scaling coefficient encoding section 403 and the fine spectrum coding parameter inputted from fine spectrum encoding section 404 and outputs the result as a first spectrum coded parameter.
- FIG. 5 to FIG. 9 are block diagrams showing a configuration of spectrum decoding section 305 .
- FIG. 5 shows a configuration for executing processing when the first layer decoded signal, all decoding scaling coefficients (for low band and high band) and all fine spectrum decoding information are received normally.
- FIG. 6 shows a configuration for executing processing when part of the fine spectrum decoding information for the high band is not received.
- FIG. 6 differs from FIG. 5 in that the output result of adder A is inputted to high band spectrum decoding section 602 .
- a spectrum for the bands to be decoded using high band fine spectrum decoding information that is not received is generated in a pseudo manner using the following method.
- FIG. 7 shows a configuration for executing processing when none of the high band fine spectrum decoding information is received (including the case where part of the low band fine spectrum decoding information is not received). This differs from FIG. 6 in that the fine spectrum decoding information is not inputted to high band spectrum decoding section 702 .
- a spectrum for the bands to be decoded using high band fine spectrum decoding information that is not received is generated in a pseudo manner using the following method.
- FIG. 8 shows a configuration for executing processing when none of the fine spectrum decoding information is not received, and further, part of the low band decoding scaling coefficients is not received. This differs from FIG. 7 in that fine spectrum decoding information is not inputted, there is no output from low band spectrum decoding section 801 , and adder A does not exist.
- a spectrum for bands to be decoded using high band fine spectrum decoding information that is not received is generated in a pseudo manner using the following method.
- FIG. 9 shows a configuration for executing processing when only high band decoding scaling coefficients are received (including the case where part of the high band decoding scaling coefficients is not received). This differs from FIG. 8 in that there is no input of the low band decoding scaling coefficients, and low band spectrum decoding section does not exist. A method for generating a high band spectrum in a pseudo manner from only the received high band decoding scaling coefficients will be described later.
- Spectrum decoding section 305 of FIG. 5 is provided with low band spectrum decoding section 501 , high band spectrum decoding section 502 , adder A and adder B.
- Low band spectrum decoding section 501 decodes the low band spectrum using low band decoding scaling coefficients inputted from scaling coefficient decoding section 302 and fine spectrum decoding information inputted from fine spectrum decoding section 303 , and outputs the result to adder A.
- a decoded spectrum is calculated by multiplying fine spectrum decoding information by decoding scaling coefficients.
- Adder A adds the decoded low band (residual) spectrum inputted from low band spectrum decoding section 501 and the first layer decoded signal (i.e. spectrum) inputted from frequency domain transform section 302 so as to obtain a decoded low band spectrum and outputs the result to adder B.
- High band spectrum decoding section 502 decodes the high band spectrum using the high band decoding scaling coefficients inputted from scaling coefficient decoding section 302 and the fine spectrum decoding information inputted from fine spectrum decoding section 303 , and outputs the result to adder B.
- Adder B adds the decoded low band spectrum inputted from adder A and the decoded high band spectrum inputted from high band spectrum decoding section 502 so as to generate a spectrum for all bands (all frequency bands combining the low band and the high band), and outputs the result as a decoded spectrum.
- FIG. 6 differs from FIG. 5 in only the operation of high band spectrum decoding section 602 .
- High band spectrum decoding section 602 decodes the high band spectrum using the high band decoding scaling coefficients inputted from scaling coefficient decoding section 302 and the high band fine spectrum decoding information inputted from fine spectrum decoding section 303 . At this time, the high band fine spectrum decoding information for part of the band is not received, and therefore the high band spectrum of the corresponding band cannot be decoded accurately. High band spectrum decoding section 602 then generates a high band spectrum in a pseudo manner using the decoded scaling coefficients, the low band decoded spectrum inputted from adder A, and the high band spectrum capable of being received and accurately decoded. A specific generating method is described in the following.
- FIG. 7 shows operation in FIG. 5 and FIG. 6 for the case where all high band fine spectrum decoding information is not received.
- high band spectrum decoding section 702 decodes high band spectrum using just high band decoding scaling coefficients inputted from scaling coefficient decoding section 302 .
- low band spectrum decoding section 701 decodes the high band spectrum using the low band decoding scaling coefficients inputted from scaling coefficient decoding section 302 and the low band fine spectrum decoding information inputted from fine spectrum decoding section 303 .
- the low band fine spectrum decoding information for some part of the band is not received. Therefore, this part of the band is not subjected to decoding processing and taken to be a zero spectrum.
- a spectrum of the corresponding band outputted via adders A and B is the first layer decoded signal (spectrum) itself.
- FIG. 8 shows operation for the case where all low band fine spectrum decoding information is not received in FIG. 7 .
- Low band spectrum decoding section 801 receives low band decoding scaling coefficients, but does not receive fine spectrum decoding information at all, and decoding processing is therefore not carried out.
- FIG. 9 shows the operation for the case where decoding scaling coefficients for the low band are not inputted at all in FIG. 8 .
- the spectrum for this band is outputted as zero.
- FIG. 9 shows the configuration of high band spectrum decoding section 902 in more detail.
- High band spectrum decoding section 902 of FIG. 10 includes amplitude adjusting section 1011 , pseudo spectrum generating section 1012 and scaling section 1013 .
- Amplitude adjusting section 1011 adjusts the amplitude of the first layer decoded signal spectrum inputted from frequency domain transform section 302 and outputs the result to pseudo spectrum generating section 1012 .
- Pseudo spectrum generating section 1012 generates a high band spectrum in a pseudo manner using the first layer decoded signal spectrum, whose amplitudes are adjusted, inputted from amplitude adjusting section 1011 , and outputs the result to scaling section 1013 .
- Scaling section 1013 scales the spectrum inputted from pseudo spectrum generating section 1012 and outputs the result to adder B.
- FIG. 11 is a schematic diagram showing an example of a series of processing of generating a high band spectrum in a pseudo manner.
- This amplitude adjustment method may be, for example, a constant multiple in a logarithmic domain ( ⁇ S, where ⁇ is an amplitude adjustment coefficient (real number) in the range of 0 ⁇ 1, and S is a logarithmic spectrum), or may be a constant ⁇ -th power (where s ⁇ , s are linear spectrum) in a linear domain.
- ⁇ S logarithmic domain
- ⁇ -th power where s ⁇ , s are linear spectrum
- the adjusting coefficients may be a fixed constant, but it is also possible to prepare a plurality of appropriate adjustment coefficients according to an index (for example, directly, a variance of spectrum amplitude occurring at a low band, or indirectly, a value of pitch gain occurring at first layer encoding section 201 ) indicating a depth of valleys of low band spectrum harmonics, and selectively use corresponding adjustment coefficients according to the index. Further, it is also possible to selectively use adjustment coefficients according to characteristics for each vowel using low band spectrum shape (envelope) information and pitch period information. Further, it is also possible to encode optimum adjustment coefficients on the encoder side as separate transmission information and transmit the encoded information.
- an index for example, directly, a variance of spectrum amplitude occurring at a low band, or indirectly, a value of pitch gain occurring at first layer encoding section 201
- an index for example, directly, a variance of spectrum amplitude occurring at a low band, or indirectly, a value of pitch gain occurring at first layer encoding section
- a high band spectrum is generated in a pseudo manner using the amplitude adjusted spectrum.
- a generating method an example of mirroring that generates a high band spectrum as a low band mirror image is shown in FIG. 11 .
- FIG. 12 shows the case where first layer spectrum information (for example, decoding LSP parameters) is inputted to amplitude adjusting section 1211 from first layer decoding section 102 .
- amplitude adjusting section 1211 decides adjustment coefficients using amplitude coefficients based on the inputted first layer spectrum information.
- First layer pitch information i.e. pitch period and/or pitch gain
- other than the first layer spectrum information may also be used to decide the adjustment coefficients.
- FIG. 13 shows the case where an amplitude adjustment coefficient is inputted separately.
- the amplitude adjustment coefficient is quantized and encoded on the encoder side, and then transmitted.
- FIG. 14 is a block diagram showing a configuration of second layer decoding section 103 according to Embodiment 2 of the present invention.
- Second layer decoding section 103 of FIG. 14 includes separating section 1401 , spectrum decoding section 1402 , extended band decoding section 1403 , spectrum decoding section 1402 B, frequency domain transform section 1404 and time domain transform section 1405 .
- Separating section 1401 separates the second layer coded parameter into a first spectrum coded parameter, an extended band coded parameter and a second spectrum coded parameter, and outputs those parameters to spectrum decoding section 1402 A, extended band decoding section 1403 and spectrum decoding section 1402 B, respectively.
- Frequency domain transform section 1404 transforms a first layer decoded signal inputted from first layer decoding section 102 to a frequency domain parameter (for example, MDCT coefficients) and outputs the result to first spectrum decoding section 1402 A as a first layer decoded signal spectrum.
- a frequency domain parameter for example, MDCT coefficients
- Spectrum decoding section 1402 A adds a quantized spectrum of the first layer coding errors, which is obtained by decoding the first spectrum coded parameter inputted from separating section 1401 , to the first layer decoded signal spectrum inputted from frequency domain transform section 1404 , and outputs the result to extended band decoding section 1403 as the first decoded spectrum.
- the first layer coding errors are improved mainly for the low band component at spectrum decoding section 1402 A.
- Extended band decoding section 1403 decodes various parameters from the extended band coded parameter inputted from separating section 1401 and decodes/generates a high band spectrum using the various decoded parameters based on the first decoded spectrum inputted from spectrum decoding section 1402 A. Extended band decoding section 1403 then outputs a spectrum for the whole band to spectrum decoding section 1402 B as the second decoded spectrum.
- Spectrum decoding section 1402 B adds a spectrum, which is the quantized coding errors of the second decoded spectrum obtained by decoding the second spectrum coded parameter inputted from separating section 1401 , to the second decoded spectrum inputted from extended band decoding section 1403 , and outputs the result to time domain transform section 1405 as the third decoded spectrum.
- Time domain transform section 1405 transforms the third decoded spectrum inputted from spectrum decoding section 1402 B to a time domain signal and outputs the result as a second layer decoded signal.
- FIG. 14 it is also possible to adopt a configuration where one or both of spectrum decoding section 1402 A and spectrum decoding section 1402 B are not present.
- the first layer decoded signal spectrum outputted from frequency domain transform section 1404 is inputted to extended band decoding section 1403 .
- the second decoded spectrum outputted from extended band decoding section 1403 is inputted to time domain transform section 1405 .
- FIG. 15 An example of a configuration of second layer encoding section 204 corresponding to second layer decoding section 103 of FIG. 14 is shown in FIG. 15 .
- the speech signal (i.e. original signal) is inputted to auditory masking calculating section 1501 and frequency domain transform section 1502 A.
- Auditory masking calculating section 1501 calculates auditory masking using the inputted speech signal and outputs the auditory masking to first spectrum encoding section 1503 , extended band encoding section 1504 and second spectrum encoding section 1505 .
- Frequency domain transform section 1502 A transforms the inputted speech signal to a frequency domain spectrum parameter (for example, MDCT coefficients), and outputs the result to first spectrum encoding section 1503 , extended band encoding section 1504 and second spectrum encoding section 1505 .
- a frequency domain spectrum parameter for example, MDCT coefficients
- Frequency domain transform section 1502 B transforms the inputted first layer decoded signal to a spectrum parameter such as MDCT and outputs the result to first spectrum encoding section 1503 .
- First spectrum encoding section 1503 encodes a differential spectrum between the input speech signal spectrum, which is inputted from frequency domain transform section 1502 , and the first layer decoded spectrum, which is inputted from frequency domain transform section 1502 B, using the auditory masking inputted from auditory masking calculating section 1501 , outputs the result as a first spectrum coded parameter, and outputs a first decoded spectrum obtained by decoding the first spectrum coded parameter to extended band encoding section 1504 .
- Extended band encoding section 1504 encodes an error spectrum between the input speech signal spectrum, which is inputted from frequency domain transform section 1502 A, and the first decoded spectrum, which is inputted from first spectrum encoding section 1503 , using the auditory masking inputted from auditory masking calculating section 1501 , outputs the result as an extended band coding parameter, and outputs the second decoded spectrum obtained by decoding the extended band coded parameter to second spectrum encoding section 1505 .
- Second spectrum encoding section 1505 encodes an error spectrum between the input speech signal spectrum, which is inputted from frequency domain transform section 1502 A, and the second decoded spectrum, which is inputted from extended band encoding section 1504 , using the auditory masking inputted from auditory masking calculating section 1501 , and outputs the result as a second spectrum coded parameter.
- separating section 1601 separates the inputted coding parameter into a coding parameter indicating scaling coefficients (i.e. scaling coefficient parameter) and a coding parameter indicating a spectrum fine structure (i.e. fine spectrum parameter), and outputs the parameters to scaling coefficient decoding section 1602 and fine spectrum decoding section 1603 , respectively.
- a coding parameter indicating scaling coefficients i.e. scaling coefficient parameter
- a coding parameter indicating a spectrum fine structure i.e. fine spectrum parameter
- Scaling coefficient decoding section 1602 decodes the inputted scaling coefficient parameter so as to obtain low band scaling coefficients and high band scaling coefficients, outputs the decoding scaling coefficients to spectrum decoding section 1604 and fine spectrum decoding section 1603 .
- Fine spectrum decoding section 1603 calculates auditory weighting of each band using the scaling coefficients inputted from scaling coefficient decoding section 1602 and obtains the number of bits allocated to fine spectrum information of each band. Fine spectrum decoding section 1603 then decodes the fine spectrum parameter inputted from separating section 1601 and obtains decoded fine spectrum information of each band, and outputs the decoded fine spectrum information to spectrum decoding section 1604 . It is also possible to use information for decoded spectrum A in calculation of auditory weighting. In this case, a configuration is adopted so that decoded spectrum A is inputted to fine spectrum decoding section 1603 .
- Spectrum decoding section 1604 then decodes decoded spectrum B from inputted decoded spectrum A, decoded scaling coefficients (for low band and high band) inputted from scaling coefficient decoding section 1602 , and decoded fine spectrum information inputted from fine spectrum decoding section 1603 .
- FIG. 16 The relationship of correspondence between FIG. 16 and FIG. 14 is described as follows.
- the coding parameter of FIG. 16 corresponds to the first spectrum coded parameter of FIG. 14
- decoded spectrum A of FIG. 16 corresponds to the first layer decoded signal spectrum of FIG. 14
- decoded spectrum B of FIG. 16 corresponds to the first decoded spectrum of FIG. 14 .
- the coding parameter of FIG. 16 corresponds to the second spectrum coded parameter of FIG. 14
- decoded spectrum A of FIG. 16 corresponds to the second decoded spectrum of FIG. 14
- decoded spectrum B of FIG. 16 corresponds to the third decoded spectrum of FIG. 14 .
- FIG. 18 An example of configuration of first spectrum encoding section 1503 , which is corresponding to spectrum decoding sections 1402 A and 1402 B of FIG. 16 , is shown in FIG. 18 .
- First spectrum encoding section 1503 shown in FIG. 18 is configured with scaling coefficient encoding section 403 , fine spectrum encoding section 404 , coding parameter multiplexing section 405 shown in FIG. 4 and spectrum decoding section 1604 shown in FIG. 16 .
- the operation thereof is the same as described in FIG. 4 and FIG. 16 , and therefore the description thereof will be omitted here. Further, if the first layer decoded spectrum of FIG.
- FIG. 18 is replaced with the second decoded spectrum, and the first spectrum coded parameter is replaced with the second spectrum coded parameter, the configuration shown in FIG. 18 is a configuration of second spectrum encoding section 1505 in FIG. 15 .
- Spectrum decoding section 1604 can be eliminated in the configuration of second spectrum encoding section 1505 .
- FIG. 17 shows a configuration of spectrum decoding sections 1402 A and 1402 B in the case of not using scaling coefficients.
- spectrum decoding sections 1402 A and 1402 B include auditory weighting and bit allocation calculating section 1701 , fine spectrum decoding section 1702 and spectrum decoding section 1703 .
- auditory weighting and bit allocation calculating section 1701 obtains auditory weighting of each band from inputted decoded spectrum A, and obtains bit allocation to each band decided according to the auditory weighting. Information of the obtained auditory weighting and bit allocation is then outputted to fine spectrum decoding section 1702 .
- Fine spectrum decoding section 1702 then decodes inputted coded parameters based on the auditory weighting and bit allocation information, which are inputted from auditory weighting and bit allocation calculating section 1701 , and obtains decoded fine spectrum information of each band, and outputs the decoded fine spectrum information to spectrum decoding section 1703 .
- Spectrum decoding section 1703 then adds fine spectrum decoding information, which is inputted from fine spectrum decoding section 1702 , to inputted decoded spectrum A, and outputs the result as decoded spectrum B.
- FIG. 17 The relationship of correspondence between FIG. 17 and FIG. 14 is described as follows.
- the coding parameter of FIG. 17 corresponds to the first spectrum coded parameter of FIG. 14
- decoded spectrum A of FIG. 17 corresponds to the first layer decoded signal spectrum of FIG. 14
- decoded spectrum B of FIG. 17 corresponds to the first decoded spectrum of FIG. 14 .
- the coding parameter of FIG. 17 corresponds to the second spectrum coded parameter of FIG. 14
- decoded spectrum A of FIG. 17 corresponds to the second decoded spectrum of FIG. 14
- decoded spectrum B of FIG. 17 corresponds to the third decoded spectrum of FIG. 14 .
- first spectrum encoded section corresponding to spectrum decoding sections 1402 A and 1402 B of FIG. 17 .
- extended band decoding section 1403 shown in FIG. 14 will be described using FIG. 19 to FIG. 23 .
- FIG. 19 is a block diagram showing a configuration of extended band decoding section 1403 .
- extended band decoding section 1403 includes separating section 1901 , amplitude adjustment section 1902 , filter state setting section 1903 , filtering section 1904 , residual spectrum shape codebook 1905 , residual spectrum gain codebook 1906 , multiplier 1907 , scale factor decoding section 1908 , scaling section 1909 and spectrum synthesizing section 1910 .
- Separating section 1901 separates the coded parameter inputted from separating section 1401 of FIG. 14 into an amplitude adjustment coefficient coding parameter, a lag coding parameter, a residual shape coding parameter, a residual gain coding parameter and a scale factor coding parameter, and outputs the parameters to amplitude adjusting section 1902 , filtering section 1904 , residual spectrum shape codebook 1905 , residual spectrum gain codebook 1906 and scale factor decoding section 1908 , respectively.
- Amplitude adjusting section 1902 decodes the coded amplitude adjustment coefficient parameter inputted from separating section 1901 , adjusts the amplitude of the first layer decoded spectrum inputted from spectrum decoding section 1402 of FIG. 14 , and outputs a first decoded spectrum, whose amplitude is adjusted, to filter state setting section 1903 .
- Amplitude adjustment is carried out using a method expressed by ⁇ S(n) ⁇ ⁇ , when, for example, the first decoded spectrum is assumed to be S(n), and the amplitude adjustment coefficient is assumed to be ⁇ .
- S(n) is spectrum amplitude in the linear domain
- n is a frequency.
- z is a variable occurring in z transform.
- z ⁇ 1 is a complex variable referred to as a delay operator.
- T is a lag for the pitch filter
- Nn is the number of valid spectrum points for the first decoded spectrum (corresponding to the upper limit frequency of the spectrum used as a filter state)
- Nw is the number of spectrum points after bandwidth extention, and a spectrum with (Nw-Nn) points is generated by this filtering processing.
- g indicates residual spectrum gain
- C[n] indicates a residual spectrum shape vector
- gC[n] is inputted from multiplier 1907 .
- Generated S[Nn to Nw] is outputted to scaling section 1909 .
- Residual spectrum shape codebook 1905 decodes the residual shape coding parameter inputted from separating section 1901 and outputs a residual spectrum shape vector corresponding to the decoding result to multiplier 1907 .
- Residual spectrum gain codebook 1906 decodes the residual gain coding parameter inputted from separating section 1901 and outputs residual gain corresponding to the decoding result to multiplier 1907 .
- Multiplier 1907 outputs result gC[n] of multiplying residual spectrum shape vector C[n] inputted from residual spectrum shape codebook 1905 by residual gain g inputted from residual spectrum gain codebook 1906 to filtering section 1904 .
- Scale factor decoding section 1908 decodes the scale factor coding parameter inputted from separating section 1901 and outputs the decoded scale factor to scaling section 1909 .
- Scaling section 1909 multiplies the scale factor inputted from scale factor decoding section 1908 by spectrum S[Nn to Nw] inputted from filtering section 1904 , and outputs the result to spectrum synthesizing section 1910 .
- Spectrum synthesizing section 1910 substitutes the first decoded spectrum inputted from spectrum decoding section 1402 A of FIG. 14 for the low band (S[0 to Nn]) and substitutes the spectrum inputted from scaling section 1909 for the high band (S[Nn to Nw]) and outputs the obtained spectrum to spectrum decoding section 1402 B of FIG. 14 as a second decoded spectrum.
- FIG. 20 a configuration of extended band decoding section 403 for the case where the spectrum differential shape coding parameter and the residual spectrum gain coding parameter cannot be received completely is shown in FIG. 20 .
- information of a coded parameter for an amplitude adjustment coefficient, a coded lag parameter and a coded scale factor parameter can be received completely.
- FIG. 20 the configuration other than for separating section 2001 and filtering section 2002 is the same as for each part of FIG. 19 and is therefore not described.
- separating section 2001 separates the coded parameter inputted from separating section 1401 of FIG. 14 into an amplitude adjustment coefficient parameter, a coded lag parameter and a coded scale factor parameter, and outputs those parameters to amplitude adjusting section 1902 , filtering section 2002 and scale factor decoding section 1908 , respectively.
- FIG. 21 a configuration of extended band decoding section 1403 for the case where the coded lag parameter can also not be received is shown in FIG. 21 .
- information of a coded parameter for an amplitude adjustment coefficient and a coded scale factor parameter can be received completely.
- filter state setting section 1903 of FIG. 20 and filtering section 2002 are substituted with pseudo spectrum generating section 2102 .
- the configuration other than for separating section 2101 and pseudo spectrum generating section 2102 is the same as for each part of FIG. 19 and is therefore not described.
- separating section 2101 separates the coding parameter inputted from separating section 1401 of FIG. 14 into a coded amplitude adjustment coefficient parameter and a coded scale factor parameter, and outputs those parameters to amplitude adjusting section 1902 and scale factor decoding section 1908 , respectively.
- Pseudo spectrum generating section 2102 generates a high band spectrum in a pseudo manner using the first decoded signal spectrum, whose amplitude is adjusted, inputted from amplitude adjusting section 1902 , and outputs the spectrum to scaling section 1909 .
- a specific method of generating a high band spectrum there are a method based on mirroring that generates a high band spectrum as a mirror image of a low band spectrum, a method of shifting the amplitude adjusted spectrum in a high band direction of the frequency axis, and a method of carrying out pitch filtering processing in a frequency axis direction on the amplitude adjusted spectrum using the pitch lag obtained from a low band spectrum. It is also possible to generate a pseudo spectrum using a noise spectrum generated in a random manner when decoded frames are determined to be unvoiced frames.
- FIG. 22 a configuration of extended band decoding section 1403 for the case where amplitude adjustment information can also not be received is shown in FIG. 22 .
- information of a coded scale factor parameter can be received completely.
- the configuration other than for separating section 2201 and pseudo spectrum generating section 2202 is the same as for each part of FIG. 19 and is therefore not described.
- separating section 2201 separates the coded scale factor parameter from the coded parameter inputted from separating section 1401 of FIG. 14 , and outputs the parameter to scale factor decoding section 1908 .
- Pseudo spectrum generating section 2202 generates a high band spectrum in a pseudo manner using the first decoded signal spectrum and outputs the spectrum to scaling section 1909 .
- a specific method of generating a high band spectrum there are a method based on mirroring that generates a high band spectrum as a mirror image of a low band spectrum, a method of shifting the amplitude adjusted spectrum in a high band direction of the frequency axis, and a method of carrying out pitch filtering processing in a frequency axis direction on the amplitude adjusted spectrum using the pitch lag obtained from a low band spectrum. It is also possible to generate a pseudo spectrum using noise spectrum generated in a random manner when decoded frames are determined to be unvoiced frames.
- the amplitude adjustment method may be, for example, a constant multiple in a logarithmic domain ( ⁇ S, where S is a logarithmic spectrum), or may be a constant ⁇ -th power (where s ⁇ , s are linear spectrum) in a linear domain.
- ⁇ S logarithmic domain
- ⁇ -th power where s ⁇ , s are linear spectrum
- coefficients typified by coefficients necessary in fitting the depth of valleys of harmonics occurring at a low band in a voiced speech to the depth of valleys of harmonics occurring at a high band as adjusting coefficients for amplitude adjustment.
- the adjusting coefficients may be a fixed constant, but it is also possible to prepare a plurality of appropriate adjusting coefficients according to an index (for example, directly, a variance value of a spectrum amplitude occurring at a low band, or indirectly, a value of pitch gain occurring at first layer encoding section 201 ) indicating a depth of valleys of low band spectrum harmonics, and selectively use the corresponding adjustment coefficients according to the index. Further, it is also possible to selectively use adjusting coefficients according to characteristics for each vowel using low band spectrum shape (envelope) information and pitch period information. More specifically, this is the same as the generation of pseudo spectrum described in Embodiment 1 and is therefore not described here.
- FIG. 23 is a schematic diagram showing a series of operations for generating a high band component in the configuration of FIG. 20 .
- amplitude adjustment of the first decoded spectrum is carried out.
- filtering processing pitch filtering
- a high band component is generated.
- scaling is carried out on the generated high band component for each scaling coefficient band so as to finally generate a high band spectrum.
- the second decoded spectrum is then generated by combining the generated high band spectrum and first decoded spectrum.
- FIG. 24 An example of a configuration for extended band encoding section 1504 corresponding to extended band decoding section 1403 of FIG. 19 is shown in FIG. 24 .
- amplitude adjusting section 2401 carries out amplitude adjustment of the first decoded spectrum inputted from first spectrum encoding section 1503 using the input speech signal spectrum inputted from frequency domain transform section 1502 A, outputs a coded parameter for the amplitude adjustment coefficient, and outputs first decoded spectrum, whose amplitude is adjusted, to filter state setting section 2402 .
- Amplitude adjusting section 2401 carries out amplitude adjustment processing so that the ratio of the maximum amplitude spectrum of the first decoded spectrum to the minimum amplitude spectrum (i.e. dynamic range) is approximated to the dynamic range of the high band of the input speech signal spectrum.
- an amplitude adjusting method there is the above-described method.
- S 1 is a spectrum before transform
- S 1 ′ is a spectrum after transform.
- amplitude adjusting section 2401 selects an amplitude adjustment coefficient ⁇ from a plurality of candidates prepared in advance so that the dynamic range of the first decoded spectrum, whose amplitude is adjusted, becomes closest to the dynamic range of the high band of the input speech signal spectrum, and outputs the coding parameter for the selected amplitude adjustment coefficient ⁇ to multiplexing section 203 .
- Filter state setting section 2402 sets the first decoded spectrum, whose amplitude is adjusted, inputted from amplitude adjusting section 2401 to the internal state of the pitch filter as with the filter state setting section 1903 of FIG. 19 .
- Lag setting section 2403 outputs lag T sequentially to filtering section 2404 while gradually changing lag T within a predetermined search range TMIN to TMAX.
- Residual spectrum shape codebook 2405 stores a plurality of residual shape vector candidates and sequentially selects and outputs residual spectrum shape vectors from all candidates or candidates restricted in advance, according to the instruction from search section 2406 .
- residual spectrum gain codebook 2407 stores a plurality of residual vector gain candidates and sequentially selects and outputs the residual spectrum vector gains from all candidates or candidates restricted in advance, according to the instruction from search section 2406 .
- Candidates for residual shape vectors outputted from residual spectrum shape codebook 2405 and candidates for residual spectrum gains outputted from residual spectrum gain codebook 2407 are multiplied by multiplying section 2408 , and the multiplication result is outputted to filtering section 2404 .
- Filtering section 2404 then carries out filtering processing using the internal state of the pitch filter set at filter state setting section 2402 , lag T outputted from lag setting section 2403 , and gain-adjusted residual spectrum shape vectors, and calculates an estimation value for the input speech signal spectrum. This operation is the same as the operation of filtering section 1904 of FIG. 19 .
- Search section 2406 decides a combination where the cross-correlation between the high band of the input speech signal spectrum (i.e. original spectrum) and the output signal of filtering section 240 becomes a maximum out of a plurality of combinations of lags, residual spectrum shape vectors and residual spectrum gains, using analysis by synthesis (AbS). At this time, the combination that gives the closest one from an auditory point of view is decided utilizing auditory masking. Further, searching is also carried out taking into consideration scaling carried out by a scale factor at a later stage. A coded parameter of lags decided by search section 2406 , coded parameter for residual spectrum shape vectors, and coding parameter for residual spectrum gains are outputted to multiplexing section 203 and extended band decoding section 2409 .
- Extended band decoding section 2409 then carries out decoding processing on the first decoded spectrum using the coded parameter for an amplitude adjustment coefficient outputted from amplitude adjusting section 2401 , the coded lag parameter outputted from search section 2406 , the coded parameter for residual spectrum shape vectors and coded parameter for residual spectrum gains, generates an estimated spectrum (that is, spectrum before scaling) for the input speech signal spectrum, and outputs the spectrum to scale factor encoding section 2410 .
- the decoding procedure is the same as for extended band decoding section 1403 of FIG. 19 (however, processing for scaling section 1909 and spectrum synthesizing section 1910 is eliminated).
- Scale factor encoding section 2410 encodes the scale factor (i.e. scaling coefficients) of the estimated spectrum most appropriate from a perceptual point of view using the high band of the input speech signal spectrum (i.e. original spectrum) outputted from frequency domain transform section 1502 A, the estimated spectrum outputted from extended band decoding section 2409 , and auditory masking, and outputs the coding parameter to multiplexing section 203 .
- scale factor i.e. scaling coefficients
- FIG. 25 is a schematic diagram showing content of a bitstream received by separating section 101 of FIG. 1 .
- bitstreams a plurality of coding parameters are time-multiplexed.
- the MSB Most Significant Bit, the most significant bit in the bitstream
- the LSB Least Significant Bit, the least significant bit in the bitstream
- the first layer decoded signal is assumed to be an output signal.
- the method for implementing the network where coding parameters are discarded preferentially in order from the LSB side is by no means limited.
- FIG. 19 a configuration is shown provided with residual spectrum shape codebook 1905 , residual spectrum gain codebook 1906 and multiplier 1907 , but a configuration where these are not adopted is also possible.
- the encoder side is capable of carrying out communication at a low bit rate without transmitting the coded residual shape vector parameter and the coded residual gain parameter.
- the decoding processing procedure in this case differs from the description using FIG. 19 in that there is no decoding processing of the residual spectrum information (shape, gain). Namely, a processing procedure is described using FIG. 20 , but the bitstream is such that the position of ( 1 ) in FIG. 25 is the LSB.
- a decoding parameter for the corresponding frame is decided using the decoding parameter decoded by the extended band coded parameters on both of the frame and the previous frame and data loss information for the received bitstream on the frame, and the second decoded spectrum is decoded.
- FIG. 26 is a block diagram showing a configuration of extended band decoding section 1403 according to Embodiment 3 of the present invention.
- amplitude adjustment coefficient decoding section 2601 decodes an amplitude adjustment coefficient from the coded amplitude adjustment coefficient parameter.
- Lag decoding section 2602 decodes a lag from the coded lag parameter.
- Decoding parameter control section 2603 decides a decoded parameter used in decoding of the second decoded spectrum of the frame, using each decoded parameter decoded by the extended band coded parameter, received data loss information and each decoded parameter of the previous frame outputted from each buffer 2604 a to 2604 e .
- Buffers 2604 a to 2604 e are buffers for storing decoded parameters on the frame, those are amplitude adjustment coefficient(s), lag(s), residual shape vector(s), residual spectrum gain(s) and scale factor(s). Other aspects of the configuration in FIG. 26 are the same as the configuration of extended band decoding section 1403 of FIG. 19 and are therefore not described.
- the decoding parameters included in the extended band coded parameters that are part of the second layer coded data of the frame are decoded by decoding sections 1908 , 2602 , 2601 , 1905 and 1906 .
- decoding parameter control section 2603 decides a decoding parameter used in decoding the second decoded spectrum of the frame, based on the received data loss information, using the decoded parameters and the parameter decoded on the previous frame.
- received data loss information is information indicating which portions of the extended band coded parameter cannot be used by extended band decoding section 1403 as a result of loss (including packet loss and the case where errors resulting from transmission errors are detected).
- the second decoded spectrum is then decoded using the decoded parameters and first decoded spectrum obtained by decoding parameter control section 2603 and the first decoded spectrum.
- This specific operation is the same as for extended band decoding section 1403 of FIG. 19 in Embodiment 2, and is therefore not described.
- decoding parameter control section 2603 Next, a first operating state of decoding parameter control section 2603 will be described below.
- decoding parameter control section 2603 substitutes a decoding parameter of the frequency band corresponding to the previous frame as the decoding parameter of the frequency band corresponding to a coding parameter that could not be obtained due to loss.
- ⁇ (n, m) amplitude adjustment coefficient of the mth frequency band of the nth frame
- a decoding parameter for the mth band of the previous frame is outputted as a decoding parameter corresponding to the lost coding parameter.
- the parameter decoded using the coded parameter for the received frame is outputted as is.
- the corresponding decoded parameter of the previous frame is used as an extended band decoded parameter for the entire band of the high frequency of the frame.
- decoding is always carried out using a decoding parameter of the previous frame at frames where loss has occurred, but another situation is also possible where decoding is carried out using the method described above only when correlation is higher than a threshold value based on correlation of a signal between the previous frame and the frame, and decoding is then carried out using a method closed within the frame in accordance with Embodiment 2 when correlation is lower than the threshold value.
- an index indicating the correlation between the signal of the previous frame and the signal of the frame there are correlation coefficients and spectrum distance between the previous frame and the frame, calculated using, for example, spectrum envelope information such as an LPC parameter obtained from the first layer coding parameter, information relating to voiced stationary of signals such as a pitch period and pitch gain parameter, first layer low band decoded signal, and the first layer low band decoded spectrum itself.
- spectrum envelope information such as an LPC parameter obtained from the first layer coding parameter, information relating to voiced stationary of signals such as a pitch period and pitch gain parameter, first layer low band decoded signal, and the first layer low band decoded spectrum itself.
- decoding parameter control section 2603 obtains a decoded parameter for the frequency band using the decoded parameter for the frequency band of the previous frame and the decoded parameter for the frequency band neighboring the frequency band of the previous frame and the frame.
- the decoded parameter is obtained in the following manner using the decoded parameter for the mth band of the previous frame ((n ⁇ 1)th frame) as a decoded parameter corresponding to the lost coded parameter and the decoded parameter for the band (the same band as for the previous frame and the frame) neighboring the frequency band of the previous frame and the frame.
- the coding parameters for the lost frames can also be decoded in the same way as for the first operating state or for the second operating state using the coded parameters for frames before and after the frame.
- an interpolated value which is an intermediate value between the decoded parameter for the previous frame and the decoded parameter for the following frame is obtained and used as a decoded parameter.
- the scalable decoding apparatus and scalable encoding apparatus according to the present invention is by no means limited to the above Embodiments 1 to 3, and various modifications thereof are possible.
- the scalable decoding apparatus and the scalable encoding apparatus according to the present invention can be provided to a communication terminal apparatus and a base station apparatus in a mobile communication system so as to make it possible to provide a communication terminal apparatus and a base station apparatus having the same operation results as described above.
- each function block used to explain the above-described embodiments is typically implemented as an LSI constituted by an integrated circuit. These may be individual chips or may partially or totally contained on a single chip.
- each function block is described as an LSI, but this may also be referred to as “IC”, “system LSI”, “super LSI”, “ultra LSI” depending on differing extents of integration.
- circuit integration is not limited to LSI's, and implementation using dedicated circuitry or general purpose processors is also possible.
- LSI manufacture utilization of a programmable FPGA (Field Programmable Gate Array) or a reconfigurable processor in which connections and settings of circuit cells within an LSI can be reconfigured is also possible.
- FPGA Field Programmable Gate Array
- mirroring is carried out after adjusting the range of variations of the original low band spectrum that is mirrored, so that it is not necessary to transmit information relating to adjustment of the range of variations.
- the present invention when lag information is not received due to transmission path errors, upon decoding of the encoded high band component, mirroring is carried out using the procedure of the first characteristic, and decoding processing is carried out for the high band component, so that it is possible to generate a spectrum having a harmonic structure at a high band without using the lag information. Further, the intensity of the harmonic structure can also be adjusted to a valid level. It is also possible to generate a pseudo spectrum using another technique in place of the mirroring.
- a bitstream is used in the order of scale factor, amplitude adjustment coefficient, lag and residual spectrum.
- a decoded signal is generated using only scale factor, amplitude adjustment coefficient and lag information.
- decoding processing is then carried out using the decoding procedure of the second characteristic.
- the present invention when the present invention is applied to a system designed so that the rate of occurrence of transmission errors and loss/discarding of coded information increases in order of scale factor, amplitude adjustment coefficient, lag and residual spectrum (that is, the scale factor is protected from errors with the highest priority, and preferentially transmitted on the transmission path), it is possible to minimize quality degradation of decoded speech due to transmission path errors. Further, the decoding speech quality gradually changes with decoding each parameter, so that it is possible to implement more fine grained scalability than in the related art.
- the extended band decoding section is provided with: a buffer for storing decoding parameters decoded from extended band coded parameters used for decoding of the previous frame; and a decoding parameter control section that decides a decoded parameter for the frame using the decoded parameters of the frame and the previous frame and using data loss information for the received bitstream for the frame, and generates a second decoded spectrum using the first decoded spectrum for the frame and the decoded parameter outputted from the decoding parameter control section.
- the decoding parameter control section may obtain the decoding parameter for the frequency band using decoding parameters for the frequency band of the previous frame and decoding parameters for the frequency band neighboring the frequency band of the previous frame and the frame.
- the scalable decoding apparatus and scalable encoding apparatus of the present invention can be applied to a mobile communication system and a packet communication system using Internet protocol.
Landscapes
- Engineering & Computer Science (AREA)
- Signal Processing (AREA)
- Physics & Mathematics (AREA)
- Computational Linguistics (AREA)
- Quality & Reliability (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Human Computer Interaction (AREA)
- Acoustics & Sound (AREA)
- Multimedia (AREA)
- Theoretical Computer Science (AREA)
- Computer Networks & Wireless Communication (AREA)
- Compression, Expansion, Code Conversion, And Decoders (AREA)
- Compression Or Coding Systems Of Tv Signals (AREA)
Applications Claiming Priority (3)
Application Number | Priority Date | Filing Date | Title |
---|---|---|---|
JP2004322954 | 2004-11-05 | ||
JP2004-322954 | 2004-11-05 | ||
PCT/JP2005/020201 WO2006049205A1 (fr) | 2004-11-05 | 2005-11-02 | Appareil de codage et de decodage modulables |
Publications (2)
Publication Number | Publication Date |
---|---|
US20080126082A1 US20080126082A1 (en) | 2008-05-29 |
US7983904B2 true US7983904B2 (en) | 2011-07-19 |
Family
ID=36319210
Family Applications (1)
Application Number | Title | Priority Date | Filing Date |
---|---|---|---|
US11/718,437 Expired - Fee Related US7983904B2 (en) | 2004-11-05 | 2005-11-02 | Scalable decoding apparatus and scalable encoding apparatus |
Country Status (8)
Country | Link |
---|---|
US (1) | US7983904B2 (fr) |
EP (1) | EP1808684B1 (fr) |
JP (1) | JP4977472B2 (fr) |
KR (1) | KR20070084002A (fr) |
CN (1) | CN101048649A (fr) |
BR (1) | BRPI0517780A2 (fr) |
RU (2) | RU2404506C2 (fr) |
WO (1) | WO2006049205A1 (fr) |
Cited By (11)
Publication number | Priority date | Publication date | Assignee | Title |
---|---|---|---|---|
US20090070120A1 (en) * | 2007-09-12 | 2009-03-12 | Fujitsu Limited | Audio regeneration method |
US20090310799A1 (en) * | 2008-06-13 | 2009-12-17 | Shiro Suzuki | Information processing apparatus and method, and program |
US20110035213A1 (en) * | 2007-06-22 | 2011-02-10 | Vladimir Malenovsky | Method and Device for Sound Activity Detection and Sound Signal Classification |
US20110282655A1 (en) * | 2008-12-19 | 2011-11-17 | Fujitsu Limited | Voice band enhancement apparatus and voice band enhancement method |
US20120158411A1 (en) * | 2005-07-11 | 2012-06-21 | Sony Corporation | Signal encoding apparatus and method, signal decoding apparatus and method, programs and recording mediums |
US20130124201A1 (en) * | 2010-06-21 | 2013-05-16 | Panasonic Corporation | Decoding device, encoding device, and methods for same |
US20140086420A1 (en) * | 2011-08-08 | 2014-03-27 | The Intellisis Corporation | System and method for tracking sound pitch across an audio signal using harmonic envelope |
US8977546B2 (en) | 2009-10-20 | 2015-03-10 | Panasonic Intellectual Property Corporation Of America | Encoding device, decoding device and method for both |
US20150334407A1 (en) * | 2012-04-24 | 2015-11-19 | Telefonaktiebolaget L M Ericsson (Publ) | Encoding and deriving parameters for coded multi-layer video sequences |
US20160173768A1 (en) * | 2013-10-01 | 2016-06-16 | Gopro, Inc. | Camera system transmission in bandwidth constrained environments |
US20220270619A1 (en) * | 2013-07-22 | 2022-08-25 | Fraunhofer-Gesellschaft Zur Foerderung Der Angewandten Forschung E.V. | Apparatus and method for encoding or decoding an audio signal with intelligent gap filling in the spectral domain |
Families Citing this family (66)
Publication number | Priority date | Publication date | Assignee | Title |
---|---|---|---|---|
EP1744139B1 (fr) * | 2004-05-14 | 2015-11-11 | Panasonic Intellectual Property Corporation of America | Dispositif de décodage et méthode pour ceux-ci |
JP4977471B2 (ja) | 2004-11-05 | 2012-07-18 | パナソニック株式会社 | 符号化装置及び符号化方法 |
CN101273403B (zh) * | 2005-10-14 | 2012-01-18 | 松下电器产业株式会社 | 可扩展编码装置、可扩展解码装置以及其方法 |
BRPI0619258A2 (pt) * | 2005-11-30 | 2011-09-27 | Matsushita Electric Ind Co Ltd | aparelho de codificação de sub-banda e método de codificação de sub-banda |
DE602006015097D1 (de) * | 2005-11-30 | 2010-08-05 | Panasonic Corp | Skalierbare codierungsvorrichtung und skalierbares codierungsverfahren |
US8352254B2 (en) * | 2005-12-09 | 2013-01-08 | Panasonic Corporation | Fixed code book search device and fixed code book search method |
CN101089951B (zh) * | 2006-06-16 | 2011-08-31 | 北京天籁传音数字技术有限公司 | 频带扩展编码方法及装置和解码方法及装置 |
CN101115124B (zh) * | 2006-07-26 | 2012-04-18 | 日电(中国)有限公司 | 基于音频水印识别媒体节目的方法和装置 |
US8532984B2 (en) * | 2006-07-31 | 2013-09-10 | Qualcomm Incorporated | Systems, methods, and apparatus for wideband encoding and decoding of active frames |
US9454974B2 (en) | 2006-07-31 | 2016-09-27 | Qualcomm Incorporated | Systems, methods, and apparatus for gain factor limiting |
US8260609B2 (en) | 2006-07-31 | 2012-09-04 | Qualcomm Incorporated | Systems, methods, and apparatus for wideband encoding and decoding of inactive frames |
US8560328B2 (en) | 2006-12-15 | 2013-10-15 | Panasonic Corporation | Encoding device, decoding device, and method thereof |
FR2911031B1 (fr) * | 2006-12-28 | 2009-04-10 | Actimagine Soc Par Actions Sim | Procede et dispositif de codage audio |
FR2911020B1 (fr) * | 2006-12-28 | 2009-05-01 | Actimagine Soc Par Actions Sim | Procede et dispositif de codage audio |
FR2912249A1 (fr) * | 2007-02-02 | 2008-08-08 | France Telecom | Codage/decodage perfectionnes de signaux audionumeriques. |
JP5294713B2 (ja) * | 2007-03-02 | 2013-09-18 | パナソニック株式会社 | 符号化装置、復号装置およびそれらの方法 |
JP4708446B2 (ja) * | 2007-03-02 | 2011-06-22 | パナソニック株式会社 | 符号化装置、復号装置およびそれらの方法 |
JP4871894B2 (ja) * | 2007-03-02 | 2012-02-08 | パナソニック株式会社 | 符号化装置、復号装置、符号化方法および復号方法 |
WO2008114078A1 (fr) * | 2007-03-16 | 2008-09-25 | Nokia Corporation | Codeur |
US9466307B1 (en) * | 2007-05-22 | 2016-10-11 | Digimarc Corporation | Robust spectral encoding and decoding methods |
CN100524462C (zh) * | 2007-09-15 | 2009-08-05 | 华为技术有限公司 | 对高带信号进行帧错误隐藏的方法及装置 |
US9872066B2 (en) * | 2007-12-18 | 2018-01-16 | Ibiquity Digital Corporation | Method for streaming through a data service over a radio link subsystem |
EP2224432B1 (fr) * | 2007-12-21 | 2017-03-15 | Panasonic Intellectual Property Corporation of America | Codeur, décodeur et procédé de codage |
JP5485909B2 (ja) * | 2007-12-31 | 2014-05-07 | エルジー エレクトロニクス インコーポレイティド | オーディオ信号処理方法及び装置 |
EP2251861B1 (fr) * | 2008-03-14 | 2017-11-22 | Panasonic Intellectual Property Corporation of America | Dispositif d'encodage et leur procédé |
EP2255534B1 (fr) * | 2008-03-20 | 2017-12-20 | Samsung Electronics Co., Ltd. | Appareil et procédé permettant d'effectuer un codage au moyen d'une extension de bande passante dans un terminal portable |
KR101424944B1 (ko) * | 2008-12-15 | 2014-08-01 | 프라운호퍼 게젤샤프트 쭈르 푀르데룽 데어 안겐반텐 포르슝 에. 베. | 오디오 인코더 및 대역폭 확장 디코더 |
JP5754899B2 (ja) | 2009-10-07 | 2015-07-29 | ソニー株式会社 | 復号装置および方法、並びにプログラム |
EP2490217A4 (fr) * | 2009-10-14 | 2016-08-24 | Panasonic Ip Corp America | Dispositif de codage, procédé de codage et procédés correspondants |
KR101309671B1 (ko) * | 2009-10-21 | 2013-09-23 | 돌비 인터네셔널 에이비 | 결합된 트랜스포저 필터 뱅크에서의 오버샘플링 |
EP2555188B1 (fr) * | 2010-03-31 | 2014-05-14 | Fujitsu Limited | Appareils et procédés d'extension de largeur de bande |
JP5850216B2 (ja) | 2010-04-13 | 2016-02-03 | ソニー株式会社 | 信号処理装置および方法、符号化装置および方法、復号装置および方法、並びにプログラム |
JP5652658B2 (ja) | 2010-04-13 | 2015-01-14 | ソニー株式会社 | 信号処理装置および方法、符号化装置および方法、復号装置および方法、並びにプログラム |
JP5609737B2 (ja) | 2010-04-13 | 2014-10-22 | ソニー株式会社 | 信号処理装置および方法、符号化装置および方法、復号装置および方法、並びにプログラム |
US8762158B2 (en) * | 2010-08-06 | 2014-06-24 | Samsung Electronics Co., Ltd. | Decoding method and decoding apparatus therefor |
JP5707842B2 (ja) | 2010-10-15 | 2015-04-30 | ソニー株式会社 | 符号化装置および方法、復号装置および方法、並びにプログラム |
US9230551B2 (en) * | 2010-10-18 | 2016-01-05 | Nokia Technologies Oy | Audio encoder or decoder apparatus |
WO2012144128A1 (fr) * | 2011-04-20 | 2012-10-26 | パナソニック株式会社 | Dispositif de codage vocal/audio, dispositif de décodage vocal/audio et leurs procédés |
JP6037156B2 (ja) | 2011-08-24 | 2016-11-30 | ソニー株式会社 | 符号化装置および方法、並びにプログラム |
JP5975243B2 (ja) | 2011-08-24 | 2016-08-23 | ソニー株式会社 | 符号化装置および方法、並びにプログラム |
JP5942358B2 (ja) | 2011-08-24 | 2016-06-29 | ソニー株式会社 | 符号化装置および方法、復号装置および方法、並びにプログラム |
TWI610296B (zh) * | 2011-10-21 | 2018-01-01 | 三星電子股份有限公司 | 訊框錯誤修補裝置及音訊解碼裝置 |
CN103366749B (zh) * | 2012-03-28 | 2016-01-27 | 北京天籁传音数字技术有限公司 | 一种声音编解码装置及其方法 |
CN103366751B (zh) * | 2012-03-28 | 2015-10-14 | 北京天籁传音数字技术有限公司 | 一种声音编解码装置及其方法 |
CN110706715B (zh) | 2012-03-29 | 2022-05-24 | 华为技术有限公司 | 信号编码和解码的方法和设备 |
US9711156B2 (en) * | 2013-02-08 | 2017-07-18 | Qualcomm Incorporated | Systems and methods of performing filtering for gain determination |
US9601125B2 (en) * | 2013-02-08 | 2017-03-21 | Qualcomm Incorporated | Systems and methods of performing noise modulation and gain adjustment |
CN108364657B (zh) | 2013-07-16 | 2020-10-30 | 超清编解码有限公司 | 处理丢失帧的方法和解码器 |
CN105745703B (zh) * | 2013-09-16 | 2019-12-10 | 三星电子株式会社 | 信号编码方法和装置以及信号解码方法和装置 |
CN105531762B (zh) | 2013-09-19 | 2019-10-01 | 索尼公司 | 编码装置和方法、解码装置和方法以及程序 |
KR101782454B1 (ko) * | 2013-12-06 | 2017-09-28 | 후아웨이 테크놀러지 컴퍼니 리미티드 | 이미지 복호화 장치, 이미지 부호화 장치, 및 부호화된 데이터 변환 장치 |
JP6593173B2 (ja) | 2013-12-27 | 2019-10-23 | ソニー株式会社 | 復号化装置および方法、並びにプログラム |
CN111370008B (zh) * | 2014-02-28 | 2024-04-09 | 弗朗霍弗应用研究促进协会 | 解码装置、编码装置、解码方法、编码方法、终端装置、以及基站装置 |
ES2878061T3 (es) * | 2014-05-01 | 2021-11-18 | Nippon Telegraph & Telephone | Dispositivo de generación de secuencia envolvente combinada periódica, método de generación de secuencia envolvente combinada periódica, programa de generación de secuencia envolvente combinada periódica y soporte de registro |
CN110875048B (zh) * | 2014-05-01 | 2023-06-09 | 日本电信电话株式会社 | 编码装置、及其方法、记录介质 |
CN106683681B (zh) | 2014-06-25 | 2020-09-25 | 华为技术有限公司 | 处理丢失帧的方法和装置 |
EP4293666A3 (fr) | 2014-07-28 | 2024-03-06 | Samsung Electronics Co., Ltd. | Procédé et appareil de codage de signal ainsi que procédé et appareil de décodage de signal |
JP2016038435A (ja) * | 2014-08-06 | 2016-03-22 | ソニー株式会社 | 符号化装置および方法、復号装置および方法、並びにプログラム |
JP6611042B2 (ja) * | 2015-12-02 | 2019-11-27 | パナソニックIpマネジメント株式会社 | 音声信号復号装置及び音声信号復号方法 |
US10825467B2 (en) * | 2017-04-21 | 2020-11-03 | Qualcomm Incorporated | Non-harmonic speech detection and bandwidth extension in a multi-source environment |
US10431231B2 (en) * | 2017-06-29 | 2019-10-01 | Qualcomm Incorporated | High-band residual prediction with time-domain inter-channel bandwidth extension |
CN110556122B (zh) * | 2019-09-18 | 2024-01-19 | 腾讯科技(深圳)有限公司 | 频带扩展方法、装置、电子设备及计算机可读存储介质 |
CN113113032B (zh) * | 2020-01-10 | 2024-08-09 | 华为技术有限公司 | 一种音频编解码方法和音频编解码设备 |
CN112309408A (zh) * | 2020-11-10 | 2021-02-02 | 北京百瑞互联技术有限公司 | 一种扩展lc3音频编解码带宽的方法、装置及存储介质 |
CN113724725B (zh) * | 2021-11-04 | 2022-01-18 | 北京百瑞互联技术有限公司 | 一种蓝牙音频啸叫检测抑制方法、装置、介质及蓝牙设备 |
CN114664319A (zh) * | 2022-03-28 | 2022-06-24 | 北京百度网讯科技有限公司 | 频带扩展方法、装置、设备、介质及程序产品 |
Citations (10)
Publication number | Priority date | Publication date | Assignee | Title |
---|---|---|---|---|
US5581652A (en) * | 1992-10-05 | 1996-12-03 | Nippon Telegraph And Telephone Corporation | Reconstruction of wideband speech from narrowband speech using codebooks |
US5774835A (en) * | 1994-08-22 | 1998-06-30 | Nec Corporation | Method and apparatus of postfiltering using a first spectrum parameter of an encoded sound signal and a second spectrum parameter of a lesser degree than the first spectrum parameter |
WO2001003124A1 (fr) | 1999-07-06 | 2001-01-11 | Telefonaktiebolaget Lm Ericsson | Etalement de la largeur de bande vocale |
US20030088423A1 (en) * | 2001-11-02 | 2003-05-08 | Kosuke Nishio | Encoding device and decoding device |
US20030158726A1 (en) * | 2000-04-18 | 2003-08-21 | Pierrick Philippe | Spectral enhancing method and device |
US6611800B1 (en) * | 1996-09-24 | 2003-08-26 | Sony Corporation | Vector quantization method and speech encoding method and apparatus |
US20050080621A1 (en) * | 2002-08-01 | 2005-04-14 | Mineo Tsushima | Audio decoding apparatus and audio decoding method |
US20050203736A1 (en) | 1996-11-07 | 2005-09-15 | Matsushita Electric Industrial Co., Ltd. | Excitation vector generator, speech coder and speech decoder |
US20060251178A1 (en) * | 2003-09-16 | 2006-11-09 | Matsushita Electric Industrial Co., Ltd. | Encoder apparatus and decoder apparatus |
US7205910B2 (en) * | 2002-08-21 | 2007-04-17 | Sony Corporation | Signal encoding apparatus and signal encoding method, and signal decoding apparatus and signal decoding method |
Family Cites Families (4)
Publication number | Priority date | Publication date | Assignee | Title |
---|---|---|---|---|
EP1405303A1 (fr) * | 2001-06-28 | 2004-04-07 | Koninklijke Philips Electronics N.V. | Systeme d'emission de signaux a large bande |
JP3926726B2 (ja) * | 2001-11-14 | 2007-06-06 | 松下電器産業株式会社 | 符号化装置および復号化装置 |
JP2003323199A (ja) * | 2002-04-26 | 2003-11-14 | Matsushita Electric Ind Co Ltd | 符号化装置、復号化装置及び符号化方法、復号化方法 |
JP3881946B2 (ja) * | 2002-09-12 | 2007-02-14 | 松下電器産業株式会社 | 音響符号化装置及び音響符号化方法 |
-
2005
- 2005-11-02 BR BRPI0517780-4A patent/BRPI0517780A2/pt not_active IP Right Cessation
- 2005-11-02 RU RU2007116937/09A patent/RU2404506C2/ru not_active IP Right Cessation
- 2005-11-02 WO PCT/JP2005/020201 patent/WO2006049205A1/fr active Application Filing
- 2005-11-02 JP JP2006542422A patent/JP4977472B2/ja not_active Expired - Fee Related
- 2005-11-02 US US11/718,437 patent/US7983904B2/en not_active Expired - Fee Related
- 2005-11-02 KR KR1020077010273A patent/KR20070084002A/ko not_active Application Discontinuation
- 2005-11-02 EP EP05805495.8A patent/EP1808684B1/fr not_active Not-in-force
- 2005-11-02 CN CNA2005800373627A patent/CN101048649A/zh active Pending
-
2010
- 2010-10-01 RU RU2010140339/09A patent/RU2434324C1/ru not_active IP Right Cessation
Patent Citations (13)
Publication number | Priority date | Publication date | Assignee | Title |
---|---|---|---|---|
US5581652A (en) * | 1992-10-05 | 1996-12-03 | Nippon Telegraph And Telephone Corporation | Reconstruction of wideband speech from narrowband speech using codebooks |
US5774835A (en) * | 1994-08-22 | 1998-06-30 | Nec Corporation | Method and apparatus of postfiltering using a first spectrum parameter of an encoded sound signal and a second spectrum parameter of a lesser degree than the first spectrum parameter |
US6611800B1 (en) * | 1996-09-24 | 2003-08-26 | Sony Corporation | Vector quantization method and speech encoding method and apparatus |
US20060235682A1 (en) | 1996-11-07 | 2006-10-19 | Matsushita Electric Industrial Co., Ltd. | Excitation vector generator, speech coder and speech decoder |
US20050203736A1 (en) | 1996-11-07 | 2005-09-15 | Matsushita Electric Industrial Co., Ltd. | Excitation vector generator, speech coder and speech decoder |
US20070100613A1 (en) | 1996-11-07 | 2007-05-03 | Matsushita Electric Industrial Co., Ltd. | Excitation vector generator, speech coder and speech decoder |
US6507820B1 (en) | 1999-07-06 | 2003-01-14 | Telefonaktiebolaget Lm Ericsson | Speech band sampling rate expansion |
WO2001003124A1 (fr) | 1999-07-06 | 2001-01-11 | Telefonaktiebolaget Lm Ericsson | Etalement de la largeur de bande vocale |
US20030158726A1 (en) * | 2000-04-18 | 2003-08-21 | Pierrick Philippe | Spectral enhancing method and device |
US20030088423A1 (en) * | 2001-11-02 | 2003-05-08 | Kosuke Nishio | Encoding device and decoding device |
US20050080621A1 (en) * | 2002-08-01 | 2005-04-14 | Mineo Tsushima | Audio decoding apparatus and audio decoding method |
US7205910B2 (en) * | 2002-08-21 | 2007-04-17 | Sony Corporation | Signal encoding apparatus and signal encoding method, and signal decoding apparatus and signal decoding method |
US20060251178A1 (en) * | 2003-09-16 | 2006-11-09 | Matsushita Electric Industrial Co., Ltd. | Encoder apparatus and decoder apparatus |
Non-Patent Citations (8)
Title |
---|
Kovesi, B. et al., "A Scalable Speech and Audio Coding Scheme with Continuous Bitrate Flexibility", Proc. of ICASSP-04, vol. 1, Mar. 17, 2004, p. 1-273-276. |
Kovesi, B. et al., "A Scalable Speech and Audio Coding Scheme with Continuous Bitrate Flexibility", Proc. of ICASSP-04, vol. 1, Mar. 17, 2004, p. I-273-276. |
Makhoul J et al., "High-Frequency Regeneration in Speech Coding Systems", International Conference on Acoustics, Speech & Signal Processing, ICASSP. Washington, Apr. 2-4, 1979; [International Conference on Acoustics, Speech & Signal Processing, ICASSP], New York, IEEE, US, vol. Conf. 4, Jan. 1, 1979, pp. 428-431, XP001122019. |
Oshikiri et al., "Pichi Filtering ni Motozuku Spectre Fugoka o Mochiita Choko Taiiki Schelable Onsei Fugoka no Kaizen", The Acoustical Society of Japan (ASJ) 2004 Nen Shuki Kenkyu Happyokai Koen Ronbunshu-I-, 2-4-13, Sep. 21, 2004, pp. 297 to 298, XP002998459. |
Oshikiri et al., "Pichi Filtering ni Motozuku Spectre Fugoka o Mochiita Choko Taiiki Schelable Onsei Fugoka no Kaizen", The Acoustical Society of Japan (ASJ) 2004 Nen Shuki Kenkyu Happyokai Koen Ronbunshu-I-, 2-4-13, Sep. 21, 2004, pp. 297 to 298. (including partial English language translation). |
Oshikiri M et al. , Efficient spectrum coding for super-wideband speech and its application to 7/10/15 KHz bandwidth scalable coders, Acoustics, Speech, and Signal Processing, 2004, Proceedings, (ICASSP '04), IEEE International Conference on Montreal, Quebec, Canada May 17-21, 2004, Piscataway, NJ, USA, IEEE, Piscataway, NJ, USA, LNKD-DOI: 10.1109/ICASSP.2004.1326027, vol. 1, May 17, 2004, pp. 481-484, XP010717670. |
U.S. Appl. No. 11/573,761 to Ehara et al., filed Feb. 15, 2007. |
U.S. Appl. No. 11/576,264 to Goto et al., filed Mar. 29, 2007. |
Cited By (22)
Publication number | Priority date | Publication date | Assignee | Title |
---|---|---|---|---|
US20120158411A1 (en) * | 2005-07-11 | 2012-06-21 | Sony Corporation | Signal encoding apparatus and method, signal decoding apparatus and method, programs and recording mediums |
US8340213B2 (en) * | 2005-07-11 | 2012-12-25 | Sony Corporation | Signal encoding apparatus and method, signal decoding apparatus and method, programs and recording mediums |
US20110035213A1 (en) * | 2007-06-22 | 2011-02-10 | Vladimir Malenovsky | Method and Device for Sound Activity Detection and Sound Signal Classification |
US8990073B2 (en) * | 2007-06-22 | 2015-03-24 | Voiceage Corporation | Method and device for sound activity detection and sound signal classification |
US8073687B2 (en) * | 2007-09-12 | 2011-12-06 | Fujitsu Limited | Audio regeneration method |
US20090070120A1 (en) * | 2007-09-12 | 2009-03-12 | Fujitsu Limited | Audio regeneration method |
US20090310799A1 (en) * | 2008-06-13 | 2009-12-17 | Shiro Suzuki | Information processing apparatus and method, and program |
US8781823B2 (en) * | 2008-12-19 | 2014-07-15 | Fujitsu Limited | Voice band enhancement apparatus and voice band enhancement method that generate wide-band spectrum |
US20110282655A1 (en) * | 2008-12-19 | 2011-11-17 | Fujitsu Limited | Voice band enhancement apparatus and voice band enhancement method |
US8977546B2 (en) | 2009-10-20 | 2015-03-10 | Panasonic Intellectual Property Corporation Of America | Encoding device, decoding device and method for both |
US20130124201A1 (en) * | 2010-06-21 | 2013-05-16 | Panasonic Corporation | Decoding device, encoding device, and methods for same |
US9076434B2 (en) * | 2010-06-21 | 2015-07-07 | Panasonic Intellectual Property Corporation Of America | Decoding and encoding apparatus and method for efficiently encoding spectral data in a high-frequency portion based on spectral data in a low-frequency portion of a wideband signal |
US20140086420A1 (en) * | 2011-08-08 | 2014-03-27 | The Intellisis Corporation | System and method for tracking sound pitch across an audio signal using harmonic envelope |
US9473866B2 (en) * | 2011-08-08 | 2016-10-18 | Knuedge Incorporated | System and method for tracking sound pitch across an audio signal using harmonic envelope |
US20150334407A1 (en) * | 2012-04-24 | 2015-11-19 | Telefonaktiebolaget L M Ericsson (Publ) | Encoding and deriving parameters for coded multi-layer video sequences |
US10609394B2 (en) * | 2012-04-24 | 2020-03-31 | Telefonaktiebolaget Lm Ericsson (Publ) | Encoding and deriving parameters for coded multi-layer video sequences |
US20220270619A1 (en) * | 2013-07-22 | 2022-08-25 | Fraunhofer-Gesellschaft Zur Foerderung Der Angewandten Forschung E.V. | Apparatus and method for encoding or decoding an audio signal with intelligent gap filling in the spectral domain |
US11922956B2 (en) * | 2013-07-22 | 2024-03-05 | Fraunhofer-Gesellschaft Zur Foerderung Der Angewandten Forschung E.V. | Apparatus and method for encoding or decoding an audio signal with intelligent gap filling in the spectral domain |
US11996106B2 (en) | 2013-07-22 | 2024-05-28 | Fraunhofer-Gesellschaft Zur Foerderung Der Angewandten Forschung E. V. | Apparatus and method for encoding and decoding an encoded audio signal using temporal noise/patch shaping |
US20160173768A1 (en) * | 2013-10-01 | 2016-06-16 | Gopro, Inc. | Camera system transmission in bandwidth constrained environments |
US9485418B2 (en) * | 2013-10-01 | 2016-11-01 | Gopro, Inc. | Camera system transmission in bandwidth constrained environments |
US9584720B2 (en) | 2013-10-01 | 2017-02-28 | Gopro, Inc. | Camera system dual-encoder architecture |
Also Published As
Publication number | Publication date |
---|---|
WO2006049205A1 (fr) | 2006-05-11 |
EP1808684B1 (fr) | 2014-07-30 |
CN101048649A (zh) | 2007-10-03 |
JP4977472B2 (ja) | 2012-07-18 |
EP1808684A1 (fr) | 2007-07-18 |
RU2434324C1 (ru) | 2011-11-20 |
KR20070084002A (ko) | 2007-08-24 |
JPWO2006049205A1 (ja) | 2008-05-29 |
US20080126082A1 (en) | 2008-05-29 |
EP1808684A4 (fr) | 2010-07-14 |
BRPI0517780A2 (pt) | 2011-04-19 |
RU2404506C2 (ru) | 2010-11-20 |
RU2007116937A (ru) | 2008-11-20 |
Similar Documents
Publication | Publication Date | Title |
---|---|---|
US7983904B2 (en) | Scalable decoding apparatus and scalable encoding apparatus | |
US8457319B2 (en) | Stereo encoding device, stereo decoding device, and stereo encoding method | |
US7769584B2 (en) | Encoder, decoder, encoding method, and decoding method | |
RU2488897C1 (ru) | Кодирующее устройство, декодирующее устройство и способ | |
US8099275B2 (en) | Sound encoder and sound encoding method for generating a second layer decoded signal based on a degree of variation in a first layer decoded signal | |
EP1939862B1 (fr) | Dispositif de codage, dispositif de décodage et son procédé | |
JP5036317B2 (ja) | スケーラブル符号化装置、スケーラブル復号化装置、およびこれらの方法 | |
JP4606418B2 (ja) | スケーラブル符号化装置、スケーラブル復号装置及びスケーラブル符号化方法 | |
JPWO2008072737A1 (ja) | 符号化装置、復号装置およびこれらの方法 | |
RU2459283C2 (ru) | Кодирующее устройство, декодирующее устройство и способ | |
WO2011058752A1 (fr) | Appareil d'encodage, appareil de décodage et procédés pour ces appareils |
Legal Events
Date | Code | Title | Description |
---|---|---|---|
AS | Assignment |
Owner name: MATSUSHITA ELECTRIC INDUSTRIAL CO., LTD., JAPAN Free format text: ASSIGNMENT OF ASSIGNORS INTEREST;ASSIGNORS:EHARA, HIROYUKI;OSHIKIRI, MASAHIRO;YOSHIDA, KOJI;REEL/FRAME:019715/0844 Effective date: 20070417 |
|
AS | Assignment |
Owner name: PANASONIC CORPORATION, JAPAN Free format text: CHANGE OF NAME;ASSIGNOR:MATSUSHITA ELECTRIC INDUSTRIAL CO., LTD.;REEL/FRAME:021835/0446 Effective date: 20081001 Owner name: PANASONIC CORPORATION,JAPAN Free format text: CHANGE OF NAME;ASSIGNOR:MATSUSHITA ELECTRIC INDUSTRIAL CO., LTD.;REEL/FRAME:021835/0446 Effective date: 20081001 |
|
STCF | Information on status: patent grant |
Free format text: PATENTED CASE |
|
FEPP | Fee payment procedure |
Free format text: PAYOR NUMBER ASSIGNED (ORIGINAL EVENT CODE: ASPN); ENTITY STATUS OF PATENT OWNER: LARGE ENTITY |
|
AS | Assignment |
Owner name: PANASONIC INTELLECTUAL PROPERTY CORPORATION OF AMERICA, CALIFORNIA Free format text: ASSIGNMENT OF ASSIGNORS INTEREST;ASSIGNOR:PANASONIC CORPORATION;REEL/FRAME:033033/0163 Effective date: 20140527 Owner name: PANASONIC INTELLECTUAL PROPERTY CORPORATION OF AME Free format text: ASSIGNMENT OF ASSIGNORS INTEREST;ASSIGNOR:PANASONIC CORPORATION;REEL/FRAME:033033/0163 Effective date: 20140527 |
|
FPAY | Fee payment |
Year of fee payment: 4 |
|
AS | Assignment |
Owner name: III HOLDINGS 12, LLC, DELAWARE Free format text: ASSIGNMENT OF ASSIGNORS INTEREST;ASSIGNOR:PANASONIC INTELLECTUAL PROPERTY CORPORATION OF AMERICA;REEL/FRAME:042386/0779 Effective date: 20170324 |
|
MAFP | Maintenance fee payment |
Free format text: PAYMENT OF MAINTENANCE FEE, 8TH YEAR, LARGE ENTITY (ORIGINAL EVENT CODE: M1552); ENTITY STATUS OF PATENT OWNER: LARGE ENTITY Year of fee payment: 8 |
|
FEPP | Fee payment procedure |
Free format text: MAINTENANCE FEE REMINDER MAILED (ORIGINAL EVENT CODE: REM.); ENTITY STATUS OF PATENT OWNER: LARGE ENTITY |
|
LAPS | Lapse for failure to pay maintenance fees |
Free format text: PATENT EXPIRED FOR FAILURE TO PAY MAINTENANCE FEES (ORIGINAL EVENT CODE: EXP.); ENTITY STATUS OF PATENT OWNER: LARGE ENTITY |
|
STCH | Information on status: patent discontinuation |
Free format text: PATENT EXPIRED DUE TO NONPAYMENT OF MAINTENANCE FEES UNDER 37 CFR 1.362 |
|
FP | Lapsed due to failure to pay maintenance fee |
Effective date: 20230719 |