WO2009068083A1 - An encoder - Google Patents

An encoder Download PDF

Info

Publication number
WO2009068083A1
WO2009068083A1 PCT/EP2007/062909 EP2007062909W WO2009068083A1 WO 2009068083 A1 WO2009068083 A1 WO 2009068083A1 EP 2007062909 W EP2007062909 W EP 2007062909W WO 2009068083 A1 WO2009068083 A1 WO 2009068083A1
Authority
WO
WIPO (PCT)
Prior art keywords
spectral
audio signal
encoded signal
dependent
spectral coefficient
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/EP2007/062909
Other languages
French (fr)
Inventor
Juha Petteri Ojanpera
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Nokia Inc
Original Assignee
Nokia Inc
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Nokia Inc filed Critical Nokia Inc
Priority to PCT/EP2007/062909 priority Critical patent/WO2009068083A1/en
Publication of WO2009068083A1 publication Critical patent/WO2009068083A1/en
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L19/00Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
    • G10L19/02Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis using spectral analysis, e.g. transform vocoders or subband vocoders
    • G10L19/032Quantisation or dequantisation of spectral components
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L19/00Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
    • G10L19/02Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis using spectral analysis, e.g. transform vocoders or subband vocoders
    • G10L19/0204Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis using spectral analysis, e.g. transform vocoders or subband vocoders using subband decomposition

Definitions

  • the present invention relates to coding, and in particular, but not exclusively to speech or audio coding.
  • Audio signals like speech or music, are encoded for example for enabling an efficient transmission or storage of the audio signals.
  • Audio encoders and decoders are used to represent audio based signals, such as music and background noise. These types of coders typically do not utilise a speech model for the coding process, rather they use processes for representing all types of audio signals, including speech.
  • Speech encoders and decoders are usually optimised for speech signals, and can operate at either a fixed or variable bit rate.
  • An audio codec can also be configured to operate with varying bit rates. At lower bit rates, such an audio codec may work with speech signals at a coding rate equivalent to a pure speech codec, At higher bit rates, the audio codec may code any signal including music, background noise and speech, with higher quality and performance.
  • the input signal is divided into a limited number of bands.
  • Each of the band signals may be quantized. From the theory of psychoacoustics it is known that the highest frequencies in the spectrum are perceptually less important than the low frequencies. This in some audio codecs is reflected by a bit allocation where fewer bits are allocated to high frequency signals than low frequency signals.
  • the encoder is responsible for adapting the received audio data in order that it may be output at a bit rate which the channel bandwidth conditions are capable of accepting. Ideally, the encoding process discards only the irrelevant information from the audio signal received and the decoder is capable of simply reversing the encoding process to obtain the decoded audio signal with little or no audible degradation.
  • One method used by the encoder to reduce to the number of bits required to encode an audio signal is that of quantization.
  • Well known perceptual quantisation schemes such as the one used in MP3 and AAC coding follow known psychoacoustic rules.
  • the psychoacoustic rules the spectral frame is divided into a series of frequency bands or regions and for each frequency band a quantization step size, also known as the scaie factor, is assigned.
  • the value of the scale factor is changed (in other words either increased or decreased) so that the quantization noise (in other words the error produced during quantization) is below a masking threshold.
  • a first frequency component has a small spectral value when compared against a neighbouring frequency component
  • all scale factors are below a masking threshold then the untrained listener w ⁇ l not perceive significant loss in quality and a 'transparent 1 quality may be achieved.
  • the scale factors have to be themselves quantized and transmitted to the receiver.
  • the high number of scale factors being using increases the bit rate associated with the scale factor coding techniques described above.
  • the encoding system is suitable for high bit rate coding the overheads introduced by transmitting scale factors to the decoder force a reduction in the quality of the encoding of the audio signal in low bit rate situations.
  • This invention proceeds from the consideration that whilst scale factor transmission is acceptable in high bit rate environments a system of audio quantization may be implemented which produces an acceptable level of quantization without needing as many parameters to be sent to the receiver which therefore allows encoded audio signals to be transmitted without suffering from over harsh coding.
  • Embodiments of the present invention aim to address the above problem.
  • an encoder for encoding an audio signal configured to: generate at least one encoded signal dependent on the audio signal; determine at least one quality indicator, wherein each quality indicator is dependent on the relative difference between an associated one of the at least one encoded signal and the audio signal; select to output one of the at least one encoded signal dependent on the associated quality indicator value.
  • the encoder may be further configured to: generate at least one spectral representation of the audio signal, wherein the spectral representation may comprise at least two sub-bands, each sub-band may comprise at least one spectral coefficient.
  • the encoder may be further configured to generate a first of the at least one spectral representation of the audio signal by applying at least one of the following: a shifted discrete fourier transform; a modified discrete cosine transform; and a discrete unitary transform:
  • the encoder may be further configured to generate further spectral representations of the audio signal by applying at least one gain factor to the at least one spectral coefficient of at least one of the at least two sub-bands.
  • the at least one gain factor may comprise at least one of the following multiplication gain factors: 0 and 0.5.
  • the encoder may be further configured to order the at least two sub-bands.
  • the encoder may be further configured to generate the further spectra! representations of the audio signal by progressively applying, dependent on the order of the at least two sub-bands, at least one gain factor to the at least one spectral coefficient of at least one of the at least two sub-bands.
  • the encoder may be further configured to generate an indicator to indicate which sub-band spectral coefficients are modified by the at least one gain factor.
  • the encoder may be further configured to order the at least two sub-bands according to a psycho acoustical model.
  • the encoder may be further configured to progressively generate the further spectral representations of the audio signal until the quality indicator is equal to or greater than a quality threshold.
  • a decoder for decoding an encoded signal configured to: partition the received encoded signal into at least a first and second part; decode the first part to generate at least one gain indicator; generate at least one frequency representation of the encoded signal dependent on the second part and the at least one gain indicator, wherein the at least one frequency representation comprises at [east one spectral coefficient.
  • the decoder may be further configured to generate at least one initial frequency representation of the encoded signal dependent on the second part of the received encoded signal.
  • the at least one gain indicator may comprise at least one deleted spectral coefficient index.
  • the decoder is preferably further configured to insert an at least one predetermined spectral coefficient value into the initial frequency representation determined by the at least one deleted spectral coefficient index.
  • the at least one predetermined spectral coefficient value is preferably a null value.
  • the at least one gain indicator may further comprise at least one spectral coefficient gain factor.
  • the decoder is preferably further configured to multiply a spectral coefficient of the initial frequency representation determined by the at least one deleted spectral coefficient index by at least one spectral coefficient gain factor.
  • a method for encoding an audio signal comprising: generating at least one encoded signal dependent on the audio signal; determining at least one quality indicator, wherein each quality indicator is dependent on the relative difference between an associated one of the at least one encoded signal and the audio signal; and selecting to output one of the at least one encoded signal dependent on the associated quality indicator value.
  • the method may further comprise: generating at least one spectral representation of the audio signal, wherein the spectral representation may comprise at least two sub-bands, each sub-band may comprise at least one spectral coefficient.
  • the method may further comprise generating a first of the at least one spectral representation of the audio signal by applying at least one of the following: a shifted discrete fourier transform; a modified discrete cosine transform; and a discrete unitary transform.
  • the method may further comprise generating further spectral representations of the audio signal by applying at least one gain factor to the at least one spectral coefficient of at least one of the at least two sub-bands.
  • the at least one gain factor may comprise at least one of the following multiplication gain factors: 0 and 0,5.
  • the method may further comprise ordering the at least two sub-bands.
  • the method ⁇ may further comprise generating the further spectral representations of the audio signal by progressively applying, dependent on the order of the at least two sub-bands, at least one gain factor to the at least one spectral coefficient of at least one of the at least two sub-bands.
  • the method may further comprise generating an indicator to indicate which sub- band spectral coefficients are modified by the at least one gain factor.
  • the method may further comprise ordering the at least two sub-bands according to a psycho acoustical model.
  • the method may further comprise progressively generating the further spectral representations of the- audio signal until the quality indicator is equal to or greater than a quality threshold.
  • a method for decoding an encoded signal comprising: partitioning the received encoded signal into at least a first and second part; decoding the first part to generate at least one gain indicator; generating at least one frequency representation of the encoded signal dependent on the second part and the at least one gain indicator, wherein the at least one frequency representation comprises at . least one spectral coefficient.
  • the method may further comprise generating at least one initial frequency representation of the encoded signal dependent on the second part of the received encoded signal.
  • the at least one gain indicator may comprise at least one deleted spectral coefficient index.
  • the method may further comprise inserting an at least one predetermined spectral coefficient value into the initial frequency representation determined by the at least one deleted spectral coefficient index.
  • the at least one predetermined spectral coefficient value is preferably a null value.
  • the at least one gain indicator may further comprise at least one spectral coefficient gain factor.
  • the method may further comprise multiplying a spectral coefficient of the initial frequency representation determined by the at least one deleted spectral coefficient index by at least one spectral coefficient gain factor.
  • An apparatus may comprise an encoder as described above.
  • An apparatus may comprise a decoder as described above.
  • An electronic device may comprise an encoder as described above.
  • An electronic device may comprise a decoder as described above.
  • a chipset may comprise an encoder as described above.
  • a chipset may comprise a decoder as described above.
  • a computer program product configured to perform a method for encoding an audio signal, comprising: generating at least one encoded signal dependent on the audio signal; determining at least one quality indicator, wherein each quality indicator is dependent on the relative difference between an associated one of the at least one encoded signal and the audio signal; selecting to output one of the at least one encoded signal dependent on the associated quality indicator value.
  • a computer program product configured to perform a method for decoding an encoded signal, comprising: partitioning the received encoded signal into at least a first and second part; decoding the first part to generate at least one gain indicator; generating at least one frequency representation of the encoded signal dependent on the second part and the at least one gain indicator, wherein the at least one frequency representation comprises at least one spectral coefficient.
  • an encoder for encoding an audio signal comprising: encoding means to generate at least one encoded signal dependent on the audio signal; processing means to determine at least one quality indicator, wherein each quality indicator is dependent on the relative difference between an associated one of the at least one encoded signal and the audio signal; and selection means to select to output one of the at least one encoded signal dependent on the associated quality indicator value.
  • a decoder for decoding an encoded audio signal, comprising: processing means to partition the received encoded signal into at least a first and second part; decoding means to decode the first part to generate at least one gain indicator; further processing means to generate at least one frequency representation of the encoded signal dependent on the second part and the at least one gain indicator, wherein the at least one frequency representation comprises at least one spectral coefficient.
  • FIG 1 shows schematically an electronic device employing embodiments of the invention
  • FIG. 2 shows schematically an audio codec system employing embodiments of the present invention
  • Figure 3 shows schematically an encoder part of the audio codec system shown in figure 2;
  • Figure 4 shows a flow diagram illustrating the operation of an embodiment of the audio encoder as shown in Figure 3 according to the present invention
  • Figure 5 shows a flow diagram illustrating the operation of the spectral modifier part of the audio encoder as shown in Figure 3 according to the present invention
  • Figure 6 shows schematically a decoder part of the audio codec system shown in Figure 2; and Figure 7 shows a flow diagram illustrating the operation of an embodiment of the audio decoder as shown in Figure 6 according to the present invention.
  • figure 1 schematic block diagram of an exemplary electronic device 10, which may incorporate a codec according to an embodiment of the invention.
  • the electronic device 10 may for example be a mobile terminal or user equipment of a wireless communication system.
  • the electronic device 10 comprises a microphone 1 1 , which is linked via an analogue-to-digital converter 14 to a processor 21.
  • the processor 21 is further linked via a digital-to-analogue converter 32 to loudspeakers 33.
  • the processor 21 is further linked to a transceiver (TX/RX) 13, to a user interface (Ul) 15 and to a memory 22.
  • the processor 21 may be configured to execute various program codes.
  • the implemented program codes comprise an audio encoding code for encoding a combined audio signal and code to extract and encode side information pertaining to the spatial information of the multiple channels.
  • the implemented program codes 23 further comprise an audio decoding code.
  • the implemented program codes 23 may be stored for example in the memory 22 for retrieval by the processor 21 whenever needed.
  • the memory 22 could further provide a section 24 for storing data, for example data that has been encoded in accordance with the invention.
  • the encoding and decoding code may in embodiments of the invention be implemented in hardware or firmware.
  • the user interface 15 enables a user to input commands to the electronic device 10, for example via a keypad, and/or to obtain information from the electronic device 10, for example via a display.
  • the transceiver 13 enables a communication with other electronic devices, for example via a wireless communication network.
  • a user of the electronic device 10 may use the microphone 1 1 for inputting speech that is to be transmitted to some other electronic device or that is to be stored in the data section 24 of the memory 22.
  • a corresponding application has been activated to this end by the user via the user interface 15.
  • This application which may be run by the processor 21 , causes the processor 21 to execute the encoding code stored in the memory 22.
  • the analogue-to-digital converter 14 converts the input analogue audio signal into a digital audio signal and provides the digital audio signal to the processor 21.
  • the processor 21 may then process the digital audio signal in the same way as described with reference to figures 2 and 3.
  • the resulting bit stream is provided to the transceiver 13 for transmission to another electronic device.
  • the coded data could be stored in the data section 24 of the memory 22, for instance for a later transmission or for a later presentation by the same electronic device 10.
  • the electronic device 10 could also receive a bit stream with correspondingly encoded data from another electronic device via its transceiver 13.
  • the processor 21 may execute the decoding program code stored in the memory 22.
  • the processor 21 decodes the received data, and provides the decoded data to the digital-to-analogue converter 32.
  • the digital-to-analogue converter 32 converts the digital decoded data into analogue audio data and outputs them via the loudspeakers 33. Execution of the decoding program code could be triggered as well by an application that has been called by the user via the user interface 15.
  • the received encoded data could also be stored instead of an immediate presentation via the loudspeakers 33 in the data section 24 of the memory 22, for instance for enabling a later presentation or a forwarding to still another electronic device.
  • FIG. 1 The general operation of audio codecs as employed by embodiments of the invention is shown in figure 2.
  • General audio coding/decoding systems consist of an encoder and a decoder, as illustrated schematically in figure 2. Illustrated is a system 102 with an encoder 104, a storage or media channel 106 and a decoder 108.
  • the encoder 104 compresses an input audio signal 1 10 producing a bit stream 112, which is either stored or transmitted through a media channel 106.
  • the bit stream 112 can be received within the decoder 108.
  • the decoder 108 decompresses the bit stream 112 and produces an output audio signal 114.
  • the bit rate of the bit stream 112 and the quality of the output audio signal 114 in relation to the input signal 110 are the main features, which define the performance of the coding system 102.
  • FIG. 3 depicts schematically an encoder according to an exemplary embodiment of the invention.
  • the encoder 104 comprises the input 203 which is arranged to receive an audio signal input.
  • the audio input 203 is connected to an input of a time to frequency domain transformer 205.
  • the time to frequency domain transformer 205 has an output which is connected to an input of a spectral coefficient orderer 213.
  • the spectral coefficient orderer 213 has an output which is connected to an input of a spectral modifier 217.
  • the spectral modifier 217 has an output which is connected to an input of an encoder 207.
  • the encoder 207 has an output which is connected to an input of a quantizer 209.
  • the quantizer 209 has an output which is connected to an input of a distortion estimator 21 1.
  • the distortion estimator 211 has an output which is connected to an input of a quality checker 215.
  • the quality checker 215 has an output which is connected to a further input of the spectral modifier 217, and a further output which is connected to a bitstream formatter 219.
  • the bitstream formatter 219 has an output which is connected to the encoder output 221.
  • the encoded audio output 221 is the signal output comprising the encoded signal. With respect to figure 4 the operation of the encoder 104 as shown in figure 3 is described in further detail.
  • the audio signal is received by the coder 104.
  • the audio signal is a digitally sampled signal.
  • the audio input may be an analogue audio signal, for example from a microphone 6, which is analogue-to-digitally converted (A-D).
  • the audio input is converted from a pulse modulation digital signal to amplitude modulation digital signal.
  • the input 203 is connected to a time-to-frequency domain transformer 205.
  • the time-to-frequency domain transformer receives the audio signal in the time domain and outputs a frequency domain representation of the audio signal.
  • the time-to-frequency domain transformer 205 may apply any suitable discrete unitary transform, of which non limiting examples may include the discrete fourier transform (DFT) or the modified discrete cosine transform (MDCT).
  • DFT discrete fourier transform
  • MDCT modified discrete cosine transform
  • the time-to-frequency domain transformer 205 may be an analysis filter bank structure in order to generate a frequency domain representation of the signal, examples of such filter bank structures may include but are not limited to quadrature mirror filter bank (QMF) and cosine modulated pseudo QMF filter banks.
  • QMF quadrature mirror filter bank
  • the output from the time to frequency domain transformer 205 may be a series of coefficient values representing the relative amplitudes of spectral elements of the signal.
  • the time to frequency domain transformation may be shown by figure 4 by step 503.
  • the spectral coefficient values output from the time-to-frequency domain transformer 205 are passed to the spectral coefficient orderer 213.
  • the spectral coefficient orderer 213 may group the spectral coefficient values into regions/sub-bands of spectral coefficients.
  • the grouping into regions/sub-bands may be dictated by a psychoacoustic model.
  • the region/sub-band groupings may be fixed or variable over time.
  • the regions/sub-band groupings may comprise an equal number of coefficients or may comprise different numbers of coefficients.
  • the spectral coefficient orderer 213 may furthermore order, or assign a relative order value to the groupings of coefficients or the coefficients according to their perceived value to the audio signal.
  • the perceived value may be determined by the application of a psychoacoustic model to the determined spectral coefficients.
  • the spectral coefficient orderer 213 may initialise a quality loop variable.
  • the spectral coefficients, ordering information and the quality loop variable may be passed to the spectral modifier 217.
  • the grouping and ordering of the spectra! coefficients and the initialization of the quality loop variable is shown in Figure 4 by step 505.
  • the spectral modifier 217 receives the grouped and ordered spectral coefficients and the quality loop value and either passes the spectral coefficients without modifying them or modifies the spectral coefficients depending on the current iteration of the quality loop (in other words the quality loop value).
  • the spectral modifier On receiving the grouped and ordered spectral coefficients and the quality loop value, the spectral modifier checks the current loop value. The checking of the loop number is shown in figure 5 by step 551.
  • the process passes to the no modification operation as described below in step 552.
  • the spectral modifier no modification operation passes the originally received spectral coefficients and any ordering information required to the encoder 207 and quantizer 209.
  • the no modification operation is shown in figure 5 by step 552. Following the no modification operation the process passes to the increase loop number operation 557.
  • the spectra! modifier 217 process passes to the downscale operation as described below in step 553.
  • the spectral modifier 217 in the downscaie operation, downscales the input samples from the least significant spectral band as indicated according to the ordering information from the spectral coefficient orderer 213.
  • the spectral modifier 217 may in a first embodiment of the invention downscale the coefficient/sample values in the least significant spectral band by multiplying each of the coefficients/samples in the least significant band by a value less than 1.
  • The may be mathematically represented by the following equation:
  • origCoef'[k] origCoef[k] ⁇ scale, 0 ⁇ k ⁇ N
  • N is the number of coefficients/spectral samples within the sub-band being downscaled and the variable scale is the value for the downscaling operation.
  • the downscaling variable scale reduces the coefficient by 3 decibels.
  • the spectral modifier downscaling operation passes the modified spectral coefficients and any ordering information required to the encoder 207 and quantizer 209.
  • the downscaling operation is shown in figure 5 by step 553. Following the downscaling operation the process passes to the increase loop number operation 557. if the spectral modifier detects that the loop is not on its initial run but has a ioop count value which is even, in other words the 2nd, 4th, 6th etc loop after the initial loop, then the spectral modifier 217 process passes to the discard operation as described below in step 555.
  • the spectral modifier 217 in the discard operation, discards or deletes the input samples from the least significant spectral band as indicated according to the ordering information from the spectral coefficient orderer 213.
  • the spectral modifier 217 may in a first embodiment of the invention discard the coefficient/sample values in the least significant spectral band by multiplying each of the coefficients/samples in the least significant band by zero.
  • The may be mathematically represented by the following equation:
  • origCoef'[k] origCoef[k] • 0, 0 ⁇ k ⁇ N
  • is the number of coefficients/spectral samples within the sub-band being discarded.
  • the spectral modifier discarding operation passes the modified spectral coefficients and any ordering information required to the encoder 207 and quantizer 209,
  • the discarding operation is shown in figure 5 by step 555. Following the discarding operation the process passes to the increase loop number operation 557.
  • the process then increases the ioop number so that a count on the number of loops have been made.
  • the loop count is initialized with a value of 0 and is incremented with a value of 1 during each increase loop number operation.
  • any suitable manner of storing the number of iterations may be used.
  • the spectral modifier 217 operation described above shows 3 possible options available to the spectral modifier 217.
  • the spectral modifier may only receive the spectral coefficients after the initial iteration of the loop and only therefore have 2 possible options: downscale and discard - the no modification option being carried out by performing an initial bypassing of the spectral modifier 217.
  • graduated spectral modification downscaling may be employed, whereby at least two downscaling operations are performed on the least significant sub-band coefficients before the sub-band coefficients are discarded.
  • the downscaling may be carried out by subtracting an amplitude from each of the spectral coefficients in the least significant spectral sub-band.
  • the spectral modifier furthermore stores a bit associated with each sub-band.
  • a value of 1 is assigned to each sub-band where the spectral coefficients have not been discarded, in other words where the coefficients are to be included for encoding and quantization.
  • the encoder receives the spectral coefficients from the spectral modifier and performs a encoding of the spectral coefficients. Any suitable encoding of the spectral coefficients may be used.
  • the encoder 207 may for a second of further loops of the quality loop only encode the modified spectral coefficients, in other words the downscaied coefficients (the discarded coefficients are not encoded).
  • the encoded spectral coefficients may then be passed to the quantizer 209.
  • the quantizer 209 assigns a quantization step size to the encoded spectral coefficient values.
  • the step size will determine the number of bits allocated to each sub-band.
  • the assignment of the quantization step may alter from iteration to iteration so that more bits may be assigned to the more significant sub-bands and fewer bits to the downscaied and the discarded sub-bands.
  • the quantizer 209 may then apply the step size to quantize the received spectral coefficients to generate a series of quantized spectral parameters.
  • the quantizer 209 may store the quantized values and perform a re-quantization or re-assignment on the encoded spectral coefficients which have been downscaled/discarded by the spectral modifier only.
  • mapping function is not directly related to the invention and many different quantization schemes may be applied.
  • quantization schemes for example, both scalar based quantization combined with entropy coding and vector based quantization combined with lattice indexing are possible implementations for quantization according to embodiments of the invention.
  • the quantizer 209 may initially scale down the received spectral coefficients.
  • the quantized spectral coefficients are output to the distortion estimator 211 .
  • the distortion estimator 211 receives the quantized spectral coefficients and calculates the distortion or the error produced by the quantization/encoding scheme on the spectral coefficients.
  • the distortion may be calculated by summing the square of the difference of the original and dequantized spectral coefficient values for all of the non-discarded spectral coefficients. This may be carried out by performing the operations shown in the following pseudo code.
  • ⁇ nmrSfb [k] 1. 0 ;
  • variable bandsTolnclude indicates whether the spectral samples in the corresponding band are quantized
  • nBands is the number of spectral bands
  • sbOffset describes the frequency offset for each spectral band
  • thrVal contains the masking threshold for each spectral band
  • nmrSfb holds the noise to masking ratio for each spectral band that was quantized.
  • the distortion estimator outputs the error or distortion value for a current frame to the quality checker 215.
  • the quality checker 215 receives the distortion or error value from the distortion estimator 211 and performs a check on the quality of the dual domain quantisation results to determine whether or not the quantization is perceptually acceptable.
  • nmrSb[k] > 1 and nmrSb[k]> m - nmrSbOld[k ⁇ or nmrSb[k]> 0 ⁇ k ⁇ nBands otherwise
  • nmrSb is the noise to masking ratio for each frequency band - which may be calculated in the within the distortion estimator 21 1 and further passed to the quality checker 215.
  • nmrSbOld describes the noise to masking ratio for a previous frame for the corresponding frequency sub-band.
  • the variable nBands is the total number of spectral bands, and nmrThr contains maximum noise to masking threshold for each spectral band.
  • the noise increase over time for each frequency band is limited to 1 OdB and the global noise limit is calculated for example as the following 32kHz sample rate with 24 frequency band example shown below.
  • nmrThrf] ⁇ 1 , 1 , 1 , 1 , 1, 12, 12, 14, 15, 16, 16, 16, 17, 17, 17, 17, 18, 18, 18, 18, 19, 19 ⁇ ;
  • variable nmrSbOld is then initialised to a unity value at start up and updated after a successful quantization according to the following equation.
  • the quality checker determines that the encoding is OK, the quantization values are passed to the bit stream formatter 219. If however, the encoding is determined by the quality checker 215 to be 'not OK', the process is passed to the quality loop spectral modification operation - where a downscaling or discarding of the least significant sub-band will free up quantization bits in a more significant sub-band - thus decreasing the global error value.
  • the step of checking the quality of the encoding is shown in Figure 4 by the step 511.
  • the bit stream formatter/multiplexer 219 inserts the received quantization bits representing the encoded quantized spectral coefficients and outputs on output 211 the bit stream of the encoded quantized audio signal.
  • the multiplexing and outputting into the bit stream is shown in Figure 4 by step 513.
  • the decoder 108 comprises an input 301 which is connected to the bit stream unpacker or demultiplexer 303. Furthermore the bit stream unpacker 303 is connected to a dequantizer/decoder 305. The dequantiser/decoder 305 is furthermore connected to a frequency-to-time domain transformer 307. The frequency-to-time domain transformer outputs a signal on the decoder output 309.
  • the bit stream unpacker demultiplexes, partitions, or unpacks the encoded bit stream 112 into a series of bit streams.
  • the quantised values are passed to the dequantizer 305.
  • the receiving of the encoded signals is shown in Figure 7 in step 401.
  • the dequantizer/decoder 305 may apply a dequantization of the demultiplexed quantized signal dependent on the quantizer used in the quantizer 409 of the encoder 104.
  • the dequantizer/decoder 305 may apply a decoding of the dequantized signal dependent of the encoding used in the encoder 407 of the encoder 104.
  • the dequantized/decoded signal is then passed to the frequency to time domain transformer 307.
  • the dequantization/decoding of the bit stream is shown in step 405 of Figure 7.
  • the dequantizer/decoder 305 furthermore is configured to restore any deleted spectral coefficients, in one embodiment of the invention the dequantizer/decoder 305 copies the previous index spectral value. In other embodiments of the invention the dequantizer/decoder 305 copies the previous index spectral value and furthermore downscales the copies spectral value by a predetermined value. As the ear's sensitivity to detect restored deleted coefficients is quite high at low frequencies in some embodiments of the invention the dequantizer/decoder 305 only restores the high frequencies.
  • the dequantizer/decoder 305 generates a deleted spectral coefficient from examining past frame and future frame spectral coefficients with the same index value. This approach would produce a better performance when applied to low frequencies.
  • the dequantizer/decoder 305 may perform the interpolation on the subband/frequency band basis to get best possible perceived output.
  • the frequency to time domain transformer receives the output of the dequantization/decoding and then transforms from the frequency domain to the time domain the received parameters in a complimentary manner to the time-to- frequency domain transformation carried out by the time to frequency domain transformer 205 in the encoder 108.
  • the frequency-to-domain transformation is shown in Figure 7 by step 407.
  • the output of the frequency-to-time domain transformer 307 is then output on the output 309.
  • embodiments of the invention operating within a codec within an electronic device 10
  • the invention as described below may be implemented as part of any variable rate/adaptive rate audio (or speech) codec.
  • embodiments of the invention may be implemented in an audio codec which may implement audio coding over fixed or wired communication paths.
  • user equipment may comprise an audio codec such as those described in embodiments of the invention above.
  • user equipment is intended to cover any suitable type of wireless user equipment, such as mobile telephones, portable data processing devices or portable web browsers.
  • elements of a public land mobile network may also comprise audio codecs as described above.
  • PLMN public land mobile network
  • the various embodiments of the invention may be implemented in hardware or special purpose circuits, software, logic or any combination thereof.
  • some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software which may be executed by a controller, microprocessor or other computing device, although the invention is not limited thereto.
  • firmware or software which may be executed by a controller, microprocessor or other computing device, although the invention is not limited thereto.
  • While various aspects of the invention may be illustrated and described as block diagrams, flow charts, or using some other pictorial representation, it is well understood that these blocks, apparatus, systems, techniques or methods described herein may be implemented in, as non-limiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof.
  • the embodiments of the invention may be implemented as a chipset, in other words a series of integrated circuits communicating among each other.
  • the chipset may comprise microprocessors arranged to run code, application specific integrated circuits (ASICs), or programmable digital signal processors for performing the operations described above.
  • ASICs application specific integrated circuits
  • programmable digital signal processors for performing the operations described above.
  • the embodiments of this invention may be implemented by computer software executable by a data processor of the mobile device, such as in the processor entity, or by hardware, or by a combination of software and hardware. Further in this regard it should be noted that any blocks of the logic flow as in the Figures may represent program steps, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions.
  • the memory may be of any type suitable to the local technical environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory and removable memory.
  • the data processors may be of any type suitable to the local technical environment, and may include one or more of general purpose computers, special purpose computers, microprocessors, digital signal processors (DSPs) and processors based on multi-core processor architecture, as non-limiting examples.
  • Embodiments of the inventions may be practiced in various components such as integrated circuit modules.
  • the design of integrated circuits is by and large a highly automated process.
  • Complex and powerful software tools are available for converting a logic fevel design into a semiconductor circuit design ready to be etched and formed on a semiconductor substrate.
  • Programs such as those provided by Synopsys, Inc. of Mountain View, California and Cadence Design, of San Jose, California automatically route conductors and locate components on a semiconductor chip using well established rules of design as well as libraries of pre-stored design modules.
  • the resultant design in a standardized electronic format (e.g., Opus, GDSlI, or the like) may be transmitted to a semiconductor fabrication facility or "fab" for fabrication.

Landscapes

  • Engineering & Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Spectroscopy & Molecular Physics (AREA)
  • Computational Linguistics (AREA)
  • Signal Processing (AREA)
  • Health & Medical Sciences (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Human Computer Interaction (AREA)
  • Acoustics & Sound (AREA)
  • Multimedia (AREA)
  • Compression, Expansion, Code Conversion, And Decoders (AREA)

Abstract

An encoder for encoding an audio signal configured to: generate at least one encoded signal dependent on the audio signal; determine at least one quality indicator, wherein each quality indicator is dependent on the relative difference between an associated one of the at least one encoded signal and the audio signal; select to output one of the at least one encoded signal dependent on the associated quality indicator value.

Description

An Encoder
Field of the Invention
The present invention relates to coding, and in particular, but not exclusively to speech or audio coding.
Background of the Invention
Audio signals, like speech or music, are encoded for example for enabling an efficient transmission or storage of the audio signals.
Audio encoders and decoders are used to represent audio based signals, such as music and background noise. These types of coders typically do not utilise a speech model for the coding process, rather they use processes for representing all types of audio signals, including speech.
Speech encoders and decoders (codecs) are usually optimised for speech signals, and can operate at either a fixed or variable bit rate.
An audio codec can also be configured to operate with varying bit rates. At lower bit rates, such an audio codec may work with speech signals at a coding rate equivalent to a pure speech codec, At higher bit rates, the audio codec may code any signal including music, background noise and speech, with higher quality and performance.
In some audio codecs the input signal is divided into a limited number of bands. Each of the band signals may be quantized. From the theory of psychoacoustics it is known that the highest frequencies in the spectrum are perceptually less important than the low frequencies. This in some audio codecs is reflected by a bit allocation where fewer bits are allocated to high frequency signals than low frequency signals.
Within an audio coding system encoder, the encoder is responsible for adapting the received audio data in order that it may be output at a bit rate which the channel bandwidth conditions are capable of accepting. Ideally, the encoding process discards only the irrelevant information from the audio signal received and the decoder is capable of simply reversing the encoding process to obtain the decoded audio signal with little or no audible degradation.
One method used by the encoder to reduce to the number of bits required to encode an audio signal is that of quantization. Well known perceptual quantisation schemes such as the one used in MP3 and AAC coding follow known psychoacoustic rules. In the psychoacoustic rules the spectral frame is divided into a series of frequency bands or regions and for each frequency band a quantization step size, also known as the scaie factor, is assigned.
The value of the scale factor is changed (in other words either increased or decreased) so that the quantization noise (in other words the error produced during quantization) is below a masking threshold.
For example where a first frequency component has a small spectral value when compared against a neighbouring frequency component, it is possible to quantize the first frequency component more roughly without degrading the final audio quality significantly as the human ear will not pick out the first frequency component significantly clearly as it the neighbouring frequency component would perceptually dominate. When all scale factors are below a masking threshold then the untrained listener wϋl not perceive significant loss in quality and a 'transparent1 quality may be achieved.
In practical implementations of quantization, the scale factors have to be themselves quantized and transmitted to the receiver. The high number of scale factors being using increases the bit rate associated with the scale factor coding techniques described above. Thus whilst the encoding system is suitable for high bit rate coding the overheads introduced by transmitting scale factors to the decoder force a reduction in the quality of the encoding of the audio signal in low bit rate situations.
Summary of the Invention
This invention proceeds from the consideration that whilst scale factor transmission is acceptable in high bit rate environments a system of audio quantization may be implemented which produces an acceptable level of quantization without needing as many parameters to be sent to the receiver which therefore allows encoded audio signals to be transmitted without suffering from over harsh coding.
Embodiments of the present invention aim to address the above problem.
There is provided according to a first aspect of the present invention an encoder for encoding an audio signal configured to: generate at least one encoded signal dependent on the audio signal; determine at least one quality indicator, wherein each quality indicator is dependent on the relative difference between an associated one of the at least one encoded signal and the audio signal; select to output one of the at least one encoded signal dependent on the associated quality indicator value. The encoder may be further configured to: generate at least one spectral representation of the audio signal, wherein the spectral representation may comprise at least two sub-bands, each sub-band may comprise at least one spectral coefficient.
The encoder may be further configured to generate a first of the at least one spectral representation of the audio signal by applying at least one of the following: a shifted discrete fourier transform; a modified discrete cosine transform; and a discrete unitary transform:
The encoder may be further configured to generate further spectral representations of the audio signal by applying at least one gain factor to the at least one spectral coefficient of at least one of the at least two sub-bands.
The at least one gain factor may comprise at least one of the following multiplication gain factors: 0 and 0.5.
The encoder may be further configured to order the at least two sub-bands.
The encoder may be further configured to generate the further spectra! representations of the audio signal by progressively applying, dependent on the order of the at least two sub-bands, at least one gain factor to the at least one spectral coefficient of at least one of the at least two sub-bands.
The encoder may be further configured to generate an indicator to indicate which sub-band spectral coefficients are modified by the at least one gain factor. The encoder may be further configured to order the at least two sub-bands according to a psycho acoustical model.
The encoder may be further configured to progressively generate the further spectral representations of the audio signal until the quality indicator is equal to or greater than a quality threshold.
According to a second aspect of the invention there is provided a decoder for decoding an encoded signal, configured to: partition the received encoded signal into at least a first and second part; decode the first part to generate at least one gain indicator; generate at least one frequency representation of the encoded signal dependent on the second part and the at least one gain indicator, wherein the at least one frequency representation comprises at [east one spectral coefficient.
The decoder may be further configured to generate at least one initial frequency representation of the encoded signal dependent on the second part of the received encoded signal.
The at least one gain indicator may comprise at least one deleted spectral coefficient index.
The decoder is preferably further configured to insert an at least one predetermined spectral coefficient value into the initial frequency representation determined by the at least one deleted spectral coefficient index.
The at least one predetermined spectral coefficient value is preferably a null value. The at least one gain indicator may further comprise at least one spectral coefficient gain factor.
The decoder is preferably further configured to multiply a spectral coefficient of the initial frequency representation determined by the at least one deleted spectral coefficient index by at least one spectral coefficient gain factor.
According to a third aspect of the invention there is provided a method for encoding an audio signal, comprising: generating at least one encoded signal dependent on the audio signal; determining at least one quality indicator, wherein each quality indicator is dependent on the relative difference between an associated one of the at least one encoded signal and the audio signal; and selecting to output one of the at least one encoded signal dependent on the associated quality indicator value.
The method may further comprise: generating at least one spectral representation of the audio signal, wherein the spectral representation may comprise at least two sub-bands, each sub-band may comprise at least one spectral coefficient.
The method may further comprise generating a first of the at least one spectral representation of the audio signal by applying at least one of the following: a shifted discrete fourier transform; a modified discrete cosine transform; and a discrete unitary transform.
The method may further comprise generating further spectral representations of the audio signal by applying at least one gain factor to the at least one spectral coefficient of at least one of the at least two sub-bands. The at least one gain factor may comprise at least one of the following multiplication gain factors: 0 and 0,5.
The method may further comprise ordering the at least two sub-bands.
The method may further comprise generating the further spectral representations of the audio signal by progressively applying, dependent on the order of the at least two sub-bands, at least one gain factor to the at least one spectral coefficient of at least one of the at least two sub-bands.
The method may further comprise generating an indicator to indicate which sub- band spectral coefficients are modified by the at least one gain factor.
The method may further comprise ordering the at least two sub-bands according to a psycho acoustical model.
The method may further comprise progressively generating the further spectral representations of the- audio signal until the quality indicator is equal to or greater than a quality threshold.
According to a fourth aspect of the present invention there is provided a method for decoding an encoded signal, comprising: partitioning the received encoded signal into at least a first and second part; decoding the first part to generate at least one gain indicator; generating at least one frequency representation of the encoded signal dependent on the second part and the at least one gain indicator, wherein the at least one frequency representation comprises at . least one spectral coefficient. The method may further comprise generating at least one initial frequency representation of the encoded signal dependent on the second part of the received encoded signal.
The at least one gain indicator may comprise at least one deleted spectral coefficient index.
The method may further comprise inserting an at least one predetermined spectral coefficient value into the initial frequency representation determined by the at least one deleted spectral coefficient index.
The at least one predetermined spectral coefficient value is preferably a null value.
The at least one gain indicator may further comprise at least one spectral coefficient gain factor.
The method may further comprise multiplying a spectral coefficient of the initial frequency representation determined by the at least one deleted spectral coefficient index by at least one spectral coefficient gain factor.
An apparatus may comprise an encoder as described above.
An apparatus may comprise a decoder as described above.
An electronic device may comprise an encoder as described above.
An electronic device may comprise a decoder as described above. A chipset may comprise an encoder as described above.
A chipset may comprise a decoder as described above. According to a fifth aspect of the present invention there is provided a computer program product configured to perform a method for encoding an audio signal, comprising: generating at least one encoded signal dependent on the audio signal; determining at least one quality indicator, wherein each quality indicator is dependent on the relative difference between an associated one of the at least one encoded signal and the audio signal; selecting to output one of the at least one encoded signal dependent on the associated quality indicator value.
According to a sixth aspect of the present invention there is provided a computer program product configured to perform a method for decoding an encoded signal, comprising: partitioning the received encoded signal into at least a first and second part; decoding the first part to generate at least one gain indicator; generating at least one frequency representation of the encoded signal dependent on the second part and the at least one gain indicator, wherein the at least one frequency representation comprises at least one spectral coefficient.
According to a seventh aspect of the present invention there is provided an encoder for encoding an audio signal; comprising: encoding means to generate at least one encoded signal dependent on the audio signal; processing means to determine at least one quality indicator, wherein each quality indicator is dependent on the relative difference between an associated one of the at least one encoded signal and the audio signal; and selection means to select to output one of the at least one encoded signal dependent on the associated quality indicator value. According to an eighth aspect of the present invention there is provided a decoder for decoding an encoded audio signal, comprising: processing means to partition the received encoded signal into at least a first and second part; decoding means to decode the first part to generate at least one gain indicator; further processing means to generate at least one frequency representation of the encoded signal dependent on the second part and the at least one gain indicator, wherein the at least one frequency representation comprises at least one spectral coefficient.
Brief Description of Drawings
For better understanding of the present invention, reference will now be made by way of example to the accompanying drawings in which:
Figure 1 shows schematically an electronic device employing embodiments of the invention;
Figure 2 shows schematically an audio codec system employing embodiments of the present invention;
Figure 3 shows schematically an encoder part of the audio codec system shown in figure 2; Figure 4 shows a flow diagram illustrating the operation of an embodiment of the audio encoder as shown in Figure 3 according to the present invention;
Figure 5 shows a flow diagram illustrating the operation of the spectral modifier part of the audio encoder as shown in Figure 3 according to the present invention;
Figure 6 shows schematically a decoder part of the audio codec system shown in Figure 2; and Figure 7 shows a flow diagram illustrating the operation of an embodiment of the audio decoder as shown in Figure 6 according to the present invention.
Description of Preferred Embodiments of the Invention
The following describes in more detail possible mechanisms for the provision of a low complexity monochannel or multichannel audio coding system. In this regard reference is first made to figure 1 schematic block diagram of an exemplary electronic device 10, which may incorporate a codec according to an embodiment of the invention.
The electronic device 10 may for example be a mobile terminal or user equipment of a wireless communication system.
The electronic device 10 comprises a microphone 1 1 , which is linked via an analogue-to-digital converter 14 to a processor 21. The processor 21 is further linked via a digital-to-analogue converter 32 to loudspeakers 33. The processor 21 is further linked to a transceiver (TX/RX) 13, to a user interface (Ul) 15 and to a memory 22.
The processor 21 may be configured to execute various program codes. The implemented program codes comprise an audio encoding code for encoding a combined audio signal and code to extract and encode side information pertaining to the spatial information of the multiple channels. The implemented program codes 23 further comprise an audio decoding code. The implemented program codes 23 may be stored for example in the memory 22 for retrieval by the processor 21 whenever needed. The memory 22 could further provide a section 24 for storing data, for example data that has been encoded in accordance with the invention.
The encoding and decoding code may in embodiments of the invention be implemented in hardware or firmware.
The user interface 15 enables a user to input commands to the electronic device 10, for example via a keypad, and/or to obtain information from the electronic device 10, for example via a display. The transceiver 13 enables a communication with other electronic devices, for example via a wireless communication network.
It is to be understood again that the structure of the electronic device 10 could be supplemented and varied in many ways.
A user of the electronic device 10 may use the microphone 1 1 for inputting speech that is to be transmitted to some other electronic device or that is to be stored in the data section 24 of the memory 22. A corresponding application has been activated to this end by the user via the user interface 15. This application, which may be run by the processor 21 , causes the processor 21 to execute the encoding code stored in the memory 22.
The analogue-to-digital converter 14 converts the input analogue audio signal into a digital audio signal and provides the digital audio signal to the processor 21.
The processor 21 may then process the digital audio signal in the same way as described with reference to figures 2 and 3. The resulting bit stream is provided to the transceiver 13 for transmission to another electronic device. Alternatively, the coded data could be stored in the data section 24 of the memory 22, for instance for a later transmission or for a later presentation by the same electronic device 10.
The electronic device 10 could also receive a bit stream with correspondingly encoded data from another electronic device via its transceiver 13. In this case, the processor 21 may execute the decoding program code stored in the memory 22. The processor 21 decodes the received data, and provides the decoded data to the digital-to-analogue converter 32. The digital-to-analogue converter 32 converts the digital decoded data into analogue audio data and outputs them via the loudspeakers 33. Execution of the decoding program code could be triggered as well by an application that has been called by the user via the user interface 15.
The received encoded data could also be stored instead of an immediate presentation via the loudspeakers 33 in the data section 24 of the memory 22, for instance for enabling a later presentation or a forwarding to still another electronic device.
It would be appreciated that the schematic structures described in figures 2, 3, 4 and 7 and the method steps in figures 5, 6 and 8 represent only a part of the operation of a complete audio codec as exemplarily shown implemented in the electronic device shown in figure 1.
The general operation of audio codecs as employed by embodiments of the invention is shown in figure 2. General audio coding/decoding systems consist of an encoder and a decoder, as illustrated schematically in figure 2. Illustrated is a system 102 with an encoder 104, a storage or media channel 106 and a decoder 108.
The encoder 104 compresses an input audio signal 1 10 producing a bit stream 112, which is either stored or transmitted through a media channel 106. The bit stream 112 can be received within the decoder 108. The decoder 108 decompresses the bit stream 112 and produces an output audio signal 114. The bit rate of the bit stream 112 and the quality of the output audio signal 114 in relation to the input signal 110 are the main features, which define the performance of the coding system 102.
Figure 3 depicts schematically an encoder according to an exemplary embodiment of the invention. The encoder 104 comprises the input 203 which is arranged to receive an audio signal input. The audio input 203 is connected to an input of a time to frequency domain transformer 205. The time to frequency domain transformer 205 has an output which is connected to an input of a spectral coefficient orderer 213. The spectral coefficient orderer 213 has an output which is connected to an input of a spectral modifier 217. The spectral modifier 217 has an output which is connected to an input of an encoder 207. The encoder 207 has an output which is connected to an input of a quantizer 209. The quantizer 209 has an output which is connected to an input of a distortion estimator 21 1. The distortion estimator 211 has an output which is connected to an input of a quality checker 215. The quality checker 215 has an output which is connected to a further input of the spectral modifier 217, and a further output which is connected to a bitstream formatter 219. The bitstream formatter 219 has an output which is connected to the encoder output 221.
The encoded audio output 221 is the signal output comprising the encoded signal. With respect to figure 4 the operation of the encoder 104 as shown in figure 3 is described in further detail.
The audio signal is received by the coder 104. In a first embodiment of the invention, the audio signal is a digitally sampled signal. In other embodiments of the present invention, the audio input may be an analogue audio signal, for example from a microphone 6, which is analogue-to-digitally converted (A-D). In further embodiments of the invention, the audio input is converted from a pulse modulation digital signal to amplitude modulation digital signal.
The reception of the audio signal is shown in Figure 4 by step 501.
The input 203 is connected to a time-to-frequency domain transformer 205. The time-to-frequency domain transformer receives the audio signal in the time domain and outputs a frequency domain representation of the audio signal.
In embodiments of the invention, the time-to-frequency domain transformer 205 may apply any suitable discrete unitary transform, of which non limiting examples may include the discrete fourier transform (DFT) or the modified discrete cosine transform (MDCT).
In further embodiments of the invention, the time-to-frequency domain transformer 205 may be an analysis filter bank structure in order to generate a frequency domain representation of the signal, examples of such filter bank structures may include but are not limited to quadrature mirror filter bank (QMF) and cosine modulated pseudo QMF filter banks. The output from the time to frequency domain transformer 205 may be a series of coefficient values representing the relative amplitudes of spectral elements of the signal.
The time to frequency domain transformation may be shown by figure 4 by step 503.
The spectral coefficient values output from the time-to-frequency domain transformer 205 are passed to the spectral coefficient orderer 213. The spectral coefficient orderer 213 may group the spectral coefficient values into regions/sub-bands of spectral coefficients. The grouping into regions/sub-bands may be dictated by a psychoacoustic model. The region/sub-band groupings may be fixed or variable over time. Furthermore, the regions/sub-band groupings may comprise an equal number of coefficients or may comprise different numbers of coefficients.
The spectral coefficient orderer 213 may furthermore order, or assign a relative order value to the groupings of coefficients or the coefficients according to their perceived value to the audio signal. The perceived value may be determined by the application of a psychoacoustic model to the determined spectral coefficients.
Furthermore the spectral coefficient orderer 213 may initialise a quality loop variable. The spectral coefficients, ordering information and the quality loop variable may be passed to the spectral modifier 217.
The grouping and ordering of the spectra! coefficients and the initialization of the quality loop variable is shown in Figure 4 by step 505. The spectral modifier 217 receives the grouped and ordered spectral coefficients and the quality loop value and either passes the spectral coefficients without modifying them or modifies the spectral coefficients depending on the current iteration of the quality loop (in other words the quality loop value).
With respect to figure 5, the operation of the spectral modifier 217 as used in embodiments of the invention is shown in further detail.
On receiving the grouped and ordered spectral coefficients and the quality loop value, the spectral modifier checks the current loop value. The checking of the loop number is shown in figure 5 by step 551.
If the spectral modifier detects that the loop is on its initial run, which for a first embodiment of the invention is shown by a quality loop value of O1 then the process passes to the no modification operation as described below in step 552.
The spectral modifier no modification operation passes the originally received spectral coefficients and any ordering information required to the encoder 207 and quantizer 209.
The no modification operation is shown in figure 5 by step 552. Following the no modification operation the process passes to the increase loop number operation 557.
If the spectra! modifier detects that the loop is not on its initial run but has a loop count value which is odd, in other words the 1st, 3rd, 5th etc loop after the initial loop, then the spectra! modifier 217 process passes to the downscale operation as described below in step 553.
The spectral modifier 217, in the downscaie operation, downscales the input samples from the least significant spectral band as indicated according to the ordering information from the spectral coefficient orderer 213.
The spectral modifier 217 may in a first embodiment of the invention downscale the coefficient/sample values in the least significant spectral band by multiplying each of the coefficients/samples in the least significant band by a value less than 1. The may be mathematically represented by the following equation:
origCoef'[k] = origCoef[k] ■ scale, 0 ≤ k < N
where N is the number of coefficients/spectral samples within the sub-band being downscaled and the variable scale is the value for the downscaling operation.
in a first embodiment of the invention, the downscaling variable scale reduces the coefficient by 3 decibels.
The spectral modifier downscaling operation passes the modified spectral coefficients and any ordering information required to the encoder 207 and quantizer 209.
The downscaling operation is shown in figure 5 by step 553. Following the downscaling operation the process passes to the increase loop number operation 557. if the spectral modifier detects that the loop is not on its initial run but has a ioop count value which is even, in other words the 2nd, 4th, 6th etc loop after the initial loop, then the spectral modifier 217 process passes to the discard operation as described below in step 555.
The spectral modifier 217, in the discard operation, discards or deletes the input samples from the least significant spectral band as indicated according to the ordering information from the spectral coefficient orderer 213.
The spectral modifier 217 may in a first embodiment of the invention discard the coefficient/sample values in the least significant spectral band by multiplying each of the coefficients/samples in the least significant band by zero. The may be mathematically represented by the following equation:
origCoef'[k] = origCoef[k] • 0, 0 < k < N
where Ν is the number of coefficients/spectral samples within the sub-band being discarded.
The spectral modifier discarding operation passes the modified spectral coefficients and any ordering information required to the encoder 207 and quantizer 209,
The discarding operation is shown in figure 5 by step 555. Following the discarding operation the process passes to the increase loop number operation 557.
The process then increases the ioop number so that a count on the number of loops have been made. In the example described above the loop count is initialized with a value of 0 and is incremented with a value of 1 during each increase loop number operation.
This increasing of the quality loop variable is shown in step 557 of Figure 5.
In further embodiments of the invention any suitable manner of storing the number of iterations may be used.
Furthermore the spectral modifier 217 operation described above shows 3 possible options available to the spectral modifier 217. In further embodiments of the invention the spectral modifier may only receive the spectral coefficients after the initial iteration of the loop and only therefore have 2 possible options: downscale and discard - the no modification option being carried out by performing an initial bypassing of the spectral modifier 217.
In further embodiments of the invention graduated spectral modification downscaling may be employed, whereby at least two downscaling operations are performed on the least significant sub-band coefficients before the sub-band coefficients are discarded.
In further embodiments of the invention the downscaling may be carried out by subtracting an amplitude from each of the spectral coefficients in the least significant spectral sub-band.
In embodiments of the invention, the spectral modifier furthermore stores a bit associated with each sub-band. A value of 1 is assigned to each sub-band where the spectral coefficients have not been discarded, in other words where the coefficients are to be included for encoding and quantization. The encoder receives the spectral coefficients from the spectral modifier and performs a encoding of the spectral coefficients. Any suitable encoding of the spectral coefficients may be used.
In some embodiments of the invention the encoder 207 may for a second of further loops of the quality loop only encode the modified spectral coefficients, in other words the downscaied coefficients (the discarded coefficients are not encoded).
The encoded spectral coefficients may then be passed to the quantizer 209.
The quantizer 209 assigns a quantization step size to the encoded spectral coefficient values. The step size will determine the number of bits allocated to each sub-band. The assignment of the quantization step may alter from iteration to iteration so that more bits may be assigned to the more significant sub-bands and fewer bits to the downscaied and the discarded sub-bands.
The quantizer 209 may then apply the step size to quantize the received spectral coefficients to generate a series of quantized spectral parameters.
In some embodiments of the invention the quantizer 209 may store the quantized values and perform a re-quantization or re-assignment on the encoded spectral coefficients which have been downscaled/discarded by the spectral modifier only.
The mapping function is not directly related to the invention and many different quantization schemes may be applied. For example, both scalar based quantization combined with entropy coding and vector based quantization combined with lattice indexing are possible implementations for quantization according to embodiments of the invention.
In some embodiments of the invention the quantizer 209 may initially scale down the received spectral coefficients.
The quantized spectral coefficients are output to the distortion estimator 211 .
The encoding and quantization of the sub-band coefficients is shown in Figure 4 by step 507,
The distortion estimator 211 receives the quantized spectral coefficients and calculates the distortion or the error produced by the quantization/encoding scheme on the spectral coefficients. The distortion may be calculated by summing the square of the difference of the original and dequantized spectral coefficient values for all of the non-discarded spectral coefficients. This may be carried out by performing the operations shown in the following pseudo code.
GetNMR (Pointer origCoef, Pointer DequantCoef, Pointer bandsToInclude ,
Integer nBands, Pointer sbOffεet, Pointer thrVal, Pointer nmrSb)
{ for (k = 0 , bin = 0 ; k < nBands ; k++)
{ nmrSfb [k] = 1. 0 ;
/*-- Spectral band encoded. --*/ if (bandsTolnclude [k] == 1)
{ gErr = 0. Of ; for (j = sbOf fset [k] ; j < sbOf f set [k + I ] ; j ++ , bin++)
{ tmp = OrigCoef [j ] - DequantCoef [bin] ; qErr += tmp * tmp; } nmrSfb [k] = qErr / thrVal [k] ; } }
where in the pseudo code above origCoef and DequantCoef contain the original and the dequantized spectral coefficients respectively, the variable bandsTolnclude indicates whether the spectral samples in the corresponding band are quantized, nBands is the number of spectral bands, sbOffset describes the frequency offset for each spectral band, thrVal contains the masking threshold for each spectral band, and nmrSfb holds the noise to masking ratio for each spectral band that was quantized.
As would be appreciated by the person skilled in the art, further methods of estimating the error produced by the spectral coefficient quantization values may be used to produce a distortion value.
The distortion estimator outputs the error or distortion value for a current frame to the quality checker 215.
The calculation of the distortion value for a frame can be seen in Figure 4 in step 509.
The quality checker 215 receives the distortion or error value from the distortion estimator 211 and performs a check on the quality of the dual domain quantisation results to determine whether or not the quantization is perceptually acceptable.
In one embodiment of the invention, the quality checker 215 decides the quality noise status according to the following equation: [- 1, any of B[k] = -1 , 0 < k < nBands qNoiseStatus = <
[ 1, otherwise nmrSb[k] > 1 and nmrSb[k]> m - nmrSbOld[k\or nmrSb[k]>
Figure imgf000025_0002
0 ≤ k < nBands
Figure imgf000025_0001
otherwise
where nmrSb is the noise to masking ratio for each frequency band - which may be calculated in the within the distortion estimator 21 1 and further passed to the quality checker 215. nmrSbOld describes the noise to masking ratio for a previous frame for the corresponding frequency sub-band. The variable nBands is the total number of spectral bands, and nmrThr contains maximum noise to masking threshold for each spectral band.
The above equation iimits the increase in the noise to masking ratio on a frequency band basis and also applies a global noise limit.
Thus, in embodiments of the invention, the noise increase over time for each frequency band is limited to 1 OdB and the global noise limit is calculated for example as the following 32kHz sample rate with 24 frequency band example shown below.
nmrThrf] = {1 , 1 , 1 , 1 , 1 , 1, 12, 12, 14, 15, 16, 16, 16, 16, 17, 17, 17, 17, 18, 18, 18, 18, 19, 19};
The variable nmrSbOld is then initialised to a unity value at start up and updated after a successful quantization according to the following equation.
bandsToInclude[k]= l L J Q ≤k < nBands
Figure imgf000025_0003
otherwise The quality checker 215 then makes a global check to see whether the quantization is acceptable. This decision in embodiments of the invention may be made according to the decision shown below:
I Encoding ok, qNoiseStatus = — 1 encStatus = <
[Encoding not ok, otherwise
If the quality checker determines that the encoding is OK, the quantization values are passed to the bit stream formatter 219. If however, the encoding is determined by the quality checker 215 to be 'not OK', the process is passed to the quality loop spectral modification operation - where a downscaling or discarding of the least significant sub-band will free up quantization bits in a more significant sub-band - thus decreasing the global error value.
The step of checking the quality of the encoding is shown in Figure 4 by the step 511.
Where the encoding has been successful, the bit stream formatter/multiplexer 219 inserts the received quantization bits representing the encoded quantized spectral coefficients and outputs on output 211 the bit stream of the encoded quantized audio signal.
The multiplexing and outputting into the bit stream is shown in Figure 4 by step 513.
Such a process as described above, would be advantageous as it would improve the quality of the audio signal without requiring a significant amount of non-audio signal data to be transmitted. To further assist with the understanding of the invention, the operation of the decoder 108 with respect of the embodiments of the invention are shown with respect to the decoder schematically shown in Figure 6 and the flow chart showing the operation of the decoder in Figure 7.
The decoder 108 comprises an input 301 which is connected to the bit stream unpacker or demultiplexer 303. Furthermore the bit stream unpacker 303 is connected to a dequantizer/decoder 305. The dequantiser/decoder 305 is furthermore connected to a frequency-to-time domain transformer 307. The frequency-to-time domain transformer outputs a signal on the decoder output 309.
The bit stream unpacker demultiplexes, partitions, or unpacks the encoded bit stream 112 into a series of bit streams. The quantised values are passed to the dequantizer 305.
The receiving of the encoded signals is shown in Figure 7 in step 401.
Furthermore, the demultiplexing or unpacking of the bit streams is shown in step 403 of figure 7.
The dequantizer/decoder 305 may apply a dequantization of the demultiplexed quantized signal dependent on the quantizer used in the quantizer 409 of the encoder 104.
Furthermore the dequantizer/decoder 305 may apply a decoding of the dequantized signal dependent of the encoding used in the encoder 407 of the encoder 104. The dequantized/decoded signal is then passed to the frequency to time domain transformer 307. The dequantization/decoding of the bit stream is shown in step 405 of Figure 7.
In some embodiments of the invention the dequantizer/decoder 305 furthermore is configured to restore any deleted spectral coefficients, in one embodiment of the invention the dequantizer/decoder 305 copies the previous index spectral value. In other embodiments of the invention the dequantizer/decoder 305 copies the previous index spectral value and furthermore downscales the copies spectral value by a predetermined value. As the ear's sensitivity to detect restored deleted coefficients is quite high at low frequencies in some embodiments of the invention the dequantizer/decoder 305 only restores the high frequencies.
In other embodiments of the invention the dequantizer/decoder 305 generates a deleted spectral coefficient from examining past frame and future frame spectral coefficients with the same index value. This approach would produce a better performance when applied to low frequencies. In some embodiments of the invention the dequantizer/decoder 305 may perform the interpolation on the subband/frequency band basis to get best possible perceived output.
The frequency to time domain transformer receives the output of the dequantization/decoding and then transforms from the frequency domain to the time domain the received parameters in a complimentary manner to the time-to- frequency domain transformation carried out by the time to frequency domain transformer 205 in the encoder 108.
The frequency-to-domain transformation is shown in Figure 7 by step 407. The output of the frequency-to-time domain transformer 307 is then output on the output 309.
The outputting of the audio is shown in Figure 7 by step 409.
The embodiments of the invention described above describe the codec in terms of separate encoders 104 and decoders 108 apparatus in order to assist the understanding of the processes involved. However, it would be appreciated that the apparatus, structures and operations may be implemented as a single encoder-decoder apparatus/structure/operation. Furthermore in some embodiments of the invention the coder and decoder may share some/or all common elements.
Although the above examples describe embodiments of the invention operating within a codec within an electronic device 10, it would be appreciated that the invention as described below may be implemented as part of any variable rate/adaptive rate audio (or speech) codec. Thus, for example, embodiments of the invention may be implemented in an audio codec which may implement audio coding over fixed or wired communication paths.
Thus user equipment may comprise an audio codec such as those described in embodiments of the invention above.
It shall be appreciated that the term user equipment is intended to cover any suitable type of wireless user equipment, such as mobile telephones, portable data processing devices or portable web browsers.
Furthermore elements of a public land mobile network (PLMN) may also comprise audio codecs as described above. In general, the various embodiments of the invention may be implemented in hardware or special purpose circuits, software, logic or any combination thereof. For example, some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software which may be executed by a controller, microprocessor or other computing device, although the invention is not limited thereto. While various aspects of the invention may be illustrated and described as block diagrams, flow charts, or using some other pictorial representation, it is well understood that these blocks, apparatus, systems, techniques or methods described herein may be implemented in, as non-limiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof.
For example the embodiments of the invention may be implemented as a chipset, in other words a series of integrated circuits communicating among each other. The chipset may comprise microprocessors arranged to run code, application specific integrated circuits (ASICs), or programmable digital signal processors for performing the operations described above.
The embodiments of this invention may be implemented by computer software executable by a data processor of the mobile device, such as in the processor entity, or by hardware, or by a combination of software and hardware. Further in this regard it should be noted that any blocks of the logic flow as in the Figures may represent program steps, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions.
The memory may be of any type suitable to the local technical environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory and removable memory. The data processors may be of any type suitable to the local technical environment, and may include one or more of general purpose computers, special purpose computers, microprocessors, digital signal processors (DSPs) and processors based on multi-core processor architecture, as non-limiting examples.
Embodiments of the inventions may be practiced in various components such as integrated circuit modules. The design of integrated circuits is by and large a highly automated process. Complex and powerful software tools are available for converting a logic fevel design into a semiconductor circuit design ready to be etched and formed on a semiconductor substrate.
Programs, such as those provided by Synopsys, Inc. of Mountain View, California and Cadence Design, of San Jose, California automatically route conductors and locate components on a semiconductor chip using well established rules of design as well as libraries of pre-stored design modules. Once the design for a semiconductor circuit has been completed, the resultant design, in a standardized electronic format (e.g., Opus, GDSlI, or the like) may be transmitted to a semiconductor fabrication facility or "fab" for fabrication.
The foregoing description has provided by way of exemplary and non-limiting examples a full and informative description of the exemplary embodiment of this invention. However, various modifications and adaptations may become apparent to those skilled in the relevant arts in view of the foregoing description, when read in conjunction with the accompanying drawings and the appended claims. However, all such and similar modifications of the teachings of this invention will still fall within the scope of this invention as defined in the appended claims.

Claims

Claims
1. An encoder for encoding an audio signal configured to: generate at least one encoded signal dependent on the audio signal; determine at least one quality indicator, wherein each quality indicator is dependent on the relative difference between an associated one of the at least one encoded signal and the audio signal; select to output one of the at least one encoded signal dependent on the associated quality indicator value.
2. The encoder as claimed in claim 1 , further configured to: generate at least one spectral representation of the audio signal, wherein the spectral representation comprises at least two sub-bands, each sub-band comprising at least one spectral coefficient.
3. The encoder as claimed in claim 2, further configured to generate a first of the at least one spectral representation of the audio signal by applying at least one of the following: a shifted discrete fourier transform; a modified discrete cosine transform; and a discrete unitary transform.
4. The encoder as claimed in claim 3, further configured to generate further spectral representations of the audio signal by applying at least one gain factor to the at least one spectral coefficient of at least one of the at least two sub- bands.
5. The encoder as claimed in claim 4, wherein the at least one gain factor comprises at least one of the following multiplication gain factors: 0 and 0.5.
6. The encoder as claimed in claim 3 to 5, further configured to order the at least two sub-bands.
7. The encoder as claimed in claim 6 when dependent on claim 4, further configured to generate the further spectral representations of the audio signal by progressively applying, dependent on the order of the at least two sub- bands, at least one gain factor to the at least one spectral coefficient of at least one of the at least two sub-bands.
8. The encoder as claimed in claims 5 to 7, further configured to generate an indicator to indicate which sub-band spectral coefficients are modified by the at least one gain factor.
9. The encoder as claimed in claim 3 to 5, further configured to order the at least two sub-bands according to a psycho acoustical model.
10. The encoder as claimed in claim 7, further configured to progressively generate the further spectral representations of the audio signal until the quality indicator is equal to or greater than a quality threshold.
11. A decoder for decoding an encoded signal, configured to: partition the received encoded signal into at least a first and second part; decode the first part to generate at feast one gain indicator; generate at ieast one frequency representation of the encoded signal dependent on the second part and the at least one gain indicator, wherein the at least one frequency representation comprises at least one spectral coefficient.
12. The decoder as claimed in claim 11, further configured to generate at least one initial frequency representation of the encoded signal dependent on the second part of the received encoded signal.
13. The decoder as claimed in claim 12, wherein the at least one gain indicator comprises at least one deleted spectral coefficient index.
14. The decoder as claimed in claim 13, wherein the decoder is further configured to insert an at least one predetermined spectral coefficient value into the initial frequency representation determined by the at least one deleted spectral coefficient index.
15. The decoder as claimed in claim 14, wherein the at least one predetermined spectral coefficient value is a null value.
16. The decoder as claimed in claims 13 to 15, wherein the at least one gain indicator further comprises at least one spectral coefficient gain factor.
17. The decoder as claimed in claim 16, wherein the decoder is further configured to multiply a spectral coefficient of the initial frequency representation determined by the at least one deleted spectral coefficient index by at least one spectral coefficient gain factor.
18. A method for encoding an audio signal, comprising: generating at least one encoded signal dependent on the audio signal; determining at least one quality indicator, wherein each quality indicator is dependent on the relative difference between an associated one of the at least one encoded signal and the audio signal; selecting to output one of the at least one encoded signal dependent on the associated quality indicator value.
19. The method as claimed in claim 18, further comprising: generating at least one spectral representation of the audio signal, wherein the spectral representation comprises at least two sub-bands, each sub-band comprising at least one spectral coefficient.
20. The method as claimed in claim 19, further comprising generating a first of the at least one spectral representation of the audio signal by applying at least one of the following: a shifted discrete fourier transform; a modified discrete cosine transform; and a discrete unitary transform.
21 . The method as claimed in claim 20, further comprising generating further spectral representations of the audio signal by applying at least one gain factor to the at least one spectral coefficient of at least one of the at least two sub- bands.
22. The method as claimed in claim 21 , wherein the at least one gain factor comprises at least one of the following multiplication gain factors: 0 and 0.5.
23. The method as claimed in claim 20 to 22, further comprising ordering the at least two sub-bands.
24. The method as claimed in claim 23 when dependent on claim 21 , further comprising generating the further spectral representations of the audio signal by progressively applying, dependent on the order of the at least two sub-bands, at least one gain factor to the at least one spectral coefficient of at least one of the at least two sub-bands.
25. The method as claimed in claims 22 to 24, further comprising generating an indicator to indicate which sub-band spectral coefficients are modified by the at least one gain factor.
26. The method as claimed in claims 20 to 22, further comprising ordering the at least two sub-bands according to a psycho acoustical model.
27. The method as claimed in claim 24, further comprising progressively generating the further spectral representations of the audio signal until the quality indicator is equal to or greater than a quality threshold.
28. A method for decoding an encoded signal, comprising: partitioning the received encoded signal into at least a first and second part; decoding the first part to generate at least one gain indicator; generating at least one frequency representation of the encoded signal dependent on the second part and the at least one gain indicator, wherein the at least one frequency representation comprises at least one spectral coefficient.
29. The method as claimed in claim 28, further comprising generating at least one initial frequency representation of the encoded signal dependent on the second part of the received encoded signal.
30. The method as claimed in claim 29, wherein the at least one gain indicator comprises at least one deleted spectral coefficient index.
31. The method as claimed in claim 30, further comprising inserting an at least one predetermined spectral coefficient value into the initial frequency representation determined by the at least one deleted spectral coefficient index.
32. The method as claimed in claim 31 , wherein the at least one predetermined spectral coefficient value is a null value.
33. The method as claimed in claims 30 to 32, wherein the at least one gain indicator further comprises at least one spectral coefficient gain factor.
34. The method as claimed in claim 33, further comprising multiplying a spectral coefficient of the initial frequency representation determined by the at least one deleted spectral coefficient index by at least one spectral coefficient gain factor.
35. An apparatus comprising an encoder as claimed in claims 1 to 10.
36. An apparatus comprising a decoder as claimed in claims 11 to 17.
37. An electronic device comprising an encoder as claimed in claims 1 to 10.
38. An electronic device comprising a decoder as claimed in claims 1 1 to 17.
39. A chipset comprising an encoder as claimed in claims 1 to 10.
40. A chipset comprising a decoder as claimed in claims 11 to 17.
41. A computer program product configured to perform a method for encoding an audio signal, comprising: generating at least one encoded signal dependent on the audio signal; determining at least one quality indicator, wherein each quality indicator is dependent on the relative difference between an associated one of the at least one encoded signal and the audio signal; selecting to output one of the at least one encoded signal dependent on the associated quality indicator value.
42. A computer program product configured to perform a method for decoding an encoded signal, comprising: partitioning the received encoded signal into at ieast a first and second part; decoding the first part to generate at least one gain indicator; generating at least one frequency representation of the encoded signal dependent on the second part and the at least one gain indicator, wherein the at least one frequency representation comprises at least one spectral coefficient.
43. An encoder for encoding an audio signal; comprising: encoding means to generate at least one encoded signal dependent on the audio signal; processing means to determine at least one quality indicator, wherein each quality indicator is dependent on the relative difference between an associated one of the at least one encoded signal and the audio signal; and selection means to select to output one of the at least one encoded signal dependent on the associated quality indicator value
44. A decoder for decoding an encoded audio signal, comprising: processing means to partition the received encoded signal into at least a first and second part; decoding means to decode the first part to generate at least one gain indicator; further processing means to generate at least one frequency representation of the encoded signal dependent on the second part and the at least one gain indicator, wherein the at least one frequency representation comprises at least one spectral coefficient.
PCT/EP2007/062909 2007-11-27 2007-11-27 An encoder Ceased WO2009068083A1 (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
PCT/EP2007/062909 WO2009068083A1 (en) 2007-11-27 2007-11-27 An encoder

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
PCT/EP2007/062909 WO2009068083A1 (en) 2007-11-27 2007-11-27 An encoder

Publications (1)

Publication Number Publication Date
WO2009068083A1 true WO2009068083A1 (en) 2009-06-04

Family

ID=39577669

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/EP2007/062909 Ceased WO2009068083A1 (en) 2007-11-27 2007-11-27 An encoder

Country Status (1)

Country Link
WO (1) WO2009068083A1 (en)

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US9460729B2 (en) 2012-09-21 2016-10-04 Dolby Laboratories Licensing Corporation Layered approach to spatial audio coding

Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20030115041A1 (en) * 2001-12-14 2003-06-19 Microsoft Corporation Quality improvement techniques in an audio encoder
US20070168186A1 (en) * 2006-01-18 2007-07-19 Casio Computer Co., Ltd. Audio coding apparatus, audio decoding apparatus, audio coding method and audio decoding method

Patent Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20030115041A1 (en) * 2001-12-14 2003-06-19 Microsoft Corporation Quality improvement techniques in an audio encoder
US20070168186A1 (en) * 2006-01-18 2007-07-19 Casio Computer Co., Ltd. Audio coding apparatus, audio decoding apparatus, audio coding method and audio decoding method

Non-Patent Citations (1)

* Cited by examiner, † Cited by third party
Title
PIERRICK PHILIPPE ET AL: "Wavelet Packet Filterbanks for Low Time Delay Audio Coding", IEEE TRANSACTIONS ON SPEECH AND AUDIO PROCESSING, IEEE SERVICE CENTER, NEW YORK, NY, US, vol. 7, no. 3, 1 May 1999 (1999-05-01), XP011054374, ISSN: 1063-6676 *

Cited By (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US9460729B2 (en) 2012-09-21 2016-10-04 Dolby Laboratories Licensing Corporation Layered approach to spatial audio coding
US9495970B2 (en) 2012-09-21 2016-11-15 Dolby Laboratories Licensing Corporation Audio coding with gain profile extraction and transmission for speech enhancement at the decoder
US9502046B2 (en) 2012-09-21 2016-11-22 Dolby Laboratories Licensing Corporation Coding of a sound field signal
US9858936B2 (en) 2012-09-21 2018-01-02 Dolby Laboratories Licensing Corporation Methods and systems for selecting layers of encoded audio signals for teleconferencing

Similar Documents

Publication Publication Date Title
US7275036B2 (en) Apparatus and method for coding a time-discrete audio signal to obtain coded audio data and for decoding coded audio data
EP2186088B1 (en) Low-complexity spectral analysis/synthesis using selectable time resolution
KR101161866B1 (en) Audio coding apparatus and method thereof
CA2482427C (en) Apparatus and method for coding a time-discrete audio signal and apparatus and method for decoding coded audio data
EP2054882B1 (en) Arbitrary shaping of temporal noise envelope without side-information
EP2215627B1 (en) An encoder
CN102084418B (en) Apparatus and method for adjusting spatial cue information of a multichannel audio signal
EP1852851A1 (en) An enhanced audio encoding/decoding device and method
US8352249B2 (en) Encoding device, decoding device, and method thereof
EP2133872B1 (en) Encoding device and encoding method
US20100250260A1 (en) Encoder
US9230551B2 (en) Audio encoder or decoder apparatus
WO2005096273A1 (en) Enhanced audio encoding/decoding device and method
WO2005096508A1 (en) Enhanced audio encoding and decoding equipment, method thereof
EP2212883B1 (en) An encoder
US20100292986A1 (en) encoder
US20110191112A1 (en) Encoder
WO2009022193A2 (en) Devices, methods and computer program products for audio signal coding and decoding
WO2009068083A1 (en) An encoder
US20100280830A1 (en) Decoder
US8924202B2 (en) Audio signal coding system and method using speech signal rotation prior to lattice vector quantization
HK40037190B (en) Methods and apparatus for unified speech and audio decoding qmf based harmonic transposer improvements
WO2008114078A1 (en) En encoder

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 07847434

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 07847434

Country of ref document: EP

Kind code of ref document: A1