WO2014199632A1 - 音響信号の帯域幅拡張を行う装置及び方法 - Google Patents
音響信号の帯域幅拡張を行う装置及び方法 Download PDFInfo
- Publication number
- WO2014199632A1 WO2014199632A1 PCT/JP2014/003103 JP2014003103W WO2014199632A1 WO 2014199632 A1 WO2014199632 A1 WO 2014199632A1 JP 2014003103 W JP2014003103 W JP 2014003103W WO 2014199632 A1 WO2014199632 A1 WO 2014199632A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- frequency
- spectrum
- harmonic
- unit
- low frequency
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L21/00—Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
- G10L21/02—Speech enhancement, e.g. noise reduction or echo cancellation
- G10L21/038—Speech enhancement, e.g. noise reduction or echo cancellation using band spreading techniques
- G10L21/0388—Details of processing therefor
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L19/00—Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
- G10L19/02—Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis using spectral analysis, e.g. transform vocoders or subband vocoders
- G10L19/0204—Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis using spectral analysis, e.g. transform vocoders or subband vocoders using subband decomposition
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L19/00—Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
- G10L19/02—Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis using spectral analysis, e.g. transform vocoders or subband vocoders
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L19/00—Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
- G10L19/04—Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis using predictive techniques
- G10L19/16—Vocoder architecture
- G10L19/167—Audio streaming, i.e. formatting and decoding of an encoded audio signal representation into a data stream for transmission or storage purposes
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L19/00—Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
- G10L19/04—Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis using predictive techniques
- G10L19/16—Vocoder architecture
- G10L19/18—Vocoders using multiple modes
- G10L19/24—Variable rate codecs, e.g. for generating different qualities using a scalable representation such as hierarchical encoding or layered encoding
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L21/00—Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
- G10L21/02—Speech enhancement, e.g. noise reduction or echo cancellation
- G10L21/038—Speech enhancement, e.g. noise reduction or echo cancellation using band spreading techniques
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L19/00—Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
- G10L19/02—Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis using spectral analysis, e.g. transform vocoders or subband vocoders
- G10L19/032—Quantisation or dequantisation of spectral components
- G10L19/035—Scalar quantisation
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L25/00—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
- G10L25/03—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the type of extracted parameters
- G10L25/18—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the type of extracted parameters the extracted parameters being spectral information of each sub-band
Definitions
- the present invention relates to acoustic signal processing, and more particularly to encoding and decoding of acoustic signals for bandwidth extension of acoustic signals.
- BWE bandwidth extension
- WB wideband
- SWB super-wideband
- BWE in encoding uses a decoded low frequency band signal to represent the high frequency band signal parametrically. That is, the BWE searches for and identifies a portion similar to the sub-band of the high frequency band signal in the low frequency band signal of the acoustic signal, encodes and transmits a parameter specifying the similar portion, and the receiving side
- the low frequency band signal is used to enable the high frequency band signal to be recombined.
- the amount of information of parameters to be transmitted can be reduced by directly encoding high frequency band signals and utilizing similar portions of low frequency band signals, and compression efficiency can be improved.
- FIGS. 1 and 2 The configuration of 718-SWB is shown in FIGS. 1 and 2 (see, for example, Non-Patent Document 1).
- an acoustic signal (hereinafter referred to as an input signal) sampled at 32 kHz is first downsampled to 16 kHz (101).
- the downsampled signal is G. It is encoded (102) by the 718 core encoder.
- SWB bandwidth extension is performed in the MDCT domain.
- the 32 kHz input signal is transformed 103 into the MDCT domain and processed 104 through the tonality estimator.
- a generic mode (106) or a sinusoidal mode (108) is used for first layer coding of the SWB. Higher order SWB layers are encoded using additional sinusoids (107 and 109).
- the generic mode is used when the signal of the input frame is considered as non-tone.
- G The MDCT coefficients (spectrum) of the WB signal encoded by the 718 core encoding unit are used for encoding the SWB MDCT coefficients (spectrum).
- the SWB frequency band (7-14 kHz) is divided into several sub-bands, and for all sub-bands, the most correlated part is searched from the encoded and normalized WB MDCT coefficients. Then, the gain of the highest correlation part is scaled so as to reproduce the amplitude level of the SWB sub-band, and a parametric representation (parametric representation) of the high frequency component of the SWB signal is obtained.
- Sinusoidal mode coding is used in frames classified into tones.
- the SWB signal is generated by adding a finite set of sinusoidal components to the SWB spectrum.
- the 718 core codec decodes the WB signal at a 16 kHz sampling rate (201).
- the WB signal is post-processed (202) and then upsampled to a 32 kHz sampling rate (203).
- the SWB frequency components are reconstructed by SWB bandwidth extension.
- the SWB bandwidth extension is mainly performed in the MDCT domain.
- the generic mode (204) and the sinusoidal mode (205) are used for decoding of the first layer of the SWB. Higher order SWB layers are decoded using additional sinusoidal modes (206 and 207).
- the reconstructed SWB MDCT coefficients are transformed into the time domain (208) and after post processing (209), G.
- the signal is added to the WB signal decoded by the 718 core decoder to reconstruct the time domain SWB output signal.
- ITU-T Recommendation G. 718 Amendment 2 New Annex B on superwideband scalable extension for ITU-T G. 718 and corrections to main body fixed-point C-code and description text, March 2010.
- SWB bandwidth extension of the input signal is performed either in sinusoidal mode or in generic mode.
- high frequency components are generated (obtained) by searching the most correlated part from the WB spectrum.
- this type of approach suffers from performance, especially for signals with harmonics.
- This approach does not maintain any harmonics relationship between the low frequency band harmonic components (tone components) and the replicated high frequency band tone components. This leads to an unclear spectrum which degrades the aural quality.
- the spectrum of the low frequency band signal in order to suppress auditory noise (or artifact) generated by the disturbance in the unclear spectrum or the spectrum (high frequency spectrum) of the replicated high frequency band signal It is desirable to maintain the harmonics relationship between) and the high frequency spectrum.
- the 718-SWB configuration comprises a sine wave mode.
- Sinusoidal modes encode significant tonal components using sinusoidal waves, thus maintaining a good harmonics structure.
- simply encoding the SWB component with an artificial tone signal has a problem that the resulting voice quality is not necessarily sufficiently good.
- the present invention aims to improve the coding performance for signals having harmonics (harmonics) possessed by the above-mentioned generic mode, and to maintain the fine structure of the spectrum, and to reproduce the low frequency spectrum and the high frequency replicated. It provides an efficient way to maintain the harmonics structure of the tonal components between the spectra.
- the relationship between the tone component of the low frequency spectrum and the tone component of the high frequency spectrum can be obtained by estimating the value of the frequency of the harmonic from the WB spectrum.
- the low frequency spectrum encoded at the encoder side is decoded, and the portion with the highest correlation to the subbands of the high frequency spectrum is copied to the high frequency band after being energy level adjusted according to the index information
- the frequency spectrum is replicated.
- the frequency of the tonal component in the replicated high frequency spectrum is identified or adjusted based on the value of the estimated harmonic frequency.
- the harmonics relationship between the tone component of the low frequency spectrum and the tone component of the replicated high frequency spectrum is maintained only if the estimate of the frequency of the harmonics is correct. For this reason, in order to improve the estimation accuracy, correction of the spectrum peak constituting the tone component is performed before the frequency of the harmonic is estimated.
- tone components in a high frequency spectrum reconstructed by bandwidth extension are accurately replicated to efficiently obtain good speech quality at a low bit rate.
- the figure which shows the constitution of 718-SWB decoding device Block diagram showing configuration of coding apparatus according to Embodiment 1 of the present invention Block diagram showing the configuration of the decoding apparatus according to Embodiment 1 of the present invention Diagram showing correction approach for spectral peak detection Figure showing an example of harmonic frequency adjustment method Figure showing another example of harmonic frequency adjustment method
- Block diagram showing configuration of coding apparatus according to Embodiment 2 of the present invention Block diagram showing configuration of decoding apparatus according to Embodiment 2 of the present invention
- Block diagram showing configuration of coding apparatus according to Embodiment 3 of the present invention Block diagram showing configuration of decoding apparatus according to Embodiment 3 of the present invention
- Embodiment 1 The configuration of the codec according to the present invention is shown in FIG. 3 and FIG.
- the sampled input signal is first downsampled (301).
- the down-sampled low frequency band signal (low frequency signal) is encoded by the core encoding unit (302).
- the core coding parameters are sent to the multiplexer (307) to form a bitstream.
- the input signal is converted into a frequency domain signal by a time-frequency (T / F) converter (303), and the high frequency band signal (high frequency signal) is divided into a plurality of sub bands.
- the coding unit may be an existing narrow band or wide band audio or voice codec, for example G.264. 718 is mentioned.
- the core encoding unit (302) not only simply encodes, but also includes a local decoding unit and a time-frequency conversion unit, performs local decoding and performs time-frequency of the decoded signal (combined signal)
- the transformation is performed to provide the combined low frequency signal to the energy normalization unit (304).
- the synthesized low frequency signal in the normalized frequency domain is used for bandwidth extension as follows.
- the similarity search unit (305) specifies a portion having the highest correlation with each sub-band of the high frequency signal of the input signal in the normalized low frequency synthesis number signal, and the index information which is the search result It is sent to the multiplexing unit (307).
- scale factor information of this most correlated portion and each sub-band of the high frequency signal of the input signal is estimated (306), and the encoded scale factor information is sent to the multiplexing unit (307).
- the multiplexing unit (307) integrates core coding parameters, index information and scale factor information into a bitstream.
- the demultiplexer (401) decodes the bit stream to obtain core coding parameters, index information and scale factor information.
- the core decoding unit reconstructs the combined low frequency signal using the core coding parameters (402).
- the combined low frequency signal is upsampled (403) and used for bandwidth extension (410).
- This bandwidth extension is performed as follows. That is, the low frequency identified according to the index information which energy normalizes the combined low frequency signal (404) and identifies a portion having the highest correlation with each sub-band of the high frequency signal of the input signal derived at the encoder side
- the signal is copied to the high frequency band (405), and energy level adjustment is performed according to the scale factor information in order to make it the same level as the energy level of the high frequency signal of the input signal (406).
- the frequencies of the harmonics are estimated 407 from the spectrum of the combined low frequency signal.
- the estimated harmonic frequency is used to adjust the frequency of the tone component in the spectrum of the high frequency signal (408).
- the reconstructed high frequency signal is transformed 409 from the frequency domain to the time domain and added to the upsampled composite low frequency signal to produce a time domain output signal.
- spectral peaks and spectral peak frequencies are calculated. However, spectral peaks with small amplitudes and very short intervals between spectral peak frequencies with adjacent spectral peaks are eliminated. This avoids an estimation error when calculating the value of the harmonic frequency. 1) Calculate the interval of the specified spectral peak frequency. 2) Estimate the frequency of the harmonic based on the spacing of the identified spectral peak frequency. One of the methods of estimating the frequency of harmonics is shown below.
- the estimation of the frequency of the harmonic can also be performed by the following method. 1) In the spectrum of the synthesized low frequency signal (LF), in order to estimate the frequency of the harmonics, a portion having a clear harmonics structure is selected so as to secure the reliability of the frequency of the estimated harmonics. A sharp harmonics structure is usually found in the vicinity of the cutoff frequency from 1-2 kHz for all harmonics. 2) Identify the spectrum having the largest amplitude (absolute value) and its frequency in the selected portion of the above-mentioned synthesized low frequency signal (spectrum). 3) From the spectral frequencies of this maximum amplitude spectrum, identify a set of spectral peaks that have approximately equal frequency spacing and whose absolute magnitude exceeds a predetermined threshold.
- LF synthesized low frequency signal
- the predetermined threshold value for example, a value twice the standard deviation of the spectrum amplitude of the selected part described above can be employed. 4) Calculate the interval of the above-mentioned spectrum peak frequency. 5) Estimate the frequency of the harmonic based on the interval of the above-mentioned spectral peak frequency. Also in this case, the method of equation (1) can be used to estimate the frequency of the harmonics.
- harmonic components in the spectrum of the synthesized low frequency signal may not be sufficiently encoded.
- some of the identified spectral peaks may not correspond at all to the harmonic content of the input signal.
- the interval between the spectral peak frequencies is significantly different from the average value, it is better to exclude from this calculation target.
- the spacing of spectral peak frequencies extracted in the missing harmonic portion is considered to be twice or several times the spacing of spectral peak frequencies extracted in a portion having a good harmonics structure.
- the average value of the extracted values of the intervals of the spectral peak frequency included in the predetermined range including the interval of the maximum spectral peak frequency is used as the estimated value of the frequency of the harmonic. This allows the high frequency spectrum to be properly replicated. Specifically, it consists of the following steps. 1) Identify the minimum and maximum values of the spectral peak frequency interval.
- the one with the smallest spectral peak frequency in the replicated high frequency spectrum is shifted from the largest spectral peak frequency of the synthesized low frequency signal spectrum to a frequency with a distance of Est Harmonic .
- the second smallest spectral peak frequency in the replicated high frequency spectrum shifts from the above shifted minimum spectral peak frequency to a frequency having an Est Harmonic spacing. This process is repeated until such adjustment is complete for the spectral peak frequencies of all spectral peaks in the replicated high frequency spectrum.
- the following harmonic frequency adjustment method is also possible. 1) Identify the one with the highest spectral peak frequency of the spectrum of the synthesized low frequency signal (LF). 2) Identify spectral peaks and spectral peak frequencies within the high frequency (HF) spectrum that is bandwidth expanded by bandwidth expansion. 3) Calculate the spectral peak frequency which can be taken in the HF spectrum with reference to the maximum spectral peak frequency of the synthesized low frequency signal spectrum. Each spectrum peak in the high frequency spectrum replicated by the bandwidth extension is moved to a frequency closest to each spectrum peak frequency among the calculated spectrum peak frequencies. This process is shown in FIG. As shown in FIG. 7, first, those with the largest spectral peak frequency of the synthesized low frequency spectrum and spectral peaks in the replicated high frequency spectrum are extracted.
- spectral peak frequencies that can be taken within the replicated high frequency spectrum are calculated.
- the frequency having a distance from the largest spectral peak frequency of the synthesized low frequency signal spectrum to the Est Harmonic is taken as the frequency of the spectral peak that can be taken first in the spectral peak in the replicated high frequency spectrum.
- the frequency having an interval of Est Harmonic from the first possible spectral peak frequency is taken as the frequency of the second possible spectral peak. Repeat this process as much as you can calculate in the high frequency spectrum.
- the spectral peak extracted in the replicated high frequency spectrum is shifted to the closest frequency among the possible spectral peak frequencies calculated above.
- the estimated harmonic value Est Harmonic may not correspond to an integer number of frequency bins.
- the spectral peak frequency is selected to be the frequency bin closest to the frequency derived based on Est Harmonic .
- the harmonic frequency estimation method in which the spectrum of the previous frame is used to estimate the harmonic frequency and the spectrum of the previous frame is considered so that frame transition becomes smooth when adjusting the tone component. It is also conceivable to adjust the frequency of the tone component as described above. Also, the amplitude may be adjusted so that the energy level of the original spectrum is maintained even if the frequency of the tone component is shifted. All these minor modifications are included within the scope of the present invention.
- the bandwidth extension method according to the present invention is to duplicate the high frequency spectrum using the high frequency spectrum and the composite low frequency signal spectrum having the highest correlation, and to shift the spectrum peak to the estimated harmonic frequency. . This makes it possible to maintain both the fine structure of the spectrum and the harmonics structure between the spectral peak of the low frequency band and the spectral peak of the replicated high frequency band.
- FIG. 8 Second Embodiment Embodiment 2 of the present invention is shown in FIG. 8 and FIG.
- the coding apparatus according to the second embodiment is substantially the same as the first embodiment except for the harmonic frequency estimation unit (708, 709) and the harmonic frequency comparison unit (710).
- flag information is transmitted based on the comparison result (710) of the estimated values of the two.
- flag information can be derived as in the following equation.
- the frequency of the harmonics estimated from the synthesized low frequency spectrum may be different from the frequency of the harmonics of the high frequency spectrum of the input signal.
- the harmonic structure of the low frequency spectrum is not well maintained.
- FIG. 10 Third Embodiment Embodiment 3 of the present invention is shown in FIG. 10 and FIG.
- the coding apparatus according to the third embodiment is substantially the same as the second embodiment except for the difference unit (910).
- the frequencies of the harmonics are estimated separately in the combined low frequency spectrum (908) and the high frequency spectrum (909) of the input signal.
- the difference (Diff) of the frequencies of the two estimated harmonics is calculated (910) and transmitted to the decoding device side.
- the difference value (Diff) is added to the estimated value of the frequency of the harmonic from the combined low frequency spectrum (1010), and the value of the frequency of the newly calculated harmonic is replicated Used for harmonic frequency adjustment in the high frequency spectrum.
- the frequency of the harmonics estimated from the high frequency spectrum of the input signal may be sent directly to the decoding unit. And harmonic frequency adjustment is performed using the received value of the frequency of the harmonic of the high frequency spectrum of an input signal. This makes it unnecessary to estimate the frequency of harmonics from the synthesized low frequency spectrum at the decoder side.
- the frequency of the harmonics estimated from the synthesized low frequency spectrum may differ from the frequency of the harmonics of the high frequency spectrum of the input signal, so the difference value or the high frequency spectrum of the input signal
- Embodiment 4 The fourth embodiment of the present invention is shown in FIG.
- the coding apparatus according to the fourth embodiment is the same as another conventional coding apparatus or the first, second or third embodiment.
- the frequencies of the harmonics are estimated from the combined low frequency spectrum (1103). An estimate of the frequency of this harmonic is used for harmonic injection (1104) in the low frequency spectrum.
- some low frequency spectrum harmonic components may be barely coded or not coded at all.
- an estimate of the frequency of the harmonic can be used to inject the missing harmonic component.
- the frequency can be derived using an estimate of the frequency of the harmonics.
- the amplitude may be, for example, the average value of the amplitudes of other existing spectral peaks or the average value of the amplitudes of existing spectral peaks close to the missing harmonic component on the frequency axis.
- the harmonic components generated according to this frequency and amplitude are injected to restore the missing harmonic components.
- the frequency of the harmonics is estimated using the coded LF spectrum (1103).
- 1.1 Estimate the frequency of the harmonics using the spacing of spectral peak frequencies identified in the coded low frequency spectrum.
- the value of the spacing of spectral peak frequencies derived in the missing harmonic part will be twice or several times the value of the spacing of spectral peak frequencies derived in the part maintaining good harmonics structure.
- the spacing of such spectral peak frequencies is grouped into different categories, and for each, the average spectral peak frequency spacing is estimated. The details will be described below.
- a. Identify minimum and maximum values of spectral peak frequency interval values.
- b. Identify all interval values in the following range: c.
- the average value of the values of the intervals specified in the above range is calculated as the estimated value of the frequency of the harmonic. 2.
- An estimate of the frequency of the harmonics is used to inject the missing harmonic components.
- 2.1 Split the selected LF spectrum into several regions.
- 2.2 Identify missing harmonics by using region information and estimated frequencies. For example, it is assumed that the selected LF spectrum is divided into three regions r 1 , r 2 and r 3 . Based on the region information, the harmonics are identified and the harmonics are injected. The signal characteristics for the harmonic spectrum gap between the harmonics becomes Est HarmonicLF2 in the area of Est HarmonicLF1 next, r 3 in the region of the r 1 and r 2. This information can be used to extend the LF spectrum. This is further illustrated in FIG. In FIG. In FIG.
- the synthesized low frequency spectrum may not be maintained.
- Some harmonic components may be missing, especially at low bit rates.
- By injecting the missing harmonic component in the LF spectrum not only the extension of the LF but also the harmonics characteristics of the reconstructed harmonic can be improved. As a result, it is possible to further improve the voice quality by suppressing the auditory influence due to the omission of the harmonics.
- the encoding device, the decoding device and the encoding / decoding method according to the present invention can be applied to a wireless communication terminal device, a base station device in a mobile communication system, a teleconference terminal device, a video conference terminal device, and a VOIP terminal device is there.
Landscapes
- Engineering & Computer Science (AREA)
- Physics & Mathematics (AREA)
- Computational Linguistics (AREA)
- Signal Processing (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Human Computer Interaction (AREA)
- Acoustics & Sound (AREA)
- Multimedia (AREA)
- Quality & Reliability (AREA)
- Spectroscopy & Molecular Physics (AREA)
- Compression, Expansion, Code Conversion, And Decoders (AREA)
Abstract
Description
本発明に係るコーデックの構成を図3及び図4に示す。
1)合成低周波数信号(LF)のスペクトルから、高調波の周波数を推定するための部分を選択。選択された部分は、選択された部分から推定される高調波の周波数が信頼できるものであるために、鮮明なハーモニクス構造を有するべきである。通常、全ての高調波に対して、1-2kHzからカットオフ周波数の付近において鮮明なハーモニクス構造が観察される。
2)選択された部分を人間のピッチ周波数に近い幅(100Hz~400Hz程度)の多数のブロックに分割。
3)各ブロック内において振幅が最大となるスペクトル(スペクトルピーク)、及びスペクトルピークの周波数(スペクトルピーク周波数)を探索。
4)エラー回避又は高調波の周波数の推定精度向上のために、特定したスペクトルピークに対して後処理を実施。
1)特定されたスペクトルピーク周波数の間隔を算出。
2)特定されたスペクトルピーク周波数の間隔に基づいて高調波の周波数を推定。高調波の周波数を推定する方法の一つを以下に示す。
1)合成低周波数信号(LF)のスペクトルにおいて、高調波の周波数を推定するため、推定される高調波の周波数の信頼性が担保できるよう鮮明なハーモニクス構造を有する部分を選ぶ。通常、全ての高調波に対して、1-2kHzからカットオフ周波数の付近において鮮明なハーモニクス構造が見られる。
2)上記の合成低周波数信号(スペクトル)の選択された部分の中で最大の振幅(絶対値)を有するスペクトルとその周波数を特定する。
3)この最大振幅のスペクトルのスペクトル周波数から、ほぼ等しい周波数間隔を有し、かつ振幅の絶対値が所定の閾値を越えるスペクトルピークのセットを特定する。所定の閾値としては例えば前述の選択された部分のスペクトル振幅の標準偏差の2倍の値が採用できる。
4)上記スペクトルピーク周波数の間隔を算出する。
5)上記スペクトルピーク周波数の間隔に基づいて高調波の周波数を推定する。なお、この場合にも高調波の周波数を推定するため、式(1)の方法を使用可能である。
1)スペクトルピーク周波数の間隔の最小値及び最大値を特定する。
2)帯域幅拡張により複製された高周波数スペクトル内のスペクトルピーク及びスペクトルピーク周波数を特定する。
3)合成低周波数信号スペクトルのスペクトルピークのうち、最大のスペクトルピーク周波数を基準として、スペクトルピーク周波数の間隔が高調波の周波数間隔の推定値と等しくなるように、スペクトルピーク周波数を調整する。この処理を図6に示す。図6に示すように、まず、合成低周波数信号スペクトル中で最大のスペクトルピーク周波数、及び、複製された高周波数スペクトル内のスペクトルピークを特定する。そして、複製された高周波数スペクトル内の最小のスペクトルピーク周波数を持つものを、合成低周波数信号スペクトルの最大のスペクトルピーク周波数からEstHarmonicの間隔を有する周波数にシフトする。複製された高周波数スペクトル内のスペクトルピーク周波数が2番目に小さなものは、上記のシフトされた最小のスペクトルピーク周波数からEstHarmonicの間隔を有する周波数にシフトする。複製された高周波数スペクトル内の全てのスペクトルピークのスペクトルピーク周波数についてこのような調整が完了するまでこの処理を繰り返す。
1)合成低周波数信号(LF)のスペクトルの最大のスペクトルピーク周波数を持つものを特定する。
2)帯域幅拡張により帯域幅拡張される高周波数(HF)スペクトル内のスペクトルピーク及びスペクトルピーク周波数を特定する。
3)合成低周波数信号スペクトルの最大のスペクトルピーク周波数を基準として、HFスペクトルにおいて採りうるスペクトルピーク周波数を算出する。帯域幅拡張により複製された高周波数スペクトル内の各スペクトルピークを算出されたスペクトルピーク周波数のうち各スペクトルピーク周波数に最も近い周波数へ移動する。この処理を図7に示す。図7に示すように、まず、合成低周波数スペクトルの最大のスペクトルピーク周波数を持つもの、及び、複製された高周波数スペクトル内のスペクトルピークが抽出される。そして、複製された高周波数スペクトル内で採りうるスペクトルピーク周波数が算出される。合成低周波数信号スペクトルの最大のスペクトルピーク周波数からEstHarmonicの間隔を有する周波数を、複製された高周波数スペクトル内のスペクトルピークが1番目に採りうるスペクトルピークの周波数とする。次に上記1番目の採りうるスペクトルピーク周波数からEstHarmonicの間隔を有する周波数を、2番目に採りうるスペクトルピークの周波数とする。高周波数スペクトル内で計算できる限りこの処理を繰り返す。
本発明に係る帯域幅拡張方法は、高周波数スペクトルと最も相関の高い合成低周波数信号スペクトルを用いて高周波数スペクトルを複製するとともに、スペクトルピークを推定された高調波の周波数へシフトするものである。これにより、スペクトルの微細構造、及び、低周波帯域のスペクトルピークと複製された高周波帯域のスペクトルピークとの間のハーモニクス構造の双方を維持することができる。
本発明の実施の形態2は、図8及び図9に示される。
いくつかの信号に対して、合成低周波数スペクトルから推定した高調波の周波数は、入力信号の高周波数スペクトルの高調波の周波数と異なる場合がある。特に低ビットレートでは、低周波数スペクトルのハーモニクス構造は良好に維持されない。フラグ情報を送ることによって、誤った高調波の周波数の推定値を用いたトーン成分の調整を回避することができる。
本発明の実施の形態3は、図10及び図11に示される。
いくつかの信号に対して、合成低周波数スペクトルから推定した高調波の周波数は、入力信号の高周波数スペクトルの高調波の周波数と異なる場合があるため、差分値、又は、入力信号の高周波数スペクトルから導出された高調波の周波数の値を送ることによって、受信側である復号装置で帯域幅拡張して複製した高周波数スペクトルのトーン成分の調整をより精度良く行うことができる。
本発明の実施の形態4は、図12に示される。
1.符号化されたLFスペクトルを用いて高調波の周波数を推定する(1103)。
1.1 高調波の周波数を、符号化された低周波数スペクトル内で特定されたスペクトルピーク周波数の間隔を用いて推定する。
1.2 欠落した高調波部分で導出されたスペクトルピーク周波数の間隔の値は良好なハーモニクス構造を維持している部分で導出されるスペクトルピーク周波数の間隔の値の2倍又は数倍となる。このようなスペクトルピーク周波数の間隔は、異なるカテゴリにグループ化され、それぞれに対して平均的なスペクトルピーク周波数の間隔が推定される。以下にその詳細を説明する。
a.スペクトルピーク周波数の間隔の値の最小値及び最大値を特定する。
2.1 選択されたLFスペクトルをいくつかの領域に分割する。
2.2 領域情報及び推定された周波数を用いることにより欠落した高調波を特定する。
例えば、選択されたLFスペクトルが3つの領域r1,r2,r3に分割されたとする。
領域情報に基づいて、高調波が特定され、高調波が注入される。
高調波に対する信号特性により、高調波間のスペクトルギャップは、r1及びr2の領域ではEstHarmonicLF1となり、r3の領域ではEstHarmonicLF2となる。この情報は、LFスペクトルの拡張に使用することができる。このことを更に図14に示す。図14では、LFスペクトルの領域r2に欠落した高調波成分があることが分かる。この周波数は、高調波の周波数の推定値EstHarmonicLF1を用いて導出可能である。
同様に、EstHarmonicLF2は、領域r2での欠落した高調波のトラッキング及び注入に使用される。
また、その振幅は、欠落していない全高調波成分の振幅の平均値、または欠落した高調波成分の前後に連なる高調波成分の振幅の平均値を用いることができる。又は、振幅はWBスペクトルで最小振幅を有するスペクトルピークを用いてもよい。その周波数及び振幅を用いて生成された高調波成分が欠落した高調波成分を復元するものとしてLFスペクトルに注入される。
いくつかの信号に対して、合成低周波数スペクトルは維持されない場合がある。特に低ビットレートでは、いくつかの高調波成分は欠落する可能性がある。LFスペクトルで欠落した高調波成分を注入することにより、LFの拡張のみでなく、再構成される高調波のハーモニクス特性を向上させることができる。これにより、高調波の欠落による聴感的な影響を抑圧して、音声品質を更に向上させることができる。
Claims (10)
- 音響信号を符号化する符号化装置から送信された符号化情報からコア符号化パラメータ、インデックス情報、およびスケールファクタ情報を取り出す逆多重化部と、
前記コア符号化パラメータを復号して、合成低周波数スペクトルを得るコア復号部と、
前記インデックス情報に基づき、前記合成低周波数スペクトルを用いて高周波数サブバンドスペクトルを複製するスペクトル複製部と、
前記スケールファクタ情報を用いて、前記複製された高周波数サブバンドスペクトルの振幅を調整するスペクトル包絡調整部と、を具備し、
前記合成低周波数スペクトルと前記高周波数サブバンドスペクトルとを用いて出力信号を生成する音響信号復号装置であって、
前記複製された高周波数サブバンドスペクトルにおける高調波成分の周波数を推定する高調波周波数推定部と、
前記合成低周波数スペクトルを用いて推定される高調波周波数を用いて高周波数スペクトルにおける高調波成分の周波数を調整する高調波周波数調整部と、
をさらに具備することを特徴とする音響信号復号装置。 - 前記高調波周波数推定部は、
前記合成低周波数スペクトルの中で予め選択された部分を所定数のブロックに分割する分割部と、
各ブロックにおいて、最大の振幅を有するスペクトル(スペクトルピーク)と、前記スペクトルピークの周波数を求めるスペクトルピーク特定部と、
前記特定されたスペクトルピークの周波数の間隔を算出する間隔算出部と、
前記特定されたスペクトルピークの周波数の間隔を用いて、前記高調波周波数を算出する高調波周波数算出部と、を具備する、
請求項1に記載の音響信号復号装置。 - 前記高調波周波数推定部は、
前記合成低周波数スペクトルの予め選択された部分で振幅の絶対値が最大となるスペクトルと当該スペクトルから周波数軸上でほぼ等間隔に位置し、かつ振幅の絶対値が所定の閾値以上のスペクトルを特定するスペクトルピーク特定部と、
前記特定されたスペクトルピークの周波数の間隔を算出する間隔算出部と、
前記特定されたスペクトルの周波数の間隔を用いて、前記高調波周波数を算出する高調波周波数算出部と、を具備する、
請求項1に記載の音響信号復号装置。 - 前記高調波周波数調整部は、
前記合成低周波数スペクトルにおけるスペクトルピークのうち最大周波数のものの周波数を特定する低周波数スペクトルピーク特定部と、
前記複製された高周波数サブバンドスペクトルにおける複数のスペクトルピークの周波数を特定する高周波数スペクトルピーク特定部と、
前記合成低周波数スペクトルにおけるスペクトルピークのうち最大周波数のものの周波数を基準として、前記複数のスペクトルピークの周波数の間隔が前記推定された高調波の周波数と等しくなるように、前記複数のスペクトルピークの周波数を調整する調整部と、を具備する、
請求項2に記載の音響信号復号装置。 - 前記高調波周波数調整部は、
前記合成低周波数スペクトルにおけるスペクトルピークのうち最大周波数のものの周波数を特定する低周波数スペクトルピーク特定部と、
前記複製された高周波数サブバンドスペクトルにおける複数のスペクトルピークの周波数を特定する高周波数スペクトルピーク特定部と、
前記合成低周波数スペクトルにおけるスペクトルピークのうち最大周波数のものの周波数に前記推定された高調波の周波数の整数倍の周波数を加算した周波数を、採りうるスペクトルピーク周波数として算出するスペクトルピーク周波数算出部と、
前記複製された高周波数サブバンドスペクトル内の前記複数のスペクトルピークの周波数を、前記算出された採りうるスペクトルピーク周波数のうち最も近い周波数へ調整する調整部と、を具備する、
請求項2に記載の音響信号復号装置。 - 音響信号を符号化する符号化装置から多重化して送信されたコア符号化パラメータと、インデックス情報とスケールファクタ情報とフラグ情報を逆多重化する逆多重化部と、
前記コア符号化パラメータを時間領域の低周波数信号に復号するとともに、前記復号された低周波数信号を周波数領域に変換して合成低周波数スペクトルを得るコア復号部と、
前記合成低周波数スペクトルから、前記インデックス情報に基づいて高周波数サブバンドスペクトルを再構成するスペクトル複製部と、
前記スケールファクタ情報を用いて、前記複製された高周波数サブバンドスペクトルの振幅を調整するスペクトル包絡調整部と、
前記合成低周波数スペクトルから高調波の周波数を推定する高調波周波数推定部と、
前記推定された高調波の周波数に基づいて、前記合成低周波数スペクトルから前記複製された高周波数サブバンドスペクトルにおけるトーン成分の周波数を調整する高調波周波数調整部と、
前記フラグ情報に基づいて、前記高調波周波数調整部を動作させるか否かを決定する決定部と、を具備し、
前記合成低周波数スペクトルと、前記高周波数サブバンドスペクトルを用いて出力信号を生成する、
音響信号復号装置。 - 前記推定された高調波の周波数に基づいて、前記合成低周波数スペクトルで欠落した高調波成分を特定する欠落高調波成分特定部と、
前記合成低周波数スペクトルに前記欠落した高調波成分を注入する高調波注入部と、を更に具備する、
請求項1又は6に記載の音響信号復号装置。 - 前記高調波注入部は、
欠落していない全高調波成分の振幅の平均値または周波数軸上で欠落した高調波成分の前後に位置する高調波成分の振幅の平均値を振幅とする高調波成分を生成する、
請求項7に記載の音響信号復号装置。 - 入力音響信号(以下、入力信号)を低サンプリングレートにダウンサンプリングするダウンサンプリング部と、
前記ダウンサンプリングされた信号をコア符号化パラメータへ符号化し、前記コア符号化パラメータを出力するとともに、前記コア符号化パラメータをローカルに復号し、周波数領域に変換して合成低周波数スペクトルを得るコア符号化部と、
前記合成低周波数スペクトルを正規化するエネルギ正規化部と、
前記入力信号をスペクトルに変換するとともに、前記合成低周波数スペクトルより高い周波数のスペクトルを複数のサブバンド(以下、高周波数サブバンド)に分割する時間-周波数変換部と、
前記各高周波数サブバンドに対して、前記正規化された合成低周波数スペクトルから最も相関の高い部分を特定し、特定結果をインデックス情報として出力する類似度探索部と、
前記各高周波数サブバンドと、前記合成低周波数スペクトルから特定された前記最も相関の高い部分との間のエネルギのスケールファクタを推定し、前記スケールファクタを、スケールファクタ情報として出力するスケールファクタ推定部と、
前記合成低周波数スペクトルの高調波の周波数と、前記変換された入力信号の高調波の周波数を推定する高調波周波数推定部と、
前記2つの高調波の周波数を比較して、高調波周波数調整をすべきか否かを判断し、前記判断結果をフラグ情報として出力する高調波周波数比較部と、
を具備する音響信号符号化装置。 - 入力音響信号(以下、入力信号)を低サンプリングレートにダウンサンプリングするダウンサンプリング部と、
前記ダウンサンプリングされた信号をコア符号化パラメータへ符号化し出力するとともに、前記コア符号化パラメータをローカルに復号し、周波数領域に変換して合成低周波数スペクトルを得るコア符号化部と、
前記入力信号をスペクトルに変換するとともに、前記合成低周波数スペクトルより高い周波数のスペクトルを複数のサブバンド(以下、高周波数サブバンド)に分割する時間-周波数変換部と、
前記各高周波数サブバンドに対して、前記低周波数スペクトルから最も相関の高い部分を特定し、特定結果をインデックス情報として出力する類似度探索部と、
前記各高周波数サブバンドと、前記合成低周波数スペクトルから特定された前記最も相関の高い部分との間のエネルギのスケールファクタを推定し、前記スケールファクタをスケールファクタ情報として出力するスケールファクタ推定部と、
前記合成低周波数スペクトルの高調波の周波数と、前記変換された入力信号の高調波の周波数を推定し、出力する高調波周波数推定部と、
を具備する音響信号符号化装置。
Priority Applications (15)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| RU2015151169A RU2658892C2 (ru) | 2013-06-11 | 2014-06-10 | Устройство и способ для расширения диапазона частот для акустических сигналов |
| CN201480031440.1A CN105408957B (zh) | 2013-06-11 | 2014-06-10 | 进行语音信号的频带扩展的装置及方法 |
| US14/894,062 US9489959B2 (en) | 2013-06-11 | 2014-06-10 | Device and method for bandwidth extension for audio signals |
| JP2015522543A JP6407150B2 (ja) | 2013-06-11 | 2014-06-10 | 音響信号の帯域幅拡張を行う装置及び方法 |
| ES14811296T ES2836194T3 (es) | 2013-06-11 | 2014-06-10 | Dispositivo y procedimiento para la extensión de ancho de banda para señales acústicas |
| EP20178265.3A EP3731226A1 (en) | 2013-06-11 | 2014-06-10 | Device and method for bandwidth extension for acoustic signals |
| BR122020016403-4A BR122020016403B1 (pt) | 2013-06-11 | 2014-06-10 | Aparelho de decodificação de sinal de áudio, aparelho de codificação de sinal de áudio, método de decodificação de sinal de áudio e método de codificação de sinal de áudio |
| EP14811296.4A EP3010018B1 (en) | 2013-06-11 | 2014-06-10 | Device and method for bandwidth extension for acoustic signals |
| MX2015016109A MX353240B (es) | 2013-06-11 | 2014-06-10 | Dispositivo y método para extensión de ancho de banda para señales acústicas. |
| KR1020157033759A KR102158896B1 (ko) | 2013-06-11 | 2014-06-10 | 음향 신호의 대역폭 확장을 행하는 장치 및 방법 |
| CN202010063428.6A CN111477245B (zh) | 2013-06-11 | 2014-06-10 | 语音信号解码装置和方法、语音信号编码装置和方法 |
| BR112015029574-6A BR112015029574B1 (pt) | 2013-06-11 | 2014-06-10 | Aparelho e método de decodificação de sinal de áudio. |
| US15/286,030 US9747908B2 (en) | 2013-06-11 | 2016-10-05 | Device and method for bandwidth extension for audio signals |
| US15/659,023 US10157622B2 (en) | 2013-06-11 | 2017-07-25 | Device and method for bandwidth extension for audio signals |
| US16/219,656 US10522161B2 (en) | 2013-06-11 | 2018-12-13 | Device and method for bandwidth extension for audio signals |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP2013-122985 | 2013-06-11 | ||
| JP2013122985 | 2013-06-11 |
Related Child Applications (2)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| US14/894,062 A-371-Of-International US9489959B2 (en) | 2013-06-11 | 2014-06-10 | Device and method for bandwidth extension for audio signals |
| US15/286,030 Continuation US9747908B2 (en) | 2013-06-11 | 2016-10-05 | Device and method for bandwidth extension for audio signals |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2014199632A1 true WO2014199632A1 (ja) | 2014-12-18 |
Family
ID=52021944
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/JP2014/003103 Ceased WO2014199632A1 (ja) | 2013-06-11 | 2014-06-10 | 音響信号の帯域幅拡張を行う装置及び方法 |
Country Status (11)
| Country | Link |
|---|---|
| US (4) | US9489959B2 (ja) |
| EP (2) | EP3731226A1 (ja) |
| JP (4) | JP6407150B2 (ja) |
| KR (1) | KR102158896B1 (ja) |
| CN (2) | CN105408957B (ja) |
| BR (2) | BR112015029574B1 (ja) |
| ES (1) | ES2836194T3 (ja) |
| MX (1) | MX353240B (ja) |
| PT (1) | PT3010018T (ja) |
| RU (2) | RU2688247C2 (ja) |
| WO (1) | WO2014199632A1 (ja) |
Cited By (7)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN105280189A (zh) * | 2015-09-16 | 2016-01-27 | 深圳广晟信源技术有限公司 | 带宽扩展编码和解码中高频生成的方法和装置 |
| JP2017523473A (ja) * | 2014-07-28 | 2017-08-17 | フラウンホーファー−ゲゼルシャフト・ツール・フェルデルング・デル・アンゲヴァンテン・フォルシュング・アインゲトラーゲネル・フェライン | 全帯域ギャップ充填を備えた周波数ドメインプロセッサと時間ドメインプロセッサとを使用するオーディオ符号器及び復号器 |
| CN108630212A (zh) * | 2018-04-03 | 2018-10-09 | 湖南商学院 | 非盲带宽扩展中高频激励信号的感知重建方法与装置 |
| CN108701467A (zh) * | 2015-12-14 | 2018-10-23 | 弗劳恩霍夫应用研究促进协会 | 处理经编码音频信号的装置及方法 |
| EP3435376A1 (en) | 2017-07-28 | 2019-01-30 | Fujitsu Limited | Audio encoding apparatus and audio encoding method |
| US10236007B2 (en) | 2014-07-28 | 2019-03-19 | Fraunhofer-Gesellschaft Zur Foerderung Der Angewandten Forschung E.V. | Audio encoder and decoder using a frequency domain processor , a time domain processor, and a cross processing for continuous initialization |
| US11367455B2 (en) | 2015-03-13 | 2022-06-21 | Dolby International Ab | Decoding audio bitstreams with enhanced spectral band replication metadata in at least one fill element |
Families Citing this family (21)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN103516440B (zh) * | 2012-06-29 | 2015-07-08 | 华为技术有限公司 | 语音频信号处理方法和编码装置 |
| CN103971693B (zh) * | 2013-01-29 | 2017-02-22 | 华为技术有限公司 | 高频带信号的预测方法、编/解码设备 |
| ES2836194T3 (es) * | 2013-06-11 | 2021-06-24 | Fraunhofer Ges Forschung | Dispositivo y procedimiento para la extensión de ancho de banda para señales acústicas |
| EP3128513B1 (en) | 2014-03-31 | 2019-05-15 | Fraunhofer Gesellschaft zur Förderung der Angewand | Encoder, decoder, encoding method, decoding method, and program |
| US9697843B2 (en) * | 2014-04-30 | 2017-07-04 | Qualcomm Incorporated | High band excitation signal generation |
| US10346126B2 (en) | 2016-09-19 | 2019-07-09 | Qualcomm Incorporated | User preference selection for audio encoding |
| KR102721794B1 (ko) * | 2016-11-18 | 2024-10-25 | 삼성전자주식회사 | 신호 처리 프로세서 및 신호 처리 프로세서의 제어 방법 |
| JP6769299B2 (ja) * | 2016-12-27 | 2020-10-14 | 富士通株式会社 | オーディオ符号化装置およびオーディオ符号化方法 |
| EP3396670B1 (en) * | 2017-04-28 | 2020-11-25 | Nxp B.V. | Speech signal processing |
| JP7214726B2 (ja) | 2017-10-27 | 2023-01-30 | フラウンホッファー-ゲゼルシャフト ツァ フェルダールング デァ アンゲヴァンテン フォアシュンク エー.ファオ | ニューラルネットワークプロセッサを用いた帯域幅が拡張されたオーディオ信号を生成するための装置、方法またはコンピュータプログラム |
| CN110660409A (zh) * | 2018-06-29 | 2020-01-07 | 华为技术有限公司 | 一种扩频的方法及装置 |
| US11100941B2 (en) * | 2018-08-21 | 2021-08-24 | Krisp Technologies, Inc. | Speech enhancement and noise suppression systems and methods |
| CN109243485B (zh) * | 2018-09-13 | 2021-08-13 | 广州酷狗计算机科技有限公司 | 恢复高频信号的方法和装置 |
| JP6693551B1 (ja) * | 2018-11-30 | 2020-05-13 | 株式会社ソシオネクスト | 信号処理装置および信号処理方法 |
| CN113192517B (zh) * | 2020-01-13 | 2024-04-26 | 华为技术有限公司 | 一种音频编解码方法和音频编解码设备 |
| CN113808596B (zh) * | 2020-05-30 | 2025-01-03 | 华为技术有限公司 | 一种音频编码方法和音频编码装置 |
| CN113963703B (zh) * | 2020-07-03 | 2025-05-02 | 华为技术有限公司 | 一种音频编码的方法和编解码设备 |
| CN113948094B (zh) * | 2020-07-16 | 2026-01-02 | 华为技术有限公司 | 音频编解码方法和相关装置及计算机可读存储介质 |
| CN113362837B (zh) * | 2021-07-28 | 2024-05-14 | 腾讯音乐娱乐科技(深圳)有限公司 | 一种音频信号处理方法、设备及存储介质 |
| CN114550732B (zh) * | 2022-04-15 | 2022-07-08 | 腾讯科技(深圳)有限公司 | 一种高频音频信号的编解码方法和相关装置 |
| CN116524951A (zh) * | 2023-03-30 | 2023-08-01 | 鼎道智芯(上海)半导体有限公司 | 音频处理方法和装置 |
Citations (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2003108197A (ja) * | 2001-07-13 | 2003-04-11 | Matsushita Electric Ind Co Ltd | オーディオ信号復号化装置およびオーディオ信号符号化装置 |
| JP2011100159A (ja) * | 2003-10-23 | 2011-05-19 | Panasonic Corp | スペクトル符号化装置、スペクトル復号化装置、音響信号送信装置、音響信号受信装置、およびこれらの方法 |
Family Cites Families (33)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP3246715B2 (ja) * | 1996-07-01 | 2002-01-15 | 松下電器産業株式会社 | オーディオ信号圧縮方法,およびオーディオ信号圧縮装置 |
| EP1351401B1 (en) * | 2001-07-13 | 2009-01-14 | Panasonic Corporation | Audio signal decoding device and audio signal encoding device |
| DE602004032587D1 (de) * | 2003-09-16 | 2011-06-16 | Panasonic Corp | Codierungsvorrichtung und Decodierungsvorrichtung |
| US7668711B2 (en) * | 2004-04-23 | 2010-02-23 | Panasonic Corporation | Coding equipment |
| CN101656077B (zh) * | 2004-05-14 | 2012-08-29 | 松下电器产业株式会社 | 音频编码装置、音频编码方法以及通信终端和基站装置 |
| JP4977471B2 (ja) * | 2004-11-05 | 2012-07-18 | パナソニック株式会社 | 符号化装置及び符号化方法 |
| JP4899359B2 (ja) * | 2005-07-11 | 2012-03-21 | ソニー株式会社 | 信号符号化装置及び方法、信号復号装置及び方法、並びにプログラム及び記録媒体 |
| US20070299655A1 (en) * | 2006-06-22 | 2007-12-27 | Nokia Corporation | Method, Apparatus and Computer Program Product for Providing Low Frequency Expansion of Speech |
| JP5339919B2 (ja) * | 2006-12-15 | 2013-11-13 | パナソニック株式会社 | 符号化装置、復号装置およびこれらの方法 |
| WO2009059633A1 (en) * | 2007-11-06 | 2009-05-14 | Nokia Corporation | An encoder |
| CN101471072B (zh) * | 2007-12-27 | 2012-01-25 | 华为技术有限公司 | 高频重建方法、编码装置和解码装置 |
| US8532998B2 (en) * | 2008-09-06 | 2013-09-10 | Huawei Technologies Co., Ltd. | Selective bandwidth extension for encoding/decoding audio/speech signal |
| WO2010028301A1 (en) * | 2008-09-06 | 2010-03-11 | GH Innovation, Inc. | Spectrum harmonic/noise sharpness control |
| US9037474B2 (en) | 2008-09-06 | 2015-05-19 | Huawei Technologies Co., Ltd. | Method for classifying audio signal into fast signal or slow signal |
| WO2010028292A1 (en) * | 2008-09-06 | 2010-03-11 | Huawei Technologies Co., Ltd. | Adaptive frequency prediction |
| EP2224433B1 (en) | 2008-09-25 | 2020-05-27 | Lg Electronics Inc. | An apparatus for processing an audio signal and method thereof |
| CN101751926B (zh) | 2008-12-10 | 2012-07-04 | 华为技术有限公司 | 信号编码、解码方法及装置、编解码系统 |
| ES3023486T3 (en) * | 2009-01-16 | 2025-06-02 | Dolby Int Ab | Cross product enhanced harmonic transposition |
| KR101661374B1 (ko) | 2009-02-26 | 2016-09-29 | 파나소닉 인텔렉츄얼 프로퍼티 코포레이션 오브 아메리카 | 부호화 장치, 복호 장치 및 이들 방법 |
| CN101521014B (zh) * | 2009-04-08 | 2011-09-14 | 武汉大学 | 音频带宽扩展编解码装置 |
| CO6440537A2 (es) * | 2009-04-09 | 2012-05-15 | Fraunhofer Ges Forschung | Aparato y metodo para generar una señal de audio de sintesis y para codificar una señal de audio |
| WO2011048820A1 (ja) | 2009-10-23 | 2011-04-28 | パナソニック株式会社 | 符号化装置、復号装置およびこれらの方法 |
| US20130030796A1 (en) * | 2010-01-14 | 2013-01-31 | Panasonic Corporation | Audio encoding apparatus and audio encoding method |
| US9093080B2 (en) * | 2010-06-09 | 2015-07-28 | Panasonic Intellectual Property Corporation Of America | Bandwidth extension method, bandwidth extension apparatus, program, integrated circuit, and audio decoding apparatus |
| ES2484795T3 (es) * | 2010-07-19 | 2014-08-12 | Dolby International Ab | Procesamiento de señales de audio durante la reconstrucción de alta frecuencia |
| US8924222B2 (en) | 2010-07-30 | 2014-12-30 | Qualcomm Incorporated | Systems, methods, apparatus, and computer-readable media for coding of harmonic signals |
| JP5707842B2 (ja) * | 2010-10-15 | 2015-04-30 | ソニー株式会社 | 符号化装置および方法、復号装置および方法、並びにプログラム |
| CN104916290B (zh) * | 2011-02-18 | 2018-11-06 | 株式会社Ntt都科摩 | 语音解码装置、语音编码装置、语音解码方法以及语音编码方法 |
| CN102800317B (zh) * | 2011-05-25 | 2014-09-17 | 华为技术有限公司 | 信号分类方法及设备、编解码方法及设备 |
| CN102208188B (zh) | 2011-07-13 | 2013-04-17 | 华为技术有限公司 | 音频信号编解码方法和设备 |
| CN103718240B (zh) * | 2011-09-09 | 2017-02-15 | 松下电器(美国)知识产权公司 | 编码装置、解码装置、编码方法和解码方法 |
| JP2013122985A (ja) | 2011-12-12 | 2013-06-20 | Toshiba Corp | 半導体記憶装置 |
| ES2836194T3 (es) * | 2013-06-11 | 2021-06-24 | Fraunhofer Ges Forschung | Dispositivo y procedimiento para la extensión de ancho de banda para señales acústicas |
-
2014
- 2014-06-10 ES ES14811296T patent/ES2836194T3/es active Active
- 2014-06-10 MX MX2015016109A patent/MX353240B/es active IP Right Grant
- 2014-06-10 EP EP20178265.3A patent/EP3731226A1/en active Pending
- 2014-06-10 US US14/894,062 patent/US9489959B2/en active Active
- 2014-06-10 RU RU2018121035A patent/RU2688247C2/ru active
- 2014-06-10 PT PT148112964T patent/PT3010018T/pt unknown
- 2014-06-10 CN CN201480031440.1A patent/CN105408957B/zh active Active
- 2014-06-10 JP JP2015522543A patent/JP6407150B2/ja active Active
- 2014-06-10 RU RU2015151169A patent/RU2658892C2/ru active
- 2014-06-10 KR KR1020157033759A patent/KR102158896B1/ko active Active
- 2014-06-10 BR BR112015029574-6A patent/BR112015029574B1/pt active IP Right Grant
- 2014-06-10 BR BR122020016403-4A patent/BR122020016403B1/pt active IP Right Grant
- 2014-06-10 WO PCT/JP2014/003103 patent/WO2014199632A1/ja not_active Ceased
- 2014-06-10 CN CN202010063428.6A patent/CN111477245B/zh active Active
- 2014-06-10 EP EP14811296.4A patent/EP3010018B1/en active Active
-
2016
- 2016-10-05 US US15/286,030 patent/US9747908B2/en active Active
-
2017
- 2017-07-25 US US15/659,023 patent/US10157622B2/en active Active
-
2018
- 2018-09-18 JP JP2018173725A patent/JP6773737B2/ja active Active
- 2018-09-18 JP JP2018173731A patent/JP2019008317A/ja active Pending
- 2018-12-13 US US16/219,656 patent/US10522161B2/en active Active
-
2020
- 2020-10-01 JP JP2020166633A patent/JP7330934B2/ja active Active
Patent Citations (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2003108197A (ja) * | 2001-07-13 | 2003-04-11 | Matsushita Electric Ind Co Ltd | オーディオ信号復号化装置およびオーディオ信号符号化装置 |
| JP2011100159A (ja) * | 2003-10-23 | 2011-05-19 | Panasonic Corp | スペクトル符号化装置、スペクトル復号化装置、音響信号送信装置、音響信号受信装置、およびこれらの方法 |
Cited By (26)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US11049508B2 (en) | 2014-07-28 | 2021-06-29 | Fraunhofer-Gesellschaft zur Förderung der angewandten Forschung e.V. | Audio encoder and decoder using a frequency domain processor with full-band gap filling and a time domain processor |
| JP2017523473A (ja) * | 2014-07-28 | 2017-08-17 | フラウンホーファー−ゲゼルシャフト・ツール・フェルデルング・デル・アンゲヴァンテン・フォルシュング・アインゲトラーゲネル・フェライン | 全帯域ギャップ充填を備えた周波数ドメインプロセッサと時間ドメインプロセッサとを使用するオーディオ符号器及び復号器 |
| US12080310B2 (en) | 2014-07-28 | 2024-09-03 | Fraunhofer-Gesellschaft zur Förderung der angewandten Forschung e.V. | Audio encoder and decoder using a frequency domain processor with full-band gap filling and a time domain processor |
| US11929084B2 (en) | 2014-07-28 | 2024-03-12 | Fraunhofer-Gesellschaft zur Förderung der angewandten Forschung e.V. | Audio encoder and decoder using a frequency domain processor with full-band gap filling and a time domain processor |
| US11915712B2 (en) | 2014-07-28 | 2024-02-27 | Fraunhofer-Gesellschaft zur Förderung der angewandten Forschung e.V. | Audio encoder and decoder using a frequency domain processor, a time domain processor, and a cross processing for continuous initialization |
| US11410668B2 (en) | 2014-07-28 | 2022-08-09 | Fraunhofer-Gesellschaft zur Förderung der angewandten Forschung e.V. | Audio encoder and decoder using a frequency domain processor, a time domain processor, and a cross processing for continuous initialization |
| US10236007B2 (en) | 2014-07-28 | 2019-03-19 | Fraunhofer-Gesellschaft Zur Foerderung Der Angewandten Forschung E.V. | Audio encoder and decoder using a frequency domain processor , a time domain processor, and a cross processing for continuous initialization |
| US10332535B2 (en) | 2014-07-28 | 2019-06-25 | Fraunhofer-Gesellschaft zur Förderung der angewandten Forschung e.V. | Audio encoder and decoder using a frequency domain processor with full-band gap filling and a time domain processor |
| US11842743B2 (en) | 2015-03-13 | 2023-12-12 | Dolby International Ab | Decoding audio bitstreams with enhanced spectral band replication metadata in at least one fill element |
| US11417350B2 (en) | 2015-03-13 | 2022-08-16 | Dolby International Ab | Decoding audio bitstreams with enhanced spectral band replication metadata in at least one fill element |
| US12260869B2 (en) | 2015-03-13 | 2025-03-25 | Dolby International Ab | Decoding audio bitstreams with enhanced spectral band replication metadata in at least one fill element |
| US12094477B2 (en) | 2015-03-13 | 2024-09-17 | Dolby International Ab | Decoding audio bitstreams with enhanced spectral band replication metadata in at least one fill element |
| US11664038B2 (en) | 2015-03-13 | 2023-05-30 | Dolby International Ab | Decoding audio bitstreams with enhanced spectral band replication metadata in at least one fill element |
| US11367455B2 (en) | 2015-03-13 | 2022-06-21 | Dolby International Ab | Decoding audio bitstreams with enhanced spectral band replication metadata in at least one fill element |
| CN105280189B (zh) * | 2015-09-16 | 2019-01-08 | 深圳广晟信源技术有限公司 | 带宽扩展编码和解码中高频生成的方法和装置 |
| CN105280189A (zh) * | 2015-09-16 | 2016-01-27 | 深圳广晟信源技术有限公司 | 带宽扩展编码和解码中高频生成的方法和装置 |
| US11100939B2 (en) | 2015-12-14 | 2021-08-24 | Fraunhofer-Gesellschaft Zur Foerderung Der Angewandten Forschung E.V. | Apparatus and method for processing an encoded audio signal by a mapping drived by SBR from QMF onto MCLT |
| CN108701467B (zh) * | 2015-12-14 | 2023-12-08 | 弗劳恩霍夫应用研究促进协会 | 处理经编码音频信号的装置及方法 |
| US11862184B2 (en) | 2015-12-14 | 2024-01-02 | Fraunhofer-Gesellschaft Zur Foerderung Der Angewandten Forschung E.V. | Apparatus and method for processing an encoded audio signal by upsampling a core audio signal to upsampled spectra with higher frequencies and spectral width |
| KR102625047B1 (ko) * | 2015-12-14 | 2024-01-16 | 프라운호퍼 게젤샤프트 쭈르 푀르데룽 데어 안겐반텐 포르슝 에. 베. | 인코딩된 오디오 신호를 처리하기 위한 장치 및 방법 |
| CN108701467A (zh) * | 2015-12-14 | 2018-10-23 | 弗劳恩霍夫应用研究促进协会 | 处理经编码音频信号的装置及方法 |
| KR20210054052A (ko) * | 2015-12-14 | 2021-05-12 | 프라운호퍼 게젤샤프트 쭈르 푀르데룽 데어 안겐반텐 포르슝 에. 베. | 인코딩된 오디오 신호를 처리하기 위한 장치 및 방법 |
| EP3435376A1 (en) | 2017-07-28 | 2019-01-30 | Fujitsu Limited | Audio encoding apparatus and audio encoding method |
| US10896684B2 (en) | 2017-07-28 | 2021-01-19 | Fujitsu Limited | Audio encoding apparatus and audio encoding method |
| CN108630212B (zh) * | 2018-04-03 | 2021-05-07 | 湖南商学院 | 非盲带宽扩展中高频激励信号的感知重建方法与装置 |
| CN108630212A (zh) * | 2018-04-03 | 2018-10-09 | 湖南商学院 | 非盲带宽扩展中高频激励信号的感知重建方法与装置 |
Also Published As
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| JP7330934B2 (ja) | 音響信号の帯域幅拡張を行う装置及び方法 | |
| CN105453176B (zh) | 智能间隙填充框架内使用双声道处理的音频编码器、音频解码器及相关方法 | |
| JP5970014B2 (ja) | オーディオエンコーダおよび帯域幅拡張デコーダ | |
| US9406307B2 (en) | Method and apparatus for polyphonic audio signal prediction in coding and networking systems | |
| CN101662288B (zh) | 音频编码、解码方法及装置、系统 | |
| KR101398189B1 (ko) | 음성수신장치 및 음성수신방법 | |
| JP2009515212A (ja) | オーディオ圧縮 | |
| KR20140004086A (ko) | 반대 위상의 채널들에 대한 개선된 스테레오 파라메트릭 인코딩/디코딩 | |
| WO2008072737A1 (ja) | 符号化装置、復号装置およびこれらの方法 | |
| CN104170009A (zh) | 感知音频编解码器中的谐波信号的相位相干性控制 | |
| WO2012053150A1 (ja) | 音声符号化装置および音声復号化装置 | |
| KR20160138373A (ko) | 부호화 장치, 복호 장치, 부호화 방법, 복호 방법, 및 프로그램 | |
| US20210233544A1 (en) | Perceptual audio coding with adaptive non-uniform time/frequency tiling using subband merging and the time domain aliasing reduction | |
| Lin et al. | Adaptive bandwidth extension of low bitrate compressed audio based on spectral correlation | |
| AU2015203736C1 (en) | Audio encoder and bandwidth extension decoder |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| WWE | Wipo information: entry into national phase |
Ref document number: 201480031440.1 Country of ref document: CN |
|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 14811296 Country of ref document: EP Kind code of ref document: A1 |
|
| ENP | Entry into the national phase |
Ref document number: 2015522543 Country of ref document: JP Kind code of ref document: A |
|
| WWE | Wipo information: entry into national phase |
Ref document number: MX/A/2015/016109 Country of ref document: MX |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 14894062 Country of ref document: US |
|
| ENP | Entry into the national phase |
Ref document number: 20157033759 Country of ref document: KR Kind code of ref document: A |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 122020016403 Country of ref document: BR |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 2015151169 Country of ref document: RU |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 2014811296 Country of ref document: EP |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| REG | Reference to national code |
Ref country code: BR Ref legal event code: B01A Ref document number: 112015029574 Country of ref document: BR |
|
| ENP | Entry into the national phase |
Ref document number: 112015029574 Country of ref document: BR Kind code of ref document: A2 Effective date: 20151126 |






