JP7210658B2

JP7210658B2 - Audio processing unit and method of decoding encoded audio bitstream

Info

Publication number: JP7210658B2
Application number: JP2021123219A
Authority: JP
Inventors: ヴィレモーズ，ラルス; プルンハーゲン，ヘイコ; エクストランド，ペール
Original assignee: ドルビー・インターナショナル・アーベー
Priority date: 2015-03-13
Filing date: 2021-07-28
Publication date: 2023-01-23
Anticipated expiration: 2036-03-10
Also published as: ZA201805869B; IL285643B2; ZA201905506B; IL256786A; HK1259131A1; HK1259544A1; JP2021167981A; HK1259548A1; HUE053954T2; JP2020079963A; HK1259302A1; HK1259406A1; TWI693595B; ZA201901691B; IL285643B; ES2867477T3; IL285643A; UA119808C2; JP6922017B2; DK3268956T3

Description

関連出願への相互参照
本願は2015年3月13日に出願された欧州特許出願第15159067.6号および16年3月16日に出願された米国仮出願第62/133,800号の優先権を主張するものである。各文献の内容は参照によってその全体において組み込まれる。
技術分野
本発明は、オーディオ信号処理に関する。いくつかの実施形態はオーディオ・ビットストリーム（たとえばMPEG-4 AACフォーマットをもつビットストリーム）のエンコードおよびデコードに関する。他の実施形態は、そのようなビットストリームの、eSBR処理を実行するよう構成されておらずそのようなメタデータを無視するレガシー・デコーダによるデコードに関し、あるいはそのようなメタデータを含まないオーディオ・ビットストリームのデコードに関し、それは該ビットストリームに応答してeSBR制御データを生成することによることを含む。 CROSS REFERENCE TO RELATED APPLICATIONS This application claims priority from European Patent Application No. 15159067.6 filed March 13, 2015 and US Provisional Application No. 62/133,800 filed March 16, 2016 is. The contents of each document are incorporated by reference in their entirety.
TECHNICAL FIELD The present invention relates to audio signal processing. Some embodiments relate to encoding and decoding audio bitstreams (eg, bitstreams with MPEG-4 AAC format). Other embodiments relate to decoding of such bitstreams by legacy decoders not configured to perform eSBR processing and ignoring such metadata, or to decoding audio streams that do not include such metadata. For bitstream decoding, it includes by generating eSBR control data in response to the bitstream.

典型的なオーディオ・ビットストリームは、オーディオ・コンテンツの一つまたは複数のチャネルを示すオーディオ・データ（たとえばエンコードされたオーディオ・データ）と、前記オーディオ・データまたはオーディオ・コンテンツの少なくとも一つの特性を示すメタデータとの両方を含む。エンコードされたオーディオ・ビットストリームを生成するための一つのよく知られたフォーマットは、MPEG規格ISO/IEC14496-3:2009に記載されるMPEG-4先進オーディオ符号化（AAC: Advanced Audio Coding）フォーマットである。MPEG-4規格では、AACは「advanced audio coding（先進オーディオ符号化）」を表わし、HE-AACは「high-efficiency advanced audio coding（高効率先進オーディオ符号化）」を表わす。 A typical audio bitstream represents audio data (e.g., encoded audio data) indicative of one or more channels of audio content and at least one characteristic of said audio data or audio content. including both metadata and One well-known format for generating encoded audio bitstreams is the MPEG-4 Advanced Audio Coding (AAC) format described in MPEG standard ISO/IEC14496-3:2009. be. In the MPEG-4 standard, AAC stands for "advanced audio coding" and HE-AAC stands for "high-efficiency advanced audio coding".

MPEG-4 AAC規格はいくつかのオーディオ・プロファイルを定義しており、それらのオーディオ・プロファイルがどのオブジェクトおよび符号化ツールが準拠するエンコーダまたはデコーダにおいて存在しているかを決める。これらのオーディオ・プロファイルのうちの三つは、（１）AACプロファイル、（２）HE-AACプロファイルおよび（３）HE-AAC v2プロファイルである。AACプロファイルはAAC低計算量（AAC low complexity）（または「AAC-LC」）オブジェクト型を含む。AAC-LCオブジェクト型は、若干の調整はあるがMPEG-2 AAC低計算量プロファイルに対応するものであり、スペクトル帯域複製（spectral band replication）（「SBR」）オブジェクト型もパラメトリック・ステレオ（parametric stereo）（「PS」）オブジェクト型も含まない。HE-AACプロファイルはAACプロファイルの上位集合であって、追加的にSBRオブジェクト型を含む。HE-AAC v2プロファイルはHE-AACプロファイルの上位集合であって、追加的にPSオブジェクト型を含む。 The MPEG-4 AAC standard defines several audio profiles that determine which objects and coding tools are present in a conforming encoder or decoder. Three of these audio profiles are (1) AAC profile, (2) HE-AAC profile and (3) HE-AAC v2 profile. The AAC profile includes AAC low complexity (or "AAC-LC") object types. The AAC-LC object type corresponds to the MPEG-2 AAC low-complexity profile with some adjustments, and the spectral band replication ("SBR") object type also supports parametric stereo. ) ("PS") object type. The HE-AAC profile is a superset of the AAC profile and additionally contains the SBR object type. The HE-AAC v2 profile is a superset of the HE-AAC profile and additionally contains the PS object type.

SBRオブジェクト型は、スペクトル帯域複製ツールを含む。これは、知覚的オーディオ・コーデックの圧縮効率を著しく改善する重要な符号化ツールである。SBRは受信器側で（たとえばデコーダにおいて）オーディオ信号の高周波数成分を再構成する。そのため、エンコーダは低周波数成分をエンコードして伝送するだけでよく、低データ・レートにおいてずっと高いオーディオ品質を許容する。SBRは、データ・レートを削減するために以前に打ち切りされた高調波のシーケンスを、エンコーダから得られる利用可能な帯域幅制限された信号および制御データから複製することに基づく。トーン様成分とノイズ様成分の間の比は適応的な逆フィルタリングならびにノイズおよび正弦波の任意的な追加によって維持される。MPEG-4 AAC規格では、SBRツールは、いくつかの隣り合う直交ミラー・フィルタ（QMF）サブバンドがオーディオ信号の伝送された低域部分から、デコーダにおいて生成されるオーディオ信号の高域部分にコピーされる、スペクトル・パッチング（spectral patching）を実行する。 The SBR object type contains spectral band replication tools. It is an important coding tool that significantly improves the compression efficiency of perceptual audio codecs. SBR reconstructs the high frequency components of the audio signal at the receiver side (eg at the decoder). As such, the encoder need only encode and transmit the low frequency components, allowing much higher audio quality at low data rates. SBR is based on duplicating a previously truncated sequence of harmonics to reduce the data rate from the available bandwidth limited signal and control data obtained from the encoder. The ratio between tone-like and noise-like components is maintained by adaptive inverse filtering and optional addition of noise and sinusoids. In the MPEG-4 AAC standard, the SBR tool copies several adjacent quadrature mirror filter (QMF) subbands from the transmitted lowband portion of the audio signal to the highband portion of the audio signal generated at the decoder. Performs spectral patching.

MPEG規格ISO/IEC14496-3:2009MPEG standardISO/IEC14496-3:2009

スペクトル・パッチングは、比較的低いクロスオーバー周波数をもつ音楽コンテンツのようなある種のオーディオ型については理想的ではないことがある。したがって、スペクトル帯域複製を改善するための技法が必要とされている。 Spectral patching may not be ideal for certain audio types, such as musical content with relatively low crossover frequencies. Therefore, techniques are needed to improve spectral band duplication.

第一のクラスの実施形態は、メモリと、ビットストリーム・ペイロード・フォーマット解除器と、デコード・サブシステムとを含むオーディオ処理ユニットに関する。メモリは、エンコードされたオーディオ・ビットストリーム（たとえばMPEG-4 AACビットストリーム）の少なくとも一つのブロックを記憶するよう構成される。ビットストリーム・ペイロード・フォーマット解除器は、エンコードされたオーディオ・ブロックを多重分離するよう構成される。デコード・サブシステムは、エンコードされたオーディオ・ブロックのオーディオ・コンテンツをデコードするよう構成される。エンコードされたオーディオ・ブロックは、充填要素（fill element）を含む。充填要素は、該充填要素の先頭を示す識別子と、該識別子後の充填データとをもつ。充填データは、そのエンコードされたオーディオ・ブロックのオーディオ・コンテンツに対して向上スペクトル帯域複製（eSBR: enhanced spectral band replication）処理が実行されるべきかどうかを同定する少なくとも一つのフラグを含む。 A first class of embodiments relates to an audio processing unit that includes a memory, a bitstream payload deformatter, and a decoding subsystem. The memory is configured to store at least one block of an encoded audio bitstream (eg MPEG-4 AAC bitstream). A bitstream payload de-formatter is configured to demultiplex the encoded audio blocks. A decoding subsystem is configured to decode the audio content of the encoded audio blocks. An encoded audio block contains a fill element. A filler element has an identifier indicating the beginning of the filler element and filler data after the identifier. The fill data includes at least one flag identifying whether enhanced spectral band replication (eSBR) processing should be performed on the audio content of the encoded audio block.

第二のクラスの実施形態は、エンコードされたオーディオ・ビットストリームをデコードするための方法に関する。本方法は、エンコードされたオーディオ・ビットストリームの少なくとも一つのブロックを受領し、前記エンコードされたオーディオ・ビットストリームの前記少なくとも一つのブロックの少なくともいくつかの部分を多重分離し、前記エンコードされたオーディオ・ビットストリームの前記少なくとも一つのブロックの少なくともいくつかの部分をデコードすることを含む。前記エンコードされたオーディオ・ビットストリームの前記少なくとも一つのブロックは、充填要素（fill element）を含む。充填要素は、該充填要素の先頭を示す識別子と、該識別子後の充填データとをもつ。充填データは、そのエンコードされたオーディオ・ビットストリームの前記少なくとも一つのブロックのオーディオ・コンテンツに対して向上スペクトル帯域複製（eSBR: enhanced spectral band replication）処理が実行されるべきかどうかを同定する少なくとも一つのフラグを含む。 A second class of embodiments relates to a method for decoding an encoded audio bitstream. The method includes receiving at least one block of an encoded audio bitstream, demultiplexing at least some portions of the at least one block of the encoded audio bitstream, and decoding the encoded audio bitstream. • decoding at least some portions of said at least one block of the bitstream; The at least one block of the encoded audio bitstream includes a fill element. A filler element has an identifier indicating the beginning of the filler element and filler data after the identifier. The filling data at least one identifying whether enhanced spectral band replication (eSBR) processing should be performed on the audio content of the at least one block of the encoded audio bitstream. contains two flags.

他のクラスの実施形態は、向上スペクトル帯域複製（eSBR: enhanced spectral band replication）処理が実行されるべきかどうかを同定するメタデータを含むオーディオ・ビットストリームをエンコードおよびトランスコードすることに関する。 Another class of embodiments relates to encoding and transcoding audio bitstreams that include metadata identifying whether enhanced spectral band replication (eSBR) processing should be performed.

本発明の方法のある実施形態を実行するよう構成されうるシステムの実施形態のブロック図である。1 is a block diagram of an embodiment of a system that may be configured to carry out certain embodiments of the method of the present invention; FIG. 本発明のオーディオ処理ユニットの実施形態であるエンコーダのブロック図である。1 is a block diagram of an encoder that is an embodiment of an audio processing unit of the invention; FIG. 本発明のオーディオ処理ユニットの実施形態であるデコーダと、任意的にはそれに結合された後処理器をも含むシステムのブロック図である。1 is a block diagram of a system including a decoder, which is an embodiment of an audio processing unit of the invention, and optionally also a post-processor coupled thereto; FIG. 本発明のオーディオ処理ユニットの実施形態であるデコーダのブロック図である。1 is a block diagram of a decoder, which is an embodiment of an audio processing unit of the invention; FIG. 本発明のオーディオ処理ユニットのもう一つの実施形態であるデコーダのブロック図である。FIG. 4 is a block diagram of a decoder, another embodiment of the audio processing unit of the present invention; 本発明のオーディオ処理ユニットのもう一つの実施形態のブロック図である。FIG. 4 is a block diagram of another embodiment of the audio processing unit of the present invention; 分割されたセグメントを含むMPEG-4 AACビットストリームのブロックを示す図である。Fig. 3 shows a block of an MPEG-4 AAC bitstream containing segmented segments;

請求項を含む本開示を通じて、信号またはデータ「に対して」動作を実行する（たとえば信号またはデータをフィルタリングする、スケーリングする、変換するまたは利得を適用する）という表現は、信号またはデータに対して直接的に、または信号またはデータの処理されたバージョンに対して（たとえば、予備的なフィルタリングまたは前処理を該動作の実行に先立って受けている前記信号のバージョンに対して）該動作を実行することを表わすために広義で使用される。 Throughout this disclosure, including the claims, references to performing an operation "on" a signal or data (e.g., filtering, scaling, transforming, or applying a gain to the signal or data) include perform the operation directly or on a processed version of the signal or data (e.g., on a version of the signal that has undergone preliminary filtering or preprocessing prior to performing the operation) used broadly to mean

請求項を含む本開示を通じて、「オーディオ処理ユニット」という表現は、オーディオ・データを処理するよう構成されているシステム、デバイスまたは装置を表わす広義で使用される。オーディオ処理ユニットの例は、エンコーダ（たとえばトランスコーダ）、デコーダ、コーデック、前処理システム、後処理システムおよびビットストリーム処理システム（時にビットストリーム処理ツールと称される）を含むがそれに限られない。携帯電話、テレビジョン、ラップトップおよびタブレット・コンピュータといった事実上あらゆる消費者電子装置がオーディオ処理ユニットを含む。 Throughout this disclosure, including the claims, the expression "audio processing unit" is used broadly to denote a system, device or apparatus configured to process audio data. Examples of audio processing units include, but are not limited to, encoders (eg, transcoders), decoders, codecs, pre-processing systems, post-processing systems and bitstream processing systems (sometimes referred to as bitstream processing tools). Virtually every consumer electronic device such as mobile phones, televisions, laptops and tablet computers includes an audio processing unit.

請求項を含む本開示を通じて、「結合する」または「結合される」という用語は、直接的または間接的な接続を意味するために広義で使われる。よって、第一の装置が第二の装置に結合する場合、その接続は、直接接続を通じてであってもよいし、他の装置および接続を介した間接的な接続を通じてであってもよい。さらに、他のコンポーネントの中にまたは他のコンポーネントと一緒に統合されたコンポーネントも互いに結合される。 Throughout this disclosure, including the claims, the terms "coupled" or "coupled" are used broadly to mean a direct or indirect connection. Thus, when a first device couples to a second device, the connection may be through a direct connection or through an indirect connection through other devices and connections. In addition, components integrated within or with other components are also coupled together.

〈本発明の実施形態の詳細な説明〉
MPEG-4 AAC規格は、エンコードされたMPEG-4 AACビットストリームが、該ビットストリームのオーディオ・コンテンツをデコードするためにデコーダによって適用されるべき（もし適用されるべきものがあるとして）SBR処理のそれぞれの型を示すおよび／またはそのようなSBR処理を制御するおよび／または該ビットストリームのオーディオ・コンテンツをデコードするために用いられるべき少なくとも一つのSBRツールの少なくとも一つの特性またはパラメータを示すメタデータを含むことを考えている。ここで、MPEG-4 AAC規格で記述または言及されているこの型のメタデータを表わすために「SBRメタデータ」という表現を使う。 <Detailed description of embodiments of the present invention>
The MPEG-4 AAC standard specifies that an encoded MPEG-4 AAC bitstream should be subjected to SBR processing (if any) to be applied by decoders to decode the audio content of the bitstream. Metadata indicating respective types and/or indicating at least one characteristic or parameter of at least one SBR tool to be used to control such SBR processing and/or to decode the audio content of the bitstream. I'm thinking of including Here we use the expression "SBR metadata" to denote this type of metadata described or referred to in the MPEG-4 AAC standard.

MPEG-4 AACビットストリームの最上レベルはデータ・ブロック（「raw_data_block」要素）のシーケンスであり、各データ・ブロックは、（典型的には1024または960サンプルの時間期間にわたる）オーディオ・データおよび関係した情報および／または他のデータを含む、データのセグメント（本稿では「ブロックと称される」）である。ここで、一つの（二つ以上ではない）「raw_data_block」要素を決定するまたは示すオーディオ・データ（および対応するメタデータおよび任意的には他の関係したデータ）を含むMPEG-4 AACビットストリームのセグメントを表わすために、用語「ブロック」を使う。 The top level of an MPEG-4 AAC bitstream is a sequence of data blocks ("raw_data_block" elements), each data block containing audio data (typically spanning a time period of 1024 or 960 samples) and associated A segment of data (referred to herein as a block) that contains information and/or other data. where one (but not more than two) "raw_data_block" elements of an MPEG-4 AAC bitstream containing audio data (and corresponding metadata and optionally other related data) are determined or indicated We use the term "block" to represent a segment.

MPEG-4 AACビットストリームの各ブロックは、いくつかのシンタックス要素を含むことができる（そのそれぞれも、ビットストリームにおけるデータのセグメントとして具現される）。七つの型のそのようなシンタックス要素がMPEG-4 AAC規格において定義されている。各シンタックス要素はデータ要素「id_syn_ele」の異なる値によって識別される。シンタックス要素の例は「single_channel_element()」、「channel_pair_element()」および「fill_element()」を含む。単一チャネル要素（single channel element）は、単一のオーディオ・チャネルのオーディオ・データ（モノフォニック・オーディオ信号）を含むコンテナである。チャネル対要素（channel pair element）は二つのオーディオ・チャネルのオーディオ・データ（すなわち、ステレオ・オーディオ信号）を含む。 Each block of an MPEG-4 AAC bitstream can contain several syntax elements (each of which is also embodied as a segment of data in the bitstream). Seven types of such syntax elements are defined in the MPEG-4 AAC standard. Each syntax element is identified by a different value of the data element "id_syn_ele". Examples of syntax elements include "single_channel_element()", "channel_pair_element()" and "fill_element()". A single channel element is a container containing the audio data of a single audio channel (a monophonic audio signal). A channel pair element contains audio data for two audio channels (ie, a stereo audio signal).

充填要素（fill element）は、識別子（たとえば上記の要素「id_syn_ele」の値）および「充填データ」（fill data）と称されるそれに続くデータを含む情報のコンテナである。充填要素は、歴史的には、一定レート・チャネルを通じて伝送されるべきビットストリームの瞬時ビットレートを調整するために使われてきた。各ブロックに適切な量の充填データを加えることによって、一定データ・レートが達成されうる。 A fill element is a container of information that includes an identifier (eg, the value of the element "id_syn_ele" above) and subsequent data called "fill data". Filling factors have historically been used to adjust the instantaneous bitrate of a bitstream to be transmitted over a constant rate channel. A constant data rate can be achieved by adding an appropriate amount of padding data to each block.

本発明の諸実施形態によれば、充填データは、ビットストリームにおいて伝送されることのできるデータ（たとえばメタデータ）の型を拡張する一つまたは複数の拡張ペイロードを含みうる。新しい型のデータを含む充填データをもつビットストリームを受け取るデコーダは、任意的に、該ビットストリームを受け取る装置（たとえばデコーダ）によって、該装置の機能を拡張するために使用されてもよい。このように、当業者には理解できるように、充填要素は特殊な型のデータ構造であり、オーディオ・データ（たとえばチャネル・データを含むオーディオ・ペイロード）を伝送するために典型的に使われるデータ構造とは異なる。 According to embodiments of the invention, the filler data may include one or more extension payloads that extend the types of data (eg, metadata) that can be transmitted in the bitstream. A decoder that receives a bitstream with fill data containing new types of data may optionally be used by a device (eg, a decoder) that receives the bitstream to extend the functionality of the device. Thus, as those skilled in the art will appreciate, a filler element is a special type of data structure that is typically used to carry audio data (e.g., an audio payload containing channel data). Different from structure.

本発明のいくつかの実施形態では、充填要素を識別するために使われる識別子は、0x6の値をもつ、三ビットの、最上位ビットが最初に伝送される符号なし整数（unsigned integer transmitted most significant bit first）（「uimsbf」）からなっていてもよい。一つのブロックにおいて、同じ型のシンタックス要素のいくつかのインスタンス（たとえばいくつかの充填要素）が生起してもよい。 In some embodiments of the present invention, identifiers used to identify filler elements are three-bit unsigned integers with a value of 0x6, most significant transmitted most significant bit first. bit first) ("uimsbf"). Several instances of the same type of syntax element (eg several filler elements) may occur in one block.

オーディオ・ビットストリームをエンコードするためのもう一つの規格は、MPEG統合音声音響符号化（USAC: Unified Speech and Audio Coding）規格（ISO/IEC 23003-3:2012）である。MPEG USAC規格は、スペクトル帯域複製処理（MPEG-4 AAC規格に記述されるSBR処理を含み、他の向上された形のスペクトル帯域複製処理をも含む）を使ってオーディオ・コンテンツをエンコードおよびデコードすることを記述している。この処理は、MPEG-4 AAC規格において記述されているSBRツールの集合の、拡張され、向上されたバージョンのスペクトル帯域複製ツール（本稿では時に「向上SBRツール」または「eSBRツール」と称される）を適用する。このように、eSBR（USAC規格において定義されている）はSBR（MPEG-4 AAC規格において定義されている）に対する改良である。 Another standard for encoding audio bitstreams is the MPEG Unified Speech and Audio Coding (USAC) standard (ISO/IEC 23003-3:2012). The MPEG USAC standard encodes and decodes audio content using spectral band replication processing (including SBR processing described in the MPEG-4 AAC standard, as well as other enhanced forms of spectral band replication processing). It describes that This process is an extended and enhanced version of the spectral band replication tool (sometimes referred to herein as the "enhanced SBR tool" or "eSBR tool") of the set of SBR tools described in the MPEG-4 AAC standard. ) apply. Thus, eSBR (defined in the USAC standard) is an improvement over SBR (defined in the MPEG-4 AAC standard).

本稿において、「向上SBR処理」（enhanced SBR processing）（または「eSBR処理」）という表現は、MPEG-4 AACにおいて記述または言及されていない少なくとも一つのeSBRツール（たとえば、MPEG USAC規格において記述または言及されている少なくとも一つのeSBRツール）を使うスペクトル帯域複製処理を表わすために使う。そのようなeSBRツールの例は高調波転換（harmonic transposition）、QMFパッチング追加的前処理もしくは「前置平坦化（pre-flattening）」およびサブバンド・サンプル間時間包絡整形（Temporal Envelope Shaping）または「インターTES」である。 In this document, the term "enhanced SBR processing" (or "eSBR processing") refers to at least one eSBR tool not described or mentioned in the MPEG-4 AAC (e.g. MPEG USAC standard). used to represent a spectral band duplication process that uses at least one eSBR tool (provided). Examples of such eSBR tools are harmonic transposition, QMF patching additional preprocessing or "pre-flattening" and subband-to-sample temporal envelope shaping or " Inter TES”.

MPEG USAC規格に従って生成されたビットストリーム（本稿では時にUSACビットストリームと称される）は、エンコードされたオーディオ・コンテンツを含み、典型的には、該USACビットストリームのオーディオ・コンテンツをデコードするためにデコーダによって適用されるべきスペクトル帯域複製処理のそれぞれの型を示すメタデータおよび／またはそのようなスペクトル帯域複製処理を制御するおよび／または該USACビットストリームのオーディオ・コンテンツをデコードするために用いられるべき少なくとも一つのSBRツールおよび／またはeSBRツールの少なくとも一つの特性またはパラメータを示すメタデータを含む。 A bitstream generated according to the MPEG USAC standard (sometimes referred to herein as a USAC bitstream) contains encoded audio content, and typically a decoding method is used to decode the audio content of the USAC bitstream. Metadata indicating each type of spectral band replication processing to be applied by a decoder and/or to be used to control such spectral band replication processing and/or to decode audio content of said USAC bitstream. Metadata indicative of at least one characteristic or parameter of at least one SBR tool and/or eSBR tool.

ここでは、「向上SBRメタデータ」（または「eSBRメタデータ」）という表現は、エンコードされたオーディオ・ビットストリーム（たとえばUSACビットストリーム）のオーディオ・コンテンツをデコードするためにデコーダによって適用されるべきスペクトル帯域複製処理のそれぞれの型を示すおよび／またはそのようなスペクトル帯域複製処理を制御するおよび／またはそのようなオーディオ・コンテンツをデコードするために用いられるべき少なくとも一つのSBRツールおよび／またはeSBRツールの少なくとも一つの特性またはパラメータを示すメタデータであって、MPEG-4 AAC規格において記述または言及されていないものを表わすために使う。eSBRメタデータの例は、MPEG USAC規格において記述または言及されているがMPEG-4 AAC規格では記述も言及もされていない（スペクトル帯域複製処理を示すまたは制御するための）メタデータである。このように、本稿でのeSBRメタデータは、SBRメタデータではないメタデータを表わし、本稿でのSBRメタデータはeSBRメタデータではないメタデータを表わす。 Here, the expression "enhanced SBR metadata" (or "eSBR metadata") refers to the spectrum to be applied by a decoder to decode the audio content of an encoded audio bitstream (e.g. USAC bitstream). of at least one SBR tool and/or eSBR tool to be used to indicate respective types of band duplicating processes and/or control such spectral band duplicating processes and/or decode such audio content; Used to represent metadata indicating at least one property or parameter that is not described or mentioned in the MPEG-4 AAC standard. An example of eSBR metadata is metadata (to indicate or control the spectral band duplication process) that is described or mentioned in the MPEG USAC standard but not described or mentioned in the MPEG-4 AAC standard. Thus, eSBR metadata in this document refers to metadata that is not SBR metadata, and SBR metadata in this document refers to metadata that is not eSBR metadata.

USACビットストリームは、SBRメタデータおよびeSBRメタデータ両方を含んでいてもよい。より具体的には、USACビットストリームは、デコーダによるeSBR処理の実行を制御するeSBRメタデータおよびデコーダによるSBR処理の実行を制御するSBRメタデータを含んでいてもよい。本発明の典型的な実施形態によれば、eSBRメタデータ（たとえばeSBR固有の構成設定データ）が（本発明に従って）（たとえばSBRペイロードの末尾のsbr_extension()コンテナにおいて）MPEG-4 AACビットストリームに含められる。 A USAC bitstream may contain both SBR metadata and eSBR metadata. More specifically, the USAC bitstream may include eSBR metadata that controls performance of eSBR processing by the decoder and SBR metadata that controls performance of SBR processing by the decoder. According to an exemplary embodiment of the invention, eSBR metadata (e.g. eSBR-specific configuration data) is (according to the invention) embedded in the MPEG-4 AAC bitstream (e.g. in the sbr_extension() container at the end of the SBR payload). be included.

（少なくとも一つのeSBRツールを含む）eSBRツール集合を使ったエンコードされたビットストリームのデコードの間の、デコーダによるeSBR処理の実行は、エンコードの間に打ち切りされた高調波のシーケンスの複製に基づいてオーディオ信号の高周波数帯域を再生成する。そのようなeSBR処理は典型的には、もとのオーディオ信号のスペクトル特性を再現するために、生成された高周波数帯域のスペクトル包絡を調整し、逆フィルタリングを適用し、ノイズおよび正弦波成分を加える。 During decoding of an encoded bitstream using an eSBR toolset (which includes at least one eSBR tool), the performance of eSBR processing by the decoder is based on replicating the sequence of harmonics truncated during encoding. Regenerate the high frequency range of the audio signal. Such eSBR processing typically adjusts the spectral envelope of the generated high frequency band, applies inverse filtering, and removes noise and sinusoidal components to reproduce the spectral characteristics of the original audio signal. Add.

本発明の典型的な実施形態によれば、eSBRメタデータが（たとえばeSBRメタデータである少数の制御ビットが）、エンコードされたオーディオ・ビットストリーム（たとえばMPEG-4 AACビットストリーム）のメタデータ・セグメントの一つまたは複数に含められる。エンコードされたオーディオ・ビットストリームは他のセグメント（オーディオ・データ・セグメント）において、エンコードされたオーディオ・データをも含む。典型的には、ビットストリームの各ブロックの少なくとも一つのそのようなメタデータ・セグメントが充填要素（該充填要素の先頭を示す識別子を含む）であり（または充填要素を含み）、eSBRメタデータは該識別子の後に該充填要素に含められる。 According to exemplary embodiments of the present invention, the eSBR metadata (e.g. a few control bits that are eSBR metadata) is the metadata of an encoded audio bitstream (e.g. an MPEG-4 AAC bitstream). Included in one or more of the segments. The encoded audio bitstream also contains encoded audio data in other segments (audio data segments). Typically, at least one such metadata segment of each block of the bitstream is (or contains) a filler element (including an identifier indicating the beginning of the filler element), and the eSBR metadata is It is included in the filler element after the identifier.

図１は、例示的なオーディオ処理チェーン（オーディオ・データ処理システム）のブロック図であり、該システムの要素の一つまたは複数が本発明の実施形態に従って構成されてもよい。本システムは、図のように一緒に結合された以下の要素を含む：エンコーダ１、送達サブシステム２、デコーダ３および後処理ユニット４。図示したシステムの変形においては、要素の一つまたは複数が省略され、あるいは追加的なオーディオ・データ処理ユニットが含められる。 FIG. 1 is a block diagram of an exemplary audio processing chain (audio data processing system), one or more of the elements of which may be configured in accordance with embodiments of the present invention. The system includes the following elements coupled together as shown: encoder 1, delivery subsystem 2, decoder 3 and post-processing unit 4. In variations of the illustrated system, one or more of the elements are omitted or additional audio data processing units are included.

いくつかの実装では、エンコーダ１（これは任意的には前処理ユニットを含む）は、入力としてオーディオ・コンテンツを含むPCM（時間領域）サンプルを受け入れ、該オーディオ・コンテンツを示すエンコードされたオーディオ・ビットストリーム（MPEG-4 AAC規格に準拠するフォーマットをもつ）を出力するよう構成されている。オーディオ・コンテンツを示すビットストリームのデータは本稿では時に「オーディオ・データ」または「エンコードされたオーディオ・データ」と称される。エンコーダが本発明の典型的な実施形態に従って構成される場合、エンコーダから出力されるオーディオ・ビットストリームは、オーディオ・データのほかにeSBRメタデータを（典型的には他のメタデータも）含む。 In some implementations, encoder 1 (which optionally includes a pre-processing unit) accepts as input PCM (time domain) samples containing audio content and encoded audio files representing the audio content. It is configured to output a bitstream (with a format compliant with the MPEG-4 AAC standard). Bitstream data representing audio content is sometimes referred to herein as "audio data" or "encoded audio data." When the encoder is configured according to the exemplary embodiment of the present invention, the audio bitstream output from the encoder contains eSBR metadata (and typically other metadata as well) in addition to the audio data.

エンコーダ１から出力される一つまたは複数のエンコードされたオーディオ・ビットストリームは、エンコードされたオーディオ送達サブシステム２に呈されてもよい。サブシステム２は、エンコーダ１から出力されたそれぞれのエンコードされたビットストリームを記憶および／または送達するよう構成される。エンコーダ１から出力されたエンコードされたオーディオ・ビットストリームはサブシステム２によって（たとえばDVDまたはブルーレイディスクの形で）記憶されてもよく、あるいはサブシステム２（これは伝送リンクまたはネットワークを実装してもよい）によって伝送されてもよく、あるいはサブシステム２によって記憶されかつ伝送されてもよい。 One or more encoded audio bitstreams output from encoder 1 may be presented to encoded audio delivery subsystem 2 . Subsystem 2 is configured to store and/or deliver each encoded bitstream output from encoder 1 . The encoded audio bitstream output from encoder 1 may be stored by subsystem 2 (eg in the form of a DVD or Blu-ray disc) or may be stored by subsystem 2 (which may implement a transmission link or network). ), or may be stored and transmitted by subsystem 2.

デコーダ３は、サブシステム２を介して受け取る（エンコーダ１によって生成された）エンコードされたMPEG-4 AACオーディオ・ビットストリームをデコードするよう構成される。いくつかの実施形態では、デコーダ３は、ビットストリームの各ブロックからeSBRメタデータを抽出し、ビットストリームをデコードして（抽出されたeSBRメタデータを使ってeSBR処理を実行することによることを含む）、デコードされたオーディオ・データ（たとえば、デコードされたPCMオーディオ・サンプルのストリーム）を生成するよう構成される。いくつかの実施形態では、デコーダ３は、ビットストリームからSBRメタデータを抽出し（だがビットストリームに含まれるeSBRメタデータは無視し）、ビットストリームをデコードして（抽出されたSBRメタデータを使ってSBR処理を実行することによることを含む）、デコードされたオーディオ・データ（たとえば、デコードされたPCMオーディオ・サンプルのストリーム）を生成するよう構成される。典型的には、デコーダ３は、サブシステム２から受領されたエンコードされたオーディオ・ビットストリームの諸セグメントを（たとえば非一時的な仕方で）記憶するバッファを含む。 Decoder 3 is configured to decode the encoded MPEG-4 AAC audio bitstream (produced by encoder 1) received via subsystem 2; In some embodiments, decoder 3 extracts eSBR metadata from each block of the bitstream, decodes the bitstream (including by performing eSBR processing using the extracted eSBR metadata ), configured to generate decoded audio data (eg, a stream of decoded PCM audio samples). In some embodiments, decoder 3 extracts SBR metadata from the bitstream (but ignores eSBR metadata included in the bitstream) and decodes the bitstream (using the extracted SBR metadata). (including by performing SBR processing on the device) to generate decoded audio data (eg, a stream of decoded PCM audio samples). Decoder 3 typically includes a buffer that stores (eg, in a non-transitory manner) segments of the encoded audio bitstream received from subsystem 2 .

図１の後処理ユニット４は、デコーダ３からのデコードされたオーディオ・データ（たとえばデコードされたPCMオーディオ・サンプル）のストリームを受け入れ、それに対して後処理を実行するよう構成される。後処理ユニットは、後処理されたオーディオ・コンテンツ（またはデコーダ３から受領されたデコードされたオーディオ）を一つまたは複数のスピーカーによる再生のためにレンダリングするよう構成されてもよい。 Post-processing unit 4 of FIG. 1 is configured to accept a stream of decoded audio data (eg, decoded PCM audio samples) from decoder 3 and perform post-processing thereon. The post-processing unit may be configured to render post-processed audio content (or decoded audio received from decoder 3) for playback by one or more speakers.

図２は、本発明のオーディオ処理ユニットのある実施形態であるエンコーダ（１００）のブロック図である。エンコーダ１００のコンポーネントまたは要素のいずれも、一つまたは複数のプロセスおよび／または一つまたは複数の回路（たとえばASIC、FPGAまたは他の集積回路）として、ハードウェア、ソフトウェアまたはハードウェアとソフトウェアの組み合わせにおいて、実装されてもよい。エンコーダ１００は、図のように接続された、エンコーダ１０５、詰め込み器（stuffer）／フォーマッタ段１０７、メタデータ生成段１０６およびバッファ・メモリ１０９を有する。典型的には、エンコーダ１００は、他の処理要素（図示せず）をも含む。エンコーダ１００は、入力オーディオ・ビットストリームを、エンコードされた出力MPEG-4 AACビットストリームに変換するよう構成される。 FIG. 2 is a block diagram of an encoder (100), one embodiment of the audio processing unit of the present invention. Any of the components or elements of encoder 100 may be implemented in hardware, software, or a combination of hardware and software, as one or more processes and/or one or more circuits (e.g., ASICs, FPGAs, or other integrated circuits). , may be implemented. Encoder 100 comprises encoder 105, stuffer/formatter stage 107, metadata generation stage 106 and buffer memory 109, connected as shown. Typically, encoder 100 also includes other processing elements (not shown). Encoder 100 is configured to convert an input audio bitstream into an encoded output MPEG-4 AAC bitstream.

メタデータ生成器１０６は、エンコーダ１００から出力されるべきエンコードされたビットストリームに段１０７によって含められるべきメタデータ（eSBRメタデータおよびSBRメタデータを含む）を生成する（および／または段１０７に素通しにする）よう結合され、構成される。 Metadata generator 106 generates metadata (including eSBR metadata and SBR metadata) to be included by stage 107 in the encoded bitstream to be output from encoder 100 (and/or passes through stage 107). combined and configured in such a way that

エンコーダ１０５は、入力オーディオ・データを（たとえばそれに対して圧縮を実行することにより）エンコードし、結果として得られるエンコードされたオーディオを、段１０７から出力されるべきエンコードされたビットストリームに含めるために、段１０７に呈するよう結合され、構成される。 Encoder 105 encodes the input audio data (eg, by performing compression on it) and the resulting encoded audio for inclusion in the encoded bitstream to be output from stage 107. , are coupled and configured to present on stage 107 .

段１０７は、エンコーダ１０５からのエンコードされたオーディオおよび生成器１０６からのメタデータ（eSBRメタデータおよびSBRメタデータを含む）を多重化して、段１０７から出力されるべきエンコードされたビットストリームを生成するよう構成される。好ましくは、エンコードされたビットストリームが本発明の実施形態の一つによって規定されるフォーマットをもつようにする。 Stage 107 multiplexes the encoded audio from encoder 105 and the metadata from generator 106 (including eSBR metadata and SBR metadata) to produce an encoded bitstream to be output from stage 107. configured to Preferably, the encoded bitstream has a format defined by one of the embodiments of the invention.

バッファ・メモリ１０９は、段１０７から出力されたエンコードされたオーディオ・ビットストリームの少なくとも一つのブロックを（たとえば非一時的な仕方で）記憶するよう構成される。その後、エンコードされたオーディオ・ビットストリームのブロックのシーケンスがバッファ・メモリ１０９から、エンコーダ１００からの出力として、送達システムに呈される。 Buffer memory 109 is configured to store (eg, in a non-transitory manner) at least one block of the encoded audio bitstream output from stage 107 . The sequence of blocks of encoded audio bitstream is then presented from buffer memory 109 as output from encoder 100 to the delivery system.

図３は、本発明のオーディオ処理ユニットの実施形態であるデコーダ（２００）を含み、任意的にはそれに結合された後処理器（３００）をも含むシステムのブロック図である。デコーダ２００のコンポーネントまたは要素のいずれも、一つまたは複数のプロセスおよび／または一つまたは複数の回路（たとえばASIC、FPGAまたは他の集積回路）として、ハードウェア、ソフトウェアまたはハードウェアとソフトウェアの組み合わせにおいて、実装されてもよい。デコーダ２００は、図のように接続された、バッファ・メモリ２０１、ビットストリーム・ペイロード・フォーマット解除器（パーサー）２０５、オーディオ・デコード・サブシステム２０２（時に「コア」デコード段または「コア」デコード・サブシステムと称される）、eSBR処理段２０３および制御ビット生成段２０４を有する。典型的には、デコーダ２００は、他の処理要素（図示せず）をも含む。 FIG. 3 is a block diagram of a system including a decoder (200), which is an embodiment of the audio processing unit of the invention, and optionally also a post-processor (300) coupled thereto. Any of the components or elements of decoder 200 may be implemented in hardware, software, or a combination of hardware and software, as one or more processes and/or one or more circuits (e.g., ASICs, FPGAs, or other integrated circuits). , may be implemented. Decoder 200 includes a buffer memory 201, a bitstream payload deformatter (parser) 205, an audio decoding subsystem 202 (sometimes a "core" decoding stage or "core" decoding stage), connected as shown. subsystem), an eSBR processing stage 203 and a control bit generation stage 204 . Typically, decoder 200 also includes other processing elements (not shown).

バッファ・メモリ（バッファ）２０１は、デコーダ２００によって受領されるエンコードされたMPEG-4 AACオーディオ・ビットストリームの少なくとも一つのブロックを（たとえば非一時的な仕方で）記憶する。デコーダ２００の動作において、ビットストリームのブロックのシーケンスがバッファ２０１からフォーマット解除器２０５に呈される。 A buffer memory (buffer) 201 stores (eg, in a non-transitory manner) at least one block of an encoded MPEG-4 AAC audio bitstream received by decoder 200 . In operation of decoder 200 , a sequence of blocks of the bitstream are presented from buffer 201 to deformatter 205 .

図３実施形態の変形（またはのちに述べる図４の実施形態）では、デコーダではないAPU（たとえば図６のAPU ５００）が、図３または図４のバッファ２０１によって受領されるのと同じ型のエンコードされたオーディオ・ビットストリーム（たとえばMPEG-4 AACオーディオ・ビットストリーム）（すなわち、eSBRメタデータを含むエンコードされたオーディオ・ビットストリーム）の少なくとも一つのブロックを（たとえば非一時的な仕方で）記憶するバッファ・メモリ（たとえばバッファ２０１と同一のバッファ・メモリ）を含む。 In a variation of the FIG. 3 embodiment (or the FIG. 4 embodiment discussed below), a non-decoder APU (eg, APU 500 of FIG. 6) uses the same type of data received by buffer 201 of FIG. 3 or FIG. Store (e.g., in a non-transitory manner) at least one block of an encoded audio bitstream (e.g., an MPEG-4 AAC audio bitstream) (i.e., an encoded audio bitstream containing eSBR metadata) buffer memory (eg, the same buffer memory as buffer 201).

再び図３を参照するに、フォーマット解除器２０５は、ビットストリームの各ブロックを多重分離して、それからSBRメタデータ（量子化された包絡データを含む）およびeSBRメタデータを（典型的には他のメタデータも）抽出し、少なくとも前記eSBRメタデータおよび前記SBRメタデータをeSBR処理段２０３に呈するとともに、典型的にはさらに他の抽出されたメタデータをデコード・サブシステム２０２に（任意的には制御ビット生成器２０４にも）呈するよう結合され、構成される。フォーマット解除器２０５は、ビットストリームの各ブロックからオーディオ・データを抽出し、抽出されたオーディオ・データをデコード・サブシステム（デコード段）２０２に呈するようにも結合され、構成される。 Referring again to FIG. 3, the deformatter 205 demultiplexes each block of the bitstream and then removes SBR metadata (including quantized envelope data) and eSBR metadata (typically other metadata) and presents at least said eSBR metadata and said SBR metadata to eSBR processing stage 203, and typically further other extracted metadata to decoding subsystem 202 (optionally is also coupled to and configured to provide control bit generator 204). Deformatter 205 is also coupled and configured to extract audio data from each block of the bitstream and present the extracted audio data to decoding subsystem (decoding stage) 202 .

図３のシステムは任意的には、後処理器３００をも含む。後処理器３００はバッファ・メモリ（バッファ）３０１と、バッファ３０１に結合された少なくとも一つの処理要素を含む他の処理要素（図示せず）とを含む。バッファ３０１は、デコーダ２００から後処理器３００によって受領されたデコードされたオーディオ・データの少なくとも一つのブロック（またはフレーム）を（たとえば非一時的な仕方で）記憶する。後処理器３００の処理要素は、バッファ３０１から出力されたデコードされたオーディオのブロック（またはフレーム）のシーケンスを受領し、デコード・サブシステム２０２（および／またはフォーマット解除器２０５）から出力されたメタデータおよび／またはデコーダ２００の段２０４から出力された制御ビットを使って適応的に処理するよう結合され、構成される。 The system of FIG. 3 optionally also includes post-processor 300 . Post-processor 300 includes a buffer memory (buffer) 301 and other processing elements (not shown) including at least one processing element coupled to buffer 301 . Buffer 301 stores (eg, in a non-transitory manner) at least one block (or frame) of decoded audio data received by post-processor 300 from decoder 200 . The processing elements of post-processor 300 receive the sequence of blocks (or frames) of decoded audio output from buffer 301, and the metadata output from decode subsystem 202 (and/or deformatter 205). coupled and configured to adaptively process using data and/or control bits output from stage 204 of decoder 200;

デコーダ２００のオーディオ・デコード・サブシステム２０２は、パーサー２０５によって抽出されたオーディオ・データをデコードして（そのようなデコードは「コア」デコード動作と称されてもよい）、デコードされたオーディオ・データを生成し、デコードされたオーディオ・データをeSBR処理段２０３に呈するよう構成される。デコードは周波数領域で実行され、典型的には逆量子化とそれに続くスペクトル処理（spectral processing）を含む。典型的には、サブシステム２０２における処理の最終段が、デコードされた周波数領域オーディオ・データに周波数領域から時間領域への変換を適用し、そのためサブシステムの出力は時間領域のデコードされたオーディオ・データである。段２０３は、（パーサー２０５によって抽出された）SBRメタデータおよびeSBRメタデータによって示されるSBRツールおよびeSBRツールを、デコードされたオーディオ・データに適用して（すなわち、SBRおよびeSBRメタデータを使ってデコード・サブシステム２０２の出力に対してSBRおよびeSBR処理を実行して）、デコーダ２００から（たとえば後処理器３００に）出力される完全にデコードされたオーディオ・データを生成するよう構成される。典型的には、デコーダ２００は、フォーマット解除器２０５から出力されるフォーマット解除されたオーディオ・データおよびメタデータを記憶するメモリ（サブシステム２０２および段２０３によってアクセス可能）を含み、段２０３はSBRおよびeSBR処理の間に必要に応じてオーディオ・データおよびメタデータ（SBRメタデータおよびeSBRメタデータを含む）にアクセスするよう構成される。段２０３におけるSBR処理およびeSBR処理は、コア・デコード・サブシステム２０２の出力に対する後処理であると考えられてもよい。任意的に、デコーダ２００は、最終的なアップミックス・サブシステム（これは、フォーマット解除器２０５によって抽出されたPSメタデータおよび／またはサブシステム２０４において生成された制御ビットを使って、MPEG-4 AAC規格において定義されているパラメトリック・ステレオ（「PS」）ツールを適用しうる）をも含む。アップミックス・サブシステムは、段２０３の出力に対してアップミックスを実行して、デコーダ２００から出力される、完全にデコードされた、アップミックスされたオーディオを生成するよう結合され、構成される。あるいはまた、後処理器３００が（たとえばフォーマット解除器２０５によって抽出されたPSメタデータおよび／またはサブシステム２０４において生成された制御ビットを使って）デコーダ２００の出力に対してアップミックスを実行するよう構成される。 Audio decoding subsystem 202 of decoder 200 decodes the audio data extracted by parser 205 (such decoding may be referred to as a "core" decoding operation) to produce decoded audio data and present the decoded audio data to the eSBR processing stage 203 . Decoding is performed in the frequency domain and typically includes inverse quantization followed by spectral processing. Typically, the final stage of processing in subsystem 202 applies a frequency-domain to time-domain transform to the decoded frequency-domain audio data, so that the output of the subsystem is time-domain decoded audio data. Data. Stage 203 applies the SBR and eSBR tools indicated by the SBR and eSBR metadata (extracted by parser 205) to the decoded audio data (i.e., using the SBR and eSBR metadata It is configured to perform SBR and eSBR processing on the output of decoding subsystem 202) to produce fully decoded audio data that is output from decoder 200 (eg, to post-processor 300). Typically, decoder 200 includes a memory (accessible by subsystem 202 and stage 203) that stores the unformatted audio data and metadata output from deformatter 205, stage 203 being SBR and It is configured to access audio data and metadata (including SBR metadata and eSBR metadata) as needed during eSBR processing. The SBR and eSBR processing in stage 203 may be considered post-processing on the output of core decode subsystem 202 . Optionally, the decoder 200 uses the final upmix subsystem (which uses the PS metadata extracted by the deformatter 205 and/or the control bits generated in the subsystem 204 to create an MPEG-4 Parametric Stereo ("PS") tool defined in the AAC standard may be applied). An upmix subsystem is coupled and configured to perform an upmix on the output of stage 203 to produce fully decoded, upmixed audio output from decoder 200 . Alternatively, post-processor 300 performs an upmix on the output of decoder 200 (eg, using PS metadata extracted by deformatter 205 and/or control bits generated in subsystem 204). Configured.

フォーマット解除器２０５によって抽出されたメタデータに応答して、制御ビット生成器２０４は制御データを生成してもよい。制御データは、デコーダ２００内で（たとえば最終的なアップミックス・サブシステムにおいて）使われてもよく、および／またはデコーダ２００の出力として（たとえば後処理で使うために後処理器３００に）呈されてもよい。入力ビットストリームから抽出されたメタデータに応答して（任意的には制御データにも応答して）、段２０４は、eSBR処理段２０３から出力されたデコードされたオーディオ・データが特定の型の後処理を受けるべきであることを示す制御ビットを生成し（後処理器３００に呈し）てもよい。いくつかの実装では、デコーダ２００は、入力ビットストリームからフォーマット解除器２０５によって抽出されたメタデータを後処理器３００に呈するよう構成され、後処理器３００は、デコーダ２００から出力されたデコードされたオーディオ・データに対して、前記メタデータを使って後処理を実行するよう構成される。 In response to the metadata extracted by deformatter 205, control bits generator 204 may generate control data. The control data may be used within decoder 200 (eg, in the final upmix subsystem) and/or presented as an output of decoder 200 (eg, to post-processor 300 for use in post-processing). may In response to metadata extracted from the input bitstream (and optionally in response to control data), stage 204 determines whether the decoded audio data output from eSBR processing stage 203 is of a particular type. A control bit may be generated (presented to post-processor 300) indicating that it should undergo post-processing. In some implementations, the decoder 200 is configured to present the metadata extracted by the deformatter 205 from the input bitstream to the post-processor 300, which applies the decoded metadata output from the decoder 200. It is configured to perform post-processing on the audio data using said metadata.

図４は、本発明のオーディオ処理ユニットのもう一つの実施形態であるオーディオ処理ユニット（「APU」）（２１０）のブロック図である。APU ２１０は、eSBR処理を実行するよう構成されていないレガシー・デコーダである。APU ２１０のコンポーネントまたは要素のいずれも、一つまたは複数のプロセスおよび／または一つまたは複数の回路（たとえばASIC、FPGAまたは他の集積回路）として、ハードウェア、ソフトウェアまたはハードウェアとソフトウェアの組み合わせにおいて、実装されてもよい。APU ２１０は、図のように接続された、バッファ・メモリ２０１、ビットストリーム・ペイロード・フォーマット解除器（パーサー）２１５、オーディオ・デコード・サブシステム２０２（時に「コア」デコード段または「コア」デコード・サブシステムと称される）およびSBR処理段２１３を有する。典型的には、APU ２１０は、他の処理要素（図示せず）をも含む。 FIG. 4 is a block diagram of an audio processing unit (“APU”) (210), which is another embodiment of the audio processing unit of the present invention. APU 210 is a legacy decoder that is not configured to perform eSBR processing. Any of the components or elements of APU 210 may be implemented in hardware, software, or a combination of hardware and software, as one or more processes and/or one or more circuits (e.g., ASICs, FPGAs, or other integrated circuits). , may be implemented. APU 210 includes a buffer memory 201, a bitstream payload de-formatter (parser) 215, an audio decoding subsystem 202 (sometimes a "core" decoding stage or "core" decoding stage), connected as shown. subsystem) and SBR processing stage 213 . Typically, APU 210 also includes other processing elements (not shown).

APU ２１０の要素２０１および２０２は、（図３の）デコーダ２００の同じ番号を付された要素と同一であり、それらについての上記の記述は繰り返さない。APU ２１０の動作においては、APU ２１０によって受領されるエンコードされたオーディオ・ビットストリーム（MPEG-4 AACビットストリーム）のブロックのシーケンスはバッファ２０１からフォーマット解除器２１５に呈される。 Elements 201 and 202 of APU 210 are identical to the same numbered elements of decoder 200 (of FIG. 3) and the above description thereof will not be repeated. In operation of APU 210 , a sequence of blocks of encoded audio bitstream (MPEG-4 AAC bitstream) received by APU 210 is presented from buffer 201 to deformatter 215 .

フォーマット解除器２１５は、ビットストリームの各ブロックを多重分離して、それからSBRメタデータ（量子化された包絡データを含む）を、典型的には他のメタデータも抽出するが、本発明の任意の実施形態によりビットストリームに含まれることがありうるeSBRは無視するよう結合され、構成される。フォーマット解除器２１５は、少なくとも前記SBRメタデータをSBR処理段２１３に呈するよう構成される。フォーマット解除器２１５は、ビットストリームの各ブロックからオーディオ・データを抽出し、抽出されたオーディオ・データをデコード・サブシステム（デコード段）２０２に呈するようにも結合され、構成される。 De-formatter 215 de-multiplexes each block of the bitstream and extracts SBR metadata (including quantized envelope data) from it, and typically other metadata as well, but the optional are combined and configured to ignore any eSBR that may be included in the bitstream according to the embodiment of . The deformatter 215 is configured to present at least said SBR metadata to the SBR processing stage 213 . Deformatter 215 is also coupled and configured to extract audio data from each block of the bitstream and present the extracted audio data to decoding subsystem (decoding stage) 202 .

デコーダ２００のオーディオ・デコード・サブシステム２０２は、フォーマット解除器２１５によって抽出されたオーディオ・データをデコードして（そのようなデコードは「コア」デコード動作と称されてもよい）、デコードされたオーディオ・データを生成し、デコードされたオーディオ・データをSBR処理段２１３に呈するよう構成される。デコードは周波数領域で実行される。典型的には、サブシステム２０２における処理の最終段が、デコードされた周波数領域オーディオ・データに周波数領域から時間領域への変換を適用し、そのためサブシステムの出力は時間領域のデコードされたオーディオ・データである。段２１３は、（フォーマット解除器２１５によって抽出された）SBRメタデータによって示されるSBRツールをデコードされたオーディオ・データに適用して（だがeSBRツールは適用しない）（すなわち、SBRメタデータを使ってデコード・サブシステム２０２の出力に対してSBR処理を実行して）、APU ２１０から（たとえば後処理器３００に）出力される完全にデコードされたオーディオ・データを生成するよう構成される。典型的には、APU ２１０は、フォーマット解除器２１５から出力されるフォーマット解除されたオーディオ・データおよびメタデータを記憶するメモリ（サブシステム２０２および段２１３によってアクセス可能）を含み、段２１３はSBR処理の間に必要に応じてオーディオ・データおよびメタデータ（SBRメタデータを含む）にアクセスするよう構成される。段２１３におけるSBR処理は、コア・デコード・サブシステム２０２の出力に対する後処理であると考えられてもよい。任意的に、APU ２１０は、最終的なアップミックス・サブシステム（これは、フォーマット解除器２１５によって抽出されたPSメタデータを使って、MPEG-4 AAC規格において定義されているパラメトリック・ステレオ（「PS」）ツールを適用しうる）をも含む。アップミックス・サブシステムは、段２１３の出力に対してアップミックスを実行して、APU ２１０から出力される、完全にデコードされた、アップミックスされたオーディオを生成するよう結合され、構成される。あるいはまた、後処理器が（たとえばフォーマット解除器２１５によって抽出されたPSメタデータおよび／またはAPU ２１０において生成された制御ビットを使って）APU ２１０の出力に対してアップミックスを実行するよう構成される。 Audio decoding subsystem 202 of decoder 200 decodes the audio data extracted by deformatter 215 (such decoding may be referred to as a "core" decoding operation) to produce the decoded audio • configured to generate data and present the decoded audio data to the SBR processing stage 213; Decoding is performed in the frequency domain. Typically, the final stage of processing in subsystem 202 applies a frequency-domain to time-domain transform to the decoded frequency-domain audio data, so that the output of the subsystem is time-domain decoded audio data. Data. Stage 213 applies the SBR tools (but no eSBR tools) indicated by the SBR metadata (extracted by deformatter 215) to the decoded audio data (i.e., using the SBR metadata It is configured to perform SBR processing on the output of the decoding subsystem 202) to produce fully decoded audio data that is output from the APU 210 (eg, to the post-processor 300). Typically, APU 210 includes memory (accessible by subsystem 202 and stage 213) that stores the unformatted audio data and metadata output from deformatter 215, which stage 213 performs SBR processing. configured to access audio data and metadata (including SBR metadata) as needed during The SBR processing in stage 213 may be considered post-processing on the output of core decode subsystem 202 . Optionally, the APU 210 performs a final upmix subsystem (which uses the PS metadata extracted by the deformatter 215 to convert parametric stereo (" PS”) tool may be applied). An upmix subsystem is coupled and configured to perform an upmix on the output of stage 213 to produce fully decoded, upmixed audio output from APU 210 . Alternatively, the post-processor is configured to perform upmixing on the output of APU 210 (eg, using PS metadata extracted by deformatter 215 and/or control bits generated at APU 210). be.

エンコーダ１００、デコーダ２００およびAPU ２１０のさまざまな実装が、本発明の方法の異なる実施形態を実行するよう構成される。 Various implementations of encoder 100, decoder 200 and APU 210 are configured to perform different embodiments of the method of the present invention.

いくつかの実施形態によれば、（eSBRメタデータをパースしたりeSBRメタデータが関係する何らかのeSBRツールを使ったりするよう構成されていない）レガシー・デコーダがeSBRメタデータを無視するが、それでもビットストリームをeSBRメタデータやeSBRメタデータが関係する何らかのeSBRツールを使うことなく、典型的にはデコードされたオーディオ品質におけるいかなる有意なペナルティもなしに可能な限りデコードできるように、eSBRメタデータが（たとえば、eSBRメタデータである少数の制御ビットが）エンコードされたオーディオ・ビットストリーム（たとえばMPEG-4 AACビットストリーム）に含められる。しかしながら、ビットストリームをパースしてeSBRメタデータを識別し、該eSBRメタデータに応答して少なくとも一つのeSBRツールを使うよう構成されたeSBRデコーダは、少なくとも一つのそのようなeSBRツールを使うことの恩恵を享受する。したがって、本発明の実施形態は、向上されたスペクトル帯域複製（eSBR）制御データまたはメタデータを、後方互換な仕方で効率的に伝送する手段を提供する。 According to some embodiments, legacy decoders (not configured to parse eSBR metadata or use any eSBR tools that involve eSBR metadata) ignore eSBR metadata, but still If the eSBR metadata is ( For example, a few control bits that are eSBR metadata) are included in the encoded audio bitstream (eg MPEG-4 AAC bitstream). However, an eSBR decoder that is configured to parse a bitstream to identify eSBR metadata and use at least one eSBR tool in response to the eSBR metadata may be unable to use at least one such eSBR tool. enjoy the benefits. Accordingly, embodiments of the present invention provide a means of efficiently transmitting enhanced spectral band replication (eSBR) control data or metadata in a backward compatible manner.

典型的には、ビットストリーム中のeSBRメタデータは、（MPEG USAC規格において記述されており、ビットストリームの生成の際にエンコーダによって適用されていてもいなくてもよい）次のeSBRツールのうちの一つまたは複数を示す（たとえば、次のeSBRツールのうちの一つまたは複数の、少なくとも一つの特性またはパラメータを示す）：
・高調波転換；
・QMFパッチング追加的前処理（前置平坦化（pre-flattening））；および
・サブバンド・サンプル間時間包絡整形（Temporal Envelope Shaping）または「インターTES」。
たとえば、ビットストリームに含まれるeSBRメタデータは、（MPEG USAC規格および本開示において記述される）パラメータ：harmonicSBR[ch]、sbrPatchingMode[ch]、sbrOversamplingFlag[ch]、sbrPitchInBins[ch]、sbrPitchInBins[ch]、bs_interTes、bs_temp_shape[ch][env]、bs_inter_temp_shape_mode[ch][env]およびbs_sbr_preprocessingの値を示してもよい。 Typically, eSBR metadata in a bitstream is generated using one of the following eSBR tools (described in the MPEG USAC standard, which may or may not have been applied by the encoder during bitstream generation): Indicate one or more (e.g., indicate at least one characteristic or parameter of one or more of the following eSBR tools):
- harmonic conversion;
- QMF patching additional pre-processing (pre-flattening); and - sub-band inter-sample Temporal Envelope Shaping or "inter-TES".
For example, the eSBR metadata included in the bitstream can be represented by parameters (described in the MPEG USAC standard and this disclosure): harmonicSBR[ch], sbrPatchingMode[ch], sbrOversamplingFlag[ch], sbrPitchInBins[ch], sbrPitchInBins[ch] , bs_interTes, bs_temp_shape[ch][env], bs_inter_temp_shape_mode[ch][env] and bs_sbr_preprocessing.

ここで、Xが何らかのパラメータであるとして記法X[ch]は、そのパラメータがデコードされるべきエンコードされたビットストリームのオーディオ・コンテンツのあるチャネル（「ch」）に関することを表わす。簡単のため、時に表現[ch]を略し、関連するパラメータがオーディオ・コンテンツのあるチャネルに関することを前提とする。 Here, the notation X[ch], where X is some parameter, denotes that the parameter relates to a certain channel ("ch") of the audio content of the encoded bitstream to be decoded. For simplicity, we sometimes omit the expression [ch] and assume that the relevant parameters relate to a certain channel of audio content.

ここで、Xが何らかのパラメータであるとして記法X[ch][env]は、そのパラメータがデコードされるべきエンコードされたビットストリームのオーディオ・コンテンツのあるチャネル（「ch」）のSBR包絡（「env」）に関することを表わす。簡単のため、時に表現[env]および[ch]を略し、関連するパラメータがオーディオ・コンテンツのあるチャネルのSBR包絡に関することを前提とする。 where the notation X[ch][env] is the SBR envelope ("env ”). For simplicity, we sometimes omit the expressions [env] and [ch] and assume that the relevant parameters relate to the SBR envelope of a certain channel of audio content.

前記したように、MPEG USACは、USACビットストリームが、デコーダによるeSBR処理の実行を制御するeSBRメタデータを含むことを考えている。eSBRメタデータは、以下の一ビットのメタデータ・パラメータを含む：harmonicSBR；bs_interTES；およびbs_pvc。 As mentioned above, MPEG USAC contemplates that the USAC bitstream contains eSBR metadata that controls the execution of eSBR processing by decoders. eSBR metadata includes the following one-bit metadata parameters: harmonicSBR; bs_interTES; and bs_pvc.

パラメータharmonicSBRは、SBRについての高調波パッチング（harmonic patching）（高調波転換（harmonic transposition））の使用を示す。具体的には、harmonicSBR＝0は、MPEG-4 AAC規格の4.6.18.6.3節に記載される非高調波（non-harmonic）スペクトル・パッチングを示し；harmonicSBR＝1は、（MPEG USAC規格の7.5.3または7.5.4節に記載される、eSBRにおいて使われる型の）高調波SBRパッチングを示す。高調波SBRパッチングは、非eSBRスペクトル帯域複製（すなわち、eSBRでないSBR）によれば使われない。本開示を通じて、スペクトル帯域複製の基本形としてはスペクトル・パッチング（spectral patching）といい、スペクトル帯域複製の向上された形としては高調波転換（harmonic transposition）という。 The parameter harmonicSBR indicates the use of harmonic patching (harmonic transposition) for SBR. Specifically, harmonicSBR=0 indicates non-harmonic spectral patching as described in section 4.6.18.6.3 of the MPEG-4 AAC standard; Fig. 7 shows harmonic SBR patching (of the type used in eSBR, described in Section 7.5.3 or 7.5.4). Harmonic SBR patching is not used with non-eSBR spectral band replication (ie, SBR that is not eSBR). Throughout this disclosure, the basic form of spectral band replication is referred to as spectral patching, and the enhanced form of spectral band replication is referred to as harmonic transposition.

パラメータbs_interTESの値は、eSBRのインターTESツールの使用を示す。 The value of the parameter bs_interTES indicates the use of eSBR's interTES tool.

パラメータbs_pvcの値は、eSBRのPVCツールの使用を示す。 The value of parameter bs_pvc indicates the use of eSBR's PVC tool.

エンコードされたビットストリームのデコードの間、（ビットストリームによって示されるオーディオ・コンテンツの各チャネル「ch」についての）デコードのeSBR処理段の間の高調波転換の実行が、以下のeSBRメタデータ・パラメータによって制御される：sbrPatchingMode[ch]；sbrOversamplingFlag[ch]；sbrPitchInBinsFlag[ch]およびsbrPitchInBins[ch]。 During decoding of an encoded bitstream, the performance of harmonic conversion between the eSBR processing stages of decoding (for each channel "ch" of the audio content indicated by the bitstream) is determined by the following eSBR metadata parameters: sbrPatchingMode[ch]; sbrOversamplingFlag[ch]; sbrPitchInBinsFlag[ch] and sbrPitchInBins[ch].

sbrPatchingMode[ch]の値は、eSBRにおいて使われる転換器（transposer）の型を示す。sbrPatchingMode[ch]＝1はMPEG-4 AAC規格の4.6.18.6.3節に記載される非高調波パッチングを示し；sbrPatchingMode[ch]＝0は、MPEG USAC規格の7.5.3または7.5.4節に記載される高調波SBRパッチングを示す。 The value of sbrPatchingMode[ch] indicates the type of transposer used in eSBR. sbrPatchingMode[ch]=1 indicates non-harmonic patching as described in clause 4.6.18.6.3 of the MPEG-4 AAC standard; sbrPatchingMode[ch]=0 indicates clause 7.5.3 or 7.5.4 of the MPEG USAC standard. shows harmonic SBR patching as described in .

sbrOversamplingFlag[ch]の値は、MPEG USAC規格の7.5.3節に記載されるDFTベースの高調波SBRパッチングと組み合わせたeSBRにおける信号適応的な周波数領域オーバーサンプリングの使用を示す。このフラグは転換器において利用されるDFTのサイズを制御する。1はMPEG USAC規格の7.5.3.1節に記載されるように有効にされた信号適応的な周波数領域オーバーサンプリングを示し；0はMPEG USAC規格の7.5.3.1節に記載されるように無効にされた信号適応的な周波数領域オーバーサンプリングを示す。 The value of sbrOversamplingFlag[ch] indicates the use of signal adaptive frequency domain oversampling in eSBR combined with DFT-based harmonic SBR patching as described in section 7.5.3 of the MPEG USAC standard. This flag controls the size of the DFT used in the converter. 1 indicates signal adaptive frequency domain oversampling enabled as described in Section 7.5.3.1 of the MPEG USAC Standard; 0 indicates disabled as described in Section 7.5.3.1 of the MPEG USAC Standard. signal adaptive frequency domain oversampling.

sbrPitchInBinsFlag[ch]の値は、sbrPitchInBins[ch]パラメータの解釈を制御する。1はsbrPitchInBins[ch]における値が有効であり、0より大きいことを示し；0はsbrPitchInBins[ch]の値が0に設定されていることを示す。 The value of sbrPitchInBinsFlag[ch] controls the interpretation of the sbrPitchInBins[ch] parameter. 1 indicates that the value in sbrPitchInBins[ch] is valid and greater than 0; 0 indicates that the value of sbrPitchInBins[ch] is set to zero.

sbrPitchInBins[ch]の値は、SBR高調波転換器におけるクロス積の項の付加（addition）を制御する。値sbrPitchInBins[ch]は[0,127]の範囲内の整数値であり、コア符号化器のサンプリング周波数に対して作用する1536ラインのDFTについての周波数ビンにおいて測られる距離を表わす。 The value of sbrPitchInBins[ch] controls the addition of cross product terms in the SBR harmonic converter. The value sbrPitchInBins[ch] is an integer value in the range [0,127] and represents the distance measured in frequency bins for the 1536-line DFT operating on the sampling frequency of the core encoder.

MPEG-4 AACビットストリームが、（単一のSBRチャネルではなく）チャネルどうしが結合されていないSBRチャネル対を示す場合、該ビットストリームは（高調波または非高調波転換について）上記のシンタックスの二つのインスタンスを示す。sbr_channel_pair_element()の各チャネルについて一つのインスタンスである。 If the MPEG-4 AAC bitstream exhibits a pair of SBR channels whose channels are not combined (rather than a single SBR channel), then the bitstream is represented in the above syntax (for harmonic or non-harmonic conversion). Show two instances. One instance for each channel of sbr_channel_pair_element().

eSBRツールの高調波転換は典型的には、比較的低いクロスオーバー周波数におけるデコードされた音楽信号の品質を改善する。高調波転換はデコーダにおいてDFTベースまたはQMFベースの高調波転換によって実装されるべきである。非高調波転換（すなわち、レガシーのスペクトル・パッチングまたはコピー）は典型的には発話信号を改善する。よって、特定のオーディオ・コンテンツをエンコードするためにどの型の転換が好ましいかについての判断における出発点は、発話／音楽検出に依存して転換方法を選択することである。ここで、音楽コンテンツに対しては高調波転換が用いられ、発話コンテンツに対してはスペクトル・パッチングが用いられる。 Harmonic conversion in eSBR tools typically improves the quality of decoded music signals at relatively low crossover frequencies. Harmonic conversion should be implemented by DFT-based or QMF-based harmonic conversion in the decoder. Non-harmonic conversion (ie, legacy spectral patching or copying) typically improves speech signals. Thus, the starting point in determining which type of transformation is preferable for encoding a particular audio content is to choose the transformation method depending on the speech/music detection. Here, harmonic conversion is used for musical content and spectral patching is used for speech content.

eSBR処理の間の前置平坦化の実行は、bs_sbr_preprocessingとして知られる一ビットのeSBRメタデータ・パラメータの値によって制御される。それは、前置平坦化がこの単一のビットの値に依存して実行されるか、実行されないという意味においてである。MPEG-4 AAC規格の4.6.18.6.3節に記載されるSBR QMFパッチング・アルゴリズムが使われるとき、高周波数信号のスペクトル包絡の形における不連続がその後の包絡調整器（該包絡調整器は前記eSBR処理の別の段階を実行する）に入力されるのを避けようとして、前置平坦化の段階が実行されてもよい（bs_sbr_preprocessingパラメータによって示されるとき）。前置平坦化は典型的には、その後の包絡調整段の動作を改善し、結果として、知覚される高域信号がより安定することになる。 The performance of pre-flattening during eSBR processing is controlled by the value of a one-bit eSBR metadata parameter known as bs_sbr_preprocessing. That is in the sense that pre-flattening is performed or not depending on the value of this single bit. When the SBR QMF patching algorithm described in section 4.6.18.6.3 of the MPEG-4 AAC standard is used, a discontinuity in the shape of the spectral envelope of the high-frequency signal is detected by the subsequent envelope adjuster. A stage of pre-flattening may be performed (when indicated by the bs_sbr_preprocessing parameter) in an attempt to avoid being input to (performing another stage of eSBR processing). Pre-flattening typically improves the operation of subsequent envelope adjustment stages, resulting in a more stable perceived high-pass signal.

デコーダにおけるeSBR処理の間のサブバンド・サンプル間時間包絡整形（inter-subband sample Temporal Envelope Shaping）（「インターTES」ツール）の実行は、デコードされているUSACビットストリームのオーディオ・コンテンツの各チャネル（「ch」）の各SBR包絡（「env」）についての以下のeSBRメタデータ・パラメータによって制御される：bs_temp_shape[ch][env]およびbs_inter_temp_shape_mode[ch][env]。 Performing inter-subband sample Temporal Envelope Shaping (“Inter-TES” tool) during eSBR processing in the decoder is performed on each channel of the audio content of the decoded USAC bitstream ( 'ch')) controlled by the following eSBR metadata parameters for each SBR envelope ('env'): bs_temp_shape[ch][env] and bs_inter_temp_shape_mode[ch][env].

インターTESツールは、包絡調整器の後にQMFサブバンド・サンプルを処理する。この処理段階は、包絡調整器の時間的粒度より細かい時間的粒度をもって、より高い周波数帯域の時間的包絡を整形する。SBR包絡における各QMFサブバンド・サンプルに利得因子を適用することによって、インターTESは、諸QMFサブバンド・サンプルの間で時間的包絡を整形する。 The InterTES tool processes the QMF subband samples after the envelope adjuster. This processing stage shapes the temporal envelope of the higher frequency band with a finer temporal granularity than that of the envelope adjuster. Inter-TES shapes the temporal envelope between QMF subband samples by applying a gain factor to each QMF subband sample in the SBR envelope.

パラメータbs_temp_shape[ch][env]は、インターTESの使用を合図するフラグである。パラメータbs_inter_temp_shape_mode[ch][env]は、インターTESにおけるパラメータγの値を（MPEG USAC規格において定義されているように）示す。 Parameter bs_temp_shape[ch][env] is a flag that signals the use of inter-TES. Parameter bs_inter_temp_shape_mode[ch][env] indicates the value of parameter γ in inter TES (as defined in the MPEG USAC standard).

MPEG-4 AACビットストリームに上述したeSBRツール（高調波転換、前置平坦化およびインターTES）を示すeSBRメタデータを含めるための全体的なビットレート要求は、毎秒数百ビットのオーダーであると期待される。本発明のいくつかの実施形態によれば、eSBR処理を実行するために必要とされる差分の制御データが伝送されるだけだからである。この情報は（のちに説明するように）後方互換な仕方で含められるので、レガシー・デコーダはこの情報を無視できる。したがって、eSBRメタデータを含めることに関連するビットレートに対する悪影響は、次のことを含むいくつかの理由により、無視できる：
・（eSBRメタデータを含めることに起因する）ビットレート・ペナルティーは、eSBR処理を実行するために必要とされる差分の制御データだけが伝送される（SBR制御データのサイマルキャストではない）ので、全ビットレートの非常に小さな割合であること；
・SBRに関係した制御情報のチューニングは典型的には転換の詳細には依存しないこと；および
・（eSBR処理の間に用いられる）インターTESツールは、転換された信号のシングルエンドの後処理を実行すること。 The overall bitrate requirement for including eSBR metadata indicating the eSBR tools mentioned above (harmonic conversion, preflattening and inter-TES) in an MPEG-4 AAC bitstream is on the order of hundreds of bits per second. Be expected. This is because, according to some embodiments of the present invention, only the differential control data required to perform eSBR processing is transmitted. This information is included in a backwards compatible manner (as explained below) so that legacy decoders can ignore this information. Therefore, the negative impact on bitrate associated with including eSBR metadata is negligible for several reasons, including:
The bitrate penalty (due to the inclusion of eSBR metadata) is that only differential control data needed to perform eSBR processing is transmitted (no simulcast of SBR control data) be a very small percentage of the total bitrate;
- The tuning of SBR-related control information is typically independent of the details of the conversion; to run.

このように、本発明の諸実施形態は、向上されたスペクトル帯域複製（eSBR）制御データまたはメタデータを後方互換な仕方で効率的に伝送する手段を提供する。eSBR制御データのこの効率的な伝送は、ビットレートに対して明確な悪影響なしに、本発明の諸側面を用いるデコーダ、エンコーダおよびトランスコーダにおけるメモリ要求を軽減する。さらに、本発明の実施形態に従ってeSBRを実行することに関連する複雑さおよび処理要求も軽減される。SBRデータが処理される必要があるのは一度だけであり、eSBRが後方互換な仕方でMPEG-4 AACコーデックに統合されるのではなくMPEG-4 AACにおける完全に別個のオブジェクト型として扱われるとしたらそうであるようにサイマルキャストされる必要がないからである。 Thus, embodiments of the present invention provide a means of efficiently transmitting enhanced spectral band replication (eSBR) control data or metadata in a backward compatible manner. This efficient transmission of eSBR control data reduces memory requirements in decoders, encoders and transcoders employing aspects of the present invention, without any appreciable adverse impact on bitrate. Furthermore, the complexity and processing demands associated with performing eSBR according to embodiments of the present invention are also reduced. SBR data only needs to be processed once, and eSBR is treated as a completely separate object type in MPEG-4 AAC rather than being integrated into the MPEG-4 AAC codec in a backwards compatible manner. This is because it does not need to be simulcast as it would be if

次に、図７を参照して、本発明のいくつかの実施形態に従ってeSBRメタデータが含められるMPEG-4 AACビットストリームのブロック（raw_data_block）の要素を記述する。図７は、MPEG-4 AACビットストリームのブロック（raw_data_block）の図であり、そのセグメントのいくつかを示している。 Referring now to FIG. 7, the elements of an MPEG-4 AAC bitstream block (raw_data_block) in which eSBR metadata is included according to some embodiments of the present invention are described. FIG. 7 is a block diagram (raw_data_block) of an MPEG-4 AAC bitstream showing some of its segments.

MPEG-4 AACビットストリームのブロックは、オーディオ・プログラムについてのオーディオ・データを含む、少なくとも一つのsingle_channel_element()（たとえば図７に示される単一チャネル要素）および／または少なくとも一つのchannel_pair_element()（図７には特定的に示していないが、存在しうる）を含んでいてもよい。ブロックは、プログラムに関係したデータ（たとえばメタデータ）を含むいくつかのfill_element（たとえば図７の充填要素１および／または充填要素２）をも含んでいてもよい。各single_channel_element()は、単一チャネル要素の先頭を示す識別子（たとえば図７の「ID1」）を含み、マルチチャネル・オーディオ・プログラムの異なるチャネルを示すオーディオ・データを含むことができる。各channel_pair_elementはチャネル対要素の先頭を示す識別子（図７には示さず）を含み、プログラムの二つのチャネルを示すオーディオ・データを含むことができる。 A block of an MPEG-4 AAC bitstream contains at least one single_channel_element() (eg, the single channel element shown in FIG. 7) and/or at least one channel_pair_element() (see FIG. 7) containing audio data for an audio program. 7 may include (although not specifically shown, may be present). A block may also contain a number of fill_elements (eg, fill element 1 and/or fill element 2 in FIG. 7) that contain program-related data (eg, metadata). Each single_channel_element( ) contains an identifier (eg, “ID1” in FIG. 7) that indicates the beginning of a single channel element and can contain audio data that indicates different channels of a multi-channel audio program. Each channel_pair_element contains an identifier (not shown in FIG. 7) that indicates the beginning of the channel pair element and can contain audio data that indicates two channels of the program.

MPEG-4 AACビットストリームのfill_element（本稿では充填要素と称される）は、充填要素の先頭を示す識別子（たとえば図７の「ID2」）を含み、識別子の後に充填データを含む。識別子ID2は、0x6の値をもつ、三ビットの、最上位ビットが最初に伝送される符号なし整数（「uimsbf」）からなっていてもよい。充填データは、extension_payload()要素（本稿では時に拡張ペイロードと称される）を含むことができる。そのシンタックスはMPEG-4 AAC規格の表4.57に示されている。拡張ペイロードのいくつかの型が存在し、extension_typeパラメータを通じて識別される。このパラメータは、四ビットの、最上位ビットが最初に伝送される符号なし整数（「uimsbf」）である。 A fill_element (referred to herein as a fill element) of an MPEG-4 AAC bitstream includes an identifier (eg, "ID2" in FIG. 7) that indicates the beginning of the fill element, followed by fill data. The identifier ID2 may consist of a 3-bit unsigned integer ("uimsbf") with a value of 0x6, with the most significant bit transmitted first. Fill data can include an extension_payload() element (sometimes referred to herein as an extension payload). Its syntax is shown in Table 4.57 of the MPEG-4 AAC Standard. Several types of extension payloads exist, identified through the extension_type parameter. This parameter is a 4-bit unsigned integer ("uimsbf") with the most significant bit transmitted first.

充填データ（たとえばその拡張ペイロード）は、SBRオブジェクトを示す充填データのセグメントを示すヘッダまたは識別子（たとえば図７の「ヘッダ１」）を含むことができる（すなわち、ヘッダが、MPEG-4 AAC規格においてsbr_extension_data()と称される「SBRオブジェクト」型を初期化する）。たとえば、スペクトル帯域複製（SBR）拡張ペイロードは、ヘッダにおけるextension_typeフィールドについての値「1101」または「1110」をもって識別され、識別子「1101」はSBRデータを用いた拡張ペイロードを同定し、「1110」はSBRデータの正しさを検証するための巡回冗長検査（CRC）をもつSBRデータを用いた拡張ペイロードを同定する。 The filler data (eg, its extension payload) may include a header or identifier (eg, "header 1" in FIG. 7) that indicates a segment of the filler data that indicates an SBR object (i.e., if the header is specified in the MPEG-4 AAC standard as Initialize an "SBR Object" type called sbr_extension_data()). For example, a Spectrum Band Replication (SBR) extension payload is identified with a value of '1101' or '1110' for the extension_type field in the header, where the identifier '1101' identifies an extension payload using SBR data, and '1110' Identify an extension payload using SBR data with a cyclic redundancy check (CRC) to verify the correctness of the SBR data.

ヘッダが（たとえばextension_typeフィールドが）SBRオブジェクト型を初期化するとき、ヘッダにはSBRメタデータ（本稿では時に「スペクトル帯域複製データ」と称され、MPEG-4 AAC規格ではsbr_data()と称される）が後続し、該SBRメタデータには少なくとも一つのスペクトル帯域複製拡張要素（たとえば、図７の充填要素１の「SBR拡張要素」）が後続することができる。そのようなスペクトル帯域複製拡張要素（ビットストリームのセグメント）は、MPEG-4 AAC規格ではsbr_extension()コンテナと称される。スペクトル帯域複製拡張要素は任意的に、ヘッダ（たとえば、図７の充填要素１の「SBR拡張ヘッダ」）を含む。 When the header initializes the SBR object type (e.g. the extension_type field), the header contains SBR metadata (sometimes referred to in this document as "spectrum band replication data" and in the MPEG-4 AAC standard as sbr_data() ), and the SBR metadata may be followed by at least one spectral band duplication extension element (eg, “SBR extension element” in fill element 1 of FIG. 7). Such spectrum band replication extension elements (bitstream segments) are called sbr_extension( ) containers in the MPEG-4 AAC standard. The spectral band duplication extension element optionally includes a header (eg, "SBR Extension Header" in filler element 1 of FIG. 7).

MPEG-4 AAC規格は、スペクトル帯域複製拡張要素がプログラムのオーディオ・データのためのPS（パラメトリック・ステレオ）データを含むことができることを考えている。MPEG-4 AAC規格は、充填要素の（たとえばその拡張ペイロードの）ヘッダが（図７の「ヘッダ１」のように）SBRオブジェクト型を初期化し、充填要素のスペクトル帯域複製拡張要素がPSデータを含むとき、充填要素（たとえばその拡張ペイロード）がスペクトル帯域複製データbs_extension_idパラメータを含むことを考えている。このパラメータの値（すなわちbs_extension_id＝2）はPSデータが充填要素のスペクトル帯域複製拡張要素に含まれることを示す。 The MPEG-4 AAC standard allows for spectral band replication extension elements to contain PS (parametric stereo) data for program audio data. The MPEG-4 AAC standard specifies that the header of the filler element (e.g. its extension payload) initializes the SBR object type (as in "Header 1" in Figure 7) and the Spectral Band Duplication extension element of the filler element carries the PS data. When included, it is contemplated that the filler element (eg its extension payload) includes the spectrum band duplication data bs_extension_id parameter. A value of this parameter (ie bs_extension_id=2) indicates that the PS data is included in the spectral band replication extension element of the filler element.

本発明のいくつかの実施形態によれば、eSBRメタデータ（たとえば向上スペクトル帯域複製（eSBR）処理がそのブロックのオーディオ・コンテンツに対して実行されるかどうかを示すフラグ）が充填要素のスペクトル帯域複製拡張要素に含められる。たとえば、そのようなフラグは図７の充填要素１に含められ、フラグは充填要素１の「SBR拡張要素」のヘッダ（充填要素１の「SBR拡張ヘッダ」）の後に現われる。任意的に、そのようなフラグおよび追加的なeSBRメタデータがスペクトル帯域複製拡張要素において、スペクトル帯域複製拡張要素のヘッダの後に（たとえば図７における充填要素１のSBR拡張要素において、SBR拡張ヘッダ後に）含められる。本発明のいくつかの実施形態によれば、eSBRメタデータを含む充填要素はbs_extension_idパラメータをも含む。そのパラメータの値（たとえばbs_extension_id＝3）は、充填要素にeSBRメタデータが含まれ、当該ブロックのオーディオ・コンテンツに対してeSBR処理が実行されるべきであることを示す。 According to some embodiments of the present invention, eSBR metadata (e.g., a flag indicating whether enhanced spectral band replication (eSBR) processing is to be performed on the audio content of that block) is specified in the spectral band of the filler element. Included in the replication extension element. For example, such a flag is included in filler element 1 of FIG. 7, where the flag appears after the "SBR Extension Header" of filler element 1 ("SBR Extension Header" of filler element 1). Optionally, such flags and additional eSBR metadata are provided in the spectral band duplication extension element after the header of the spectral band duplication extension element (e.g., in the SBR extension element of fill element 1 in FIG. 7, after the SBR extension header ) are included. According to some embodiments of the invention, a filler element containing eSBR metadata also contains a bs_extension_id parameter. The value of that parameter (eg bs_extension_id=3) indicates that the filler element contains eSBR metadata and eSBR processing should be performed on the audio content of that block.

本発明のいくつかの実施形態によれば、eSBRメタデータは、充填要素のスペクトル帯域複製拡張要素（SBR拡張要素）以外のMPEG-4 AACビットストリームの充填要素（たとえば図７の充填要素２）に含められる。これは、SBRデータまたはCRCをもつSBRデータをもつextension_payload()を含む充填要素は、他のいかなる拡張型の他のいかなる拡張ペイロードをも含まないからである。したがって、eSBRメタデータが自分自身の拡張ペイロードに記憶される実施形態では、eSBRメタデータを記憶するために別個の充填要素が使われる。そのような充填要素は、充填要素の先頭を示す識別子（たとえば図７の「ID2」）を含み、該識別子の後に充填データを含む。充填データは、extension_payload()要素（本稿では時に拡張ペイロードと称される）を含むことができる。そのシンタックスはMPEG-4 AAC規格の表4.57に示されている。充填データ（たとえばその拡張ペイロード）は、eSBRオブジェクトを示すヘッダ（たとえば図７の充填要素２の「ヘッダ２」）を含むことができ（すなわち、ヘッダが、向上スペクトル帯域複製（eSBR）オブジェクト型を初期化する）、充填データ（たとえばその拡張ペイロード）は、前記ヘッダ後にeSBRメタデータを含む。たとえば、図７の充填要素２はそのようなヘッダ（「ヘッダ２」）を含み、該ヘッダ後に、eSBRメタデータ（すなわち、向上スペクトル帯域複製（eSBR）処理がそのブロックのオーディオ・コンテンツに対して実行されるかどうかを示す、充填要素２内の「フラグ」）をも含んでいる。任意的には、ヘッダ２後に、図７の充填要素２の充填データに追加的なeSBRメタデータも含められる。本段落で述べている実施形態では、ヘッダ（たとえば図７のヘッダ２）は、MPEG-4 AAC規格の表4.57において指定されている通常の値のうちの一つではなく、eSBR拡張ペイロードを示す識別情報値をもつ（よって、ヘッダのextension_typeフィールドが充填データがeSBRメタデータを含むことを示す）。 According to some embodiments of the present invention, the eSBR metadata is a filler element of an MPEG-4 AAC bitstream other than the spectral band replication extension element (SBR extension element) of the filler element (eg, filler 2 in FIG. 7). be included in This is because a filler element containing extension_payload() with SBR data or SBR data with CRC does not contain any other extension payload of any other extension type. Therefore, in embodiments where the eSBR metadata is stored in its own extension payload, a separate filler element is used to store the eSBR metadata. Such filler elements include an identifier (eg, "ID2" in FIG. 7) that indicates the beginning of the filler element, followed by the filler data. Fill data can include an extension_payload() element (sometimes referred to herein as an extension payload). Its syntax is shown in Table 4.57 of the MPEG-4 AAC Standard. The filler data (eg, its extension payload) may include a header (eg, "header 2" of filler element 2 in FIG. 7) that indicates an eSBR object (i.e., the header indicates an enhanced spectral band replication (eSBR) object type). initialization), the filler data (eg its extension payload) contains the eSBR metadata after the header. For example, filler element 2 in FIG. 7 contains such a header (“header 2”), after which eSBR metadata (i.e., the enhanced spectral band replication (eSBR) processing is applied to the audio content of that block). It also contains a "flag" in filler element 2) that indicates whether or not it is executed. Optionally, additional eSBR metadata is also included in the filler data for filler element 2 in FIG. 7 after header 2 . In the embodiment described in this paragraph, the header (e.g. Header 2 in Figure 7) indicates an eSBR extension payload rather than one of the normal values specified in Table 4.57 of the MPEG-4 AAC standard. It has an identity value (thus the extension_type field in the header indicates that the filler data contains eSBR metadata).

第一のクラスの実施形態では、本発明は、オーディオ処理ユニット（たとえばデコーダ）であって：
エンコードされたオーディオ・ビットストリームの少なくとも一つのブロック（たとえばMPEG-4 AACビットストリームの少なくとも一つのブロック）を記憶するよう構成されたメモリ（たとえば図３または図４のバッファ２０１）と；
前記メモリに結合され、前記ビットストリームの前記ブロックの少なくとも一部を多重分離するよう構成されているビットストリーム・ペイロード・フォーマット解除器（たとえば、図３の要素２０５または図４の要素２１５）と；
前記ビットストリームの前記ブロックのオーディオ・コンテンツの少なくとも一つの部分をデコードするよう結合され、構成されたデコード・サブシステム（たとえば図３の要素２０２および２０３または図４の要素２０２および２１３）とを有し、前記ブロックは、
充填要素を含み、該充填要素の先頭を示す識別子（たとえば、MPEG-4 AAC規格の表4.85の値0x6をもつid_syn_ele識別子）と、該識別子後の充填データとを含み、前記充填データは：
前記ブロックのオーディオ・コンテンツに対して（たとえば前記ブロックに含まれるスペクトル帯域複製データおよびeSBRメタデータを使って）向上スペクトル帯域複製（eSBR）処理が実行されるべきかどうかを同定する少なくとも一つのフラグを含む、
オーディオ処理ユニットである。 In a first class of embodiments, the invention is an audio processing unit (e.g. decoder) which:
a memory (eg, buffer 201 of FIG. 3 or 4) configured to store at least one block of an encoded audio bitstream (eg, at least one block of an MPEG-4 AAC bitstream);
a bitstream payload de-formatter (e.g., element 205 of FIG. 3 or element 215 of FIG. 4) coupled to the memory and configured to demultiplex at least a portion of the blocks of the bitstream;
a decoding subsystem (e.g., elements 202 and 203 of FIG. 3 or elements 202 and 213 of FIG. 4) coupled and configured to decode at least one portion of audio content of said block of said bitstream. and said block is
including a filler element, an identifier indicating the beginning of the filler element (e.g., an id_syn_ele identifier with value 0x6 in Table 4.85 of the MPEG-4 AAC Standard) and filler data after the identifier, said filler data being:
At least one flag identifying whether enhanced spectral band replication (eSBR) processing should be performed on the audio content of the block (eg, using spectral band replication data and eSBR metadata included in the block). including,
Audio processing unit.

前記フラグは、eSBRメタデータであり、前記フラグの例はsbrPatchingModeフラグである。前記フラグのもう一つの例はharmonicSBRフラグである。これらのフラグはいずれも、基本形のスペクトル帯域複製または向上した形のスペクトル複製のどちらが前記ブロックのオーディオ・データに対して実行されるべきかを示す。基本形のスペクトル複製はスペクトル・パッチングであり、向上した形のスペクトル帯域複製は高調波転換である。 Said flag is eSBR metadata and an example of said flag is the sbrPatchingMode flag. Another example of said flag is the harmonicSBR flag. Both of these flags indicate whether a basic form of spectral band duplication or an enhanced form of spectral band duplication should be performed on the block of audio data. The basic form of spectral replication is spectral patching, and the enhanced form of spectral band replication is harmonic conversion.

いくつかの実施形態では、前記充填データは追加的なeSBRメタデータ（すなわち、前記フラグ以外のeSBRメタデータ）をも含む。 In some embodiments, the fill data also includes additional eSBR metadata (ie, eSBR metadata other than the flag).

前記メモリは、エンコードされたオーディオ・ビットストリームの前記少なくとも一つのブロックを（たとえば非一時的な仕方で）記憶するバッファ・メモリ（たとえば、図４のバッファ２０１の実装）であってもよい。 The memory may be a buffer memory (eg, an implementation of buffer 201 of FIG. 4) that stores (eg, in a non-transient manner) the at least one block of an encoded audio bitstream.

eSBRメタデータを含むMPEG-4 AACビットストリームのデコードの間のeSBRデコーダによる（eSBR高調波転換、前置平坦化およびインターTESツールを使う）eSBR処理（前記eSBRメタデータがこれらのeSBRツールを示す）の実行の複雑さは、（示されるパラメータを用いた典型的なデコードについて）以下のようになると推定される：
●高調波転換（16kbps、14400/28800Hz）
○DFTベース：3.68WMOPS（weighted million operations per second［加重百万演算毎秒］）；
○WMFベース：0.98WMOPS；
●QMFパッチング前処理（前置平坦化）：0.1WMOPS；
●サブバンド・サンプル間時間的包絡整形（インターTES）：高々0.16WMOPS
過渡成分については、DFTベースの転換が典型的にはQMFベースの転換よりよい性能を発揮することがわかっている。 eSBR processing (using eSBR harmonic conversion, preflattening and inter-TES tools) by an eSBR decoder during decoding of an MPEG-4 AAC bitstream containing eSBR metadata (the eSBR metadata indicates these eSBR tools) ) is estimated to be (for a typical decoding with the parameters shown):
●Harmonic conversion (16kbps, 14400/28800Hz)
○ DFT-based: 3.68 WMOPS (weighted million operations per second);
○ WMF base: 0.98 WMOPS;
QMF patching pretreatment (pre-flattening): 0.1 WMOPS;
Sub-band inter-sample temporal envelope shaping (inter-TES): 0.16 WMOPS at most
For transient components, DFT-based transformations are found to typically perform better than QMF-based transformations.

本発明のいくつかの実施形態によれば、eSBRメタデータを含む（エンコードされたオーディオ・ビットストリームの）充填要素は、eSBRメタデータが充填要素に含まれることおよび当該ブロックのオーディオ・コンテンツに対してeSBR処理が実行されるべきであることを合図する値（たとえばbs_extension_id＝3）をもつパラメータ（たとえばbs_extension_idパラメータ）および／または充填要素のsbr_extension()コンテナがPSデータを含むことを合図する値（たとえばbs_extension_id＝2）をもつパラメータ（たとえば同じbs_extension_idパラメータ）をも含む。たとえば、下記の表１に示されるように、値bs_extension_id＝2をもつそのようなパラメータは、充填要素のsbr_extension()コンテナがPSデータを含むことを合図してもよく、値bs_extension_id＝3をもつそのようなパラメータは、充填要素のsbr_extension()コンテナがeSBRメタデータを含むことを合図してもよい。 According to some embodiments of the present invention, a filler element (of an encoded audio bitstream) containing eSBR metadata indicates that the eSBR metadata is included in the filler element and for the audio content of the block. A parameter (e.g. bs_extension_id parameter) with a value (e.g. bs_extension_id = 3) signaling that eSBR processing should be performed on the Also include parameters with bs_extension_id=2) (eg same bs_extension_id parameter). For example, as shown in Table 1 below, such a parameter with value bs_extension_id=2 may signal that the filler element's sbr_extension() container contains PS data, and has value bs_extension_id=3. Such parameters may signal that the filler element's sbr_extension() container contains eSBR metadata.

本発明のいくつかの実施形態によれば、eSBRメタデータおよび／またはPSデータを含む各スペクトル帯域複製拡張要素のシンタックスは下記の表２に示されるとおりである（ここで、sbr_extension()はスペクトル帯域複製拡張要素であるコンテナを表わし、bs_extension_idは上記の表１で述べたとおりであり、ps_dataはPSデータを表わし、esbr_dataはeSBRメタデータを表わす）。

According to some embodiments of the present invention, the syntax of each spectral band replication extension element containing eSBR metadata and/or PS data is as shown in Table 2 below (where sbr_extension() is bs_extension_id is as described in Table 1 above, ps_data represents PS data, and esbr_data represents eSBR metadata).

ある例示的実施形態では、上記の表２で言及されているesbr_data()は以下のメタデータ・パラメータの値を示す。
１．上記の一ビットのメタデータ・パラメータharmonicSBR；bs_interTES；およびbs_sbr_preprocessing；
２．デコードされるべきエンコードされたビットストリームのオーディオ・コンテンツの各チャネル（「ch」）について、上記のパラメータ：sbrPatchingMode[ch]；sbrOversamplingFlag[ch]；sbrPitchInBinsFlag[ch]；およびsbrPitchInBins[ch]のそれぞれ；および
３．デコードされるべきエンコードされたビットストリームのオーディオ・コンテンツの各チャネル（「ch」）の各SBR包絡（「env」）について、上記のパラメータ：bs_temp_shape[ch][env]；およびbs_inter_temp_shape_mode[ch][env]のそれぞれ。

In one exemplary embodiment, the esbr_data() referred to in Table 2 above indicates the values of the following metadata parameters.
1. the above one-bit metadata parameters harmonicSBR; bs_interTES; and bs_sbr_preprocessing;
2. For each channel ("ch") of the audio content of the encoded bitstream to be decoded, each of the above parameters: sbrPatchingMode[ch]; sbrOversamplingFlag[ch]; sbrPitchInBinsFlag[ch]; and sbrPitchInBins[ch]; and 3. For each SBR envelope ("env") for each channel ("ch") of the audio content of the encoded bitstream to be decoded, the above parameters: bs_temp_shape[ch][env]; and bs_inter_temp_shape_mode[ch][ env] respectively.

たとえば、いくつかの実施形態では、esbr_data()は、これらのメタデータ・パラメータを示すために、表３に示されるシンタックスを有していてもよい。 For example, in some embodiments, esbr_data() may have the syntax shown in Table 3 to indicate these metadata parameters.

表３では、中央の列における数字は左の列における対応するパラメータのビット数を示す。

In Table 3, the numbers in the middle column indicate the number of bits of the corresponding parameter in the left column.

上記のシンタックスは、高調波転換のような向上した形のスペクトル帯域複製の、レガシー・デコーダへの拡張としての効率的な実装を可能にする。具体的には、表３のeSBRデータは、向上した形のスペクトル帯域複製を実行するために必要とされるパラメータであって、ビットストリームにおいてすでにサポートされていたりビットストリームにおいてすでにサポートされているパラメータから直接導入可能であったりするものではないもののみを含む。向上した形のスペクトル帯域複製を実行するために必要とされる他のすべてのパラメータおよび処理データは、ビットストリームにおいてすでに定義されている位置にある既存のパラメータから抽出される。これは、向上スペクトル帯域複製のために使われる処理メタデータの全部を単純に送信する代替的な（より非効率的な）実装とは対照的である。 The above syntax allows efficient implementation of enhanced forms of spectral band duplication, such as harmonic conversion, as an extension to legacy decoders. Specifically, the eSBR data in Table 3 are the parameters required to perform an enhanced form of spectral band duplication that are already supported in the bitstream or already supported in the bitstream. includes only material that is not or is not directly retrievable from All other parameters and processing data required to perform the enhanced form of spectral band replication are extracted from existing parameters at locations already defined in the bitstream. This is in contrast to alternative (more inefficient) implementations that simply transmit all of the processing metadata used for enhanced spectral band duplication.

たとえば、MPEG-4 HE-AACまたはHE-AAC-v2準拠デコーダは、高調波転換のような向上した形のスペクトル帯域複製を含むよう拡張されてもよい。この向上した形のスペクトル帯域複製は、デコーダによってすでにサポートされている基本形のスペクトル帯域複製に加えてのものである。MPEG-4 HE-AACまたはHE-AAC-v2準拠デコーダのコンテキストでは、この基本形のスペクトル帯域複製は、MPEG-4 AAC規格の4.6.18節において定義されているQMFスペクトル・パッチングSBRツールである。 For example, MPEG-4 HE-AAC or HE-AAC-v2 compliant decoders may be extended to include enhanced forms of spectral band duplication such as harmonic conversion. This enhanced form of spectral band replication is in addition to the basic form of spectral band replication already supported by the decoder. In the context of MPEG-4 HE-AAC or HE-AAC-v2 compliant decoders, this basic form of spectral band replication is the QMF spectral patching SBR tool defined in section 4.6.18 of the MPEG-4 AAC standard.

向上した形のスペクトル帯域複製を実行するとき、拡張されたHE-AACデコーダは、ビットストリームのSBR拡張ペイロードにすでに含まれているビットストリーム・パラメータの多くを再利用しうる。再利用されうる具体的なパラメータは、たとえば、マスター周波数帯域テーブルを決定するさまざまなパラメータを含む。これらのパラメータは、bs_start_freq（マスター周波数テーブル・パラメータの先頭を決定するパラメータ）、bs_stop_freq（マスター周波数テーブルの終わりを決定するパラメータ）、bs_freq_scale（オクターブ当たりの周波数帯域の数を決定するパラメータ）およびbs_alter_scale（周波数帯域のスケールを変更するパラメータ）を含む。再利用されうるパラメータは、ノイズ帯域テーブルを決定するパラメータ（bs_noise_bands）およびリミッター帯域テーブル・パラメータ（bs_limiter_bands）をも含む。 When performing an enhanced form of spectral band duplication, the enhanced HE-AAC decoder may reuse many of the bitstream parameters already included in the bitstream's SBR extension payload. Specific parameters that can be reused include, for example, various parameters that determine the master frequency band table. These parameters are bs_start_freq (the parameter that determines the start of the master frequency table parameter), bs_stop_freq (the parameter that determines the end of the master frequency table), bs_freq_scale (the parameter that determines the number of frequency bands per octave) and bs_alter_scale (the parameters that change the scale of the frequency band). Parameters that can be reused also include parameters that determine the noise band table (bs_noise_bands) and limiter band table parameters (bs_limiter_bands).

前記の数多くのパラメータに加えて、他のデータ要素も、本発明の実施形態に従って向上した形のスペクトル帯域複製を実行するときに、拡張されたHE-AACデコーダによって再利用されてもよい。たとえば、包絡データおよびノイズ・フロア・データは、bs_data_envおよびbs_noise_envデータから抽出されて、向上した形のスペクトル帯域複製の間に使われてもよい。 In addition to many of the parameters described above, other data elements may also be reused by the enhanced HE-AAC decoder when performing an enhanced form of spectral band duplication according to embodiments of the present invention. For example, envelope data and noise floor data may be extracted from bs_data_env and bs_noise_env data and used during an enhanced form of spectral band replication.

本質的には、これらの実施形態は、SBR拡張ペイロードにおいてレガシーのHE-AACまたはHE-AAC v2デコーダによってすでにサポートされている構成設定パラメータおよび包絡データを、できるだけ追加的な伝送データを必要とせずに向上した形のスペクトル帯域複製を可能にするために、活用する。よって、向上した形のスペクトル帯域複製をサポートする拡張されたデコーダは、すでに定義されたビットストリーム要素（たとえばSBR拡張ペイロード内のもの）に頼り、向上した形のスペクトル帯域複製をサポートするために必要とされるパラメータのみを（充填要素拡張ペイロード内に）追加することによって、非常に効率的な仕方で生成されうる。このデータ削減特徴は、新たに追加されたパラメータを拡張コンテナのようなリザーブされたデータ・フィールドに配置することと組み合わさって、ビットストリームが向上した形のスペクトル帯域複製をサポートしないレガシー・デコーダと後方互換であることを保証することによって、向上した形のスペクトル帯域複製をサポートするデコーダを作り出すことへの障壁を実質的に軽減する。 In essence, these embodiments use the configuration parameters and envelope data already supported by legacy HE-AAC or HE-AAC v2 decoders in the SBR extension payload with as little additional transmission data as possible. to enable an improved form of spectral band duplication. Thus, an enhanced decoder that supports an improved form of spectral band duplication relies on the already defined bitstream elements (e.g., those in the SBR extension payload) to support an improved form of spectral band duplication. can be generated in a very efficient manner by adding (into the filler element extension payload) only the parameters that This data reduction feature, combined with placing the newly added parameters into reserved data fields such as the extension container, allows the bitstream to be used with legacy decoders that do not support an enhanced form of spectral band duplication. By ensuring backward compatibility, the barrier to creating decoders that support an improved form of spectral band duplication is substantially reduced.

いくつかの実施形態では、本発明は、エンコードされたビットストリーム（たとえばMPEG-4 AACビットストリーム）を生成するためにオーディオ・データをエンコードする段階を含む方法である。該生成は、eSBRメタデータをエンコードされたビットストリームの少なくとも一つのブロックの少なくとも一つのセグメントに含め、オーディオ・データを前記ブロックの少なくとも一つの他のセグメントに含めることによることを含む。典型的な実施形態では、本方法は、エンコードされたビットストリームの各ブロックにおいてオーディオ・データをeSBRメタデータと多重化する段階を含む。eSBRデコーダにおける前記エンコードされたビットストリームの典型的なデコードでは、デコーダはeSBRメタデータをビットストリームから抽出し（これはeSBRメタデータおよびオーディオ・データをパースして多重分離することによることを含む）、eSBRメタデータを、オーディオ・データを処理してデコードされたオーディオ・データのストリームを生成するために使う。 In some embodiments, the invention is a method that includes encoding audio data to produce an encoded bitstream (eg, MPEG-4 AAC bitstream). The generating includes by including eSBR metadata in at least one segment of at least one block of the encoded bitstream and audio data in at least one other segment of said block. In an exemplary embodiment, the method includes multiplexing audio data with eSBR metadata in each block of the encoded bitstream. In typical decoding of said encoded bitstream in an eSBR decoder, the decoder extracts the eSBR metadata from the bitstream (this includes by parsing and demultiplexing the eSBR metadata and audio data). , uses the eSBR metadata to process the audio data to produce a stream of decoded audio data.

本発明のもう一つの側面は、eSBRメタデータを含まないエンコードされたオーディオ・ビットストリーム（たとえばMPEG-4 AACビットストリーム）のデコードの間に、（たとえば高調波転換、前置平坦化またはインターTESとして知られるeSBRツールの少なくとも一つを使って）eSBR処理を実行するよう構成されたeSBRデコーダである。そのようなデコーダの例について、図５を参照して述べる。 Another aspect of the invention is that during decoding of encoded audio bitstreams (e.g. MPEG-4 AAC bitstreams) that do not contain eSBR metadata (e.g. harmonic conversion, preflattening or inter-TES An eSBR decoder configured to perform eSBR processing using at least one of the eSBR tools known as eSBR tools. An example of such a decoder is described with reference to FIG.

図５のeSBRデコーダ（４００）は、図のように接続された、バッファ・メモリ２０１（これは図３および図４のメモリ２０１と同一）と、ビットストリーム・ペイロード・フォーマット解除器２１５（これは図４のフォーマット解除器２１５と同一）と、オーディオ・デコード・サブシステム２０２（時に「コア」デコード段または「コア」デコード・サブシステムと称され、図３のコア・デコード・サブシステム２０２と同一）と、eSBR制御データ生成サブシステム４０１と、eSBR処理段２０３（これは図３の段２０３と同一）とを含む。典型的には、デコーダ４００は他の処理要素（図示せず）も含む。 The eSBR decoder (400) of FIG. 5 comprises a buffer memory 201 (which is identical to the memory 201 of FIGS. 3 and 4) and a bitstream payload deformatter 215 (which is a 4) and an audio decoding subsystem 202 (sometimes referred to as the "core" decoding stage or "core" decoding subsystem, identical to the core decoding subsystem 202 of FIG. 3). ), an eSBR control data generation subsystem 401, and an eSBR processing stage 203 (which is identical to stage 203 in FIG. 3). Decoder 400 typically also includes other processing elements (not shown).

デコーダ４００の動作においては、デコーダ４００によって受領されたエンコードされたオーディオ・ビットストリーム（MPEG-4 AACビットストリーム）のブロックのシーケンスがバッファ２０１からフォーマット解除器２１５に呈される。 In operation of decoder 400 , a sequence of blocks of encoded audio bitstream (MPEG-4 AAC bitstream) received by decoder 400 is presented from buffer 201 to deformatter 215 .

フォーマット解除器２１５は、ビットストリームの各ブロックを多重分離して、それからSBRメタデータ（量子化された包絡データを含む）を、典型的には他のメタデータも抽出するよう結合され、構成される。フォーマット解除器２１５は、少なくとも前記SBRメタデータをeSBR処理段２０３に呈するよう構成される。フォーマット解除器２１５は、ビットストリームの各ブロックからオーディオ・データを抽出し、抽出されたオーディオ・データをデコード・サブシステム（デコード段）２０２に呈するようにも結合され、構成される。 De-formatter 215 is coupled and configured to de-multiplex each block of the bitstream and extract therefrom the SBR metadata (including the quantized envelope data), and typically other metadata as well. be. The deformatter 215 is configured to present at least said SBR metadata to the eSBR processing stage 203 . Deformatter 215 is also coupled and configured to extract audio data from each block of the bitstream and present the extracted audio data to decoding subsystem (decoding stage) 202 .

デコーダ４００のオーディオ・デコード・サブシステム２０２は、フォーマット解除器２１５によって抽出されたオーディオ・データをデコードして（そのようなデコードは「コア」デコード動作と称されてもよい）、デコードされたオーディオ・データを生成し、デコードされたオーディオ・データをeSBR処理段２０３に呈するよう構成される。デコードは周波数領域で実行される。典型的には、サブシステム２０２における処理の最終段が、デコードされた周波数領域オーディオ・データに周波数領域から時間領域への変換を適用し、そのためサブシステムの出力は時間領域のデコードされたオーディオ・データである。段２０３は、（フォーマット解除器２１５によって抽出された）SBRメタデータおよびサブシステム４０１において生成されたeSBRメタデータによって示されるSBRツール（およびeSBRツール）を、デコードされたオーディオ・データに適用して（すなわち、SBRおよびeSBRメタデータを使ってデコード・サブシステム２０２の出力に対してSBRおよびeSBR処理を実行して）、デコーダ４００から出力される完全にデコードされたオーディオ・データを生成するよう構成される。典型的には、デコーダ４００は、フォーマット解除器２１５（および任意的にはサブシステム４０１）から出力されるフォーマット解除されたオーディオ・データおよびメタデータを記憶するメモリ（サブシステム２０２および段２０３によってアクセス可能）を含み、段２０３はSBRおよびeSBR処理の間に必要に応じてオーディオ・データおよびメタデータにアクセスするよう構成される。段２０３におけるSBR処理は、コア・デコード・サブシステム２０２の出力に対する後処理であると考えられてもよい。任意的に、デコーダ４００は、最終的なアップミックス・サブシステム（これは、フォーマット解除器２１５によって抽出されたPSメタデータを使って、MPEG-4 AAC規格において定義されているパラメトリック・ステレオ（「PS」）ツールを適用しうる）をも含む。アップミックス・サブシステムは、段２０３の出力に対してアップミックスを実行して、APU ２１０から出力される、完全にデコードされた、アップミックスされたオーディオを生成するよう結合され、構成される。 Audio decoding subsystem 202 of decoder 400 decodes the audio data extracted by deformatter 215 (such decoding may be referred to as a "core" decoding operation) to produce the decoded audio • configured to generate data and present the decoded audio data to the eSBR processing stage 203; Decoding is performed in the frequency domain. Typically, the final stage of processing in subsystem 202 applies a frequency-domain to time-domain transform to the decoded frequency-domain audio data, so that the output of the subsystem is time-domain decoded audio data. Data. Stage 203 applies the SBR tools (and eSBR tools) indicated by the SBR metadata (extracted by deformatter 215) and the eSBR metadata generated in subsystem 401 to the decoded audio data. configured to produce fully decoded audio data output from decoder 400 (i.e., perform SBR and eSBR processing on the output of decoding subsystem 202 using SBR and eSBR metadata). be done. Typically, decoder 400 has a memory (accessed by subsystem 202 and stage 203) that stores unformatted audio data and metadata output from deformatter 215 (and optionally subsystem 401). possible), and stage 203 is configured to access audio data and metadata as needed during SBR and eSBR processing. The SBR processing in stage 203 may be considered post-processing on the output of core decode subsystem 202 . Optionally, the decoder 400 uses the final upmix subsystem (which uses the PS metadata extracted by the deformatter 215 to convert parametric stereo (" PS”) tool may be applied). An upmix subsystem is coupled and configured to perform an upmix on the output of stage 203 to produce fully decoded, upmixed audio output from APU 210 .

図５の制御データ生成サブシステム４０１は、デコードされるべきエンコードされたオーディオ・ビットストリームの少なくとも一つの属性を検出し、検出段階の少なくとも一つの結果に応答してeSBR制御データ（これは、本発明の他の実施形態に従って、エンコードされたオーディオ・ビットストリームに含まれている型のうちいずれかの型のeSBRメタデータであってもく、それを含んでいてもよい）を生成するよう結合され、構成される。eSBR制御データは、段２０３に呈されて、ビットストリームの特定の属性（または複数の属性の組み合わせ）を検出したときに個々のeSBRツールまたはeSBRツールの組み合わせの適用を惹起するおよび／またはそのようなeSBRツールの適用を制御する。たとえば、高調波転換を使ったeSBR処理の実行を制御するために、制御データ生成サブシステム４０１のいくつかの実施形態は：ビットストリームが音楽を示すまたは示さないことを検出することに応答してsbrPatchingMode[ch]パラメータを設定する（そして設定されたパラメータを段２０３に呈する）ための音楽検出器（たとえば、通常の音楽検出器の単純化されたバージョン）；ビットストリームによって示されるオーディオ・コンテンツにおける過渡成分の存在または不在を検出することに応答してsbrOversamplingFlag[ch]パラメータを設定する（そして設定されたパラメータを段２０３に呈する）ための過渡検出器；および／またはビットストリームによって示されるオーディオ・コンテンツのピッチを検出することに応答してsbrPitchInBinsFlag[ch]およびsbrPitchInBins[ch]パラメータを設定する（そして設定されたパラメータを段２０３に呈する）ためのピッチ検出器を含むことになる。本発明の他の側面は、この段落および前段落において述べた本発明のデコーダのいずれかの実施形態によって実行されるオーディオ・ビットストリーム・デコード方法である。 The control data generation subsystem 401 of FIG. 5 detects at least one attribute of the encoded audio bitstream to be decoded and, in response to at least one result of the detection step, eSBR control data (which is according to another embodiment of the invention, to produce eSBR metadata of any type contained in the encoded audio bitstream). configured. eSBR control data is presented to stage 203 to trigger the application of individual eSBR tools or combinations of eSBR tools and/or such when certain attributes (or combinations of attributes) of the bitstream are detected. control the application of various eSBR tools. For example, to control the execution of eSBR processing using harmonic conversion, some embodiments of control data generation subsystem 401: In response to detecting that the bitstream indicates or does not indicate music: A music detector (e.g., a simplified version of a regular music detector) for setting the sbrPatchingMode[ch] parameter (and presenting the set parameter to stage 203); a transient detector for setting the sbrOversamplingFlag[ch] parameter (and presenting the set parameter to stage 203) in response to detecting the presence or absence of a transient; A pitch detector will be included for setting the sbrPitchInBinsFlag[ch] and sbrPitchInBins[ch] parameters (and presenting the set parameters to stage 203) in response to detecting the pitch of the content. Another aspect of the invention is an audio bitstream decoding method performed by any of the embodiments of the decoder of the invention described in this paragraph and the previous paragraph.

本発明の諸側面は、本発明のAPU、システムまたはデバイスのいずれかの実施形態が実行するよう構成される（たとえばプログラムされる）型のエンコードまたはデコード方法を含む。本発明の他の側面は、本発明の方法のいずれかの実施形態を実行するよう構成された（たとえばプログラムされた）システムまたはデバイスならびに本発明の方法のいずれかの実施形態もしくはその段階を実装するためのコードを（たとえば非一時的な仕方で）記憶するコンピュータ可読媒体（たとえばディスク）を含む。たとえば、本発明のシステムは、プログラム可能な汎用プロセッサ、デジタル信号プロセッサまたはマイクロプロセッサが、本発明の方法の実施形態またはその段階を含む多様な動作のいずれかをデータに対して実行するようソフトウェアもしくはファームウェアを用いてプログラムされたおよび／または他の仕方で構成されたものであるまたはそれを含むことができる。そのような汎用プロセッサは、入力装置、メモリおよび処理回路を含むコンピュータ・システムが、それに呈されるデータに応答して本発明の方法の実施形態（またはその段階）を実行するようプログラムされた（および／または他の仕方で構成された）ものであってもよく、あるいはそれを含んでいてもよい。 Aspects of the present invention include encoding or decoding methods of the type that any embodiment of the APU, system or device of the present invention is configured (eg, programmed) to perform. Other aspects of the invention include systems or devices configured (eg programmed) to perform any embodiment of the method of the invention and implementing any embodiment of the method of the invention or steps thereof. computer readable medium (eg, disk) that stores (eg, in a non-transitory manner) code for For example, the system of the present invention may comprise software or software such that a programmable general purpose processor, digital signal processor or microprocessor performs any of a variety of operations on data, including embodiments of the method of the present invention or steps thereof. It may be or include programmed with firmware and/or otherwise configured. Such a general-purpose processor is a computer system comprising input devices, memory and processing circuitry programmed to perform an embodiment (or steps thereof) of the method of the present invention in response to data presented thereto ( and/or otherwise configured).

本発明の実施形態は、ハードウェア、ファームウェアまたはソフトウェアまたは両者の組み合わせにおいて（たとえばプログラム可能な論理アレイとして）実装されてもよい。特に断わりのない限り、本発明の一部として含まれるアルゴリズムまたはプロセスは、いかなる特定のコンピュータまたは他の装置にも本来的に関係していることはない。特に、さまざまな汎用機械が、本稿の教示に従って書かれたプログラムと一緒に使われてもよいし、あるいは要求される方法段階を実行するよう、より特化した装置（たとえば集積回路）を構築するほうが便利であることもありうる。このように本発明は、一つまたは複数のプログラム可能なコンピュータ・システム（たとえば、図１の要素または図２のエンコーダ１００（またはそのある要素）または図３のデコーダ２００（またはそのある要素）または図４のデコーダ２１０（またはそのある要素）または図５のデコーダ４００（またはそのある要素）のいずれかの実装）上で実行される一つまたは複数のコンピュータ・プログラムにおいて実装されてもよい。各コンピュータ・システムは少なくとも一つのプロセッサと、少なくとも一つのデータ記憶システム（揮発性および不揮発性メモリおよび／または記憶要素を含む）と、少なくとも一つの入力装置またはポートと、少なくとも一つの出力装置またはポートとを有する。プログラム・コードは、本稿に記載される機能を実行して出力情報を生成するために、入力データに適用される。出力情報は、既知の仕方で一つまたは複数の出力装置に加えられる。 Embodiments of the invention may be implemented in hardware, firmware or software, or a combination of both (eg, as a programmable logic array). Unless specified otherwise, the algorithms or processes included as part of the invention are not inherently related to any particular computer or other apparatus. In particular, various general-purpose machines may be used with programs written in accordance with the teachings herein, or construct more specialized apparatus (eg, integrated circuits) to perform the required method steps. It may be more convenient. Thus, the present invention may be implemented in one or more programmable computer systems (e.g., elements of FIG. 1 or encoder 100 (or some element thereof) of FIG. 2 or decoder 200 (or some element thereof) of FIG. 3 or 4 (or some element thereof) or any implementation of decoder 400 (or some element thereof) of FIG. 5). Each computer system has at least one processor, at least one data storage system (including volatile and nonvolatile memory and/or storage elements), at least one input device or port, and at least one output device or port. and Program code is applied to input data to perform the functions described herein and generate output information. The output information is applied to one or more output devices in known fashion.

そのような各プログラムは、コンピュータ・システムと連絡するためにいかなる所望されるコンピュータ言語（機械語、アセンブリーまたは高レベルの手続き型、論理的またはオブジェクト指向のプログラミング言語を含む）で実装されてもよい。いずれにせよ、言語はコンパイルまたはインタープリットされる言語でありうる。 Each such program may be implemented in any desired computer language (including machine language, assembly or a high level procedural, logical or object oriented programming language) to communicate with a computer system. . In any case, the language can be a compiled or interpreted language.

たとえば、コンピュータ・ソフトウェア命令シーケンスによって実装されるとき、本発明の実施形態のさまざまな機能および段階は、好適なデジタル信号処理ハードウェアにおいて走るマルチスレッド・ソフトウェア命令シーケンスによって実装されてもよく、その場合、実施形態のさまざまな装置、段階および機能はソフトウェア命令の諸部分に対応しうる。 For example, when implemented by computer software instruction sequences, the various functions and steps of embodiments of the invention may be implemented by multi-threaded software instruction sequences running on suitable digital signal processing hardware, where , the various devices, steps and functions of the embodiments may correspond to portions of software instructions.

そのような各コンピュータ・システムは、好ましくは、汎用または特殊目的のプログラム可能なコンピュータによって読み取り可能な記憶媒体またはデバイス（たとえば半導体メモリもしくはメディアまたは磁気もしくは光学式メディア）に記憶され、またはダウンロードされる。該記憶媒体またはデバイスがコンピュータ・システムによって読まれるときに、本稿に記載される手順を実行するようコンピュータを構成し、動作させるためである。本発明のシステムは、コンピュータ・プログラムをもって構成された（すなわちコンピュータ・プログラムを記憶している）コンピュータ可読記憶媒体として実装されてもよい。ここで、そのように構成された記憶媒体はコンピュータ・システムに、本稿に記載される機能を実行するよう、特定のあらかじめ定義された仕方で動作させる。 Each such computer system is preferably stored on or downloaded from a general purpose or special purpose programmable computer readable storage medium or device (e.g., semiconductor memory or medium or magnetic or optical medium). . To configure and operate a computer to perform the procedures described herein when the storage medium or device is read by a computer system. The system of the present invention may be implemented as a computer-readable storage medium configured with (ie, having stored thereon) a computer program. Here, the storage medium so configured causes the computer system to operate in a specific predefined manner to perform the functions described herein.

本発明のいくつかの実施形態を記述してきた。にもかかわらず、本発明の精神および範囲から外れることなくさまざまな修正がなしうることは理解されるであろう。上記の教示に照らして本発明の数多くの修正および変形が可能である。付属の請求項の範囲内で、本発明は、本稿に具体的に記述されている以外の仕方で実施されうることは理解される。請求項に含まれる参照符号があったとしても、単に例解目的のためであり、いかなる仕方であれ請求項を解釈したり限定したりするために使われるべきではない。 A number of embodiments of the invention have been described. Nevertheless, it will be understood that various modifications can be made without departing from the spirit and scope of the invention. Many modifications and variations of the present invention are possible in light of the above teachings. It is understood that within the scope of the appended claims, the invention may be practiced otherwise than as specifically described herein. Any reference signs contained in the claims are for illustrative purposes only and shall not be used to interpret or limit the claims in any way.

いくつかの態様を記載しておく。
〔態様１〕
エンコードされたオーディオ・ビットストリームの少なくとも一つのブロックを記憶するよう構成されたバッファと；
前記バッファに結合され、前記エンコードされたオーディオ・ビットストリームの前記少なくとも一つのブロックの少なくとも一部を多重分離するよう構成されたビットストリーム・ペイロード・フォーマット解除器と；
前記ビットストリーム・ペイロード・フォーマット解除器に結合され、前記エンコードされたオーディオ・ビットストリームの前記少なくとも一つのブロックの少なくとも一部をデコードするよう構成されたデコード・サブシステムとを有するオーディオ処理ユニットであって、前記エンコードされたオーディオ・ビットストリームの前記少なくとも一つのブロックは：
充填要素を含み、該充填要素は、該充填要素の先頭を示す識別子と、該識別子の後の充填データとをもち、前記充填データは：
前記エンコードされたオーディオ・ビットストリームの前記少なくとも一つのブロックのオーディオ・コンテンツに対して基本形のスペクトル帯域複製または向上された形のスペクトル帯域複製のどちらが実行されるべきかどうかを同定する少なくとも一つのフラグを含む、
オーディオ処理ユニット。
〔態様２〕
前記基本形のスペクトル帯域複製はスペクトル・パッチングを含み、前記向上された形のスペクトル帯域複製は高調波転換を含み、前記充填データはさらに向上スペクトル帯域複製メタデータを含み、前記向上スペクトル帯域複製メタデータは、スペクトル・パッチングおよび高調波転換両方のために使われる一つまたは複数のパラメータを含まない、態様１記載のオーディオ処理ユニット。
〔態様３〕
スペクトル・パッチングおよび高調波転換両方のために使われる前記一つまたは複数のパラメータが、充填要素の拡張ペイロードに含まれる、態様２記載のオーディオ処理ユニット。
〔態様４〕
スペクトル・パッチングおよび高調波転換両方のために使われる前記一つまたは複数のパラメータが、マスター周波数帯域テーブルを定義する一つまたは複数のパラメータを含む、態様２または３記載のオーディオ処理ユニット。
〔態様５〕
スペクトル・パッチングおよび高調波転換両方のために使われる前記一つまたは複数のパラメータが、包絡スケール因子またはノイズ・フロア・スケール因子を含む、態様２または３記載のオーディオ処理ユニット。
〔態様６〕
当該オーディオ処理ユニットがオーディオ・デコーダであり、前記識別子が、0x6の値をもつ、三ビットの、最上位ビットが最初に伝送される符号なし整数である、態様１ないし５のうちいずれか一項記載のオーディオ処理ユニット。
〔態様７〕
前記充填データが拡張ペイロードを含み、前記拡張ペイロードがスペクトル帯域複製拡張データを含み、前記拡張ペイロードは、「1101」または「1110」の値をもつ、四ビットの、最上位ビットが最初に伝送される符号なし整数を用いて同定され、任意的には、
前記スペクトル帯域複製拡張データは：
任意的なスペクトル帯域複製ヘッダ、
前記ヘッダの後のスペクトル帯域複製データおよび
前記スペクトル帯域複製データの後のスペクトル帯域複製拡張要素を含み、前記第一のフラグは、前記スペクトル帯域複製拡張要素に含まれる、
態様１ないし６のうちいずれか一項記載のオーディオ処理ユニット。
〔態様８〕
前記エンコードされたオーディオ・ビットストリームの前記少なくとも一つのブロックは、第一の充填要素および第二の充填要素を含み、前記第一の充填要素にはスペクトル帯域複製データが含まれ、前記第二の充填要素には前記第一のフラグが含まれるが、スペクトル帯域複製データは含まれない、態様１ないし７のうちいずれか一項記載のオーディオ処理ユニット。
〔態様９〕
前記向上された形のスペクトル帯域複製処理は高調波転換を含み、前記基本形のスペクトル帯域複製処理はスペクトル・パッチングを含み、前記第一のフラグの一つの値は、前記エンコードされたオーディオ・ビットストリームの前記少なくとも一つのブロックのオーディオ・コンテンツに対して前記向上された形のスペクトル帯域複製処理が実行されるべきであることを示し、前記第一のフラグの別の値は、前記エンコードされたオーディオ・ビットストリームの前記少なくとも一つのブロックのオーディオ・コンテンツに対してスペクトル・パッチングが実行されるべきであるが前記高調波転換は実行されるべきではないことを示す、態様１ないし８のうちいずれか一項記載のオーディオ処理ユニット。
〔態様１０〕
前記スペクトル帯域複製拡張要素が、前記第一のフラグ以外の向上スペクトル帯域複製メタデータを含み、前記向上スペクトル帯域複製メタデータが前置平坦化を実行するかどうかを示すパラメータを含む、態様７記載のオーディオ処理ユニット。
〔態様１１〕
前記スペクトル帯域複製拡張要素が、前記第一のフラグ以外の向上スペクトル帯域複製メタデータを含み、前記向上スペクトル帯域複製メタデータがサブバンド・サンプル間時間的包絡整形を実行するかどうかを示すパラメータを含む、態様７記載のオーディオ処理ユニット。
〔態様１２〕
前記第一のフラグを使って向上スペクトル帯域複製処理を実行するよう構成された向上スペクトル帯域複製処理サブシステムをさらに有する、態様１ないし１１のうちいずれか一項記載のオーディオ処理ユニット。
〔態様１３〕
前記少なくとも一つのフラグが前記向上された形のスペクトル帯域複製処理を示す場合、第二のフラグが、信号適応的な周波数領域オーバーサンプリングが有効にされるか無効にされるかを同定する、態様１ないし１２のうちいずれか一項記載のオーディオ処理ユニット。
〔態様１４〕
エンコードされたオーディオ・ビットストリームをデコードする方法であって：
エンコードされたオーディオ・ビットストリームの少なくとも一つのブロックを受領する段階と；
前記エンコードされたオーディオ・ビットストリームの前記少なくとも一つのブロックの少なくとも一部を多重分離する段階と；
前記エンコードされたオーディオ・ビットストリームの前記少なくとも一つのブロックの少なくとも一部をデコードする段階とを含み、
前記エンコードされたオーディオ・ビットストリームの前記少なくとも一つのブロックは：
充填要素を含み、該充填要素は、該充填要素の先頭を示す識別子と、該識別子の後の充填データとをもち、前記充填データは：
前記エンコードされたオーディオ・ビットストリームの前記少なくとも一つのブロックのオーディオ・コンテンツに対して基本形のスペクトル帯域複製または向上された形のスペクトル帯域複製処理のどちらが実行されるべきかを同定する少なくとも一つのフラグを含む、
方法。
〔態様１５〕
前記識別子が、0x6の値をもつ、三ビットの、最上位ビットが最初に伝送される符号なし整数である、態様１４記載の方法。
〔態様１６〕
前記基本形のスペクトル帯域複製はスペクトル・パッチングを含み、前記向上された形のスペクトル帯域複製は高調波転換を含み、前記充填データはさらに向上スペクトル帯域複製メタデータを含み、前記向上スペクトル帯域複製メタデータは、スペクトル・パッチングおよび高調波転換両方のために使われる一つまたは複数のパラメータを含まない、態様１４または１５記載の方法。
〔態様１７〕
前記充填データが拡張ペイロードを含み、前記拡張ペイロードがスペクトル帯域複製拡張データを含み、前記拡張ペイロードは、「1101」または「1110」の値をもつ、四ビットの、最上位ビットが最初に伝送される符号なし整数を用いて同定され、任意的には、
前記スペクトル帯域複製拡張データは：
任意的なスペクトル帯域複製ヘッダ、
前記ヘッダの後のスペクトル帯域複製データおよび
前記スペクトル帯域複製データの後のスペクトル帯域複製拡張要素を含み、前記第一のフラグは、前記スペクトル帯域複製拡張要素に含まれる、
態様１４ないし１６のうちいずれか一項記載の方法。
〔態様１８〕
前記向上された形のスペクトル帯域複製処理が高調波転換であり、前記基本形のスペクトル帯域複製はスペクトル・パッチングであり、前記第一のフラグのある値は前記エンコードされたオーディオ・ビットストリームの前記少なくとも一つのブロックのオーディオ・コンテンツに対して前記向上された形のスペクトル帯域複製処理が実行されるべきであることを示し、前記第一のフラグの別の値は前記エンコードされたオーディオ・ビットストリームの前記少なくとも一つのブロックのオーディオ・コンテンツに対してスペクトル・パッチングが実行されるべきであるが前記高調波転換は実行されるべきではないことを示す、態様１４ないし１７のうちいずれか一項記載の方法。
〔態様１９〕
前記スペクトル帯域複製拡張要素が、前記第一のフラグ以外の向上スペクトル帯域複製メタデータを含み、前記向上スペクトル帯域複製メタデータが前置平坦化を実行するかどうかを示すパラメータを含む、または、
前記スペクトル帯域複製拡張要素が、前記第一のフラグ以外の向上スペクトル帯域複製メタデータを含み、前記向上スペクトル帯域複製メタデータがサブバンド・サンプル間時間的包絡整形を実行するかどうかを示すパラメータを含む、
態様１７または１８記載の方法。
〔態様２０〕
前記第一のフラグおよび前記第二のフラグを使って向上スペクトル帯域複製処理を実行する段階をさらに含み、前記向上スペクトル帯域複製は高調波転換を含む、態様１４ないし１９のうちいずれか一項記載の方法。
〔態様２１〕
前記エンコードされたオーディオ・ビットストリームがMPEG-4 AACビットストリームである、態様１４ないし２０のうちいずれか一項記載の方法または態様１ないし８のうちいずれか一項記載のオーディオ処理ユニット。 Some aspects are described.
[Aspect 1]
a buffer configured to store at least one block of an encoded audio bitstream;
a bitstream payload de-formatter coupled to said buffer and configured to demultiplex at least a portion of said at least one block of said encoded audio bitstream;
a decoding subsystem coupled to the bitstream payload de-formatter and configured to decode at least a portion of the at least one block of the encoded audio bitstream. and said at least one block of said encoded audio bitstream is:
A filler element having an identifier indicating the beginning of the filler element and filler data following the identifier, the filler data comprising:
At least one flag identifying whether a basic form of spectral band duplication or an enhanced form of spectral band duplication is to be performed on the audio content of the at least one block of the encoded audio bitstream. including,
audio processing unit.
[Aspect 2]
said base form of spectral band replication comprising spectral patching, said enhanced form of spectral band replication comprising harmonic conversion, said filling data further comprising enhanced spectral band replication metadata, said enhanced spectral band replication metadata does not include one or more parameters used for both spectral patching and harmonic conversion.
[Aspect 3]
3. The audio processing unit of aspect 2, wherein the one or more parameters used for both spectral patching and harmonic conversion are included in an expansion payload of a filler element.
[Aspect 4]
4. The audio processing unit of aspects 2 or 3, wherein the one or more parameters used for both spectral patching and harmonic conversion include one or more parameters defining a master frequency band table.
[Aspect 5]
4. The audio processing unit of aspects 2 or 3, wherein the one or more parameters used for both spectral patching and harmonic conversion include an envelope scale factor or a noise floor scale factor.
[Aspect 6]
6. Any one of aspects 1-5, wherein the audio processing unit is an audio decoder and the identifier is a 3-bit unsigned integer with a value of 0x6, the most significant bit being transmitted first. Audio processing unit as described.
[Aspect 7]
The filling data includes an extension payload, the extension payload includes spectral band duplication extension data, the extension payload has a value of '1101' or '1110' and is of four bits, the most significant bit being transmitted first. is identified using an unsigned integer that is
Said spectral band replication extension data is:
an optional spectral band duplication header,
spectrum band replication data after the header; and a spectrum band replication extension element after the spectrum band replication data, wherein the first flag is included in the spectrum band replication extension element.
7. An audio processing unit according to any one of aspects 1-6.
[Aspect 8]
The at least one block of the encoded audio bitstream includes a first filler element and a second filler element, the first filler element including spectral band replication data, and the second filler element. 8. The audio processing unit of any one of aspects 1-7, wherein a filler element includes the first flag but does not include spectral band replication data.
[Aspect 9]
The enhanced form of spectral band duplication includes harmonic conversion, the base form of spectral band duplication includes spectral patching, and a value of the first flag is equal to the encoded audio bitstream. , another value of the first flag indicating that the enhanced form of spectral band replication processing should be performed on the audio content of the at least one block of the encoded audio any of aspects 1-8, indicating that spectral patching should be performed on the audio content of said at least one block of a bitstream, but said harmonic conversion should not be performed; An audio processing unit according to any preceding claim.
[Aspect 10]
8. Aspect 7. The aspect 7, wherein the spectral band replication extension element includes enhanced spectral band replication metadata other than the first flag, and includes a parameter indicating whether the enhanced spectral band replication metadata performs pre-flattening. audio processing unit.
[Aspect 11]
wherein the spectral band replication extension element includes enhanced spectral band replication metadata other than the first flag, and a parameter indicating whether the enhanced spectral band replication metadata performs subband-to-sample temporal envelope shaping. 8. The audio processing unit of aspect 7, comprising:
[Aspect 12]
12. The audio processing unit of any one of aspects 1-11, further comprising an enhanced spectral band replication processing subsystem configured to perform an enhanced spectral band replication processing using the first flag.
[Aspect 13]
A second flag identifies whether signal adaptive frequency domain oversampling is enabled or disabled when said at least one flag indicates said enhanced form of spectral band replication processing. 13. An audio processing unit according to any one of claims 1-12.
[Aspect 14]
A method for decoding an encoded audio bitstream comprising:
receiving at least one block of an encoded audio bitstream;
demultiplexing at least a portion of the at least one block of the encoded audio bitstream;
decoding at least a portion of the at least one block of the encoded audio bitstream;
The at least one block of the encoded audio bitstream is:
A filler element having an identifier indicating the beginning of the filler element and filler data following the identifier, the filler data comprising:
At least one flag identifying whether a basic form of spectral band duplication or an enhanced form of spectral band duplication is to be performed on the audio content of the at least one block of the encoded audio bitstream. including,
Method.
[Aspect 15]
15. The method of aspect 14, wherein the identifier is a 3-bit, most significant bit transmitted first unsigned integer with a value of 0x6.
[Aspect 16]
said base form of spectral band replication comprising spectral patching, said enhanced form of spectral band replication comprising harmonic conversion, said filling data further comprising enhanced spectral band replication metadata, said enhanced spectral band replication metadata 16. The method of aspect 14 or 15, wherein does not include one or more parameters used for both spectral patching and harmonic conversion.
[Aspect 17]
The filling data includes an extension payload, the extension payload includes spectral band duplication extension data, the extension payload has a value of '1101' or '1110' and is of four bits, the most significant bit being transmitted first. is identified using an unsigned integer that is
Said spectral band replication extension data is:
an optional spectral band duplication header,
spectrum band replication data after the header; and a spectrum band replication extension element after the spectrum band replication data, wherein the first flag is included in the spectrum band replication extension element.
17. The method of any one of aspects 14-16.
[Aspect 18]
The enhanced form of spectral band duplication is harmonic conversion, the base form of spectral band duplication is spectral patching, and a value of the first flag is at least Another value of the first flag indicates that the enhanced form of spectral band replication processing should be performed on a block of audio content, and another value of the first flag indicates that the encoded audio bitstream is 18. The method of any one of aspects 14-17, indicating that spectral patching should be performed on the at least one block of audio content, but that the harmonic conversion should not be performed. Method.
[Aspect 19]
the spectral band replication extension element includes enhanced spectral band replication metadata other than the first flag, and includes a parameter indicating whether the enhanced spectral band replication metadata performs pre-flattening; or
wherein the spectral band replication extension element includes enhanced spectral band replication metadata other than the first flag, and a parameter indicating whether the enhanced spectral band replication metadata performs subband-to-sample temporal envelope shaping. include,
19. A method according to aspect 17 or 18.
[Aspect 20]
20. The method of any one of aspects 14-19, further comprising performing an enhanced spectral band replication process using the first flag and the second flag, wherein the enhanced spectral band replication includes harmonic conversion. the method of.
[Aspect 21]
21. The method of any one of aspects 14-20 or the audio processing unit of any one of aspects 1-8, wherein the encoded audio bitstream is an MPEG-4 AAC bitstream.

Claims

An audio processing unit for decoding an encoded audio bitstream, the audio processing unit:
a bitstream payload de-formatter configured to demultiplex the encoded audio bitstream;
a decoding subsystem coupled to the bitstream payload deformatter and configured to decode the encoded audio bitstream, the encoded audio bitstream comprising:
A filler element having an identifier indicating the beginning of the filler element and filler data following the identifier, the filler data comprising:
at least one flag identifying whether a basic form of spectral band duplication or an enhanced form of spectral band duplication is to be performed on the audio content of the encoded audio bitstream;
The base form of spectral band replication comprises spectral patching, the enhanced form of spectral band replication comprises harmonic conversion, and one value of the flag is the enhanced form of spectral band replication for the audio content. indicating that spectral band replication should be performed, another value of said flag indicating that said base form spectral band replication should be performed on said audio content but said harmonic conversion should not be performed; indicate that it should not
said spectral patching copies several adjacent quadrature mirror filter (QMF) subbands from a low pass portion of an audio signal to a high pass portion of said audio signal;
audio processing unit.

2. The audio processing unit of claim 1, wherein the fill data further includes enhancement spectral band replication metadata.

3. The audio processing unit of claim 2, wherein the enhanced spectral band replication metadata is included in an extension payload of a filler element.

4. Audio processing unit according to claim 2 or 3, wherein the enhanced spectral band replication metadata comprises one or more parameters defining a master frequency band table.

4. Audio processing unit according to claim 2 or 3, wherein the encoded audio bitstream further comprises spectral band replication metadata, said spectral band replication metadata comprising envelope data or noise floor data .

A method for decoding an encoded audio bitstream comprising:
demultiplexing the encoded audio bitstream;
decoding the encoded audio bitstream;
The encoded audio bitstream is:
A filler element having an identifier indicating the beginning of the filler element and filler data following the identifier, the filler data comprising:
at least one flag identifying whether a basic form of spectral band duplication or an enhanced form of spectral band duplication is to be performed on the audio content of the encoded audio bitstream;
The base form of spectral band replication comprises spectral patching, the enhanced form of spectral band replication comprises harmonic conversion, and one value of the flag is the enhanced form of spectral band replication for the audio content. indicating that spectral band replication should be performed, another value of said flag indicating that said base form spectral band replication should be performed on said audio content but said harmonic conversion should not be performed; indicate that it should not
said spectral patching copies several adjacent quadrature mirror filter (QMF) subbands from a low pass portion of an audio signal to a high pass portion of said audio signal;
Method.

7. The method of claim 6, wherein the identifier is a 3-bit unsigned integer with a value of 0x6, with the most significant bit transmitted first.

8. A method according to claim 6 or 7, wherein said filling data further comprises enhanced spectral band replication metadata.

A computer readable storage medium storing program instructions which, when executed by a processor, cause the processor to perform the method of any one of claims 6 to 8.

An apparatus for decoding an encoded audio bitstream, the apparatus comprising:
a memory configured to store program instructions;
a processor coupled to the memory and configured to execute the program instructions;
The program instructions, when executed by the processor, cause the processor to perform the method of any one of claims 6-8,
Device.