EP4736160A1 - Spatial coding of object-based audio - Google Patents
Spatial coding of object-based audioInfo
- Publication number
- EP4736160A1 EP4736160A1 EP24742379.1A EP24742379A EP4736160A1 EP 4736160 A1 EP4736160 A1 EP 4736160A1 EP 24742379 A EP24742379 A EP 24742379A EP 4736160 A1 EP4736160 A1 EP 4736160A1
- Authority
- EP
- European Patent Office
- Prior art keywords
- audio
- matrix
- spatial coding
- coding method
- objects
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L19/00—Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
- G10L19/008—Multichannel audio signal coding or decoding using interchannel correlation to reduce redundancy, e.g. joint-stereo, intensity-coding or matrixing
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F3/00—Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
- G06F3/16—Sound input; Sound output
- G06F3/165—Management of the audio stream, e.g. setting of volume, audio stream path
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S3/00—Systems employing more than two channels, e.g. quadraphonic
- H04S3/008—Systems employing more than two channels, e.g. quadraphonic in which the audio signals are in digital form, i.e. employing more than two discrete digital channels
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S7/00—Indicating arrangements; Control arrangements, e.g. balance control
- H04S7/30—Control circuits for electronic adaptation of the sound field
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L19/00—Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
- G10L19/04—Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis using predictive techniques
- G10L19/16—Vocoder architecture
- G10L19/18—Vocoders using multiple modes
- G10L19/20—Vocoders using multiple modes using sound class specific coding, hybrid encoders or object based coding
Landscapes
- Engineering & Computer Science (AREA)
- Physics & Mathematics (AREA)
- Multimedia (AREA)
- Signal Processing (AREA)
- Acoustics & Sound (AREA)
- Theoretical Computer Science (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Human Computer Interaction (AREA)
- General Physics & Mathematics (AREA)
- General Engineering & Computer Science (AREA)
- General Health & Medical Sciences (AREA)
- Mathematical Physics (AREA)
- Computational Linguistics (AREA)
- Compression, Expansion, Code Conversion, And Decoders (AREA)
- Stereophonic System (AREA)
Abstract
A spatial coding method and audio system configured to reduce the complexity of an audio scene via audio-object clustering. In at least some examples, the spatial coding method is implemented using only a limited set of basic matrix operations, which tends to significantly reduce the associated computational complexity. For example, the spatial coding method employs a cost-matrix construction approach, under which the object inter-product matrix is constructed first and then a plurality of cost matrices is derived therefrom by decimation and addition, with no advanced computational operations, such as multiplications or divisions, being needed. At least some embodiments can beneficially be used for reduction, simplification, or compression of complex audio content, with minimal impact on the audio quality, such that the audio content can be distributed through transmission systems that do not possess sufficient bandwidth to timely deliver all of the original audio-object data to the end users.
Description
SPATIAL CODING OF OBJECT-BASED AUDIO 1. Cross-Reference to Related Applications [0001] This application claims the benefit of priority from PCT International Application PCT/CN2023/104047 filed on 29 June 2023, and U.S. Provisional Application. No.63/558,254 filed on 27 February 2024, each of which is incorporated herein by reference in its entirety. 2. Field of the Disclosure [0002] Various example embodiments relate generally to audio signal processing and, more specifically but not exclusively, to spatial coding of object-based audio for rendering with bandwidth-constrained playback systems at runtime. 3. Background [0003] Some spatial audio formats include both audio beds and audio objects. Herein, the term “audio beds” refers to audio channels that are meant to be reproduced as originating from predefined, fixed locations. The term “audio objects” refers to individual audio elements that may exist for a defined duration in time but also have spatial information of each object, such as position, size, etc. During audio transmission, audio beds and audio objects can be sent separately and then used by a spatial audio rendering system to recreate the audio scene in accordance with the artistic intent. In some examples, the audio rendering system can have a variable number of speakers or headphones. In recent years, various spatial audio formats are becoming progressively more popular with users both for music creation and for interactive entertainment content, such as gaming and eXtended Reality (XR) content. BRIEF SUMMARY OF SOME SPECIFIC EMBODIMENTS [0004] Disclosed herein are various embodiments of a spatial coding method and audio system configured to reduce the complexity of an audio scene via audio-object clustering. In at least some examples, the spatial coding method is implemented using only a limited set of basic matrix operations, which tends to significantly reduce the associated computational complexity. For example, the spatial coding method employs a streamlined cost-matrix construction approach, under which the object inter-product matrix is constructed first and then a plurality of cost matrices is derived therefrom by decimation and addition, with no advanced computational operations, such as multiplications or divisions, being needed. At least some embodiments described herein can beneficially be used for reduction, simplification, or compression of
complex audio content, with minimal impact on the audio quality, such that the audio content can be distributed through transmission systems that do not possess sufficient bandwidth to timely deliver all of the original audio-object data to the end users. [0005] According to an example embodiment, a spatial coding method for object-based audio comprises: selecting N cluster seeds from L audio objects of an audio scene based on perceptually weighted energies of the L audio objects, where N < L; and obtaining N clusters corresponding to the N cluster seeds by applying a respective one of L gain vectors to each of the L audio objects, each of the L gain vectors being determined via minimization of a respective one of L cost functions configured to substantially preserve one or more selected metrics of the audio scene for rendering in an audio system upon replacement of the L audio objects by the N clusters. [0006] According to another example embodiment, an audio system for object-based audio comprises: at least one processor; and at least one memory including program code; and wherein the at least one memory and the program code are configured to, with the at least one processor, cause the audio system at least to: select N cluster seeds from L audio objects of an audio scene based on perceptually weighted energies of the audio objects, where N < L; and obtain N clusters corresponding to the N cluster seeds by applying a respective one of L gain vectors to each of the L audio objects, each of the L gain vectors being determined via minimization of a respective one of L cost functions configured to substantially preserve one or more selected metrics of the audio scene for rendering in the audio system upon replacement of the L audio objects by the N clusters. [0007] According to yet another example embodiment, provided is a non-transitory computer-readable medium storing instructions that, when executed by an electronic processor, cause the electronic processor to perform operations comprising the above spatial coding method. BRIEF DESCRIPTION OF THE DRAWINGS [0008] Other aspects, features, and benefits of various disclosed embodiments will become more fully apparent, by way of example, from the following detailed description and the accompanying drawings, in which: [0009] FIG.1 is a block diagram of an audio system in which various embodiments can be practiced.
[0010] FIG.2 is a block diagram illustrating a spatial coding method that can be implemented in the audio system of FIG.1 according to some examples. [0011] FIG.3 is a flowchart illustrating a method of identifying cluster seeds that can be used in the spatial coding method of FIG.2 according to some examples. [0012] FIG.4 is a flowchart illustrating a method of generating the output clusters that can be used in the spatial coding method of FIG.2 according to some examples. [0013] FIG.5 is a block diagram of an example computing device, one or more instances of which can be used to implement the audio system of FIG.1 according to various examples. DETAILED DESCRIPTION [0014] For interactive entertainment content, transmitting the original object-based audio signal, which may contain hundreds of individual objects, can be challenging because endpoints that support object-based audio typically have limitations with respect to the maximum number of audio objects that can be supported, e.g., due to the limited amounts of computational resources and/or memory. In such cases, application of efficient object-based audio-scene management is desirable. [0015] In some cases, “spatial coding” may be used to reduce the complexity of an audio scene. Some spatial coding methods employ clustering techniques that aim to reduce the number of input objects and/or beds to a smaller set of output objects (hereafter referred to as clusters) with minimal impact on the audio quality. [0016] Audio objects can be individual sound elements or collections of sound elements that are perceived to emanate from a particular physical location or locations in the listening environment. Such objects can be static (that is, stationary) or dynamic (that is, moving). Audio objects are controlled by metadata that define the position of the sound source at a given time, along with other functions. When objects are played back, they are rendered according to the positional metadata using the speakers that are present, rather than necessarily being output to a predefined physical channel. A track in a session can be an audio object, and standard panning data may be analogous to positional metadata. While the use of audio objects provides control over discrete effects, other aspects of a soundtrack may work more effectively in a channel-based environment. For example, many ambient effects or reverberation may benefit from being fed to arrays of speakers rather than individual drivers.
[0017] An adaptive audio system typically extends beyond speaker feeds as a means for distributing spatial audio and uses advanced model-based audio descriptions to tailor playback configurations that suit individual needs and system constraints so that audio can be rendered specifically for individual configurations. The spatial effects of audio signals are important for providing an immersive experience for the listener. Sounds that are meant to emanate from a specific region of a viewing screen or room may be played through speaker(s) located at the corresponding relative location. Thus, an important audio metadatum of a sound event in a model-based description is position, although other parameters, such as size, volume, orientation, velocity and, acoustic dispersion, can also be described. [0018] As stated above, in a representative example, audio content may comprise several bed channels and a plurality of individual audio objects that are combined during rendering to create a spatially diverse and immersive audio experience. In a cinema theater environment characterized by a relatively large processing bandwidth, a very large number of beds and objects can be created and accurately rendered. However, as cinema or other complex audio content is produced for distribution and reproduction in home or personal listening environments, the relatively limited processing bandwidth of such devices may prevent optimal rendering or playback of this content. For example, typical transmission media used for consumer applications include Blu-ray disc, broadcast (cable, satellite, or terrestrial), mobile (3G, 4G, 5G), and over the top (OTT) or Internet distribution. These media channels may impose significant limitations on the bandwidth available to digitally transmit the bed and object information of the corresponding audio content. As such, some embodiments described herein are directed to mechanisms that enable reduction, simplification, or streamlining of complex audio content so that it can be distributed through transmission systems that may not possess large enough bandwidth to timely deliver all of audio bed and object data. [0019] In an example embodiment, an audio system has a coding component configured to reduce the bandwidth demand of object-based audio content through object clustering and/or other simplifications that may be based, inter alia, on perceptual importance of different objects. An object clustering process executed by the coding component uses certain information about the objects, such as spatial position, content type, temporal attributes, object width, and loudness, e.g., as described in more detail below, to reduce the spatial complexity of the audio scene by grouping some objects into object clusters that replace the original objects with minimal impact on the audio quality. One purpose of this process is to reduce the spatial complexity of the audio scene through reduction of the number of individual audio elements (e.g., beds and objects) to be specified to the reproduction device. The specified objects still retain enough spatial information
so that the difference between the original content and the output rendered the reproduction device is perceptually insignificant. [0020] In some examples, an audio scene simplification process facilitates the rendering of object-plus-bed content in reduced-bandwidth channels or coding systems based on pertinent information about the objects. In various examples, the audio scene simplification process can reduce the number of objects by clustering (i) two or more audio objects and/or (ii) one or more audio objects with one or more audio beds. Upon such clustering, object clusters replace the collection of individual waveforms and metadata elements of constituent objects with a single, substantially perceptually equivalent waveform and metadata set. The clustering process may utilize an error metric that is based on the amount of distortion to determine the clusters. In various examples, the clustering process may be performed synchronously or may be event- driven, e.g., using auditory scene analysis (ASA) and/or audio event boundary detection. In some examples, the audio scene simplification process may utilize side knowledge of endpoint rendering algorithms and/or devices to inform certain aspects of the clustering process. For example, different clustering schemes may be utilized for speakers versus headphones or other audio drivers, or different clustering schemes may be utilized for lossless versus lossy coding, etc. [0021] Herein, the terms “clustering,” “grouping,” and “combining” may be used interchangeably to describe combinations of objects and/or beds (or channels) configured to reduce the amount of data in a unit of adaptive audio content for transmission and rendering in an audio playback system. The terms “compression” and “reduction” may be used to refer to an act of performing audio scene simplification, e.g., via such clustering. The terms “clustering,” “grouping,” and “combining” throughout this description are not limited to a strictly unique assignment of an audio object or bed to a single cluster only. In some cases, an audio object or bed may be distributed over more than one output bed or cluster using weights or gain vectors that determine the relative contribution of an object or bed signal to the output cluster or output bed signal. [0022] FIG.1 is a block diagram of an audio system 100 in which various embodiments can be practiced. The audio system 100 includes an audio encoder 120 and an audio decoder 140 connected via a bandwidth-limited communication channel 130. The encoder 120 receives input signals 104, 112 and processes the received input signals to generate an encoded bitstream 132. The decoder 140 receives the encoded bitstream 132 via the communication channel 130 and decodes the received bitstream to generate output audio signals 142. The output audio signals
142 are applied to an audio rendering component 150 that operates to render and playback the audio content represented by the audio signals 142. [0023] The input signals 112 received by the encoder 120 are generated by a spatial coding (e.g., clustering) component 110 in response to input signals 102. The spatial coding component 110 applies spatial coding to the input signals 102 to simplify the audio scene represented by the input signals 102, e.g., via clustering. The input signals 104 bypass the spatial coding component 110 and are applied directly to the encoder 120. [0024] In various examples, the audio rendering component 150 may include any professional or consumer-grade audio system, such as a home theater (e.g., including an A/V receiver, a soundbar, a Blu-ray player, etc.), one or more E-media devices (e.g., a computer, a tablet, a mobile phone equipped with headphones and/or speakers, etc.), a TV set, and a sound reproduction system. In some examples, the audio rendering component 150 provides an audio environment for playback of audio or audio/visual content using a plurality of speakers and suitable playback devices. In some examples, the audio rendering component 150 may represent any environment in which a listener is experiencing playback of the audio content, such as a cinema, a concert hall, an outdoor theater, a home theater or room, a listening booth, a car, a game console, a headset device, a public address (PA) system, or other audio playback environment. [0025] FIG.2 is a block diagram illustrating a spatial coding method 200 that can be implemented in the spatial coding component 110 of the audio system 100 according to some examples. The method 200 includes first and second blocks 210, 220 of processing operations. The first block 210 is a cluster-seed selection block in which input audio objects 202 are evaluated to identify a set 212 of seed objects for output clusters 222. Once the set 212 is identified, the cluster positions are determined, and the associated metadata are generated. The second block 220 is a cluster generation block configured to calculate the object to cluster gains and generate the corresponding clusters 222 from the set 212. [0026] In some examples, operations of the first block 210 include evaluating the input audio objects 202 and selecting perceptually most-important objects as cluster seeds. Both the object loudness and spatial position may be considered as factors in the process of determining the object’s importance. Operations of the second block 220 include calculating the object-to-cluster gains and applying the gains to the input audio objects 202 to generate the corresponding output clusters 222. In some examples, the object-to-cluster gains are calculated in the second block
220 based on minimizing a cost function, which may be constricted to consider the position correctness, distance preservation, and amplitude preservation jointly. [0027] Unlike typical linear media, such as music or video with audio (A/V), interactive content is typically generated at runtime as the user interacts with the corresponding virtual scene. In such content, various audio objects may be triggered at any time, while their position, orientation, and loudness are determined at runtime. Hence, relatively low computational complexity is a desired attribute for the corresponding audio scene management method. [0028] Various embodiments disclosed herein provide a relatively low complexity spatial- coding approach to implementing the method 200, which may be referred to as “Spatial Coding Lite.” In some embodiments, the first block 210 of operations employs a novel perceptual importance calculation method described in more detail below. This calculation method takes loudness and spatial position into account. Unlike some previously used methods that employ non-linear psycho-acoustic models, the disclosed perceptual importance calculations of the block 210 can be implemented exclusively via linear operations, which tends to significantly reduce the computational complexity. [0029] In some examples, operations of the first block 210 include selecting N “most important” objects from the input audio objects 202, where the number N is a fixed algorithm parameter. In various use cases, the number N may depend on certain characteristics of the audio system 100. For example, different values of N can be pre-determined for different bandwidth ranges of the communication channel 130. Then, in different deployments, the bandwidth of the corresponding communication channel 130 may be tested, and the number N may then be set accordingly by selecting the corresponding one of the pre-determined values of N. [0030] For an audio object 202 with the spectrum s, the Threshold-in-Quiet filter, ^^ ^^ ^^ ^^, can be applied to the spectrum to get the perceptually weighted spectrum ^^^ as follows: ^^̃ ൌ ^^ ^^ ^^ ^^ ⊗ ^^ (1) where ⊗ denotes the element-wise product of two vectors. In one example, the Threshold-in- Quiet filter ^^ ^^ ^^ ^^ has a spectral shape that represents the frequency-dependent absolute hearing threshold (AHT) or auditory threshold of an average human ear (with normal hearing) under the conditions in which no other sound is present. The AHT can be reported in reference to the RMS sound pressure of 20 micropascals, representing 0 dB SPL (Sound Pressure Level) and corresponding to a sound intensity of 0.98 pW/m2 at 1 atmosphere and 25°C. The AHT is
frequency-dependent and reflects the fact that the ear’s sensitivity is at its best at frequencies between 2 kHz and 5 kHz, where the AHT reaches as low as −9 dB SPL. The AHT progressively increases toward lower frequencies and toward higher frequencies, which causes the Threshold-in-Quiet filter ^^ ^^ ^^ ^^ to have an approximately U-shaped spectral profile in the range between 20 Hz and 20 kHz. Then, the perceptually weighted energy, e, of the audio object 202 can be obtained as: ^^ ൌ ^^̃ ^^̃∗ (2) where * denotes the conjugate transpose. [0031] Suppose there are L (>N) input audio objects 202 in the current audio frame, the objects having the perceptually weighted energies ^^^, … , ^^^ and positional metadata ^^^, … , x^, respectively. Here, the positional metadata ^^^ , ^^ ∈ ^1, … , ^^^ are three-dimensional (3D) vectors representing the objects’ locations in Cartesian form. The relative spatial distance for every audio object pair can be represented by the distance matrix D: 0 ‖x^ െ xଶ‖ଶ ⋯ ‖x^ െ xଶ‖ଶ As is evident
[0032] Next, a masking matrix, F, is calculated by taking the inverse of each object pair distance. In one example, the masking matrix F ^^ ^^^,^൧ is computed using the following formula: ^^ ൌ ^ ^^^,^൧ ൌ ^1 െ ^ ^ ^^,^൧ (4) where ^ ^ ^^,^ is obtained by
distance matrix D to the interval [0,1]. [0033] Using the distance matrix D and the masking matrix F, cluster seeds for the output clusters 222 are iteratively identified one-by-one until the cluster count reaches the number N. In each iteration, the object with the highest perceptual importance in the remaining pool of (previously non-selected) audio objects 202 is selected as a corresponding cluster seed. The relative object importance in that iteration is measured using a linear transform, r, of perceptually weighted energy of all audio objects 202 computed as follows: ^^ ൌ ^^ ^^ (5) where ^^ ൌ ^ ^^^, … , ^^^ ^ୃ. The ^^∗-th object having the highest perceptual importance is selected in the iteration as the corresponding cluster seed, where: ^^∗ ൌ argm ^ax ^^ (6)
When the object is selected as the cluster seed, the perceptually weighted energy vector is updated for the next iteration as follows: ^^ ᇱ ൌ ^^ ⊗ ^ 1 െ ^^^ ∗^ (7) where the vector ^^^ ∗ is the ^^∗-th column of the masking matrix F, and ⊗ denotes the element- wise product of two vectors. The object selections and the perceptually weighted energy updates are repeated until the number of cluster seeds reaches N. [0034] FIG.3 is a flowchart of a method 300 of identifying the set 212 of cluster seeds that can be used in the block 210 of the method 200 according to some examples. The method 300 includes the spatial coding component 110 receiving the input audio objects 202 and corresponding metadata (in a block 302). The method 300 also includes the spatial coding component 110 computing (in a block 304) the perceptually weighted energy vector ^^ ൌ ^ ^^^, … , ^^^^ୃ corresponding to the input audio objects 202 received in the block 302. In various examples, the energy vector computations of the block 304 may be performed in accordance with Eqs. (1)-(2). The method 300 also includes the spatial coding component 110 computing (in a block 306) the distance matrix D and the masking matrix F corresponding to the input audio objects 202 received in the block 302. In various examples, the matrix computations of the block 306 may be performed in accordance with Eqs. (3)-(4). [0035] The method 300 also includes a processing loop including blocks 308-314, wherein the spatial coding component 110 iteratively identifies cluster seeds for the set 212. More specifically, operations of the block 308 include the spatial coding component 110 incrementing by one the iteration index, n. The initial value of the iteration index is set to n=0. Operations of the block 310 include the spatial coding component 110 identifying a next cluster seed based on the perceptually weighted energy vector and the masking matrix F. In various examples, the cluster-seed identification of the block 310 may be performed in accordance with Eqs. (5)-(6). For the first instance of the block 310 (i.e., for n=1), the perceptually weighted energy vector ^^ computed in the block 304 is used. For any subsequent instance of the block 310 (i.e., for n>1), the perceptually weighted energy vector ^^ᇱ computed in the preceding instance of the block 312 is used. Operations of the block 312 include the spatial coding component 110 computing the perceptually weighted energy vector ^^ᇱ by updating the previously computed perceptually weighted energy vector and masking matrix. In various examples, the updates of the block 312 may be performed in accordance with Eq. (7). Operations of the decision block 314 control the exit from the processing loop 308-314 and include the spatial coding component 110 comparing the iteration index n with the number N. When n<N (“No” at the decision block 314), the
processing of the method 300 is directed back to the block 308. When n=N (“Yes” at the decision block 314), the method 300 is terminated. [0036] The following part of this specification introduces how to construct a cost function that can be used in the block 220 of the method 200 according to various examples. [0037] After the set 212 of cluster seeds is identified, e.g., via the method 300, a respective cost matrix needs to be constructed for each input object 202 for calculation of the object-to- cluster gains. More specifically, for a given object 202 and the set 212 of N cluster seeds, the object-to-cluster gains, denoted by ^^ ൌ ^ ^^^, … , ^^ே ^ୃ, can be obtained by minimizing an overall cost ^^ including various sub-costs. Each sub-cost can be considered as a function of the object metadata, clusters’ metadata, and gain vector, where the gain vector is yet unknown. In some examples, the overall cost function ^^ can be expressed using a generalized quadratic form with respect to the gain vector ^^, e.g., as follows: ^^ ൌ ^^ୃ ^^ ^^ ^ ^^ୃ ^^ ^ ^^ (8) In one example, the sub-cost components of the overall cost function ^^ include the metrics of effective position correctness ^^^, object-to-cluster distance ^^ௗ, and amplitude preservation ^^^ defined as follows: ^^ ൌ ^^^ ^ ^^ௗ ^ ^^^ (9) where ^^^, ^^^
N times), respectively; and the matrices ^^, ^^^,^ represent the diagonal matrix and m-by-n all-ones matrix, respectively. By comparing Eq. (8) with Eq. (9) into which Eqs. (10)-(12) are substituted, we find that: ^^^ ൌ ^^ୃ ^^^ ^^ (13) where ^^^ ൌ ^ ^^^ െ ^^^ ^^ ^^^ െ ^^^ ^ୃ (14) [0038] According to the above description, the total number of cost matrices ^^^ is equal to the number of input audio objects 202. The size of each cost matrix is equal to N2. In at least some examples, the corresponding computational load on the processing element(s) of the spatial coding component 110 can be relatively high, which may result in detrimentally high latency times for the audio system 100. Example embodiments described below are directed at reducing
the computational complexity associated with the direct coding of Eqs. (8)-(14) for implementation in the spatial coding component 110. [0039] Suppose there are L input audio objects 202 in the current audio frame. The respective spatial positions of the L input audio objects 202 are given by a set of L 3D position vectors ^^^, where ^^ ൌ 1, … , ^^. We define an object inter-product matrix, X, representing the inter products of various object position vector pairs as follows: ‖ ‖ଶ é x^ x^x ⋯ x^x^ ‖x ‖ଶ ù ^^ ଶ ⋯ xଶx^ ú As is evident from symmetric matrix, computing of
and conclusively show below that the object inter-product matrix X contains all information needed for obtaining the L cost matrices for the L input audio objects 202. In other words, all matrices used for the coding of Eqs. (8)-(14) can be derived from the matrix X. [0040] For the method 300, one needs to obtain the distance matrix D in accordance with Eq. (3). Toward that goal, we first construct an object norm matrix, Y, by replicating the norm vectors as follows: ‖x ‖ଶ ‖x ‖ଶ ⋯ ‖x ‖ଶ é ^ ^ ^ ⋯ As is evident from
matrix. With the matrices X and Y calculated, the distance matrix D can be obtained as follows: ^^ ൌ ^^ െ ^^ ^^ ^ ^^ୃ (17) The following chain of derivations provides a proof of validity for Eq. (17): 0 ‖x^ െ xଶ‖ଶ ⋯ ‖x ଶ ^ െ xଶ‖ ൌ
ൌ ^^ െ ^^ ^^ ^ ^^ୃ (18) [0041] In the method 300, the cluster seeds for the set 212 are identified by iteratively selecting perceptually “most important” objects one-by-one. Therefore, after the processing loop 308-314 is executed N times, the source object indices of the N input audio objects selected as cluster seeds are known. Hereafter, we use {N} to denote the index set containing all source object indices of the cluster seeds. The index set {N} is a subset of 1, … , ^^. Using the index set {N}, we obtain a cluster inter-product matrix, U, by selecting the rows and columns of the object inter-product matrix X having indices that belong to the index set {N}. The resulting cluster inter-product matrix U can be represented as follows: ‖ ‖ଶ é u^ u^u ⋯ u^u u u ‖u ‖ଶ ⋯ ù ^^ ଶ ^ ଶ uଶu [0042] For an ^^-th
an object-cluster inter-product matrix, ^^^, by selecting the ^^ ∈ ^^-th rows and ^^-th column of the matrix X and replicating N times. The resulting object-cluster inter-product matrix ^^^ can be represented as follows: éx^u^ x^u^ ⋯ x^u^ ⋯ ù [0043] For the ^^-th
also obtain a self-product matrix ^^^ in a similar manner. The self-product matrix ^^^ can be represented as follows: éx^x^ x^x^ ⋯ x^x^ ù
[0044] With the , ^^ can as follows: ^^ ^^ ൌ ^^ െ ^^^ െ ^^^ ^ ^^^ (22) The following chain of derivations provides a proof of validity for Eq. (22): ^ ^^ ^ୃ ^ ^^ ^ୃ ^ ^^ ^ୃ é u^ െ x^ u^ െ x^ u^ െ x^ uଶ െ x^ ⋯ u^ െ x^ uே െ x^ ú ൌ ୃú û
éu^u^ u^u ⋯ u^u x^u x^u ⋯ x^u u u u ù é ^ ù ൌ ê ଶ ^ ଶu ⋯ uଶu ú െêx^u^ x^u ⋯ x^u ú െ [0045]
222 that can be used in the block 220 of the method 200 according to some examples. The method 400 uses the above-presented correspondences between different matrices to beneficially reduce the computational load compared to that associated with the direct coding of Eqs. (8)-(14). In various examples, the method 400 is used in conjunction with the method 300. For example, the set 212 of cluster seeds identified with the method 300 and the corresponding source object indices are used as inputs to the method 400. [0046] The method 400 includes receiving identification of the cluster-seed objects (in a block 402). As already indicated above, the set 212 of cluster-seed objects used as an input to the method 400 can be determined with the method 300. Operations of the block 402 also include receiving the metadata corresponding to the set 212 and the set {N} of source object indices of the set 212 with respect to the original set of the input audio objects 202. [0047] The method 400 also includes the spatial coding component 110 constructing the cluster inter-product matrix U (in a block 404). In various examples, the matrix U is constructed in accordance with Eq. (19). As already indicated above, the matrix U can be constructed using a first subset of matrix elements of the object inter-product matrix X, which is selected based on the index set {N} as described above in reference to Eq. (19). In some examples, the object inter-product matrix X is computed in the block 306 of the method 300, wherein the distance matrix D is computed in accordance with Eqs. (15)-(17). In such examples, the object inter- product matrix X is saved in the memory in the block 306 of the method 300 and is retrieved from the memory for the operations of the block 404. In some other examples, the object inter- product matrix X can be computed in the block 404 in accordance with Eq. (15). [0048] The method 400 also includes the spatial coding component 110 constructing a set { ^^^} of object-cluster inter-product matrices and a set { ^^^} of self-product matrices (in a block 406). Each of the sets { ^^^} and { ^^^} has L respective matrices. In various examples, for each l,
the matrix ^^^ is constructed in accordance with Eq. (20), and the matrix ^^^ is constructed in accordance with Eq. (21). As already indicated above, the matrix ^^^ is constructed using a respective second subset of matrix elements of the object inter-product matrix X, which is selected based on the index l and the index set {N} as described above in reference to Eq. (20). The matrix ^^^ is similarly constructed using a respective third subset of matrix elements of the object inter-product matrix X, which is selected based on the index l and the index set {N} as described above in reference to Eq. (21). [0049] The method 400 also includes the spatial coding component 110 computing a set { ^^ ^^} of cost matrices (in a block 408). The set { ^^ ^^} has L matrices. In various examples, for each l, the corresponding cost matrix ^^ ^^ is computed in accordance with Eq. (22) using the matrix U constructed in the block 404 and the corresponding matrices ^^^ and ^^^ constructed in the block 406. [0050] The method 400 also includes the spatial coding component 110 determining a set { ^^^} of gain vectors (in a block 410). In various examples, the l-th gain vector ^^^ is determined in the block 410 by minimizing the corresponding cost function ^^. Recall that the metric of position correctness ^^^ of the cost function ^^ is determined in accordance with Eq. (13) using the corresponding cost matrix ^^ ^^, which has been computed in the block 408. [0051] The method 400 also includes the spatial coding component 110 obtaining the output clusters (in a block 412). In various examples, the l-th output cluster of the output clusters 222 is determined by applying the corresponding gain vector ^^^ determined in the block 410 to the input audio objects 202. Operations of the block 412 also include generating the metadata for the output clusters 222. Upon completion of the operations of the block 412, the method 400 is terminated. [0052] FIG.5 is a block diagram of an example computing device 500 according to various examples. In some examples, the computing device 500 is configured to perform at least some operations of the methods 200, 300, and 400. In some examples, two or more instances of the computing device 500 are used in the audio system 100. [0053] The computing device 500 of FIG.5 is illustrated as having a number of components, but any one or more of these components may be omitted or duplicated, as suitable for the application and setting. In some embodiments, some or all of the components included in the computing device 500 may be attached to one or more motherboards and enclosed in a housing. In some embodiments, some of those components may be fabricated onto a single system-on-a-
chip (SoC) (e.g., the SoC may include one or more electronic processing devices 502 and one or more storage devices 504). Additionally, in various embodiments, the computing device 500 may not include one or more of the components illustrated in FIG.5, but may include interface circuitry for coupling to the one or more components using any suitable interface (e.g., a Universal Serial Bus (USB) interface, a High-Definition Multimedia Interface (HDMI) interface, a Controller Area Network (CAN) interface, a Serial Peripheral Interface (SPI) interface, an Ethernet interface, a wireless interface, or any other appropriate interface). For example, the computing device 500 may not include a display device 510, but may include display device interface circuitry (e.g., a connector and driver circuitry) to which an external display device 510 may be coupled. [0054] The computing device 500 includes a processing device 502 (e.g., one or more processing devices). As used herein, the terms “electronic processor device” and “processing device” interchangeably refer to any device or portion of a device that processes electronic data from registers and/or memory to transform that electronic data into other electronic data that may be stored in registers and/or memory. In various embodiments, the processing device 502 may include one or more digital signal processors (DSPs), application-specific integrated circuits (ASICs), central processing units (CPUs), graphics processing units (GPUs), server processors, or any other suitable processing devices. [0055] The computing device 500 also includes a storage device 504 (e.g., one or more storage devices). In various embodiments, the storage device 504 may include one or more memory devices, such as random-access memory (RAM) devices (e.g., static RAM (SRAM) devices, magnetic RAM (MRAM) devices, dynamic RAM (DRAM) devices, resistive RAM (RRAM) devices, or conductive-bridging RAM (CBRAM) devices), hard drive-based memory devices, solid-state memory devices, networked drives, cloud drives, or any combination of memory devices. In some embodiments, the storage device 504 may include memory that shares a die with the processing device 502. In such an embodiment, the memory may be used as cache memory and include embedded dynamic random-access memory (eDRAM) or spin transfer torque magnetic random-access memory (STT-MRAM), for example. In some embodiments, the storage device 504 may include non-transitory computer readable media having instructions thereon that, when executed by one or more processing devices (e.g., the processing device 502), cause the computing device 500 to perform any appropriate ones of the methods disclosed herein below or portions of such methods.
[0056] The computing device 500 further includes an interface device 506 (e.g., one or more interface devices 506). In various embodiments, the interface device 506 may include one or more communication chips, connectors, and/or other hardware and software to govern communications between the computing device 500 and other computing devices. For example, the interface device 506 may include circuitry for managing wireless communications for the transfer of data to and from the computing device 500. The term “wireless” and its derivatives may be used to describe circuits, devices, systems, methods, techniques, communications channels, etc., that may communicate data via modulated electromagnetic radiation through a nonsolid medium. The term does not imply that the associated devices do not contain any wires, although in some embodiments they might not. Circuitry included in the interface device 506 for managing wireless communications may implement any of a number of wireless standards or protocols, including but not limited to Institute for Electrical and Electronic Engineers (IEEE) standards including Wi-Fi (IEEE 802.11 family), IEEE 802.16 standards, Long-Term Evolution (LTE) project along with any amendments, updates, and/or revisions (e.g., advanced LTE project, ultramobile broadband (UMB) project (also referred to as “3GPP2”), etc.). In some embodiments, circuitry included in the interface device 506 for managing wireless communications may operate in accordance with a Global System for Mobile Communication (GSM), General Packet Radio Service (GPRS), Universal Mobile Telecommunications System (UMTS), High Speed Packet Access (HSPA), Evolved HSPA (E-HSPA), or LTE network. In some embodiments, circuitry included in the interface device 506 for managing wireless communications may operate in accordance with Enhanced Data for GSM Evolution (EDGE), GSM EDGE Radio Access Network (GERAN), Universal Terrestrial Radio Access Network (UTRAN), or Evolved UTRAN (E-UTRAN). In some embodiments, circuitry included in the interface device 506 for managing wireless communications may operate in accordance with Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Digital Enhanced Cordless Telecommunications (DECT), Evolution-Data Optimized (EV-DO), and derivatives thereof, as well as any other wireless protocols that are designated as 3G, 4G, 5G, and beyond. In some embodiments, the interface device 506 may include one or more antennas (e.g., one or more antenna arrays) configured to receive and/or transmit wireless signals. [0057] In some embodiments, the interface device 506 may include circuitry for managing wired communications, such as electrical, optical, or any other suitable communication protocols. For example, the interface device 506 may include circuitry to support communications in accordance with Ethernet technologies. In some embodiments, the interface device 506 may support both wireless and wired communication, and/or may support multiple
wired communication protocols and/or multiple wireless communication protocols. For example, a first set of circuitry of the interface device 506 may be dedicated to shorter-range wireless communications such as Wi-Fi or Bluetooth, and a second set of circuitry of the interface device 506 may be dedicated to longer-range wireless communications such as global positioning system (GPS), EDGE, GPRS, CDMA, WiMAX, LTE, EV-DO, or others. In some other embodiments, a first set of circuitry of the interface device 506 may be dedicated to wireless communications, and a second set of circuitry of the interface device 506 may be dedicated to wired communications. [0058] The computing device 500 also includes battery/power circuitry 508. In various embodiments, the battery/power circuitry 508 may include one or more energy storage devices (e.g., batteries or capacitors) and/or circuitry for coupling components of the computing device 500 to an energy source separate from the computing device 500 (e.g., to AC line power). [0059] The computing device 500 also includes a display device 510 (e.g., one or multiple individual display devices). In various embodiments, the display device 510 may include any visual indicators, such as a heads-up display, a computer monitor, a projector, a touchscreen display, a liquid crystal display (LCD), a light-emitting diode display, or a flat panel display. [0060] The computing device 500 also includes additional input/output (I/O) devices 512. In various embodiments, the I/O devices 512 may include one or more data/signal transfer interfaces, audio I/O devices (e.g., microphones or microphone arrays, speakers, headsets, earbuds, alarms, etc.), audio codecs, video codecs, printers, sensors (e.g., thermocouples or other temperature sensors, humidity sensors, pressure sensors, vibration sensors, etc.), image capture devices (e.g., one or more cameras), human interface devices (e.g., keyboards, cursor control devices, such as a mouse, a stylus, a trackball, or a touchpad), etc. [0061] Depending on the specific embodiment, various components of the interface devices 506 and/or I/O devices 512 can be configured to output suitable control signals, receive suitable control/telemetry signals, and receive and transmit data streams. In some examples, the interface devices 506 and/or I/O devices 512 include one or more analog-to-digital converters (ADCs) for transforming received analog signals into a digital form suitable for operations performed by the processing device 502 and/or the storage device 504. In some additional examples, the interface devices 506 and/or I/O devices 512 include one or more digital-to-analog converters (DACs) for transforming digital signals provided by the processing device 502 and/or the storage device 504 into an analog form suitable for being transmitted through a communication channel.
[0062] According to an example embodiment disclosed above, e.g., in the summary section and/or in reference to any one or any combination of some or all of FIGS.1-5, provided is an audio system for object-based audio, the audio system comprising: at least one processor; and at least one memory including program code; and wherein the at least one memory and the program code are configured to, with the at least one processor, cause the audio system at least to: select N cluster seeds from L audio objects of an audio scene based on perceptually weighted energies of the audio objects, where N < L; and obtain N clusters corresponding to the N cluster seeds by applying a respective one of L gain vectors to each of the L audio objects, each of the L gain vectors being determined via minimization of a respective one of L cost functions configured to substantially preserve one or more selected metrics of the audio scene for rendering in the audio system upon replacement of the L audio objects by the N clusters. [0063] According to another example embodiment disclosed above, e.g., in the summary section and/or in reference to any one or any combination of some or all of FIGS.1-5, provided is a spatial coding method for object-based audio, the method comprising: selecting N cluster seeds from L audio objects of an audio scene based on perceptually weighted energies of the L audio objects, where N < L; and obtaining N clusters corresponding to the N cluster seeds by applying a respective one of L gain vectors to each of the L audio objects, each of the L gain vectors being determined via minimization of a respective one of L cost functions configured to substantially preserve one or more selected metrics of the audio scene for rendering in an audio system upon replacement of the L audio objects by the N clusters. [0064] In some embodiments of the above method, each of the L gain vectors has N respective components. [0065] In some embodiments of any of the above methods, the selecting is further based on pairwise distances of the L audio objects. [0066] In some embodiments of any of the above methods, the one or more selected metrics include one or more of a metric of position correctness, a metric of object-to-cluster distance, and a metric of amplitude preservation. [0067] In some embodiments of any of the above methods, the selecting includes applying a threshold-in-quiet filter to frequency spectra of the L audio objects. [0068] In some embodiments of any of the above methods, the selecting includes iteratively identifying the N cluster seeds one by one.
[0069] In some embodiments of any of the above methods, the method further comprises selecting a value of N based on one or more parameters of the audio system. [0070] In some embodiments of any of the above methods, the method further comprises computing a first L ^L matrix whose matrix elements represent inter products of position vectors of different pairs of the audio objects. [0071] In some embodiments of any of the above methods, the method further comprises computing a second L ^L matrix whose matrix elements represent squared norms of the position vectors of individual ones of the L audio objects, the second L ^L matrix being a rank-1 matrix. [0072] In some embodiments of any of the above methods, the selecting includes computing a distance matrix using a linear combination of the first L ^L matrix and the second L ^L matrix, the distance matrix having non-diagonal matrix elements that represent squared norms of position vector differences between different object pairs selected from the L audio objects and further having all zero diagonal matrix elements. [0073] In some embodiments of any of the above methods, the selecting further includes adding to the linear combination a transposed version of the second L ^L matrix. [0074] In some embodiments of any of the above methods, the method further comprises: constructing a first N ^N matrix using a first subset of the matrix elements of the first L ^L matrix; for each of the L audio objects, constructing a respective second N ^N matrix using a respective second subset of the matrix elements of the first L ^L matrix; and for each of the L audio objects, constructing a respective third N ^N matrix using a respective third subset of the matrix elements of the first L ^L matrix. [0075] In some embodiments of any of the above methods, the method further comprises, for each of the L audio objects, computing a respective cost matrix using a linear combination of the first N ^N matrix, the respective second N ^N matrix, and the respective third N ^N matrix. [0076] In some embodiments of any of the above methods, the computing said respective cost matrix further comprises subtracting from the linear combination a transposed version of the respective second N ^N matrix. [0077] In some embodiments of any of the above methods, the obtaining comprises computing the respective one of the L cost functions using the respective cost matrix.
[0078] In some embodiments of any of the above methods, the audio system comprises an audio rendering component configured to generate sound corresponding to the N clusters. [0079] In some embodiments of any of the above methods, the audio system comprises: a spatial coding component configured to perform the selecting and further configured to perform the obtaining; and an audio encoder configured to generate a bitstream having encoded therein the N clusters; and wherein the spatial coding method further comprises transmitting the bitstream over a communication channel. [0080] In some embodiments of any of the above methods, the audio system further comprises: an audio decoder configured to decode the bitstream received over the communication channel to recover the N clusters; and an audio rendering component configured to generate sound corresponding to the recovered N clusters. [0081] A non-transitory computer-readable medium storing instructions that, when executed by an electronic processor, cause the electronic processor to perform operations comprising any one of the above methods. [0082] With regard to the processes, systems, methods, heuristics, etc. described herein, it should be understood that, although the steps of such processes, etc. have been described as occurring according to a certain ordered sequence, such processes could be practiced with the described steps performed in an order other than the order described herein. It further should be understood that certain steps could be performed simultaneously, that other steps could be added, or that certain steps described herein could be omitted. In other words, the descriptions of processes herein are provided for the purpose of illustrating certain embodiments and should in no way be construed so as to limit the claims. [0083] Accordingly, it is to be understood that the above description is intended to be illustrative and not restrictive. Many embodiments and applications other than the examples provided would be apparent upon reading the above description. The scope should be determined, not with reference to the above description, but should instead be determined with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled. It is anticipated and intended that future developments will occur in the technologies discussed herein, and that the disclosed systems and methods will be incorporated into such future embodiments. In sum, it should be understood that the application is capable of modification and variation.
[0084] All terms used in the claims are intended to be given their broadest reasonable constructions and their ordinary meanings as understood by those knowledgeable in the technologies described herein unless an explicit indication to the contrary is made herein. In particular, use of the singular articles such as “a,” “the,” “said,” etc. should be read to recite one or more of the indicated elements unless a claim recites an explicit limitation to the contrary. [0085] The Abstract of the Disclosure is provided to allow the reader to quickly ascertain the nature of the technical disclosure. It is submitted with the understanding that it will not be used to interpret or limit the scope or meaning of the claims. In addition, in the foregoing Detailed Description, it can be seen that various features are grouped together in various embodiments for the purpose of streamlining the disclosure. This method of disclosure is not to be interpreted as reflecting an intention that the claimed embodiments incorporate more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive subject matter lies in fewer than all features of a single disclosed embodiment. Thus, the following claims are hereby incorporated into the Detailed Description, with each claim standing on its own as a separately claimed subject matter. [0086] While this disclosure includes references to illustrative embodiments, this specification is not intended to be construed in a limiting sense. Various modifications of the described embodiments, as well as other embodiments within the scope of the disclosure, which are apparent to persons skilled in the art to which the disclosure pertains are deemed to lie within the principle and scope of the disclosure, e.g., as expressed in the following claims. [0087] Some embodiments may be implemented as circuit-based processes, including possible implementation on a single integrated circuit. [0088] Some embodiments can be embodied in the form of methods and apparatuses for practicing those methods. Some embodiments can also be embodied in the form of program code recorded in tangible media, such as magnetic recording media, optical recording media, solid state memory, floppy diskettes, CD-ROMs, hard drives, or any other non-transitory machine-readable storage medium, wherein, when the program code is loaded into and executed by a machine, such as a computer, the machine becomes an apparatus for practicing the patented invention(s). Some embodiments can also be embodied in the form of program code, for example, stored in a non-transitory machine-readable storage medium including being loaded into and/or executed by a machine, wherein, when the program code is loaded into and executed by a machine, such as a computer or a processor, the machine becomes an apparatus for
practicing the patented invention(s). When implemented on a general-purpose processor, the program code segments combine with the processor to provide a unique device that operates analogously to specific logic circuits. [0089] Unless explicitly stated otherwise, each numerical value and range should be interpreted as being approximate as if the word “about” or “approximately” preceded the value or range. [0090] The use of figure numbers and/or figure reference labels in the claims is intended to identify one or more possible embodiments of the claimed subject matter in order to facilitate the interpretation of the claims. Such use is not to be construed as necessarily limiting the scope of those claims to the embodiments shown in the corresponding figures. [0091] Although the elements in the following method claims, if any, are recited in a particular sequence with corresponding labeling, unless the claim recitations otherwise imply a particular sequence for implementing some or all of those elements, those elements are not necessarily intended to be limited to being implemented in that particular sequence. [0092] Reference herein to “one embodiment” or “an embodiment” means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the disclosure. The appearances of the phrase “in one embodiment” in various places in the specification are not necessarily all referring to the same embodiment, nor are separate or alternative embodiments necessarily mutually exclusive of other embodiments. The same applies to the term “implementation.” [0093] Unless otherwise specified herein, the use of the ordinal adjectives “first,” “second,” “third,” etc., to refer to an object of a plurality of like objects merely indicates that different instances of such like objects are being referred to, and is not intended to imply that the like objects so referred-to have to be in a corresponding order or sequence, either temporally, spatially, in ranking, or in any other manner. [0094] Unless otherwise specified herein, in addition to its plain meaning, the conjunction “if” may also or alternatively be construed to mean “when” or “upon” or “in response to determining” or “in response to detecting,” which construal may depend on the corresponding specific context. For example, the phrase “if it is determined” or “if [a stated condition] is detected” may be construed to mean “upon determining” or “in response to determining” or
“upon detecting [the stated condition or event]” or “in response to detecting [the stated condition or event].” [0095] Also, for purposes of this description, the terms “couple,” “coupling,” “coupled,” “connect,” “connecting,” or “connected” refer to any manner known in the art or later developed in which energy is allowed to be transferred between two or more elements, and the interposition of one or more additional elements is contemplated, although not required. Conversely, the terms “directly coupled,” “directly connected,” etc., imply the absence of such additional elements. [0096] As used herein in reference to an element and a standard, the term compatible means that the element communicates with other elements in a manner wholly or partially specified by the standard and would be recognized by other elements as sufficiently capable of communicating with the other elements in the manner specified by the standard. The compatible element does not need to operate internally in a manner specified by the standard. [0097] The functions of the various elements shown in the figures, including any functional blocks labeled as “processors” and/or “controllers,” may be provided through the use of dedicated hardware as well as hardware capable of executing software in association with appropriate software. When provided by a processor, the functions may be provided by a single dedicated processor, by a single shared processor, or by a plurality of individual processors, some of which may be shared. Moreover, explicit use of the term “processor” or “controller” should not be construed to refer exclusively to hardware capable of executing software, and may implicitly include, without limitation, digital signal processor (DSP) hardware, network processor, application specific integrated circuit (ASIC), field programmable gate array (FPGA), read only memory (ROM) for storing software, random access memory (RAM), and nonvolatile storage. Other hardware, conventional and/or custom, may also be included. Similarly, any switches shown in the figures are conceptual only. Their function may be carried out through the operation of program logic, through dedicated logic, through the interaction of program control and dedicated logic, or even manually, the particular technique being selectable by the implementer as more specifically understood from the context. [0098] As used in this application, the terms “circuit,” “circuitry” may refer to one or more or all of the following: (a) hardware-only circuit implementations (such as implementations in only analog and/or digital circuitry); (b) combinations of hardware circuits and software, such as (as applicable): (i) a combination of analog and/or digital hardware circuit(s) with
software/firmware and (ii) any portions of hardware processor(s) with software (including digital signal processor(s)), software, and memory(ies) that work together to cause an apparatus, such as a mobile phone or server, to perform various functions); and (c) hardware circuit(s) and or processor(s), such as a microprocessor(s) or a portion of a microprocessor(s), that requires software (e.g., firmware) for operation, but the software may not be present when it is not needed for operation.” This definition of circuitry applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term circuitry also covers an implementation of merely a hardware circuit or processor (or multiple processors) or portion of a hardware circuit or processor and its (or their) accompanying software and/or firmware. The term circuitry also covers, for example and if applicable to the particular claim element, a baseband integrated circuit or processor integrated circuit for a mobile device or a similar integrated circuit in server, a cellular network device, or other computing or network device. [0099] It should be appreciated by those of ordinary skill in the art that any block diagrams herein represent conceptual views of illustrative circuitry embodying the principles of the disclosure. Similarly, it will be appreciated that any flow charts, flow diagrams, state transition diagrams, pseudo code, and the like represent various processes which may be substantially represented in computer readable medium and so executed by a computer or processor, whether or not such computer or processor is explicitly shown. [00100] “BRIEF SUMMARY OF SOME SPECIFIC EMBODIMENTS” in this specification is intended to introduce some example embodiments, with additional embodiments being described in “DETAILED DESCRIPTION” and/or in reference to one or more drawings. “BRIEF SUMMARY OF SOME SPECIFIC EMBODIMENTS” is not intended to identify essential elements or features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter.
Claims
CLAIMS What is claimed is: 1. A spatial coding method for object-based audio, the method comprising: selecting N cluster seeds from L audio objects of an audio scene based on perceptually weighted energies of the L audio objects, where N < L; and obtaining N clusters corresponding to the N cluster seeds by applying a respective one of L gain vectors to each of the L audio objects, each of the L gain vectors being determined via minimization of a respective one of L cost functions configured to substantially preserve one or more selected metrics of the audio scene for rendering in an audio system upon replacement of the L audio objects by the N clusters.
2. The spatial coding method of claim 1, wherein each of the L gain vectors has N respective components.
3. The spatial coding method of claim 1 or 2, wherein the selecting is further based on pairwise distances of the L audio objects.
4. The spatial coding method of any preceding claim, wherein the one or more selected metrics include one or more of a metric of position correctness, a metric of object-to-cluster distance, and a metric of amplitude preservation.
5. The spatial coding method of any preceding claim, wherein the selecting includes applying a threshold-in-quiet filter to frequency spectra of the L audio objects.
6. The spatial coding method of any preceding claim, wherein the selecting includes iteratively identifying the N cluster seeds one by one.
7. The spatial coding method of any preceding claim, further comprising selecting a value of N based on one or more parameters of the audio system.
8. The spatial coding method of any preceding claim, further comprising computing a first L ^L matrix whose matrix elements represent inter products of position vectors of different pairs of the audio objects.
9. The spatial coding method of claim 8, further comprising computing a second L ^L matrix whose matrix elements represent squared norms of the position vectors of individual ones of the L audio objects, the second L ^L matrix being a rank-1 matrix.
10. The spatial coding method of claim 9, wherein the selecting includes computing a distance matrix using a linear combination of the first L ^L matrix and the second L ^L matrix, the distance matrix having non-diagonal matrix elements that represent squared norms of position vector differences for different object pairs selected from the L audio objects and further having all zero diagonal matrix elements.
11. The spatial coding method of claim 10, wherein the selecting further includes adding to the linear combination a transposed version of the second L ^L matrix.
12. The spatial coding method of claim 8, further comprising: constructing a first N ^N matrix using a first subset of the matrix elements of the first L ^L matrix; for each of the L audio objects, constructing a respective second N ^N matrix using a respective second subset of the matrix elements of the first L ^L matrix; and for each of the L audio objects, constructing a respective third N ^N matrix using a respective third subset of the matrix elements of the first L ^L matrix.
13. The spatial coding method of claim 12, further comprising: for each of the L audio objects, computing a respective cost matrix using a linear combination of the first N ^N matrix, the respective second N ^N matrix, and the respective third N ^N matrix.
14. The spatial coding method of claim 13, wherein the computing said respective cost matrix further comprises subtracting from the linear combination a transposed version of the respective second N ^N matrix.
15. The spatial coding method of claim 13 or 14, wherein the obtaining comprises computing the respective one of the L cost functions using the respective cost matrix.
16. The spatial coding method of any preceding claim, wherein the audio system comprises an audio rendering component configured to generate sound corresponding to the N clusters.
17. The spatial coding method of any preceding claim, wherein the audio system comprises: a spatial coding component configured to perform the selecting and further configured to perform the obtaining; and an audio encoder configured to generate a bitstream having encoded therein the N clusters; and wherein the spatial coding method further comprises transmitting the bitstream over a communication channel.
18. The spatial coding method of claim 17, wherein the audio system further comprises: an audio decoder configured to decode the bitstream received over the communication channel to recover the N clusters; and an audio rendering component configured to generate sound corresponding to the recovered N clusters.
19. A non-transitory computer-readable medium storing instructions that, when executed by an electronic processor, cause the electronic processor to perform operations comprising the method of any one of claims 1 to 18.
20. An audio system for object-based audio, the audio system comprising: at least one processor; and at least one memory including program code; and wherein the at least one memory and the program code are configured to, with the at least one processor, cause the audio system at least to: select N cluster seeds from L audio objects of an audio scene based on perceptually weighted energies of the audio objects, where N < L; and obtain N clusters corresponding to the N cluster seeds by applying a respective one of L gain vectors to each of the L audio objects, each of the L gain vectors being determined via minimization of a respective one of L cost functions configured to substantially preserve one or more selected metrics of the audio scene for rendering in the audio system upon replacement of the L audio objects by the N clusters.
Applications Claiming Priority (3)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN2023104047 | 2023-06-29 | ||
| US202463558254P | 2024-02-27 | 2024-02-27 | |
| PCT/US2024/034469 WO2025006265A1 (en) | 2023-06-29 | 2024-06-18 | Spatial coding of object-based audio |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP4736160A1 true EP4736160A1 (en) | 2026-05-06 |
Family
ID=91924610
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP24742379.1A Pending EP4736160A1 (en) | 2023-06-29 | 2024-06-18 | Spatial coding of object-based audio |
Country Status (3)
| Country | Link |
|---|---|
| EP (1) | EP4736160A1 (en) |
| CN (1) | CN121666616A (en) |
| WO (1) | WO2025006265A1 (en) |
Family Cites Families (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US9805725B2 (en) * | 2012-12-21 | 2017-10-31 | Dolby Laboratories Licensing Corporation | Object clustering for rendering object-based audio content based on perceptual criteria |
| WO2015017037A1 (en) * | 2013-07-30 | 2015-02-05 | Dolby International Ab | Panning of audio objects to arbitrary speaker layouts |
| CN105895086B (en) * | 2014-12-11 | 2021-01-12 | 杜比实验室特许公司 | Metadata-preserving audio object clustering |
| WO2018017394A1 (en) * | 2016-07-20 | 2018-01-25 | Dolby Laboratories Licensing Corporation | Audio object clustering based on renderer-aware perceptual difference |
-
2024
- 2024-06-18 CN CN202480051047.2A patent/CN121666616A/en active Pending
- 2024-06-18 EP EP24742379.1A patent/EP4736160A1/en active Pending
- 2024-06-18 WO PCT/US2024/034469 patent/WO2025006265A1/en not_active Ceased
Also Published As
| Publication number | Publication date |
|---|---|
| WO2025006265A1 (en) | 2025-01-02 |
| CN121666616A (en) | 2026-03-13 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| CN104956695B (en) | It is determined that the method and apparatus of the renderer for spherical harmonics coefficient | |
| EP3028273A1 (en) | Processing spatially diffuse or large audio objects | |
| WO2015138856A1 (en) | Low frequency rendering of higher-order ambisonic audio data | |
| US20240119946A1 (en) | Audio rendering system and method and electronic device | |
| US11122386B2 (en) | Audio rendering for low frequency effects | |
| EP4295587B1 (en) | Clustering audio objects | |
| EP4736160A1 (en) | Spatial coding of object-based audio | |
| CN114128312B (en) | Audio rendering for low frequency effects | |
| RU2852710C2 (en) | Audio object clustering | |
| WO2025193580A1 (en) | Binaural determination of direction to an audio object | |
| WO2025137504A1 (en) | Stereo recording with closely spaced microphones | |
| CN116965062A (en) | Cluster audio objects | |
| CN121260169A (en) | Audio processing method, electronic equipment, storage medium and chip | |
| KR20180024612A (en) | A method and an apparatus for processing an audio signal |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: UNKNOWN |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE |
|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE |